Hi, A lot happened this week.
First, we hosted a dinner during Snowflake Summit in San Francisco with data and AI leaders from companies like Google, NVIDIA, and Snowflake. These are people who’ve been building data systems long enough to know where AI succeeds and where it breaks.
We also had a few milestones worth celebrating.
Genloop is now #1 on LiveSQLBench with a 68.15% score. For context, leading agents from OpenAI and Anthropic are below 50%.
Our new product demo is live. It shows how a business user can get answers directly from their data with trust and governance.
Let’s get into it.
Shipped at Genloop
Connect. Chat. Save. That’s It.
The old way: seat licenses per user, data team builds the semantic layer, builds the dashboard, every investigation goes into their queue. You wait days for a business insight.
Three steps to get started with Genloop now.
Connect your warehouse
Chat with your data
Save what matters as a Liveboard
That gets you free dashboards and a pocket analyst that leads benchmarks globally. In minutes.
The Business User Problem. And How We’re Solving It.
There are two kinds of agentic analytics products being built right now.
One is built for analysts. They know the schema, they can review the SQL. Useful. Incremental. Claude or ChatGPT already do it.
The harder problem is the business user. They come with intent, not schema knowledge. “Why did revenue drop in the northeast last quarter?” No idea which tables join, no idea how region is modeled, no idea if sandbox accounts are excluded. It’s an understanding and trust problem.
Making it trustable means the system knows your business: your KPIs, your policies, your edge cases. When it returns an answer, it shows its work. A non-technical user can tell a confident answer from a confidently wrong one, and get a human in the loop when they need it.
That’s what we’ve been building. The new demo shows how it shapes the entire experience.
Genloop Corner
#1 on LiveSQLBench
Spider 2.0 was one of the most demanding text-to-SQL benchmarks out there. LiveSQLBench is another, and we’re first on that too.
Genloop’s Agent Sentinel V1 scored 68.15% on LiveSQLBench, a live benchmark built on real end-user databases with tasks that reflect how SQL problems actually look in production. No advance look at what’s coming, live leaderboard, no gaming it.
The gap:
Genloop: 68.15%
o3-mini (OpenAI): 50%
Claude Opus 4.6: 49.60%
On the 180 read-only tasks specifically, Sentinel scored 162/180. That’s 90.5%. We don’t manipulate source data, so that’s the number that matters.
First on Spider 2. First on LiveSQLBench. Same story, different benchmark.
This Week’s Read
What’s Actually Behind Genloop
People keep asking what the core pillars of Genloop are. How we’re hitting 97% on Spider 2 with a 5-minute time to value.
Five pillars: Living Context Graph, Deterministic Reasoning, Compounding Learning, Decision Intelligence, and Governance. Each one addresses a specific reason why “add AI to your BI” keeps failing in production. If you’ve ever wondered why the gap between Genloop and everything else on the benchmarks is that wide, this is where the answer is.
What Agentic Analytics Actually Needs: The Science Behind Genloop
Most AI Rollouts Don’t Die From One Resistance. They Die From Two.
Senior engineers resist because handing reasoning to AI feels like an identity ask, not an efficiency one. Non-technical users try it once, get a wrong answer, and go back to emailing the data team. They always had a working alternative.
Two different problems. One rollout plan. That’s why most adoption programs fail, and why the middle always moves faster than anyone expects.
What Anthropic Got Right and Wrong About Agentic Analytics
Anthropic just wrote 5,000 words on how they do agentic analytics in-house. Central admission: pointing Claude directly at a warehouse doesn’t work. Cold start accuracy sits at 21%. After months of senior data work: 95%. Without active maintenance, it drifts back to 65% in a single month.
They built it in-house because they had the time, team, and money to do it. Most companies don’t. That’s the gap this piece unpacks.
That’s it for Issue #2.
SF had a good week on sidelines of the Snowflake Summit. Good people, great conversations, and a benchmark win to go with it.
More of this coming. Next issue we’ll get into what we’re seeing inside enterprise deployments. Try Genloop free today!
Thanks for reading.
Ayush
CEO, Genloop








