Glean Leverages Model Routing to Cut Enterprise AI Costs
As frontier model expenses climb, enterprise AI platform Glean is using dynamic model routing and its Waldo search agent to dramatically lower operational costs for large organizations.

Enterprise AI assistant provider Glean has reached $300 million in annual recurring revenue, marking a threefold increase over 15 months. Valued at $7.2 billion following a $150 million Series F funding round last June, the company is seeing massive demand for its model routing capabilities. According to CEO Arvind Jain, organizations are turning to routing to curb the spiraling costs of frontier models like GPT or Claude Opus. To combat this, Glean's system dynamically directs queries to the most cost-effective model, or bypasses large language models entirely for simple tasks.
The financial impact of this approach is significant. Glean engineering lead Tony Gentilcore reported that the platform is four times more cost-effective than Claude Code, averaging $0.45 per task compared to $1.84 for Claude Cowork. This efficiency is partly driven by Waldo, an agentic search model introduced in April. Waldo acts as a front-end filter that breaks down queries and gathers necessary information first, which reduces latency by 50% and token usage by 25%. This ensures that expensive frontier models are only queried when absolutely necessary.
This routing strategy aligns with a rapid shift toward open-weight models like Kimi K3 and Qwen3.8-Max. Jain noted that while enterprise interest in open-source AI was negligible last year, high costs over the past three months have made open-weight alternatives a core strategy. This shift is playing out across massive deployments, including Zillow, which has seen 80% adoption among its 7,000 employees, and Booking.com, which adopted Glean company-wide.
For AI practitioners, Glean's architecture changes how LLM applications are managed at scale. Instead of relying on a single provider, developers can use Glean's automatic routing mode or set administrative limits. Behind the scenes, Glean refines its routing by running parallel tests on a fraction of real-world traffic, using "AI-based judges" to evaluate whether the router made the most cost-effective choice. This continuous feedback loop allows practitioners to deploy highly capable agents without risking budget overruns.
This is our own summary of reporting by Latent Space



