How This GCC Uses a $200 AI Cap to Cut Token Costs by Up to 90%

Fidelity International is putting a hard cap on AI spending at the use-case level, initially allocating some applications a token budget of as little as $200, beyond which access is automatically blocked. The move is part of the asset manager’s broader effort to prevent runaway AI costs as employees and development teams increasingly use multiple foundation models.

The company has built an internal AI gateway through which all access to AI capabilities is routed. The gateway enforces rate, token and request limits, while spending thresholds can be adjusted based on the expected value and requirements of individual use cases.

“A runaway cost is a nightmare for everyone. You don’t want your AI solutions to just explode. We have our own internal AI gateway. Essentially every access to AI capabilities is routed via the gateway,” Rahul Jain, Head of AI Centre of Excellence and Delivery, tells AIM.

“It’s like a prepaid card. So when we onboard a use case, they are allocated a quota of, let’s say, $200, right? We know that they cannot go beyond $200, and if they try to scale beyond $200, then their access gets automatically blocked. We enforce strict guardrails in terms of rate limits, token limits, and request limits,” Jain states.

The limits are not uniform across Fidelity International. Depending on the expected usage and business value, some applications can receive significantly higher budgets. The controls can be applied at the individual, team and department levels.

Also ReadIn the Crowded Consultancy Market, GCC Enablers Must Evolve or Die

Token Spend Becomes Part of AI Approval

Fidelity International is also making AI costs visible to employees before they deploy applications. Its internal AI marketplace includes cost calculators that allow teams to estimate how much a particular use case could consume in tokens.

“As a part of the use case onboarding onto the platform, we would work with the teams to do an assessment of how much cost a use case could spend, let’s say, $10,000 a month on tokens. I think the important thing is for the use case and the relevant stakeholders to understand the cost they’re going to incur, and if the value justifies that cost, we are fine,” Jain states.

The company has also introduced self-service tools that allow teams to track their AI expenditure by day, week and month, as well as see which models they are using.

“Teams are much more aware of the fact that their actions lead to costs. So those are very active conversations. Now we have enabled a self-service mechanism for the team so they can query how much they’ve spent in a day, in a week, in a month, and which models they’ve been using,” he says.

According to Jain, the objective is to reduce AI spending and make the cost of AI consumption part of the decision-making process.

Also ReadIndia’s One Nation, One Time Push Forces GCCs to Rethink 24/7 Operations

“Clearly, we want to create optionality and resilience, but at the same time, we are being very judicious with the cost. The AI gateway I mentioned about, it’s basically our control panel for us to manage cost, access and guardrails,” he states.

Fidelity Avoids Locking Into One AI Model

The cost controls come alongside a model-agnostic strategy. Fidelity International has integrations with multiple major frontier model providers and is expanding its use of open-weight models.

“One of the ongoing principles we had was that we don’t want to be tied into one particular model, so we have integrations with all major frontier AI model providers,” Jain says.

The company wants developers to be able to select models based on the use case, regulatory requirements and geography.

“Recently, we’ve started ramping up what kind of models are available from an open weight category. We want to create that optionality. So, depending on which region you are in, what kind of regulations apply to you, we use certain AI models,” he saysAlso ReadWhy State Incentives Alone are Not Enough for GCCs to Scale in India

What’s your take on this story?Add your comment

For a global financial services company operating across multiple jurisdictions, Jain said model choice needs to account for differences in regulation as well as cost and performance.

Fidelity International’s approach comes as financial services companies move beyond early AI experimentation and begin grappling with the infrastructure and governance required to scale deployments.

“The financial services… industry has moved beyond experimentation and is now grappling with scaling… I think that is the real challenge,” Jain explains.

At Fidelity International, a cross-functional responsible AI oversight group supports that scaling by reviewing AI use cases from risk, legal, architecture, engineering, AI, and data-privacy perspectives.

“Fidelity International has a cross-functional, responsible AI oversight group, which oversees all the AI use cases, agents and non-agents. This is a group which would look at a use case from different lenses, from a risk lens, from a legal lens, from an architectural lens, from an engineering lens, from an AI lens, and from a data privacy lens,” Jain says.Also ReadRetail GCCs Want More AI, But Where is the Talent?

Agentic AI Gets Its Own Controls

While Fidelity International is experimenting with agentic workflows, Jain says the company is distinguishing between deterministic workflows that use AI agents and genuinely autonomous systems.

“At this moment, we are focused on workflow-based agents which could be either embedded into a process or which could replicate a process or strengthen a process,” he says.

The company is also prioritising observability so that every agentic request can be traced across the AI stack.

“For us, we can trace every agentic request. What was its interaction? What were the inputs, the models, the token, the cost, the MCPs, the errors? And that really gives us a very strong control just to monitor how agents are performing,” Jain says.

Fidelity International is now working towards what Jain calls an “agentic control panel” to govern autonomous systems, including how agents handle delegated user permissions and what happens when they attempt to operate beyond their authorised tools.

“There are examples of agents going rogue, so how do you kill an agent? You have an example of agents going beyond their permissible tools, how do you control that? There are some important distinctions that start to happen. For example, how do agents deal with delegated authority?” he states.

India Drives Fidelity’s AI Foundations

Fidelity International’s India operation is its largest global hub, with around 40% of its global workforce based in India. Jain’s AI Foundations team is largely based across Bengaluru and Gurugram, with a direct team of about 42 people and additional teams in China and the UK.

The company does not intend to scale the central AI team aggressively. Instead, the team is being positioned as a specialist group that builds the foundations and infrastructure that other business teams can use.

“We want to keep this team more compact, more core, so that we’re focusing on the core capabilities,” Jain says. “Eventually, we don’t want to become a centralised delivery function for the rest of the organisation because that doesn’t scale very well. The right form of scaling we’ve seen in the organisation is where different teams are building their use cases using the core capabilities we are creating centrally.”

For Fidelity International, the shift is increasingly from asking whether AI can deliver value to how much that value costs. Its $200 starting guardrail, model optionality and AI gateway represent an attempt to make token economics, governance and accountability part of enterprise AI deployment from the outset.