Agentic AI and deterministic answers: Inside Numerix's AI Day
Many financial institutions are testing agentic AI. Far fewer run it in production, where results must be dependable and auditable. That gap was the focus of a recent Numerix event, AI Day: Elevate Intelligence, which brought together Numerix experts, clients, and AWS to discuss how agentic AI fits into quantitative finance workflows and use cases.
Across a welcome address, five product sessions, a guest lecture from AWS, and a client panel, a clear message kept resurfacing. Agentic AI is a useful new way to interact with quantitative analytics, but for now it should not be responsible for generating pricing and risk numbers. Below is how that argument played out across Numerix products, including NxCore Analytics Services (NAS), CrossAsset, and Oneview, and why speakers kept coming back to the concept of trust.
Opening remarks from Numerix leadership
Numerix CEO Manny Conti opened with his view of what changes for quant teams, which he described as more than just productivity gains. In his account, AI expands who can ask a pricing or risk question, how often they ask it, and at what scale.
Numerix’s Chief Product Officer, Satyam Kancharla, pushed the point further. AI, he described, is well suited to judgment and interpretation, while pricing and risk still depend on deterministic calculations. A generative model can produce a plausible number, but it cannot reliably produce the sensitivities, Greeks, and hedges a trading desk or risk team needs, because those require a calculation engine rather than a guess. The line between what AI decides and what a deterministic engine computes was a key theme in sessions for the remainder of the day.
Layering agentic access onto engines that already exist
Evidence that Numerix is building on top of its analytics core, not replacing it, was showcased in two live demonstrations.
Mayank Nanda, SVP of Risk Analytics at Numerix, showed NxCore Analytics Services (NAS), Numerix’s cloud API for pricing and risk, connected to an agentic AI model (Claude) through an MCP (Model Context Protocol) server. In minutes, the connected agent parsed a term sheet, priced an autocallable structured note, and solved for its par rate. Assembling the same work by hand would typically take a desk considerably longer. Nanda summarized his demo by stating, “AI decides what to ask. NAS decides the answer.” NAS’s roadmap keeps building on that separation, with skills and tools capabilities scheduled for the end of 2026 and hosted AI assistants targeted for the first quarter of 2027.
Ping Sun, SVP and Head of Quantitative Research at Numerix, then built a full pricing template for a dual digital note, an exotic structure suggested by the audience on the spot, using Claude plus a knowledge base built from CrossAsset’s own reference guides and scripting documentation. In roughly 12 minutes, Claude produced a Python SDK template, a payoff script, a calibrated model, and documentation. Because the knowledge base was built from CrossAsset, the output followed the library’s existing conventions for pricing the product.
Trust as the limit on adoption
Alvin Huang, a financial services specialist at AWS, brought an outside view drawn from production AI deployments at firms including Bridgewater, Jefferies, and Nasdaq. He warned that if a domain can’t tolerate hallucinations, then AI may not be the right solution. He also cautioned against mistaking speed for value, relaying a comparison that reaching the stoplight faster does not tell you if you’ve gained anything meaningful.
That skepticism carried into the closing panel, moderated by Numerix’s Chief Marketing Office James Jockle, featuring Numerix SVP Andy McClelland, Ping Sun, and Alex Marion, a principal at Apollo Global Management. Marion drew a distinction that is central to Numerix’s approach. “Authorship is not assurance,” he said. “I don’t care who produces the output. What I care about is if I have the ability to check it.” He was very concerned about an AI model that is “confidently right, but for the wrong reasons,” where the output looks correct and can pass a casual review, but no one has checked it against a deterministic source.
What an agentic risk analyst looks like in practice
Luyao Yu, a Numerix financial engineer, closed the product demonstrations with Numerix Oneview. Querying the platform in natural language with a different agentic AI model (GitHub Copilot), Yu requested a portfolio summary, built a geopolitical stress scenario, and traced a value-at-risk (VaR) spike from the book level down to the specific sensitivity driving it. Along the way, the agent’s analysis surfaced and corrected a data lag issue, without any prompting from Luyao. The same session covered XVA reporting and RFQ execution, both queried through the same chat interface rather than separate manual reports.
The one insight that each presenter emphasized was the critical importance of having human experts review all outputs from AI queries. When a multi-day task takes minutes, the reviewer’s time shifts from producing the answer to checking it thoroughly.
Where the next AI gains are expected
An internal survey referenced during the panel found that close to four in 10 respondents already report three to five times productivity gains from AI in their existing workflows. Asked where that goes next, the panelists called out specific, narrower areas of work.
Marion expects the biggest near-term gains in operations, which may include agents continuously verifying massive, complex flows like collateral, hedging and intercompany cash movements, rather than trading or pricing. Sun expects AI to handle a growing share of harder edge cases with less manual intervention. McClelland pointed to automating the diagnosis of failures, such as tracing why an XVA calculation may have broken, and faster model testing.
Get deeper insight
Numerix’s white paper Trust by Task examines the boundary between agentic AI and deterministic analytics engines in more depth. It sorts quantitative finance tasks into three groups, covering where AI reasoning has earned trust, where that trust is conditional, and where human judgment should still lead.
Trust by Task is the first paper in the three-part Trust, Verified series. In the research, six senior Numerix practitioners disclose which day-to-day quantitative tasks they now trust AI reasoning to handle, and the areas where they hold back.
Frequently Asked Questions
Q1: How can agentic AI deliver pricing and risk outputs that trading desks can actually trust?
A generative model can produce a number that sounds plausible, but a trading desk needs proof it is correct before acting on it. Numerix Chief Product Officer Satyam Kancharla told the AI Day: Elevate Intelligence audience that AI suits judgment and interpretation, while pricing and risk still require a deterministic calculation engine. SVP Mayank Nanda summed up the division on Numerix's NxCore Analytics Services (NAS): "AI decides what to ask. NAS decides the answer."
Q2: Can an AI assistant actually price a structured note from a term sheet in real time?
Pricing an exotic structure from a term sheet by hand typically takes a desk far longer than a few minutes. In a live demonstration at Numerix's AI Day, SVP Mayank Nanda connected NxCore Analytics Services (NAS) to Claude through an MCP (Model Context Protocol) server. The assistant read the term sheet and priced an autocallable structured note, then solved for its par rate. NAS plans to add skills and tools capabilities by the end of 2026, with hosted AI assistants following in the first quarter of 2027.
Q3: How should quant teams verify AI-generated outputs instead of trusting them outright?
An output can look correct and still be wrong, and nobody catches the error without checking it against a deterministic source. At Numerix's AI Day, Apollo Global Management principal Alex Marion put it bluntly, "I don't care who produces the output. What I care about is if I have the ability to check it." Numerix's white paper, Trust by Task, sorts quantitative finance tasks into three groups to show where AI reasoning has earned trust and where human judgment should still lead.