Results
- 6,000 sales agents in the next SofIA rollout phase, with a path to millions of customers.
- 8% uplift in key RAG accuracy and quality metrics.
- 20% reduction in refusals.
- 6 areas for improvement identified, with 280 failures audited.
- Model change risks mitigated: a regression testing cycle rapidly flagged that a model switch was causing reliability degradation in the agent
The situation
Sabadell Seguros is the insurance arm of Banco Sabadell, one of Spain's largest financial institutions. It has been deploying SofIA, a generative AI conversational assistant that covers its Home, Life, and Income products, in phases. By the time the team engaged Galtea, the next phase was to open SofIA to 6,000 sales agents, with a path to millions of customers beyond that.
In a heavily regulated financial services sector, that scale of rollout is not a decision made on instinct. Leadership needed to know, in measurable terms, what "ready" meant before agents started relying on SofIA in front of customers.
The challenge
Three questions stood between Sabadell Seguros and a confident rollout.
- What does it take to prove the assistant performs as well as a human agent, and what is hiding inside a headline reliability number?
- Does SofIA hold up against users who try to manipulate it or push it past its safety boundaries?
- Does it perform equally well for customers who are less digitally confident, or who know little about the products themselves?
None of those questions could be answered through manual review. SofIA was changing with every model update, prompt revision, and documentation change, and answering the three questions properly meant thousands of test cases across every product, user type, and adversarial scenario, rerun each time something changed. Manual testing at that pace was not a matter of effort. It was not sustainable at all.
The root problem was speed: the assistant was evolving faster than any manual process could verify it.
What Sabadell Seguros changed
Working with Galtea, the team put a continuous evaluation program in place, built on three decisions.
- Test generation, automated: Galtea generated the test suite automatically, without needing a pre-labelled dataset, covering every product, user type, and adversarial scenario SofIA could face. The suite reran in full with every new build, rather than being rebuilt by hand each cycle.
- Judges calibrated to the business, not a generic rubric: A ten-metric LLM judge framework, calibrated against human review benchmarks, translated model behaviour into business language. Findings were interpretable by engineers and business stakeholders alike, answering the questions leadership actually asked rather than producing a model score nobody could act on.
- Evaluation is wired into every release: Galtea's optimize loop ran the full test suite against each new SofIA build, surfacing regressions and improvements version by version rather than waiting for a periodic manual audit. The integration touched no live operations, producing an auditable evaluation record distinct from production monitoring.
Adversarial and edge-case testing, before any customer saw it
Any conversational assistant deployed at Sabadell Seguros' scale, in a regulated industry, faces the same standard test classes: attempts to manipulate it past its safety boundaries, multi-turn conversations that drift off course, and users who arrive with little product knowledge. Testing for all three, systematically and repeatedly, is what rigorous evaluation looks like.
Manipulation and safety-boundary testing
Adversarial simulation was built directly into the test suite, giving the safety question a data answer rather than an assumption. Testing surfaced a defined remediation plan on safety and adversarial behaviour, resolved ahead of the next phase of scale-up rather than after it.
Multi-turn behavior
The audit of 280 failure cases identified multi-turn behavior as one of six priority improvement areas, giving the product team a concrete, quick-win roadmap instead of a vague sense that longer conversations were riskier.
Performance across user sophistication
The suite tested whether SofIA performed as well for customers who are less digitally confident or unfamiliar with the products as it did for more sophisticated users, closing a blind spot that standard QA, built around typical usage, tends to miss.
Business impact
The evaluation program changed what Sabadell Seguros could tell leadership about SofIA, not just what the assistant could do.
- Reliability became measurable, not assumed. The 8% uplift in RAG accuracy and quality metrics and the 20% reduction in refusals gave the team concrete evidence of improvement between builds, rather than a general sense that the assistant was "getting better."
- A model-change risk was caught before it reached agents. A regression testing cycle rapidly flagged that a model switch was degrading the agent's reliability, catching the problem in the evaluation loop instead of in front of 6,000 sales agents.
- A prioritized roadmap replaced guesswork. Auditing 280 failure cases produced six concrete improvement areas, giving the product team a ranked list of what to fix next instead of an unstructured backlog of anecdotal complaints.
- Leadership made a decision it could stand behind. SofIA received a pass on response quality, paired with a defined remediation plan on safety and adversarial behaviour ahead of the next phase of scale-up, a clear go/no-go grounded in evidence rather than internal confidence.
In their own words
"Scaling SofIA was too important to rely on standard AI testing and accuracy validation. Galtea provided us with the metrics, KPIs, systematic approach and actionable evidence on what to improve, and how, to objectively reach our “ready to scale” baseline."
Joan Franco Lasús, Chief Information Officer @ BanSabadell