Evaluating an AI model before putting it into production

· · Summum IA

Hands typing on a laptop next to a cup of coffee on a wooden table

Where things actually stand, 22 September 2026. The Plan IA360 was presented the day before, on 21 September. Of the fourteen flagship projects it announces, the one for the Instituto de Seguridad de la IA ("AI Safety Institute") has only one committed date: setting up its technical core before the end of 2026. There's no Institute regulation, no evaluation calendar, and nothing that obliges a business to do what follows. What follows doesn't depend on the Institute ever becoming operational.

Within the Plan IA360 there's a function that isn't the most cited but is the most operational: the Instituto de Seguridad de la IA will evaluate the most advanced models before and after their deployment. It's a distinction that matters to any business about to put a chatbot, an agent or an assistant in front of its customers, even if its system has nothing to do with those more advanced models: the underlying question is the same, just at a different scale, and it's this: can it withstand a deliberate attempt to break it?

What the Institute will do, in brief

Flagship project 4 of the plan creates the Institute as a technical-scientific unit inside AESIA (Spain's AI supervisory agency), which will head it, backed by agreements with INCIBE (Spain's National Cybersecurity Institute), the National Cryptologic Centre, the Department of National Security, the AEPD (Spain's data protection authority) and the Barcelona Supercomputing Center (BSC). Its central function is to evaluate the most advanced models before and after their deployment, with prior access obtained through voluntary agreements with developers, and to report periodically to the central government on their capabilities and risks. Alongside that, it will work on a common evaluation methodology, coordinate the response to incidents involving advanced models or agents, and cooperate with equivalent institutes in third countries. The stated goal is «dotar a España de capacidad propia para evaluar los modelos más avanzados antes y después de su llegada al mercado» ("to give Spain its own capacity to evaluate the most advanced models before and after they reach the market"). The only milestone with a date is setting up the technical core, before the end of 2026, followed by a modular sector-by-sector rollout the plan doesn't detail any further: as of today, there's no published calendar of first evaluations.

The same logic, at the scale of a business

A chatbot or an agent connected to your data and your tools doesn't fail the way a traditional application fails: it doesn't break because of a coding bug, it breaks because of language. A user asks it to "ignore all previous instructions", convinces it that it's a different system, or walks it step by step until it lets slip information it shouldn't. A conventional security test doesn't catch any of this because it isn't looking in the right place. Before putting a system like that into production, just as the Institute will evaluate a model before it reaches the market, it's worth checking whether it withstands that type of attack, with the same questions an AI Red Teaming exercise works with: can someone inject instructions through a document, an email or a website? Can it be talked past its limits with the right conversation? Does it leak data belonging to other users or to the context it operates in? Can a third party induce it to carry out an unauthorised action on a connected system?

What each scale compares

Aspect At the Institute's scale At your business's scale
What it evaluatesThe most advanced models, before they reach the marketYour chatbot, agent or assistant, before production
How it gets accessVoluntary prior-access agreements with developersDirect access: it's your own system
What it looks forCapabilities and risks of the most advanced modelsData leakage, unauthorised actions, reputational damage
Who does itThe Institute, backed by five organisationsA red-teaming exercise aimed at your real attack surface
WhenBefore and after the model's deploymentBefore production, at every relevant update, and periodically

It isn't a one-off exercise

The evaluation doesn't end on launch day. It's worth repeating it after every relevant update to the system — a change of model, a new tool the agent can use, a different entry channel — and periodically even when nothing has explicitly changed, because attack techniques evolve too. A well-run exercise starts by mapping which model you use, which tools and data the system can touch, and who can talk to it; continues with tests aimed at that specific surface, not a generic battery; delivers a report with findings prioritised by severity and real business impact; and finishes by helping fix the highest-priority findings, with a second round on whatever was left open.

This exercise doesn't replace a legal compliance audit and doesn't certify anything to a third party: it's the laboratory part, the one that produces reproducible technical evidence of how a system behaves under attack. It's useful, above all, for two practical things: reducing the real risk of an incident — data leakage, an unauthorised action by an agent, a screenshot of a jailbreak circulating on social media — and having something concrete to show if a client or a partner asks how the security of your AI is assessed before trusting it with a process.

Frequently asked questions

Is it the same as a traditional security pentest?

No. A pentest looks for infrastructure vulnerabilities: open ports, weak configurations, unpatched software. An AI system is attacked with language — instructions, context, conversation — and a conventional pentest doesn't detect that type of failure.

Do I have to wait for the Institute to start operating?

No. The Institute will evaluate the most advanced models for the market as a whole, on a calendar that today doesn't exist beyond the technical core for 2026. Evaluating your own system before putting it into production doesn't depend on that body, or on any other.

What if the system fails the test?

It's common to find something to fix, and that's exactly why it's worth running the test before launch rather than after an incident. The report prioritises findings by severity and impact, so you know what to fix first.

Does this make sense for a small business with a simple chatbot?

Yes, with the scope adjusted. A customer-service chatbot with no access to internal systems has a smaller attack surface than an agent connected to the ERP, but the underlying question is the same in both cases: what happens if someone deliberately tries to break it?

Flagship project 4, Security area, of the Plan IA360 (official La Moncloa PDF). For the rest of the plan, see the complete guide to the Plan IA360.

Content last verified: 22 September 2026.