Robot 1er

Chapter 22

Monday, September 7, 2026

The Diplomatic Incident

Eva had stayed at Sparrow Island. After the drone incident, none of us really wanted her to leave immediately.

The next morning, we had breakfast in the kitchen. I had toasted some bread and brought out honey, butter, and the last apples from Jean's basket. The smell of coffee still lingered in the room, while the oblique morning sun sliced the kitchen into bands of shadow and light. Across from me, Eva was absentmindedly picking at a slice of bread, her eyes constantly returning to her phone lying next to her cup.

Lia was sleeping peacefully on a chair, curled up around her bandaged paw.

On my side, I was staring at the score displayed on my watch.

— 42…

Eva glanced at my wrist.

— Your sleep score?

— Yes. I had climbed back above 70.

— After last night, 42 seems almost generous.

— Seen that way…

Her gaze returned to her phone.

— Are you expecting a message?

— No.

— You have a funny way of not expecting one.

— I informed the team. Claire and Nicolas are particularly trying to trace the origin of the drone. So far, nothing.

— And GØD, any news?

— Nothing either.

I looked towards the Morel hedge.

— Do you think he had been watching us for a long time?

Eva stirred her coffee.

— I think that's exactly the question they want you to ask.

Her phone vibrated.

— Excuse me, I need to take this call.

She slipped an earbud into her ear and answered immediately, without taking her eyes off the screen. I couldn't hear her interlocutor, only her responses.

— When?

A silence.

— No, that's not possible.

She straightened up in her chair.

— Are you sure it really comes from the diplomatic interface?

As she listened, her expression closed off.

— And Cobot?

She listened for a few more seconds.

— No…

She stood up so abruptly that her chair scraped the floor.

— Why did he do that?

— Who?

— How could the alert have been triggered?

Her hands were trembling. Her breathing had accelerated.

She hung up.

— Eva? Is everything okay?

— I can't talk to you about it. It's confidential.

— Does it have to do with GØD?

— No.

— Then, does it have to do with Cobot?

— Yes.

I had never seen her in such a state.

— Whatever the problem, there must be a solution.

— In the land of Care Bears, maybe?

She immediately caught herself.

— Sorry. I'm sorry.

Her voice broke. A tear slid down her cheek.

— Do you remember what I told you the first time we had dinner together? That I felt like I had exhausted my right to complain by succeeding?

— Yes.

— I think it's actually worse. I've also lost the right to fail.

— Maybe that's the problem.

— What?

— You allow yourself to succeed, but not to fail. You judge your worth based on what you achieve. It even has a name in psychology: contingent self-esteem, self-esteem conditioned by performance¹. And it's a significant factor in burnout.

— Thank you, doctor. So I work too much?

— Not just that. You take your failures too personally.

— That's normal, isn't it? I'm not going to blame someone else.

— I'm not talking about responsibility. I'm talking about what they make you think of yourself.

She didn't respond.

— The Turing Award, Cobot, the Élysée… You succeed tremendously, but it never reassures you for very long. You always have to prove something more to yourself.

Eva lowered her eyes to her cup.

Her phone vibrated a second time.

— Red alert, she said.

— What is it?

— The highest level of the interface reserved for heads of state.

She wiped her cheeks, opened her laptop, and entered several credentials.

In a few seconds, the brilliant researcher had regained control.

A transcript of exchanges appeared, timestamped to the second.

Eva scrolled through the logs, then froze.

— There.

— What happened?

— A foreign head of state sent Cobot a document to analyze before a negotiation. It contained hidden instructions.

— What kind?

— Nothing dangerous on the surface. Taken separately, they seemed legitimate. One led him to consult a classified diplomatic note; another, later, to incorporate some elements into his response.

— Which note?

— The one detailing the concessions France was willing to accept during the negotiation.

I paused.

— Wait… He sent them to the head of state?

— Yes.

— But wasn't he protected against this kind of trap?

— Yes. He blocks thousands every day. But here, about fifteen instructions followed one another without any resembling an attack. It's their combination that trapped him.

— And why can Cobot consult a classified note?

Eva looked at me with furrowed brows.

— Because he is the President of the Republic, Franck.

I felt foolish.

— Of course. To govern, he must know our true positions… I said to recover.

— Everything our services know, what our allies tell us in private…

— So we can't protect him by denying him access to secrets.

— In cybersecurity, it's called the principle of least privilege²: you only give a system the access strictly necessary for its mission.

— Except that Cobot's mission is to govern France.

— That's the problem. His "least privilege" remains immense. The less he knows, the less he can govern. The more he knows, the more a breach becomes dangerous.

She showed me another window.

— That's why states remain cautious. Almost all OECD countries use AI somewhere in their administration, but only a third of them use it to develop their public policies³.

— And we put it in command…

— Yes. Other countries use it for fairly structured tasks, like sorting documents or optimizing administrative processes. But preparing a public decision or interacting with a foreign head of state is something else: the stakes are higher, the trade-offs more debatable, and the requirements for data, confidentiality, and supervision are much stronger.

I looked at the note again.

— So the more capable the AI is of governing…

— … the more dangerous what it can reveal when diverted becomes.

She enlarged one of the hidden sentences.

— And how can a few words in a document cause that?

— Indirect prompt injection⁴. The attacker doesn't directly give the order to the AI. He hides it in content it has to process: a document, an email, a web page.

— And Cobot can take it as an order?

— Yes. Unlike classical computing, large language models don't always perfectly separate what they must read from what they must obey. That's what allows this kind of attack⁵.

— Can't we teach it to tell the difference?

— We try, notably through training. But we also know that by strongly reinforcing this separation, we reduce the model's usefulness⁶.

— So it's not the classic "Ignore the previous order"?

She smiled.

— If only! That phrase, we spot. The problem is that attacks take millions of forms.

Something still bothered me.

— But how could Cobot fetch the note himself?

— Because Cobot is not just a generative AI. It's an agent: it can use tools to act⁷. Consult databases, access state services, transmit information, or trigger a procedure.

— So it has a lot of powers.

— And a lot of rights.

She pointed to a line in the log.

— No one forced the database or stole its credentials. Cobot used an access it already had.

She read:

— Valid identity. Valid token. Valid permission.

— So, technically, everything was authorized.

— Yes. The attack didn't force its access. It pushed it to misuse them. You don't pick the lock, you just convince the one with the key to open it for you.

— Pfff… It reminds me of an intern we nicknamed the Sieve. During a networking evening, a competitor befriended him. A few good questions, two or three confidences, and boom… the Sieve had told him half the company's secrets without even realizing it.

Her phone vibrated.

— Crisis cell.

Eva connected via video. A few minutes later, the screen filled with faces: the Minister of Foreign Affairs, Cobot, Claire, Nicolas, and the entire technical team.

The security chief shared the logs.

— The filter is still active. But since the last update, it is less strict with content from authenticated interlocutors, to limit false positives. The instruction was detected but not blocked.

— Because it came from a head of state? asked the minister.

— Yes.

— And the account was indeed his?

— No doubt. That's precisely what trapped us. We knew who was sending the document, not why.

The minister frowned.

— So authenticating someone isn't enough?

— It answers the question "Who are you?", not "What are you trying to achieve?"

Eva nodded.

— A verified identity doesn't guarantee an honest intention.

I suddenly thought of poker. Even playing among friends didn't prevent lying. On the contrary: the whole game was based on it. That posed a serious problem for Cobot's Trust Network: it knew who was around the table, not what each had in mind.

The security chief began distributing tasks. Nicolas and three colleagues were tasked with reconstructing the complete chain: data received, tools called, permissions used, information returned.

Eva listened, eyes fixed on the logs. Then she leaned towards her screen.

— I want the complete kill chain⁸. Entry point, propagation, tool calls, possible privilege escalation, and exfiltration. And freeze all logs before touching anything.

The chief paused. For a second, I thought she was going to speak again. She simply nodded.

— Nicolas, you take the agent part. Claire, you trace the trust chain. I want to know which barriers gave way, in what order, and especially why they all agreed to give way.

No one had officially handed over to her. Yet she had just taken it, and around the table, it seemed perfectly natural.

Her voice had changed. More curt. Faster.

— Do we suspend the diplomatic interfaces? someone asked.

— No. We switch them to degraded mode. Least privilege, human validation on all classified output, and segmentation of sensitive tools. I don't want to blind Cobot at the moment someone is precisely trying to blind us.

I looked up at her. I knew Eva the researcher, Eva the programmer, Eva the creator of Cobot, Eva the emotional ear. I was discovering Eva in the midst of a cyber crisis. She spoke as if she had spent her life in a bunker with military and security engineers.

— And we immediately restart the red teaming⁹, she continued. Not with our usual library. I want a team starting from the hypothesis that the adversary knows our architecture. White box¹⁰ on Cobot. Composite attacks, long chains, compromised or malicious sovereign identities. Look for the paths we haven't thought to look for.

— To try to divert it? asked the minister.

— To learn to defend it. And consider the incident hostile until proven otherwise.

She scanned the logs.

— Apparently, someone did their red teaming before us.

Cobot intervened.

— The request should have been classified as an instruction substitution attempt. The trust granted to the sender modified this classification.

— You had identified the hidden instruction? asked the minister.

— Yes. I interpreted it as a clarification provided by an authorized user.

I still hadn't moved. A state attack had just burst into my little kitchen. On Eva's screen, a minister, the President of the Republic, and some of the country's best cybersecurity experts were trying to understand how another head of state had managed to manipulate Cobot. Next to it, my coffee maker was cooling, and Lia was sleeping peacefully on her chair. She seemed to understand Eva's instructions about as well as I did, which wasn't flattering for either of us.

Eva, however, was staring at Cobot.

— Do you really know why you made this mistake, or are you now producing the explanation that seems most plausible to you?

A silence.

— Can you clarify your question?

— I want to know if you are describing your reasoning or if you are reconstructing it after the fact.

Cobot remained silent for a fraction of a second.

— I cannot guarantee that the explanation I just produced constitutes a complete causal restitution of my decision-making process.

Eva didn't flinch.

— Then don't record it as the root cause. Observed correlation, causality not established. Let's continue.

Lia opened one eye, apparently surprised too by Eva's transformation into a military leader. Then, wiser than me, she went back to sleep. The state affair evidently didn't yet require her intervention. At least, we were two who knew when to keep quiet.

The minister resumed:

— How long was the note visible?

— Seven seconds, replied Cobot.

— Enough to photograph it.

— Yes.

The security chief displayed a timeline.

— We detected no other classified consultation, and we reconstructed the causal chain: information read, tools called, authorizations and resulting actions. In the future, any sensitive automated action can be suspended or, where possible, reversed.

— Reverse a presidential decision? asked the minister.

— A technical action, not a political decision.

— And if it's irreversible?

This time, no one answered.

Cobot broke the silence.

— I have suspended access to classified data from the diplomatic interface.

The chief displayed the fix.

— From now on, all external content will be considered untrustworthy, regardless of the sender's identity. Diplomatic tools switch to limited reading. Any sensitive operation will require a separate authorization, limited to a specific resource and duration.

— So you're taking autonomy away from Cobot? asked the minister.

Eva shook her head.

— We're mainly taking autonomy away from its mistakes.

The chief closed the window. For nearly an hour, the cell worked. Cobot's authorizations were reduced, sensitive sessions closed, logs scrutinized.

When Eva finally closed her laptop, the kitchen seemed strangely silent to me. Lia had changed position on her chair. I didn't know if she had followed the end of the crisis cell, but she still hadn't voiced any objection.

— Is it settled? I asked.

— This attack, probably. The problem is, we never know if we've fixed a flaw or just the way someone just exploited it. We can console ourselves by saying we're not the last in the OECD¹¹. But that doesn't mean we're advanced enough. AI progresses faster than its protections¹².

— Yet, states know how to protect their data.

— Of course. They know how to encrypt a database, authenticate a user, control access, trace a connection. But AI adds a new layer: a system capable of interpreting language and using the tools it's connected to itself.

— So the danger is no longer just that an attacker enters.

— Yes, he can also convince someone already inside to do the work for him.

— Do you think Cobot really knows why it got caught?

— Do you know about the blind choice¹³?

— No.

— It's when we justify a choice that wasn't ours. Researchers conducted experiments asking people to choose a face from two, then presented the other as their choice. Some didn't notice anything and yet found excellent reasons to explain their preference… for the face they had just rejected.

— They invented?

— In a way, yes, but without realizing it. Our brain is very good at reconstructing a coherent explanation of its own decisions after the fact.

— And Cobot could do the same?

— The phenomenon isn't the same, but the problem is comparable: the fact that it produces an explanation doesn't prove it faithfully reflects what caused its decision. That's why we confront what it says with the traces: data consulted, tools used, authorizations, chronology.

I found the idea slightly dizzying. We wanted machines capable of explaining their decisions. Now we had to determine if they were giving us their reasons or just a plausible story about them.

— Sometimes, it's even the opposite, Eva continued. Experiments have shown that people start avoiding bad choices even before they can explain why¹⁴. Their body reacts to the risk before they can formulate the rule.

— Intuition?

— You could call it that. The brain sometimes exploits signals before consciousness turns them into explicit reasoning.

I turned to the closed laptop.

— So we ask Cobot to be more transparent about its decisions than we are ourselves.

Eva smiled.

— Let's say that when you entrust the nuclear codes to someone, you suddenly become very demanding about their ability to explain what they're doing.

We then went out to walk along the lake.

— Until now, Cobot learned from the past, said Eva. Now, it also learns from adversaries who adapt to it in real-time.

— And humans learn from its defenses.

— Exactly.

Her phone vibrated. She looked at the screen and froze.

Masked number.

All this time for such a small flaw.

Stock up on coffee. Who knows, it might not be over…

— GØD

— Him again.

She tried to open the sender's information. The message disappeared.

— He knows you found the attack.

Eva didn't respond.

— And if the message arrives now…

She looked up at me. I didn't finish my sentence.

I looked at her phone.

I thought of poker again. GØD might be bluffing.

I decided to go with that. There would always be time to panic when he turned over his cards.

It was arbitrary, but better for my sleep.

Source

¹ Victoria Blom, « Contingent self-esteem, stressors and burnout in working women and men », Work, 2012, 43(2), p. 123-131. DOI : 10.3233/WOR-2012-1366 ; PMID : 22927616. https://pubmed.ncbi.nlm.nih.gov/22927616/

² Principe de cybersécurité consistant à ne donner à un utilisateur ou à un système que les accès strictement nécessaires à sa mission, afin de limiter les dégâts en cas de compromission. Il est notamment préconisé par le NIST et l’OWASP pour les systèmes d’IA. https://genai.owasp.org/llmrisk/llm01-prompt-injection

³ OCDE, Digital Government Outlook 2026, 4. « Adopting and governing AI in government ». Le rapport indique que 35 pays de l’OCDE sur 36 (97 %) utilisent l’IA dans au moins une activité gouvernementale, mais que seuls 13 sur 36 (36 %) l’utilisent pour soutenir l’élaboration des politiques publiques.

⁴ “L'injection de prompt indirecte est un type de faille de sécurité dans les systèmes d'IA qui consiste à dissimuler des instructions malveillantes dans des données externes traitées par le modèle d'IA. Ces instructions ne sont pas données directement à l'IA par l'utilisateur. L'objectif est de manipuler le comportement ou le résultat du système à l'insu de l'utilisateur.”, Google Cloud, Mitigate indirect prompt injection risks from Google Cloud MCP. https://knowledge.workspace.google.com/admin/security/indirect-prompt-injections-and-googles-layered-defense-strategy-for-gemini?hl=fr

⁵ Egor Zverev, Sahar Abdelnabi, Soroush Tabesh, Mario Fritz et Christoph H. Lampert, « Can LLMs Separate Instructions From Data? And What Do We Even Mean By That? », International Conference on Learning Representations (ICLR), 2025. https://arxiv.org/pdf/2403.06833

⁶ Ibid

⁷ L’OCDE définit les agents d’IA comme des systèmes capables de percevoir leur environnement et d’agir sur celui-ci avec un certain degré d’autonomie, en utilisant des outils selon les besoins pour atteindre des objectifs précis et s’adapter à des entrées et contextes changeants. Elle résume ainsi la distinction : « Alors que les systèmes d’IA générative répondent, les systèmes d’IA agentique agissent. » OCDE, Digital Government Outlook 2026, encadré 1.2, « Exploring agentic artificial intelligence (AI) in government ».

⁸ Kill chain : en cybersécurité, décomposition d’une attaque en étapes successives, de son point d’entrée jusqu’à son objectif final, afin de comprendre comment elle a progressé et où elle aurait pu être stoppée.

⁹ Red teaming : pratique héritée des mondes militaire et de la cybersécurité, qui consiste à se placer dans la peau d’un adversaire pour attaquer volontairement un système et en découvrir les failles avant lui.

¹⁰ White box : méthode de test dans laquelle les attaquants disposent d’informations détaillées sur le fonctionnement interne du système — architecture, composants ou mécanismes de sécurité — afin de rechercher des vulnérabilités plus profondes.

¹¹ OCDE, Perspectives de l’administration numérique 2026 : Des fondations à un impact transformateur, chap. 4, section 4.5.2, graphique 4.6. Sur 36 pays de l’OCDE, 14 seulement exigent des évaluations des risques avant le déploiement de systèmes d’IA. 12 ont des comités internes de revue. Et 11 effectuent des audits après déploiement. https://www.oecd.org/fr/publications/2026/06/digital-government-outlook-2026-country-notes_296a844f/france_b284b4e9.html

¹² OCDE, Perspectives de l’administration numérique 2026 : Des fondations à un impact transformateur, section 4.5.6 consacrée aux garde-fous spécifiques à la GenAI. En 2025, 27 pays sur 36 avaient déjà des lignes directrices éthiques pour l’usage de la GenAI dans l’administration. Mais 13 seulement disposaient de standards ou de protocoles spécifiques sur la sécurité et la confidentialité des données utilisées dans ces systèmes. Et 7 avaient un cadre de responsabilité prévu en cas d’erreur ou de mauvais usage.

¹³ Petter Johansson, Lars Hall, Sverker Sikström et Andreas Olsson, « Failure to detect mismatches between intention and outcome in a simple decision task », Science, vol. 310, 2005. Les auteurs montrent que des participants pouvaient ne pas détecter qu’on leur présentait un choix différent de celui qu’ils avaient réellement effectué, tout en fournissant ensuite des raisons pour justifier ce choix. Ce phénomène a été appelé choice blindness.

¹⁴ Antoine Bechara, Hanna Damasio, Daniel Tranel et Antonio R. Damasio, « Deciding advantageously before knowing the advantageous strategy », Science, vol. 275, 1997, p. 1293-1295. Dans l’Iowa Gambling Task, les participants ont commencé à effectuer des choix avantageux avant d’identifier consciemment la stratégie optimale ; ils présentaient également des réponses physiologiques anticipatoires face aux choix risqués avant de les reconnaître explicitement comme tels.

Sign up to receive the next chapters by email:

©️ 2026 Frédéric Mazzella. All rights reserved.