Lab notebook · AI

The lab , in the open.

Real cases told as model cards : context, eval, metrics. What we would ask of anyone who says "we did AI".

EVAL · 2026-04

RAG over 14,200 technical documents

91% accuracy with 100% citation traceability. Engineer onboarding reduced from 12 to 4 weeks.

91%
Precision
100%
Appointment
EVAL · 2026-03

Vision: welding defects

99.2% recall on critical defects, 2.1% false positives. Line throughput +14%.

99,2%
Recall
+14%
Throughput
NOTE · 2026-03

The inference cost that kills projects

Why the metric everyone forgets sinks AI projects within six months, and how we measure it from day one.

€/1k
From sprint 1
NOTE · 2026-02

"You don't need AI here"

Three cases where the honest answer was a process, a training programme or a better ERP. Selling what is not needed destroys trust.

3
Cases

All articles

23 June 2026

AI Agents in Production: A Reliability Guide

An agent that works in a demo isn't necessarily reliable: taking it into production requires its own identity, least-privilege permissions, reproducible evaluation and human oversight. This guide covers the architecture, the threats and the 90-day plan to get there.

Read article
June 25, 2026

AI Voice Agents: Privacy and Quality

An AI voice agent is not a chatbot that talks: it's telephony, transcription, reasoning and speech synthesis chained together, and each layer adds latency, errors and data to protect. This guide explains how to design it with privacy by design, latency measured in percentiles and human escalation always available.

Read article
26 June 2026

AI Act and AI agents: technical obligations

Calling an application an “agent” does not change its legal obligations: what matters is the organisation's role, the purpose and the risk of the system. This technical guide explains how to classify the use case, assign responsibilities and prepare inventory, traceability and oversight before going into production.

Read article
28 June 2026

n8n AI Automation: A Secure Architecture

Taking an n8n workflow with AI to production takes far more than a correct run: protected credentials, validated inputs, idempotent retries and human approval on risky actions. This guide walks through the architecture, the controls and the 60-day plan to do it securely.

Read article
30 June 2026

AI Copilot for Advisory Firms: Uses and Limits

An AI copilot in an advisory firm speeds up document classification, regulatory research and drafts of emails and reports, but it must never file, send or decide on behalf of the professional. This guide covers use cases, architecture, permissions, evaluation and common mistakes for a well-governed rollout.

Read article
5 July 2026

Sovereign AI for SMBs: Local, Private, or Hybrid

Sovereign AI isn't a choice between cloud and on-premise — it's about proving who controls the data, the keys, and the ability to walk away. This guide helps SMBs choose between a public API, private cloud, on-premise, or hybrid setup based on the real risk of each use case.

Read article
6 July 2026

LLMOps for SMEs: evaluate and monitor

LLMOps doesn't require a massive platform: with a clear inventory, systematic evaluation and controlled versioning, an SME can deploy its AI applications with security, observability and cost under control.

Read article
8 July 2026

MCP for Business: Connecting Agents with Control

Connecting AI agents to email, files or databases through MCP multiplies your attack surface if the server isn't treated as a privileged API. Here's how to roll it out with inventory, authorization, least-privilege permissions, validation and human approval.

Read article
10 July 2026

Intelligent Invoice Processing With Traceability

Intelligent invoice processing turns every document into a verifiable proposal, never a blind accounting entry or payment: here's how to keep validation, traceability, duplicate control and human review at every stage of the workflow.

Read article
13 July 2026

How to Measure AI ROI in an SME

AI ROI isn't proven by counting prompts or minutes saved — it's proven by comparing a real process before and after, using a baseline, total cost, quality, adoption and risk. A practical guide to designing the pilot, doing the maths properly and deciding whether to scale on data, not promises.

Read article
010203Next →