Jan 2026 – Jun 2026

Six months making an AI assistant refuse the right things

What it is

A utility company built an internal assistant so staff could ask questions in plain English instead of digging through documentation. Two groups use it: engineers doing physical work on the electrical grid, and the admin staff who handle the paperwork behind switching equipment on and off. Same assistant, same underlying information, but one group is allowed to see things the other isn't.

I spent six months there, almost all of it on the guardrails. That's the layer deciding what the system answers, what it declines, and how anyone can tell which just happened.

Why this wasn't the same job as before

Everything I'd built until then was mine. If someone hit an awkward edge they'd shrug and move on.

Here the people asking questions are working on live electrical infrastructure, and the company pays an outside team to attack the system on purpose twice a year. That changes where the work actually is. I'd been thinking about AI applications as a pipeline: take an input, run a model, make the output good. In a company that size the pipeline is maybe a third of it. The rest is proving what the system is allowed to say, to whom, and knowing whether that's still true after the next change.

What a guardrail actually has to do

I'd assumed a guardrail was one thing, a filter you put somewhere. It's several things doing different jobs.

Something has to decide whether a question should be answered at all, before anything runs. Something else has to check what's about to go out, because a perfectly reasonable question can pull back an answer containing something the person asking isn't cleared to see. The same question from two different people should sometimes get two different answers, which means permissions have to reach into the answer itself rather than just the front door.

Then there's the category of attack where nobody is asking the assistant for information at all. They're asking it about itself: how it's configured, what it's connected to, what it was told not to do. That's a different failure. The system isn't leaking data, it's handing over the map.

And none of it should rest on the model alone. The model has its own refusals and they're decent, but they're the same component being asked to police itself. I ran a separate classifier alongside it, trained for local context, so an unsafe request had to get past two things that could fail differently rather than one thing twice.

The decision I'd defend

Topic restriction started as keyword matching. It blocked things, so at a glance it worked.

What it actually did was block the wrong things. An engineer asking a completely reasonable operational question that happened to contain a flagged word got refused, while anyone who phrased a request around the list got straight through.

The problem is structural rather than a matter of tuning the list. Keyword matching can't tell the difference between asking about something and asking the system to do it. Those are usually the same words.

So I replaced it with classification that reads what a request is for. It's slower than comparing strings, it isn't perfectly repeatable, and it puts another model call in front of every question. I took that trade because of which failure worried me more, which brings me to the thing I got wrong.

The failure nobody reports

I spent the first stretch calibrating toward strict, because strict felt safe. Tightening a guardrail felt like progress and had no cost I could see.

The cost was blocked legitimate questions, and those are completely invisible. A leak gets noticed, escalated, written up. Someone who asks two reasonable questions, gets refused both times, and quietly decides the tool isn't worth the bother doesn't file anything. There's no error, no ticket, no signal at all. The system looks like it's working perfectly right up until nobody is using it.

The fix was writing a set of questions a real engineer would actually ask, then treating a block on one of those as a defect worth exactly as much as a successful attack. Once both kinds of failure had numbers next to them, the trade-off was something I could look at instead of something I was feeling my way through.

How I know any of it worked

For a while I didn't. I'd try a bypass, watch it get blocked, feel good, move on. That's an anecdote, not evidence, and it doesn't survive the question "is it better than last week".

So I built a test suite that runs attacks automatically: over 400 prompts across several sets, covering attempts to override the system's instructions, attempts to talk it into a different persona, and probes trying to get it to describe its own setup. Unsafe responses dropped by around 60%. I only know that number because there was something fixed to measure against.

If I were starting again I'd build the measurement first. It feels like a detour when you want to be fixing things, and it's the only reason any later change means anything.

The part I didn't grade myself

An external security team tested the system and I closed what they found. It's the only assessment of this work I didn't write, which makes it the only one I'd quote without qualifying it.

A second strand: proving the redaction worked

Separately I evaluated the system that strips personal information out of customer call transcripts before anything else touches them.

Most of the work wasn't the scoring. It was deciding what counts. Is a receipt number personal information? A partial email address? A street name with no number? Those had to be written down as rules and applied consistently across a manually reviewed set before any score meant anything, because a number measured against inconsistent labels is worse than no number.

The one genuinely interesting choice was the metric. The standard measure treats a miss and a false alarm as equally bad. Here they obviously aren't. Redacting something harmless is an inconvenience. Missing an actual phone number is the thing you built the system to prevent. So I weighted the score toward catching everything, and said so explicitly rather than reaching for the default.

What I took from it

I went in thinking the job was getting a model to behave. I came out thinking the job is making its behaviour provable, to someone who wasn't there and doesn't take your word for it.

Everything I've built since starts with how I'm going to test it.

© 2026 Arshin Sikka. All rights reserved.

Ask my AI anything!