Skip to main contentSkip to navigationSkip to search

We taught an AI our brand voice. Here's what we learned

By Magnus Lindgren
Magnus LindgrenDeveloper

Most brand guidelines get written, get approved, and then sit on SharePoint while everyone tries to remember if it's "website" or "web site". Not because we don't take them seriously – but because applying them consistently across every LinkedIn post, every web page, every proposal is grinding work, and grinding work always falls off the bottom of the list. 

That gap between what the guideline says and what gets published was the problem I set out to close. Over the past few weeks I've built a brand compliance agent for Comprend in Optimizely Opal. Four specialised agents, one Workflow. 

I came at this as a developer, not a content strategist. That matters. To a developer, an agent looks a lot like a function: takes an input, does one job, returns an output. The same things that separate good code from bad – narrow scope, predictable behaviour, something you can test – separate reliable agents from unreliable ones. That instinct shaped every design decision. 

Real LinkedIn copy that scored 2.8 out of 5 on first review came back at 4.0 – ready to publish – after one automated pass. About 90 seconds, end to end. But honestly? The score lift wasn't the most useful thing this pilot produced. 

Why brand review is the right place to start 

I didn't start with the technology. A colleague handed me the brief – build something that reviews copy against our brand guidelines — and the first question was where the real bottleneck actually sat. 

Every piece of copy gets checked against the same kind of rules. Different reviewers focus on different things – one cares about tone, another about terminology, a third about whether the em-dashes are correct – so the same text comes back with three different sets of notes. And the senior people doing the reviewing are the people we'd rather have on strategic work. 

Classic bottleneck. Also the kind of work an agent should be good at: defined inputs, defined outputs, rules already written down. 

What we built – four agents, one Workflow 

The pilot uses four specialised agents in Optimizely Opal, each with one job. 

The terminology checker flags forbidden phrases, missing abbreviation expansions, and outdated product names. Rule-based and boring on purpose. 

The brand reviewer scores copy on a scale of 1–5 across Comprend's four tone-of-voice pillars: Confident, Clear, Approachable, Playful. It returns structured feedback with direct quotes – not "this opening is generic" but "the sentence 'in today's rapidly evolving digital landscape' is a hollow opening." That specificity is what makes it useful. 

The copy rewriter rewrites in Comprend's voice, calibrated to channel. The copy cleaner does surgical cleanup – em-dashes, abbreviations, sentence case, weasel words. 

All four are wrapped in one Opal Workflow. An editor types @brand-compliance and pastes their copy. The brand reviewer runs first. If the text scores 4.0 or higher, the Workflow stops. If it scores below, the rewriter and cleaner run automatically, then the brand reviewer runs again on the improved version. 

That's the full loop – from raw draft to brand-compliant copy. Without a human reviewer in the middle. 

Treat agents like functions, not employees 

This is where being a developer shaped the whole thing. 

My first rewriter tried to do everything. Rewrite for tone. Fix abbreviations. Remove weasel words. Restructure sentences. It followed about half its own instructions on a good day, and the failure mode was the worst kind: output that looked plausible but was unreliable in ways I couldn't predict. Sometimes it fixed the em-dashes. Sometimes it didn't. Sometimes it kept the credential-leading paragraph I'd told it three different ways to remove. 

Any developer recognises that pattern. It's what happens when a function tries to do too much. You can't reason about the output. You can't write a meaningful test. When something breaks, you can't isolate which step broke it. 

Splitting the work was the fix. One agent for voice. One agent for rules. Each became a unit I could actually debug. 

The mental shift that made it click: stop thinking about agents as colleagues you're delegating to. Start thinking about them as functions in a pipeline. Functions have one job. They have predictable inputs and outputs. They compose into bigger things. That's the model. 

The instinct to give an agent more context, more capability, more responsibility is the wrong instinct. It's the same instinct that makes junior developers write 400-line "god objects" – bloated single units of code that try to handle everything at once – instead of clean, composable functions. Feels like more capability. Produces worse software. 

What broke, and how we fixed it 

The first version of the brand reviewer was way too lenient. It gave a known-good text 2.8 and a weaker text 4.0. Worse than no agent at all. 

The fix was three changes. First, define what every score on the 1–5 scale actually means – without that, the model averages out to three on everything. Second, make the agent count violations and tie the count to the score: six or more concrete, quotable violations means the score can't go above 3.7. Third, set a clear publication threshold. The brand owner asked for it strict – 4.0 out of 5 is the bar. 

Even with all that, calibration kept slipping. The agent was consistent. It just wasn't always consistent with us – and that turned out to be the most interesting problem. 

The lesson I didn't expect 

The agent was scoring strictly against our written brand guideline. The brand team were scoring against something else – a feel for the voice that lives in their heads, calibrated from years of publishing. Those two scoring systems didn't always agree. 

So I asked the brand team for three short paragraphs: examples of Comprend at its best. I added them to the agent's instructions as gold-standard references, so it had something concrete to compare against rather than just rules. 

The difference showed up immediately in the output. Instead of flagging "this opening is generic," the agent started citing our own reference paragraphs directly: "the opening should deliver a sharp observation, not a hollow superlative." 

It still doesn't land on every text the same way the brand team would. Fair enough – the brand team don't always agree with each other either. But the disagreement is now legible. 

What changes now 

The agent doesn't replace the editor. It catches the stuff an editor misses when they're tired. The em-dash that should be an en-dash. The abbreviation never spelled out. The credential-leading sentence everyone reads past because they've read it a hundred times. It takes the repetitive layer off the senior reviewer's desk so they can spend time on judgement instead. 

There's a lot still to build – and we're looking forward to it. The agent needs more reference texts to handle the full range of Comprend formats. The next step is moving the trigger from a manual chat command into our content marketing platform, so a review runs automatically when content is marked ready. After that, the model is repeatable: brief validators, headline generators, terminology checkers for other brands across the Aura Group (our parent company). 

The real finding wasn't the score improvement. It was this: a brand guideline tells you what to avoid. It doesn't tell you what good actually sounds like. Once we added reference examples to the agent's instructions, the feedback got sharper and the results got more consistent. If you're building something similar, start there – get your gold standards written down before you write your rules. 

 

If you're building AI-driven content workflows and want to compare notes, get in touch – we're always up for a chat. 

Contact us

Do you wish to exchange more thoughts with us on how to thrive and grow from within? Join us at our next Comprend day or get it touch now.

James HandslipManaging director, UK
Gabriella BjörnbergManaging director, Stockholm
Kimmo KanervaExecutive director