brand identity

Voice Rubric

Voice rubrics are the structured spec that makes brand voice survive AI scale. They replace vague descriptions with lists of measurable behaviors complete with targets and scoring criteria. A typical rubric might specify average sentence length of 15 words maximum 25. It might require leading with the answer in the first eight words of every piece. It bans filler openings like in a world where or it is important to note that. It sets minimums for active voice percentage and concrete noun usage. These rules get packaged as voice tokens in your brand system. They drop directly into prompt packs for different use cases. Marketing copy gets one version. Product UI strings get another. Support documentation gets its own. Automated evals scan generated text for compliance on each rule and return a composite score. The editor reviews weekly dashboards of those scores and adjusts the rubric when scores trend down. This turns voice from an opinion into part of the token graph. Linear built their entire writing AI around such a rubric. Anthropic runs voice evals against theirs on every public facing string. The approach scales the human director role from approving individual pieces to governing the rules that generate thousands.

A voice rubric is not the decorative voice page that still appears in most brand books. Those pages list adjective pairs and stop. They read like creative writing assignments written by people for people who already share context. They fail the moment you hand them to an AI because models need numbers not moods. A rubric is not a static artifact. It must be versioned tested and tuned against real output data. It is not a replacement for judgment but a scaffold that makes judgment scalable. Brands that treat voice as prose instead of spec are the ones whose AI campaigns look off in ways they cannot fix quickly. The 2024 Klarna AI ads showed inconsistent tone across markets because their guidance lacked measurable targets. Coca Cola saw similar issues in Create Real Magic where consumer prompts produced text that ranged from on brand to completely alien. The classic voice document created that vulnerability. The rubric removes it by making every rule something an eval can check in milliseconds.

The best concrete example comes from Linear. Their voice rubric contains twelve specific rules with associated scores. Rule one: Sentences average fewer than 18 words. Score drops one point for every two words over. Rule two: No sentence begins with and or but. Rule three: Every paragraph must contain at least one concrete number or example. Rule four: Ban list of corporate speak terms including delve journey leverage nexus and ecosystem. The prompt pack starts with the full rubric text then adds six few shot examples of copy that scored nine or higher and three that scored below five with explanations of the failures. Their eval system combines regex checks for sentence length and banned words with an LLM judge that scores the more qualitative rules. Outputs below 8.0 get flagged and sent back through the model with a correction prompt that references the specific failed rules. This loop has kept Linear voice consistent across product updates that number in the hundreds per year. Anthropic uses a similar but more sophisticated setup. Their rubric focuses on three pillars: truthfulness clarity and directness. Specific rules include never qualifying a factual statement with seems like or appears to and always providing concrete citations or examples for technical claims. They trained a smaller evaluator model on thousands of human scored examples from their own content team. The system runs on every blog post help article and interface string. Stripe documentation follows a rubric that demands developer empathy metrics such as explaining every technical term on first use and including runnable code examples in 70 percent of explanatory sections. Before and after reveals the difference. Bad copy reads In todays fast paced digital landscape we are delighted to delve into our latest innovation that we believe will revolutionize the way you interact with our platform. Rubric fixed copy reads Our new dashboard loads 40 percent faster. Click here to try it. These examples prove the pattern. The rubric integrates with the token system the prompt packs and the eval loop to create the closed system that stops drift cold. Without it even strong visual identities like Heinz in 2022 would fail at scale in 2026 because training data luck runs out.

Implement a voice rubric the day your AI writing volume exceeds manual review capacity. That moment arrived for most teams years ago. Use it for social copy product messaging email sequences blog drafts and ad variations. Embed the full rubric in every relevant prompt pack. Run it through your brand evals before any asset ships. Let the editor focus on tuning rules and curating edge cases instead of line editing. The leverage multiplies. Skip the full rubric only if you write everything by hand or if voice is not central to your brand differentiation. In the latter case a short list of principles might work. But few brands in the current environment can claim voice is not central. The cautionary tales keep multiplying. Brands without rubrics watch their AI output slowly slide into generic patterns that erode trust. The ones with strong rubrics ship faster with tighter consistency. The editor role becomes governance not production exactly as described in the four part system of tokens prompts evals and editor.

Voice rubrics turn vague brand personality into scorable rules that keep AI outputs on voice at ten thousand a day scale.

Related terms

Keep exploring