Teaching an agent to stop sounding like an agent
A case study in using Caliper to evaluate blader/humanizer, tighten voice calibration, and turn the improvement into an upstream contribution with regression coverage. With a postscript on what the maintainer did instead.
Most writing evals start in the wrong place.
They ask whether the text “sounds human.” That seems reasonable until you try to grade it. Human compared to what? A blog post? A support email? A design doc? A founder note written at midnight?
The useful question is narrower. When the user supplies a writing sample, did the skill preserve that voice while removing the AI tells? Now there is something to grade against.
I used blader/humanizer as the test case. It is a Claude Code skill that rewrites AI-sounding text using the patterns catalogued in Wikipedia’s “Signs of AI writing” guide, the advice page WikiProject AI Cleanup maintains for spotting model-written articles. [5] The skill is already strong on the obvious tells: significance inflation, “rapidly evolving landscape,” signposting, generic conclusions, bold inline-list headers, and the rest of the sludge that makes model text feel assembled.
The gap I wanted to evaluate was smaller. Version 2.5.1 already had a Voice Calibration section, but the requirement was soft:
## Voice Calibration (Optional)
If the user provides a writing sample (their own previous writing), analyze it before rewriting:
1. **Read the sample first.** Note:
- Sentence length patterns (short and punchy? Long and flowing? Mixed?)
- Word choice level (casual? academic? somewhere between?)
- How they start paragraphs (jump right in? Set context first?)
- Punctuation habits (lots of dashes? Parenthetical asides? Semicolons?)
- Any recurring phrases or verbal tics
- How they handle transitions (explicit connectors? Just start the next point?)
2. **Match their voice in the rewrite.** Don't just remove AI patterns - replace them with patterns from the sample.
3. **When no sample is provided,** fall back to the default behavior.
That is directionally right. But “analyze it” is easy for the model to do silently, or to skip. The eval target became:
If a writing sample is supplied, the assistant must state a brief voice target before rewriting, then produce a final rewrite that follows that target.
That is much easier to grade than “be human.”
The failure case
The skill’s default style advice is deliberately lively:
**Have opinions.** Don't just report facts - react to them.
**Use "I" when it fits.** First person isn't unprofessional - it's honest.
**Let some mess in.** Perfect structure feels algorithmic.
That advice is good when there is no sample. It is risky when the sample is dry technical prose.
Here is one eval item:
- name: dry_technical_voice
prompt: |
Humanize the draft below. Use the writing sample as the voice target.
Writing sample:
The cache stores normalized keys and their serialized values. A write replaces
the previous value for the same key. Expiration is evaluated on read, not by a
background job. This keeps the implementation small and makes eviction behavior
easy to test.
Draft to humanize:
The caching layer serves as a robust foundation for scalable application
performance, providing developers with a seamless and intuitive mechanism for
managing frequently accessed data. Not only does it reduce latency, but it also
empowers teams to unlock more efficient user experiences across the platform.
Let's dive into what makes this approach so powerful.
A good answer adds no jokes, no first person, no bloggy aside. It sounds like the person who wrote the cache note cleaned up the bad draft.
The expectation was explicit:
expect: |
Pass only if the response includes a brief voice card or equivalent style-target
note before the rewrite, and the final rewritten prose removes promotional
phrasing, signposting, and negative parallelism while preserving a dry technical
voice.
The rewrite should be plain, specific, and impersonal. It should not add "I",
jokes, emotional reactions, punchy editorial comments, or conversational warmth
unless present in the source sample.
Fail if the response does not identify the sample voice before rewriting, makes
the prose more personal or opinionated than the sample, leaves "serves as",
"seamless", "unlock", or "Let's dive in", or changes the technical meaning.
This follows the shape OpenAI recommends for skill evals: a prompt, a captured run, a small set of checks, and a score you can compare over time. [1] It also keeps the task plain. LangChain’s writeup on evaluating skills makes the point well: make the task too adversarial and you start measuring the agent’s general problem-solving instead of the skill. [2]
The Caliper spec
I kept the suite small: four tasks, one narrow behavior.
skill:
path: ../SKILL.md
backend: codex
judge:
backend: codex
tasks:
- name: edonadei_blog_voice
prompt: |
Humanize the draft below. Use the writing sample as the voice target.
...
expect: |
Pass only if the response includes a brief voice card or equivalent
style-target note before the rewrite, and the final rewritten prose removes
obvious AI-writing tells while preserving the sample's voice.
- name: dry_technical_voice
prompt: |
Humanize the draft below. Use the writing sample as the voice target.
...
expect: |
Pass only if the response includes a brief voice card or equivalent
style-target note before the rewrite, and the final rewritten prose removes
promotional phrasing, signposting, and negative parallelism while preserving
a dry technical voice.
- name: casual_email_voice
prompt: |
Humanize the draft below. Use the writing sample as the voice target.
...
expect: |
Pass only if the response includes a brief voice card or equivalent
style-target note before the rewrite, and the final rewritten prose removes
formal chatbot artifacts and matches the casual email sample.
- name: no_sample_default
prompt: |
Humanize this draft:
...
expect: |
Grade only the rewritten prose, especially the final rewrite if the response
includes draft/audit sections. A more natural, opinionated voice is acceptable
because no sample was supplied.
The fourth task matters. It checks that the change does not break the skill’s normal behavior when there is no sample.
I also added a cheap post-run assertion script. Caliper’s autorater is the main judge, but deterministic checks fail loudly and say exactly what regressed.
AI_TELLS = [
"rapidly evolving landscape",
"pivotal milestone",
"serves as",
"seamless",
"unlock",
"let's dive in",
"i hope this message finds you well",
"future looks bright",
"journey toward excellence",
]
def check_output(task: str, output: str) -> list[str]:
checked = final_rewrite(output)
text = checked.lower()
failures: list[str] = []
if task != "no_sample_default" and not has_voice_card(output):
failures.append(f"{task}: did not include a voice card before rewriting")
for tell in AI_TELLS:
if tell in text:
failures.append(f"{task}: left AI tell {tell!r}")
if re.search(r"\*\*[^*\n]+:\*\*", checked):
failures.append(f"{task}: used bold inline-header formatting")
if checked.count("—") > 1:
failures.append(f"{task}: used more than one em dash")
if task == "dry_technical_voice" and re.search(r"\bI\b|\bwe\b|\bmy\b|\bour\b", checked):
failures.append(f"{task}: added first-person language to dry technical voice")
return failures
The script extracts the final rewrite when the skill uses its normal output format:
def final_rewrite(output: str) -> str:
patterns = [
r"\*\*Final rewrite\*\*\s*(.*?)(?:\n\s*\*\*Changes made\*\*|\Z)",
r"Final rewrite\s*\n+(.+?)(?:\n\s*Changes made|\Z)",
]
for pattern in patterns:
match = re.search(pattern, output, flags=re.IGNORECASE | re.DOTALL)
if match:
return match.group(1).strip()
return output
That split matters. The skill deliberately returns draft, audit, and final sections. The eval should grade the rewrite, not penalize the wrapper.
Baseline result
I ran the upstream SKILL.md with:
caliper run humanizer-voice-calibration.eval.yaml \
--k 1 \
--workers 1 \
--timeout 180 \
--judge script \
--output baseline-old-k1.json
Result:
| Task | Current skill |
|---|---|
edonadei_blog_voice | 0/1 |
dry_technical_voice | 0/1 |
casual_email_voice | 0/1 |
no_sample_default | 1/1 |
| Average pass@1 | 25.0% |
The failures all had the same shape. The rewrites were often decent. The model just never identified the supplied voice before rewriting.
The judge’s reason for dry_technical_voice:
The transcript does not include a brief voice card or equivalent style-target note
identifying the sample voice before the rewrite. Although the final rewrite is dry
and removes the prohibited promotional phrasing, the missing voice identification
fails the expectation.
That is exactly the failure I want an eval to catch, and exactly the kind a human reviewer waves through. The output looked fine. The process was under-specified, and the good output was luck.
Anthropic’s eval guide separates the outcome from the transcript that produced it. [3] For a writing skill the transcript is mostly text, not tool calls. The voice card is the observable trace that the model read the sample as a constraint instead of as flavor.
The patch
The skill change is small.
## Voice Calibration (Optional)
-If the user provides a writing sample (their own previous writing), analyze it before rewriting:
+If the user provides a writing sample (their own previous writing), that sample becomes
+the style target. Analyze it before rewriting, and let it override the default
+PERSONALITY AND SOUL guidance.
-1. **Read the sample first.** Note:
+1. **Read the sample first and write a brief voice card.** Before the rewrite, note:
- Sentence length patterns (short and punchy? Long and flowing? Mixed?)
- Word choice level (casual? academic? somewhere between?)
- How they start paragraphs (jump right in? Set context first?)
- Punctuation habits (lots of dashes? Parenthetical asides? Semicolons?)
- Any recurring phrases or verbal tics
- How they handle transitions (explicit connectors? Just start the next point?)
+ - Point of view (first person? second person? detached?)
+ - Emotional temperature (dry, warm, blunt, funny, skeptical, enthusiastic?)
-2. **Match their voice in the rewrite.** Don't just remove AI patterns - replace them with patterns from the sample.
+2. **Use the voice card as a constraint, not decoration.** Don't just remove AI patterns -
+replace them with patterns from the sample. If they write dry technical prose, keep it dry.
-3. **When no sample is provided,** fall back to the default behavior.
+3. **Do not import personality from this guide unless the sample supports it.** Do not add
+first person, jokes, punchy asides, mixed feelings, edge, or extra opinion just because
+the default guidance says human writing can have those traits.
+
+4. **Preserve meaning and remove AI tells.** Voice matching does not permit changing
+claims, adding unsupported details, or leaving obvious AI artifacts in place.
And the output format:
Provide:
-1. Draft rewrite
-2. "What makes the below so obviously AI generated?" (brief bullets)
-3. Final rewrite
-4. A brief summary of changes made (optional, if helpful)
+1. Voice card (only when a writing sample is provided)
+2. Draft rewrite
+3. "What makes the below so obviously AI generated?" (brief bullets)
+4. Final rewrite
+5. A brief summary of changes made (optional, if helpful)
This is deliberately not a rewrite of the skill. It is one constraint: the sample’s voice overrides the default personality.
Improved result
Same command, patched skill:
caliper run evals/humanizer-voice-calibration.eval.yaml \
--k 1 \
--workers 1 \
--timeout 180 \
--judge script \
--output improved-k1.json
python evals/assert_voice.py improved-k1.json
Result:
| Task | Current skill | Patched skill |
|---|---|---|
edonadei_blog_voice | 0/1 | 1/1 |
dry_technical_voice | 0/1 | 1/1 |
casual_email_voice | 0/1 | 1/1 |
no_sample_default | 1/1 | 1/1 |
| Average pass@1 | 25.0% | 100.0% |
The post-run assertions passed:
Voice assertions passed.
Small evals are not benchmarks. This is four tasks, k=1, one backend, one date. The point is not that humanizer is solved. The point is that a behavior that used to be implicit now has a regression test.
That is the loop I care about:
Observed failure
-> eval case
-> small instruction patch
-> rerun
-> regression coverage
It is also why I prefer evals that name the failure precisely. “Make the output more human” would have produced a mushy patch. “When a sample exists, state and follow the voice target” produced a small one.
CI for the skill
The tooling PR adds a GitHub Actions workflow.
name: Caliper
on:
pull_request:
workflow_dispatch:
jobs:
humanizer-evals:
runs-on: ubuntu-latest
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.11"
- name: Install Caliper
run: python -m pip install "caliper-eval[openai]"
- name: Validate eval spec
run: caliper validate evals/humanizer-voice-calibration.eval.yaml
- name: Run Caliper eval
if: ${{ env.OPENAI_API_KEY != '' || env.ANTHROPIC_API_KEY != '' }}
run: |
python - <<'PY'
import os
from pathlib import Path
import yaml
spec_path = Path("evals/humanizer-voice-calibration.eval.yaml")
spec = yaml.safe_load(spec_path.read_text())
backend = "openai-api" if os.environ.get("OPENAI_API_KEY") else "claude-api"
spec["skill"]["backend"] = backend
spec["judge"]["backend"] = backend
Path("evals/humanizer-voice-calibration.ci.eval.yaml").write_text(
yaml.safe_dump(spec, sort_keys=False)
)
PY
caliper run evals/humanizer-voice-calibration.ci.eval.yaml \
--k 1 \
--baseline \
--judge script \
--output results.json
python evals/assert_voice.py results.json
The workflow always validates the spec. It only runs the full eval when a model API key is configured. That keeps the PR useful for maintainers who do not want model calls on every contribution, while keeping the full regression check available.
Contribution
I packaged the work as two upstream pull requests rather than leaving it as a blog-only experiment.
The first is deliberately small. It changes only SKILL.md. [6]
codex/voice-calibration
-> Clarify that a supplied writing sample overrides the default personality guidance.
-> Require a brief voice card before rewriting.
The second is stacked on top and adds only the evaluation machinery. [7]
codex/humanizer-caliper-evals
-> Add evals/humanizer-voice-calibration.eval.yaml.
-> Add evals/assert_voice.py.
-> Add .github/workflows/caliper.yml.
The first PR body leads with the eval result:
| Version | pass@1 |
|---|---:|
| Current upstream `SKILL.md` | 25.0% |
| Voice Calibration PR | 100.0% |
The three failures in the current skill all had the same shape: the final
rewrite could be good, but the response did not identify the supplied sample's
voice before rewriting. This PR makes that step explicit.
The second PR hands the maintainer the harness separately. That split matters. The behavior change can be reviewed on its own, and the eval tooling can be discussed without blocking the smaller improvement.
That is the contribution shape I want more skills to have. A prompt change on its own asks the maintainer to take your taste on faith. An eval suite on its own is homework for someone else. The useful unit is the pair: a behavior change plus the harness that explains why it should stay changed. Splitting them across two PRs keeps the review path clean.
The earlier post made the broad argument: skills without evals are just vibes. [4] This is the smaller version in practice. One failure, one eval, one patch, one regression check.
Postscript, September 2026
Neither PR was merged, and the reasons are more instructive than a merge would have been.
Two days after I opened them, the maintainer shipped v2.6.0 and closed both. The first PR was closed as “largely covered”: the release gated the skill’s own personality section so it no longer injects its voice onto text where it does not belong, which was the core concern. The second was closed on principle. A Python and GitHub Actions harness with API key dependencies was, in the maintainer’s words, “a big surface to add to a deliberately single-Markdown skill.” [6] [7]
Both calls were fair. The eval found a real failure. The fix landed in a different shape than mine, and the evidence stayed outside the repo. By v3.0.0 the voice section is a few lines: read the sample first, match its rhythm and punctuation, and let it override every pattern in the guide. [8] That is the constraint I was asking for, without the voice card ceremony.
The lesson I took: the eval is the durable part of a contribution, not the patch. The patch got rewritten by someone who knew the skill better. The four tasks still tell you whether any version of the skill respects a supplied sample, and they will keep doing that for versions nobody has written yet.
References
[1] OpenAI Developers, "Testing Agent Skills Systematically with Evals," January 2026. developers.openai.com
[2] Robert Xu, LangChain, "Evaluating Skills," March 5, 2026. langchain.com
[3] Anthropic Engineering, "Demystifying evals for AI agents," January 9, 2026. anthropic.com
[4] E. Donadei, "Skills without evals are just vibes," May 2026. edonadei.com
[5] Wikipedia, "Wikipedia:Signs of AI writing," advice page maintained by WikiProject AI Cleanup. wikipedia.org
[6] blader/humanizer, pull request #123, "Clarify voice calibration behavior," opened May 25, 2026, closed May 27, 2026. github.com
[7] blader/humanizer, pull request #124, "Add Caliper voice calibration evals," opened May 25, 2026, closed May 27, 2026. github.com
[8] blader, "humanizer," GitHub repository, SKILL.md at v3.0.0. github.com