How I use stylometry and a corpus to make AI writing sound like me
Most AI writing does not sound generic because the model is incapable of good writing. It sounds generic because the model starts from nowhere.
That is the part founders usually skip. We open an agent, paste a rough idea, and ask the model to “write it in my voice” or “make it sound less AI.” Sometimes we add a few style words: direct, thoughtful, founder-led, practical, conversational. The output gets cleaner, but it still feels like it came from the same place as everyone else’s LinkedIn post.
The problem is not only the prompt. The problem is that the prompt is being asked to carry the whole system. It has to understand the idea, infer the audience, know the company, remember past claims, detect tone, avoid banned language, and somehow imitate a person it has barely seen.
That is too much work for one instruction.
The better way is to give the AI a corpus, keep that corpus available in memory, and let the system derive the right style brief for the specific writing run. Then the draft uses that brief together with durable writing guidelines. The prompt still matters, but it is no longer the only thing holding the piece together.
I do not think of stylometry as a static voice file that gets attached to every draft. I think of it as something the system derives from the relevant corpus each time it writes. The voice is not hardcoded. It is reconstructed from context.
The prompt is too small
A prompt can describe a voice, but it cannot replace one. Voice is not a list of adjectives. It is the pattern created by hundreds of small choices: how you open a point, how much context you give, what examples you reach for, what you cut, what you refuse to say, and how you end when you actually have something to say.
This is why “write in my voice” prompts usually collapse into average internet prose. The model has no evidence, so it fills the gap with the nearest public pattern. The result can be fluent and polished, but it does not feel owned by anyone.
A real voice has constraints. It has habits, edges, old arguments, and phrases it would never use. It knows that one sentence sounds like something you would say and another sounds like a sales deck. You cannot reliably prompt that into existence every time. The writing system has to retrieve the right context before it writes.
The four parts are different
The system works better when the pieces are separated. I find it useful to think in four layers: corpus, memory, guidelines, and the runtime style brief.
The corpus is the raw source material. It includes approved posts, rough notes, emails, essays, transcripts, product updates, sales notes, edits, and anti-examples. It is not the voice by itself. It is the evidence of how someone actually thinks and writes.
The memory layer keeps that material available. It is where the agent can retrieve relevant examples, product facts, previous decisions, edits, and writing constraints before a draft starts. Without that retrieval step, every writing session starts from a blank page again.
Guidelines are the durable rules. They say what should stay true across runs: who the reader is, what claims are allowed, what tone is too polished, what words to avoid, what product facts cannot drift, and what the company should never overclaim.
The stylometry brief is created for the run. The agent pulls the relevant corpus and memory, studies the examples for that specific topic and channel, then creates a temporary writing brief: how this piece should open, how technical it should get, what rhythm fits, what examples belong, and what kind of ending will feel natural.
That brief does not need to become a permanent asset every time. Some lessons are worth saving later as edits, facts, or guideline updates, but the brief itself is mostly working context. It exists to make this draft better.
What belongs in the corpus
Most people only want to feed AI their best writing. That is useful, but incomplete. Finished posts show the public voice, but rough notes show the thinking underneath it. Edits show taste. Anti-examples show boundaries.
The highest-signal material is often a before-and-after edit. If a draft says “unlock powerful workflows” and I change it to “run the weekly sales follow-up without opening five tabs,” that correction teaches the system something concrete. It shows the direction of taste: less abstract, more operational, closer to the actual work.
I like the corpus to contain a few different layers:
- approved writing that still feels right
- raw notes and transcripts that capture how I think before polish
- before-and-after edits that show preference
- anti-examples and banned phrases
- product and messaging facts that should not drift
The product facts are not a side detail. A model can sound convincing and still be wrong. If it does not know what is live, what is planned, who the product is for, and what claims are off-limits, it will eventually write something smooth and false. That is worse than a rough draft.
Why runtime stylometry matters
Static style guides are useful, but they are blunt. They usually say things like “be clear,” “be human,” “be practical,” and “avoid jargon.” Those are fine as guardrails, but they do not tell the agent how to write this specific piece.
A founder writing a reflective essay does not sound exactly like the same founder writing a launch post. A teardown does not sound like a cold email. A post about product pain does not sound like a hiring note. Same person, different job.
That is why I prefer deriving the style brief from the situation. For a given run, the agent should ask what topic this is, what channel it is for, which examples are closest, what the person has approved before, and what the current product truth allows us to say.
The answer becomes the stylometry for that draft. Not a permanent mask. Not a frozen persona. A context-specific reading of the corpus.
This makes the writing feel more natural because it respects the way people actually write. We do not use one exact voice everywhere. We stay recognizable while changing shape for the moment.
My actual workflow
When I want AI help with a piece, I try not to start with “write a post.” I start with retrieval. In the memory layer, the agent can pull the relevant writing samples, product facts, guidelines, and prior edits before it drafts anything.
Then the agent derives a short style brief from what it found. Not a generic one. A specific one for the piece in front of us: open with the concrete problem, stay founder-level instead of technical-doc-level, avoid launch-copy words, and use the product mechanism only after the reader understands the pain.
Only after that does the drafting happen. At that point, the prompt is not doing all the work by itself. The agent has the corpus for evidence, a memory layer for retrieval, guidelines for durable constraints, and a runtime style brief for the shape of this draft.
That narrower box is where most of the quality comes from.
A simple setup founders can copy
You do not need an elaborate system to start. A folder, database, or memory layer is enough if it is organized around how you actually write.
I would start with six sections:
- Approved writing: posts, essays, emails, and updates you would publish again.
- Raw voice: transcripts, notes, memos, and rough ideas before they become content.
- Durable guidelines: concrete rules about audience, tone, claims, structure, and language to avoid.
- Banned language: phrases that make you sound like everyone else.
- Edits and rewrites: before-and-after examples that show taste.
- Product facts: audience, positioning, live features, planned features, and claims to avoid.
The exact tool matters less than the behavior. The AI should be able to retrieve the right examples at writing time, derive the right brief for that run, and draft against the same durable constraints every time. You should not have to paste your whole company context into every prompt.
The human part is still the editor
A corpus does not remove the human from the process. It moves the human to the right place.
The human should not spend all their time fixing the same generic first draft. The human should define taste, make judgment calls, and turn good edits into reusable memory. If I remove a sentence because it sounds like SaaS copy, that is a signal. If I replace a vague claim with a concrete mechanism, that is a signal. If I change the ending because it summarizes instead of landing the point, that is a signal too.
Most people skip this loop. They edit the draft, publish the piece, and the system learns nothing. Then the next draft repeats the same mistakes. The better habit is to save meaningful corrections back into the corpus or guidelines so the next run starts from a slightly better place.
That is how the writing system compounds.
The AI should not preserve every habit
There is one trap in all of this: your corpus will contain things you should probably outgrow. Old positioning, repeated phrases, lazy openings, stale arguments, typos, and habits that were fine in short posts but do not work in long-form writing.
So the goal is not perfect mimicry. The goal is useful continuity. I want the agent to sound like me on a good day, not like an average of everything I have ever written.
That is why the guidelines matter. The corpus gives examples, but examples can conflict. Memory retrieves context, but not every retrieved thing deserves equal weight. The runtime style brief gives the draft shape, but the durable guidelines decide what should win when the corpus is messy.
Corpus gives the model evidence. The memory layer makes the evidence available. Runtime stylometry turns the relevant evidence into a writing brief. Guidelines give the draft judgment.
You need all four.
The test I use
After a draft is done, I ask three simple questions:
- Could anyone have written this?
- Does it contain a real mechanism?
- Would I say this to a founder in a normal conversation?
If anyone could have written it, it is too generic. If there is no real mechanism, it is probably just a polished opinion. If I would not say it in conversation, the prose is performing instead of explaining.
That is the problem with a lot of AI writing. It performs intelligence. It sounds like writing, but there is no lived pattern underneath it.
A corpus fixes that by giving the system actual examples. The memory layer fixes the blank-page problem by making those examples retrievable. Runtime stylometry turns the examples into a useful brief for the moment. Guidelines keep the draft honest.
The point is not to make AI “more human” by adding randomness or typos. The point is to stop asking the model to guess what your writing should sound like.
Founders should use AI for content. It is too useful not to. But if the workflow is only prompts, the output will keep sounding prompt-shaped.
Build the corpus. Keep it in memory. Let the system derive the style brief for the run. Draft with guidelines. Save the edits that teach the system something.
That is the shift: stop asking AI to sound like you from a blank page. Give it somewhere real to start.
