CG Common Ground | Common Ground Systems
What we did and what it producedActive

The work, decision by decision

The pipeline, in the order it runs

  1. Texts. Read a person's own sent messages directly from their device's message store and decode the modern message format, which most extraction scripts get wrong and silently return empty.
  2. Email, fast pass. Pull from the mail client's own index, which holds a usable preview of most messages even before the full body has been downloaded.
  3. Email, deep pass. A slow background pass that requests full message bodies for a sample weighted toward older years, because recent email is often already AI-drafted and teaches the system the wrong voice.
  4. Dictation. Read a dictation tool's own local history, including the raw spoken words with fillers intact, the tool's own cleaned-up version, and, most valuably, every place the person corrected the cleaned-up version by hand afterward. Those corrections are the richest signal in the whole pipeline, because they show exactly what a machine gets wrong about this specific person.
  5. Recordings. Where meeting or call transcripts exist, work out which speaker is the person being modeled and flag any recording where that call is uncertain for a human check.
  6. Count, then check the counting. Build the frequency and stylistic-distance statistics, then check the held-out accuracy before trusting the scorer in any one register. If the scorer cannot tell a person's own writing from AI-generated text in a register it was supposedly trained on, the scorer is broken there, not the writing.
  7. Read, stratified. Have separate reading passes work through stratified samples (one per register, one per era) against a fixed brief, then synthesize a voice guide and a mind profile from what they find, and run the strongest available model as an adversary whose only job is to try to break every claim in the result.
  8. Test. Build a blind test of matched pairs, the person's own real writing against a generated draft on the same task, generated without the generator ever seeing the real original. The person takes the test cold. Their hit rate is the score.

What it produced, and what it caught

The first full run read one partner's personal messages, email, dictation history and available call transcripts. It produced a short voice guide (lexicon, sentence rhythm, punctuation habits, what the person never says), an exemplar bank of tagged real passages, a mind profile with every claim tied to evidence, a scorer, and a ten-pair blind test.

It also caught its own early mistakes, which is the point of building it as a measured system rather than a described one. The first version of the mind profile carried more than a dozen wrong claims and several leaks of content that should have stayed private, all caught by the adversarial pass built specifically to find them. The first version of the scorer called the large majority of the person's own real emails AI-written, which is a scorer bug, not a fact about the person's email, and it had to be proven against real text before anyone trusted a single number it produced. And a specific word the team assumed was an AI tell turned out to be a word this particular person genuinely uses; the rule became count before you ban.

What we kept, replaced and installed

We kept the instinct to use a person's real writing as the source of truth. We replaced the adjective-list style guide and the paste-a-few-examples-into-the-prompt approach, because both produce writing that reads as generically competent and is easy to catch as AI, especially once the model has to write something the examples did not cover. We installed the four-layer system (fingerprint, retrieved exemplars, contrastive rules, adversarial mind profile) plus the scorer and the blind test as the standing method, with three output tiers: everything stays local to the person's own machine as full canon; a style-and-work layer with no personal content can go to internal collaborators; and a public-safe tier, work-only and free of anything that cannot be checked, can power an outward-facing voice.

What it costs, and what we would watch

The extraction itself runs in minutes except for the slow email deep pass, which runs for hours in the background. Reading and synthesis take roughly an hour of coordinated agent time per run. The real cost is discipline: every claim in the mind profile has to trace to a count or a quote, and every scorer result has to be proven on real text before it is trusted, because a broken checking system that looks fine is worse than no checking system at all. What we would watch on any future run: whether an assumed AI tell is actually this person's own word before banning it, and whether a register's held-out accuracy is real before anyone drafts against it.

The result, in short

A four-layer voice system (a counted fingerprint, a retrieved exemplar bank, a contrastive rule layer, and an adversarially checked mind profile), a scorer, and a blind test now sit behind every drafting surface that carries a partner's name. The first run caught its own early errors, a broken scorer and a wrongly assumed AI tell, before either shipped.

A slice of the project list

A few related projects.