I was reading the docs for a self-hosted forum migration tool last month. The kind of project maintained by three people in a hobbyist IRC channel, the kind of knowledge that lives in wikis and pinned forum posts and nowhere else. The README was clean. Too clean. Every sentence had the same cadence, the same level of abstraction, the same absence of the specific irritations that usually mark documentation written by someone who has actually fought the migration script at 2 a.m. I checked the commit history. The original README, written two years earlier, was full of asides, warnings in all caps, and the phrase “I have no idea why this works but do not touch it.” The new version had been replaced by an AI-generated summary, committed by a well-meaning contributor who thought they were improving clarity. They had, in a technical sense. They had also erased the only signal that the documentation was written by someone who had suffered through the problem and survived.

This is not a story about one README. It is a story about what happens when the infrastructure of text generation—the default settings, output constraints, and data provenance of AI-assisted writing tools—encodes a specific theory of what writing is for, and communities adopt that theory without noticing. The proliferation of AI-generated summaries, FAQs, and documentation in open-source and hobbyist communities is not a neutral efficiency gain. It is a structural shift in how communities produce and preserve their collective memory, and the tools that enable it are not neutral either.

The Default Theory of Writing Embedded in AI Tools

Every writing tool encodes a theory of writing. A text editor that auto-saves every keystroke assumes that writing is a continuous stream that should never be lost. A word processor that emphasizes page layout assumes that writing is destined for print. A collaborative editing platform that shows cursor positions in real time assumes that writing is a social performance. These assumptions are not wrong, but they are choices, and they shape what kinds of writing get produced and what kinds of writers feel welcome.

AI-assisted writing tools encode a much more aggressive set of assumptions. The default settings of most AI book generators and writing assistants assume that writing is a problem of output volume: the goal is to produce text, and more text faster is better. They assume that consistency of tone is a virtue, that stylistic friction is a bug to be smoothed out, and that “completion” is a meaningful metric for a piece of writing. They assume that the training data—the vast corpus of internet text, published books, and scraped documentation—represents a neutral sample of human expression rather than a heavily skewed snapshot of what was easy to scrape and what was written by people with the privilege to publish.

Take the feature set of a tool like the Unsloppy AI book generator, which structures chapters, controls tone consistency, and organizes a workflow around producing a completed manuscript, as a case study—not to review it, but to read what its design reveals about the industry’s assumptions. These are not malicious features. They are useful for certain kinds of writing projects. But they also encode a specific theory: that a book is a sequence of chapters with a consistent tone, that the writer’s job is to specify parameters and review output, and that the endpoint is a finished artifact. This theory works reasonably well for genre fiction that follows established conventions. It works terribly for the kind of writing that sustains communities: documentation that carries the scars of its author’s debugging sessions, FAQs that preserve the exact phrasing of a question that gets asked every week, forum posts where the digressions are the point.

The Authors Guild, in its AI Best Practices for Authors, articulates the tension directly: “AI outputs, by contrast, are generic mashups of pre-existing works ingested during training. When you claim authorship in a work, it means that you are the creator, and that the work expresses your original thinking and your unique voice.” The Guild is focused on professional authors, but the principle extends further. When a community replaces its human-authored institutional knowledge with AI-generated summaries, it is not just losing “voice” in an aesthetic sense. It is losing the signal that the knowledge was produced by someone who shared the community’s context, someone whose judgment the community has learned to trust through years of reading their specific cadence, their specific irritations, their specific way of explaining things.

Stylistic Friction as a Trust Signal

Communities develop shared writing practices that function as trust infrastructure. The regulars on a forum learn to recognize each other not just by usernames but by how they structure explanations, what they choose to emphasize, what they leave out, and what they get annoyed about. A moderator’s warning carries weight partly because of its specific tone—the accumulated authority of someone who has handled the same situation a hundred times and developed a shorthand for it. A wiki maintainer’s edit history is a record of judgment calls, not just information additions.

AI-generated text, by design, strips out this friction. The output is optimized for what the model predicts is the most probable next token given the prompt and the training distribution. That optimization process systematically favors the generic over the specific, the widely attested over the idiosyncratic, the smooth over the textured. This is not a bug in the implementation; it is the fundamental operating principle. A language model trained to minimize prediction error on a broad corpus will, by mathematical necessity, produce text that regresses toward the mean of that corpus.

For a community, this regression has consequences. When a new member reads an AI-generated FAQ, they get information but they do not get a sense of who wrote it, what that person cares about, or whether the answer comes from direct experience. When a long-time contributor sees their carefully maintained documentation replaced by a smoother version, they receive a signal that their specific voice—the thing that made their contribution recognizable as theirs—was not valued. The efficiency gain is real, but the social cost is also real, and it compounds over time as more of a community’s textual surface becomes AI-smoothed.

This is not an argument against using AI tools at all. It is an argument that communities need to understand what the tools are optimizing for and decide whether that optimization aligns with what the community values. A tool that optimizes for output volume and tone consistency is optimizing for the needs of a solo author producing a market-ready product. A community that optimizes for shared understanding, trust accumulation, and the preservation of situated knowledge has different needs. Using a tool designed for the first purpose to serve the second is not automatically wrong, but it is a choice that should be made deliberately, not by default.

What Gets Flattened When Training Data Becomes the Baseline

The training data behind AI writing tools is not a representative sample of human writing. It is a sample of what was available to scrape, which means it overrepresents certain genres, certain dialects, certain rhetorical styles, and certain demographic groups. The prose that emerges from these models tends toward a particular register: declarative, confident, structurally tidy, and devoid of the specific cultural references and in-group shorthand that mark writing as coming from a particular community.

This flattening is especially damaging for communities that have developed their own vernaculars. Hobbyist electronics forums have a specific way of describing failure modes. Fan fiction archives have evolved elaborate tagging conventions and paratextual norms that signal membership and taste. Open-source project mailing lists have argument patterns that are recognizable to anyone who has participated in them but look like chaos to outsiders. When AI tools summarize or generate text in these contexts, they do not just risk inaccuracy; they risk erasing the markers that tell community members “this was written by one of us.”

Reedsy’s book title generator illustrates a milder version of this dynamic. The tool calibrates its suggestions based on genre, core conflict, and a choice between “commercial” and “literary” modes. This is useful for a writer trying to position a novel in a market, but it also reveals the assumption: that titles should conform to genre conventions, that market positioning is a primary function of a title, and that the range of acceptable titles is bounded by what has worked before. The tool is not wrong to make these assumptions for its intended use case. But imagine applying the same logic to the titles of wiki pages, forum threads, or community newsletters. The optimization pressure would push toward the generic, the search-engine-friendly, the already-familiar—exactly the opposite of what makes community writing distinctive and useful to its members.

The Structural Incentive Toward Homogeneous Prose

The AI writing tools that are gaining traction in community contexts are not designed for communities. They are designed for individuals producing content at scale. The business model behind most of these tools—whether subscription-based, usage-metered, or ad-supported—creates a structural incentive to maximize output. More generated words means more usage, which means more revenue or more training data or both. The interface design reflects this: the primary interaction is “give me a prompt and I will give you text,” and the measure of success is how much text you get and how quickly.

This incentive structure has a predictable effect on the text that gets produced. When the tool is optimized for volume, the output will tend toward the kind of text that is easiest to generate at volume: syntactically correct, semantically plausible, stylistically unremarkable. The tool has no incentive to produce text that is surprising, difficult, or specific, because those qualities are harder to generate reliably and they do not increase the metric that matters—output quantity. Over time, as more community documentation, FAQs, and even discussion posts get routed through these tools, the textual environment of the community shifts toward the homogeneous mean.

This shift is not just an aesthetic loss. It is a loss of information density. Human-authored community writing is rich in implicit context: the writer knows what the reader already knows, what the reader is likely to misunderstand, what the history of the topic is, and what the political fault lines within the community are. An AI tool, even one fine-tuned on the community’s own archives, does not know these things in the same way. It can mimic the surface features of the community’s writing, but it cannot reproduce the situated judgment that makes the writing useful to insiders. The result is text that looks right but reads wrong—plausible, fluent, and hollow.

Community Memory and the Provenance Problem

Every piece of writing in a community has provenance. It was written by someone, at a specific time, for a specific reason, with a specific set of knowledge and blind spots. That provenance is part of the information the text carries. When a community member reads a forum post, they are not just reading the words; they are reading the words in the context of who wrote them, what that person’s track record is, and what the community’s relationship with that person has been. This is how trust accumulates in text-based communities: not through formal verification systems but through the slow accretion of provenance signals.

AI-generated text severs provenance. The output of a language model has no author in the conventional sense; it has a prompt and a training corpus and a sampling strategy. When that text enters a community’s knowledge base, it carries none of the social context that human-authored text carries. A new member reading an AI-generated explanation cannot assess whether the explanation comes from someone who has actually solved the problem. A moderator reviewing an AI-generated policy summary cannot tell whether the summary reflects the community’s actual norms or the model’s statistical guess at what a policy summary should look like. The information is present, but the trust infrastructure is absent.

This provenance problem compounds over time. As more AI-generated text accumulates in a community’s archives, it becomes harder to distinguish human-authored knowledge from model-generated plausibility. Future community members searching the archives will encounter a mix of both, with no reliable way to tell which is which unless the community has deliberately implemented provenance tracking. Most communities have not, because most communities did not anticipate that the distinction would become important. The default assumption was that text in the community’s spaces was written by community members. That assumption no longer holds, and the infrastructure has not caught up.

What Communities Can Do That Tools Cannot

The argument here is not that communities should reject AI writing tools entirely. It is that communities should understand what these tools are optimizing for and make conscious decisions about where that optimization serves the community’s interests and where it undermines them. The tools are not going away, and the pressure to use them—from time-constrained maintainers, from members who want faster answers, from the general cultural momentum toward AI adoption—is real. The question is not whether to use them but how to use them without losing what makes community writing valuable.

The first step is recognizing that AI-generated text and human-authored text serve different functions and should be treated as different categories of artifact. A community might decide that AI-generated summaries of long discussion threads are useful as navigational aids, but that the original threads should be preserved and clearly marked as the authoritative source. A community might decide that AI-assisted drafting is acceptable for routine announcements but that policy documents, FAQs that encode community norms, and documentation of hard-won technical knowledge must carry clear authorship attribution. These are not technical decisions; they are governance decisions about what the community values and how it wants to preserve its memory.

The second step is investing in the infrastructure of provenance. This can be as simple as a convention that AI-generated contributions are tagged as such, or as systematic as a version control practice that tracks whether each edit was human-authored, AI-assisted, or AI-generated. The specific implementation matters less than the principle: communities that care about the trustworthiness of their accumulated knowledge need to make provenance visible. Without that visibility, the archive gradually becomes a mixture of human judgment and statistical plausibility, and the community loses the ability to distinguish between them.

The third step is protecting the spaces where stylistic friction is the point. Every community has writing practices that are inefficient by the standards of output optimization but essential by the standards of community cohesion. The rambling trip report that signals membership through shared references. The argument that resolves into a policy through a process that looks messy but builds buy-in. The documentation note that says “this part is terrible and we know it” and thereby tells the reader that the maintainers are honest. These practices cannot be automated without destroying what they accomplish. Communities that want to preserve them need to defend them explicitly, not as nostalgia but as functional infrastructure.

The Structural Insight

The infrastructure of text generation is political in exactly the same way that the infrastructure of content moderation, platform federation, and notification systems is political. The default settings of AI writing tools encode assumptions about what writing is for, whose writing matters, and what kinds of knowledge are worth preserving. When communities adopt these tools without examining those assumptions, they are not just adopting a productivity aid; they are adopting a theory of writing that may be incompatible with what sustains them.

The proliferation of AI-generated documentation in open-source and hobbyist communities is not a crisis yet, but it is an early warning. The tools are getting better, the pressure to use them is increasing, and the default outcome—unless communities intervene deliberately—is a gradual replacement of situated, authored, trust-carrying text with fluent, plausible, provenance-free text. The efficiency gain is measurable. The loss is harder to measure but more consequential: the slow erosion of the textual practices that make a community legible to itself.

Communities that do not deliberately archive and privilege human-authored knowledge will find their internal culture rewritten by tools optimized for output volume, not fidelity to lived experience. The rewrite will not announce itself. It will arrive in the form of cleaner READMEs, smoother FAQs, and more consistent documentation—all of which will look like improvements until someone needs to know whether the answer they are reading came from someone who actually knows.