Milan Bogojevic Blog

How to Write Content That AI Models Actually Cite?

8 min read updated 16 August 2026

For years the goal was simple: rank first, get the click, turn the visitor into a customer.

Now there's another layer between you and that visitor. Someone asks an assistant a question, the assistant reads through a handful of sources and stitches together an answer. If you're in that answer, you exist. If you're not, you don't, no matter how well you'd have ranked in a classic search.

That doesn't change everything. But it changes one important thing: you're no longer writing just for a person who reads. You're writing for a system that extracts. And those two readers have different habits.

This is about those habits, and about how to write something that can actually be lifted out and cited.

Why a model decides to cite someone

There's no public rulebook, but a few patterns show up consistently in how these systems behave.

A model doesn't rank your page the way a search engine does. It looks for claims it can pull out and drop into an answer with as little risk of being wrong as possible. That means it isn't looking for the best text. It's looking for the most usable one.

Usable, in this context, breaks down into a few concrete things.

The claim is self-contained. You can lift it out of the paragraph and it still makes sense. If your key sentence reads "as we explained earlier, this depends on several factors," there's nothing to lift.

The claim is specific. A number, a deadline, a condition, a name. "The appeal deadline is fifteen days from receiving the decision" is citable. "The deadline is fairly short" is not.

The context sits next to the claim, not three paragraphs above it. Models work with chunks. If a sentence says "in that case, a fee applies" and "that case" was defined two sections earlier, the chunk is useless on its own, and it gets skipped.

The source reads as reliable. There's an author, a date, the claims are backed up, and what's written lines up with what's written elsewhere. A model won't cite an isolated claim that contradicts everything else it's seen, even if that claim happens to be true.

That's roughly it. Nothing mystical. You write so your sentence can be safely retold.

What actually changes in how you write

The answer goes at the top, not the bottom

Classic writing builds tension: intro, context, development, conclusion. That's good for reading and bad for extraction.

A better layout: state the question, answer it in two or three sentences, then everything else. A reader who knows what they want gets it immediately. A reader who wants depth keeps going. And a system pulling text finds a complete claim right at the top of the section, exactly where it's cheapest to grab.

Don't do this as a dry summary bolted onto the top of the article. Do it inside every section. The first sentence carries the point, the rest is development.

Headings should ask or state, not label

"Pricing" is a label. "How much does it cost to register a company in Kosovo" is a question a real person actually types.

The second one carries meaning on its own. When a model scans a document's structure, the heading is the cheapest way for it to figure out what a section covers. A label tells it nothing.

Same logic applies to claims as headings: "Registration takes three to seven days" beats "Processing time" as a heading, every time.

One idea, one paragraph

A paragraph carrying three different points can't be cited for any of them cleanly. Split it up.

This happens to help human readers too, which is a rare case where the two interests line up without a trade-off.

Numbers instead of adjectives

This is the single biggest win available.

Compare "the process is fairly quick and doesn't cost too much" with "the process takes five to ten business days and costs around a hundred euros." The first sentence has nothing that can be transferred anywhere. The second is a finished answer.

A rule worth keeping: every section should carry at least one checkable number, date, or name. If it doesn't, the section probably isn't saying anything concrete.

Define terms where you use them

If you're writing about something with a name, write the one sentence that says what it is, even if you assume everyone already knows.

That sentence often ends up being exactly what gets cited, because a model treats it as solid ground for building an answer. It costs you one line.

Say what doesn't work and where the limits are

This sounds like the opposite of good marketing, and it performs better than good marketing.

Text that says "this doesn't apply to companies registered before 2020" or "this approach won't help if the problem is your data" gives the system precisely what it needs for a precise answer, and it gives a human reader a reason to trust you.

Content that's entirely upbeat reads as promotional and gets treated more cautiously. That doesn't mean you can't take a position. It means the position needs an edge to it.

The technical side, briefly

A few things that still matter, even if they get talked about less now.

The page has to be readable without running scripts. If content only appears after everything loads in the browser, whatever's reading the page on the model's side never sees it. Old advice, new consequence.

Structured data still helps, mainly because it removes ambiguity. Who's the author, when it was published, when it was last updated, which organization is behind it. That's not a ranking trick. It's making something already on the page unambiguous.

Show the publish date and last-updated date. When an answer depends on how current something is, and a lot of answers do, undated content loses to dated content.

Give the author a real name and a bio. Not for the sake of form, but because reliability gets judged partly by who's standing behind the text.

Check what you're actually allowing. Different systems respect different access rules, and it's possible to be blocking exactly the ones you'd want reading you. Worth sitting down once and deciding deliberately who gets to read your content, instead of leaving it to whatever the default settings happened to be.

How to tell if you're actually being cited

Standard analytics don't help much here, since a citation often doesn't generate a click. But there are a few practical checks.

Ask directly. Build a list of twenty questions your content answers, phrased the way a real person would ask them, not the way you'd phrase them internally. Run them through a few assistants. Note who gets mentioned, who gets cited, and what gets retold accurately.

Do this monthly and keep the results in the same spreadsheet. The trend matters more than any single result, since answers vary run to run.

Watch how you get paraphrased, not just whether you get mentioned. If a model misstates your claim, the cause is almost always in the text. Something was ambiguous, the context sat too far away, or the claim leaned on something said earlier. That's the most useful feedback you'll get from this whole exercise.

Track whatever traffic does arrive from these systems, however small the number. It's usually small, but it tends to be higher intent. Someone arrives after already getting an answer and wants more.

And keep tracking classic search. It hasn't gone anywhere, and for most sites it still brings in most of the traffic. This is an additional layer, not a replacement.

What not to do

A few things already circulating as "tactics" that carry more risk than payoff.

Hidden text meant only for machines. Old trick, same ending. It gets discovered, and when it does, it takes the whole site down with it.

Piling up questions and answers with nothing behind them. Thirty questions with a one-line empty answer doesn't make you a source, it makes you noise. Ten questions with real answers works better.

Publishing large volumes of generated text. If your content is assembled from the same sources a model already has access to, you have nothing to offer. What gets cited is what exists nowhere else: your data, your experience, your numbers.

Forgetting that a human still reads this. Text optimized purely for extraction ends up choppy, dry, and nobody shares it. And if nobody shares it, over time it loses the trust signal that would have made it worth citing in the first place.

What actually pays off

If you boil all of this down to one thing, it's this: have something others don't, and say it clearly and specifically.

Your own field data. Your pricing. Your process, laid out step by step with actual timelines. What you tried that didn't work. A number from your last project.

That's the only content a system can't assemble from other sources, and that's the only reason it has to name you at all.

Everything else, structure, headings, dates, structured data, exists to make that one thing easy to find and safe to pass along. Useful, but only if there's something worth passing along underneath it.

Next in AI SEO & GEO