publishing.co.uk
AI Search

Author Website SEO for AI Search: What Your Site Can and Cannot Do


In brief

Your author website is almost never the source an AI cites when it recommends a book. Our index of 155,569 citations puts Goodreads, Wikipedia and YouTube at the top and no author-owned domain in the top ten. The site earns its keep by settling who you are, holding the facts nobody else gets right, and catching readers who arrive already sold. Start with robots.txt: if the search crawlers are disallowed, nothing else you do to the site matters.

Last reviewed by Robert Prime — August 2026


Your author website is almost never the source an AI cites when it recommends a book. Goodreads, Wikipedia and YouTube are.

We log every source the engines quote when a reader asks them for something to read. As of 4 August 2026 the index holds 155,569 citations across 8,947 domains and 981 audited titles, and ranked by citations the top of it runs goodreads.com (10,116), en.wikipedia.org (8,598), youtube.com (6,615), fivebooks.com (4,506), penguinrandomhouse.com (3,631) and reddit.com (2,359). Amazon sits 20th, on 862.

That head is thinner than the ranking makes it sound. Goodreads, the most-cited domain of all, takes 6.5% of citations. The whole top ten adds up to 28%. The other 72% is spread across roughly 8,900 domains, which is why “get cited” is a harder instruction than it looks.

No author-owned domain appears in the top ten. We do not classify domains by ownership, so there is no honest total to give you for the category. What we can see is that the highest-placed author site anywhere in the index is alisonwearing.com, fifth in the memoir genre on 349 citations, about 0.2% of everything logged. An author’s own site can be cited. It is not what the engines reach for.

The index covers ChatGPT, Claude, Gemini and Perplexity, the four engines that show their sources; runs against our simulated Amazon shopping assistant are kept separate and excluded from every figure above. It is our own measurement, on our sample, with our prompts, and the method is published with the findings. One number from it reframes the exercise: at the June 2026 snapshot, when the index covered 916 audited titles, 27% of them had never been named by any engine, for any reader question, and only 3% were recommended reliably by name. Those two figures have not been recut since.

If you are building a website in order to be the thing ChatGPT quotes, you have misunderstood the job. The site still matters. Most of what gets sold to authors as “AI SEO” does nothing about the reasons why.

Google ranks pages. Answer engines assemble answers.

They share an input and do different jobs. Google hands you ten links and lets you choose. An answer engine writes one paragraph and footnotes it, so it needs only a handful of sources, and it picks the ones that let it state a fact without hedging.

Google SearchAnswer engines
Unit of competitionThe pageThe claim inside the page
How many winTen blue links, plus featuresA handful of cited sources per answer
What it rewardsRelevance, links, page experienceThe clearest, most checkable statement of a fact
How you measure itSearch Console: impressions, clicks, positionSearch Console shows impressions in AI features, with no clicks, CTR or query data

When a reader asks an assistant for a recommendation, the assistant wants a source that has already made the comparison: a list, a review aggregator, a reference page. Your site talks about one author, so it is the wrong shape for that question. The general mechanics are in the AI book discovery guide.

Google is blunt about the entry ticket. To be shown as a supporting link in AI Overviews or AI Mode, a page needs to be indexed and eligible to appear with a snippet, and Google states there are no additional requirements beyond that. It also describes a query fan-out technique, where one question becomes many background searches, which spreads the citations across more sources than a results page does.

What your website is actually for now

The site has three jobs and none of them is being the citation.

The first is identity. Two authors share a name: one writes Cornish mysteries, the other wrote a physiotherapy textbook. Every engine has to decide whether they are the same person, and it decides from whatever it can find. If your site answers that plainly, and agrees with your Amazon Author Central bio and your Goodreads profile, the guessing stops. If your site is silent and the profiles disagree with each other, engines merge people who should stay separate.

The second is holding the facts of last resort. Series reading order. Which books are standalones. Pen names. Which edition carries the new epilogue. Where the UK paperback is actually available. We see these come back wrong in the AI Discovery Score reports we run, often enough that we look for them every time. The index logs which sources were cited rather than whether the answer was right, so take that as experience and not as a measurement. A page stating them plainly in text is cheap insurance either way.

The third is being somewhere to land. Zero-click search means fewer visits, and the ones you get arrive knowing something about the book already. We have not measured what that is worth in sales, so we will not claim it converts better. The page should carry the buy links and an email signup, because the assistant keeps the click and you will not get many. Our author email list guide covers what to offer in exchange.

What does get cited, and how to get near it

The domains at the top of our index are mostly open to you, and none of them is yours.

  • Goodreads. The most cited domain in the index at 10,116 citations, and the profile most authors half-finish. Claim it, list every title, get the series numbering right, let reviews accumulate. Our Goodreads for authors guide covers it properly.
  • Amazon. Twentieth in the index on 862 citations, 0.6% of the total. It is not where the engines learn about books; it is where the reader who has already decided goes to buy. A complete Amazon Author Central profile with a real bio, photo and every title attached earns its keep on that second job.
  • Wikipedia and Wikidata. en.wikipedia.org is the second-most-cited domain in the index. Wikipedia is not something you write about yourself, and most authors will never meet the notability bar. Wikidata is its machine-readable sibling and a far lower hurdle, but we have not measured what an item does to an author’s citations, so treat it as housekeeping rather than a lever.
  • YouTube. Third on 6,615 citations: interviews, podcast clips and reader reviews that name your book. Video hosted on your own domain does nothing here.
  • Curated lists. fivebooks.com is fourth and crimereads.com ninth. Somebody else’s best-of list carries weight your own list of your own books never will. Pitch the sites that already publish them.
  • Reddit. Sixth on 2,359 citations, 1.5% of the index. Web-wide it is reported as one of the most-cited domains anywhere; the published estimates of its share vary so wildly between engines and methods that we will not quote one. For books it is a slow, unbuyable payoff and a long way from where the citations are.

What actually helps on your own site

Should you block AI crawlers on your author website?

The first thing we check when an author sends us a URL is the robots.txt, and it is the fault we find most often. Fixing it costs nothing.

Every major assistant runs separate crawlers for separate purposes. Blocking the wrong one takes you out of the answers while doing nothing you intended.

CrawlerRun byPurposeEffect of blocking it
GooglebotGoogleCrawls for SearchRemoves you from Search, and therefore from AI Overviews and AI Mode
Google-ExtendedGoogleManages whether content Google crawls may be used for training future generations of Gemini modelsContent not used to train future Gemini models. Does not affect Search inclusion and is not a ranking signal in Google Search
OAI-SearchBotOpenAIIndexes content for ChatGPT searchYour site will not appear in ChatGPT’s search answers
GPTBotOpenAICollects content for model trainingContent is not used to train OpenAI models
ChatGPT-UserOpenAIFetches a page on a user’s requestOpenAI notes robots.txt rules may not apply to user-triggered fetches
Claude-SearchBotAnthropicIndexes content for Claude’s searchAnthropic says this may reduce your visibility in search results
Claude-UserAnthropicFetches pages in response to a user questionAnthropic says this may reduce visibility for user-directed web search
ClaudeBotAnthropicTrainingFuture material excluded from training datasets
PerplexityBotPerplexityIndexes for Perplexity resultsRemoved from Perplexity’s index
Perplexity-UserPerplexityUser-initiated fetchPerplexity says this one generally ignores robots.txt

Three checks, in this order.

  1. robots.txt. Open yourdomain.com/robots.txt and confirm Googlebot, OAI-SearchBot, Claude-SearchBot and PerplexityBot are not disallowed. OpenAI says a robots.txt change takes around 24 hours to register on its side, so do not judge it the same afternoon.
  2. Cloudflare, if you sit behind it. Cloudflare says its AI traffic options “are live now, and can be configured by all existing customers in their zone Settings”, free tier included. The defaults change on 15 September 2026, and only for new arrivals: “For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.” The line that matters most covers crawlers doing both jobs at once: “Multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training.” Block training on a Cloudflare zone, through the new controls or the legacy Block AI bots service, and you can take Googlebot out with it. An owner who wants no change to training crawlers that also crawl for search “can easily mark this in their Security settings any time leading up to September 15”.
  3. Snippet directives. Google’s nosnippet rule removes snippet eligibility, and a page that cannot show a snippet cannot appear in AI Overviews or AI Mode. Some SEO plugins ship with a max-snippet cap already set. The documented value for no cap is -1.

An about page that reads as a fact sheet

Most author about pages open with atmosphere: coffee, a dog, the Yorkshire weather. Fine for a human reader, and it gives a machine nothing to record.

Write the first two sentences so they state, in plain text, your full name as it appears on your books, what you write, and one verifiable credential. Then the corroboration: links to the profiles that already exist for you. Amazon Author Central, Goodreads, your publisher’s page, ALLi or Society of Authors membership, a Wikidata entry if you have one, your LinkedIn. Schema.org’s sameAs property exists to make that same-person claim in markup, but the plain links do most of the work, because they are what an engine follows when it wants a second source.

Consistency holds the whole thing together. Same name spelling, same bio facts, same book list. If Goodreads says four books and your website says six, an engine cannot tell which is stale, and its safest move is to commit to neither.

Google’s position deserves stating plainly, because plenty of people are being sold the opposite. Structured data is not required for its generative AI features, and there is “no special schema.org structured data that you need to add”. Google still recommends it as part of ordinary SEO, because it keeps you eligible for rich results.

So do it in an afternoon, then stop thinking about it. Markup lets a machine state who you are without guessing; it is not a lever on whether you get quoted. If you followed our guide to getting ChatGPT to recommend your book, this is that advice with the ceiling made explicit.

The current vocabulary is schema.org version 30.0, released 19 March 2026. The types that earn their place:

PageTypeProperties to set
About / author pagePerson inside a ProfilePagename, sameAs, knowsAbout, jobTitle, award, description, image, url
Individual book pageBookauthor, name, isbn, bookFormat, bookEdition, numberOfPages, inLanguage, datePublished, publisher, workExample
Blog postsArticleauthor (pointing at your Person), datePublished, headline
Publisher / imprint, if you have oneOrganizationname, url, sameAs

Google lists an author page among the valid uses of ProfilePage, with mainEntity as the required property. knowsAbout gets skipped most often, and it is the property that states the topics you genuinely have expertise in.

Google’s Book rich result, the one with the buy and borrow buttons, is closed to individual authors: the documentation limits it to book providers with a wide selection of books, who register interest and are onboarded, with no guarantee of participation. Mark up Book anyway for clarity, but do not expect that result. FAQ rich results have gone altogether. Google’s documentation says the feature stopped appearing in Search on 7 May 2026 and the documentation was removed on 15 June 2026, so FAQPage is still parsed and the expanded box under your listing no longer exists.

Write the facts as sentences

Assistants extract claims, and a claim has to be stated to be extracted. “Available now in all good bookshops” is not a claim. “The Hollow Tide is a 312-page crime novel published in paperback on 14 March 2026, and is the second book in the Marlow series after Salt and Ash” is seven checkable claims.

Put a short factual block on every book page: title, series and number, format, page count, ISBN, publication date, publisher or imprint, genre, and one line on the reader it is for. Put publication history and awards on the about page as dated statements.

This breaks in two familiar ways. Facts that live only inside a cover image or a video are not text and do not get extracted; Google’s AI guidance asks for content in textual form, with media supporting it. And facts that appear only once JavaScript has run are a gamble: Google documents a rendering stage, while the published crawler documentation from OpenAI and Anthropic describes fetching content and says nothing about executing JavaScript. View your own page source and search for your ISBN. If it is not in there, assume some engines never see it.

What does not help

A genre page about your own books. A page titled “Best Cosy Crime Novels UK” on an author’s own domain, listing that author’s own three titles, is not a curated list. Readers can tell, and so can an engine cross-referencing it against the lists that are.

Anyone promising you a ranking inside ChatGPT. There is no ranking. No engine publishes or sells position. What can honestly be measured is how often you are named and cited across repeated prompts, which is what our index does, and that buys you a baseline and a re-test rather than a control panel.

llms.txt. A proposed plain-text file at your domain root that summarises your site for machines. As of August 2026 no major AI company treats it as a control or a ranking input. Google’s AI optimisation guide of 15 May 2026 says llms.txt files “will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them”, and OpenAI governs crawler access through robots.txt while publishing an llms.txt for its own developer docs, which is where much of the confusion started. Write one if you like; it takes ten minutes and is harmless. An “llms.txt package” sold as an AI visibility service is a text file with an invoice attached.

Amazon’s shopping assistant does not read your website

Amazon announced on 13 May 2026 that Rufus was being renamed Alexa for Shopping, folding the assistant into Alexa+ and moving it into the main search bar. Amazon’s UK announcement page still calls it Rufus and says nothing about the rename, so a reader on amazon.co.uk in August 2026 may well be looking at the old name on the same assistant. Amazon still describes it as trained on “Amazon’s extensive product catalog, customer reviews, community Q&As, and information from across the web”. The catalogue is the centre of gravity, and your website is not in it. That puts the levers on the detail page:

  • Title, bullets and description, written as statements about who the book is for rather than adjectives about how good it is
  • A+ Content, including the text inside the image modules
  • Category and keyword selection
  • Reviews and answered questions
  • A complete Author Central profile with a real bio and photo

We sell A+ Content design from £89, so that list is not disinterested advice. You can build A+ Content yourself in KDP for nothing but time, and for a single title on a modest budget that is usually the right call. Our A+ Content guide walks through the DIY version, and Amazon Rufus goes deeper on the mechanics.

When an author website is worth £199, and when it is not

We build author websites at £199 one-off plus £15 a month: a designed site with home, books, about and contact pages, up to 15 books with covers and buy links, your own web address registered to you, one revision round before launch, hosting and support on the care plan. Author Pro, for catalogues up to 40 books with dedicated series pages, starts from the same £199 and is quoted individually depending on how many books you have. No blog, no shop, no page builder. The constraint is deliberate.

Do not buy it if you have one book, no email list, and a cover you are not confident about. The £199 does more for you spent on the cover. Our author website essentials guide covers the do-it-yourself route honestly, and Carrd will get a debut author a decent single page for a fraction of that.

Buy the website when you have three or more titles, a series that needs its order explained, or a launch coming and nowhere for the traffic to land except an Amazon listing.

Either way, run the free KDP readiness audit first. If the book files themselves are wrong, formatting is the problem to solve before the website is.

Frequently asked questions

Does ChatGPT read my author website?

Sometimes, and it is rarely the page it ends up quoting. OAI-SearchBot indexes sites for ChatGPT’s search answers and ChatGPT-User fetches a page when someone asks about it directly, so an unblocked author site is readable. On open questions of the what-should-I-read-next kind, the sources that come back are overwhelmingly Goodreads, Wikipedia, YouTube and best-of lists.

Can I see AI visibility in Google Search Console?

Partly: Search Console’s Generative AI performance report shows how often links to your site were shown in a generative AI feature on Google Search, segmented by page, country, device and date, with no clicks, CTR or query data. It is still a limited rollout, so it may not be in your account yet. You can see that you appeared without learning what it was worth.

How do I find out whether ChatGPT knows about my book?

Ask the four engines that show their sources (ChatGPT, Claude, Gemini and Perplexity) the same question a reader would ask, along the lines of “what are some good [your genre] novels by British authors”, and see whether you are named without being prompted. Then ask each one directly about your title and check the facts in the answer. Wrong publication dates, invented series order and merged author identities are the usual failures, and they normally trace back to disagreements between your website, Goodreads and Amazon. Our free AI Discovery Score runs two engines in about 90 seconds; the £29.99 report runs five, those four plus a simulated Amazon shopping assistant, and gives you the verbatim answers. We sell it, so weigh that accordingly.

External references

Robert Prime — Founder of publishing.co.uk

About the Author

Robert Prime

Robert Prime is a best-selling self-published author, veteran eCommerce strategist, and the founder of publishing.co.uk. With over 25 years of experience in digital business he brings a battle-tested perspective to the publishing industry. After experiencing firsthand the archaic, headache-inducing process of formatting a KDP-compliant book for his own best-seller, Google. Panic. Repeat., Robert built publishing.co.uk to solve the problem for other authors. He is also a co-owner of the LoveReading.co.uk network (the UK’s leading book discovery platforms), founder of the Amazon growth agency MrPrime.com, and a member of the Forbes Business Council.