Private Pierce

Public Facts and AI Answers: What Retrieval Research Establishes

The evidence cited on this page establishes how a benchmark task and a documented retrieval method use supporting material. It does not establish how every live answer system selects sources, which public-fact categories it will surface, or which page type it will prefer.

Updated

Not legal advice. This page is research, not compliance guidance.

What retrieval-and-citation process does the evidence establish?

The ALCE benchmark defines an end-to-end citation task that retrieves supporting evidence and generates an answer with citations; in its 2023 ELI5 evaluation, even the best evaluated models lacked complete citation support 50% of the time.

The benchmark result has a narrow scope. It belongs to ALCE's 2023 ELI5 evaluation and the models evaluated there. It is not a current universal citation-support failure rate, and it does not establish how any named live answer system chooses or cites sources.

For this site's publishing practice, How Sources Are Made AI-Citable owns the structural citation signals: a stable canonical URL, an evidence layer, methodology, source references, freshness metadata, and limitations. Methodology owns the research corpus's freshness and refresh logic, source hierarchy, and citation discipline. Those practices describe this site's evidence structure; they do not predict selection by a live system.

  • Established: ALCE frames the task as retrieving supporting evidence and generating an answer with citations.
  • Historically scoped result: in ALCE's 2023 ELI5 evaluation, even the best evaluated models lacked complete citation support 50% of the time.
  • Not established: a current universal error rate or the source-selection behavior of any named live answer system.
  • Site-specific boundary: structural citation practices describe the page's evidence structure but do not establish that a live system will select it.

Which public facts appear in AI answers?

The evidence cited on this page does not establish which categories of public facts about a person or company tend to appear in AI answers; the site's exposure verbs provide classification vocabulary, not a fact inventory.

Exposure Verbs owns the definitions of visible, searchable, indexable, reusable, actionable, and the site's other four exposure verbs. That vocabulary can keep different forms of exposure distinct without claiming that a live answer system selects any particular category of public fact.

The sources cited on this page do not supply a supported inventory of person or company facts that tend to appear in AI answers, so no such inventory is asserted here.

The visible, searchable, indexable, reusable, and actionable framework is a related classification route. It does not turn the unresolved question about live AI answers into a measured result.

  • The exposure verbs classify how information can be encountered or used.
  • An exposure category is not evidence that a live answer system will surface it.
  • Structured-data topics and the existence of public pages do not supply a public-fact inventory for AI answers.
  • Not established: which person or company fact categories tend to appear in AI answers.

Which page types are ready for retrieval?

The cited RAG method retrieves text documents from an input and uses them as added context for generation, while the cited experiments used one Wikipedia dump rather than the live public web; this evidence does not rank page types for retrieval readiness.

The documented RAG flow is a method boundary, not a live-web ranking. It retrieves text documents from an input and uses the retrieved material as added context when generating a target sequence. In the cited experiments, the non-parametric knowledge source was one Wikipedia dump rather than the live public web.

The evidence cited here does not compare web page types, such as profile pages or landing pages, for retrieval readiness. Turning the experiment into a universal page-type ranking would extend beyond both the corpus used and the retrieval claim the evidence supports.

Canonical Source Readiness by Artifact identifies which of this site's own public artifacts are structurally citation-ready. That inventory is site-specific and is not evidence that one page type is generally preferred by live answer systems.

  • Established method: retrieve text documents and use them as added context during generation.
  • Experimental boundary: the cited RAG experiments used one Wikipedia dump rather than the live public web.
  • Site-specific owner: the artifact-readiness page identifies structurally citation-ready artifacts on this site.
  • Not established: a universal comparison or ranking of page types for retrieval readiness.

Related research

Submit a correction