Grounding markup language

GCML structures and annotates documents to enhance LLM data retrieval, addressing inefficiencies in unstructured data sources and improving response accuracy and efficiency by providing structured metadata for LLMs.

WO2025248126A1PCT designated stage Publication Date: 2025-12-04ICS AI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/065068
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-30
Filing Date
2025-05-30
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing approaches for grounding large language models (LLMs) face challenges in retrieving relevant and reliable data from unstructured and unlabelled sources, particularly when dealing with large documents, leading to inefficiencies in computing resources and inaccurate responses due to the unstructured nature of the grounding data.

Method used

The implementation of a Grounding Content Markup Language (GCML) that structures and annotates documents with metadata, including sections, importance levels, reliability data, and usage instructions, to enhance the retrieval of relevant data for LLMs, complementing retrieval-augmented generation techniques.

Benefits of technology

GCML provides structured grounding data that improves the accuracy and efficiency of LLM responses by ensuring relevance and reliability, reducing computational overhead and minimizing hallucinations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025065068_04122025_PF_FP_ABST
    Figure EP2025065068_04122025_PF_FP_ABST
Patent Text Reader

Abstract

A computer system includes at least one processor and a memory storing at least one markup language document The document includes grounding data for use in grounding a generative AI model, the markup language document defining one or more of: a plurality of sections of the grounding data; an importance level of each of the plurality of sections of the grounding data; reliability data indicating the reliability of the grounding data; and usage instructions indicating an intended use of the grounding data. The system receives a first input such as a user query and generate a second input for a generative artificial intelligence, AI, model including the first input and at least some of the markup language document to act as grounding data in providing a response to the first input. The disclosure also provides a tool for applying the markup language, either manually or using a generative AI model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] GROUNDING MARKUP LANGUAGE

[0002] Generative artificial intelligence (Al) models have recently been made widely available. For example, large language models (LLMs), which are one example type of generative Al model, receive input in the form of natural language (referred to as a prompt) and provide output in natural language. LLMs are large in the sense that they are trained on a huge quantity of data, encompassing many diverse data sets. They are also large in the sense that they have billions of parameters. The huge training data set and parameter set gives such models the ability to be employed in a wide range of tasks, including question answering, translation, summarisation, and classification.

[0003] Whilst users may interact with LLMs by authoring and submitting their own prompts, such models may also be integrated into other applications which make use of LLMs to carry out tasks. The application may make use of template prompts (also referred to as "metaprompts" or "system prompts"), which are filled with data by the application and then submitted by the application. In some examples, the application may then take action based on the received response (e.g. send an email, create a file, kick off a process etc).

[0004] In the context of LLMs, grounding has been adopted to improve the accuracy, relevance and reliability of results. Grounding involves the use of specific contextual information and / or personalised data (hereinafter grounding data) from which the Al model can draw insights when being executed to perform a task. For example, a grounded Al model can leverage company information, such as employee names etc, medical patient-related information, trade information to support financial trading and / or student contextual information for education. The primary aim of grounding is to guide the LLM to generate responses relating to the grounding data which has been provided but yet leveraging the huge and diverse range of knowledge upon which it has been trained. The use of grounding data may also prevent or reduce hallucinations, in which the LLM invents plausible data rather than providing accurate responses. The grounding data is typically appended to or otherwise included in a prompt.

[0005] Whilst in some contexts the grounding data may be available in structured data stores (e.g. relational databases storing employee or medical information), in many contexts the grounding data may comprise unstructured data sources or data sources in a wide range of formats. These may for example include unstructured text documents.

[0006] Summary

[0007] Some approaches have been employed for retrieving relevant grounding data from text or multimedia sources, such as retrieval-augmented generation (RAG). These approaches rely on semantic similarity between the input data and the grounding data to retrieve suitable grounding data for inclusion in a prompt. However, difficulties arise because of unstructured and unlabelled nature of the grounding data. Furthermore, where large documents (e.g. text documents with hundreds or thousands of pages) are used for grounding, the whole document may not be relevant. It may also not be desirable to include all of the document, either because the surfeit of information does not assist the LLM in providing accurate responses or because it unnecessarily takes up computing resource or network bandwidth in cases where the LLM is hosted remotely.

[0008] In order to address these issues, the disclosure provides markup languages, which can be applied to documents comprising grounding data to provide structure thereto. The markup languages may provide a mechanism for assisting retrieval of relevant data from documents, such that relevant parts of the grounding data can be readily identified for inclusion in prompts, complementing or supplanting existing RAG techniques. Furthermore, the presence of the markup itself imparts structure to the grounding data sources. The markup languages are generally referred to herein as grounding content markup language (GCML).

[0009] The disclosure also provides tools for applying the GCMLto unstructured documents, and systems and methods that make use of documents in GCML format to ground LLMs at inference time. In this context "unstructured" may simply refer to documents without GCML applied thereto, and may encompass documents that already have some internal structure present (e.g. HTML documents).

[0010] According to a first aspect disclosed herein, there is provided a computer system comprising: at least one processor; a memory storing at least one markup language document, the document comprising grounding data for use in grounding a generative Al model, the markup language document defining one or more of: a plurality of sections of the grounding data; an importance level of each of the plurality of sections of the grounding data; reliability data indicating the reliability of the grounding data; and usage instructions indicating an intended use of the grounding data; the memory storing instructions which when executed by the processor cause the processor to: receive a first input; and generate a second input for a generative Al model including the first input and at least some of the markup language document to act as grounding data in providing a response to the first input.

[0011] The system may be configured to select the at least one markup language document from amongst a plurality of markup language documents based on the first input. The system may be configured to retrieve one or more sections of the at least one markup language document based on the first input, and include them in the second input. Alternatively, the system may be configured to include substantially all of the markup language document in the second input. The second input may be a prompt.

[0012] The plurality of sections may include headings and paragraphs. The paragraphs may comprise an attribute defining an importance thereof.

[0013] The reliability data may include a verification status, indicating whether the grounding data has been verified. The reliability data may include a confidence score, indicative of a confidence in the accuracy of the grounding data. The reliability data may include a source citation.

[0014] The usage instructions may indicate an intended use of training or answer generation. The usage instructions may indicate that the grounding data is not suitable for particular uses, such as legal advice. The usage instruction may indicate potential bias in the data.

[0015] The markup language document may define one or more of a title of the document, a classification of the document, an author of the document, a publication date of the document or a language of the document.

[0016] The markup language document may define contact information in relation to a person or organisation referred to in the grounding data. The markup language document may define quotations and / or citations. The quotations may comprise one or more attributes including whether the quotation is direct, paraphrased or historical, a speaker, a timestamp and a context.

[0017] The markup language document may define subjective material in the grounding data. The markup language document may include one or more of an author, stance, relevant expertise and publication context associated with the subjective material. The markup language document may define parts of the grounding data that relate to disputed claims.

[0018] The markup language document may define parts of the grounding data relating to instructions and / or procedures, comprising ordered steps.

[0019] The markup language document may define parts of the grounding data relating to formal rules, regulations or standards. The markup language document may define the authority responsible for such rules, regulations or standards, and / or data relating to exceptions or penalties for non- compliance.

[0020] The markup language document may define parts of the grounding data related to exclusion, limitations or contraindications in clinical or research contexts.

[0021] The markup language document may define parts of the grounding data comprising numerical data and / or statistical findings. The markup language document may define parts of the grounding data comprising warnings, alerts and / or precautions. The markup language document may define parts of the grounding data relating conditional logic and / or temporal sequences.

[0022] The markup language document may be an XML-based document. By XML-based, it may be meant that the document complies with a suitable XML standard and / or comprises tags in the style of an XML document. The markup language may be defined in an XML schema.

[0023] The generative Al model may be configured to receive input text, for example a prompt. The generative Al model may be configured to output text. The generative Al model may be a large language model. The generative Al model may be configured to receive input in other modalities in addition to text, including audio, images and video. The generative Al model may be configured to generate audio, images or video.

[0024] References herein to the markup language document "defining" material may refer to the markup language document including suitable mechanisms for marking parts of the document as comprising that material. For example, the markup language document may comprise tags delimiting parts of the document, and optionally attributes associated with the tags.

[0025] According to a second aspect of the disclosure there is provided a computer system comprising at least one processor and memory, the memory storing instructions which when executed by the processor cause the processor to implement: a tagging tool configured to apply markup language to grounding data, wherein the markup language defines one or more of: a plurality of sections of the grounding data; an importance level of each of the plurality of sections of the grounding data; reliability data indicating the reliability of the grounding data; and usage instructions indicating an intended use of the grounding data.

[0026] The tagging tool may comprise a user interface configured to receive user input to apply the markup language.

[0027] The tagging tool may be configured to generate an input for a generative Al model, the input comprising: a definition of the markup language, an untagged document, and instructions that cause the generative Al model to apply the markup language to the untagged document. The tagging tool may be configured to provide the input to the generative Al model, and in response, receive a tagged document in the markup language. The tagging tool may be configured to store the tagged document.

[0028] Further optional features of the computer system of the second aspect are defined above in relation to the first aspect, and may be combined in any combination. Furthermore, the systems of the first and second aspect may be combined. That is to say, the disclosure extends to a system that applies markup language to documents and then uses the marked-up documents.

[0029] According to a third aspect of the disclosure, there is provided a computer-implemented method, comprising: storing at least one markup language document, the document comprising grounding data for use in grounding a generative Al model, the markup language document defining one or more of: a plurality of sections of the grounding data; an importance level of each of the plurality of sections of the grounding data; reliability data indicating the reliability of the grounding data; and usage instructions indicating an intended use of the grounding data; the memory storing instructions which when executed by the processor cause the processor to: receiving a first input; and generate a second input for a generative Al model including the first input and at least some of the second markup language document to act as grounding data in providing a response to the first input.

[0030] Further optional features of the method of the third aspect are defined above in relation to the first aspect, and may be combined in any combination.

[0031] According to a fourth aspect of the disclosure, there is provided a computer-implemented method comprising: using a tagging tool to apply markup language to grounding data, wherein the markup language defines one or more of: a plurality of sections of the grounding data; an importance level of each of the plurality of sections of the grounding data; reliability data indicating the reliability of the grounding data; and usage instructions indicating an intended use of the grounding data.

[0032] Further optional features of the method of the fourth aspect are defined above in relation to the second aspect, and may be combined in any combination. The methods of the third and fourth aspects may be combined.

[0033] The disclosure extends to computer-readable media and computer program products corresponding to any preceding aspect, or storing documents according to any of the markup languages disclosed herein. The disclosure extends to combinations of any of the aspects defined herein.

[0034] Brief Description of the Drawings To assist understanding of the present disclosure and to show how embodiments may be put into effect, reference is made by way of example to the accompanying drawings in which:

[0035] Figure 1 is a schematic block diagram of an exemplary architecture for delivering an Al generated interaction with a user.

[0036] Figure 2 is a schematic block diagram of an exemplary user device.

[0037] Figure 3 is a schematic block diagram of an exemplary display.

[0038] Figure 4 is a schematic diagram of a elements of an example markup language.

[0039] Figures 5 - 13 are schematic diagrams illustrating example tags that may be incorporated into GCML.

[0040] Figure 14 is a schematic illustrating a smart grounding process.

[0041] Figure 15 is a flow chart illustrating a process of applying GCML to unstructured documents.

[0042] Figure 16 is a flow chart illustrating a process of employing GCML documents in a RAG setting.

[0043] Detailed Description of Embodiments

[0044] In overview, examples of the disclosure provide a markup language that assist in grounding LLMs or other generative models. The markup language imparts structure to the grounding data sources to assist retrieval and / or to guide the LLM in correctly interpreting the grounding data, and may also include metadata that can be used to ensure the grounding data is put to the correct use.

[0045] "Grounding" in the context of Al typically refers to the method through which Al systems ensure their responses are based on relevant, specific and contextually appropriate knowledge or data. More particularly, grounding typically relates to the inclusion of data in inputs to Al models (e.g. in prompts). Such data provides either contextual background that leads the model to provide more relevant or accurate responses, or that in some circumstances comprises data on which the output of the model should be based. In the latter, this may take the form of question answering or summarisation of the data included in the prompt.

[0046] A major challenge with grounding content is that it typically fails to perform optimally because it was originally created for human consumption, in forms such as text and visual formats. The techniques herein address this issue, by providing a markup language that designed to provide suitable structure to documents. These techniques complement existing methods like the RAG (Retrieval- Augmented Generation) approach, either by enabling the retrieval of various parts of the documents according to their tags in the markup, or by providing the marked-up document to the LLM such that it can interpret the tags therein, thereby directing the LLM to produce output based on the marked- up document.

[0047] Figure 1 is a schematic diagram illustrating an example environment in which examples of the disclosure may operate. It illustrates an exemplary architecture for delivering an Al generated interaction with a user. Reference number 300 denotes a computer device which may be used by a user. The computer device may be any suitable computer device, including but not limited to smart phones, tablets, laptops, desktop computers, wearable devices or any other form of computer device that allows a user to interact with it to perform computing functions. The user device 300 has a network interface for communicating with a backend infrastructure 302 via a communication network 304. The communication network may be any type of communication network enabling such communication to be achieved including without limitation wired and wireless networks, wireless networks including for example telecommunications networks and / or Wi-Fi networks.

[0048] The backend infrastructure 302 may comprise one or more server, the or each server executing a large language model (LLM) capable of delivering an output based on an input received from the user device. Each LLM may also have access to grounding data sources specific to the task to be implemented by the LLM. The LLMs may be capable of understanding and utilising the markup languages disclosed herein. The LLMs may include generative Al models which are capable of receiving a natural language input and generating an output, which may be in natural language or in another form. The servers may also include machine learning models which are capable of receiving a defined input and generating an output. In other words, the servers may store trained machine learning models that are task-specific, such as speech-to-text or translation models. The models may also include models trained to mimic the output of an LLM, such as a small language model (SLM). Hereinafter "LLM" is used as a shorthand for all of these possible models, but it will be appreciated that the disclosure extends to all such models.

[0049] Generative Al models have recently been made widely available. Such generative Al models currently take the form of large language models (LLMs), so-called because they are trained on a huge quantity of data, encompassing many diverse data sets. Open Al has developed GPT-3 (generative pre-trained transformer 3). Another manifestation, BERT (bi-directional encoder representations from transformers) has been developed by Google. Such LLMs currently use a transformer architecture. These LLMs have the ability to use natural language prompts, which may be employed to ask the LLM to solve a task. The LLM uses a transformer deep learning network to implement the trained language model. The LLM has been trained on a very large dataset (for example in the order of billions of tokens). It is a generative model that can generate text, data or code in receipt of a prompt.

[0050] Reference number 306 denotes a document repository which may contain documents having the markup language discussed herein. The document repository 306 is shown connected to the communication network 304. Note that this is highly schematic. In certain examples, the document repository 306 may be connected to the user device 300 via a communication network which is entirely separate to the communication network 304 which is used to connect the user device to the server. In particular, the user device 300 and the repository 306 may form part of an entity's internal infrastructure, where the entity is for example a company or corporation, or other type of organisation, which is designated to carrying out a particular function. In one example, that function is the routing of calls in a call centre or other organisation.

[0051] Figure 2 is a schematic block diagram of an exemplary user device 300. The user device 300 comprises a processor 400 and a memory 402. The processor 400 may be any suitable type of computing device, appropriate to the nature of the user device. For example, it could be implemented by a single core processor, a multi core processor or a computing cluster. If the user device 300 is a smart phone, it is likely that the processor 400 will be embedded in a single casing with the other elements of the smart phone and could be a single core or multi core processor. The memory 402 may include volatile memory and non-volatile memory. The non-volatile memory may include read-only memory, for example for storing computer instructions to be executed by the processor 400 to operate the computing device 300. The volatile memory may include for example RAM for holding data for carrying out operations of the user device. It will be appreciated that any suitable kinds of memory may be utilised, as suitable to the nature of the user device 300. As mentioned, operations of the user device 300 are carried out by the processor 400 executing instructions which have been stored in the memory 402. The user device 300 also includes a network interface 404 which enables it to communicate with the communication network 404 via a protocol suited to the communication network (be it wired or wireless etc).

[0052] The user device 300 also includes a user interface 406. The user interface is provided to enable or facilitate interaction between a user of the user device 300 and the processor 400. To that end, the user interface 406 includes an input mechanism which allows a user to enter information into the device, and an output mechanism which allows a user to perceive output from the user device 300. The input mechanism may include a touch screen, buttons, controllers such as a mouse or any other form enabling a user to input to the user device. The input and / or output mechanisms may additionally or alternatively include audio input and output mechanisms (e.g. microphones, speakers etc) or haptic inputs and outputs. Broadly speaking, any suitable input or output mechanisms for interacting with the device 100 may be incorporated. One form of output is to provide a display 408.

[0053] It will be appreciated that while the infrastructure which has been described assumes that the LLMs will be provided in the backend infrastructure 302, it is also possible to have one or more LLM being executed locally by the processor 400 at the user device.

[0054] In certain applications, processing user input may not be done for the purpose of displaying results to the user as indicated in Figure 2 in the display 408, but instead may be used to control other functions and / or make decisions to implement further operations based on the user input.

[0055] Figure 3 is a schematic block diagram of an exemplary display 408 illustrating an application display panel opened to allow the user to interact with the LLM provided in the backend infrastructure 302 by interacting with forms displayed supporting a question-and-answer format interaction illustrated. To that end, the display panel includes arrays of LLM provided text 500a and user input fields 500b into which a user can enter text or select from presented options to input to an LLM. The user input into 500b may for example be natural language text. A user may for example type text using a keyboard, which may form part of the user interface 406 or may be a separate accessory device. User input is provided to the processor 400 for processing and the processor 400 causes the input to be transmitted to the backend infrastructure via the network interface 404.

[0056] Markup Language As discussed herein, the disclosure provides a markup language applicable to content to be used as grounding data. In general, the markup language provides a scheme for annotating documents to impart structure and metadata therein. For example, the markup language may provide tags that are used to demarcate sections and provide metadata in respect of said sections, the demarcation and the metadata providing useful information to the LLM and / or a system retrieving the grounding data (or parts thereof) for inclusion in prompts to the LLM.

[0057] There now follows a discussion of various examples of the structure of the markup languages. In each example described, the mark-up language is based on XML (extensible markup language). That is to say, the markup language is defined by an XML schema. The schema defines constraints on the structure and content of the markup language - for example which elements can reside in which other elements, which attributes are and are not legal to have on a particular element, and so on. It will be appreciated that this is merely one convenient means of defining a suitable markup language, and various other options are within the scope of the disclosure. For example, other markup languages may be employed.

[0058] Figure 4 summarises the elements of an example GCML (grounding content markup language), 600 including document declaration and metadata 610, content description 620, body 630, annotations for contextual and factual accuracy 640, usage instructions 650, and references and further reading 660. Annotations 640 include a fact check, source used for verification, confidence score in the verification, and instructions for duplicate content handling. Usage instructions include how the content is intended to be used, situations where the content should not be used, and warnings of potential biases in the content.

[0059] GCML is tailored specifically for use by LLMs, and involves providing a language that provides structured, clear, and explicit instructions about content. The primary goal of GCML is to enhance understanding, reduce errors like hallucinations, and provide contextual clues to the LLM about how to interpret and use the information.

[0060] Expanding now on the example GCML schema set out in Figure 4, there follows an example of a marked-up document according to the schema.

[0061] The document declaration 610 and content description 620 takes the following form:

[0062] #### 1. **Document Declaration and Metadata**

[0063] Start with a declaration that identifies the document as GCML, followed by comprehensive metadata to help the LLM contextualize the content. '"xml

[0064] <GCML version="1.0">

[0065] <Header>

[0066] <Title>Example Content Title< / Title>

[0067] <SourcelD>unique_identifier_12345< / SourcelD>

[0068] <ContentType>Article< / ContentType>

[0069] <Author>Author Name< / Author>

[0070] <PublicationDate>2024-01-01< / PublicationDate>

[0071] <Language>English< / Language>

[0072] <Keywords>Machine Learning, Al, Data Ethics< / Keywords>

[0073] < / Header>

[0074] In this section, the initial tag <GCML> root element declares the document type and version. In other words, it provides a declaration that the document complies with the GCML schema.

[0075] The <Header> section contains metadata that helps contextualize the document. <Title> names the content, <SourcelD> provides a unique identifier, <ContentType> classifies the document - for example setting out the type of the document, and <Author>, <PublicationDate>, and <Language> describe its origin. <Keywords> offers thematic tags to assist in indexing or retrieval.

[0076] The body 630 section contains the actual content, annotated to indicate its structure and importance, which can guide the LLM's attention and interpretation mechanisms. It takes the following form:

[0077] #### 2. **Content Description**

[0078] This section contains the actual content, annotated to indicate its structure and importance, which can guide the LLM's attention and interpretation mechanisms.

[0079] '"xml

[0080] <Body>

[0081] <Section importance="high">

[0082] <Heading>lntroduction to Machine Learning< / Heading>

[0083] <Paragraph> Machine learning is a subset of artificial intelligence...

[0084] < / Paragraph>

[0085] < / Section>

[0086] <Section importance="medium">

[0087] <Heading>Historical Background< / Heading>

[0088] <Paragraph>

[0089] The concept of machine learning was first proposed...

[0090] < / Paragraph>

[0091] < / Section>

[0092] < / Body>

[0093] The <Body> tag encloses the main content. Each <Section> tag demarcates a particular section of the document, and is optionally annotated with an importance attribute to signal its relevance. Inside each section, <Heading> defines the title of the subsection, and <Paragraph> represents a paragraph containing the actual text. Although a single paragraph is shown in each section, it will be appreciated that in reality, multiple sections are present.

[0094] The annotations for contextual and factual accuracy 640 can help the LLM understand the reliability of the information and how it should be treated to minimize hallucinations, and takes the following form:

[0095] #### 3. **Annotations for Contextual and Factual Accuracy**

[0096] Annotations can help the LLM understand the reliability of the information and how it should be treated to minimize hallucinations.

[0097] '''xml

[0098] <FactCheck status="verified">

[0099] <Source citation="https: / / verifiedsource.com / articlel23">Article on Verified Source< / Source> <ConfidenceScore>High< / ConfidenceScore>

[0100] < / FactCheck>

[0101] <DuplicateContentHandling action="ignore">

[0102] This content has been identified in multiple sources. Prioritize latest version.

[0103] < / DuplicateContentHandling> The <FactCheck> tag provides metadata about the factual accuracy of the content, with a status attribute indicating verification. <Source> includes a citation URL and a human-readable label, while <ConfidenceScore> expresses the reliability level. <DuplicateContentHandling> gives instructions on how to treat repeated content, with an action attribute guiding the LLM's behaviour.

[0104] The usage instruction 650 provides specific instructions can guide the LLM on how to utilize the content effectively, including handling biases or potential misuse. It takes the following form:

[0105] #### 4. **Usage Instructions**

[0106] Specific instructions can guide the LLM on how to utilize the content effectively, including handling biases or potential misuse.

[0107] '"xml

[0108] <Usage>

[0109] <lntendedUse>Training, Answer Generation< / lntendedUse>

[0110] <NotSuitableFor>Sensitive Issues, Legal Advice< / NotSuitableFor>

[0111] <BiasWarning type="gender, race">

[0112] Potential bias in examples used. Use with caution.

[0113] < / BiasWarning>

[0114] < / Usage>

[0115] The <Usage> section outlines how the content should be used. <lntendedUse> specifies acceptable applications, while <NotSuitableFor> warns against inappropriate contexts. <BiasWarning> flags potential biases, with a type attribute identifying the nature of the concern, helping ensure responsible use.

[0116] The references and further reading section 660 provides links or references to additional materials that can aid in deeper understanding or verification. It takes the following form:

[0117] #### 5. **References and Further Reading**

[0118] Provide links or references to additional materials that can aid in deeper understanding or verification. '"xml

[0119] <References>

[0120] <Reference url="https: / / moreinfo.com / more_details">Further Details Here< / Reference> < / References>

[0121] < / GCML>

[0122] The <References> section lists external sources for further reading or verification. Each <Reference> includes a url attribute for the link and a label for display.

[0123] Turning now to Figures 5 to 13, a variety of other suitable tags are discussed that may be incorporated into GCML.

[0124] Figure 5 illustrates tags used to demarcate contact and reference information.

[0125] The <contact_info> tag is used to provide structured contact details for individuals or organizations. It includes attributes such as type (e.g., "person" or "organization") and role (e.g., "author"), which help clarify the entity's function in the document. For individuals, nested tags like <name>, <email>, <phone>, <affiliation>, and <orcid> offer precise metadata that can be used for attribution, correspondence, or academic referencing. For organizations, additional tags like <url> and <address> (with sub-elements for street, city, and country) provide institutional context and location.

[0126] The <reference> tag is designed to cite external sources. It includes a type attribute (e.g., "external jink") and child elements such as <url>, <title>, <access_date>, and <status>. This structure ensures that references are both human-readable and machine-actionable, supporting traceability and verification of claims.

[0127] Figure 6 illustrates tags used to demarcate quotations and citations. The <quotation> tag captures direct, paraphrased, or historical quotes. It includes a type attribute to distinguish the nature of the quote, and may include a source_ref to link to a citation, a confidence score to indicate reliability, and a <timestamp> for temporal context. For direct quotes, <speaker> and <context> provide attribution and situational framing, while <content> holds the quoted material. Paraphrased quotes include <original_context> to indicate where the interpretation originated. Historical quotes may include <temporal_context> and <verification_status> to clarify authenticity. This structure allows for nuanced representation of quoted material, helping systems distinguish between verified statements, paraphrased summaries, and potentially apocryphal historical attributions.

[0128] Figure 7 illustrates tags used to demarcate opinions and subjective content. The <opinion> tag is used to mark subjective statements, such as from experts or commentators. It includes a confidence score and a biasjndicator to help assess the reliability and neutrality of the opinion. Nested elements like <author>, <stance>, and <expertise_relevance> provide context about the speaker and their qualifications. <publication_context> situates the opinion within a specific outlet or format, while <content> contains the actual statement.

[0129] The <subjective_assessment> tag is used for evaluative judgments, such as peer reviews or quality ratings. It includes the assessor's role, the type of assessment, a defined scale, and a numerical <value>. The <rationale> explains the reasoning behind the score, offering interpretability.

[0130] The <controversy> tag captures debates or disputed claims. It includes a <positions> container with multiple <position> elements, each detailing a stance, number of supporters, and strength of evidence. This structure is ideal for representing polarized viewpoints and guiding nuanced interpretation.

[0131] Figure 8 illustrates tags suitable for application to instructions and procedures. The <instruction> tag outlines procedural steps, such as medical or technical protocols. It includes a type (e.g., "procedural") and complexity level. The <title> names the procedure, while <prerequisite> lists any conditions that must be met beforehand. The <steps> container holds ordered <step> elements, each with an order and critical flag to indicate importance. <duration> gives an estimated time, and <safety_warnings> includes <warning> elements with severity levels to highlight risks.

[0132] The <dosage_instruction> tag is tailored for medical contexts. It specifies a drug and condition, and includes dosage details like <starting_dose>, <maximum_dose>, and <timing>. <contraindications> lists conditions under which the drug should not be used. This structured format supports safe and context-aware medication guidance. The <recipe> tag generalizes procedural instructions for scientific or laboratory preparations. It includes a <title>, a list of <ingredient> elements with quantities, and a <method> section with sequential <step> instructions. This format is ideal for reproducible protocols in research or clinical settings.

[0133] Figure 9 illustrates tags suitable for encoding formal rules, regulations, and compliance standards. The <regulation> tag captures legal requirements, including attributes like type, jurisdiction, and status. Nested elements such as <title>, <authority>, and <effective_date> provide context, while <requirement> outlines the specific rule. Optional sections like <penalties> and <exemptions> detail consequences for non-compliance and any exceptions. Similarly, the <policy> tag is used for institutional rules, including a list of <rule> elements and <exceptions>, making it ideal for internal governance. The <compliance_standard> tag documents adherence to frameworks like HIPAA, including verification methods and audit history.

[0134] Figure 10 illustrates tags suitable for defining exclusions, limitations, and contraindications in clinical or research contexts. The <exclusion> tag specifies disqualifying criteria for study participants, with a <rationale> and an <impact_on_generalizability> to explain the implications. The <limitation> tag describes methodological weaknesses, including severity, impact, and mitigation strategies. The <contraindication> tag identifies medical conditions where a treatment is unsafe, and includes <alternative_treatments> to suggest safer options. The <scope_limitation> tag defines the boundaries of applicability for findings or recommendations, including geographic, temporal, and population scopes.

[0135] Figure 11 illustrates tags suitable for representing numerical data and statistical findings. The <statistical_data> tag structures descriptive metrics, each with a <value>, <unit>, and optional statistical elements like <confidence_interval>, <sample_size>, and <p_value>. The <comparison> tag is used to show changes over time or between conditions, with <baseline> and <followup> values and a <change> section that quantifies differences using absolute and relative metrics, along with statistical significance.

[0136] Figure 12 illustrates tags suitable for issuing warnings, alerts, and general precautions. The <warning> tag conveys critical safety information, including a <title>, <content>, and timing-related elements such as <onset> and <monitoring_required>. The <alert> tag is used for urgent notifications like product recalls, detailing the affected <product>, the <issue>, and required <action>. The <precaution> tag offers general advisories, including applicability and monitoring recommendations, helping users interpret content with appropriate caution.

[0137] Figure 13 illustrates tags suitable for encoding conditional logic and temporal sequences. The <conditional_statement> tag models decision-making rules, with <condition>, <then>, and <else> elements, and support for nested <additional_conditions>. This is useful for adaptive protocols or clinical decision trees. The <temporal_sequence> tag structures events over time, with ordered <event> elements and <decision_point> logic based on outcomes. This format is ideal for representing treatment plans, workflows, or experimental timelines.

[0138] Figure 14 illustrates an outline of a first process of applying GCML to unstructured documents.

[0139] Firstly, at Step 1 the user maps non-optimised grounding content (i.e. untagged raw grounding data) and adds meta data to allow an orchestrator to optimise. This may involve a suitable user interface for applying the tags. At Step 2, the orchestrator then transforms the original content into GCML format (i.e. by applying the tags). At Step 3, a human-in-the-loop performs content validation or improvement, again using a suitable user interface. This optimised grounding source according to the schema (GCML) is stored at Step 4, ready for use at Step 5. The process may be carried out by the system described with respect to Figure 1, or any other suitable computer system. It will be appreciated that a wide variety of tagging tools exist for applying tags to documents using a suitable user interface, including tagtog (https: / / docs.tagtog.com / ) , Prodigy (https: / / prodi. y / ), INCEpTION (https: / / inception-project. ithub.io / ) and various XML editors, which are generally described in https: / / en.wikipedia.org / wiki / XML editor.

[0140] Figure 15 illustrates an outline of a second process of applying GCML to unstructured documents, in which an LLM is employed to generate the GCML. In a first step S1501, an input unstructured document received. In a second step, S1502, a prompt is generated for the LLM that comprises instructions for applying GCML and the content of the unstructured document. The instructions may comprise a definition of the GCML schema, as well as instructions which explain the purpose of tags in the schema and information about when they ought to be applied. In some examples, the instructions may comprise examples of documents to which GCML has been applied - i.e. "shots" in the paradigm of one or few shot learning. In step S1503, the prompt is provided to the LLM. In step S1504, the LLM responds with a GCML document based on the input document, which may then be stored, e.g. in repository 306. The process may be carried out by the system described with respect to Figure 1, or any other suitable computer system. Figure 16 illustrates a process of employing GCML documents in a RAG setting. In a first step S1601, an input is received. This may take the form of a query (e.g. a question) or instructions in natural language. In some examples, the instructions are provided by a user, for example via a suitable user interface. In other examples, the instructions are generated by a computer - for example forming part of a programmatically generated prompt intended for an LLM. Either way, the input is used in step S1602 to retrieve one or more relevant grounding documents from a document store, said documents in GCML format. RAG is well known, and so a detailed explanation is omitted for the sake of brevity and clarity, but it will be understood that suitable techniques for selecting relevant documents based on an input such as a query exist. In step S1603, a prompt is generated including the input and the retrieved document. In some examples, the prompt also includes a description of GCML and / or other suitable instructions that assist the LLM in interpreting the GCML. In step S1604, the prompt is input to the LLM, which provides a response in step S1605 that is based on the grounding data. The process may be carried out by the system described with respect to Figure 1, or any other suitable computer system.

[0141] In some examples, the process of Figure 16 is extended so as to enable the retrieval of only certain parts of the GCML document based on the input. For example, where the input is a query in relation to facts rather than opinions, the system may exclude parts of the document tagged as opinions.

[0142] A computer system comprises execution hardware which may be configured to execute the method steps disclosed herein. The term execution hardware encompasses any form / combination of hardware configured to execute the relevant method steps. The execution hardware may take the form of one or more processors, which may be programmable or non-programmable, or a combination of programmable and non-programmable hardware may be used. Examples of suitable programmable processors include general purpose processors based on an instruction set architecture, such as CPUs, GPUs / accelerator processors etc. Such general-purpose processors typically execute computer readable instructions held in memory coupled to or internal to the processor and carry out the relevant steps in accordance with those instructions. Other forms of programmable processors include field programmable gate arrays (FPGAs) having a circuit configuration programmable through circuit description code. Examples of non-programmable processors include application specific integrated circuits (ASICs). Code, instructions etc. may be stored as appropriate on transitory or non-transitory media (examples of the latter including solid state, magnetic and optical storage device(s) and the like). It will be appreciated that references herein to large language models (LLMs) are not limited to text- only models, but extend to multimodal models capable of processing and generating content across multiple modalities, including but not limited to text, image, audio, and video. In certain embodiments, grounding data— such as structured markup, metadata, or contextual annotations— may be incorporated into prompts used for generating or interpreting non-textual outputs. For example, GCML-tagged content may be used to guide the generation of images, video sequences, or audio narratives by providing semantic structure and contextual constraints. Furthermore, multimodal prompts may be employed, wherein textual grounding data is combined with other input types, such as spoken instructions, visual references, or environmental cues, to influence model behaviour. Such configurations are within the scope of the present disclosure and may be implemented using any suitable multimodal architecture or interface.

[0143] Although certain embodiments described herein illustrate the application of GCML tagging and related methods being performed on a user device, it will be understood that such methods may alternatively or additionally be executed on one or more remote servers or cloud-based systems. In such configurations, user interfaces for tagging, reviewing, or interacting with GCML-annotated content may be served to the user device via a web interface, application programming interface (API), or other networked communication means. Furthermore, the system architecture may be distributed such that the application of GCML tags (e.g., during content authoring or ingestion) is performed by a first system or service, while the inference-time utilization of GCML (e.g., by a language model or content delivery engine) is performed by a second, distinct system. These and other variations in deployment architecture are within the scope of the present disclosure.

[0144] It will be appreciated that the above embodiments have been disclosed by way of example only. Other variants or use cases may become apparent to a person skilled in the art once given the disclosure herein. The scope of the present disclosure is not limited by the above-described embodiments, but only by the accompanying claims.

Claims

Claims1. A computer system comprising: at least one processor; a memory storing at least one markup language document, the document comprising grounding data for use in grounding a generative artificial intelligence, Al, model, the markup language document defining one or more of: a plurality of sections of the grounding data; an importance level of each of the plurality of sections of the grounding data; reliability data indicating the reliability of the grounding data; and usage instructions indicating an intended use of the grounding data; the memory storing instructions which when executed by the processor cause the processor to: receive a first input; and generate a second input for a generative Al model including the first input and at least some of the markup language document to act as grounding data in providing a response to the first input.

2. The computer system of claim 1, wherein the at least one markup language document is selected from among a plurality of second markup language documents based on the first input, and wherein the processor is configured to retrieve one or more sections of the at least one markup language document based on the first input and to include the retrieved one or more sections in the second input.

3. The computer system of claim 1, wherein the processor is configured to include substantially all of the markup language document in the second input.

4. The computer system of any preceding claim, wherein the plurality of sections comprises headings and paragraphs.

5. The computer system of any preceding claim, wherein the reliability data includes a verification status indicating whether the grounding data has been verified.

6. The computer system of any preceding claim, wherein the usage instructions indicate that the grounding data is not suitable for particular uses.

7. The computer system of any preceding claim, wherein the markup language document defines one or more of: quotations and / or citations; subjective material in the grounding data; parts of the grounding data relating to instructions and / or procedures, comprising ordered steps; parts of the grounding data related to exclusion, limitations or contraindications in clinical or research contexts; and parts of the grounding data comprising numerical data and / or statistical findings.

8. The computer system of any preceding claim, wherein the markup language document is an XML-based document.

9. The computer system of any preceding claim, wherein the generative Al model is a large language model.

10. A computer system comprising at least one processor and memory, the memory storing instructions which when executed by the processor cause the processor to implement: a tagging tool configured to apply markup language to grounding data, wherein the markup language defines one or more of: a plurality of sections of the grounding data; an importance level of each of the plurality of sections of the grounding data; reliability data indicating the reliability of the grounding data; and usage instructions indicating an intended use of the grounding data.

11. The tagging tool of claim 10, comprising a user interface configured to receive user input to apply the markup language.

12. The tagging tool of claim 10, further configured to: generate an input for a generative Al model, the input comprising: a definition of the markup language, an untagged document, and instructions that cause the generative Al model to apply the markup language to the untagged document;provide the input to the generative Al model, and in response, receive a tagged document in the markup language; and store the tagged document.

13. A computer-implemented method, comprising: storing at least one markup language document, the document comprising grounding data for use in grounding a large language model, the markup language document defining one or more of: a plurality of sections of the grounding data; an importance level of each of the plurality of sections of the grounding data; reliability data indicating the reliability of the grounding data; and usage instructions indicating an intended use of the grounding data; the memory storing instructions which when executed by the processor cause the processor to: receiving a first input; and generate a second input for a generative Al model including the first input and at least some of the second markup language document to act as grounding data in providing a response to the first input.

14. A computer-implemented method comprising: using a tagging tool to apply markup language to grounding data, wherein the markup language defines one or more of: a plurality of sections of the grounding data; an importance level of each of the plurality of sections of the grounding data; reliability data indicating the reliability of the grounding data; and usage instructions indicating an intended use of the grounding data.

15. A computer-readable storage medium storing instructions, which when executed by a processor, cause the processor to carry out the method of claim 13 or 14.

16. A computer-readable storage medium storing the markup document of any of claims 1 to 9.