Use of large language model in source-to-pay platform
The method and platform improve Source-to-Pay systems by chunking data and determining user intent to enhance LLM functionality, addressing capacity and accuracy issues, and enabling flexible, user-friendly GenAI integration.
Patent Information
- Application Number
- PCT/EP2025/073279
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-14
- Filing Date
- 2025-08-13
- Publication Date
- 2026-02-19
AI Technical Summary
Existing Source-to-Pay platforms face challenges in deploying Generative Artificial Intelligence (GenAI) functionality due to limited processing capacity and unconstrained creativity, leading to inaccurate information generation, and require unique implementation for different enterprise maturity levels.
A computer-implemented method and platform that utilize a large language model (LLM) by dividing context data into manageable chunks, determining user intent, and selecting appropriate processing workflows to generate responses, while supporting configurability and flexibility for non-expert users.
Enhances user interaction and functionality, enabling efficient utilization of LLMs with improved accuracy and adaptability, reducing the need for extensive coding and accelerating adoption across diverse enterprise implementations.
Smart Images

Figure EP2025073279_19022026_PF_FP_ABST
Abstract
Description
E4327.WO PATENTUSE OF LARGE LANGUAGE MODEL IN SOURCE-TO-PAY PLATFORMBACKGROUND OF THE INVENTIONField of the Invention
[0001] The present application relates to improving functionality and user interaction in Source-to-Pay platforms using Large Language Models (LLM).Description of the Related Technology
[0002] Source-to-Pay platforms assist commodity managers, enterprise buyers, and account payable professionals to conduct market research, manage supplier relationships, author and negotiate contracts, search for products, compare prices, and electronically purchase and order products, and validate tax compliance and pay invoices. Such Source-to-Pay systems may include an electronic catalog with information from multiple suppliers. Source-to-Pay platforms also allow buyers to set up new suppliers and to automate various aspects of the associated product searching, purchasing, invoicing and payments functions. This requires Source-to-Pay platforms to manage and deploy large amounts of data whilst attempting to simplify these processes for buyers in order to prevent overwhelm and mistakes.
[0003] Generative Artificial Intelligence (GenAI) functionality may be implemented using Large Language Models (LLM). However, whilst offering many potential benefits, LLM suffers from a number of challenges which limit its applicability to Source-to-Pay platforms and services. In the context of sourcing and procurement processes where large amounts and disparate types of information must be considered, it can be challenging to deploy GenAI functionality. Another issue is unconstrained or unchecked “creativity” which can lead to the generation of false information. In addition, different enterprises have different levels of maturity in their sourcing and procurement processes leading to different and unique implementation requirements.SUMMARY
[0004] According to a first aspect, there is provided a computer-implemented method of using a large language model (LLM) in a Source-to-Pay platform. The method comprises receiving a user query associated with a context data, estimating a number of tokens associated with the context data and dividing the context data into a plurality of context data parts eachE4327.WO PATENT having an estimated number of tokens below a maximum number of tokens per transaction associated with the LLM, determining a user intent from the user query and selecting one of a plurality of predetermined processing workflows based on the user intent, processing the context data parts according to the selected workflow to generate one or more large language model (LLM) prompts each comprising a respective context data part, forwarding the one or more LLM prompts to the LLM and receiving a respective response, and using at least some content from the one or more responses to generate an answer to the user query.
[0005] This method improves configurability and simplifies user interaction, enabling non-expert users to navigate and more fully utilize a Source-to-Pay platform using LLM’s whilst managing technical implementation challenges such as limited processing capacity of LLM’s as well as their generation of inaccurate or misdirected information. The method also is compatible with a very configurable and flexible application environment where customers can extend the data model with new data fields, extend a user interface with new pages, new tabs, and new controls, and govern the overall application experience with custom data visibility, data processing and navigation logic. As such, this method allows for supporting configurable Generative Al use cases unique to each implementation which in turn may accelerate adoption and create maximum value for our customers without forcing an enterprise’s R&D team or software developers to become the bottleneck for coding and releasing new use cases.
[0006] According to a second aspect, there is provided a Source-to-Pay platform for use with a large language model (LLM), the platform having a processor and memory comprising processor readable instructions which when executed on the processor, cause the processor to: receive a user query from a user device, the user query associated with a context data, determine a user intent parameter using the user query, estimate a number of tokens for processing the context data in a transaction associated with the LLM, responsive to determining that the estimated number of tokens exceeds a maximum number of tokens per transaction associated with the LLM, (also known as the context window of the LLM) divide the context data into a plurality of context data parts, each context data part having an estimated number of tokens below the maximum number of tokens, select one of a plurality of predetermined processing workflows based on the user intent parameter, process the context data parts according to the selected processing workflow to generate one or more LLM prompts comprising respective context data parts, transmit the one or more LLM prompt to the LLM and receiving respective response data from the LLM, process at least some of the receivedE4327.WO PATENT response data to generate user response data, and transmit a user response message to the user device, the user response message based on the user response data.
[0007] Corresponding systems and computer program products are also provided.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Various features of the present disclosure will be apparent from the detailed description which follows, taken in conjunction with the accompanying drawings, which together illustrate, features of the present disclosure, and wherein:
[0009] Fig. 1 is a schematic diagram of a Source-to-Pay system utilizing Large Language Models (LLM), according to an example.
[0010] Fig. 2 is a schematic diagram of a Source-to-Pay platform for managing user queries and interacting with the LLM, according to an example.
[0011] Fig. 3 is a flowchart of a method of answering a user query in a Source-to-Pay platform using an LLM, according to an example.
[0012] Fig. 4a - 4b illustrate processes for managing data using different workflows, according to an example.
[0013] Fig. 5a - 5c are flowcharts illustrating different workflows, according to an example.
[0014] Fig. 6 is a schematic of part of a Source-to-Pay platform for answering user queries, according to an example.
[0015] Fig. 7 is a screenshot of a user interface, according to an example.DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS
[0016] Examples address some of the limitations of using GenAI functionality implemented using LLM in a Source-to-Pay platform. Examples allow for improving the functionality of the Source-to-Pay platform, improving user experience and enabling user interaction even with limited experience or knowledge of the platform. Further, examples may allow for user-specific configurability of the Source-to-Pay platform using no-code solutions.
[0017] Fig. 1 is a schematic diagram of a Source-to-Pay system 100 according to an example and which includes a Large Language Model (LLM) which is used to help respond to user queries. The Source-to-Pay system 100 comprises a Source-to-Pay platform 105, a LLM 110 and one or more users 115. The LLM 110 may be an external LLM such as Open Al ChatGPT™, Anthropic Claude™ or Google Genini™, and / or may be an internal LLM provided as part of the Source-to-Pay platform 105 such as Llama-3 open-sourced by FacebookE4327.WO PATENT or Mixtral open-sourced by Mistral. A specialized LLM agent (also referred to as a “processing agent”) may be part of the Source-to-Pay platform 105 and be configured to engage one or more other LLM agents or LLMs (external 110 or internal), possibly via an internal LLM, to co-ordinate the generation of a response to a user query. The users may be employees of a purchasing company that subscribes to or otherwise uses the Source-to-Pay platform to source and purchase products and services offered by suppliers that are available on the platform. In other examples, users may be external users such as suppliers responding to requests for proposals, redlining contracts, acknowledging an order, or sending an invoice.
[0018] As used herein, a “processing agent” or “LLM agent” refers to a software component within the Source-to-Pay platform 105 that coordinates and manages interactions with one or more LLMs to accomplish specific tasks. An LLM agent may implement one or more processing workflows and may be configured to select other LLM agents or appropriate LLMs, manage data flow between processing steps, and handle error recovery. The LLM agent may be considered a function that an associated LLM can call in order to complete a task.
[0019] A “processing workflow” or “workflow” refers to a predetermined sequence of operations and decision points that define how context data is processed to generate responses to specific types of user queries. Workflows may be implemented by agents and may involve sequential or parallel processing of data chunks using one or more LLMs.
[0020] In examples, the system 100 may support multi -LLM scenarios to optimize processing efficiency and response quality. Different LLMs 110 may be coordinated by one or more LLM agents, with each agent selected based on task type, with specialized agents implementing workflows that utilize specific LLMs assigned to specific intents such as translation, summarization, or contract analysis. A master agent, also referred to herein as an “orchestrator agent”, within the platform 105 (that is, within an orchestration layer of the platform (Figure 2, 231)) may coordinate multiple specialized agents, each responsible for different aspects of a complex workflow. A selection process performed by the system 100 may consider each LLM’s performance characteristics for specific tasks, directing more complex operations to more capable models while using faster, more efficient models for simpler tasks. The system 100 may implement context window optimization by routing larger chunks to LLMs with larger context windows, while smaller chunks or simpler tasks may be processed by smaller, faster LLMs. This adaptive selection mechanism ensures optimal processing efficiency while maintaining response quality. The selection mechanism may be implemented by the aforementioned orchestrator LLM agent of platform 105 that may be configured to identify and use an appropriate predefined sequence of LLMs to provide aE4327.WO PATENT response to a user query. In some examples, the predefined sequence may be an ordered list of multiple LLMs or other LLM agents to be used to generate a response to a user query. The predefined sequence may depend on the task that is requested within a user query. As an example, at a first step, the orchestrator LLM agent may use a reasoning LLM to determine the intent of a user query, next, the orchestrator LLM agent may use an extraction LLM to extract requested text from a large document of text based on the determined intent and a web agent to find another source of information relating to the extracted text, then, the orchestrator LLM agent may use an analytic LLM to analyse the extracted text and the web-retrieved information per the user query.
[0021] The system 100 may also implement robust fallback mechanisms when working with multiple LLMs. For example, if a primary LLM fails to produce a satisfactory response or encounters an error, the system can automatically route the request to an alternative LLM. Where there is an error accessing an LLM, for example if the LLM is unavailable due to a number of calls or tokens being above a threshold set by the provider of the LLM, a request may be resent to the LLM after a period of time. In addition, responses from a first LLM may be evaluated by another LLM (a “reviewer” LLM) and scored, where feedback is then provided to the first LLM on how to improve future responses. This review and assessment loop may occur several times to iteratively improve the first LLM. In addition, the platform 105 may also conduct A / B testing between different LLMs to continuously optimize performance based on real -world usage patterns. For organizations with varying data sensitivity requirements, the system 100 may support hybrid processing approaches, using internal LLMs for sensitive data processing while leveraging external LLMs for non-sensitive information. In some workflows, responses from multiple LLMs may be combined to generate more comprehensive answers, leveraging the unique strengths of each model.
[0022] The Source-to-Pay platform 105 comprises a processor 120 and memory 125 which may be configured to provide: a user interface 135 for interfacing with a user; an LLM Service to interface with an external LLM 110 and / or an internal LLM 145. The processor 120 and memory 125 may provide resources for implementing the internal LLM 145. The memory 125 may comprise Random Access Memory (RAM) and / or storage such as a hard disk or a solid-state storage device.
[0023] The user interface 135 may be a Graphical User Interface (GUI) which provides a structured way for a user to interact with the Source-to-Pay platform, for example using predetermined forms for completing. The GUI may provide a chat capability in which the user can interact with the Source-to-Pay platform using a series of questions and answers. The GUIE4327.WO PATENT may also be configurable to allow different users to add new functions. In some examples, an audio user interface may be provided.
[0024] The memory 125 comprises processor readable instructions 130 to answer user queries using the LLM 110 and / or 145. The instructions 130 include a sequence of instructions 160 - 170 which may be carried out by the processor 120. At 160, a user query is received which is associated with context data. The user query may be provided in natural language via a chat service and / or may be provided using structured data input using a form. The context data is data associated with the user query and may include for example a document referred to by the user or input by the user, for example by “cutting and pasting” part of a document, “pasting” the full document into chat history or linked to a document already stored in the Source-to-Pay platform 105. The context data may also include the chat history, results from related Internet searches, information from user documents and information collected from other services such as automated workflows. As such, the instructions 130 may cause the processor 120 to implement a Retrieval -Augmented Generation (RAG) process whereby the context data (as described above) is a secondary data source that is input to the LLM 110 and / or 145 to ensure the output generated by the LLM 110 / 145 is more accurate and highly relevant to the user.
[0025] At 162, an estimate of the number of tokens associated with the context data is determined. If above a limit for the LLM 110, 145 to be used, the context data is divided into a plurality of context data parts, each having an estimated number of tokens below a maximum number of tokens per LLM transaction. Tokens are the basic units of data processed by LLMs and may correspond approximately to a word, part of a word or even a character. The process of tokenization takes a user query and breaks that query into a sequence of tokens - this process may be different for different LLMs. An LLM transaction comprises a prompt (a question for the LLM to answer and which may include background data and / or the content of a document) as well as a response from the LLM. The total number of tokens involved in an LLM transaction includes the tokens in the prompt and in some cases the response. Each type of LLM will have its own maximum number of tokens per transaction. If the tokens in the prompt (and those estimated for a response) are above the limit for the LLM, then some action to address this must be taken or the LLM will simply return an error.
[0026] In this example, if the estimated number of tokens exceeds the token limit for the LLM to be queried (called the contextual window), then the context data is divided into a plurality of context data parts which will be further assessed and / or used in multiple promptsE4327.WO PATENT to the LLM in order to answer the user query whilst managing the LLM’s contextual window. The context data parts are also referred to herein as “chunks” or chunks of context data.
[0027] For token estimation, the system 100 employs specialized tokenization algorithms tailored to each supported LLM. While the system leverages tools like tiktoken for OpenAI models, it also implements custom tokenization logic for other LLMs such as Claude and Llama-3. For a given LLM, the token estimation algorithm accounts for both user prompt tokens, context data tokens, and expected response tokens, applying model-specific multipliers to estimate response length based on prompt characteristics. For example, summarization queries typically generate responses approximately 20% the size of the input, while translation maintains roughly equivalent token counts. The token estimation algorithm may also account for the number of tokens associated with one or more of the following, for a given LLM: system prompts, chat history, text content from relevant documents or from the chunks of context data. As a result, the algorithm may take into account the different context windows and response lengths for different LLMs to compute the maximum number of tokens left for input tokens.
[0028] The system 100 determines optimal chunk sizes through an adaptive algorithm that considers multiple factors: the specific LLM’s context window, the document type, and the intended operation. The chunking algorithm incorporates analysis of different document types, with customized chunk sizes and overlap configurations for contracts, invoices, questionnaires, and other document types. For example, contract documents require different chunking strategies than invoices due to their different structural characteristics. In one example, the adaptive algorithm is executed by an orchestration function 231 described in relation to Figure 2. In one example, a specific LLM’s context window (e.g. number of tokens) may be determined (e.g. 8000 tokens) and compared to the number of tokens in a document (e.g., 20,000 tokens). It is determined that chunking is required and at least three chunks are to be generated (since 8000 multiplied by 2 is 16,000 which is not large enough to cover document size). Next, the type of document may be considered. For a contract, semantic segmentation may be required so that the document is divided by clause or paragraph, depending on the size of the paragraphs it may be determined that 6000 tokens per chunk is appropriate and leaves room for a response by the LLM. The intended use may then be considered and used to determine a suitable overlapping of the chunks, for example, 600-1200 token overlaps. In this example, where 1200 tokens of overlap are used 4 chunks would be required for the document.
[0029] To maintain contextual coherence between chunks, the system 100, and specifically, the orchestration function 231 of Figure 2, implements an adaptive overlapping mechanism where the degree of overlap is dynamically determined, and may be determined inE4327.WO PATENT this way for each input document. For tasks requiring high contextual continuity like summarization, chunks may overlap by 10-20% of their content. For more isolated processing like translation, minimal overlap is used. This approach increases redundancy and connectivity between chunks, ensuring that information is not lost at chunk boundaries.
[0030] The system 100 employs sophisticated contract segmentation techniques to identify natural document divisions such as clauses, sections, and paragraphs and a hierarchical structure of the same. This hierarchical structure is stored in a graph representation that preserves the document’s organization. In some examples, the system 100 may ensure chunks are divided only at boundaries such as paragraph endings or clause boundaries rather than midsentence, which would disrupt semantic coherence. Other criteria could be topic boundaries to ensure each chunk relates to a single topic, named entity boundaries so that chunks are not split in the middle of names (e.g., company name), or time-based boundaries to ensure chunks contain a time based sequence (e.g. meeting minutes, project plans) rather than diving the sequence across multiple chunks. The graph-based representation enables efficient chunk retrieval during processing workflows and improves semantic search capabilities by maintaining relationships between chunks, allowing the system to find relevant sections within broader topics.
[0031] At 164, the user’s intent is determined from the user query and in some cases the context data. The user’s intent may be determined from a user selection such as button on the GUI, by identifying predetermined words in the user query, and / or by providing the user query to the LLM in a predetermined prompt asking for the user’s intent, this may be accompanied by some of the context data such as chat history. Determining user intent may be implemented by attempting to match the user query and / or context data with one or more of a plurality of predetermined user intent parameters, or else returning an error indicating that more input from the user may be required.
[0032] The system 100 employs sophisticated Natural Language Processing (NLP) techniques to determine user intent. Rather than simple keyword matching, the system 100 may use the LLM itself to extract intent by feeding the LLM with a taxonomy of expected intent types within the Source-to-Pay domain. The taxonomy may be fed to the LLM during inference of the model and possibly as part of an initial prompt that includes a user query. For instance, when processing a user query, the system 100, specifically, the orchestration function 231 of Figure 2, may generate an intent detection prompt for the LLM 110 that includes the user query, relevant context data such as chat history, and a structured schema of possible intent parameters.E4327.WO PATENT
[0033] The intent detection system (e.g., the orchestration function 231, the find intent function 613 of Figure 6 and the intent detection function 633 of figure 6) implements a function-calling approach where an LLM agent utilizing the LLM 110 may act as a decisionmaking engine. In such a scenario, the LLM agent associated with the system 105 and / or the LLM 110 (such as the aforementioned orchestrator LLM agent of the orchestration layer 231) may analyze the user input and identify which API or function should be called to fulfill the user’s request. For example, if a user asks to “summarize section 4 of a document,” the LLM agent may identify this as requiring two sequential operations: (1) extracting section 4 from the document, and (2) summarizing the extracted content and send appropriate prompts to one of more LLMs 110. This enables the system 100 to handle complex, multi-step intents by mapping them to appropriate processing workflows.
[0034] When user intent is ambiguous, the system 100 (for example, the orchestration function 231, intent detection function and the LLM 110) may implement a “Loop Until” fallback mechanism. This “Loop Until” process may be implemented by the orchestrator LLM agent associated with system 100 and involves the master LLM agent determining that mandatory information is missing and then engaging in a conversation with the user (for example, by initiating a chat agent) by dynamically generating clarifying questions to collect additional information until the missing information is obtained and sufficient clarity is achieved. The system 100 may maintain a schema defining what information is required for each intent type, categorized as mandatory or optional. In examples, user intent may be determined to be ambiguous when the system 100 cannot map the user query to a known intent type or when mandatory information is missing.
[0035] For example, if a user indicates they want to create a contract with Amazon starting on a specific date, the system 100 will extract as much information as possible from this initial prompt (contract type: NDA, party: Amazon, effective date: specified date). The system 100, via the orchestrator LLM agent, then identifies what mandatory information is still missing and generates questions to obtain this information. If multiple Amazon entities exist in the system 100, the agent will ask the user to select the correct one. This approach makes the platform accessible to users unfamiliar with the system’s architecture while ensuring all necessary information is captured.
[0036] The intent detection system improves over time through a combination of function calling, inference mechanisms, and action-taking capabilities. The system’s “Loop Until” mechanism has evolved to not only collect required information but also to take appropriate actions based on that information, creating a more autonomous and intelligentE4327.WO PATENT interaction model. This approach is particularly valuable in intake management scenarios, where the system 100 first determines the user’s intent, identifies what they want to accomplish (e.g., order products, sign a contract), and then guides them through the appropriate process with a schema of required information categorized as mandatory or optional.
[0037] After determining a user’s intent at 164, one of a plurality of predetermined processing workflows may then be selected by an LLM agent within the platform 105 (e.g., the orchestrator LLM agent) based on the determined user intent parameter. The agent implements the selected workflow by coordinating the necessary sequence of operations. Some examples of processing workflows include: translate document X, summarize document X, ask a question of document X such as identify risks in a contract, compare document X with document Y. Each workflow defines the specific steps and logic required to process the user query, while the orchestrator LLM agent executes these steps by managing LLM interactions, data flow, and response aggregation. The workflows may gather data, follow a pre-determine algorithm or procedure, leverage a specific internal or external tool (such as an internal library or an external search engine API), and prepare LLM specific prompts based on user intent and the user query. This may include populating a template prompt using the gathered data. The gathered data may include responses from previous prompts, for example those relating to other chunks of context data. The objective of the workflow (or method) is to ground the LLM with factual and relevant data, minimize LLM hallucination, and optimize the quality of the response to the original user request. The system 100, specifically the orchestrator LLM agent, employs a sophisticated decision tree algorithm to select the appropriate workflow based on the determined user intent parameter. Each workflow may be technically represented as a directed acyclic graph (DAG) of operations, with nodes representing discrete processing steps and edges representing the flow of data between steps. The workflow selection logic considers multiple factors including the user intent, document type, estimated processing complexity, and available system resources.
[0038] Workflows may be stored in a registry service that is a part of the platform 105 and may be integrated with the orchestration function 231 and a request builder. The registry service may maintain metadata about each workflow including its input requirements, output specifications, and performance characteristics. The registry may implement a versioning system that allows multiple versions of the same workflow to coexist, with rules for selecting the appropriate version based on context. New custom workflows can be registered through a standardized API that validates the workflow definition against a schema and ensures compatibility with the existing system architecture.E4327.WO PATENT
[0039] When a user intent is detected, the system 100 may map this intent to a specific endpoint signature (e.g., translate, summarize) which is then used by a request builder to construct the appropriate API call with all required parameters, as described above in relation to the intent detection system. For example, the / find intent endpoint returns the signature of the next endpoint to call, and from this signature, the request builder may build a request and send it to a web service with all necessary parameters (text, prompt, etc.). This architecture enables flexible workflow selection and execution while maintaining a clean separation between intent detection and workflow implementation.
[0040] The request builder (Figure 6, 639) may operate as a technical interface component that may be utilized by LLM agents to format and transmit requests to LLMs. While the herein described LLM agents handle high-level workflow coordination and decisionmaking, the request builder may focus specifically on the technical aspects of API call construction, parameter formatting, and request transmission. An LLM agent may determine what type of request is needed and provide the necessary data, while the request builder handles the technical implementation of formatting this into the appropriate API call structure for the target LLM.
[0041] At 166, the context data parts or chunks are processed according to the selected processing workflow to generate one or more prompts each comprising a respective chunk. This ensures that the LLM’s token processing limit per transaction is not exceeded. In a simple example, the chunks of a large document are added to respective prompts which ask the LLM to translate the content. For example, each prompt may include:“Translate the following content into French:[Chunk #]”Where # is the number of the chunk of the large document.
[0042] In a more complex example, the LLM may be prompted to summarize a first chunk, with the processing workflow receiving the summary response and adding this to the content of the next chunk and asking the LLM to summarize the combination of content. This is because individually and separately summarizing chunks may remove any context or association between them resulting in a series of disparate summaries which cannot be further combined. By including the summary of one chunk with the full content of the next chunk in the next prompt, the relationship between the two chunks is maintained and so the combined summary will have greater value and meaning to a user. Other types of workflows may also rely on mapping the response from one prompt to a next prompt. For example, a query relating to the expiration of a contract involve may prompts involving summarizing chunks to identifyE4327.WO PATENT a relevant contract term and then combining this summary with prompts for identifying an Effective Date in other chunks. Another example might include the combination of multiple processing workflows such as extracting a section from a contract about environmental protection, then translating that section in the language of the supplier country, then searching the internet to find environmental protection regulations in this country, and then checking any compliance issues between the translated section and the local regulation.
[0043] In yet another example, the number of chunks to be queried may be filtered depending on user intent. This may avoid misleading answers and / or reduce the amount of processing required by the LLM. For example, if a user query relates to the risks associated with a document, the chunks of the document may be semantically filtered to find a subset of relevant chunks, such as those referring to indemnification. A prompt may then be generated for each of the subset of chunks.
[0044] At 168, the prompts are forwarded to the LLM and respective responses received. Depending on the workflow, the prompts may be sent in a sequence that awaits the response from an earlier prompt.
[0045] At 170, the content of at least one response is used to answer the user query. The answer may be provided in the GUI and / or may be provided in a new document to which the user is referred. In the simple translation workflow noted above, the received responses are simply combined in chunk order. However, in the summary workflow, only the content from the final response may be included in the answer to the user’s query. Various other postprocessing steps may be employed depending on the selected workflow.
[0046] Fig. 2 is a schematic diagram of a Source-to-Pay system according to an example. The Source-to-Pay system 200 includes a Source-to-Pay platform 205 which may be used by a user 215 to answer a procurement related user query using an external LLM 210. The user query 222 will be associated with context data such as chat history or a document provided by the user 224 as well as data stored within the platform such as contract documents and supplier catalogues 233. If required, the context data may be divided into chunks 236 for use with LLM prompts.
[0047] The Source-to-Pay platform 205 also comprises an embeddings service 238 and a vector database 241. The embeddings service 238 generates embeddings for chunks 236 as needed. The embeddings may be vector representations of words or groups of words in a chunk 236 which may be generated by analyzing the words in order to encode them with a meaning represented by a vector. The vector database 241 may comprise a number of semantic vectors against which the embeddings of a chunk may be compared. This may be useful whenE4327.WO PATENT semantically filtering chunks 236 to identify only chunks relevant to the user query. For example, when the user query asks a question about a large document, the vectors associated with the chunks of that document may be compared with vectors in the vector database 241 which relate to the user query. In one example, if the user asks to identify part of a contract document associated with “termination”, the vectors of each chunk may be compared with a vector in the vector database which is associated with “termination”. Only relevant chunks may then be considered further. As the document may use a different word than “termination” to mean the same or similar concept, the use of vector-based searching improves the likelihood of finding all relevant parts of the document.
[0048] The embeddings service 238 implements sophisticated techniques for generating and storing vector representations of text. The system 100 may process words into chunks according to standardized chunk sizes, capturing semantic meaning through vector representations. Both document content and user questions may be processed into chunks and embedded using the same methodology to ensure consistent semantic comparison.
[0049] The system 100 may utilize embedding services provided by various LLM providers (such as OpenAI) and implement a hybrid search approach that combines multiple embedding models to improve search accuracy. This hybrid approach allows the system 100 to leverage the strengths of different embedding models for different types of content and queries. For example, some models are tuned or optimized for different types of input such as legal language, conversational language or general purpose text so a combination of these models would ensure broader semantic coverage. The system 100 may also implement validation techniques to verify search results, ensuring that retrieved chunks are truly relevant to the user’s query.
[0050] For semantic filtering, the system 100 may employ vector similarity algorithms that calculate the distance between vectors in high-dimensional space, for example, cosine similarity or Euclidean distance. When a user asks a question about a document (e.g., identifying contract terms related to “termination”), the system 100 may generate a semantic vector for the query term and compare it against vectors for all document chunks. The system 100 may use configurable similarity thresholds that vary based on the query type and document characteristics, filtering out chunks that fall below these thresholds.
[0051] The vector database 241 may be implemented as a specialized storage system optimized for high-dimensional vector operations. It may support efficient nearest-neighbor searches and maintain indices that accelerate similarity calculations. The integration between the embeddings service 238 and the vector database 241 may be implemented through aE4327.WO PATENT standardized API that handles vector storage, retrieval, and comparison operations, allowing the system 100 to efficiently identify semantically relevant content across large document collections.
[0052] The system 100 may employ a multi-stage approach for processing complex queries that involve multiple chunks. For document question-answering workflows, the system 100 may first identify relevant candidate chunks using vector similarity, process each chunk individually, and then aggregate the results using weighted relevance scoring. In some examples, a reranking algorithm may be used that ranks the most relevant candidate chunks and thereby refines the initial set of candidate chunks to be processed, which improves the overall quality of the generated output. The algorithm may re-evaluate, and re-order chunks identified as being relevant ensuring that the LLM receives the most pertinent and accurate information. This approach ensures comprehensive answers while prioritizing the most relevant information. For multi-step workflows (such as extracting contract terms, translating them, and checking compliance), the system 100 may maintain context between steps by passing relevant metadata and intermediate results through a workflow state management system, ensuring coherent end-to-end processing.
[0053] The Source-to-Pay platform 205 may also comprise a store of scripts for different use cases. For example, different scripts may be associated with different processing workflows depending on the intention of the user. For example, as noted above, a user query to summarize a document will require different processing than a query to translate a document. The Source-to-Pay platform 205 may be customized by adding new scripts corresponding to different processing workflows or additional functions such as guided user-interface workflows.
[0054] A prompt engineering function 245 is employed to generate and forward prompts to the LLM. In some cases, the prompts may be LLM-specific. The prompts will also depend on the processing workflow selected based on the user intent. The prompts may include chunks of context data together with predetermined or template questions. In some cases, a prompt may be used to determine user intent by including the user query and optionally some of the context data, such as chat history.
[0055] The prompt engineering function 245 may employ a library of specialized prompt templates optimized for different workflows and LLM providers. For translation workflows, prompts may include clear language-pair specifications and formatting instructions. Summarization workflows may use templates that specify desired summary length, focus areas, and style guidelines. Question-answering workflows may include context-E4327.WO PATENT setting instructions that improve the relevance and accuracy of responses. Each workflow may be considered a processing tool available to various LLM agents, for example, LLM agents that may be executed as part of the platform 105 or associated with external LLMs 110.
[0056] Prompts may be dynamically constructed based on multiple contextual factors. The system 100 may analyze the document structure, user intent, and previous interactions (e.g., chat history) to generate customized prompts. For example, when processing legal documents, prompts include specialized legal terminology instructions to ensure accurate interpretation. User-specific information such as company size, location, and industry may be incorporated to provide more relevant and personalized responses.
[0057] To optimize token usage, the system 100 may employ several techniques: instruction compression (using concise language in prompt instructions), dynamic exemplar selection (including only the most relevant examples based on the specific task), contextual pruning (removing irrelevant portions of context data before including in prompts), and template specialization (maintaining LLM-specific prompt templates that account for the unique characteristics of each LLM).
[0058] The system may handle LLM-specific API signature through a provider abstraction layer that translates generic generative Al requests into LLM provider-specific API format. This includes adjusting temperature settings, response format specifications and other parameters based on the specific LLM being used. For complex tasks, the system 100 may select more advanced LLMs (prioritizing quality over speed), while for simpler tasks, it may use faster, more efficient models.
[0059] A post processing function 247 is employed to select and / or combine the responses from the LLM into an answer 226 to the user query 222. This post processing will be dependent on the workflow selected. Where the response is the user intent, this may then be used to determine which processing workflow to select in order to process the chunks of context data associated with the user query.
[0060] The post-processing function (PPF) 247 may implement sophisticated algorithms for combining responses from multiple LLM calls. For translation workflows, the system 100, for example the PPF 247, may maintain chunk order and seamlessly concatenates responses, ensuring proper flow between translated sections. For summarization workflows, the system 100 may implement a recursive summarization approach that preserves key information across chunks by incorporating previous summaries into subsequent prompts. An index may capture the sequence of chunks in a document through 233, 238, 241. The indexE4327.WO PATENT may be used by both 231 (orchestration) to follow a recursive summarization approach (sequential processing) for summaries and by PPF 247 (post-processing) to preserve the order necessary to reassemble the translated chunks (concurrent processing via map-reduce).
[0061] The system 100 may include robust error handling and recovery mechanisms to address LLM failures or hallucinations. Each LLM response undergoes validation against taskspecific criteria. For factual responses, the system 100 may employ verification against trusted data sources. When inconsistencies or potential hallucinations are detected, the system 100 can regenerate prompts with additional constraints, try alternative LLMs, flag responses for human review, or fall back to more conservative response strategies that rely on verified information.
[0062] For workflows requiring structured output, the system 100 implements specialized post-processing algorithms that clean and format LLM responses, these algorithms may be executed via a “review and validation” agent that may be part of the orchestration layer 231. For example, an output parser can extract specific data points from free-text responses, normalize terminology, and convert responses into structured formats like JSON, SQL, HTML, or XML. This ensures downstream systems can reliably process the LLM outputs regardless of variations in response format or style. In addition, unnecessary SQL or HTML may be removed and / or JSON formatting changed to comply with the required specifications. For example, when a user requests SQL code to make a query, the output parser may remove extraneous text (such as “Here is your SQL:” prefixes) and ensure the returned code strictly complies with SQL syntax requirements, making it directly executable by database systems.
[0063] The platform comprises a controller or orchestration function 231 which manages the various aspects of the platform 205 in order to answer the user query 222 using the LLM 210. The orchestration function 231 determines whether the context data and estimated response is too large for an LLM transaction and determines how to divide the context data into smaller chunks. The orchestration function also determines the user intent from the user query (and in some cases some or all of the context data), which as noted previously may involve forwarding a predetermined prompt including the user query to the LLM (or a different specialized user intent determination LLM). The orchestration function 231 then controls how the chunks are processed and incorporated into prompts to be forwarded to the LLM, and how responses from the LLM are handled in order to provide the answer 226.
[0064] The orchestration function 231 may implement or coordinate with multiple LLM agents, each specialized for different types of workflows, or may implement the aforementioned orchestrator LLM agent that interacts with multiple other LLM agents. For example, a translation agent may be responsible for implementing translation workflows inE4327.WO PATENT association with one or more translation-focused LLMs, while a summarization agent handles summarization workflows in association with one or more summarization-focused LLMs. These LLM agents (including the orchestrator LLM agent) operate under the control of the orchestration function 231 and may share common services such as the embeddings service 238 and vector database 241. Each agent maintains its own logic for managing LLM interactions specific to its designated workflow types, while the orchestration function 231 provides overall coordination and resource management.
[0065] Fig. 3 is a flowchart of a method of using an LLM to answer a user query in a Source-to-Pay platform, such as the Source-to-Pay platforms of Fig. 1 or Fig. 2. The method 300 may be implemented using any suitable hardware such as a processor, memory, databases and user interfaces. At 305, the method comprises receiving a user query associated with context data. The user query may be provided as part of an interactive chat session, entered as part of a form or selected as a pre-configured function within a GUI. The user query may include or reference other data such as one or more documents and ask questions such as “summarize document X”, “compare documents X and Y”, find products that match criteria A, B, C.
[0066] The user query will be associated with context data which is used to generate prompts for the LLM. The context data may include the data included in or referred to in the user query such as specific documents. The context data may include chat history which may be useful in clarifying user intent. The context data may also include historical or category information about the user such as the marketplace they usually interact with. The context data may include Internet search results relating to the chat history, historical or category information.
[0067] At 310, the method detects whether there is a skills match. A skill as used herein is a preconfigured input workflow that assists a user with performing a predefined action, such as adding a new supplier. A user may not know how best to do this, and a skill can be defined as a workflow requesting specific types of information to enable better query answers and / or to simply add new data to the platform. A skill may differ from a processing workflow as it may involve a conversational interaction with a user to collect input necessary to execute a task. The method may detect a potential skills match by querying the user intent with the LLM using the original user query and context data. Alternatively, trigger words corresponding to a skill may be detected in the user query. If a skills match is detected, the method moves to 315 where the user is offered the skill. If accepted, the user is then guided through the associated input workflow to gather useful additional information, such as supplier address and productE4327.WO PATENT category. This may be implemented using a question-and-answer chat session or a form provided in a GUI. If there is no match, the method moves to 320.
[0068] At 320, the method estimates the number of tokens required for an LLM transaction involving the context data. Various tools are available to estimate token numbers, such as tiktoken at https : / / ithub . com / openai / tiktoken If the estimated number of tokens exceeds a per transaction token limit for the LLM, the method moves to 325. If the estimated number of tokens doesn’t exceed a per transaction token limit for the LLM, the method moves to 330.
[0069] The system 100 may employ various specialized token counting tools beyond tiktoken to support different LLMs. For OpenAI models (GPT-3.5, GPT-4), the system uses tiktoken, a Python library that supports different encoding schemes specific to various OpenAI models. For Anthropic’s Claude models, the system leverages Anthropic’s proprietary tokenizers available through their API, which provide token count estimates specific to Claude models. The system also incorporates HuggingFace Tokenizers, a general-purpose Python library supporting multiple tokenization algorithms and models, particularly useful for open- source LLMs.
[0070] For more complex document processing scenarios, the system may utilize LangChain’s TokenTextSplitter, which not only estimates tokens but also splits text accordingly based on configurable parameters for different LLM tokenization schemes. For models requiring subword tokenization, including some Google models, the system may employ Google’s SentencePiece, a C++ library with Python bindings that implements subword units like Byte-Pair-Encoding (BPE). This comprehensive approach to token counting ensures accurate estimation across the diverse range of LLMs supported by the platform.
[0071] At 325, the context data associated with the user query is divided into context data parts or chunks. For example, a large document may be divided into smaller sections each having less than a predetermined number of words for example. Each of these chunks may be included in separate LLM prompts and correspond with an estimate LLM transaction token size lower than the LLM transaction token limit. In order to improve operation of the method, documents may also be divided based on other considerations such as separating one chunk from another only at the end of a paragraph, section or page, rather than in the middle of a sentence for example. Chunks may overlap by some quantity of words to ensure contextual connectivity and relevance between them. A segmentation service may be used to segment a contract document into clauses and the document may then be divided into a plurality of chunks each having whole numbers of clauses.E4327.WO PATENT
[0072] At 330 the method determines the user intent, that is what the user query is likely to want to achieve. Examples include summarize a document or content pasted into the user query, translate, compare, and / or ask questions of a document. The intent may be determined using an appropriate predetermined prompt template and adding the user query and some of the context data, such as chat history and / or whether a skill has been used to complete the user query. Other approaches may parse the user query in order to identify key words. If a user intent cannot be identified, an error message may be returns asking the user to reformulate the user query. Determining the user intent may be similar to detecting a skills match in 310, but this is used for a different purpose.
[0073] This may be implemented by trying to match user intent with one or more of a plurality of predetermined user intent parameters. In an example, these user intent parameters may be provided in an LLM prompt together with the user query and asking whether the user query corresponds to one or more of the provided user intent parameters. For example, if the user query is to “Summarize section 4 of a document”, it could mean (1) extract section 4 of a document, and then (2) summarize the content of the extracted section of the document. In this case, the intent detection leads to the sequential execution of two agents or two processing workflows applying two different methods, one specialized in extracting text from a longer document and another one specialized in summarizing a piece of text. If this is not found, an error message may be generating asking the user for additional explanation or to reformulate the user query.
[0074] At 335, the method selects a processing workflow based on the determined user intent (e.g. the selected user intent parameter). For example, if the response from the LLM prompt is that the user intent is to summarize a document, the method may select a summarize processing workflow. A number of predetermined workflows may be available to select from, for example translate a document, summarize a document, ask a question of a document. Each workflow may be provided as a script to be called in order to carry out a sequence of actions such as generate a first LLM prompt with a predetermined question and a first chunk. A subsequent action may include using the response to generate a new LLM prompt including a next chunk. Some processing workflows may filter the chunks in order to only use a sub-set of these in the prompts. Some examples processing workflows are described in more detail below.
[0075] At 340, an LLM prompt is generated for a respective chunk according to the workflow. As previously noted, this may include a chunk, a predetermined question, as well as other information gathered about the user such as the size of their company, location, or other information that may be useful in answering the particular user query.E4327.WO PATENT
[0076] At 345, the method forwards the prompt to the LLM. This may be implemented using a webservice having pre-established access to the LLM. At 350, the method receives a response from the LLM.
[0077] At 355, the method determines whether there are more chunks to process and if so, returns to 340. If there are no more chunks, the method moves to 360.
[0078] At 360, the method generates an answer to the user query and provides this to the user, for example content displayed via a GUI or as a continuation of a chat session, or as a link to a newly generated document. How each of the LLM responses are used to generate the answer will depend on the processing workflow selected.
[0079] Fig. 4a illustrates a “translate” processing workflow in which a first chunk (chunkl) is sent to the LLM together with an instruction to translate to a specific language (Translation!). A second chunk (chunk2) is sent separately to the LLM together with the same instruction and a translation of that chunk is received (Translation2). Once translations of all chunks are received from the LLM, these translations (Translation!, Translation2) are simply combined in the order in which the chunks were divided, and this combination provided to the user as an answer to the user query.
[0080] Fig. 4b illustrates a “summarize” processing workflow in which a first chunk (chunkl) is sent to the LLM together with an instruction to summarize, with the LLM providing a summary of this first chunk (Summary 1). With the second, and any subsequent, chunks (chunk2), the LLM prompt includes the chunk itself (chunk2) together with the previously received summary (Summary!), as well as the instruction to summarize. This will return a summary (Summary2) of the combination of chunk2 and Summary 1. Similarly, if a third chunk (chunk3) is present, this would be included in a new prompt together with Summary2 and the instruction to summarize; and returning an updated summary (Summary3). The most recently received summary is then provided to the user as an answer to the user query.
[0081] The chunks are processed in different ways depending on the user intent. Translations of chunks can be handled independently and then simply combined as information in one chunk is not needed for translating another chunk. However independently summarizing separate chunks and then trying to combine them is unlikely to provide a summary of the whole document as the contents of one chunk or part of the document may be related to and / or influence the meaning of another chunk.
[0082] Methods corresponding to these workflows are illustrated in Fig. 5a and 5b. Fig. 5a illustrates a method of translating a document. The method 500 may be implemented as a script which is called when an appropriate user intent is detected. This script may be providedE4327.WO PATENT with the Source-to-Pay platform, or may be a custom developed script generated by a user and which may be related to a specific workflow routinely used by that user.
[0083] At 505, the method generates a first LLM prompt for a first chunk using a processing workflow template. The template may be a simple command such as “translate the following content into French”, after which the content of the chunk is inserted.
[0084] At 510, the method receives the translation from the LLM. At 515, the method checks if any more chunks need processing and if so returns to 505. Otherwise, at 520, the method combines the received translations in chunk order and presents the result to the user.
[0085] Fig. 5b illustrates a method of summarizing a document. The method 530 may be implemented as a script which is called when an appropriate user intent is detected. This script may be provided with the Source-to-Pay platform, or may be a custom developed script generated by a user and which may be related to a specific workflow routinely used by that user.
[0086] At 535, the method generates a first LLM prompt for a first chunk using a workflow template. The template may be a simple command such as “summarize the following content”, after which the content of the chunk is inserted.
[0087] At 540, the method receives the summary of the first chunk. At 545, the method generates a new prompt using the received summary and the next chunk. This may be implemented using a workflow template having the same simple command such as “summarize the following content” after which the content from the received summary and the next chunk are inserted.
[0088] At 550, the method receives a summary of the content of the previous summary combined with the next chunk. At 555, the method checks if any more chunks need processing and if so returns to 545. Here the content of the next chunk is combined with the content of the last received summary and the LLM is prompted to summarize the combined content.
[0089] At 560, once all chunks have been processed, the method simply provides the latest received summary as the output.
[0090] Fig. 5c illustrates a method of asking a question of a document. The method 570 may be implemented as a script which is called when an appropriate user intent is detected. This script may be provided with the Source-to-Pay platform, or may be a custom developed script generated by a user and which may be related to a specific workflow routinely used by that userE4327.WO PATENT
[0091] At 575, the method determines embeddings for each chunk. The embeddings may be at the chunk, sentence and / or word level. Embeddings may be used from many sources, such as OpenAI or other third-party sources as well as proprietary models.
[0092] At 580, the method performs a vector search on each chunk based on user intent. For example, if the user query asks to identify risks in a large contract document or documents, the vector search may filter the associated chunks so that only a sub-set of those chunks having meanings (vectors) associated with risk are further processed. The search may filter out chunks which do not have any vectors within a given distance of a vector associated with the word “risk”.
[0093] At 585, the method generates and forwards an LLM prompt using the subset of chunks. For example, a template instruction may ask “Identify [user intent] in the following content”, and the method then inserts the user intent or question such as “risk” into the instruction followed by the content of the chunk.
[0094] At 590, the method outputs the response received from the LLM. In some examples this may be implemented simply by providing directly the final output from the LLM response. In some other examples, different types of post-processing may be employed. For example, the LLM response may be cleaned up to comply with strict coding rules in HTML, JSON, SQL, or XML so that it can be executed programmatically by another process or system.
[0095] Fig. 6 illustrates a functional schematic of a Source-to-Pay platform according to an example. The functional blocks may be implemented in the Source-to-Pay platform of Fig. 1 and 2. The functions described in relation to Figure 6 may be implemented by specific LLM agents that are controlled by the orchestration layer 231 (and the orchestrator agent).
[0096] The Source-to-Pay platform 600 comprises a document extraction function 603 which extracts text from any document imported into the platform, for example using Optical Character Recognition (OCR). The extracted text may be stored as data within the platform along with other data useful for operating the platform. The collective data of the platform is represented by the platform state 609 and may additionally include chat history for different users of the platform and well as data generated by the platform, such as document chunks and embeddings.
[0097] A chat endpoint 606 is used to check whether the text of the document is too long to be included in an LLM prompt. This may be implemented by estimating the number of tokens in an LLM transaction involving the document; for example, using Tiktoken. An intent detection function 633 determines a user intent in received user queries or prompts. This function transforms the data in the user query into one or a number of predetermined user intentE4327.WO PATENT parameters, and associating these two data in the state 609. This may be implemented in an LLM-based find intent function 613 which uses a predetermined processing workflow which generates an LLM prompt using a template prompt with the user query text inserted. This may ask the LLM to determine whether the user intent can be allocated to one or more of a plurality of predetermined user intent parameters. Examples include “translate document”, summarize document”, “ask document a question”, “add new supplier”, etc. The workflow may then return the user intent parameter 636 based on the response from the LLM to the prompt, or indicate an error if the LLM is unable to allocate the user intent to one of the predetermined parameters.
[0098] Depending on the user intent, the platform may use a request builder 639 to generate an appropriate LLM prompt, or perform some other action such as run a script to facilitate user entry of additional information such as enabling the set-up of a new supplier by obtaining all of the required data from the user, either using a predetermined form or as a guided chat sequence. In an example, the / find intent endpoint sends back the signature of the next endpoint to call (for example translate or summarize). From this signature the request builder builds a request and sends this back to a Web service with all the parameters (e.g., text, prompt...)
[0099] The request builder may utilize one of a number of scripts, workflows or templates 616, 619, 623, 626, 629 to generate one or more prompts for the LLM. Summarize and translate workflows have been previously described. In an example, the endpoint / document / ingest extracts text, splits by chunk and converts to embeddings, then sends back these chunks with embeddings to the Buyer (user interface). The endpoint Document / chat by chunk asks the same questions for all the chunks. For example, if the user asks for an obligation — > the endpoint ensures that this question is asked for all the chunks before summarizing all the results
[0100] The LLM prompts are forwarded to an LLM, and responses received. The workflows may select one of the responses or combine all of the responses to generate an answer to the user query.
[0101] An output parser 643 may be used to interpret one or more of the LLM responses. In an example, the output parser may be able to parse the response to ensure it complies strictly to a predetermined coding format such as JSON, SQL, HTML, or XML. This improves handling of LLM responses which sometimes contain extra-text, for example if the user requests a SQL to make a query, the LLM response can respond with ‘Here is your SQL + then SQL’ . The output parser may implement logic to provide a strict code output that can be readily executed in a programmatic manner by another process..E4327.WO PATENT
[0102] The platform 600 may use an external embeddings service to generate semantic vectors for some of the data in the state 609, for example document chunks. This embeddings information may be used by some of the workflows 619 - 629 to help generate LLM prompts and / or to filter the data to be included in the prompts. In an example, document chunks may be filtered into a sub-set based on the detected user intent 636 and only sub-set of chunks used to generate LLM prompts. If the user asks a question of the document such as identify contract terms related to termination, then embeddings may be used to identify only those contract chunks relating to termination, for example by generating a semantic vector for “termination” and searching for any vectors associated with document chunks that are within a predetermined distance of the termination vector.
[0103] Fig. 7 illustrates a screenshot illustrating an example GUI which may be used by a configurator to add a skill or user-configured function for a Source-to-Pay platform. Once defined, the user-configured function may simply be added to a screen of the GUI as a enduser actionable button, without the need for any recoding of the platform, or coding to implement the user-configured function. This type of additional functionality may be implemented, for example, in any of the Source-to-Pay platforms of Fig. 1, 2 or 6.
[0104] In order to add a skill, in an example the configurator adds a description and keywords or other triggers. These may be used to match and suggest the skill (once configured) to an end-user based on user inputs to the Source-to-Pay platform. The configurator then adds a sequence of steps for the skill to execute when selected. The steps may be chosen from a predetermined set with the ability to add content, such as display a message and ask a question.
[0105] A skill configuration system may be built on a technical architecture that integrates LLMs directly into the Source-to-Pay platform 105. Such a system may allow configurators to attach LLM requests to specific UI element events (such as a click on a button or a page refresh) with custom LLM scripts. These scripts may include specialized prompts (e.g., “you are an expert lawyer, summarize what you are sent to...”), output format specifications (e.g., “your output should look like [THIS]”), and user input prompts. The result may be displayed in a specific part of the UI (an output field specified in the script), creating a seamless and configurable integration between the LLM and the platform interface.
[0106] The skill configuration system supports sophisticated output handling mechanisms. Initially designed to write output to a single field (e.g., “IVA response”), the system has evolved to support structured JSON outputs that can populate multiple fields across the UI with a single LLM interaction. This enables one-click actions that update multiple related fields simultaneously, significantly improving user efficiency. A fallback mechanismE4327.WO PATENT ensures that the value for any field control that could not be found on the page (either because they don’t exist, or they were deleted, or they are hidden given the user’s role and access controls) can be displayed in a dynamic chat panel.
[0107] User-defined skills are represented in the system as sequences of configurable steps stored in a database. Each skill has a unique identifier (GUID), descriptive label, detailed description, and a set of trigger words that activate the skill when detected in user input. A skill execution engine within the platform 105, specifically in a user interface and orchestration layer, may process these steps sequentially, with each step having a specific type (e.g., display message, ask question, execute LLM call, retrieve database information) and associated configuration parameters. Question steps can be configured to accept various input types (text, date, file upload, selection from a list) and can be marked as mandatory or optional.
[0108] Skills are triggered through three primary mechanisms: user explicit invocation (when a user navigates to a specific page or clicks a button); chat system invocation (when the chat system triggers a skill based on a user’s chat messages) or implicit invocation (when the system 100 detects relevant trigger words in natural language input). The skill execution engine may maintain session state in memory, storing variables and user inputs throughout the conversation flow. This allows the system to automatically determine which skill to launch based on previously provided information, creating a more natural and contextual interaction experience. In relation to the chat system mechanism, user intent may be inferred by a chat assistant (e.g., a chat bot) that is configured to identify and launch a skill to address a user’s request (e.g., to launch a user support assistant that may be able to provide answers relating to a Help Center of the platform 105). The chat assistant may first identify that a skill is clearly indicated in a message or prompt from the user. In launching a new skill, the chat assistant may provide context data associated with and previously provided by the user to the new skill. This may be completed through the launching of a function calling mechanism that maps the user prompt to the various user-input variables defined in the new skill so that the user does not need to provide the information again. In this way, a new skill is launched on demand in response to user intent rather than relying on a user to navigate a skills menu and make a selection themselves.
[0109] The skill execution engine may implement a “Loop Until” capability that enables agentic behavior within skills. When provided with an input schema (a set of required information), the system 100 can dynamically assess what information is missing, generate appropriate questions to collect this information, and continue the conversation until all required data is obtained. For example, when creating a contract, the system 100 can extractE4327.WO PATENT information from uploaded documents, identify missing fields, and generate questions to complete the necessary information. This approach enables complex workflows like checking for duplicate contracts or validating supplier information without requiring explicit scripting of every possible interaction path.
[0110] In an example, a configured skill may be used to receive a constrained user input (e.g. using forms and / or chat questions / answers) and from this an optimal LLM prompt is generated. The LLM response may then be further processed by the skill and returned to the user. In another example, a platform configurator or administrator may wish to add a new function “add supplier” which guides non-expert end-users of the platform through a defined workflow for adding a supplier to ensure that all appropriate information is captured and to simplify the process for the non-expert end-user. The configured function or skill defines a series of steps for a user to enter information and may include other features such as using a template to define a LLM prompt which includes specific types of information in a specific order with a request to provide the answer in a useful format.
[0111] Example skills code is provided below:"skills": [{"skill_guid": "EE49D548-B680-ED11-B8E2-00155DB48D0E","skill_label": "Create Contract","skill_desc": "This skill allows a non casual user to be guided to create a new contract with a new document from a template or from a uploaded attachment.","trigger_words": ["create contract","contract","agreement","legal"]},{"skill_guid": "F049D548-B680-ED11-B8E2-00155DB48D0E","skilljabel": "Training demo",E4327.WO PATENT"skill_desc": "Demo for the various skill step types available.","trigger_words": ["demo type"]},{"skiljguid": "F149D548-B680-ED11-B8E2-00155DB48D0E","skilljabel": "Guided Buying","skill_desc": "This is very simple intake form to guide a user to the Shopping page or a new Purchase Request.","trigger_words": ["buy"," items"," good"," item"," services"]},{"skiljguid": "F349D548-B680-ED11-B8E2-00155DB48D0E","skilljabel": "Add Supplier","skill_desc": "IVA can help you create a new supplier that will not be a duplicate of an existing one.","trigger_words": ["supplier"," vendor"," distributor"," merchant", contractor"]E4327.WO PATENT},{"skill_guid": "6A50E383-1 F07-EE11-B8E8-00155DB48D0E","skill_label": "Legal Assistant","skill_desc": "As a legal assistant, I can complete or generate clause language for you.","trigger_words": ["clause","completion","complete","legal","agreement","contract"]},{"skill_guid": "2E6611 B3-036E-EE11-B8EE-00155DB48D0E","skilljabel": "Category Monitor","skill_desc": "I can search the internet to build market intelligence for a particular commodity, including price fluctuations, supply disruption, regulatory changes, compliance risk. This market intelligence should be useful for a category manager to develop appropriate risk mitigation and action plans.","trigger_words": ["market intelligence"," commodity strategy"," price index"," supply disruption events"]}]
[0112] The ability for a configurator or an administrator to configure their own bespoke functions, without writing code, improves the usability and functionality of the underlyingE4327.WO PATENT platform, making it better suited to individual end-user needs and organizations depending on their own implementation objectives, unique processes, and maturity. The added configuration ability may be implemented without the need for any coding by the configurator or the administrator.
[0113] At least some aspects of the embodiments described herein with reference to Figures 1-7 comprise computer processes performed in processing systems or processors. However, in some examples, the invention also extends to computer programs, particularly computer programs on or in a carrier, adapted for putting the invention into practice. The program may be in the form of non-transitory source code, object code, a code intermediate source and object code such as in partially compiled form, or in any other non-transitory form suitable for use in the implementation of processes according to the invention. The carrier may be any entity or device capable of carrying the program. For example, the carrier may comprise a storage medium, such as a solid-state drive (SSD) or other semiconductor-based RAM; a ROM, for example a CD ROM or a semiconductor ROM; a magnetic recording medium, for example a floppy disk or hard disk; optical memory devices in general; etc.
[0114] In the preceding description, for purposes of explanation, numerous specific details of certain examples are set forth. Reference in the specification to "an example" or similar language means that a particular feature, structure, or characteristic described in connection with the example is included in at least that one example, but not necessarily in other examples.
[0115] The above examples are to be understood as illustrative. It is to be understood that any feature described in relation to any one example may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the examples, or any combination of any other of the examples. Furthermore, equivalents and modifications not described above may also be employed.
Claims
E4327.WO PATENTCLAIMS1. A computer-implemented method of using a large language model (LLM) in a Source-to-Pay platform; the method comprising: receiving, by a processor of a server computer, a user query from a user device, the user query associated with a context data; determining, by the processor of the server computer, a user intent parameter using the user query; estimating, by the processor of the server computer, a number of tokens for processing the context data in a transaction associated with the LLM; responsive to the processor of the server computer determining that the estimated number of tokens exceeds a maximum number of tokens per transaction associated with the LLM, the processor of the server computer dividing the context data into a plurality of context data parts, each context data part having an estimated number of tokens below the maximum number of tokens; selecting, by the processor of the server computer, one or more of a plurality of predetermined processing workflows based on the user intent parameter; processing, by the processor of the server computer, the context data parts according to the selected processing workflow to generate one or more LLM prompts comprising respective context data parts; transmitting, by the processor of the server computer, the one or more LLM prompt to the LLM and receiving respective response data from the LLM; processing, by the processor of the server computer, at least some pf the received response data to generate user response data; transmitting, by the processor of the server computer, a user response message to the user device, the user response message based on the user response data.
2. The method of claim 1, wherein the estimating a number of tokens associated with the context data comprises estimating the number of tokens for the LLM responding to the user query.E4327.WO PATENT3. The method of claim 1 or claim 2, wherein the selected processing workflow causes the processor of the server computer to process the response data associated with one LLM prompt and adds the processed response data to another LLM prompt.
4. The method of any preceding claim, wherein the selected processing workflow causes the processor of the server computer to receive first response data for a first LLM prompt for a first context data part and to add at least a part of the first response data to a second LLM prompt comprising a second context data part.
5. The method of any preceding claim, wherein the processor of the server computer to: semantically filter the plurality of context data parts to identify a subset of context data parts depending on the user intent and the context data parts, and generate LLM prompts for the subset.
6. The method of any preceding claim, wherein the semantic filtering comprises determining embeddings for each context data part and performing a vector search on the embeddings of each context data parts to determine the subset of context data parts, the vector search based on the user query.
7. The method of any preceding claim, wherein the processor of the server computer to: generate a set of instructions based on a template and the selected processing workflow; add the set of instructions to the one or more LLM prompts.
8. The method of any preceding claim, wherein the processor of the server computer to: present to the user one or more questions and receive respective user responses; process the one or more user responses to generate additional user query data and adding the additional user query data at least one of the one or more LLM prompts.E4327.WO PATENT9. The method of any preceding claim, wherein the processor of the server computer to determine the user intent parameter: obtaining content data associated with the user; forwarding a user intent prompt to the LLM, the user intent prompt comprising the user query and at least some of the obtained content data; receiving response data from the LLM and using the response data to determine the user intent parameter.
10. The method of any preceding claim, wherein the processor of the server computer to generate the user query using a user input workflow to prompt a user to input a predetermined sequence of query data.
11. A Source-to-Pay platform for use with a large language model (LLM), the platform having a processor and memory comprising processor readable instructions which when executed on the processor, cause the processor to: receive a user query from a user device, the user query associated with a context data; determine a user intent parameter using the user query; estimate a number of tokens for processing the context data in a transaction associated with the LLM; responsive to determining that the estimated number of tokens exceeds a maximum number of tokens per transaction associated with the LLM, divide the context data into a plurality of context data parts, each context data part having an estimated number of tokens below the maximum number of tokens; select one or more of a plurality of predetermined processing workflows based on the user intent parameter; process the context data parts according to the selected processing workflow to generate one or more LLM prompts comprising respective context data parts; transmit the one or more LLM prompts to the LLM and receiving respective response data from the LLM; process at least some of the received response data to generate user response data;E4327.WO PATENT transmit a user response message to the user device, the user response message based on the user response data.
12. The Source-to-Pay platform of claim 11, wherein the instructions, when executed on the processor, cause the processor to process the response data associated with one LLM prompt and add the processed response data to another LLM prompt.
13. The Source-to-Pay platform of claim 12 or claim 11, wherein the instructions, when executed on the processor, cause the processor to receive first response data for a first LLM prompt for a first context data part and to add at least a part of the first response data to a second LLM prompt comprising a second context data part.
14. The Source-to-Pay platform of any of claims 11 to 13, wherein the instructions, when executed on the processor, cause the processor to: semantically filter the plurality of context data parts to identify a subset of context data parts depending on the user intent and the context data parts, and generate LLM prompts for the subset.
15. The Source-to-Pay platform of claim 14, wherein the semantic filtering comprises determining embeddings for each context data part and performing a vector search on the embeddings of each context data parts to determine the subset of context data parts, the vector search based on the user query.
16. The Source-to-Pay platform of any of claims 11 to 15, wherein the instructions, when executed on the processor, cause the processor to: generate a set of instructions based on a template and the selected processing workflow; and add the set of instructions to the one or more LLM prompts.
17. The Source-to-Pay platform of any of claims 11 to 16, wherein the instructions, when executed on the processor, cause the processor to: display one or more questions to the user and receive respective user responses;E4327.WO PATENT process the one or more user responses to generate additional user query data and adding the additional user query data at least one of the one or more LLM prompts.
18. The Source-to-Pay platform of any of claims 11 to 17, wherein the instructions, when executed on the processor, cause the processor to: obtain content data associated with the user; forward a user intent prompt to the LLM, the user intent prompt comprising the user query and at least some of the obtained content data; receive response data from the LLM and use the response data to determine the user intent parameter.
19. The Source-to-Pay platform of any of claims 11 to 18, wherein the instructions, when executed on the processor, cause the processor to generate the user query using a user input workflow to prompt a user to input a predetermined sequence of query data.
20. A non-transitory computer-readable medium storing a program for using a large language model (LLM) in a Source-to-Pay platform, the computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to: receive a user query from a user device, the user query associated with a context data; determine a user intent parameter using the user query; estimate a number of tokens for processing the context data in a transaction associated with the LLM; responsive to determining that the estimated number of tokens exceeds a maximum number of tokens per transaction associated with the LLM, divide the context data into a plurality of context data parts, each context data part having an estimated number of tokens below the maximum number of tokens; select one of a plurality of predetermined processing workflows based on the user intent parameter; process the context data parts according to the selected processing workflow to generate one or more LLM prompts comprising respective context data parts;E4327.WO PATENT transmit the one or more LLM prompt to the LLM and receiving respective response data from the LLM; process at least some pf the received response data to generate user response data; transmit a user response message to the user device, the user response message based on the user response data.
Citation Information
Patent Citations
Text reduction and analysis interface to a text generation modeling system
US11861320B1
Systems and methods for orchestration of parallel generative artificial intelligence pipelines
US12039263B1
Enterprise generative artificial intelligence architecture
US20240202225A1