Method and device for generating a response to a request using a large language model, wherein the sources used are controlled

The method addresses the lack of data source control in large language models by using vector comparisons and certificates to ensure secure, compliant data usage, enhancing response relevance and ecosystem control.

WO2026052543A1PCT designated stage Publication Date: 2026-03-12ORANGE SA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Current large language models lack control over the use of data sources, leading to indiscriminate data usage without compliance with security, confidentiality, or regulatory constraints, limiting their implementation in closed user ecosystems and preventing secure data searches.

Method used

A method and device for selecting data sources by comparing request vectors with data usage constraints, using certificates to authenticate data providers, and ensuring compatibility with client declarations, thereby controlling data usage and compliance with regulations.

Benefits of technology

Ensures secure and compliant data usage for large language models, allowing controlled data access within closed ecosystems and improving response relevance by selecting data sources that align with client requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025074741_12032026_PF_FP_ABST
    Figure EP2025074741_12032026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for generating a response to a request transmitted by a client agent (Agt) associated with a large language model (LLM), the method comprising: - receiving a request from the client agent (Agt), and further comprising a certificate associated with a declaration relating to the request, the request further comprising an identifier of a vector representative of a source datum, the source datum contributing to the formulation of the response, and additionally representative of a declaration of the client agent relating to the request; - obtaining, from at least one data provider (Fourn), data relating to the source datum associated with the received vector identifier; - transmitting a response to the request to the client agent (Agt), the response being formulated by means of the source datum.
Need to check novelty before this filing date? Find Prior Art

Description

Description Title: Method and device for generating a response to a query technical field

[0001] This disclosure falls within the domain of artificial intelligence based on a large language model and is specifically aimed at ensuring certification of the data used and control of the use of this data to respond to an intention or request from a client agent. Previous technique

[0002] Large language models (LLMs) are increasingly used in generative artificial intelligence techniques and aim in particular to produce a text response to an intention, which can be called either a question or a query.

[0003] One of the key features of a large language model is its ability to utilize a wide range of data sources, either learned or received with the query, and to generate a response to a query it hasn't previously encountered. The large language model can thus be trained on a vast array of source data. Furthermore, for example, when dealing with sensitive data requiring usage control or data relevant to a real-time function, data can be sent along with a request received from a client agent. This data is then used by the large language model to formulate a response, thereby minimizing the risk of hallucinations. It's also important to note that training a large language model requires significant resources and represents an increasingly substantial cost.

[0004] A large language model is therefore trained to respond to a wide variety of intentions using these diverse and numerous data sources.

[0005] Thus, in response to a request, the large language model uses indiscriminately the different data sources to which it has access to produce a response to the request received from a client agent.

[0006] According to the previous technique, there is no control over the use of different data sources. Thus, regardless of the request received, the large language model uses any data source, and consequently any data from that source, to construct and provide a response to the client agent that issued the request. Known techniques therefore do not allow for searching and using data sources that comply with certain constraints regarding security, disclosure, confidentiality, or any other statement associated with the request. Control over the use of data from a source used by a large model The language is not defined in the current specifications. As a result, data can be used and reused without any disclosure controls or guarantees of reliability. In particular, it is not possible, using known techniques, to restrict the use of data to situations where its context of use complies with specific conditions, such as those imposed by regulations or commercial agreements, thus limiting data use to certain actors or a particular context.

[0007] The techniques known in this way do not allow the implementation of a large language model for a closed group of users in an ecosystem (consortium, commercial domain, interest group, etc.) while controlling the participants, the data sources used, and the use of this data.

[0008] Current techniques do not allow for the automation or tracking of the use of external databases or unlearned documents during the training of the large language model to generate a response to an intention.

[0009] These techniques also do not allow for the secure search of data sources from different actors and / or the search for data sources that comply with regulations or contractual agreements. Summary of the invention

[0010] The purpose of this disclosure is to address all or part of the limitations of prior art solutions, including those described above, by proposing a solution that enables control over the use of data used to develop a response to a customer agent's intent.

[0011] To this end, a method is proposed for selecting at least one source data point made available by a data provider, said source data being adapted for developing a response to an intent of a client agent associated with a large language model, the method comprising: - Receipt from the client agent of a request and a certificate associated with a statement from the client agent regarding the request, - Calculation of a vector representing the received request and declaration, - A selection of at least one source data by means of a comparison between a vector representing the source data, used for the development of a response to the query, and also representing a constraint on the use of the source data, and the calculated vector.

[0012] The selection process allows for the selection of source data by choosing a data provider whose statements are compatible with those of an agent requesting a response to a query. This compatible source data can then contribute to The development of the response, for example by an LLM (Learning Management Tool), in response to a request or prompt. Source data corresponds to data (text, video, image, etc.) made available by a data provider and / or data owner for use in formalizing or contributing to the development of a response to a request or prompt from a client agent. Source data can thus be made available, that is, used or exploited to generate a response to a request or prompt. Source data can be external, notably found in a database external to the IT domain (company, operator offering, organization, etc.) in which the selection process is implemented, or internal data from a database internal to the IT domain. Selecting source data thus allows, by extension, the selection of a data provider whose statements are compatible with a request.The values ​​of vectors, calculated to represent an agent's request on the one hand, and of source data on the other, allow for the identification of a data provider offering source data whose characteristics are compatible with the request. The vector calculation includes the declaration relating to the request, and on the other hand, the vector representing the source data represents a constraint on the use of that source data. This embodiment provides an enriched value for the vector, and the comparison of vectors allows for the comparison of both the request and source data, as well as the declarations of the request and the data usage constraint.In particular, when the calculated vector and the representative vectors of data sources are trusted—that is, calculated by trusted connectors or servers—a simple comparison of the vectors is sufficient to select a data provider offering the source data. The values ​​of the vectors, calculated to represent an agent's request, and of the data sources, allow for the identification of a data provider whose statements associated with a given source data are compatible with the requirements specified in the request statements transmitted with the request. A certificate, such as a digital certificate, is an identifier or electronic document that can prove the authenticity of a user, terminal, or even a server. For identity verification, a certificate contains information such as organizational details (name and address), a public key with a corresponding private key used to encrypt the text from the certificate owner, and other information such as the name of the certification authority and a digital signature to prove the authenticity of the certificate issuer. The certificate is therefore associated with a declaration (context of the request, organizational environment, level of confidentiality required for the response to the request, search domain for the request, etc.). identity or information about the sender of the request...).

[0013] According to one aspect of the invention, the representative vector of the source data is previously obtained from a data provider making the source data available in accordance with the usage constraint.

[0014] Thus, representative vectors of source data can be obtained from data providers. This source data, along with its usage constraints, can be obtained, for example, from a data owner or source. This source data is only used if it aligns with the declarations (regulations, commercial agreements, etc.) of a client request, thereby protecting data dissemination and preventing the use of data from these providers if the declarations between the client request and the usage constraints of this source data are incompatible.

[0015] In another aspect, a certificate attesting to the identity and authenticity of the data provider is also obtained with the representative vector of the source data.

[0016] The security of the process can be improved not only by obtaining a representative vector of a source data and a constraint on the use of this source data but also by accompanying this vector with a certificate attesting to the authenticity of the data provider transmitting this vector, this certificate can then be transmitted to the LLM function generating a response to the request.

[0017] According to another aspect, the selection of at least one source data is carried out in accordance with a comparison operation implemented between the calculated vector and at least one vector included in a database comprising vector identifiers representative of the statements associated with source data.

[0018] A database, for example managed by a proxy server also called a broker, containing vector identifiers and possibly the values ​​of these vectors, can be advantageously used to compare the calculated vector of the request with the representative vectors of the source data. This database can thus be used as soon as a request is received from an agent. The comparison can be adapted to determine a greater or lesser number of source data points. This comparison can advantageously take into account the response or acknowledgment message received from client agents, which evaluates the LLM responses and thus assesses the different source data points used to generate those responses.

[0019] According to another aspect, a selection of at least one second source data point is made, this at least one second source data point being adapted to contribute to developing said response and requiring from the client agent a certificate associated with a declaration compatible with a constraint of using at least one second source data.

[0020] For example, by using a database containing vectors, it is possible to select an alternative or second source data source whose declarations or usage constraints do not precisely match the declarations of the received request, but which could potentially contribute more optimally or completely to generating a response to the request. However, using this second source data source, which is initially incompatible with the declaration associated with the received request, requires a different certificate, or an updated certificate, from the client agent. The agent can then be asked to issue this new or updated certificate, which includes a supplementary declaration compatible with the declaration or usage constraint of the second source data source. This can lead to the calculation of a vector associated with this request and an associated declaration, based on the associated certificate.This improves the response to the query while maintaining the correspondence of the statements between the query and the second source data.

[0021] According to another aspect, the selection process further includes sending to a large language model the received request, the certificate associated with a statement from the client agent relating to the request, and the identifier of the vector representing an owner of the selected source data and / or the data provider making the source data available.

[0022] The selection process, in one embodiment, involves requesting a Data Link Management (DLM) system, which is responsible for obtaining the data necessary to formulate a response to the query by contacting the data provider(s) determined to be compatible with the query. The certificate identifying the requesting agent can be used so that the LLM transmits the response directly or indirectly to the requesting agent. Given that several source datasets may be managed by different data owners, the vectors representing these source datasets can advantageously be based on certificates associated with these data owners and / or the data provider making them available to respond to a query whose declaration complies with the usage constraints of these source datasets.

[0023] Depending on a particular aspect, depending on the selection process, the large language model to which the request is transmitted is selected based on contractual or regulatory information present in the previously transmitted client agent's declaration and / or a characteristic relating to the previously obtained data provider.

[0024] Given that several LLMs can potentially be used, the LLM The person in charge of developing a response may be the one who corresponds to constraints, for example, of a regulatory or contractual type, relating to a client agent or a data provider.

[0025] According to another aspect of the selection process, the request received from the client agent also includes proof of the validity of the certificate associated with a statement from the client agent regarding the request.

[0026] Particularly in an open or unadministered context, the request may also advantageously include proof of the validity of the certificate associated with a statement from the client agent relating to the request, relating to the statement, this validity being able to then be checked by the entities in charge of selecting a data provider and / or responding to the request.

[0027] According to another more specific aspect, the selection process also includes a verification of the proof of validity of at least one certificate received by means of an identification information of a decryption algorithm present in the message received from the client agent.

[0028] The entity in charge of verifying the validity of the certificate(s) may use an identification information from a decryption algorithm present in the received request.

[0029] According to another aspect of the selection process, the comparison between the vector representing the source data, said source data being used for the development of a response to the query and also representing a constraint on the use of the source data, and the calculated vector is carried out in accordance with a filtering operation associated with the constraint on use and / or the statement of the client agent relating to the query.

[0030] Different vector comparison rules can be used depending, for example, on the constraints of using the source data and the declaration associated with the query. Thus, for example, for constraints related to a security context, a comparison operation may be stricter than for a commercial context. It is therefore advantageous to be able to adapt the comparison rule to the type of metadata associated with a source data point, for example.

[0031] The different aspects of the selection process that have just been described can be implemented independently of each other or in combination with each other.

[0032] The invention also relates to a device for selecting at least one source data point made available by a data provider, said source data point being suitable for developing a response to an intention of a client agent associated with a large language model, said selection device comprising a processor coupled to a memory in which program instructions are stored for execution by the processor for a process comprising: - Receipt from the client agent of a request and a certificate associated with a statement from the client agent regarding the request, - Calculation of a vector representing the received request and declaration, - A selection of at least one source data by means of a comparison between a vector representing the source data, used for the development of a response to the query, and also representing a constraint on the use of the source data, and the calculated vector.

[0033] This selection device is capable of implementing, in all its modes of realization, the selection process that has just been described.

[0034] The invention also relates to a system for selecting at least one source data item made available by a data provider, said source data being adapted for developing a response to an intention of a client agent associated with a large language model, said selection system comprising: A selection device such as the one described above, A client agent, adapted to communicate with the selection device and configured to issue the request and a certificate associated with a statement from the client agent relating to the request.

[0035] The invention also relates to a computer program product comprising a set of program code instructions which, when executed by at least one processor, configure said at least one processor to implement the method of selecting at least one source data according to any one of the implementation modes of this disclosure.

[0036] The invention also relates to a computer-readable recording medium on which is recorded a set of program code instructions which, when executed by at least one processor, configure said at least one processor to implement a method for selecting at least one source data according to any one of the implementation modes of this disclosure.

[0037] Furthermore, a method is proposed for generating a response to a request transmitted by a client agent associated with a large language model, the method comprising: - a receipt of a request from the client agent, and further including a certificate associated with a declaration relating to the request, said request further including an identifier of a vector representing source data, said source data contributing to the preparation of the response, and furthermore representative of a statement from the client agent regarding the request, - obtaining the source data associated with the received vector identifier from at least one data provider, - a transmission to the client agent of a response to the request, said response being developed using said source data.

[0038] According to one embodiment of the invention, an LLM can thus obtain an identifier of a vector associated with source data that can contribute to formulating a response to a request. This source data has been selected because of its compatibility with a declaration or metadata associated with a request issued by the agent. The LLM agent can then develop a response using source data whose usage and dissemination constraints are compatible with the declaration (specific to a service, a usage constraint, a regulation, or a context associated with the request) relating to the received request. The LLM can then transmit the response generated from the selected source data, which is compatible with the request, to the client agent.Data providers could thus have a guarantee regarding the use of the source data they receive and store, and only source data compatible with conditions corresponding to the requirements and constraints of the request is requested and subsequently used to generate a response. The received request includes a certificate associated with a declaration relating to the request, for example, in connection with a regulation or contractual agreement. This allows the LLM (Logical Lifecycle Management) system to ensure the compatibility of the source data selected and made available by a service provider with the client agent's declarations. The declaration may thus correspond, but is not limited to, a control over the use or disclosure of data and / or a security feature related to the request.

[0039] According to one aspect of the process, the response to the request sent to the client agent includes a certificate associated with the source data used to develop said response.

[0040] Transmitting one or more certificates associated with the source data used to generate the response allows the client agent to be assured of the authenticity of the received response and potentially to identify and authenticate the provider or even the owner of the source data. The transmitted certificate can also relate to declarations associated with the source data, assuring the client agent of the correspondence between the request issued and the source data selected to generate the response.

[0041] According to one aspect of the generation process, the request is received from of a mediation agent, the said request originating from the client agent having been certified by the mediation agent.

[0042] The request can advantageously originate from a mediation agent or broker, which may include a selection mechanism. This mechanism can manage both requests from client agents and proposals from service providers regarding source data, thereby selecting source data, and consequently data providers, that are compatible with the client agents' requests. This facilitates, improves the reliability of, and accelerates the LLM's task in developing a response. In this way, the mandated entity only manages vector identifiers without having access to the source data, thus limiting potential problems of unwanted disclosure of this source data.

[0043] According to another aspect of the generation process, the received request also includes a certificate associated with an identity of the client agent who issued the request.

[0044] The request received from the client agent advantageously includes a certificate associated with an identity of the client agent, thus allowing the large language model, on the one hand, to authenticate the client agent issuing the request.

[0045] According to a particular aspect of the generation process, the client agent to which the response to the request is transmitted is determined by means of the certificate associated with the identity of the client agent received.

[0046] The client agent to which the response to the request is transmitted is determined from the certificate associated with the identity of the client agent received, which allows on the one hand for the possibility that the request may not be received directly from the client agent and on the other hand for the identity of the client agent to be certified.

[0047] According to one aspect of the invention, the generation process further comprises obtaining from a second data provider a second source data, said second source data not corresponding to the declaration relating to the request received and requiring from the client agent a certificate associated with a declaration with the second source data managed by the second data provider.

[0048] The LLM may solicit a data provider whose source data can contribute to generating a more relevant response to the received query, corresponding to a query, but not corresponding to a statement, such as a disclosure or security or sharing constraint, obtained with the initially received query.

[0049] According to one aspect of the generation process, the received request also includes a certificate associated with the source data and / or the data provider and / or a source data owner.

[0050] The received request may, in one embodiment, include a certificate associated with the source data and / or the data provider, thus ensuring the authenticity of the source data and / or the entity that transmitted it and / or the owner of the source data, notably by verifying the certificate's validity. This information is also useful when the single vector cannot be fully utilized by the large language model to identify and guarantee the authenticity of the source data. The received certificate(s) thus enable the LLM to determine which data provider to contact to obtain the source data needed to generate the response to the request.

[0051] According to yet another aspect of the generation process, the response transmitted to the client agent also includes a certificate of the large language model.

[0052] The response sent to the customer agent may, in one embodiment, include a certificate attesting to the authenticity of the LLM that generated the response, which may be important particularly in cases where the LLM is specific to a regulatory or organizational context.

[0053] According to one aspect of the generation process, at least one of the certificates transmitted with the response also includes proof of validity of at least one of the certificates established by a server of the large language model.

[0054] The customer agent can thus ensure the validity of the certificates (from the data provider, the source data, etc.) received with the response and ensure that the data providers and the source data are properly authenticated and validated.

[0055] According to one aspect of the generation process, the verification of the validity of at least one certificate is carried out using a decryption algorithm determined from information present jointly with the proof of validity.

[0056] The decryption algorithms therefore do not need to be stored by the LLM, and thus the LLM can verify any certificate, including those of data providers or data sources that have never provided data before. The transmitted information can thus correspond to a public key used to verify validity by decrypting certificate validity information.

[0057] According to one aspect of the invention, the generation process further includes receiving from the client agent a validation message of the response received.

[0058] The validation message received from the client agent allows us to validate the data providers and source data that contributed to developing the response and, if necessary, to confirm that the provider and possibly the source data have control over the use of the source data in accordance with the requirements issued by the client agent.

[0059] According to one aspect of the invention, the generation process further includes a record of the operations performed for the generation of the response and / or including a record of the entities that contributed to this generation.

[0060] Recording operations (exchanges, interactions with other entities) and entities (data provider, source data, data owner, mediation entity, client agent)—for example, in a Distributed Hash Table (DHT), Distributed Ledger Technology (DLT), or Data Clearing House—allows for tracing information related to response generation and subsequently identifying, or even validating, an entity's participation in generating the response. This record can also be used for auditing and / or reuse in subsequent requests.

[0061] The different aspects of the generation process that have just been described can be implemented independently of each other or in combination with each other.

[0062] The invention further relates to a method for obtaining a response to a query, generated from source data, said source data contributing to the generation of the response, made available by a data provider, the method implemented by a client agent associated with a large language model comprising - the issuance of a request, which also includes a certificate associated with a statement from the client agent regarding the request, - a response to the request, said response being developed using the source data associated with a vector representing the source data and also representing a constraint on the use of the source data.

[0063] The invention also relates to a device for generating a response to a request transmitted by a client agent associated with a large language model, said selection device comprising a processor coupled to a memory in which program instructions intended to be executed by the processor are stored, comprising: - a receipt of a request from the client agent, and further including a certificate associated with a statement relating to the request, said request further including an identifier of a vector representing source data, said source data contributing to the preparation of the response, and further representing a statement from the client agent relating to the request, - obtaining the source data associated with the received vector identifier from at least one data provider, - a transmission to the client agent of a response to the request, said response being developed using said source data.

[0064] This generation device is capable of implementing, in all its embodiments, the generation process that has just been described.

[0065] The invention also relates to a device for obtaining a response to a query, generated from source data, said source data contributing to the generation of the response, made available by a data provider, said obtaining device comprising a processor coupled to a memory in which program instructions intended to be executed by the processor are stored, comprising - the issuance of a request including a certificate associated with a statement from the client agent relating to the request, - receiving a response to the request, said response being developed using the source data associated with a vector representing the source data and also representing a constraint on the use of the source data.

[0066] The invention also relates to a system for generating a response to a request transmitted by a client agent associated with a large language model, said system comprising: - A generation device as described above, - A retrieval device as described previously.

[0067] The invention also relates to a computer program product comprising a set of program code instructions which, when executed by at least one processor, configure said at least one processor to implement the method of generating a response according to any one of the implementation modes of this disclosure.

[0068] The invention also relates to a computer-readable recording medium on which is recorded a set of program code instructions which, when executed by at least one processor, configure said at least one processor to implement a method of generating a response according to any of the implementation modes of this disclosure.

[0069] The invention also relates to a computer program product comprising a set of program code instructions which, when executed by at least one processor, configure said at least one processor to implement the method of obtaining a response according to any one of the implementation modes of this disclosure.

[0070] The invention also relates to a computer-readable recording medium on which a set of program code instructions is recorded which, when executed by at least one processor, configure said at least one processor to implement a method for obtaining a response according to any of the implementation methods of this disclosure. Brief description of the drawings

[0071] The invention will be better understood upon reading the following description, given by way of non-limiting example, and made with reference to the figures which represent: [Fig. 1] a schematic representation of a communication architecture in which is implemented a process for selecting at least one data provider, a process for generating a response to a request and a process for obtaining a response to a request according to an embodiment. [Fig. 2] a diagram illustrating the main steps of a process for selecting at least one data provider, a process for generating a response to a query, and a process for obtaining a response to a query according to another embodiment, - [Fig. 3] an infrastructure comprising a generation device, a selection device and a transmission device according to an embodiment of the invention as shown in an example, [Fig. 4] a representation of a device for selecting source data according to another example, [Fig. 5] a representation of a device for generating a response to a query according to another example, [Fig. 6] a representation of a device for obtaining a response to a request according to another example.

[0072] In these figures, identical references from one figure to another designate identical or analogous elements. For clarity, the elements shown are not to scale unless otherwise indicated. Description of the implementation methods

[0073] More generally, it should be noted that the implementation and realization methods considered above have been described as non-limiting examples, and that other variants are therefore conceivable.

[0074] The following description presents an architecture consisting of a single large language model, a single client agent, a single data owner, and a single data provider, but the processes can be implemented interchangeably in an architecture comprising several of these entities.

[0075] We first refer to [Fig 1] which describes a schematic representation of a communication architecture in which is implemented a process for selecting at least one data provider, a process for generating a response to a request and a process for obtaining a response to a request according to an embodiment.

[0076] The architecture of [Fig. 1] includes a corresponding Agt agent with a large language model. The LLM Agt client agent can interact with a person or a computer entity, such as a robot, not shown in [Fig. 1]. This person or computer entity can pose a question to an artificial intelligence system, which the LLM Agt client agent will then clarify, specify, and, in this case, complete. To this end, the LLM client agent has a wallet that allows it to add one or more statements to a request transmitted by the Agt client agent. More specifically, the Agt client agent can transmit a certificate associated with the client agent's identity, attesting to the authenticity of the Agt client agent.This can also be contractual or regulatory information, for which the associated certificate is attached to a request submitted by the agent, attesting to the validity of their declaration. For example, the client agent Agt can indicate in a declaration that they belong to a domain (Human Resources, Accounting, Legal, etc.) and therefore require a response related to that domain or a related domain, and / or a response with a high level of security, and / or one specific to an entity or consortium of companies. Any type of declaration that can be used to filter and tailor the response to a request is possible. Naturally, the client agent can issue multiple declarations in their request. It should be noted that the client agent Agt also interacts with a Mandate entity (for agent), a data provider (Fourn), and a large language model (LLM), whose roles and functions will be described below.

[0077] The architecture of [Fig. 1] therefore includes a Mandate entity, also called a Broker. The Mandate entity acts as an intermediary between the various entities involved in the selection, generation, and retrieval processes described in the different embodiments of the invention. In particular, it performs the parameterization function (also called Embedding Function (EF1)) of the requests and declarations transmitted by the client agent Agt. This parameterization function consists of representing the request received from the client agent Agt and the associated declaration as a vector. The vector thus integrates the various declarations, such as contractual or regulatory information, transmitted by the client agent in a certified manner. The Mandate entity also manages a vector database, Vectl, containing the requests received from the client agent parameterized as vectors.This vector database also interfaces with a maintained Vect2 vector database. by the data provider Fourn. This interface allows the proxy server to match a request and a statement from a client agent Agt with source data, and a statement from a data provider with the statements of the source data owner, appropriate to respond to the request, according to statements compatible between the request from the client agent Agt and the source data, itself associated with an accreditation, i.e., a certified statement, made available by the data provider and the source data owner. The statement associated with the source data corresponds to a constraint on the use of the source data. The Vectl database applies matching rules defined within the ecosystem specific to the architecture of the [Fig.[1] The Mandate entity also includes a server Srv1, also called a connector, whose main function is to ensure interaction and communication with the other entities in [Fig. 1]. Servers named Srv2, Srv3, and Srv4 are found in the data provider Fourn, the data owner Prop, and the large language model LLM, respectively. The functions of these different servers are relatively comparable. In main, - They ensure control over data usage and sharing by applying usage policies, - They ensure transit to other servers or to an application container and manage messages between these servers and the container, - They verify the received attestations and present themselves to another server with the verifiable attestations concerning them involved in a transaction between a client agent and a large LLM language model, - they orchestrate application containers such as the EF function, a VECT vector database and a large LLM language model.

[0078] The servers Srv1, Srv2, Srv3, and Srv4 also record data exchange interactions between entities, and more specifically between the servers themselves. These records, or logs, allow for the identification of problems in case of malfunctions or conflicts, and even enable billing for data exchanges. The servers also optionally perform certificate verification for declarations of source data, applications, or requests.

[0079] The architecture includes a large Language Modeling (LLM) function, corresponding to a neural network trained on a large dataset and responsible for developing a response to a request issued by a client agent, in collaboration with the Mandate, Supplier, and Proprietor entities. As mentioned above, the LLM includes a Srv4 server, also known as connector, whose missions have been indicated above.

[0080] The data provider Fourn is responsible for maintaining a database of Vect2 vectors associated with one or more source data points, which may come from distinct data sources, retrieved and managed by a Prop data owner. The source data points can be text files, images, videos, or any type of data that can contribute to generating a response to a query. The Fourn provider therefore configures the source data points and the associated declarations for the data owner and the data provider using an EF2 function in the form of vectors, which it maintains in a Vect2 database. The Fourn provider also maintains a record of the data sources in a Stock2 file or database. The Stock2 database stores the data owners' source data points, which can be exchanged, subject to usage controls, by the data providers.A calculated vector thus represents the owner of the selected source data and / or the data provider making the source data available. This calculated vector will then be compared to a vector calculated from a query and a statement associated with that query, and an identifier of this vector, if it corresponds to a received query, can be transmitted to an LLM in order to retrieve the source data associated with that vector.

[0081] The data provider Fourn obtains the source data from the data owner(s) Prop, who maintains in a Certif file or database the certificates of the declarations associated with the source data, the information for which it stores in a Stocks storage space.

[0082] The architecture of [Fig. 1] also includes a DID (Decentralized Identification) entity. This DID entity is designed to verify, and thus prove and validate, the exchanged certificates used by the various entities to certify a submitted declaration.

[0083] The exchanges and interactions between these different entities, according to a mode of implementation, will now be described in accordance with the numerical indications of the exchanges or operations of the entities of [Fig. 1],

[0084] During step E10, one or more data sources from the Stocks storage space, containing source data potentially usable by the LLM large language model and associated declarations, are parameterized as vectors by the EF2 function, also known as the EF2 connector. These data sources are thus enriched with certified declarations, called accreditations, added by the Prop entity. These accreditations can be of various types and can, for example, correspond to security characteristics, Confidentiality, sharing, compliance with a technical or administrative rule, compliance with a commercial domain or contractual exchange area, or any type of information that qualifies data sources and restricts or allows the sharing of data from these sources. Thus, source data may be supplemented by a declaration indicating that it is financial data, HR data, data specific to an entity (Operator A, for example), Top Secret security level data, or any type of declaration that restricts the use of this source data to certain requests or queries. Using its Certif file, the Prop entity attests to the validity of the communicated data sources and accreditations by adding, via the Srv3 server, and transmitting the certificate(s) associated with the declarations of this source data to the Fourn entity.Proof of the validity of these certificates can also be sent to the Fourn entity.

[0085] During this E10 step, the Supplier entity, via the Srv2 server, verifies the validity of the received certificate(s), for example, by querying the DID entity. To this end, if a proof of validity is transmitted, the Supplier entity can also verify the validity of the proof, for example, by using a decryption algorithm determined from a credential present in the certificate(s). The Supplier entity stores the vector(s) parameterized by the EF2 parameterization function in the Vect2 database. A vector is therefore parameterized using a given source identifier and also based on the accreditation added by the Prop entity to the source data.

[0086] During step E20, the Vectl database aggregates the vectorized data from the various Fourn entities, and in particular from the Vect2 database of the Fourn entity. In [Fig. 1], only one Fourn entity is represented, but since the Mandat entity acts as an agent for the client agent Agt, the Vectl database is designed to aggregate vectors from different providers and, consequently, from a set of source data from multiple data sources. The Mandat entity thus synchronizes the vector databases from different data sources, supplemented by the certificates associated with the declarations linked to this source data, validated by the Srv2 server of the Fourn entity. One of the declarations concerns the identification information of the Fourn data provider (name, IP address, email address, FQDN, etc.).The Mandate entity only has access to the vectors, which represent the source data of the data sources and the accreditations linked to that source data. However, the Mandate entity does not have access to the source data itself, thus limiting the unwanted dissemination of source data that could potentially conflict with the accreditation requirements of that source data. The Mandate entity can also validate the certificates associated with the source data declarations, for example, by using the DID entity. and can also verify that the evidence, possibly transmitted, is valid, for example by using a decryption algorithm determined from a credential present in the certificate(s). It should be noted that steps 10 and 20 can be repeated several times so that the Mandate entity updates the various available data sources, which can be used to formulate a response to a request from a client agent.

[0087] During an E30 step, the Agt client agent transmits a request message to the Mandate entity, including an intent and one or more certificates associated with one or more statements from the Agt client agent. One of these statements relates to identifying information about the Agt client agent (name, IP address, email address, FQDN, etc.). The client agent can also add one or more other certificates from the CertifO certificate database, associated with other statements relating to security characteristics, compliance with a business and / or organizational environment, a service or authorization, or other characteristics that may influence the data sources used to formulate a response to the request.Thus, through the certificates associated with the statements, the Agt client agent can indicate that they belong to an organization, that they require a response regarding HR data, that they require financial data, or that they are accredited with a Top Secret security level. The source data, respectively HR, financial, or Top Secret, can then be transmitted to them depending on the statement submitted. The Agt client agent can therefore transmit several certificates associated with distinct statements. Alternatively, the Agt client agent may present, in addition to the certificate(s), proof of the validity of the certificate(s). Specifically, the Agt client agent may transmit proof of the validity of the certificate associated with the Agt client agent's identity and possibly a credential, including, for example, a public key, present in the certificate(s) that allows the proof of validity to be decrypted.Upon receipt of the request message, as well as the certificate(s), the Mandat mediation entity verifies the validity of the certificate(s), for example by requesting the DI D entity or by using verification tools in its possession, for example by ensuring that a key present in the certificate is still valid, and possibly verifies that the proof provided is valid, for example by using the decryption algorithm possibly transmitted by the Agt client agent.

[0088] During an E40 step, the Mandate entity, and more specifically the Srv1 function of the Mandate entity, validates the certificate (in the rest of the document we will refer to a certificate but it may be a plurality of certificates) associated with the request received from the client agent Agt.

[0089] During this step E40, the EF1 function of the Mandat entity calculates a The vector representing the request and the received certificate. The EF1 parameterization function thus allows the request and the received certificate to be represented as a vector loaded into the Mandat entity's vector database, Vectl. Having received the request from the client agent Agt and knowing of vectors associated with source data linked to accreditations, the Mandat agent performs a search in step E50 for data source vectors close to the request vector using its Vectl vector database.The Mandate entity can thus initially restrict the search for source data based on the accreditations of the source data and the certified declarations of the Agt client agent, thereby ensuring an initial sorting of the available source data. It can then perform a semantic comparison to select only the representative vectors of source data that can contribute to answering the request issued by the Agt client. For example, and without limiting this to this scenario, if an Agt client agent transmits a request requiring "operator A" information, only source data from data sources that can be used in the "operator A" context, according to their accreditation, can be used to generate a response to the request.The purpose of the selection carried out by the Mandat agent is to match the client agent's statements with the accreditations of the data sources while offering a suitable response to the request transmitted by the client and thus limit the hallucinations in response to the client's request.

[0090] During this E40 step, the Mandate entity identifies, through vector comparison, potentially source data that could contribute to developing a more relevant response to the Agt customer agent's request, but whose credentials do not correspond, or do not fully correspond, to the statements associated with the request. This identification can then be used to request a new or updated statement from the Agt customer agent, which must be transmitted to the Mandate entity and the LLM function to obtain a response developed from this source data, thereby improving the quality and relevance of the response to the request.If the Agt customer agent cannot provide a new statement for this source data, because he is not authorized or does not belong to the organization for which the source data can be made available, for example, he will not be able to obtain a response to the query based on this data.

[0091] During an E50 step, the Mandate entity transmits to the LLM language model function a request, originating from the client agent Agt, and further including a certificate associated with a statement from the client agent, said request further including an identifier of a vector associated with source data managed by at least one data provider. The request may include more than one vector, particularly if more than one source data point compatible with the declaration transmitted with the Agt client agent's request is identified. It should be noted that the LLM function requested by the Mandate entity, in cases where multiple LLM functions can be requested, is determined based on information present in a declaration from a Supplier entity and / or the Agt client agent. Thus, the selected LLM function may be determined based on regulatory and / or contractual constraints of the Supplier entity and / or the Agt client agent, these constraints being obtained from one or more declarations associated with the Agt client agent's request and / or source data transmitted by the Supplier entity.

[0092] During step E60, the LLM function obtains from the data provider Fourn source data associated with the received vector identifier and a certificate associated with at least one data provider. The data provider to be requested, as well as information about the source data to be requested, is obtained from information transmitted by the Mandate entity to the LLM function, specifically in one or more certificates issued by the Mandate entity. The LLM function may also transmit certificates associated with the declarations accompanying the request from the client agent Agt, certificates associated with the source data obtained from the Mandate entity (possibly obtained during step E50), and a certificate associated with the LLM function itself, transmitted by the Srv4 server. This allows the data provider Fourn to verify the identity and authenticity of the LLM function requesting the source data.The certificates associated with the Mandate entity and the client agent can also be transmitted to attest to their identities. It should be noted that, alternatively, certificates are not specifically required to attest to identity. This can occur when the EF functions are certified, and the calculated vectors are then considered sufficiently reliable to guarantee that the information retrieved from these vectors is sufficient to attest to the respective identities of the entities and functions. Thus, during this E60 step, the data provider Fourn can authenticate the various pieces of information whose certificates are transmitted, for example, via the Srv2 server or connector, and possibly by using the DID entity.

[0093] Once the received certificates are validated, the Provider entity transmits, during step E60, the source data identified by the vector identifier. Typically, if the LLM function has sent three vector identifiers, the Provider entity(ies), determined by the Mandate entity, transmits three source data items to the LLM function in response to the identifiers. The Provider entity also transmits the certificates associated with the data sources and possibly proofs of validity for this source data.

[0094] During step E70, the LLM function generates a response to the request message received from the client agent Agt, via the Mandate entity. The source data, corresponding for example to data chunks, is used by the LLM function to generate the most relevant response to the client agent's request, based on data whose credentials are compatible with the statements of the client agent Agt and possibly with the statements of the source data provider(s), Fourn.

[0095] During step E80, the LLM function sends the response generated during step E70 to the client agent Agt. This response may be accompanied by certificates associated with the source data used to generate the response and presented by the Fourn entity, the certificate(s) relating to the statements accompanying the request received from the client agent Agt, and also a certificate associated with the LLM function itself. This allows verification of the LLM function's authenticity and confirms that the response is indeed relevant to the request transmitted during step E30. The presence of certificates associated with the source data provides assurance regarding the data used to generate the response to the received request. However, in a simplified mode, the response can be sent to the client agent without the certificates. This can occur, for example, when the client agent Agt and the LLM interact based on a pre-existing trust relationship.The response may also include the certificate associated with the Mandate entity that performed the source data selection. Certificate verification is carried out by the Agt client agent, possibly using the DID entity service and a decryption algorithm that may be present with the certificate(s). It should be noted that the transmitted response can be sent directly to the Agt client agent, for example, using information identifying the Agt client agent present in the request message transmitted during step E30 and then step E50. Alternatively, as in one example, the response can be transmitted by the LLM function to the Agt client agent via the Mandate entity, allowing for a more efficient response to routing or even filtering constraints in the entities involved in the response generation process.The response transmitted during step E80, as an example, also includes information about source data not used to generate the transmitted response but which could be used to develop a more relevant response, provided the client agent (Agt) submits a statement consistent with the security clearance associated with the source data. The LLM function can thus transmit information about the required statement, such as "Operator B, Defense Confidential," or "Confidential HR Data," requiring a statement from the client agent consistent with this information. The client agent (Agt) can then obtain a response based on this data. additional source, and if it has the possibility or the right, transmits a new request and the missing statement to the LLM function, possibly via the Mandate entity, to obtain a response elaborated with this additional source data.

[0096] During step E90, the Agt customer agent, having received the response from the LLM function, analyzes it, particularly with regard to accreditation compliance and the semantics of the response to the request. Optionally, the agent then transmits feedback on the response to the Mandate entity and possibly the LLM function, for example, using a satisfaction percentage or appropriate labeling. This allows the Mandate entity and potentially the LLM function to evaluate the responses provided to requests, validate the information and statements submitted by the various entities involved in preparing the response, and improve the responses provided. The Agt customer agent thus contributes to improving the response generation process.

[0097] The various processing steps performed by the different entities during the generation of the response are potentially recorded in a file or set of files, such as a distributed hashed table (DHT), a permissioned or permissionless distributed ledger (DLT), or a data exchange center. This allows for subsequent use in case of problems, malfunctions, or billing needs related to actions performed by the different entities. This recording action can be carried out by the Srv1, Srv2, Srv3, and Srv4 servers of the entities contributing to the generation of the response, potentially requiring control over the usage associated with source data.

[0098] We then refer to [Fig. 2] which describes the main steps of a process for selecting at least one data provider, a process for generating a response to a query and a process for obtaining a response to a query according to another embodiment.

[0099] During an F10a Disp step, a Data Owner provides a Data Producer with source data from various data sources and presents certificates associated with declarations, corresponding to accreditations, linked to this source data. The source data can be text, images, videos, or a mix of these data types. An example of the information provided by the Data Owner is shown below (the terms in the example below are not in a specific language and represent a valid encoding syntax in any language, including French). { @context": [ "product id": "orange mobile service", "authorization": "secret" L "proof": { "type": "Ed25519Signature2018", "created": "2021-03-19T15:30:15Z", "jws": "eyJhb...JQdBw", "proofPurpose": "assertionMethod", "verificationMethod": "did:example:issuer#keys-l" } } L "id": "ebc6flc2", "holder": "did:example:holder", "proof": { "type": "Ed25519Signature2018", "created": "2021-03-19T15:30:15Z", "challenge": "n-0S6_WzA2Mj", "domain": "https: / / client.example.org / cb", "jws": "eyJhbG.JAoDA", "proofPurpose": "authentication", "verificationMethod": "did:example:holder#key-l" } }

[0100] For example, in this example, we can see that the metadata corresponding to accreditations specific to the source data are three in number and are as follows: credentialSubject": { "Company ID": "Orange", "product id": "orange mobile service", "authorization": "secret"}

[0101] The source data, for which the description is given above, is specific to an operator and more particularly to a mobile service and is of type "secret" indicating that it cannot be used for any type of intention, but only for requests for which a declaration is at least of type secret.

[0102] During an F10b Valid step, the data provider, specifically a server or connector such as the Srv2 server, verifies the certificates related to the declarations associated with the data source, received previously. It can use a DID entity, shown in [Fig. 1], for this purpose. If the certificates are valid, they are added to the description of a data source. In the example below, if the certificates associated with a data source 'metadata': {'accreditation' { 'Company ID': Orange, 'product id': 'orange mobile service' 'secret authorization': 'secret'}

[0103] During an F10c Vect step, the data provider, and more specifically the EF2 (Embedding Function), vectorizes the source data along with the declarations relating to the data source and saves the vector in a database such as the Vect2 vector database in [Fig. 1]. The vector therefore provides information not only about the source data but also about the property or properties associated with that source data. The declaration(s) allow for certified control of access to the vectorized source data. Since the source data can be quite large, for example, if it is a configuration file for a piece of equipment, the EF2 function vectorizes blocks (or chunks) of the source data. Thus, for a single source data item, several vectors can be parameterized.

[0104] A vector of a chunk of source data corresponds for example to id': '573387acd058e615000b5cb5en' 'vector': [0.1 , 0.8, 0.2, ..., 0.7], As shown in the example below, the vector identifier also includes declarations associated with the source and may include other information such as the version of the EF2 function that vectorized the data source or the type of the source data, as in the example below, a "manual" type for this source. { 'id': '573387acd058e615000b5cb5en', 'vector': [0.1, 0.8, 0.2, ..., 0.7], 'metadata': { 'accreditation' { 'company ID': Orange, 'product ID': 'Orange mobile service', 'secret authorization': 'secret', 'embedding model: 'embedding function version ID'}, 'source type': 'manual', 'connector ID'}

[0105] During an F20 Aggr step, the database, such as the Vectl database, of a broker entity aggregates the vectorized data from the various supplier entities, including the Vect2 database, which contains the vector representing the data source and its certified declarations, as described above. The broker entity only contains vector identifiers and not the source data itself, thus preventing a potential leak of this source data.

[0106] During an F30 Intent step, the principal entity receives a request from a client agent, such as the Agt agent. The request may include one or more statements associated with the request. These statements specify the business, organizational, contractual, regulatory, authorization, etc., context of the request and, consequently, of the source data to be used to generate the response. The principal entity verifies the certificates associated with the statements in the request and then vectorizes—that is, calculates a vector—during an F40 step. This vector includes the statements whose certificates have been validated, for example, by the EF1 function. An example of a vector of a request received by a principal entity and its certified statements is shown below: { ' query' 'intent vector' : [0.1, 0.8, 0.2, ... , 0.7] , 'metadata' : { 'accreditation' { 'company ID' : 'company certificate ID' , 'product ID' : 'product certificate ID' 'secret authorization' : 'secret certificate ID' 'embedding model version' : 'version certificate ID' 'company ID 2' : 'company certificate ID 2' , 'robot agent ID' : 'robot certificate ID'} 'product type' : 'manual'.

[0107] In the example, the declarations to the query are as follows: 'company ID': 'company certificate ID', 'product ID': 'product certificate ID' 'secret clearance': 'secret certificate id' 'embedding model version': 'version certi ficate id', 'company 2 id': 'company certi ficate id 2', 'robot agent id': 'robot certi ficate id'

[0108] During this F40 Search step, the principal entity, via the Vectl database, searches for vectors from data sources compatible with the received query. The search aims to find one or more vectors of a source data source close to the search vector(s) while ensuring that the accreditations associated with the query match the accreditations associated with a source data source. The result of a query identifies the k (this number can be defined by the search tool) vectors closest to the query. In the example below, 3 vectors are identified. The result of the search with the k=3 closest vectors in the vector database corresponds, for example, to: { ' results ' : [ { ' matches ' : [ { ' id ' : ' 5733a6424776f41900660 f51it ' , ' score ' : 106 . 119217, 'values': []}, {'id': '56e7906100c9c71500d772d7it', 'score': 117.300644, 'values': []}, {'id': '573387acd058e615000b5cb5en', 'score': 121. 347099 , ' values ​​': [ ]} ] , ' namespace ': ' '} ]}.

[0109] The score associated with a vector gives a level of correspondence between the data source and its statements on the one hand, and the query and its statements on the other.

[0110] The following operations (equal to, not equal to, greater than, greater than or equal to, less than, less than or equal to, in a list, not in a list) can be used to compare the metadata associated with the representative vectors of the request and the source data, and their respective declarations. The operation used may differ for different declarations. Thus, for a "secret authorization" declaration, there must be an exact match between the request declaration and the source data declaration, whereas, for example, for a company declaration, the operation might consist of searching in a list, knowing that for a request from mobile operator A, a list of mobile operators associated with a source data point might be suitable.

[0111] During an F50 Trans step, the principal entity transmits the search result, including the identifiers of the vectors associated with the source data resulting from the comparison, to an LLM function. The principal entity also transmits the previously received request from the client agent, along with a certificate associated with the client agent's declaration (or certificates associated with declarations if there are multiple declarations). The LLM function chosen by the principal entity is the one that best meets the constraints. Regulatory and contractual obligations of the source data provider(s) and the client agent identified in the declarations accompanying the received request. The LLM function can be selected by the Mandate entity because it is compatible with the certificate constraints of other entities.

[0112] During an F60 Obt step, using the obtained vector identifiers, the LLM function retrieves the source data corresponding to these vector identifiers from the data provider. The LLM function also presents the certificate associated with the client agent's declaration (or certificates if there are multiple certificates), and possibly the certificates associated with the source data, and possibly a certificate associated with the LLM function itself.

[0113] During an F70 Gen step, the LLM function generates a response to the query using the source data obtained during the F60 step and the query data obtained during the F30 step. The LLM function can analyze the source data(s), and perform pre-processing of this source data so that the generated response best matches the query received from the client agent.

[0114] Once the response is generated, the LLM function transmits the generated response to the client agent that issued the request during an F80 Emet step. To perform this transmission, the LLM function can advantageously use a client agent identifier declared with the received request. Along with the response, the LLM function transmits the certificate associated with the declaration accompanying the client agent's request, and possibly the certificate associated with the declaration associated with the data source, as well as a certificate associated with the LLM itself. This transmission preferably occurs between connectors or servers, such as the Srv4 server and a corresponding SrvO server of the client agent.In the event that during step E50, the Mandate entity transmitted a vector identifier of a source data whose declaration does not correspond to a declaration of the request but may contribute to a more relevant response, the LLM function may indicate during step E80 the missing declaration to the request which may allow the client agent to obtain a more relevant response and also corresponding to the new declaration to be provided by the client agent to the LLM function possibly via the mandate entity, by transmitting an updated request or a new request.

[0115] During an F90 Ack step, after verifying the validity of the received certificate(s) using the SrvO server and possibly the DID entity, the client agent validates the received response with the LLM function and possibly the proxy entity, allowing them to validate the data source selected to generate the response and thus validate the comparison operation between the vectors representing the request and the source data, these vectors also representing the respective declarations. associated with the intent and the source data. If the validation is not completely satisfactory, the principal entity can modify the comparison operation for one or more declarations in the parameterized vectors. The message transmitted during step E90 can optionally include the new declaration or the new modified query with the new declaration, as described above.

[0116] During an F 100 step, the various operations performed by the entities, and more particularly by the SrvO,... ,Srv4 servers are recorded in a distributed hash table (DHT) or a distributed ledger based on DLT technology for example (Distributed ledger technology (DLT)) or a data exchange center (Data Clearing House).

[0117] We then refer to [Fig 3] which presents an infrastructure comprising a generation device, a selection device and a transmission device according to an embodiment of the invention.

[0118] In this [Fig. 3], the entities Agt, LLM, Mandat, DID, Fourn, and Prop are represented as described in [Fig. 1]. These entities exchange messages, such as those described in [Fig. 1] and [Fig. 2], using a communication network Res. The Res network can correspond to any communication network that allows the transmission and reception of data, such as a public or private IP network.

[0119] The client agent entity Agt also includes a mechanism for obtaining a response to a request, Disp Obt, as described in [Fig. 6], implementing a method for obtaining a response to a request in all the described embodiments, including those in [Fig. 1] and [Fig. 2]. The agent entity Mandat includes a selection mechanism, Disp Sel, as described in [Fig. 4], implementing a selection method in all the described embodiments, including those in [Fig. 1] and [Fig. 2]. It should be noted that the selection mechanism, Disp Sel, in another embodiment, particularly where the Mandat entity and the Supplier entity are co-located or form a single entity, could be instantiated in the Supplier entity. The LLM function further includes a mechanism for generating a response, Disp Gen, as described in [Fig. 6].5], implementing a method for generating a response in all the embodiments described, including those of [Fig. 1] and [Fig. 2].

[0120] We then refer to [Fig 4] which presents a Disp Sel selection device for source data, adapted to implement the source data selection process, according to a particular realization.

[0121] The Disp Sel device includes a data processing module comprising a storage space 401, for example memory (MEM), a processing unit 402, equipped for example with a microprocessor (PROC), and controlled by a program computer (PGR) 403 whose instructions are configured to implement the selection process as described above in relation to [Fig. 1] and [Fig. 2],

[0122] At initialization, the code instructions of computer program 403 are, for example, loaded into memory 401 before being executed by the processor of processing unit 402. The microprocessor of processing unit 402 implements, according to the instructions of computer program 403, the steps of the selection process described above with reference to [Fig. 1] and [Fig. 2].

[0123] To this end, in addition to memory 401 and processor 402, the device includes communication means 404, enabling it to exchange messages with other devices. These communication means include, for example, the Res network, an Ethernet network interface, WiFi, 3G, 4G, 5G, etc. The communication means 404 allow the Disp Sel device, in particular, to exchange data with generation and retrieval devices and other entities in [Fig. 1]. These communication means 404 notably allow the client agent to receive a request and a certificate associated with a statement from the client agent regarding the request.

[0124] The Disp Sel device includes a 405 calculation module configured to calculate a representative vector of the received request and statement.

[0125] The Disp Sel device further includes a selection module 406 configured to select at least one source data by means of a comparison between a vector representing the source data, used for the development of a response to the query and also representing a constraint on the use of the source data, and the calculated vector.

[0126] We then refer to [Fig 5] which presents a Disp Gen response generation device adapted to implement the response generation process, according to a particular embodiment.

[0127] The Disp Gen device includes a data processing module comprising a storage space 501, for example a memory (MEM), a processing unit 502, equipped for example with a microprocessor (PROC), and controlled by a computer program (PGR) 503 whose instructions are configured to implement the response generation process as described previously in relation to [Fig. 1] and [Fig. 2],

[0128] At initialization, the code instructions of computer program 503 are, for example, loaded into memory 501 before being executed by the processor of processing unit 502. The microprocessor of processing unit 502 implements, according to the instructions of computer program 503, the steps of the generation process described above with reference to [Fig. 1] and [Fig. 2].

[0129] To this end, in addition to memory 501 and processor 502, the device includes communication means 504, enabling it to exchange messages with other devices. These communication means include, for example, the Res network, an Ethernet network interface, WiFi, 3G, 4G, 5G, etc. The communication means 504 allow the Disp Gen device, in particular, to exchange data with selection and retrieval devices and other entities in [Fig. 1]. The communication means 404 are specifically configured to allow the reception of a request from the client agent, and which also includes a certificate associated with a statement relating to the request. This request also includes an identifier of a vector representing source data and, furthermore, representing a statement from the client agent relating to the request.The 404 communication methods are also configured to allow the transmission of a response to the request to the client agent.

[0130] The Disp Gen device includes a 505 retrieval module configured to obtain from at least one data provider the source data associated with the received vector identifier.

[0131] The Disp Gen device further includes a 506 processing module configured to process a response using said source data, and possibly also includes a certificate associated with the source data.

[0132] We then refer to [Fig 6] which presents a device for obtaining a response, adapted to implement the process of obtaining a response, according to a particular realization.

[0133] The Disp Obt device includes a data processing module comprising a storage space 601, for example a memory (MEM), a processing unit 602, equipped for example with a microprocessor (PROC), and controlled by a computer program (PGR) 603 whose instructions are configured to implement the method of obtaining a response as described previously in relation to [Fig. 1] and [Fig. 2],

[0134] At initialization, the code instructions of computer program 603 are, for example, loaded into memory 601 before being executed by the processor of processing unit 602. The microprocessor of processing unit 602 implements, according to the instructions of computer program 603, the steps of the obtaining process described above with reference to [Fig. 1] and [Fig. 2].

[0135] To achieve this, in addition to the 601 memory and 602 processor, the device includes 604 communication means, enabling it to exchange messages with other devices. These communication means include, for example, the Res network, an Ethernet network interface, WiFi, 3G, 4G, 5G, etc. The 604 communication means allow in particular the Disp Obt device to exchange data with the selection and generation devices and other entities of the [Fig. 1],

[0136] The Disp Obt device includes a 605 transmission module configured to send a request and further includes a certificate associated with a client agent statement relating to the request.

[0137] The Disp Obt device further includes a 606 receiver configured to receive a response to the request, said response being developed using the source data associated with a vector representing the source data and also representing a constraint on the use of the source data, said response possibly also including a certificate associated with the source data.

[0138] The Disp Obt device also includes an optional 607 validation module allowing validation of the response received by comparing it to the request transmitted and by validating one or more certificates received.

Claims

Demands

1. A method for generating a response to a request transmitted by a client agent (Agt) associated with a large language model (LLM), the method implemented by a generation device comprising a processor, the method comprising: - a receipt (E50, F50) of a request from the client agent, and further including a certificate associated with a statement relating to the request, said request further including an identifier of a vector representing source data, said source data contributing to the preparation of the response, and further representing a statement from the client agent relating to the request, - obtaining (E60, F60) from at least one data provider (Fourn) the source data associated with the received vector identifier, - a transmission (E80, F80) to the client agent (Agt) of a response to the request, said response being developed (E70, F70) using said source data.

2. A generation method, according to claim 1, wherein the response to the request transmitted to the client agent includes a certificate associated with the source data used to develop said response.

3. A generation method, according to claim 1 or claim 2, wherein the request is received from a mediation agent (Mandate), said request from the client agent having been certified by the mediation agent.

4. A generation method, according to any one of the preceding claims, wherein the received request further includes a certificate associated with an identity of the client agent who issued the request.

5. A generation method according to claim 3, wherein the client agent (Agt) to which the response to the request is transmitted is determined by means of the certificate associated with the identity of the client agent received.

6. A generation method according to any one of the preceding claims, further comprising obtaining from a second data provider a second source data, said second source data not corresponding to the statement relating to the received request and requesting from the client agent a certificate associated with a statement with the second source data managed by the second data provider.

7. A generation method, according to any one of the preceding claims, wherein the received request further includes a certificate associated with the source data and / or the data provider and / or an owner of the source data.

8. A generation method, according to any one of the preceding claims, wherein the response transmitted to the client agent further includes a certificate of the large language model.

9. A generation method, according to any one of the preceding claims, wherein at least one of the certificates transmitted with the response further includes proof of validity of at least one of the certificates established by a server of the large language model.

10. A generation method according to claim 8, wherein the verification of the validity of at least one certificate is carried out using a decryption algorithm determined from information present jointly with the proof of validity.

11. A method of generating, according to any one of the preceding claims, further comprising receiving from the client agent a validation message of the response received.

12. A generation method according to any one of the preceding claims, further comprising a record of the operations performed for the generation of the response and / or comprising a record of the entities that contributed to this generation.

13. A method for obtaining a response to a query, developed from source data made available by a data provider (Fourn), said source data contributing to the development of the response, the method implemented by a client agent (Agt) associated with a large language model comprising - an issuance (E30, F30) of a request and further including a certificate associated with a statement from the client agent (Agt) relating to the request, - a reception (E80, F80) response to the request, said response being developed using the source data associated with a vector representing the source data and also representing a constraint on the use of the source data.

14. A device for generating a response to a request transmitted by a client agent (Agt) associated with a large language model (LLM), said selection device (Disp Gen) comprising a processor (PROC 502) coupled to a memory (MEM 501) in which program instructions (PGR 503) are stored for execution by the processor, comprising: - a receipt (E50, F50) of a request from the client agent, and further including a certificate associated with a statement relating to the request, said request further including an identifier of a vector representing source data, said source data contributing to the preparation of the response, and further representing a statement from the client agent relating to the request, - obtaining (E60, F60) from at least one data provider the source data associated with the received vector identifier, - a transmission (E80, F80) to the client agent of a response to the request, said response being developed (E70, F70) using said source data.

15. A device for obtaining (Disp Obt) a response to a query, generated from source data, said source data contributing to the generation of the response, made available by a data provider, said device for obtaining (Disp Obt) comprising a processor (PROC 602) coupled to a memory (MEM 601) in which program instructions (PGR 603) are stored for execution by the processor, comprising - an issuance (E30, F30) of a request and further including a certificate associated with a statement from the client agent relating to the request, - a reception (E80, F80) of a response to the request, said response being developed using the source data associated with a vector representing the source data and also representing a constraint on the use of the source data.

16. A system for generating a response to a request transmitted by a client agent associated with a large language model, said system comprising: - A generation device (Disp Gen) according to claim 14, - A obtaining device (Disp Obt) according to claim 15.

17. Program product comprising program code instructions for implementing a generation method according to any one of claims 1 to 12 when executed by a processor.

18. Program product comprising program code instructions for implementing a method of obtaining according to claim 13 when executed by a processor.

Citation Information

Patent Citations

  • Large model cue word verification processing method and device

    CN117131529A

  • Techniques for securing, accessing, and interfacing with enterprise resources

    WO2024091682A1