Method for selecting a provider for generating a response by a large language model (LLM)

The method addresses data usage control in LLMs by using vector-based data selection and certificates to ensure compliant and secure data sourcing, improving response reliability and regulatory compliance.

WO2026052541A1PCT designated stage Publication Date: 2026-03-12ORANGE SA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Current techniques for large language models (LLMs) lack control over the use of data sources, leading to indiscriminate data usage without compliance with security, confidentiality, or regulatory constraints, limiting their implementation in closed ecosystems and preventing secure data searches and automated training with external databases.

Method used

A method and device for selecting source data points using vector representation and comparison with usage constraints, accompanied by certificates for authenticity, ensuring compatibility with client requests and regulatory/commercial agreements, and a system for generating responses using trusted data providers.

Benefits of technology

Ensures secure and compliant data usage for LLM responses, enhancing reliability and compliance with regulations, allowing controlled data sourcing and improved response quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025074738_12032026_PF_FP_ABST
    Figure EP2025074738_12032026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for selecting at least one source datum, the source datum being suitable for generating a response to a request from a client agent (Agt) associated with a large language model (LLM), the method comprising: - receiving (E30), from the client agent (Agt), a request and a certificate associated with a declaration of the client agent relating to the request; - calculating (E40) a vector representative of the received request and declaration; - selecting (E40) the at least one source datum by means of a comparison between a vector representative of the source datum used for formulating a response to the request and furthermore representative of a usage constraint for the source datum, and the calculated vector.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR SELECTING A SUPPLIER FOR RESPONSE GENERATION USING A LARGE LANGUAGE MODEL (LLM)

[0001] This disclosure falls within the domain of artificial intelligence based on a large language model and is specifically aimed at ensuring certification of the data used and control of the use of this data to respond to an intention or request from a client agent.

[0002] Large language models (LLMs) are increasingly used in generative artificial intelligence techniques and aim in particular to produce a text response to an intention, which can be called either a question or a query.

[0003] One of the key features of a large language model is its ability to utilize a wide range of data sources, either learned or received with the query, and to generate a response to a query it hasn't previously encountered. The large language model can thus be trained on a vast array of source data. Furthermore, for example, when dealing with sensitive data requiring usage control or data relevant to a real-time function, data can be sent along with a request received from a client agent. This data is then used by the large language model to formulate a response, thereby minimizing the risk of hallucinations. It's also important to note that training a large language model requires significant resources and represents an increasingly significant cost.

[0004] A large language model is therefore trained to respond to a wide variety of intentions using these diverse and numerous data sources.

[0005] Thus, in response to a request, the large language model uses indiscriminately the different data sources to which it has access to produce a response to the request received from a client agent.

[0006] According to the previous technique, there is no control over the use of different data sources. Thus, regardless of the request received, the large language model uses any data source, and consequently any data from that source, to construct and provide a response to the client agent that issued the request. Known techniques therefore do not allow for searching and using data sources that comply with certain constraints regarding security, disclosure, confidentiality, or any other statement associated with the request. Control over the use of data from a source by a large language model is not defined in the current specifications. As a result, data can be used and reused without any disclosure control and without any guarantee of reliability.In particular, it is not possible, according to known techniques, to use data only if the context of use of this data complies with certain conditions specific to a regulation or a commercial agreement for example, thus limiting the use of data to certain actors or to a particular context.

[0007] The techniques known in this way do not allow the implementation of a large language model for a closed group of users in an ecosystem (consortium, commercial domain, interest group, etc.) while controlling the participants, the data sources used, and the use of this data.

[0008] Current techniques do not allow for the automation or tracking of the use of external databases or unlearned documents during the training of the large language model to generate a response to an intention.

[0009] These techniques also do not allow for the secure search of data sources from different actors and / or the search for data sources that comply with regulations or contractual agreements.

[0010] The purpose of this disclosure is to address all or part of the limitations of prior art solutions, including those described above, by proposing a solution that enables control over the use of data used to develop a response to a customer agent's intent.

[0011] To this end, a method is proposed for selecting at least one source data point, said source data being adapted for developing a response to an intention of a client agent associated with a large language model, the method implemented by a selection device comprising a processor, the method comprising:

[0012] - Receipt from the client agent of a request and a certificate associated with a statement from the client agent regarding the request,

[0013] - Calculation of a vector representing the request and the received statement, - Selection of at least one source data by means of a comparison between a vector representing the source data, used for the development of a response to the request, and also representative of a constraint on the use of the source data, and the calculated vector.

[0014] The selection process allows for the selection of source data, for example, by selecting a data provider whose statements are compatible with those of an agent requesting a response to a query. This compatible source data can then contribute to the development of the response, for example, via a Language Management (LLM) tool, in response to a query or prompt. Source data is data (text, video, image, etc.) made available by a data provider and / or data owner for use in formalizing or contributing to the development of a response to a query or prompt from a client agent. The source data can thus be made available, that is, used or exploited to generate a response to a query or prompt.The source data can be external, notably found in a database external to the IT domain (company, operator offering, organization, etc.) in which the selection process is implemented, or data from a database internal to the IT domain. Selecting source data thus allows, by extension, the selection of a data provider whose declarations are compatible with a query. The values ​​of vectors, calculated to represent an agent's query, and of source data, enable the identification of a data provider offering source data whose characteristics are compatible with the query. The vector calculation includes the declaration related to the query, and the vector representing the source data represents a constraint on the use of that source data.This implementation allows for enriched vector values ​​and vector comparisons, enabling comparisons of both the query and source data, as well as query declarations and data usage constraints. In particular, when the calculated vector and the representative vectors of data sources are trusted—that is, calculated by trusted connectors or servers—vector comparison alone allows for the selection of source data and potentially a source data provider, which may be compatible with the query requirements.The values ​​of vectors, calculated to represent an agent's request on the one hand, and of data sources on the other, make it possible to determine, for example, a data provider whose statements associated with a source data are compatible with the requirements formulated in the statements relating to the request transmitted with the request. A certificate, such as a digital certificate, is an identifier or electronic document that can prove the authenticity of a user, a terminal, or even a server. For the identity verification of an entity, a certificate contains information such as organizational information (name and address), a public key corresponding to a private key used to encrypt the text from the certificate owner, and other information such as the name of the certification authority and a digital signature to prove the authenticity of the certificate issuer.The certificate is therefore associated with a declaration (context of the request, organizational environment, level of confidentiality required for the response to the request, search domain for the request, identity or information about the sender of the request…).

[0015] According to one aspect of the invention, the representative vector of the source data is previously obtained from a data provider making the source data available in accordance with the usage constraint.

[0016] Thus, representative vectors of source data can be obtained from data providers. This source data, along with its usage constraints, can be obtained, for example, from a data owner or source. This source data is only used if it aligns with the declarations (regulations, commercial agreements, etc.) of a client request, thereby protecting data dissemination and preventing the use of data from these providers if the declarations between the client request and the usage constraints of this source data are incompatible.

[0017] In another aspect, a certificate attesting to the identity and authenticity of the data provider is also obtained with the representative vector of the source data.

[0018] The security of the process can be improved not only by obtaining a representative vector of a source data and a constraint on the use of this source data but also by accompanying this vector with a certificate attesting to the authenticity of the data provider transmitting this vector, this certificate can then be transmitted to the LLM function generating a response to the request.

[0019] According to another aspect, the selection of at least one source data is carried out in accordance with a comparison operation implemented between the calculated vector and at least one vector included in a database comprising vector identifiers representative of the statements associated with source data.

[0020] A database, for example managed by a proxy server also called a broker, containing vector identifiers and possibly the values ​​of these vectors, can be advantageously used to compare the calculated vector of the request with the representative vectors of the source data. This database can thus be used as soon as a request is received from an agent. The comparison can be adapted to determine a greater or lesser number of source data points. This comparison can advantageously take into account the response or acknowledgment message received from client agents, which evaluates the LLM responses and thus assesses the different source data points used to generate those responses.

[0021] According to another aspect, a selection of at least one second source data is made, this at least one second source data being adapted to contribute to developing said response and requiring from the client agent a certificate associated with a declaration compatible with a constraint of use of at least one second source data.

[0022] For example, by using a database containing vectors, it is possible to select an alternative or second source data source whose declarations or usage constraints do not precisely match the declarations of the received request, but which could potentially contribute more optimally or completely to generating a response to the request. However, using this second source data source, which is initially incompatible with the declaration associated with the received request, requires a different certificate, or an updated certificate, from the client agent. The agent can then be asked to issue this new or updated certificate, which includes a supplementary declaration compatible with the declaration or usage constraint of the second source data. This can lead to the calculation of a vector associated with this request and an associated declaration, based on the associated certificate.This improves the response to the query while maintaining the correspondence of the statements between the query and the second source data.

[0023] According to another aspect, the selection process further includes sending to a large language model the received request, the certificate associated with a statement from the client agent relating to the request, and the identifier of the vector representing an owner of the selected source data and / or the data provider making the source data available.

[0024] The selection process, in one embodiment, involves requesting a Data Link Management (DLM) system, which is responsible for obtaining the data necessary to formulate a response to the query by contacting the data provider(s) determined to be compatible with the query. The certificate identifying the requesting agent can be used so that the LLM transmits the response directly or indirectly to the requesting agent. Given that several source datasets may be managed by different data owners, the vectors representing these source datasets can advantageously be based on certificates associated with these data owners and / or the data provider making them available to respond to a query whose declaration complies with the usage constraints of these source datasets.

[0025] Depending on a particular aspect, depending on the selection process, the large language model to which the request is transmitted is selected based on contractual or regulatory information present in the previously transmitted client agent's declaration and / or a characteristic relating to the previously obtained data provider.

[0026] Given that several LLMs may potentially be used, the LLM in charge of developing a response may be the one that corresponds to constraints, for example, of a regulatory or contractual type, relating to a client agent or a data provider.

[0027] According to another aspect of the selection process, the request received from the client agent also includes proof of the validity of the certificate associated with a statement from the client agent regarding the request.

[0028] Particularly in an open or unadministered context, the request may also advantageously include proof of the validity of the certificate associated with a statement from the client agent relating to the request, relating to the statement, this validity being able to then be checked by the entities in charge of selecting a data provider and / or responding to the request.

[0029] According to another more specific aspect, the selection process also includes a verification of the proof of validity of at least one certificate received by means of an identification information of a decryption algorithm present in the message received from the client agent.

[0030] The entity in charge of verifying the validity of the certificate(s) may use an identification information from a decryption algorithm present in the received request.

[0031] According to another aspect of the selection process, the comparison between the vector representing the source data, said source data being used for the development of a response to the query, and also representing a constraint on the use of the source data, and the calculated vector is carried out in accordance with a filtering operation associated with the constraint on use and / or the client agent's statement relating to the query.

[0032] Different vector comparison rules can be used depending, for example, on the constraints of using the source data and the declaration associated with the query. Thus, for example, for constraints related to a security context, a comparison operation may be stricter than for a commercial context. It is therefore advantageous to be able to adapt the comparison rule to the type of metadata associated with a source data point, for example.

[0033] The different aspects of the selection process that have just been described can be implemented independently of each other or in combination with each other.

[0034] The invention also relates to a device for selecting at least one source data item, said source data item being adapted for generating a response to an intention of a client agent associated with a large language model, said selection device comprising a processor coupled to a memory in which program instructions intended to be executed by the processor are stored, by a method comprising:

[0035] - Receipt from the client agent of a request and a certificate associated with a statement from the client agent regarding the request,

[0036] - Calculation of a vector representing the request and the received statement, - Selection of at least one source data by means of a comparison between a vector representing the source data, used for the development of a response to the request, and also representative of a constraint on the use of the source data, and the calculated vector.

[0037] This selection device is capable of implementing, in all its modes of realization, the selection process that has just been described.

[0038] The invention also relates to a system for selecting at least one source data, said source data being adapted for developing a response to an intention of a client agent associated with a large language model, said selection system comprising: A selection device as described above, A client agent, adapted to communicate with the selection device and configured to issue the request and a certificate associated with a statement of the client agent relating to the request.

[0039] The invention also relates to a computer program product comprising a set of program code instructions which, when executed by at least one processor, configure said at least one processor to implement the method of selecting at least one source data according to any one of the implementation modes of this disclosure.

[0040] The invention also relates to a computer-readable recording medium on which is recorded a set of program code instructions which, when executed by at least one processor, configure said at least one processor to implement a method for selecting at least one source data according to any one of the implementation modes of this disclosure.

[0041] Furthermore, a method is proposed for generating a response to a request transmitted by a client agent associated with a large language model, the method comprising: - receiving a request from the client agent, and further comprising a certificate associated with a statement relating to the request, said request further comprising an identifier of a vector representing source data and further representing a statement from the client agent relating to the request,

[0042] - obtaining the source data associated with the received vector identifier from at least one data provider,

[0043] - a transmission to the client agent of a response to the request, said response being developed using said source data, and also including a certificate associated with the source data.

[0044] According to one embodiment of the invention, an LLM can thus obtain an identifier of a vector associated with source data that can contribute to formulating a response to a request. This source data has been selected because of its compatibility with a declaration or metadata associated with a request issued by the agent. The LLM agent can then generate a response using source data whose usage and dissemination constraints are compatible with the declaration (specific to a service, a usage constraint, a regulation, or a context associated with the request) relating to the received request. The LLM can then transmit the response, along with the certificates of the source data used to generate the response and compatible with the declarations associated with the request, to the client agent.Data providers thus have a guarantee regarding the use of the source data they receive and store. Only source data compatible with the requirements and constraints of the request is requested and subsequently used to generate a response. The received request includes a certificate associated with a declaration related to the request, for example, in connection with a regulation or contractual agreement. This allows the LLM (Logical Lifecycle Management) system to verify the compatibility of the source data selected and made available by a service provider with the client agent's declarations. The declaration may, but is not limited to, a control over data usage or disclosure and / or a security feature related to the request.

[0045] According to one aspect of the generation process, the request is received from a mediation agent, the said request from the client agent having been certified by the mediation agent.

[0046] The request can advantageously originate from a mediation agent or broker, which may include a selection mechanism. This mechanism can manage both requests from client agents and proposals from service providers regarding source data, thereby selecting source data, and consequently data providers, that are compatible with the client agents' requests. This facilitates, improves the reliability of, and accelerates the LLM's task in developing a response. In this way, the mandated entity only manages vector identifiers without having access to the source data, thus limiting potential problems of unwanted disclosure of this source data.

[0047] According to another aspect of the generation process, the received request also includes a certificate associated with an identity of the client agent who issued the request.

[0048] The request received from the client agent advantageously includes a certificate associated with an identity of the client agent, thus allowing the large language model, on the one hand, to authenticate the client agent issuing the request.

[0049] According to a particular aspect of the generation process, the client agent to which the response to the request is transmitted is determined by means of the certificate associated with the identity of the client agent received.

[0050] The client agent to which the response to the request is transmitted is determined from the certificate associated with the identity of the client agent received, which allows on the one hand for the possibility that the request may not be received directly from the client agent and on the other hand for the identity of the client agent to be certified.

[0051] According to one aspect of the invention, the generation process further comprises obtaining from a second data provider a second source data, said second source data not corresponding to the declaration relating to the request received and requiring from the client agent a certificate associated with a declaration with the second source data managed by the second data provider.

[0052] The LLM may solicit a data provider whose source data can contribute to generating a more relevant response to the received query, corresponding to a query, but not corresponding to a statement, such as a disclosure or security or sharing constraint, obtained with the initially received query.

[0053] According to one aspect of the generation process, the received request also includes a certificate associated with the source data and / or the data provider and / or a source data owner.

[0054] The received request may, in one embodiment, include a certificate associated with the source data and / or the data provider, thus ensuring the authenticity of the source data and / or the entity that transmitted it and / or the owner of the source data, notably by verifying the certificate's validity. This information is also useful when the single vector cannot be fully utilized by the large language model to identify and guarantee the authenticity of the source data. The received certificate(s) thus enable the LLM to determine which data provider to contact to obtain the source data needed to generate the response to the request.

[0055] According to yet another aspect of the generation process, the response transmitted to the client agent also includes a certificate of the large language model.

[0056] The response sent to the customer agent may, in one embodiment, include a certificate attesting to the authenticity of the LLM that generated the response, which may be important particularly in cases where the LLM is specific to a regulatory or organizational context.

[0057] According to one aspect of the generation process, at least one of the certificates transmitted with the response also includes proof of validity of at least one of the certificates established by a server of the large language model.

[0058] The customer agent can thus ensure the validity of the certificates (from the data provider, the source data, etc.) received with the response and ensure that the data providers and the source data are properly authenticated and validated.

[0059] According to one aspect of the generation process, the verification of the validity of at least one certificate is carried out using a decryption algorithm determined from information present jointly with the proof of validity.

[0060] The decryption algorithms therefore do not need to be stored by the LLM, and thus the LLM can verify any certificate, including those of data providers or data sources that have never provided data before. The transmitted information can thus correspond to a public key used to verify validity by decrypting certificate validity information.

[0061] According to one aspect of the invention, the generation process further includes receiving from the client agent a validation message of the response received.

[0062] The validation message received from the client agent allows us to validate the data providers and source data that contributed to developing the response and, if necessary, to confirm that the provider and possibly the source data have control over the use of the source data in accordance with the requirements issued by the client agent.

[0063] According to one aspect of the invention, the generation process further includes a record of the operations performed for the generation of the response and / or including a record of the entities that contributed to this generation.

[0064] Recording operations (exchanges, interactions with other entities) and entities (data provider, source data, data owner, mediation entity, client agent), for example in a Distributed Hash Table (DHT), Distributed Ledger Technology (DLT), or Data Clearing House, makes it possible to trace information related to response generation and subsequently identify, or even validate, an entity's participation in generating the response. This record can also be used for auditing and / or reuse in a subsequent request.

[0065] The different aspects of the generation process that have just been described can be implemented independently of each other or in combination with each other.

[0066] The invention further relates to a method for obtaining a response to a query, developed from source data made available by a data provider, the method implemented by a client agent associated with a large language model comprising

[0067] - the issuance of a request, which also includes a certificate associated with a statement from the client agent regarding the request,

[0068] - a response to the request, said response being developed using the source data associated with a vector representing the source data and also representing a constraint on the use of the source data, said response also including a certificate associated with the source data.

[0069] The invention also relates to a device for generating a response to a request transmitted by a client agent associated with a large language model, said selection device comprising a processor coupled to a memory in which program instructions intended to be executed by the processor are stored, comprising:

[0070] - a receipt of a request from the client agent, and further including a certificate associated with a statement relating to the request, said request further including an identifier of a vector representing source data and further representing a statement from the client agent relating to the request,

[0071] - obtaining the source data associated with the received vector identifier from at least one data provider,

[0072] - a transmission to the client agent of a response to the request, said response being developed using said source data, and also including a certificate associated with the source data.

[0073] This generation device is capable of implementing, in all its embodiments, the generation process that has just been described.

[0074] The invention also relates to a device for obtaining a response to a query, generated from source data provided by a data supplier, said obtaining device comprising a processor coupled to a memory in which program instructions intended to be executed by the processor are stored, comprising

[0075] - the issuance of a request, which also includes a certificate associated with a statement from the client agent regarding the request,

[0076] - a receipt of a response to the request, said response being developed using the source data associated with a vector representing the source data and also representing a constraint on the use of the source data, said response also including a certificate associated with the source data.

[0077] The invention also relates to a system for generating a response to a request transmitted by a client agent associated with a large language model, said system comprising: - A generation device as described above,

[0078] - A retrieval device as described previously.

[0079] The invention also relates to a computer program product comprising a set of program code instructions which, when executed by at least one processor, configure said at least one processor to implement the method of generating a response according to any one of the implementation modes of this disclosure.

[0080] The invention also relates to a computer-readable recording medium on which is recorded a set of program code instructions which, when executed by at least one processor, configure said at least one processor to implement a method of generating a response according to any of the implementation modes of this disclosure.

[0081] The invention also relates to a computer program product comprising a set of program code instructions which, when executed by at least one processor, configure said at least one processor to implement the method of obtaining a response according to any one of the implementation modes of this disclosure.

[0082] The invention also relates to a computer-readable recording medium on which is recorded a set of program code instructions which, when executed by at least one processor, configure said at least one processor to implement a method for obtaining a response according to any of the embodiments of this disclosure.

[0083] The invention will be better understood upon reading the following description, given by way of non-limiting example, and made with reference to the figures which represent: a schematic representation of a communication architecture in which is implemented a method of selecting at least one data provider, a method of generating a response to a request as well as a method of obtaining a response to a request according to an embodiment.a diagram illustrating the main steps of a process for selecting at least one data provider, a process for generating a response to a query and a process for obtaining a response to a query according to another embodiment, an infrastructure comprising a generation device, a selection device and a transmission device according to an embodiment of the invention according to an example, a representation of a device for selecting source data according to another example, a representation of a device for generating a response to a query according to another example, a representation of a device for obtaining a response to a query according to another example.

[0084] In these figures, identical references from one figure to another designate identical or analogous elements. For clarity, the elements shown are not to scale unless otherwise indicated.

[0085] More generally, it should be noted that the implementation and realization methods considered above have been described as non-limiting examples, and that other variants are therefore conceivable.

[0086] The following description presents an architecture consisting of a single large language model, a single client agent, a single data owner, and a single data provider, but the processes can be implemented interchangeably in an architecture comprising several of these entities.

[0087] We refer first of all to the one which describes a schematic representation of a communication architecture in which is implemented a process of selecting at least one data provider, a process of generating a response to a request as well as a process of obtaining a response to a request according to an embodiment mode.

[0088] The architecture includes an Agt agent corresponding to a large language model. The LLM Agt client agent can interact with a person or a computer entity, such as a robot, not represented on the system. This person or computer entity can ask a question to an artificial intelligence system, which the LLM Agt client agent will then clarify, specify, and, in this case, complete. To this end, the LLM client agent has a wallet that allows it to add one or more statements to a request transmitted by the Agt client agent. More specifically, the Agt client agent can transmit a certificate associated with the client agent's identity, attesting to the authenticity of the Agt client agent. It can also be contractual or regulatory information, for which the associated certificate is attached to a request transmitted by the agent, attesting to the validity of its statement.For example, the Agt client agent can indicate in a statement that it belongs to a domain (Human Resources, Accounting, Legal, etc.) and therefore requests a response related to that domain or a related domain, and / or a response with a high level of security, and / or one specific to a particular entity or consortium of companies. Any type of statement that can be used to filter and tailor the response to a query is possible. Naturally, the client agent can issue multiple statements in its query. It should be noted that the Agt client agent also interacts with a Mandat entity (for agent), a Fourn data provider, and a large LLM language model, whose roles and functions will be described below.

[0089] The architecture therefore includes a Mandate entity, also called a Broker. The Mandate entity acts as an intermediary between the various entities involved in the selection, generation, and acquisition processes described in the different embodiments of the invention. In particular, it performs the parameterization function (also called Embedding Function (EF1)) of the requests and declarations transmitted by the client agent Agt. This parameterization function consists of representing the request received from the client agent Agt and the associated declaration as a vector. The vector thus integrates the various declarations, such as contractual or regulatory information, transmitted by the client agent in a certified manner. The Mandate entity also manages a vector database, Vect1, containing the requests received from the client agent parameterized as vectors.This vector database also interfaces with a vector database, Vect2, maintained by the data provider Fourn. This interface allows the proxy server to match a request and a statement from an Agt client agent with source data, and a statement from a data provider with the statements of the source data owner, appropriate for responding to the request, according to statements compatible between the Agt client agent's request and the source data. The source data itself is associated with an accreditation, that is, a certified statement, made available by the data provider and the source data owner. The statement associated with the source data corresponds to a constraint on the use of the source data. The Vect1 database applies matching rules defined within the ecosystem specific to the architecture.The Mandate entity also includes a server, Srv1, also called a connector, whose main function is to ensure interaction and communication with the other entities. Servers named Srv2, Srv3, and Srv4 are found within the data provider, the data owner, and the large language model, LLM, respectively. The functions of these different servers are relatively similar. Primarily, they ensure usage control and data sharing by applying usage policies.

[0090] - They ensure transit to other servers or to an application container and manage messages between these servers and the container,

[0091] - They verify the received attestations and present themselves to another server with the verifiable attestations concerning them involved in a transaction between a client agent and a large LLM language model,

[0092] - they orchestrate application containers such as the EF function, a VECT vector database and a large LLM language model.

[0093] The servers Srv1, Srv2, Srv3, and Srv4 also record data exchange interactions between entities, and more specifically between the servers themselves. These records, or logs, allow for the identification of problems in case of malfunctions or conflicts, and even enable billing for data exchanges. The servers also optionally perform certificate verification for declarations of source data, applications, or requests.

[0094] The architecture includes a large Language Modeling (LLM) function, corresponding to a neural network trained on a large dataset and responsible for developing a response to a request issued by a client agent, in collaboration with the Mandate, Supplier, and Proprietor entities. As mentioned above, the LLM includes an Srv4 server, also called a connector, whose functions have been described above.

[0095] The data provider Fourn is responsible for maintaining a database of Vect2 vectors associated with one or more source data points, which may come from distinct data sources, retrieved and managed by a Prop data owner. The source data points can be text files, images, videos, or any type of data that can contribute to generating a response to a query. The Fourn provider therefore configures the source data points and the associated declarations for the data owner and the data provider using an EF2 function in the form of vectors, which it maintains in a Vect2 database. The Fourn provider also maintains a record of the data sources in a Stock2 file or database. The Stock2 database stores the data owners' source data points, which can be exchanged, subject to usage controls, by the data providers.A calculated vector thus represents the owner of the selected source data and / or the data provider making the source data available. This calculated vector will then be compared to a vector calculated from a query and a statement associated with that query, and an identifier of this vector, if it corresponds to a received query, can be transmitted to an LLM in order to retrieve the source data associated with that vector.

[0096] The data provider Fourn obtains the source data from the data owner(s) Prop, who maintains in a Certif file or database the certificates of the declarations associated with the source data, the information of which it stores in a Stock3 storage space.

[0097] The architecture also includes a DID (Decentralized Identification) entity. This DID entity is designed to verify, prove, and validate the exchanged certificates used by the various entities to certify a submitted declaration.

[0098] The exchanges and interactions between these different entities, according to a mode of implementation, will now be described in accordance with the numerical indications of the exchanges or operations of the entities of the.

[0099] During step E10, one or more data sources from the Stock3 storage space, containing source data potentially usable by the LLM large language model and associated declarations, are parameterized as vectors by the EF2 function, also known as the EF2 connector. These data sources are thus enriched with certified declarations, called accreditations, added by the Prop entity. These accreditations can be of various types and may, for example, correspond to characteristics of security, confidentiality, sharing, compliance with a technical or administrative rule, compliance with a commercial domain or contractual exchange area, or any type of information that qualifies the data sources and restricts or allows the sharing of data from these sources.Source data can be supplemented with a declaration indicating that it is financial data, HR data, data specific to an entity (e.g., Operator A), data with a Top Secret security level, or any other type of declaration restricting the use of this source data to specific requests or queries. Using its Certif file, the Prop entity attests to the validity of the provided data sources and accreditations by adding, via the Srv3 server, and transmitting the certificate(s) associated with the source data declarations to the Fourn entity. Proof of the validity of these certificates can also be transmitted to the Fourn entity.

[0100] During this E10 step, the Supplier entity, via the Srv2 server, verifies the validity of the received certificate(s), for example, by querying the DID entity. To this end, if a proof of validity is transmitted, the Supplier entity can also verify the validity of the proof, for example, by using a decryption algorithm determined from a credential present in the certificate(s). The Supplier entity stores the vector(s) parameterized by the EF2 parameterization function in the Vect2 database. A vector is therefore parameterized using identification information from a source data item and also based on the accreditation added by the Prop entity to the source data.

[0101] During step E20, the Vect1 database aggregates the vectorized data from the various Fourn entities, and in particular from the Vect2 database of the Fourn entity. While only one Fourn entity is represented, the Mandat entity acts as an agent for the client agent Agt. Therefore, the Vect1 database is designed to aggregate vectors from different providers and, consequently, from a set of source data from multiple data sources. The Mandat entity thus synchronizes the vector databases from different data sources, supplemented by the certificates associated with the declarations linked to this source data, validated by the Srv2 server of the Fourn entity. One of the declarations concerns the identification information of the Fourn data provider (name, IP address, email address, FQDN, etc.).The Mandate entity only has access to the vectors, representing the source data of the data sources and the credentials associated with that source data. However, the Mandate entity does not have access to the source data itself, thus limiting the unwanted dissemination of source data that could potentially conflict with the credential requirements of that source data. The Mandate entity can also validate the certificates associated with the source data declarations, for example, by using the DID entity, and can also verify that the evidence, if transmitted, is valid, for example, by using a decryption algorithm determined from a credential present in the certificate(s). It should be noted that steps 10 and 20 can be repeated several times so that the Mandate entity updates the various available data sources that can be used to formulate a response to a client agent's request.

[0102] During an E30 step, the Agt client agent transmits a request message to the Mandate entity, including an intent and one or more certificates associated with one or more statements from the Agt client agent. One of these statements relates to identifying information about the Agt client agent (name, IP address, email address, FQDN, etc.). The client agent can also add one or more other certificates from the Certif0 certificate database, associated with other statements relating to security characteristics, compliance with a business and / or organizational environment, a service or authorization, or other characteristics that may influence the data sources used to formulate a response to the request.Thus, through the certificates associated with the statements, the Agt client agent can indicate that they belong to an organization, that they require a response regarding HR data, that they require financial data, or that they are accredited with a Top Secret security level. The source data, respectively HR, financial, or Top Secret, can then be transmitted to them depending on the statement submitted. The Agt client agent can therefore transmit several certificates associated with distinct statements. Alternatively, the Agt client agent may present, in addition to the certificate(s), proof of the validity of the certificate(s). Specifically, the Agt client agent may transmit proof of the validity of the certificate associated with the Agt client agent's identity and possibly a credential, including, for example, a public key, present in the certificate(s) that allows the proof of validity to be decrypted.Upon receipt of the request message, as well as the certificate(s), the Mandat mediation entity verifies the validity of the certificate(s), for example by requesting the DID entity or by using verification tools in its possession, for example by ensuring that a key present in the certificate is still valid, and possibly verifies that the proof provided is valid, for example by using the decryption algorithm possibly transmitted by the Agt client agent.

[0103] During an E40 step, the Mandate entity, and more specifically the Srv1 function of the Mandate entity, validates the certificate (in the rest of the document we will refer to a certificate but it may be a plurality of certificates) associated with the request received from the client agent Agt.

[0104] During step E40, the EF1 function of the Mandat entity calculates a vector representing the request and the received certificate. The EF1 parameterization function thus allows the request and the received certificate to be represented as a vector loaded into the Mandat entity's vector database, Vect1. Having received the request from the client agent Agt and knowing about vectors associated with source data linked to accreditations, the Mandat agent performs a search during step E50 for data source vectors close to the request vector using its Vect1 vector database.The Mandate entity can thus initially restrict the search for source data based on the accreditations of the source data and the certified declarations of the Agt client agent, thereby ensuring an initial sorting of the available source data. It can then perform a semantic comparison to select only the representative vectors of source data that can contribute to answering the request issued by the Agt client. For example, and without limiting this to this scenario, if an Agt client agent transmits a request requiring "operator A" information, only source data from data sources that can be used in the "operator A" context, according to their accreditation, can be used to generate a response to the request.The purpose of the selection carried out by the Mandat agent is to match the client agent's statements with the accreditations of the data sources while offering a suitable response to the request transmitted by the client and thus limit hallucinations in response to the client's request.

[0105] During this E40 step, the Mandate entity identifies, through vector comparison, potentially source data that could contribute to developing a more relevant response to the Agt customer agent's request, but whose credentials do not correspond, or do not fully correspond, to the statements associated with the request. This identification can then be used to request a new or updated statement from the Agt customer agent, which must be transmitted to the Mandate entity and the LLM function to obtain a response developed from this source data, thereby improving the quality and relevance of the response to the request.If the Agt customer agent cannot provide a new statement for this source data, because he is not authorized or does not belong to the organization for which the source data can be made available, for example, he will not be able to obtain a response to the query based on this data.

[0106] During step E50, the Mandate entity transmits a request to the LLM (Language Model Language) function. This request originates from the client agent Agt and includes a certificate associated with a statement from the client agent. The request also includes an identifier for a vector associated with source data managed by at least one data provider. The request may include more than one vector, particularly if more than one source data compatible with the statement transmitted with the Agt client agent request is identified. It should be noted that the LLM function requested by the Mandate entity, in cases where multiple LLM functions can be requested, is determined based on information present in a statement from a Supplier entity and / or the Agt client agent.Thus, the selected LLM function can be determined based on regulatory and / or contractual constraints of the Fourn entity and / or the Agt customer agent, these constraints being obtained in one or more statements associated with the Agt customer agent request and / or source data transmitted by the Fourn entity.

[0107] During step E60, the LLM function obtains from the data provider Fourn the source data associated with the received vector identifier and a certificate associated with at least one data provider. The data provider to be requested, as well as information about the source data to be requested, is obtained from information transmitted by the Mandate entity to the LLM function, specifically in one or more certificates issued by the Mandate entity. The LLM function may also transmit certificates associated with the declarations accompanying the request from the client agent Agt, certificates associated with the source data obtained from the Mandate entity (possibly obtained during step E50), and a certificate associated with the LLM function itself, transmitted by the Srv4 server. This allows the data provider Fourn to verify the identity and authenticity of the LLM function requesting the source data.The certificates associated with the Mandate entity and the client agent can also be transmitted to attest to their identities. It should be noted that, alternatively, certificates are not specifically required to attest to identity. This can occur when the EF functions are certified, and the calculated vectors are then considered sufficiently reliable to guarantee that the information retrieved from these vectors is sufficient to attest to the respective identities of the entities and functions. Thus, during this E60 step, the data provider Fourn can authenticate the various pieces of information whose certificates are transmitted, for example, via the Srv2 server or connector, and possibly by using the DID entity.

[0108] Once the received certificates are validated, the Provider entity transmits, during step E60, the source data identified by the vector identifier. Typically, if the LLM function has sent three vector identifiers, the Provider entity(ies), determined by the Mandate entity, transmits three source data items to the LLM function in response to the identifiers. The Provider entity also transmits the certificates associated with the data sources and possibly proofs of validity for this source data.

[0109] During step E70, the LLM function generates a response to the request message received from the client agent Agt, via the Mandate entity. The source data, corresponding for example to data chunks, is used by the LLM function to generate the most relevant response to the client agent's request, based on data whose credentials are compatible with the statements of the client agent Agt and possibly with the statements of the source data provider(s), Fourn.

[0110] During step E80, the LLM function sends the response generated during step E70 to the client agent Agt. This response includes the certificates associated with the source data used to generate the response and presented by the Fourn entity, the certificate(s) relating to the declarations accompanying the request received from the client agent Agt, and a certificate associated with the LLM function itself. This certificate verifies the authenticity of the LLM function and confirms that the response is indeed relevant to the request transmitted during step E30. The response may also include the certificate associated with the Mandat entity that performed the source data selection. The certificate verification is performed by the client agent Agt, possibly using the DID entity service and a decryption algorithm that may be present with the certificate(s).It should be noted that the transmitted response can be sent directly to the Agt client agent, for example, using information identifying the Agt client agent present in the request message transmitted during step E30 and then step E50. Alternatively, as one example, the response can be transmitted by the LLM function to the Agt client agent via the Mandate entity, allowing for a more efficient response to routing or even filtering constraints in the entities involved in the response generation process. The response transmitted during step E80, as one example, also includes information relating to source data not used in the development of the transmitted response but which could allow for the development of a more relevant response, provided that the Agt client agent transmits a statement compatible with the accreditation associated with the source data.The LLM function can thus transmit information about the required declaration, such as "Operator B," "Defense Confidential," or "Confidential HR Data," requiring a declaration from the client agent that is consistent with this information. The client agent, if they wish to obtain a response based on this additional source data, and if they have the capability or the right to do so, transmits a new request and the missing declaration to the LLM function, possibly via the Mandate entity, to obtain a response prepared with this additional source data.

[0111] During step E90, the Agt customer agent, having received the response from the LLM function, analyzes it, particularly with regard to accreditation compliance and the semantics of the response to the request. Optionally, the agent then transmits feedback on the response to the Mandate entity and possibly the LLM function, for example, using a satisfaction percentage or appropriate labeling. This allows the Mandate entity and potentially the LLM function to evaluate the responses provided to requests, validate the information and statements submitted by the various entities involved in preparing the response, and improve the responses provided. The Agt customer agent thus contributes to improving the response generation process.

[0112] The various processing steps performed by the different entities during the generation of the response are potentially recorded in a file or set of files, such as a distributed hash table (DHT), a permissioned or permissionless distributed ledger (DLT), or a data exchange center. This allows for subsequent use in case of problems, malfunctions, or billing needs related to actions performed by the different entities. This recording action can be carried out by the Srv1, Srv2, Srv3, and Srv4 servers of the entities contributing to the generation of the response, potentially requiring control over the usage associated with source data.

[0113] We then refer to the one that describes the main steps of a process for selecting at least one data provider, a process for generating a response to a query, and a process for obtaining a response to a query according to another embodiment.

[0114] During an F10a Disp step, a Data Owner provides a Data Producer with source data from various data sources and presents certificates associated with declarations, corresponding to accreditations, linked to this source data. The source data can be text, images, videos, or a mix of these data types. An example of the information provided by the Data Owner is shown below (the terms in the example below are not in a specific language and represent a valid encoding syntax in any language, including French).

[0115] {

[0116] "@context": [

[0117] "https: / www.w3.org / 2018 / credentials / v1"

[0118] ],

[0119] "kind": [

[0120] "VerifiablePresentation"

[0121] ],

[0122] "verifiableCredential": [

[0123] {

[0124] "@context": [

[0125] "https: / www.w3.org / 2018 / credentials / v1",

[0126] "https: / www.w3.org / 2018 / credentials / examples / v1"

[0127] ],

[0128] "id": "https: / example.com / credentials / 1872",

[0129] "type": [

[0130] "VerifiableCredential",

[0131] "IDCardCredential"

[0132] ],

[0133] "issuer": {

[0134] "id": "did:example:issuer"

[0135] },

[0136] "issuanceDate": "2010-01-01T19:23:24Z",

[0137] "credentialSubject": {

[0138] "id entreprise": "Orange",

[0139] "id produit": "service mobile orange",

[0140] "habilitation": "secret"

[0141] },

[0142] "proof": {

[0143] "type": "Ed25519Signature2018",

[0144] "created": "2021-03-19T15:30:15Z",

[0145] "jws": "eyJhb...JQdBw",

[0146] "proofPurpose": "assertionMethod",

[0147] "verificationMethod": "did:example:issuer#keys-1"

[0148] }

[0149] }

[0150] ],

[0151] "id": "ebc6f1c2",

[0152] "holder": "did:example:holder",

[0153] "proof": {

[0154] "type": "Ed25519Signature2018",

[0155] "created": "2021-03-19T15:30:15Z",

[0156] "challenge": "n-0S6_WzA2Mj",

[0157] "domain": "https: / client.example.org / cb",

[0158] "jws": "eyJhbG...IAoDA",

[0159] "proofPurpose": "authentication",

[0160] "verificationMethod": "did:example:holder#key-1"

[0161] }

[0162] }

[0163] For example, in this example, we can see that the metadata corresponding to accreditations specific to the source data are three in number and are as follows:

[0164] credentialSubject": {

[0165] "Company ID": "Orange",

[0166] "product id": "orange mobile service",

[0167] "authorization": "secret"}

[0168] The source data, for which the description is given above, is specific to an operator and more particularly to a mobile service and is of type "secret" indicating that it cannot be used for any type of intention, but only for requests for which a declaration is at least of type secret.

[0169] During an F10b Valid step, the data provider, specifically a server or connector such as the Srv2 server, verifies the certificates related to the declarations associated with the data source, received previously. It can use a DID entity, presented in the [document / section / etc.]. If the certificates are valid, they are added to the description of a data source. In the example below, if the certificates associated with a data source

[0170] 'metadata': { 'accreditation'

[0171] {

[0172] 'Company ID': Orange,

[0173] 'product ID': 'orange mobile service'

[0174] 'secret clearance': 'secret'

[0175] }

[0176] During an F10c Vect step, the data provider, and more specifically the EF2 (Embedding Function), vectorizes the source data along with the declarations relating to the data source and saves the vector in a database such as the Vect2 vector database. The vector therefore provides information not only about the source data but also about the property or properties associated with that source data. These declarations allow for certified control of access to the vectorized source data. Given that the source data can be quite large, for example, if it is a configuration file for a piece of equipment, the EF2 function vectorizes blocks (or chunks) of the source data. Thus, for a single source data item, several vectors can be parameterized.

[0177] A vector of a chunk of source data corresponds, for example, to

[0178] id': '573387acd058e615000b5cb5en' 'vector': [0.1, 0.8, 0.2, ..., 0.7].

[0179] As shown in the example below, the vector identifier also includes declarations associated with the source and may include other information such as the version of the EF2 function that vectorized the data source or the type of the source data, as in the example below, a "manual" type for this source.

[0180] 'id': '573387acd058e615000b5cb5en',

[0181] 'vector': [0.1, 0.8, 0.2, ..., 0.7],

[0182] 'metadata': { 'accreditation'

[0183] {

[0184] 'Company ID': Orange,

[0185] 'product ID': 'orange mobile service'

[0186] 'secret clearance': 'secret'

[0187] 'embedding model: 'embedding function version id'

[0188] }

[0189] 'Source type': 'manual'

[0190] 'id connector'

[0191] }

[0192] During an F20 Aggr step, the database, such as the Vect1 database, of a broker entity aggregates the vectorized data from the various supplier entities, including the Vect2 database, which contains the vector representing the data source and its certified declarations, as described above. The broker entity only contains vector identifiers and not the source data itself, thus preventing a potential leak of this source data.

[0193] During an F30 Intent step, the agent entity receives a request from a client agent, such as agent Agt. The request may include one or more statements associated with the request. These statements specify the business, organizational, contractual, regulatory, authorization, etc., context of the request and, consequently, of the source data to be used to generate the response. The agent entity verifies the certificates associated with the statements in the request and then vectorizes—that is, calculates a vector—in an F40 step. This vector includes the statements whose certificates have been validated, for example, by the EF1 function. An example of a vector of a request received by an agent entity and its certified statements is shown below:

[0194] {'query'

[0195] 'intent vector': [0.1, 0.8, 0.2, ..., 0.7],

[0196] 'metadata': { 'accreditation'

[0197] {

[0198] 'company ID': 'company certificate ID',

[0199] 'product ID': 'product certificate ID'

[0200] 'secret authorization': 'secret certificate id'

[0201] 'embedding model version': 'version certificate id'

[0202] 'company ID 2': 'company certificate ID 2',

[0203] 'robot agent id': 'robot certificate id'

[0204] }

[0205] 'Product type': 'Manual'

[0206] In the example, the declarations to the query are as follows: 'id company': 'id company certificate',

[0207] 'product ID': 'product certificate ID'

[0208] 'secret authorization': 'secret certificate id'

[0209] 'embedding model version': 'version certificate id'

[0210] 'company ID 2': 'company certificate ID 2', 'robot agent ID': 'robot certificate ID'

[0211] During this F40 Search step, the principal entity, using the Vect1 database, searches for vectors from data sources compatible with the received query. The search aims to find one or more vectors from a source data source that are close to the vectors being searched, while ensuring that the credentials associated with the query match the credentials associated with the source data. The query results identify the k vectors (this number can be defined by the search tool) that are closest to the query. In the example below, 3 vectors are identified. The result of the search with the k=3 closest vectors in the vector database corresponds, for example, to:

[0212] {'results': [{'matches': [{'id': '5733a6424776f41900660f51it',

[0213] 'score': 106.119217,

[0214] 'values': []},

[0215] {'id': '56e7906100c9c71500d772d7it',

[0216] 'score': 117.300644,

[0217] 'values': []},

[0218] {'id': '573387acd058e615000b5cb5en',

[0219] 'score': 121.347099,

[0220] 'values': []}],

[0221] 'namespace': ''}]}

[0222] The score associated with a vector gives a level of correspondence between the data source and its statements on the one hand, and the query and its statements on the other.

[0223] The following operations (equal to, not equal to, greater than, greater than or equal to, less than, less than or equal to, in a list, not in a list) can be used to compare the metadata associated with the representative vectors of the request and the source data, and their respective declarations. The operation used may differ for different declarations. Thus, for a "secret authorization" declaration, there must be an exact match between the request declaration and the source data declaration, whereas, for example, for a company declaration, the operation might consist of searching in a list, knowing that for a request from mobile operator A, a list of mobile operators associated with a source data point might be suitable.

[0224] During an F50 Trans step, the principal entity transmits the search result, including the identifiers of the vectors associated with the source data resulting from the comparison, to an LLM function. The principal entity also transmits the previously received request from the client agent, along with a certificate associated with the client agent's declaration (or certificates associated with declarations if there are multiple declarations). The LLM function chosen by the principal entity is the one that best meets the regulatory and contractual constraints of the source data provider(s) and the client agent identified in the declarations accompanying the received request. The principal entity may select the LLM function because it is compatible with the certificate constraints of other entities.

[0225] During an F60 Obt step, using the obtained vector identifiers, the LLM function retrieves the source data corresponding to these vector identifiers from the data provider. The LLM function also presents the certificate associated with the client agent's declaration (or certificates if there are multiple certificates), and possibly the certificates associated with the source data, and possibly a certificate associated with the LLM function itself.

[0226] During an F70 Gen step, the LLM function generates a response to the query using the source data obtained during the F60 step and the query data obtained during the F30 step. The LLM function can analyze the source data(s), and perform pre-processing of this source data so that the generated response best matches the query received from the client agent.

[0227] Once the response is generated, the LLM function transmits it, in an F80 Emet step, to the client agent that issued the request. To perform this transmission, the LLM function can advantageously use a client agent identifier declared with the received request. Along with the response, the LLM function transmits the certificate associated with the declaration accompanying the client agent's request, and possibly the certificate associated with the declaration associated with the data source, as well as a certificate associated with the LLM itself. This transmission preferably occurs between connectors or servers, such as the Srv4 server and a corresponding Srv0 server of the client agent.In the event that during step E50, the Mandate entity transmitted a vector identifier of a source data whose declaration does not correspond to a declaration of the request but may contribute to a more relevant response, the LLM function may indicate during step E80 the missing declaration to the request which may allow the client agent to obtain a more relevant response and also corresponding to the new declaration to be provided by the client agent to the LLM function possibly via the mandate entity, by transmitting an updated request or a new request.

[0228] During an F90 Ack step, after verifying the validity of the received certificate(s) using the Srv0 server and possibly the DID entity, the client agent validates the received response with the LLM function and possibly the proxy entity. This allows them to validate the data source selected to generate the response and thus validate the comparison operation between the vectors representing the request and the source data. These vectors also represent the respective declarations associated with the intent and the source data. If the validation is not entirely satisfactory, the proxy entity can modify the comparison operation for one or more declarations in the parameterized vectors. The message transmitted during the E90 step can optionally include the new declaration or the new request modified with the new declaration, as described above.

[0229] During an F100 step, the various operations performed by the entities, and more particularly by the servers Srv0,…,Srv4 are recorded in a distributed hash table (DHT) or a distributed ledger based on DLT technology for example (Distributed ledger technology (DLT)) or a data exchange center (Data Clearing House).

[0230] We then refer to the one that presents an infrastructure comprising a generation device, a selection device and a transmission device according to an embodiment of the invention.

[0231] This diagram represents the entities Agt, LLM, Mandat, DID, Fourn, and Prop as described in the diagram. These entities exchange messages, such as those described in the diagrams, using a communication network called Res. The Res network can be any communication network that allows the transmission and reception of data, such as a public or private IP network.

[0232] The client agent entity Agt also includes a mechanism for obtaining a response to a request, Disp Obt, as described in [reference], implementing a method for obtaining a response to a request in all the described embodiments, including those in [reference] and [reference]. The agent entity Mandat includes a selection mechanism, Disp Sel, as described in [reference], implementing a selection method in all the described embodiments, including those in [reference] and [reference]. It should be noted that the selection mechanism, Disp Sel, in another embodiment, particularly where the Mandat and Supplier entities are co-located or form a single entity, could be instantiated in the Supplier entity. The LLM function further includes a response generation mechanism, Disp Gen, as described in [reference], implementing a response generation method in all the described embodiments, including those in [reference] and [reference].

[0233] We then refer to the one that presents a Disp Sel selection device for source data, adapted to implement the process of selecting source data, according to a particular realization.

[0234] The Disp Sel device includes a data processing module comprising a storage space 401, for example a memory (MEM), a processing unit 402, equipped for example with a microprocessor (PROC), and controlled by a computer program (PGR) 403 whose instructions are configured to implement the selection process as described above in relation to the set.

[0235] At initialization, the code instructions of computer program 403 are, for example, loaded into memory 401 before being executed by the processor of the processing unit 402. The microprocessor of the processing unit 402 implements, according to the instructions of computer program 403, the steps of the selection process described above with reference to the set.

[0236] To this end, in addition to memory (401) and a processor (402), the device includes communication means (404) that allow it to exchange messages with other devices. These communication means include, for example, the Res network, an Ethernet network interface, Wi-Fi, 3G, 4G, 5G, etc. The communication means (404) enable the Disp Sel device to exchange data with generation and acquisition devices and other entities. Specifically, these communication means (404) allow the client agent to receive a request and a certificate associated with a statement from the client agent regarding the request.

[0237] The Disp Sel device includes a 405 calculation module configured to calculate a representative vector of the received request and statement.

[0238] The Disp Sel device further includes a selection module 406 configured to select at least one source data by means of a comparison between a vector representing the source data, used for the development of a response to the query and also representing a constraint on the use of the source data, and the calculated vector.

[0239] We then refer to the one that presents a response generation device adapted to implement the response generation process, according to a particular realization.

[0240] The Disp Gen device includes a data processing module comprising a storage space 501, for example a memory (MEM), a processing unit 502, equipped for example with a microprocessor (PROC), and controlled by a computer program (PGR) 503 whose instructions are configured to implement the process of generating a response as described above in relation to the set.

[0241] At initialization, the code instructions of computer program 503 are, for example, loaded into memory 501 before being executed by the processor of the processing unit 502. The microprocessor of the processing unit 502 implements, according to the instructions of computer program 503, the steps of the generation process described above with reference to the.

[0242] To this end, in addition to memory (501) and a processor (502), the device includes communication means (504) that allow it to exchange messages with other devices. These communication means include, for example, the Res network, an Ethernet network interface, Wi-Fi, 3G, 4G, 5G, etc. The communication means (504) specifically allow the Disp Gen device to exchange data with selection and retrieval devices and other entities. The communication means (404) are configured to receive a request from the client agent, which includes a certificate associated with a statement about the request. This request also includes an identifier of a vector representing source data and a statement from the client agent about the request.The 404 communication methods are also configured to allow the transmission of a response to the request to the client agent.

[0243] The Disp Gen device includes a 505 retrieval module configured to obtain from at least one data provider the source data associated with the received vector identifier.

[0244] The Disp Gen device further includes a 506 processing module configured to process a response using said source data, and further includes a certificate associated with the source data.

[0245] We then refer to the one that presents a device for obtaining a response, adapted to implement the process of obtaining a response, according to a particular realization.

[0246] The Disp Obt device includes a data processing module comprising a storage space 601, for example a memory (MEM), a processing unit 602, equipped for example with a microprocessor (PROC), and controlled by a computer program (PGR) 603 whose instructions are configured to implement the method of obtaining a response as described above in relation to the set.

[0247] At initialization, the code instructions of computer program 603 are, for example, loaded into memory 601 before being executed by the processor of processing unit 602. The microprocessor of processing unit 602 implements, according to the instructions of computer program 603, the steps of the obtaining process described above with reference to the.

[0248] For this purpose, in addition to the memory 601 and the processor 602, the device includes communication means 604, enabling it to exchange messages with other devices. These communication means are, for example, the Res network, an Ethernet network interface, WiFi, 3G, 4G, 5G, etc. The communication means 604 allow, in particular, the Disp Obt device to exchange data with the selection and generation devices and other entities of the.

[0249] The Disp Obt device includes a 605 transmission module configured to send a request and further includes a certificate associated with a client agent statement relating to the request.

[0250] The Disp Obt device further includes a 606 receiver configured to receive a response to the request, said response being elaborated using the source data associated with a vector representing the source data and further representing a constraint on the use of the source data, said response further including a certificate associated with the source data.

[0251] The Disp Obt device also includes an optional 607 validation module allowing validation of the response received by comparing it to the transmitted request and by validating one or more received certificates.

Claims

Method for selecting at least one source data, said source data being adapted for the development of a response to a request from a client agent (Agt) associated with a large language model (LLM), the method implemented by a selection device comprising a processor, the method comprising: - a reception (E30, F30) from the client agent (Agt) of a request and a certificate associated with a statement from the client agent relating to the request, - a calculation (E40, F40) of a vector representing the request and the received statement, - a selection (E40, F40) of at least one source data by means of a comparison between a vector representing the source data, used for the development of a response to the request, and also representing a constraint on the use of the source data, and the calculated vector. Selection method, according to claim 1, wherein the representative vector of the source data is previously obtained (E20, F20) from a data provider making the source data available in accordance with the usage constraint. Selection method, according to claim 2, wherein a certificate attesting to the identity and authenticity of the data provider is further obtained with the representative vector of the source data. Selection method, according to any one of the preceding claims, wherein the selection of at least one source data is carried out in accordance with a comparison operation implemented between the calculated vector and at least one vector included in a database comprising vector identifiers representative of the statements associated with source data. Selection method, according to any one of claims 1 to 4, wherein a selection of at least one second source data is carried out, this at least one second source data being adapted to contribute to developing said response and requiring from the client agent a certificate associated with a declaration compatible with a constraint on the use of the at least one second source data. Selection method, according to any one of the preceding claims, further comprising sending to a large language model the received request, the certificate associated with a client agent statement relating to the request and the identifier of the vector representing an owner of the selected source data and / or the data provider making the source data available. Selection method, according to claim 6, wherein the large language model to which the request is transmitted is selected based on contractual or regulatory information present in the previously transmitted customer agent declaration and / or a previously obtained data provider characteristic. Selection method, according to any one of the preceding claims, wherein the request received from the client agent further includes proof of the validity of the certificate associated with a statement from the client agent relating to the request. Selection method, according to claim 8, further comprising a verification of the proof of validity of at least one certificate received by means of an identification information of a decryption algorithm present in the message received from the client agent. Selection method, according to any one of the preceding claims, wherein the comparison between the vector representing the source data, said source data being used for the development of a response to the query, and further representing a constraint on the use of the source data, and the calculated vector, is carried out in accordance with a filtering operation associated with the constraint on the use and / or the client agent statement relating to the query. Selection device (Disp Sel) of at least one source data, said source data being adapted for the development of a response to a request from a client agent associated with a large language model, said selection device (Disp Sel) comprising a processor (PROC 402) coupled to a memory (MEM 401) in which are stored program instructions (PGR 403) intended to be executed by the processor for a process comprising: - a reception (E30, F30) from the client agent of a request and a certificate associated with a declaration from the client agent relating to the request, - a calculation (E40, F40) of a vector representing the request and the declaration received, - a selection (E40, F40) of at least one source data by means of a comparison between a vector representing the source data, used for the development of a response to the request, and also representing a constraint on the use of the source data, and the calculated vector. A selection system for at least one source data, said source data being adapted for the development of a response to a request from a client agent associated with a large language model, said selection system comprising: A selection device (Disp Sel) according to claim 11, A client agent (Agt), adapted to communicate with the selection device and configured to issue the request, and a certificate associated with a statement from the client agent relating to the request. Product program comprising program code instructions for implementing a selection process according to any one of claims 1 to 10 when executed by a processor. Recording medium readable by a selection device on which the program according to claim 13 is recorded.

Citation Information

Patent Citations

  • Large model cue word verification processing method and device

    CN117131529A

  • Techniques for securing, accessing, and interfacing with enterprise resources

    WO2024091682A1