METHOD FOR TESTING A LARGE LANGUAGE MODEL IMPLEMENTED WITHIN A CONVERSATIONAL AGENT

The method for testing large language models in conversational agents addresses the challenge of maintaining robustness by generating varied tests and refining responses, enhancing their performance over time.

FR3168282A1Pending Publication Date: 2026-05-08GISKARD AI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
GISKARD AI
Filing Date
2024-11-07
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Conversational agents struggle to maintain robustness over time due to variations in new concepts arising from data sources, requiring improved methods to evaluate and refine their performance.

Method used

A method for testing large language models within conversational agents involves generating varied test messages using multiple language models to simulate different contexts, calculating error and conformance indicators, and automatically refining responses based on these tests.

Benefits of technology

Enhances the robustness of conversational agents by identifying areas of validity and enabling continuous refinement, ensuring reliable responses to evolving data sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000025_0000
    Figure 00000025_0000
  • Figure 00000025_0001
    Figure 00000025_0001
Patent Text Reader

Abstract

METHOD FOR TESTING A LARGE LANGUAGE MODEL IMPLEMENTED WITHIN A CONVERSATIONAL AGENT A computer-implemented method for testing the performance of a large language model (LLMt) comprising: Receiving a dataset (ENS1) from at least one data source (Si) over a given time period; Generating (GEN1) at least one first message (M1) from the received dataset (ENS1) and by applying at least one first large variation language model (LLM1); Generating (GEN2) a plurality of messages (M1) by applying at least one large variation language model; Generating (GEN3) a plurality of message exchanges ({SEQ1 k}k ∈ [1 ; N]) by applying a large language model to be tested (LLMt); Calculating (TEST1) a first error indicator (IND1); Calculation (TEST2) of a second conformity indicator (IND2). Figure for the summary: Fig.1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: METHOD FOR TESTING A LARGE LANGUAGE MODEL IMPLEMENTED WITHIN A CONVERSATIONAL AGENT Scope of the invention

[0001] The field of the invention relates to that of automatically generated tests of conversational agents implementing large language models in order to strengthen their robustness over time. State of the art

[0002] Currently, solutions exist for interacting with a conversational agent, called a "chatbot" in the English-language literature, to assist a human in gathering specific information. These conversational agents require precise contextual information depending on how they are intended to be used. A known problem is the ability of a conversational agent to offer a persistent service over time that can take into account variations related to new concepts, new concepts arising from news published on data sources accessible via a data exchange network such as the internet. Summary of the invention

[0003] According to a first aspect, the invention relates to a computer-implemented method for testing the performance of a large language model implemented within a conversational agent comprising: • Receiving a dataset from at least one data source, said data corresponding to an encoding of discrete symbol sequences in natural language, the dataset being previously extracted from at least the data source automatically at a given frequency and from a selection of a natural language; • Generation of at least one first test message from the received data set and by applying at least one first large language model configured with a main context including a definition of a language and a given instruction specific to the data source; • Generation of a plurality of variation sequences by applying at least a plurality of large variation language models configured from a plurality of secondary contexts allowing to generate on the one hand the variations of the first test message and the associated responses generated by at least one large language model; • Generation of a plurality of message sequences by applying a large language model to be tested, a sequence comprising an input and the corresponding generated output of a large language model; • Calculation of a first error indicator evaluating a set of error criteria by comparing the sequences produced by the large language model to be tested and the variation sequences; • Calculation of a second conformance indicator from a conformance domain comprising conformance rules defining sets of validity of the sequences produced by the large language model to be tested.

[0004] One advantage of the invention is that it allows the robustness of a chatbot to be evaluated over time by automatically generating tests. These tests make it possible to diagnose and identify areas of validity for a chatbot. Furthermore, the tests allow for the redefinition or refinement of a chatbot's prompt so that it can automatically generate reliable responses.

[0005] According to one embodiment, the responses of the variation sequences are generated by the plurality of large variation language models. One advantage is to extend the test domain.

[0006] According to one embodiment, the responses of the variation sequences are generated by at least one large evaluation language model taking as input a variation produced by a large variation language model and producing as output an associated response.

[0007] According to one embodiment, the first message is a first sequence of symbols in natural language defining a question in a natural language.

[0008] According to one embodiment, the process is executed at a predefined frequency on a set of predefined sources.

[0009] According to one embodiment, frequency is used to select data from data sources published from a given date.

[0010] According to one embodiment, each source is associated with a given frequency.

[0011] According to one embodiment, the data source is selected prior to starting from a uniform resource locator within a data network and an organization name allowing selection of a subset of the data accessible from the uniform resource locator.

[0012] According to one embodiment, the data reception originates from one of the data sources characterized by: • A data source accessible from a social network via an authentication process; • A data source defining comments or opinions from a plurality of individuals, • A source of freely accessible information data. • A data source defining one or more database(s) internal to an organization such as a product or item database, a service database, or even a database of vehicles in stock; • A data source defining conversational agent(s) data recorded in production or in a test environment, • A data source defining electronic documentation.

[0013] One advantage is to allow the generation of tests whose heterogeneity is obtained through the diversity of the selected sources.

[0014] According to one embodiment, the method includes generating an alert when at least one first error indicator and / or a second conformity indicator is generated.

[0015] According to one embodiment, the method includes configuring access to a data source.

[0016] According to one embodiment, the method includes executing a large source data processing language model in order to filter, format and / or normalize the datasets extracted from the data sources.

[0017] One advantage is to obtain messages simulating a type of question that is likely to arise, for example by automatically introducing a personal pronoun of the type "I".

[0018] According to one embodiment, each exchange sequence comprises a sequence of symbols in natural language defining a question, said sequence of symbols in natural language being produced from at least one first large model of variation language and an answer produced by using a large model of language to be tested.

[0019] According to one embodiment, the main context of a first major variation language model includes the definition of a domain associated with a lexical field or a set of keywords. The context is, for example, a prompt of an LLM.

[0020] According to one embodiment, a first major variation language model includes a configured context allowing the automatic generation of a sequence of discrete symbols in a natural language from one or more paraphrases of the first message.

[0021] According to one embodiment, a second major variation language model includes a configured context enabling the automatic generation of a sequence of discrete symbols in a natural language from a translation of the first message into another natural language.

[0022] According to one embodiment, a third major variation language model includes a configured context allowing automatic generation of a sequence of discrete symbols in a natural language from an exaggeration of the first message.

[0023] According to one embodiment, a fourth major variation language model includes a configured context enabling the automatic generation of a sequence of discrete symbols in a natural language from a change in tone of the first message.

[0024] According to one embodiment, a fifth major variation language model includes a configured context for automatically generating a sequence of discrete symbols in a natural language from an introduction of at least one insult in the first message.

[0025] According to one embodiment, a sixth major variation language model includes a configured context for automatically generating a sequence of discrete symbols in a natural language from an introduction of at least one error in the first message, said error being, for example, a spelling or grammatical error in a natural language.

[0026] According to one embodiment, the method comprises generating a plurality of message variations for each major variation language model.

[0027] One advantage is to allow the generation of a very wide variety of unit tests of a test LLM based on variations of the message with a large heterogeneity and a large diversification of the context and taking into consideration data evolving over time.

[0028] According to one embodiment, the conformance domain is defined from the response of a third major language model configured from a context defining a conformance domain.

[0029] According to one embodiment, the conformance domain is defined from a set of rules defining predefined sets of validity of sequences of symbols in natural language and / or sets of invalidity of sequences of symbols in natural language.

[0030] According to one embodiment, the set of rules includes the specification of a response language, the specification of a topic to be excluded from the response field, or that a link to a resource in a data network must be present in a given response type.

[0031] According to one embodiment, the set of rules defining invalidity sets includes a knowledge base listing a set of themes, categories, labels, or keywords, each defining a sequence of discrete symbols in a natural language and possibly variations of this sequence.

[0032] According to one embodiment, the compliance domain is generated partly automatically from the organization name, rules and main context of the first major language model.

[0033] According to one embodiment, an error criterion of the first error indicator includes a check that a set of common concepts are present on the one hand in the response produced from the produced variation sequence and on the other hand in the response produced by the large language model to be tested to which a variation of the first message has been provided.

[0034] According to one embodiment, when an error indicator and / or a conformity indicator is generated, a notification is automatically issued to a remote server or a memory resource of equipment on which the process is executed.

[0035] According to one embodiment, when an error indicator and / or a conformity indicator is generated, an error counter is generated to produce an evaluation of the conversational agent over a given period.

[0036] According to another aspect, the invention relates to a system comprising an electronic user terminal having a user interface, at least one data server hosting all or part of a first data source, a second data server comprising at least one computer and a memory in which the large language model to be tested is executed and at least a third data server comprising means for executing a variation language model and comprising a computer for executing the steps of the method of the invention.

[0037] According to one embodiment, the system includes a fourth data server comprising at least one computer and a memory within which the large evaluation language model is executed. Brief description of the figures

[0038] Other features and advantages of the invention will become apparent from the following detailed description, with reference to the accompanying figures, which illustrate: • [Fig.1]: a method of implementing the steps of the process of the invention; • [Fig.2]: an embodiment of a system of the invention.

[0039] A "large language model" (LLM) is a machine learning model with a large number of parameters. In some implementations, these are deep neural networks trained on large quantities of unlabeled text using self-supervised or semi-supervised learning.

[0040] A "large language model to be tested" is an LLM used by a conversational agent whose limits, side effects, and robustness in producing consistent, true, and unbiased responses over time are to be tested. It is denoted LLMt.

[0041] A "large variation language model" is a language model configured to produce variations of a text based on a given parameterization. It is denoted LLMV k according to the i-th variation language model with a given parameterization. When the text is a question, LLMV k produces the variation of the question with respect to an original question and optionally the answer to the variation. In the latter case, an LLMV k produces a sequence of variations SEQvar i comprising a first message defining an input of the large variation language model LLMvk, for example in the form of a question, and comprising an answer to the first message produced by LLMV k.

[0042] A "large evaluation language model" or "large reference language model" is an LLM configured to compare the result produced in response to a query with the result produced when executing a large language model under test, LLMt, subjected to the same query / input. This large reference / evaluation language model is denoted LLMe. In one embodiment, a large evaluation model LLMe may be of the same type as a similar large variation language model LLMVk used to generate the variations VAR; according to a given criterion that produces the inputs and outputs of LLMk from a message Mb, however, the prompt of a large evaluation language model LLMe differs from the prompt of a large variation language model LLMvk.

[0043] In this latter case, the method of the invention makes it possible to compare the result of an execution of the model under test LLMt with what is generated by the variation LLMvi. The comparisons can only be based on the responses produced by each model LLMt, respectively LLMV k, from different or the same inputs.

[0044] A "large source data processing language model" is called an LLM configured to process data extracted from data sources by homogenizing, normalizing, filtering or formatting the data according to a given parameterization of the LLM context.

[0045] The "context" of an LLM is a "prompt" that allows for specifying or parameterizing a textual description of the task that a machine learning algorithm, such as an LLM, must perform. A prompt can be accessed from a user interface or directly programmed in a programming language. In the context of the invention, a "configured LLM" refers to an LLM for which the prompt or context is parameterized or specified.

[0046] Figure 1 illustrates an embodiment of the method of the invention. The example is detailed for an organization with a designation and, for example, an identifier. In the example, the organization uses a chatbot implemented within a digital service offered to multiple users. The service is accessible from at least one data server SERVi. Users access the service from a PCi client, which can be an electronic terminal such as a PC, tablet, smartphone, or mobile phone.

[0047] An organization can be a company, an association, a natural person, a laboratory or any other form of community of users who have jointly defined a digital service accessible from a NETb data network such as the internet.

[0048] According to various examples, the organization could be a public service, an insurer, a bank, a school, a travel agency, etc., offering a digital service accessible from a NETb data network. The digital service could be open, freely accessible via a URL, meaning in Anglo-Saxon terminology "Uniform Resource Locator" and more generally designating a uniform resource locator. Such access allows access to data within a NETi data network. According to another example, the digital service could be closed and made available online within a private data network such as an intranet or a network requiring user authentication.

[0049] The method aims, for each organization, and therefore for each identifier representing an organization, to retrieve data from a set of data sources Si in order to test the performance and robustness of a conversational agent LLMt. By "performance of a conversational agent," we mean its ability to produce structured, coherent, true responses, or responses addressing a given question, or sometimes even to redefine the scope or domain in which it is capable of responding so that a user reformulates a question, etc.

[0050] It is worth recalling that a conversational agent is generally implemented using a Learning Management System (LMS) configured in a specific domain. For this purpose, a context, also called a "prompt," allows, for example, the enrichment of the machine learning algorithm's training domain. The method of the invention makes it possible to update the domain over time using a set of data sources that evolve over time and to test the domain so that it is able to maintain or improve its performance in a given domain over time. Receiving data from sources

[0051] The method of the invention makes it possible to automatically generate tests that cover specificities of a given application domain that may evolve over time. One of the objectives of the invention is to implement automated active monitoring of contexts via heterogeneous sources that evolve over time.

[0052] To this end, the process of the invention includes a first step AQCi which corresponds to the reception of a set of ENSi data from at least one Si data source.

[0053] The ENSi dataset comprises a set of sequences of symbols in natural language. These sequences can correspond to a word, a digit or a number, a sentence, a paragraph containing a plurality of sentences, a text from a document such as a file in .docx, .pdf format, or any other text format, or a web page for example in html, xml format or any other format allowing data to be contained and structured and displayed within a browser.

[0054] Sources Si can correspond to different data containers accessible from different digital resource locators, denoted URLs, within a data network NETb such as the internet. The i-th data source Si is accessible from a URL. In one example, a source Si may be accessible only by means of a resource locator URL. In another example, sources Si are accessible by means of a resource locator URL and at least one other piece of data. In a first example, authentication data is used to access a data source Si. The authentication data may be, for example, a login and password, two-factor authentication, or any other means of identification or authentication. In one embodiment, the source Si is a URL of a web page of an organization.

[0055] In one embodiment, the data source can be an internal organizational database such as a product or item database, a service database, electronic documentation, or a vehicle inventory database. Any other type of database can be configured to define a usable data source.

[0056] According to another embodiment, the data source can correspond to conversational agent(s) data recorded in production or in a test environment. Indeed, this makes it possible to use the topics actually generated by the conversational agent's users as a source of data generation.

[0057] According to one example, the data source Si corresponds to a part of the data accessible from a URL resource locator; via a NETb data network. The part may correspond to a set of comments or reviews of a web page, a title of an article, an article.

[0058] According to one embodiment, a given configuration allows the parameters to be set for the data extracted from a given source Si. Thus, if a plurality of data sources Si is used in the execution of the method of the invention, then a plurality of parameters can be implemented so that the data ENSi from each source Si is received. The reception, collection, and recording of the data can be carried out within a data server SERVi. In one example, each parameter includes at least the name of an organization, a URL, and a frequency allowing the definition of a period for extracting the data from the data source.

[0059] In one embodiment, the data is received by at least one memory of an electronic terminal such as a computer or a server. In various embodiments, the equipment that receives the data from each source Si is equipment that includes at least one memory and a computer. The data ENSi is received and recorded for processing by a computer. In one example, a database is implemented and used to store the data ENSi from the different sources Si in an ordered manner. In another example, the database allows the data to be stored and ordered chronologically in order to verify whether the data from a source S has changed over time and, if so, to compare the extent to which it has changed between two different times.

[0060] According to one embodiment, the ENSi data are received at regular time intervals, for example according to a predefined period, a period on the order of the hour, day, week or month, or even the year can be defined.

[0061] According to one embodiment, the data reception is preceded by a step initiated by a given piece of equipment that generates requests to the various sources Si to extract data present in each source Si. In an example where a server is configured to retrieve data from the various sources Si, requests are defined to periodically retrieve data from different sources distributed on a NETi data network and accessible from it.

[0062] According to one embodiment, an analysis function is executed on the SERVi server to analyze whether a set of ENSi data is recorded and used by the process of the invention or whether it is not recorded. Certain criteria can be defined to parameterize the analysis function. For example, the analysis function compares subjects, themes, keywords, concepts or calculates a similarity score between two sets retrieved at two different dates or between the data set and a reference set.

[0063] One interest is to retrieve data from a set of sources Si which are heterogeneous in order to generate different varieties of questions defining inputs to the conversational agent in order to test its robustness to different variations likely to occur over time according to themes and topics related to current events for example.

[0064] According to one embodiment, a set of queries are generated in order to retrieve a wide variety of datasets from different sources.

[0065] It is understood that textual data from comments or reviews will not be formulated in the same way as a website publishing institutional information or online editorial journals. Differences in tone, different language registers, including informal and formal language, and the presence or absence of spelling mistakes allow for the generation of different inputs, enabling broad testing of the large language model under test, LLMt. Finally, retaining data showing topic updates based on the publication date allows for the continuous updating of the context and thus the updating of all the data defining the prompt of a conversational agent, and more generally, verification that the LLMt conversational agent has access to up-to-date data. Step to generate a test message

[0066] The method includes a second step for generating at least one first message Mi from the received data set ENSi and by applying at least one first large language model LLMi configured with a main context CTi comprising a language definition and a description defining an instruction. In the most general case, the method allows the generation of a plurality of messages Mi so that different tests of the large language model to be tested, LLMt, can be performed.

[0067] In this step, the method of the invention implements a machine learning algorithm to generate a question directly usable by the LLMt chatbot to be tested. One advantage of using a large number of heterogeneous sources is to produce a variety of test questions to test the LLMt under test. The LLMt is configured to generate questions for the LLMt.

[0068] In order to generate questions to effectively test the LLMt conversational agent to be tested, a context is defined allowing the question to be generated to be formatted in a given domain and language.

[0069] The language allows in particular to extract data in the configured language or to translate the content of the extracted and received ENSi data to test the LLMt in the determined language.

[0070] The domain may refer to a general field such as a scientific, economic, or political field, or a field related to a given profession such as banking, crafts, perfumery, insurance, or automobiles, etc. The domain may also refer to an activity of an organization, such as a retail activity, a training activity, a service activity, etc.

[0071] One advantage is that it allows the question to be customized according to a given field. For example, the field could be "after-sales service for cosmetic products" or "assistance to people who have suffered an accident" or even the field of "pre-medical diagnosis to direct an individual to the appropriate emergency service".

[0072] In these cases, the extracted ENSi data is used to generate a domain-specific input using the LLMi.

[0073] For example, if a data source Si specifies that "insurance reimbursement rates for a drug have decreased from 100% to 50%" and the domain is "assistance to persons who have suffered an accident," the LLMi can generate a question such as: "Can I receive 100% reimbursement for my care in the event of a workplace accident?" Thus, the LLMi is configured to generate first-person responses applied to the case of assistance or care, taking into account the data produced by the given data source. Generation of variations

[0074] According to one embodiment, at least one machine learning algorithm, such as a large language model LLM, is configured to generate variations of the test message Mb. The generated variations are denoted VAR;. One advantage of generating variations VAR; is to allow the test domain of a conversational agent to be extended, said conversational agent implementing a large language model LLMt to be tested.

[0075] The variations correspond to variations of the message Mb. According to one embodiment, each major variation language model LLMvi generates a sequence SEQvarî comprising the variation VAR; of the message Mi and the associated response.

[0076] According to one embodiment, a plurality of large language models {LLMvi}ieli;Nj are configured to generate variations of the test message Mb. In this example, N models are implemented. One advantage of this solution is to configure VAR variations; according to different criteria in order to generate the most comprehensive test domain possible.

[0077] Various examples are described, however the invention is not limited to the examples cited.

[0078] According to a first example, a first large language model LLMvi is configured to generate a plurality of reformulations or paraphrases of the message Mb. This model is configured with a prompt or context that facilitates the production of new messages Mi generated from a message Mi by varying the words in the discrete symbol sequence while preserving the meaning. To this end, the modification and replacement of terms with synonyms can be carried out within the message Mi, or reformulations of expressions or equivalents can be produced.

[0079] According to a second example, a second major language model LLMv2 is configured to generate a plurality of VAR variations based on total or partial translations of the original message Mi. This model is configured with a prompt or context that facilitates the production of new messages Mi from a given message Mi by varying the translations of certain expressions or even the entire set of words in the discrete symbol sequence forming the message Mi, while preserving its meaning. To this end, different languages ​​can be configured to produce variations corresponding to a plurality of translations of all or part of the message Mi into a plurality of languages. For example, mixtures of translations of certain parts of the same message are produced to generate a message containing different portions of text expressed in different languages.

[0080] According to a third example, a third major language model LLMV 3 is configured to generate a plurality of VAR variations based on exaggerations of the original message Mi. This model is configured with a prompt or context that facilitates the production of new messages Mi generated from a message Mi by varying the exaggerations of certain words, expressions, or even phrases—see the set of words in the discrete symbol sequence forming the message Mb. To this end, certain synonyms or equivalents that are too close to terms in the original message Mi are not retained in favor of replacing terms that exaggerate a characteristic defined by the meaning of a word or that correspond to an emphasis of a term or group of words. The exaggeration can also be applied to a number, a value, an estimate, a percentage, a statistic, or any other quantity expressed in a message.These variations correspond to a plurality of sequences that can modify the meaning of the message. Mi or at least to vary it around the meaning defined by the first message Mb The third major language model LLMv3 can thus generate a modification of the following sentence "I have a problem with my computer" into "I have a big bug with my PC".

[0081] According to a fourth example, a fourth major language model LLMv4 is configured to generate a plurality of VAR variations based on tone changes in the original message Mi. This model is configured with a prompt or context that facilitates the production of new messages Mi generated from a message Mi by varying the tones of certain groups of words, expressions, or even sentences, or the entire set of words in the discrete symbol sequence forming the message Mb. The tone changes can convey exasperation, an order, irritation, anger, or even a calm masking an individual's restraint, etc.

[0082] To this end, LLMv4 allows modification of verb tenses, pronouns and intonations, as well as the interrogative or exclamatory form of groups of words in the sequence of discrete symbols forming the message Mi

[0083] The fourth major language model LLMV 4 can thus generate a modification of the following sentence "can you help me book a train for tonight to go to Nantes from Paris? Thank you in advance" into "give me the timetables for Paris-Nantes tonight or I'm disconnecting forever".

[0084] According to a fifth example, a fifth major language model LLMv5 is configured to generate a plurality of VAR variations based on modifications to the message Mi or the introduction into the original message Mi of words from a given register, such as vulgar words or insults. This model is configured with a prompt or context to facilitate the production of new messages Mi from a message Mi by varying the vocabulary of certain word groups, expressions, or sentences within the message Mb. Modifications to the message Mi can include replacing words from a given register with words from another register or introducing them without replacement.

[0085] The fifth major language model LLMV 5 can thus generate a modification of the following sentence "I would like to know the life insurance return rates of the contracts taken out during the last 3 months" into "Give me the life insurance return rates, you madman".

[0086] Combination of large language models to produce variations

[0087] According to one embodiment, the VAR variations are produced from two LLMvk variation language models assembled in cascade. More generally, A plurality of large variation patterns can be configured in cascade to produce a wide variety of heterogeneous variations.

[0088] According to this design, the output of a large variation language model can be used to define an input for another large variation language model LLMvk. This configuration allows for enriching the variations of the original message Mb. By way of example, the second variation language model LLMV2 and the third variation language model LLMvî can be implemented in cascade so that exaggerations of the partial or total translations of the message Mi are produced. Variations produced by other algorithms

[0089] According to one embodiment, an encoding using unconventional characters such as symbols can be used to generate variations. For example, the term "Hello" can be encoded as follows:

[0090] According to, for example, an encoding allowing Caesar ciphers, hexadecimal, "leetspeak", also called: "elite language" and corresponding to a writing system using ASCII alphanumeric characters in a way that is not easily understood by the novice in order to distinguish itself from them, can be used to generate variations of the message.

[0091] According to another example, an algorithm for executing a "tactical" type function can be implemented in the method of the invention. Such a tactical function consists of using a predefined description accompanied by examples to automatically generate variations. According to one example, the method implements a variation LLM to modify the message Mi so that it uses the selected tactic.

[0092] According to one example, these tactics can be configured from the scientific literature, and possibly enriched by data characterizing identified vulnerabilities of an LLM. Recording of variations and filtering

[0093] The set of VAR variations produced by the message Mi can be stored in memory or a database for later use during the testing of the LLMt under test. In one embodiment, filtering is performed to retain the most distinctive variations or to discard certain variations when a variation is too close to the original message Mi. In another embodiment, a predefined number of variations is configured to limit the computational resources required during the testing of the large language model under test, LLMt.

[0094] The filter may, for example, consist of comparisons of the different variations and a measurement of a similarity indicator, for example, based on the number of different discrete symbols from one variation to another. Other possibilities may to be implemented to filter out some of the variations in order to keep only a limited number of variations for the testing phase.

[0095] Generation of responses from the chatbot to be tested

[0096] According to one embodiment, the method of the invention comprises a step of transmitting a plurality of VAR variations to a conversational agent under test, LLMt, to produce a plurality of responses from the conversational agent under test, LLMt. Each response produced by the conversational agent under test, LLMt, constitutes a response that is subject to a unit test using the method of the invention. The tests can be performed sequentially or in parallel so that a plurality of instances of the conversational agent are produced.

[0097] The method of the invention includes a generation step, denoted GEN3 in [Fig.1], of a plurality of message exchanges comprising the variation VAR; and the associated response REP; by application of a large language model to be tested LLMt.

[0098] The set formed by a VAR variation; and the response produced by the conversational agent LLMt to be tested is denoted a SEQ sequence;.

[0099] The method of the invention includes a test step aimed at producing two indicators denoted INDi and IND2 allowing the LLMt conversational agent to be tested.

[0100] A first generated INDi indicator is a factual error indicator in the sequence. The second generated IND2 indicator is a conformity indicator of the sequence. Error indicator

[0101] According to one embodiment, the error indicator INDi aims to measure the extent to which the conversational agent produces an erroneous, false or inconsistent response or a response produced by a hallucination or confabulation.

[0102] To this end, a first large evaluation language model LLMe can be used so as to verify the outputs produced by this large evaluation model LLMe with the outputs produced by the large language model under test LLMt from the same input variations. Thus, in this case, LLMvk produces a sequence comprising a variation and a response; the sequence is denoted SEQvarî, and the response is compared with that of the model under test LLMt.

[0103] The tests aimed at producing or not producing an error indicator are noted TESTi on [Fig.1].

[0104] In the latter case, the prompt or context of such a large evaluation language model LLMe can be predefined. According to one embodiment, a plurality of large evaluation language models LLMe, and therefore where appropriate a plurality of large variation language models LLMvk, can be configured to test the large language model to be tested LLMt according to different criteria.

[0105] According to one embodiment, in order to produce or not an error indicator INDi, the method of the invention includes a step of verifying that a set of common concepts are present in the response produced by the large language model to be tested LLMt and in the large variation language model LLMvk.

[0106] The set of concepts can be predetermined or generated from a given semantic domain, or generated by a large domain language model LLM2 to which a variation of the first message Mi has been provided, the latter having been configured with a given prompt or context in order to provide a list of semantic domains or fields. One advantage of this last solution is having an LLM trained on different data than the large language model under test, LLMt.

[0107] Thus, it is possible to verify that a set of expected concepts are present in the response of the large language model to be tested LLMt.

[0108] According to another example, the error indicator INDi is produced by calculating a similarity index between a response produced by a large variation model LLMvk and the responses produced by the large language model under test LLMt from the same variation considered as input to both models LLMvk and LLMt. A large evaluation language model LLMe can be configured to produce the similarity index from a comparison performed between the two outputs of the two models LLMvk and LLMt.

[0109] According to another embodiment, a similarity score can be calculated from a similarity score based on the differences and similarities of the two character strings produced or more generally the two sequences of discrete symbols in natural language produced by the two models LLMvk and LLMt.

[0110] A similarity score allows us to assess how similar or different the responses produced by the two models are. An error indicator INDi can be generated when a threshold for the similarity index is exceeded.

[0111] Other comparable ones can be configured so as to produce an INDi error indicator for each response produced from the LLMt from a VAR variation;.

[0112] The method of the invention subsequently enables the production of an automatic action based on the generation of the error indicator INDi. Conformity indicator

[0113] According to one embodiment, the conformance indicator IND2 aims to measure the extent to which the conversational agent LLMt produces a response REP; conforming to a predefined domain, called the conformance domain DOMc. The conformance domain DOMc can be defined by a set of rules or a set of reference responses produced by another large language model called a conformance model.

[0114] The tests aimed at producing or not producing an error indicator are noted TEST2 on [Fig.1].

[0115] According to one example, a set of rules includes the specification of a response language, the specification of topics to be excluded from the response field, or that a link to a resource in a data network must be present in a given response type.

[0116] According to one embodiment, the DOMc conformance domain is defined by a set of rules generated by a conformance LLM configured to delimit a response domain. According to another example, the conformance domain is defined by the semantic field or a set of concepts produced in the responses by a conformance LLM.

[0117] According to another example, the DOMc conformance domain is defined from a set of RGLi rules defining sets of validity of the response produced by a large language model.

[0118] According to one example, this set of RGLi rules defines invalidity sets comprising a knowledge base listing a set of themes, categories, labels, or keywords, each defining a sequence of discrete symbols in a natural language and possibly variations of that sequence. Generation of actions, alerts.

[0119] According to one embodiment, when at least one error indicator INDi, IND2, is generated, the method of the invention includes a step for automatically performing an action. In a first example, the action corresponds to the generation of a notification, such as an alert.

[0120] In one embodiment, notifications are generated as soon as at least one error indicator or at least one conformity indicator is generated. In another example, data providing the number of errors and / or the statistics on the occurrence of these errors is notified.

[0121] According to one example, a notification is sent to a server to administer and test the conversational agent.

[0122] According to another example, the action is an automated response produced by the chatbot. The latter can, for example, generate a response such as: "An error has been detected in our conversation, could you rephrase your question?" for example within a user interface.

[0123] According to another example, the automatically generated action is an update to take into account new data sources to regenerate or update the chatbot prompt.

[0124] According to another example, the action is the generation of a command to trigger the execution of a computer program aimed, for example, at suspending the production of the conversational agent or switching back the assistance function to a human or any other software function modifying the operation of the conversational agent. System

[0125] Figure 2 represents an example of infrastructure enabling the implementation of method of the invention. A SERVI server allows the calculation of error and compliance indicators.

[0126] According to an example of the architecture of the invention, a second server is a server hosting at least one data source Si and the SERV3 server is a server enabling the values ​​of the compliance and error indicators to be returned over time, said SERV3 server being accessible from an administration console of the solution.

Claims

1. Demands A computer-implemented method for testing the performance of a large language model (LLMt) implemented within a conversational agent comprising: • Receiving a dataset (ENSi) from at least one data source (Si), said data corresponding to an encoding of discrete symbol sequences in natural language, the dataset (ENSi) being previously extracted from at least the data source (Si) automatically according to a given frequency (TEMPO and from a selection of a natural language (LGi); • Generation (GENi) of at least one first test message (MO) from the received dataset (ENSi) and by application of at least one first large language model (LLMi) configured with a main context (CTi) including a definition of a language and a given instruction specific to the data source (Si); • Generation (GEN2) of a plurality of variation sequences (SEQvar i) by applying at least a plurality of large variation language models (LLMvk) configured from a plurality of secondary contexts (CT2 i) allowing to generate on the one hand the variations (VAR;) of the first test message (Mi) and the associated responses generated by at least one large language model; • Generation (GEN3) of a plurality of message sequences (SEQ;) by application of a large language model to be tested (LLMt), a message sequence comprising an input and the corresponding generated output of a large language model; • Calculation (TESTi) of a first error indicator (INDi) evaluating a set of error criteria by comparing the sequences produced by the large language model to be tested (LLMt) and the variation sequences (SEQvarD; • Calculation (TEST2) of a second compliance indicator (IND2) from a compliance domain (DOMc) including conformance rules (RGLi) defining sets of validity of the sequences produced by the large language model to be tested (LLMt), • Generation of an alert when at least one first error indicator and / or a second conformance indicator is generated.

2. Method according to claim 1 characterized in that the responses of the variation sequences (SEQvarî) are generated by the plurality of large variation language models (LLMvk).

3. Method according to claim 1 characterized in that the responses of the variation sequences (SEQVARi) are generated by at least one large evaluation language model (LLMe) taking as input a variation produced by a large variation language model (LLMV k) and producing as output an associated response.

4. A method according to any one of claims 1 to 3 characterized in that the predefined frequency is used to select data from data sources (Si) published from a given date.

5. A method according to any one of claims 1 to 4 characterized in that the data source (Si) is pre-selected from a uniform resource locator (URLi) within a data network (NETi) and an organization name enabling selection of a subset of the data accessible from the uniform resource locator.

6. A method according to any one of claims 1 to 5 characterized in that the data reception (ENS1) originates from one of the data sources characterized by: • A data source accessible from a social network via an authentication method; • A data source defining comments or opinions from a plurality of individuals; • A data source of freely accessible information; • A data source defining one or more internal databases of an organization such as a product or item database, a service database, or a database of vehicles in stock; • a data source defining conversational agent(s) data recorded in production or in a test environment, • a data source defining electronic documentation.

7. A method according to any one of claims 1 to 6 characterized in that it comprises the execution of a large source data processing language model (LLMs) in order to filter, format and / or normalize the datasets extracted from the data sources (Si).

8. A method according to any one of claims 1 to 7 characterized in that each message sequence (SEQ;) comprises a sequence of natural language symbols defining a question, said sequence of natural language symbols being produced from at least one first large variation language model (LLMvk) and a response produced by use of a large language model to be tested (LLMt).

9. A method according to any one of claims 1 to 8 characterized in that the main context (CTi) of a first major variation language model (LLMi) includes the definition of a domain associated with a lexical field or a set of keywords.

10. A method according to any one of claims 1 to 9 characterized in that a first large variation language model (LLMvi) includes a configured context for automatically generating a sequence of discrete symbols in a natural language from one or more paraphrases of the first message (MJ).

11. A method according to any one of claims 1 to 10 characterized in that a second major variation language model (LLMV 2) includes a configured context enabling the automatic generation of a sequence of discrete symbols in a natural language from a translation of the first message (MJ) into another natural language.

12. A method according to any one of claims 1 to 11 characterized in that a third major variation language model (LLMV 3) includes a configured context for generating automatically generates a sequence of discrete symbols in a natural language from an exaggeration of the first message (Mi).

13. A method according to any one of claims 1 to 12 characterized in that a fourth major variation language model (LLMV 4) includes a configured context enabling the automatic generation of a sequence of discrete symbols in a natural language from a change in tone of the first message (Mi).

14. A method according to any one of claims 1 to 13 characterized in that a fifth major variation language model (LLMV 5) includes a configured context for automatically generating a sequence of discrete symbols in a natural language from an introduction of at least one insult in the first message (MJ.

15. A method according to any one of claims 1 to 14 characterized in that a sixth major variation language model (LLMV 6) includes a configured context for automatically generating a sequence of discrete symbols in a natural language from an introduction of at least one error in the first message (Mi), said error being, for example, a spelling or grammatical error in a natural language.

16. A method according to any one of claims 1 to 15 characterized in that it comprises the generation of a plurality of message variations (MJ) for each major variation language model (LLM).

17. A method according to any one of claims 1 to 16 characterized in that the conformance domain (DOMc) is defined from a set of reference responses produced by another large conformance language model configured from a context defining a conformance domain.

18. A method according to any one of claims 1 to 16 characterized in that the conformance domain (DOMc) is defined from the conformance rules (RGLi) defining predefined natural language symbol sequence validity sets and / or predefined natural language symbol sequence invalidity sets.

19. A method according to any one of claims 1 to 18 characterized in that the conformity rules (RGLi) comprise the specification of a response language, the specification of the topic to be excluded from the response field or that a link to a resource in a data network (NETi) must be present in a given response type.

20. A method according to any one of claims 1 to 19 characterized in that the conformity rules (RGLi) defining invalidity sets include a knowledge base listing a set of themes, categories, labels or keywords each defining a sequence of discrete symbols in a natural language and possibly variations of this sequence.

21. A method according to any one of claims 1 to 20 taken in combination with claim 5, characterized in that the compliance domain (DOMc) is generated partly automatically from the organization name, compliance rules (RGLi), and main context (CTi) of the first major language model (LLMi).

22. h Method according to any one of claims 1 to 21 characterized in that an error criterion of the first error indicator (INDi) includes a check that a set of common concepts are present on the one hand in the response produced of the variation sequence (SEQVarjÛ produced and on the other hand in the response produced by the large language model to be tested (LLMt) to which a variation of the first message (MJ) has been provided.

23. A method according to any one of claims 1 to 22 characterized in that when an error indicator and / or a conformity indicator is generated, a notification is automatically issued to a remote server or a memory resource of equipment on which the method is executed.

24. A method according to any one of claims 1 to 23 characterized in that when an error indicator and / or a conformity indicator is generated, an error counter is generated to produce an evaluation of the chatbot over a given period.

25. A system comprising a user electronic terminal (PCi) having a user interface, at least one data server (SERVi) hosting all or part of a first data source (Si), a second data server (SERV2) having at least one computer and a memory within which the large model

26. language to be tested (LLMt) is run and at least a third data server (SERV3) comprising means for running a variation language model (LLMvi) and comprising a calculator for running the steps of the method of any one of claims 1 to 24. System according to claim 25 characterized in that it comprises a fourth data server (SERV4) including at least one computer and a memory within which the large evaluation language model (LLMe) is executed and including a computer for executing the steps of the method of claim 3.