Automated Testing Method for Agent Language Ability Based on the Self-Evolution of Multilingual Models

Through the automated test method of agent language ability based on self-evolution of multilingual models, multiple voting language models are used to evaluate the agent's output results, the problems of low testing efficiency and inaccurate results in the existing technology are solved, and more efficient and accurate testing is achieved.

CN119357069BActive Publication Date: 2025-05-30BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411933345.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-30
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

The existing technology has problems of low efficiency and inaccurate test results in the agent language ability test, especially in large-scale tests, manual judgment efficiency is slow and subjective, automated tests rely on predefined templates and string matching, and cannot understand semantics, and NLP technology is complex and computing resources are consumed.

Method used

The automated test method of agent language ability based on self-evolution of multilingual models is adopted. By obtaining the output results and correlation information of agent responses, similarity search is performed, prompt words are constructed, and voting is used to evaluate the output results of agents.

Benefits of technology

It improves the efficiency and accuracy of the test, achieves higher inference accuracy and consistency, automatically completes the test, dynamically expands the sample library, and improves the generalization ability and self-evolution ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357069B_ABST
    Figure CN119357069B_ABST
Patent Text Reader

Abstract

The present invention provides an automated test method for the language ability of an agent based on the self - evolution of a multilingual model. Among them, the automated test method for the language ability of an agent based on the self - evolution of a multilingual model includes: obtaining associated information; performing similarity retrieval in a pre - constructed sample library to obtain retrieval samples that meet the set requirements; constructing prompt words; inputting the prompt words into a vote counter, where the vote counter includes multiple voting language models, and each voting language model outputs a corresponding voting result; evaluating the output result of the agent based on multiple voting results to obtain a judgment result, achieving the purpose of improving efficiency and accuracy. The output result is accompanied by a confidence score. When the confidence is high enough, the result is automatically incorporated into the sample library to continuously increase high - quality samples in a dynamically expanding manner, thereby improving the generalization ability of the voting language model and realizing the self - evolution of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to an automated test method for the language ability of an agent based on the self-evolution of a multilingual model. Background Art

[0002] General artificial intelligence aims to enable agents to be like humans, not only able to independently perceive the environment, think, make decisions, learn new skills, execute tasks, but also able to effectively cooperate with humans or other agents. At the same time, the agent can understand human emotions and follow social ethics and moral norms.

[0003] After interacting with the agent through speech and the agent gives feedback (answer / action), how to determine whether the agent's feedback is correct is an important technology for evaluating the agent.

[0004] Currently, for the test of the language ability of general artificial intelligence agents, the following 3 methods are generally adopted:

[0005] Manual judgment of results (manual testing): The tester judges whether the agent's answer is accurate, reasonable, and natural by interacting with the agent according to a predefined set of questions. This method is usually applicable to complex dialogue scenarios and questions that cannot be verified by simple rules, and can simulate the usage scenarios of real users and give the most intuitive feedback.

[0006] Automated testing based on keyword rules verification of a specific corpus use case library: This method constructs a specific corpus or use case library, predefines templates for questions and answers, and verifies whether the agent's answer meets the expectations through string matching. This method is often used to check some simple common sense questions and chatting questions. By setting a set of rules or keywords, when the agent's answer contains these keywords, it is determined that the answer is correct.

[0007] Automated testing based on traditional natural language processing technology: Automated testing based on natural language processing (NLP) technology mainly understands the agent's answer through means such as word segmentation, part-of-speech tagging, semantic analysis, syntactic analysis, intent recognition, and action mapping. Through these technologies, the test system can not only check whether the answer is accurate, but also judge whether the semantics match, whether the context is coherent, and can even handle complex multi-turn conversations and ambiguous expressions. Through semantic-level analysis, NLP technology makes the agent test no longer limited to simple string matching.

[0008] The inventor found the following disadvantages in the prior art during the implementation of this embodiment:

[0009] Manual judgment results: a) Low efficiency, especially in large-scale tests, with slow execution speed; b) Test results may be affected by the subjectivity of testers, lacking consistency and objectivity.

[0010] Automated testing based on keyword rule verification of a specific corpus: a) Unable to handle the diversity of language expressions, relying on predefined templates and unable to adapt to changes in new scenarios and new statements. b) String matching cannot understand semantics, resulting in potential misjudgments and insufficient test coverage.

[0011] Automated testing based on traditional natural language processing techniques: a) Complex to implement, requiring a large amount of computing resources and having low execution efficiency for large-scale tests. b) The effectiveness and accuracy of NLP techniques depend on the training quality of the model. If the model is not perfect, it may lead to deviations in understanding and affect the test results. Summary of the Invention

[0012] In view of the problems existing in the prior art, the present invention provides an automated testing method for the language ability of an intelligent agent based on the self-evolution of a multi-language model, which at least partially solves the problems of low efficiency and inaccurate test results existing in the prior art.

[0013] In a first aspect, an embodiment of the present disclosure provides an automated testing method for the language ability of an intelligent agent based on the self-evolution of a multi-language model, including:

[0014] Obtaining associated information based on the output result obtained by the intelligent agent in response to the user input;

[0015] Performing similarity retrieval in a pre-constructed sample library based on the output result and the associated information to obtain a retrieval sample that meets the set requirements;

[0016] Constructing a prompt word based on the output result, the associated information, and the retrieval sample;

[0017] Inputting the prompt word into a voter, the voter includes multiple voting language models, each voting language model individually votes on the input, and each voting language model outputs a corresponding voting result;

[0018] Evaluating the output result of the intelligent agent based on multiple voting results to obtain a determination result.

[0019] Optionally, the obtaining associated information based on the output result obtained by the intelligent agent in response to the user input includes:

[0020] Converting the voice input by the user into the corresponding text;

[0021] Performing intent recognition according to the text and classifying the user input according to the scenario of the recognition result.

[0022] Optionally, obtaining associated information based on the output result obtained by the agent in response to the user input includes:

[0023] Converting the voice response in the agent output result into the corresponding output text;

[0024] Obtaining the scenario information and the action sequence of the agent according to the output text.

[0025] Optionally, in the similarity retrieval in the pre-constructed sample library based on the output result and the associated information, the selection of the sample similarity includes:

[0026] A sample selector based on length, a sample selector based on maximum marginal relevance, and / or a sample selector based on semantic similarity.

[0027] Optionally, the voter includes at least three voting language models.

[0028] Optionally, in the evaluation of the output result of the agent based on multiple voting results, when the voting results of the voting language models are different, select the voting result with the same majority voting result.

[0029] Optionally, after the step of obtaining the determination result by evaluating the output result of the agent based on multiple voting results, it further includes expanding the pre-constructed sample library based on the confidence score of the judgment result output by the voter.

[0030] Optionally, expanding the pre-constructed sample library based on the confidence score of the judgment result output by the voter includes:

[0031] When the confidence is greater than the set value, store the data related to the voting result in the sample library; when the confidence is not greater than the set value, put the voting result into the annotation pool, and after manually evaluating and annotating the voting results in the annotation pool, store the data related to the evaluated and annotated voting results in the sample library.

[0032] In a second aspect, an embodiment of the present disclosure further provides an electronic device, which includes:

[0033] At least one processor; and,

[0034] A memory communicatively connected to the at least one processor; wherein,

[0035] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the methods for automatically testing the language ability of an agent based on multi-language model self-evolution in the first aspect.

[0036] In a third aspect, an embodiment of the present disclosure also provides a computer-readable storage medium storing computer instructions for causing a computer to execute the automated test method for the language ability of an agent based on the self-evolution of a multilingual model according to any one of the first aspects.

[0037] In a fourth aspect, an embodiment of the present disclosure also provides a computer program product including computer programs / instructions that, when executed by a processor, implement the automated test method for the language ability of an agent based on the self-evolution of a multilingual model according to any one of the first aspects.

[0038] The automated test method for the language ability of an agent provided by the present invention, wherein the automated test method for the language ability of an agent based on the self-evolution of a multilingual model obtains associated information, then performs similarity retrieval in a pre-constructed sample library, constructs a prompt word according to information such as the retrieved similar samples, and inputs the prompt word into a vote counter for voting. The language ability of the agent is automatically tested according to the voting result of the vote counter. Voting is performed by multiple voting language models to ensure higher reasoning accuracy and consistency, and the test is automatically completed, thereby achieving the purpose of improving efficiency and accuracy.

[0039] The output result is accompanied by a confidence score. When the confidence is high enough, the result is automatically incorporated into the sample library to continuously increase high-quality samples in a dynamically expanding manner, thereby improving the generalization ability of the voting language model and realizing the self-evolution of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. In the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.

[0041] Figure 1 It is a flowchart of an automated test method for the language ability of an agent based on the self-evolution of a multilingual model provided by an embodiment of the present disclosure;

[0042] Figure 2 It is a schematic block diagram of a device for automated test of the language ability of an agent based on the self-evolution of a multilingual model provided by an embodiment of the present disclosure;

[0043] Figure 3 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The embodiments of the present disclosure will be described in detail below in conjunction with the accompanying drawings.

[0045] It should be clear that the following uses specific specific examples to illustrate the implementation manners of the present disclosure, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope protected by the present disclosure.

[0046] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement a device and / or practice a method. Additionally, this device and / or this method can be implemented using other structures and / or functionality in addition to one or more of the aspects described herein.

[0047] It also should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. The diagrams only show the components related to the present disclosure and are not drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in its actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0048] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0049] With the rapid development of artificial intelligence technology, the capabilities of agents in language understanding and generation have received increasing attention. The proposed method for automated testing of agent language capabilities based on the self-evolution of multilingual models aims to evaluate and verify the language processing capabilities of agents in an automated manner. This method can not only save a large amount of human and time costs but also provide more objective and accurate test results.

[0050] The core of the automated test method for the language ability of an agent based on the self - evolution of a multilingual model lies in simulating the interaction process between the user and the agent and comprehensively evaluating the agent's language processing ability.

[0051] For ease of understanding, as Figure 1 shown, this embodiment discloses an automated test method for the language ability of an agent based on the self - evolution of a multilingual model, including:

[0052] Step S101: Obtain associated information based on the output result obtained when the agent responds to the user input;

[0053] Obtaining associated information based on the output result obtained when the agent responds to the user input includes:

[0054] Convert the user - input speech into the corresponding text;

[0055] Perform intent recognition based on the text and classify the user input according to the scenario of the recognition result;

[0056] The speech - to - text technology is the process of converting the user's speech input into readable text information. The key components of the speech - to - text technology are as follows:

[0057] Speech signal pre - processing: In the process of speech - to - text conversion, it is first necessary to pre - process the speech signal. This includes noise reduction, echo cancellation, gain control, etc., to improve the quality of the speech signal and ensure the accuracy of subsequent processing.

[0058] The main purpose of noise reduction is to reduce or eliminate the interference of background noise on the speech signal. In daily life, the environment is often accompanied by various noises, such as traffic noise, human voices, background sounds of electronic devices, etc. If these noises are not processed, they will be mixed with the speech signal, making it difficult for the speech recognition system to accurately recognize and transcribe the speech content

[0059] The noise reduction technology analyzes the audio input and identifies the noise pattern, then applies a filter to reduce or eliminate the unwanted noise while retaining the speech signal. In this way, the speech recognition system can more clearly capture the user's speech and improve the recognition accuracy.

[0060] Noise reduction also helps to improve the quality of the speech signal, making the speech - to - text result more accurate and reducing misunderstandings and errors caused by noise interference.

[0061] The echo cancellation technology is mainly used to process acoustic echoes, which usually occur in hands - free or conference applications when the sound of the speaker is fed back to the microphone multiple times. This kind of echo will cause the speaker to hear their own just - spoken voice within a short time, seriously affecting the call quality and the accuracy of speech recognition.

[0062] Echo cancellation technology ensures sound clarity and fidelity by identifying and eliminating audio echoes. It can intelligently identify and eliminate up to 99.9% of audio echoes, and multiple high-precision AI technologies are used to separate the near-end signal from the echo signal to ensure accurate restoration of sound quality.

[0063] In the case of multi-person voice chat applications, after enabling echo cancellation, the active duration and participation of users can be significantly improved because users can hear the other party's voice more clearly without being disturbed by echoes.

[0064] Feature extraction: The preprocessed speech signal needs to be converted into feature vectors that can represent the key information of the speech signal. Common feature extraction methods include Mel Frequency Cepstral Coefficients and Mel-spectrum energy features.

[0065] Acoustic model: The acoustic model is the core of the speech-to-text system, which is responsible for mapping the feature vectors to acoustic units (such as phonemes or letters). Deep learning technologies, especially recurrent neural networks and convolutional neural networks, play an important role in the construction of acoustic models.

[0066] Language model: The language model is used to predict the probability distribution of text sequences, which helps the acoustic model select the most likely word sequence among multiple possible candidates. The language model can be neural network-based, such as Long Short-Term Memory networks.

[0067] Post-processing: The post-processing step includes correcting the recognition results to improve accuracy. This may include word-level language model re-ranking, grammar correction, and context-related adjustments.

[0068] Intent recognition is usually based on natural language processing techniques, which analyze the user's text input and map it to predefined intent categories. This can be achieved through keyword matching, pattern recognition, or machine learning methods such as classification algorithms or deep learning models.

[0069] Entity recognition: While recognizing the intent, entity recognition is also an important step. Entity recognition aims to extract key information from the user's input, such as names, locations, dates, etc., which are crucial for understanding the user's specific needs.

[0070] Scene classification: Based on the recognized intent and entities, the user's input can be classified into different scenes. Scene classification helps the agent provide more personalized and context-related responses. For example, if the user asks about the weather, the agent can classify it into the "weather query" scene and provide the corresponding weather information.

[0071] Context Management: In multi-turn conversations, context management is crucial for maintaining the coherence and relevance of the conversation. The agent needs to maintain the conversation state to reference previous information in subsequent interactions.

[0072] Response Generation: Finally, based on the results of intent recognition and scenario classification, the agent will generate corresponding responses. This may involve retrieving information, executing commands, or providing suggestions.

[0073] According to the final use case scenarios, they are divided into common sense, casual chat, scenario, and action instruction categories, etc. Examples for different categories are shown in Tables 1 and 2:

[0074] Table 1. Example Table of Language-related Questions

[0075]

[0076] Table 2. Example Table of Action Instruction Categories

[0077]

[0078] Convert the voice reply in the agent's output result into the corresponding output text;

[0079] Obtain the scenario information and the agent's action sequence based on the output text.

[0080] For common sense questions, in order to verify whether the answer is correct, the current scenario information also needs to be obtained.

[0081] For action instruction questions, in order to verify whether the answer and the behavior are correct, the current scenario information and the agent's action sequence also need to be obtained.

[0082] After the agent responds to the user input, relevant associated information related to the output result needs to be obtained. This information may include, but is not limited to:

[0083] The context information of the user input; the semantic information of the agent's output result; the sentiment tendency of the agent's output result; the grammatical structure of the agent's output result.

[0084] The purpose of obtaining this information is to more accurately evaluate the agent's output result and provide the necessary data support for subsequent similarity retrieval.

[0085] Step S102: Conduct similarity retrieval in the pre-constructed sample library based on the output result and the associated information to obtain retrieval samples that meet the set requirements;

[0086] The sample library records the user's questions, answers, and for each question, the answer, action, and the corresponding context information, whether the result of each answer action is correct, the thinking process of the result, and the confidence in the result.

[0087] Optionally, in the similarity retrieval in the pre-built sample library based on the output result and the associated information, the selection of sample similarity includes:

[0088] A sample selector based on length, a sample selector based on maximum marginal relevance, and / or a sample selector based on semantic similarity.

[0089] Based on the output result of the agent and the associated information, similarity retrieval is performed in the pre-built sample library. The samples contained in the sample library can be manually annotated or automatically generated by machine learning methods. The purpose of similarity retrieval is to find samples similar to the agent's output result for further comparison and analysis.

[0090] The construction of the sample library needs to consider the following aspects:

[0091] Diversity of samples: The sample library needs to contain various types of language samples to cover different language scenarios;

[0092] Representativeness of samples: The samples need to be able to represent various language situations that the agent may encounter;

[0093] Annotation quality of samples: The samples need to have accurate annotations for subsequent similarity retrieval and evaluation.

[0094] According to the current question, answer, and context information, the 10 most relevant samples can be retrieved from the sample library. For Chinese texts, the samples selected by the semantic similarity sample selector have higher quality and better final effects.

[0095] Step S103: Construct a prompt word based on the output result, associated information, and retrieved samples;

[0096] Fuse the current question answering action to be judged, context information, and the reference answers selected from the sample library into the prompt word template to generate the final prompt word.

[0097] After obtaining the similar samples, it is necessary to construct prompt words based on the output result, associated information, and retrieved samples. Prompt words are key information used to guide the decision-making of the voter, and they can be keywords, phrases, or entire sentences. The construction of prompt words needs to consider the following factors:

[0098] Relevance of prompt words: The prompt words need to be highly relevant to the agent's output result and associated information;

[0099] Coverage of prompt words: The prompt words need to cover all key information of the agent's output result as much as possible;

[0100] Conciseness of prompt words: The prompt words need to be concise and clear for the voter to understand and process.

[0101] Step S104: Input the prompt into the vote counter, where the vote counter includes multiple voting language models, and each voting language model individually votes on the input and outputs a corresponding voting result.

[0102] Input the constructed prompt into the vote counter, which contains multiple voting language models. Each model individually votes on the input and outputs a corresponding voting result. The voting language model can be a machine learning model based on different algorithms, such as: support vector machine, random forest, neural network, and deep learning model. Each voting language model has its own characteristics and advantages. By integrating multiple models, the accuracy and robustness of the decision-making can be improved.

[0103] Step S105: Evaluate the output result of the intelligent agent based on multiple voting results to obtain a judgment result.

[0104] Before conducting automated testing, a series of evaluation criteria need to be set. These criteria can be:

[0105] Accuracy criterion: The degree of matching between the output result of the intelligent agent and the expected result;

[0106] Consistency criterion: Whether the responses of the intelligent agent in different contexts are consistent;

[0107] Naturalness criterion: Whether the responses of the intelligent agent conform to human language habits;

[0108] Reliability criterion: Whether the responses of the intelligent agent are stable and reliable.

[0109] Optionally, the vote counter includes at least three voting language models.

[0110] In the evaluation of the output result of the intelligent agent based on multiple voting results, when the voting results of the voting language models are different, select the voting result with the majority of the same voting results.

[0111] Suppose there are differences in the judgment results of three voting language models for the same question and answer. Then, select the results of the two voting language models with the same judgment results as the final result. Even if a certain voting language model makes a mistake, one can still rely on the majority of the correct versions to ensure normal operation, thus achieving the effect of fault tolerance and improving the accuracy rate. It is found that after adopting three voting language models, the accuracy rate can be increased by about 10% compared to using a single voting language model alone.

[0112] Through the voting mechanism composed of multiple voting language models, comprehensively evaluate the answers and behaviors of the intelligent agent, output the final judgment result, and determine whether it is correct.

[0113] After the step of evaluating the output result of the agent based on multiple voting results to obtain a judgment result, it further includes expanding a pre-constructed sample library based on the confidence score of the judgment result output by the voter.

[0114] The expanding of the pre-constructed sample library based on the confidence score of the judgment result output by the voter includes:

[0115] When the confidence is greater than the set value, the data related to the voting result is stored in the sample library; when the confidence is not greater than the set value, the voting result is put into the annotation pool, and after manual evaluation and annotation of the voting results in the annotation pool, the data related to the evaluated and annotated voting results is stored in the sample library. The set value is set as needed and can be set to 95 in a specific scenario, or 90, etc. The confidence range is between 0 and 100, and the higher the value, the higher the confidence of the model in the prediction result. The data related to the voting result includes questions, answers, context, reasoning logic, and judgment results, etc.

[0116] Through the method of this embodiment, the sample library has the ability of self-expansion and self-optimization without manual intervention, realizing true self-evolution. Through experiments by the inventor, 20 pieces of data were used in the cold start stage. As the running time increased and the scenarios and Q&A data became continuously rich, the sample library self-accumulated approximately 500 pieces of data, significantly improving the accuracy of the entire test and verification system. Compared with the traditional model system, it greatly reduces the dependence on the initial data volume, can quickly converge based on a small amount of data, achieves significant effects, and realizes more efficient data utilization and model generalization ability.

[0117] Based on multiple voting results, a comprehensive evaluation is carried out on the output result of the agent. The contents of the evaluation include:

[0118] Accuracy of the agent's output result: Evaluate whether the agent can accurately understand the user's intention and give a correct response;

[0119] Consistency of the agent's output result: Evaluate whether the agent can maintain a consistent response in different contexts;

[0120] Naturalness of the agent's output result: Evaluate whether the agent's response is natural and fluent, conforming to human language habits;

[0121] Reliability of the agent's output result: Evaluate whether the agent's response is stable and reliable, not affected by external factors.

[0122] The voter can adopt a model generated by deep learning algorithms, neural network algorithms, etc., which has a certain understanding and reasoning ability for questions, etc., and can refer to existing ones such as Gpt, Tongyi Qianwen, and Gemini.

[0123] Deep learning and neural network algorithms have made significant progress in simulating human language understanding and reasoning abilities. The development of these technologies has made it possible to build intelligent systems capable of handling complex language tasks.

[0124] Deep learning algorithms are a series of algorithms based on artificial neural networks. These algorithms learn complex patterns in data by mimicking the way the human brain processes information. Deep learning models usually consist of multiple layers, each layer composed of multiple neurons that can learn different levels of features in the data. The core of deep learning algorithms lies in automatically extracting features without the need to artificially design feature extraction rules.

[0125] The development of neural network algorithms can be traced back to the 1940s when scientists tried to simulate the neuron networks of the human brain. Over time, neural network algorithms have experienced several ups and downs. In the early 21st century, with the improvement of computing power and the availability of big data, deep learning algorithms began to make breakthrough progress in fields such as image recognition, speech recognition, and natural language processing.

[0126] The application of deep learning algorithms in language understanding mainly focuses on the field of natural language processing. NLP is a branch of artificial intelligence, and its key applications are as follows:

[0127] Text classification: Deep learning models can be trained to identify the sentiment of text or classify text into predefined categories.

[0128] Machine translation: Deep learning models, especially those based on the attention mechanism, have significantly improved the accuracy and fluency of machine translation.

[0129] Question answering systems: Deep learning models are able to understand questions and extract answers from given texts, which is very useful in question answering systems.

[0130] Text summarization: Deep learning models can automatically generate summaries of texts, which is very valuable for processing large amounts of information.

[0131] Dialogue systems: Deep learning models can be used to build chatbots and virtual assistants that can understand user inputs and provide appropriate responses.

[0132] Deep learning models learn hierarchical features of data through multi-layer structures. In text processing, this usually means that the model first learns the embedded representation of words, then the semantics of sentences, and finally the theme of the entire document. This hierarchical representation enables the model to capture the complexity and diversity of language.

[0133] Word Embedding: Word embedding is a vector representation that maps words into a high-dimensional space, and these vectors capture the semantic information of words. For example, the distance between "cat" and "dog" in the vector space may be closer than that between "cat" and "chair".

[0134] Sentence and Paragraph Representations: Based on word embeddings, deep learning models can further learn the representations of sentences and paragraphs. These representations can capture the structure and semantic information of sentences.

[0135] Document and Topic Representations: For longer texts such as articles or books, deep learning models can learn document-level representations that can capture the theme and structure of the entire document.

[0136] The attention mechanism is an important concept in deep learning models. It allows the model to focus on the most relevant parts for the current task when processing sequential data. In machine translation, the attention mechanism can help the model pay attention to relevant words in the source language when translating a word. In text summarization, the attention mechanism can help the model determine which sentences are crucial for generating the summary.

[0137] Context understanding is a key ability of deep learning models when dealing with language tasks. By using models such as Transformer, long-distance dependencies and context information can be better captured, which is crucial for understanding the overall meaning of a sentence or paragraph. The Transformer model achieves this through the self-attention mechanism, which allows the model to consider the entire input sequence when processing a word.

[0138] Many modern NLP models learn general language representations through pre-training on large-scale datasets and then fine-tune on specific tasks, which enhances the model's ability to understand problems in specific domains. The pre-training stage allows the model to learn rich language patterns, while the fine-tuning stage enables the model to adapt to specific tasks.

[0139] Models used to solve specific inference tasks, such as question-answering systems and dialogue systems, can generate answers by analyzing questions and context information, showing a certain degree of reasoning ability. This reasoning ability is achieved through the language patterns and structures learned by the model, rather than by simulating the logical reasoning process of humans.

[0140] The automated test process disclosed in this embodiment includes the following steps:

[0141] User Input Simulation: Simulate the interaction between the user and the agent and input test samples;

[0142] Agent Response: The agent processes the user input and gives a response;

[0143] Associated information acquisition: Obtain the associated information of the agent's response;

[0144] Similarity retrieval: Conduct similarity retrieval in the sample library to find similar samples;

[0145] Prompt construction: Construct prompts based on the response results and similar samples;

[0146] Voter decision-making: Input the prompts into the voter for decision-making;

[0147] Result evaluation: Evaluate the output of the agent according to the voting results.

[0148] Result analysis and optimization:

[0149] After completing the automated test, it is necessary to analyze the test results. The content of the analysis includes:

[0150] Statistical analysis of test results: Statistically analyze the performance of the agent on various evaluation criteria;

[0151] Comparative analysis of test results: Compare the performance of the agent on different test samples;

[0152] Problem diagnosis of test results: Diagnose the problems that occur in the agent during the test and propose improvement suggestions;

[0153] Optimization of the test process: Optimize the test process according to the test results to improve the accuracy and efficiency of the test.

[0154] The automated test method for the language ability of an agent based on the self-evolution of a multilingual model disclosed in this embodiment can comprehensively evaluate the language processing ability of the agent. By simulating the interaction between the user and the agent, it automatically obtains associated information, conducts similarity retrieval, constructs prompts, and uses a voter for decision-making, and finally evaluates the output results of the agent. This method not only improves the efficiency and objectivity of the test, but also can monitor the performance of the agent in real time and make dynamic adjustments. However, this method also faces challenges such as sample library construction, selection of voting language models, interpretation of test results, and security and privacy protection. Future research and practice need to explore these aspects in depth to further improve the accuracy and reliability of the automated test method for the language ability of an agent based on the self-evolution of a multilingual model.

[0155] This embodiment has the following advantages:

[0156] 1. Discriminate the test results of the agent's language ability through a voting language model: In the test of the agent's language ability, an automated test result discrimination is carried out using a voting language model, which effectively improves the intelligence and adaptability of the discrimination, and also significantly expands the range of statements supported for verification, making the test coverage wider and the verification statements richer.

[0157] 2. Voting mechanism of multiple voting language models: By combining multiple voting language models to form a voting device, the accuracy of reasoning is significantly improved. Compared with a single model, the voting method has stronger fault tolerance and consistency in complex language understanding and reasoning tasks. In the automated test of the language ability of a general artificial intelligence agent, a voting device composed of multiple voting language models is used to judge the test results, thus significantly improving the accuracy and robustness of the judgment.

[0158] 3. Self-expanding sample library mechanism: A high-confidence output mechanism based on the voting language model is designed, and a confidence score is attached to the output result. When the confidence is high enough, the system automatically incorporates the judgment result into the sample library, thus realizing the self-expansion and dynamic enrichment of the sample library. Without manual intervention, the generalization ability and adaptability can be gradually enhanced as the scenario and Q&A content grow, achieving continuous evolution. Automatic expansion mechanism of the voting language model sample with high-confidence results: Based on the voting language model, a discrimination result is given along with a confidence level, and the high-confidence results are directly incorporated into the sample library, enabling the sample library to achieve self-expansion, dynamic growth, and continuous optimization in actual application scenarios, realizing intelligent sample library management without manual intervention, thereby effectively enhancing the generalization ability and evolution speed of the system.

[0159] As Figure 2 shown, this embodiment also discloses an automated test device for the language ability of an intelligent agent that self-evolves based on a multi-language model, including:

[0160] An information module for obtaining associated information based on the output result obtained by the intelligent agent in response to the user input;

[0161] A retrieval module for performing similarity retrieval in a pre-constructed sample library based on the output result and the associated information to obtain a retrieval sample that meets the set requirements;

[0162] A construction module for constructing a prompt word based on the output result, the associated information, and the retrieval sample;

[0163] A voting result module for inputting the prompt word into a voting device, where the voting device includes multiple voting language models, each voting language model individually votes on the input, and each voting language model outputs a corresponding voting result;

[0164] A judgment module for evaluating the output result of the intelligent agent based on multiple voting results to obtain a judgment result.

[0165] The obtaining of the associated information based on the output result obtained by the intelligent agent in response to the user input includes:

[0166] Converting the voice input by the user into the corresponding text;

[0167] Perform intent recognition based on the text, and classify the user input according to the scenarios of the recognition results.

[0168] Obtain associated information based on the output result obtained by the agent in response to the user input, including:

[0169] Convert the voice response in the agent output result into the corresponding output text.

[0170] Obtain scenario information and the action sequence of the agent according to the output text.

[0171] In the similarity retrieval in the pre-constructed sample library based on the output result and the associated information, the selection of sample similarity includes:

[0172] A sample selector based on length, a sample selector based on maximum marginal relevance, and / or a sample selector based on semantic similarity.

[0173] The vote counter includes at least three voting language models.

[0174] In the evaluation of the output result of the agent based on multiple voting results, when the voting results of the voting language models are different, select the voting result with the same majority of voting results.

[0175] After the step of evaluating the output result of the agent based on multiple voting results to obtain the judgment result, it further includes expanding the pre-constructed sample library based on the confidence score of the judgment result output by the vote counter.

[0176] The expansion of the pre-constructed sample library based on the confidence score of the judgment result output by the vote counter includes:

[0177] When the confidence is greater than the set value, store the data related to the voting result in the sample library; when the confidence is not greater than the set value, put the voting result into the annotation pool, and after manually evaluating and annotating the voting results in the annotation pool, store the data related to the evaluated and annotated voting results in the sample library.

[0178] The electronic device disclosed in this embodiment includes a memory and a processor. The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0179] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In one embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the electronic device executes all or part of the steps of the method for automatically testing the intelligent agent language ability based on multi-language model self-evolution in the foregoing embodiments of the present disclosure.

[0180] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, this embodiment may also include well-known structures such as communication buses, interfaces, etc., and these well-known structures should also be included in the protection scope of the present disclosure.

[0181] Such as Figure 3 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. It shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present disclosure. Figure 3 The shown electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0182] Such as Figure 3 As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the electronic device are also stored. The processing device, ROM, and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0183] Generally, the following devices may be connected to the I / O interface: an input device including, for example, a sensor or a visual information acquisition device; an output device including, for example, a display screen; a storage device including, for example, a magnetic tape, a hard disk, etc.; and a communication device. The communication device may allow the electronic device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 3 The shown electronic device has various devices, but it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.

[0184] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, all or part of the steps of the method for automatically testing the language ability of an agent based on self-evolution of a multi-language model according to embodiments of the present disclosure are performed.

[0185] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.

[0186] The computer-readable storage medium disclosed in this embodiment stores non-temporary computer-readable instructions. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the method for automatically testing the language ability of an agent based on self-evolution of a multi-language model according to the foregoing embodiments of the present disclosure are performed.

[0187] The above-mentioned computer-readable storage medium includes but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).

[0188] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.

[0189] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-disclosed specific details are only for illustrative and facilitating understanding purposes, rather than limitations, and the above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0190] In this disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.

[0191] In addition, as used herein, "or" in a listing of items beginning with "at least one" indicates a disjunctive listing, so that for example, a listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Further, the term "exemplary" does not mean that the examples described are preferred or better than other examples.

[0192] It should also be noted that in the systems and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.

[0193] Various changes, substitutions, and alterations to the technologies described herein can be made without departing from the teachings defined by the appended claims. In addition, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Accordingly, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.

[0194] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0195] The foregoing description has been presented for purposes of illustration and description. In addition, the description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. An automated testing method for the language ability of an intelligent agent based on self-evolution of a multilingual model, characterized in that: include: Obtaining relevant information based on the output results obtained by the agent in response to the user input; Based on the output results and the associated information, similarity search is performed in a pre-built sample library to obtain search samples that meet the set requirements; Construct prompt words based on output results, related information and retrieval samples; Input the prompt word into a voting machine, wherein the voting machine includes a plurality of voting language models, each voting language model votes on the input separately, and each voting language model outputs a corresponding voting result; Based on multiple voting results, the output results of the intelligent agent are evaluated to obtain a judgment result; The step of obtaining the associated information based on the output result obtained by the agent in response to the user input includes: Convert the voice response in the agent's output into the corresponding output text; Obtain scene information and the action sequence of the agent based on the output text; After the step of evaluating the output results of the intelligent agent based on the multiple voting results to obtain the judgment result, the method further includes expanding the pre-built sample library based on the confidence score of the judgment result output by the voter; The confidence score based on the determination result output by the voter is used to expand the pre-built sample library, including: When the confidence is greater than the set value, the data related to the voting results will be stored in the sample library; when the confidence is not greater than the set value, the voting results will be placed in the annotation pool. After manually evaluating and annotating the voting results in the annotation pool, the data related to the evaluated and annotated voting results will be stored in the sample library. The set value is 95.

2. The method for automated testing of agent language ability based on multilingual model self-evolution according to claim 1, characterized in that: The step of obtaining the associated information based on the output result obtained by the agent in response to the user input includes: Convert the user's input voice into corresponding text; Perform intent recognition based on text and classify user input based on the scenario of the recognition result; According to the final use case scenario, it is divided into common sense category, chat category, scenario category and action instruction category.

3. The method for automated testing of agent language ability based on multilingual model self-evolution according to claim 1, characterized in that: In the similarity search in the pre-built sample library based on the output result and the associated information, the sample similarity selection includes: A length-based sample selector, a maximum marginal correlation-based sample selector, and / or a semantic similarity-based sample selector.

4. The method for automated testing of agent language ability based on multilingual model self-evolution according to claim 1, characterized in that: The voting device includes at least three voting language models.

5. The method for automated testing of agent language ability based on multilingual model self-evolution according to claim 4, characterized in that: In the evaluation of the output result of the intelligent agent based on multiple voting results, when the voting results of the voting language models are different, the voting results with the same majority are selected.

6. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the automated testing method for the language ability of an intelligent agent based on self-evolution of a multi-language model as described in any one of claims 1-5.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the method for automated testing of the language ability of an intelligent agent based on self-evolution of a multi-language model as described in any one of claims 1-5.

8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the method for automatically testing the language ability of an intelligent agent based on self-evolution of a multi-language model as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Corpus tag determination method and device, electronic equipment and storage medium

    CN115203405A

  • Contract review report generation method based on large model integrated optimization

    CN118940732A