Multi-agent simulation confrontation decision generation method based on large language model

By employing a multi-agent simulation adversarial decision generation method, the problems of repetitive queries and historical response optimization in question-answering systems are solved, enabling dynamic allocation of computing resources and high-quality response generation, thereby improving the user interaction experience.

CN121765052APending Publication Date: 2026-03-31SHANGHAI SHUZAI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing question-answering systems based on large language models struggle to identify patterns of repeated user inquiries, lack mechanisms for reflecting on the quality of historical responses, and have fixed methods for allocating computing resources, making it difficult to balance response efficiency and quality and unable to adapt to demanding scenarios.

Method used

By employing a multi-agent simulation adversarial decision generation method, a semantic analyzer is used to identify repetitive query behavior, and an intelligent agent set is dynamically configured to perform simulation adversarial decision-making to generate high-quality responses.

Benefits of technology

It effectively identifies and quantifies repeated user queries, optimizes historical response content, adaptively configures computing resources, generates high-quality text that meets user needs, and improves the accuracy and diversity of the question-and-answer system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765052A_ABST
    Figure CN121765052A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent simulation confrontation decision generation method based on a large language model, and relates to the technical field of machine questions and answers. The method comprises the steps of receiving a target dialogue text input by a user, and obtaining a dialogue record where the target dialogue text is located; and performing semantic comparison on the target dialogue text and the plurality of historical dialogue texts to obtain a plurality of text semantic consistency degrees. And when the text semantic consistency is greater than or equal to a consistency threshold, configuring a target agent set. According to the target dialogue text, multiple rounds of simulation confrontation decision making are carried out through the target agent set, and multiple candidate reply texts are generated. And performing cross scoring and accumulation on the candidate reply texts through each selected agent, determining a target reply text, and sending the target reply text to the user side. The problems that a traditional question-answering system lacks reflection on repeated questions, reply homogenization is serious, resource allocation is rigid and the like are solved, and the accuracy, diversity and user experience of question-answering interaction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine question answering technology, specifically to a multi-agent simulation adversarial decision generation method based on a large language model. Background Technology

[0002] With the rapid development of artificial intelligence technology, intelligent question-answering systems based on large language models have been widely applied in various fields such as intelligent customer service, online education, technical support, and medical consultation. In these application scenarios, users often repeatedly ask the same question or related topics, hoping to obtain clearer, more comprehensive, or deeper answers. Therefore, how to enable the system to proactively identify users' repeated questioning behavior and generate more targeted and informative responses based on the quality of historical interactions has become a real need to improve the interactive experience and service efficiency.

[0003] However, existing question-answering systems based on large language models generally suffer from several technical shortcomings. First, current technologies lack an effective mechanism for identifying user question patterns, making it difficult to determine whether the current question is a semantic repetition of a previous question, thus failing to trigger differentiated response strategies. Second, they lack the ability to reflect on the quality of their past responses. When previous responses were inadequate, subsequent repeated questions from users fail to prompt the system to correct and optimize the content, leading to a predicament of repeatedly generating similar, low-quality responses. Furthermore, the current technologies allocate computational resources in a relatively fixed manner, unable to dynamically adjust processing depth and resource investment based on the actual complexity of the question, the intensity of user follow-up inquiries, and historical performance. This results in a difficulty in balancing response efficiency and quality, limiting the adaptability of question-answering systems in demanding scenarios. Summary of the Invention

[0004] This invention addresses the technical problems of existing technologies, such as difficulty in identifying repeated user inquiry patterns, lack of reflection mechanisms on the quality of historical responses, susceptibility to repeatedly generating similar content, and difficulty in dynamically configuring computing resources. It provides a multi-agent simulation adversarial decision generation method based on a large language model.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: This invention provides a multi-agent simulation adversarial decision generation method based on a large language model, comprising: The system receives the target dialogue text input by the user through the user terminal and obtains the dialogue record where the target dialogue text is located. The dialogue record includes multiple historical dialogue groups, and each historical dialogue group includes the historical dialogue text input by the user and the historical reply text received by the user. The target dialogue text is compared with multiple historical dialogue texts to obtain multiple text semantic consistency scores. When there is a text semantic consistency score greater than or equal to the consistency score threshold, the target intelligent agent set is configured according to the target dialogue text, multiple historical dialogue texts and multiple historical response texts. Based on the target dialogue text, a response simulation adversarial decision is made through the target intelligent agent set to generate a target response text, which is then sent to the user terminal.

[0006] The beneficial effects of this invention are: Compared to existing technologies, this invention firstly effectively identifies users' repeated questioning behavior through semantic analysis and quantifies the intensity of their follow-up inquiries, thereby proactively triggering a response optimization mechanism. Secondly, it introduces an evaluation of the quality of historical responses, identifying and avoiding the repetition of previously existing defects by analyzing the homogeneity of response content. Thirdly, based on the dynamic analysis results of the dialogue context, it adaptively configures the scale and type of multiple agents, achieving precise on-demand allocation of computing resources. Finally, by organizing multiple lightweight agents for simulated adversarial and collaborative optimization, it generates high-quality text that both meets the user's current needs and transcends the limitations of historical responses, thus systematically improving the accuracy, diversity, and user satisfaction of responses in various question-and-answer interaction scenarios. Attached Figure Description

[0007] Figure 1 A flowchart illustrating a multi-agent simulation adversarial decision generation method based on a large language model provided by the present invention; Figure 2 This is a schematic diagram illustrating the configuration process of the first intelligent agent configuration coefficients provided by the present invention. Detailed Implementation

[0008] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0009] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0010] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0011] Example 1, as Figure 1 As shown, this embodiment of the invention provides a multi-agent simulation adversarial decision generation method based on a large language model, including: S10: Receive the target dialogue text input by the user through the user terminal, and obtain the dialogue record where the target dialogue text is located. The dialogue record includes multiple historical dialogue groups, and each historical dialogue group includes the historical dialogue text input by the user and the historical reply text received by the user. First, the system receives the target dialogue text input by the user through the client. This target dialogue text is the specific question or statement that the user hopes to receive a response to. Then, it retrieves the dialogue record containing this target dialogue text. Specifically, this dialogue record covers a series of pre-stored historical interaction sequences associated with the current user or the current conversation topic, including multiple historical dialogue groups. Each historical dialogue group includes the historical dialogue text input by the user and the historical response text received by the user.

[0012] Specifically, each historical dialogue group is a structured data unit that fully records the content of a historical interaction between the two parties. Each historical dialogue group clearly contains two components: first, the historical dialogue text entered by the user at that time, i.e., the historical question or statement; and second, the historical response text received by the user in that historical interaction, generated and fed back by the question-and-answer system.

[0013] The resulting dialogue transcript provides a complete informational foundation for subsequent analysis, including the target dialogue text and its historical context.

[0014] S20: Compare the target dialogue text with multiple historical dialogue texts to obtain multiple text semantic consistency scores. When there is a text semantic consistency score greater than or equal to the consistency score threshold, configure the target intelligent agent set according to the target dialogue text, multiple historical dialogue texts and multiple historical reply texts. Specifically, the target dialogue text is compared with multiple historical dialogue texts to obtain multiple text semantic consistency scores, including: Text feature extraction is performed on the target dialogue text and multiple historical dialogue texts respectively to obtain the target text feature vector and multiple historical text feature vectors; Extract a first historical text feature vector from the plurality of historical text feature vectors, and input the first historical text feature vector and the target text feature vector into a semantic analyzer to obtain the first text semantic consistency. Following the method of obtaining the first text semantic consistency degree corresponding to the first historical text feature vector, the text semantic consistency degrees corresponding to the other historical text feature vectors are obtained, resulting in multiple text semantic consistency degrees.

[0015] First, text feature extraction processing is performed on the aforementioned target dialogue text and the multiple historical dialogue texts extracted from the dialogue records. Specifically, this text feature extraction process converts each text segment into a numerical representation in a high-dimensional space, i.e., a feature vector. After processing, a target text feature vector representing the current target dialogue text and multiple historical text feature vectors representing each historical dialogue text are generated.

[0016] Secondly, the first vector is extracted from multiple historical text feature vectors and used as the first historical text feature vector for analysis. This first historical text feature vector and the target text feature vector are then input into a pre-trained semantic analyzer. This semantic analyzer performs in-depth comparison and calculation of the semantic information contained in the first historical text feature vector and the target text feature vector, outputting a quantified similarity score, which is the semantic consistency score of the first text. This semantic consistency score characterizes the degree of semantic similarity between the first historical dialogue text and the current target dialogue text. The higher the semantic consistency score, the more similar the two texts are semantically, meaning the closer the user's current query is to the content of the historical query; conversely, a lower score indicates a greater semantic difference between the two texts.

[0017] Specifically, the construction steps of the semantic analyzer include: Collect a set of historical dialogue records. For each historical dialogue record in the set, extract all historical dialogue texts. Pair the extracted historical dialogue texts to construct a set of sample feature vector pairs, where each sample feature vector pair includes two text feature vectors. Semantic consistency is labeled for each sample feature vector pair in the sample feature vector pair set to obtain a sample semantic consistency set; Using the set of sample feature vector pairs as input features and the set of sample semantic consistency as supervision labels, supervised training is performed until convergence to obtain the semantic analyzer.

[0018] First, a historical dialogue record set is collected. This set includes multiple complete, completed dialogue sessions, typically derived from real or simulated interaction logs in specific application scenarios, obtained through data collection and anonymization. For each historical dialogue record in this set, all historical dialogue text is extracted. Next, text feature extraction processing is performed on the extracted historical dialogue text, converting each segment of text into a corresponding text feature vector. Based on this, all text feature vectors are paired to construct a sample feature vector pair set. Specifically, each sample unit in this sample feature vector pair set includes two text feature vectors, which may come from different rounds within the same historical dialogue record or from different historical dialogue records, thus simulating various possible scenarios of text semantic comparison in reality.

[0019] Secondly, the constructed set of sample feature vector pairs is labeled. Specifically, the labelers or the labeling system determine the degree of semantic consistency between the two text feature vector pairs based on the actual semantic content of the original texts they correspond to, and assign a quantified semantic consistency label to each sample feature vector pair. All labeled labels constitute the sample semantic consistency set, which serves as a supervision signal and provides a clear target for model learning.

[0020] Finally, using the sample feature vector pair set as input features and the sample semantic consistency set as corresponding supervision labels, supervised learning is employed for training until convergence, yielding a semantic analyzer. This semantic analyzer possesses the ability to accurately infer the semantic consistency between any two input text feature vectors. The convergence condition is set based on performance evaluation metrics during training. For example, convergence is defined as follows: the model is considered to have reached convergence when the loss function value on the independent validation set no longer decreases within 20 consecutive training epochs, or when the decrease is below a predefined threshold such as 0.001.

[0021] For example, since the semantic similarity judgment between two text feature vectors needs to capture deep semantic associations and complex patterns, and deep neural network models have shown powerful capabilities in semantic understanding and relation modeling, a deep neural network model can be selected to build this semantic analyzer.

[0022] Specifically, this semantic analyzer employs a neural network architecture based on Siamese networks or interactive matching. Its core structure mainly includes an input layer, a multi-layer feature interaction and fusion layer, and a similarity output layer. The input layer receives a sample feature vector pair consisting of two text feature vectors. The feature interaction and fusion layer uses stacked multi-layer fully connected networks or attention mechanism layers to deeply compute the interaction information between the two input vectors and progressively abstract joint features representing their semantic relationship. In this layer, each neural network layer uses the ReLU activation function to introduce non-linear modeling capabilities, and a Dropout layer is added after some network layers with a dropout rate set to 0.2 to 0.3 to enhance the model's generalization ability and prevent overfitting. The output layer uses a fully connected layer in conjunction with a Sigmoid activation function to map the fused high-level features to a continuous value between 0 and 1, which represents the text semantic consistency predicted by the model.

[0023] During training, the key hyperparameters were set as follows: learning rate was set to 0.0005, number of training epochs was set to 200, and batch size was set to 128. The learning rate was set to balance training stability and convergence speed; the number of training epochs ensured that the model had sufficient iterations to learn complex semantic matching patterns; and the batch size was chosen to consider both training efficiency and gradient estimation stability.

[0024] The specific training process employs a supervised learning approach. The previously constructed set of sample feature vector pairs serves as the input sample set, and its corresponding labeled set of sample semantic consistency scores serves as the supervised label set. The input sample set and the label set are randomly divided into a training set, a validation set, and a test set in an 8:1:1 ratio.

[0025] Furthermore, using the sample feature vector pairs in the training set as input and the corresponding semantic consistency labels as supervision signals, the network weight parameters of the semantic analyzer are iteratively updated through backpropagation and the Adam optimizer. During training, the mean squared error loss function is used to quantify the deviation between the model's predicted semantic consistency and the true labeled values. Simultaneously, a validation set is used to monitor the training process. When the loss function value on the validation set no longer decreases within 20 consecutive training epochs, and the Pearson correlation coefficient between the predicted and true values ​​reaches and stabilizes above 0.98 on the validation set, the model is considered to have converged, and the training process is terminated, thus obtaining a fully trained semantic analyzer. This semantic analyzer can accurately infer semantic consistency from any two input text feature vectors.

[0026] Furthermore, the feature vectors of the first historical text and the target text are jointly input into the trained semantic analyzer. The semantic analyzer calculates and outputs the semantic consistency score of the first text. This semantic consistency score directly represents the degree of semantic agreement between the first historical text and the current target text.

[0027] Finally, the processing method described above for the first historical text feature vector is iteratively applied to all other historical text feature vectors. Specifically, each historical text feature vector is paired with the target text feature vector in turn and input into the semantic analyzer for calculation. After traversing all historical texts, multiple text semantic consistency scores are obtained. These multiple text semantic consistency scores reflect the semantic association strength between the current target dialogue text and each historical dialogue text in the dialogue session, providing accurate data for subsequent determination of whether duplicate queries exist.

[0028] Furthermore, after obtaining the semantic consistency of multiple texts, the semantic consistency of multiple texts is analyzed and judged according to the preset consistency threshold, and the subsequent response generation strategy is determined accordingly.

[0029] Specifically, firstly, each obtained semantic consistency score is compared with a pre-set consistency threshold. The consistency threshold is a critical value used to define whether semantic similarity reaches the standard of duplication or high relevance. The setting of this threshold depends on the performance calibration of the semantic analyzer and the sensitivity requirements of the actual application scenario regarding question repetition. For example, if semantic consistency uses a 0-1 scale, where 1 represents complete consistency, analysis of historical interaction data might determine that when the consistency exceeds 0.85, the two questions can be considered substantially duplicated in their core intent. Therefore, the consistency threshold can be set to 0.85 for example. This consistency threshold can be adjusted accordingly for scenarios with different rigor requirements; for example, it can be set to 0.9 in domains requiring extremely high precision, and 0.8 in general domains.

[0030] Then, based on the comparison results, different branch processes are executed. Specifically, if at least one text semantic consistency is greater than or equal to a preset consistency threshold, the current target dialogue text is determined to constitute a repeated inquiry or in-depth follow-up question to a historical question. At this time, a multi-agent collaborative optimization mechanism will be activated. Subsequent processes will dynamically configure and activate the target agent set based on the current target dialogue text, the historical dialogue texts corresponding to those with semantic consistency exceeding the consistency threshold, and the historical response texts associated with those historical dialogue texts, to perform simulated adversarial decision-making, aiming to generate optimized content that surpasses the quality of historical responses.

[0031] Conversely, if the semantic consistency of all texts is below the consistency threshold, it indicates that the current target dialogue text is a semantically novel question with no significant connection to the historical dialogue content. In this case, it is determined that there is no need to initiate the subsequent complex multi-agent simulation adversarial process, and instead a standard or simplified response generation path is adopted. For example, a single benchmark large language model can be directly called to generate the response, effectively saving computational resources while ensuring response efficiency.

[0032] Specifically, when there is a text semantic consistency score greater than or equal to the consistency score threshold, a target intelligent agent set is configured based on the target dialogue text, multiple historical dialogue texts, and multiple historical response texts, including: Configure the first agent configuration coefficient based on the target dialogue text and the plurality of historical dialogue texts, and configure the second agent configuration coefficient based on the plurality of historical reply texts; The configuration coefficients of the first and second agents are weighted to obtain the comprehensive configuration coefficients of the agents. Connect to the intelligent agent swarm and obtain the total number of intelligent agents in the intelligent agent swarm; The number of agent configurations is determined based on the comprehensive agent configuration coefficient and the total number of agents. Then, multiple selected agents are randomly selected from the agent group based on the number of agent configurations to configure the target agent set.

[0033] Specifically, when the semantic consistency of text is greater than or equal to the consistency threshold, it indicates that the current user's question and the historical dialogue content are semantically repetitive or highly related. The user may not be satisfied with the previous response or may be continuously seeking deeper information. Therefore, a target intelligent agent set needs to be dynamically configured based on the current target dialogue text, multiple historical dialogue texts corresponding to the semantic consistency exceeding the threshold, and multiple associated historical response texts. This target intelligent agent set is a temporary collaborative computing unit composed of multiple lightweight, functionally specialized agents. It can conduct multi-angle simulation and adversarial optimization in parallel to address the identified repetitive questioning patterns and potential defects in historical responses, thereby generating high-quality answers that surpass single models or historical responses in terms of relevance, novelty, and explanatory depth.

[0034] The configuration process of the target intelligent agent set includes several steps. First, two independent configuration coefficients are generated based on the user's question characteristics and the quality of historical responses, including the first intelligent agent configuration coefficient and the second intelligent agent configuration coefficient. Then, they are weighted and fused into a comprehensive intelligent agent configuration coefficient. Next, the number of intelligent agents to be mobilized is determined by combining the total number of intelligent agents. Finally, the target intelligent agent set to be specifically executed for this task is selected from the intelligent agent group through a random selection mechanism.

[0035] First, such as Figure 2As shown, configuring the first agent's configuration coefficients based on the target dialogue text and the multiple historical dialogue texts includes: The multiple historical dialogue texts and target dialogue texts are sorted in chronological order to obtain the current dialogue text sequence, and the total number of text sequences is determined. Each dialogue text in the current dialogue text sequence has a position label. Based on the current dialogue text sequence, the semantic consistency of multiple texts is sorted to obtain a semantic consistency sequence; Find the first text semantic consistency score that is greater than or equal to the consistency score threshold from the semantic consistency score sequence, and determine the position number of the corresponding dialogue text in the current dialogue text sequence as the first superposition position number; Obtain the position number of the target dialogue text in the current dialogue text sequence as the target position number, and calculate the difference between the target position number and the first out-of-place position number as the dialogue backtracking number; The ratio of the number of dialogue rollbacks to the total number of text sequences is used to obtain the configuration coefficient of the first agent.

[0036] First, multiple historical dialogue texts and the current target dialogue text are sorted according to their chronological order to form an ordered sequence of current dialogue texts. The total number of dialogue texts in this current dialogue text sequence is then counted, i.e., the total number of text sequences. Each dialogue text in the current dialogue text sequence is assigned a position number starting from 1 and increasing sequentially.

[0037] Secondly, based on the sorted current dialogue text sequence, the previously calculated semantic consistency scores of multiple texts are simultaneously sorted to form a semantic consistency score sequence. This semantic consistency score sequence reflects the trend of semantic similarity changes between the user's continuous questions from historical dialogues to the current dialogue and historical content.

[0038] Then, from the obtained semantic consistency sequence, find the first text with a semantic consistency value greater than or equal to the consistency threshold. Determine the dialogue text corresponding to this semantic consistency value, and obtain the position label of this dialogue text in the current dialogue text sequence. Record this position label as the first superposition label. This first superposition label identifies the exact position where the user first raises a question highly similar to the current target question in the current session.

[0039] Furthermore, the positional label of the current target dialogue text itself is obtained in the current dialogue text sequence and used as the target positional label. The difference between the target positional label and the first surpassing positional label is calculated, and this difference is defined as the dialogue backtracking number. Specifically, the dialogue backtracking number intuitively represents the number of dialogue rounds experienced from the user's first asking of a similar question to the current follow-up question. The larger the dialogue backtracking number, the more rounds of other dialogues the user experienced after initially asking a similar question before returning to the question. This usually means that the previous responses failed to effectively resolve the user's confusion, and the problem complexity or user dissatisfaction is high. Therefore, more intelligent agent resources need to be configured for in-depth analysis and collaborative optimization to generate more fundamental responses.

[0040] Finally, the ratio of the number of backtracking responses to the total number of text sequences is calculated, and this ratio is used as the first agent configuration coefficient. This first agent configuration coefficient normalizes and quantifies the intensity of repeated user queries. A higher first agent configuration coefficient indicates that the user spends a higher proportion of the total dialogue rounds asking the same question, reflecting that past responses may not have met the user's needs. Therefore, more agent resources need to be configured for in-depth optimization and collaborative decision-making to generate higher-quality responses. Conversely, a lower first agent configuration coefficient indicates a lower intensity of repeated user queries, and resource investment can be reduced accordingly.

[0041] In summary, this mechanism enables the adaptive adjustment of the system's computing resource allocation strategy based on the objective intensity of the user's follow-up questions.

[0042] Secondly, the second agent is configured with coefficients based on the multiple historical reply texts, including: Based on the semantic consistency sequence, historical dialogue texts corresponding to texts with semantic consistency exceeding the consistency threshold are filtered to identify multiple dialogue texts that exceed the threshold. Based on the multiple out-of-limit dialogue texts, obtain the corresponding historical reply texts to form an out-of-limit reply text set; Keyword extraction is performed on each reply text in the set of reply texts that exceed the limit, resulting in multiple sets of reply keywords; Calculate the intersection and union of all reply keywords in the multiple reply keyword sets to obtain the number of intersection elements and the number of union elements; The configuration coefficient of the second agent is obtained based on the ratio of the number of intersection elements to the number of union elements.

[0043] First, based on the sorted semantic consistency sequence, all texts with semantic consistency scores exceeding the consistency threshold are filtered out. The corresponding historical dialogue texts for each of these semantic consistency scores are then located and identified as multiple over-limit dialogue texts. Over-limit dialogue texts represent all past questions in the current conversation history that are judged to be highly semantically similar to the current target question.

[0044] Furthermore, based on multiple over-limit dialogue texts, the historical response texts generated for each question at that time are retrieved from the dialogue records, forming an over-limit response text set. This over-limit response text set contains all the responses the system has given in the past when faced with similar questions.

[0045] Secondly, keyword extraction is performed on each historical reply text in the over-limit reply text set. Specifically, a keyword set consisting of its core words is generated for each reply text, resulting in multiple reply keyword sets that correspond one-to-one with the over-limit reply text. This keyword extraction process can be implemented using mature keyword extraction algorithms in the field of natural language processing. For example, a word frequency-inverse document frequency statistical method can be used, combined with part-of-speech filtering rules, to identify nouns and technical terms as keywords; or a graph model method can be used to construct a word co-occurrence network from the text and select key nodes as keywords based on node centrality indicators.

[0046] Furthermore, a comprehensive analysis is performed on all sets of response keywords. The intersection and union of all response keywords across multiple sets are calculated, and the number of elements in the intersection and union are counted respectively. The ratio of the number of elements in the intersection to the number of elements in the union is then calculated and used as the configuration coefficient for the second agent.

[0047] Specifically, this ratio essentially measures the similarity of multiple response texts, representing the overall homogenization of all historical responses exceeding the limit at the keyword level. Specifically, the higher the configuration coefficient of this second agent, the greater the overlap of core keywords among multiple historical responses to similar questions. This indicates highly homogenized response content, lacking diversity and novel perspectives, and potentially failing to comprehensively cover different aspects of the problem or meet the user's deeper needs. Therefore, when homogenization is detected, more agents need to be configured to break stereotypes through multi-angle simulation and adversarial mechanisms, generating optimized responses with richer content and more diverse perspectives.

[0048] Conversely, if the configuration coefficient of the second agent is low, it indicates that the historical responses have good diversity and the system is currently performing well in handling similar problems. Therefore, the additional agent resources invested to combat homogenization can be reduced accordingly.

[0049] In summary, this mechanism enables adaptive adjustment of computational resource allocation for optimization strategies based on the specific quality defects in historical responses.

[0050] Furthermore, after obtaining the configuration coefficients of the first and second agents, it is necessary to perform a weighted fusion of the first and second agent configuration coefficients to obtain a comprehensive agent configuration coefficient that integrates the two factors of the intensity of user follow-up behavior and the evaluation of the quality defects of historical responses.

[0051] Specifically, the configuration coefficients of the first and second intelligent agents are weighted to obtain the comprehensive configuration coefficients of the intelligent agents, including: Obtain preset weighting coefficients, wherein the preset weighting coefficients include a first weighting coefficient and a second weighting coefficient; Based on the first weight coefficient and the second weight coefficient, the first agent configuration coefficient and the second agent configuration coefficient are weighted and summed to obtain the comprehensive agent configuration coefficient.

[0052] First, preset weighting coefficients need to be obtained. These preset weighting coefficients include two specific values: a first weighting coefficient and a second weighting coefficient. The first weighting coefficient is used to adjust the contribution of the first agent's configuration coefficient in the final decision, and it is mainly related to the resource requirements reflected by the intensity of user follow-up inquiries. The second weighting coefficient is used to adjust the contribution of the second agent's configuration coefficient in the final decision, and it is mainly related to the optimization requirements reflected by the homogeneity of historical responses.

[0053] For example, the first weighting coefficient can be set to 0.6, and the second weighting coefficient can be set to 0.4. This setting means that in the comprehensive decision-making process, the consideration of the user's current follow-up behavior accounts for 60%, while the reflection on the quality of the system's historical responses accounts for 40%. Specifically, the weighting allocation can be adjusted according to different application scenarios to emphasize different aspects of immediate response and long-term optimization. For example, in customer service scenarios that require rapid response, the first weighting coefficient can be increased, while in consultation scenarios that emphasize the rigor and accumulation of answers, the second weighting coefficient can be increased.

[0054] Furthermore, based on the obtained first and second weighting coefficients, the configuration coefficients of the first and second agents are weighted respectively, and the weighted results are summed. The formula for the weighted summation is: Comprehensive agent configuration coefficient = First agent configuration coefficient × First weighting coefficient + Second agent configuration coefficient × Second weighting coefficient. Through this calculation, the comprehensive agent configuration coefficient is finally obtained.

[0055] The intelligent agent comprehensive configuration coefficient is a quantitative indicator that combines the urgency of user needs in the current dialogue context with the degree of deficiencies in the system's historical performance. Its value directly determines the total amount of intelligent agent resources that need to be mobilized subsequently, providing a precise and unified basis for dynamic resource allocation.

[0056] Furthermore, after obtaining the comprehensive configuration coefficients of the agents, it is also necessary to connect the agent swarm to obtain the total number of agents in the swarm. Specifically, this agent swarm is a pre-built, dynamically invoked distributed computing resource pool, comprising multiple fully functional and independent agents. Each agent encapsulates a lightweight language generation model and a two-dimensional response evaluator, possessing the complete capability to independently execute from text understanding to candidate response generation and evaluation, and reserves standardized communication interfaces for collaborative invocation.

[0057] Specifically, the group of intelligent agents includes multiple intelligent agents, and the construction steps of each intelligent agent include: A lightweight language generation model is obtained by compressing and simplifying the parameters of a pre-trained large language model to serve as a response generator. Construct a response evaluator, which includes a relevance evaluation branch and a difference evaluation branch. The relevance evaluation branch is used to evaluate the degree of matching between the candidate response text and the target dialogue text, and the difference evaluation branch is used to evaluate the degree of difference between the candidate response text and the set of response texts that exceed the limit. The response generator and response evaluator are integrated and encapsulated to form a single intelligent agent.

[0058] First, based on a pre-trained large language model, model compression and parameter simplification techniques, such as knowledge distillation, weight pruning, or quantization, are used to reduce the model size and improve inference speed while preserving its core language understanding and generation capabilities as much as possible. The resulting lightweight language generation model serves as one of the core components of the agent, namely the response generator. This response generator is specifically responsible for quickly generating a candidate response text based on the input target dialogue text. The pre-trained large language model is a foundational AI model with powerful general language understanding and generation capabilities, such as the GPT series, LLaMA series, or ERNIE series models.

[0059] Secondly, a response evaluator is constructed. Specifically, this evaluator is a modular structure with two-dimensional evaluation capabilities, containing two functionally independent branches. The first branch is a relevance evaluation branch, which assesses the semantic matching and relevance between candidate response texts and the current target dialogue text, ensuring that the response content closely addresses the user's question. The second branch is a difference evaluation branch, which assesses the degree of difference between candidate response texts and the historical set of responses that exceed the limits, aiming to encourage the generation of answers that go beyond historical responses, provide new information or perspectives, and avoid content homogenization.

[0060] Finally, the lightweight language generation model, i.e., the response generator, which has been constructed as described above, and the response evaluator, which has dual-dimensional evaluation capabilities, are integrated into a tightly cooperating software unit through software encapsulation technology. This unit provides a unified calling interface to the outside world and realizes a closed loop of generation and evaluation processes internally, thus forming a fully functional, independently operating single intelligent agent. A large number of such intelligent agents together constitute an intelligent agent group that can be dynamically invoked by the system.

[0061] Furthermore, the specific number of agents to be configured is determined based on the agent comprehensive configuration coefficient and the total number of agents. Specifically, the agent comprehensive configuration coefficient is multiplied by the total number of agents, and the result is rounded to obtain the final number of agents required, i.e., the agent configuration number. Based on this agent configuration number, a corresponding number of agents are randomly selected from the agent group, and the selected agents are combined to form the target agent set for performing this response generation task.

[0062] In summary, through the above series of configuration steps, the specific number of intelligent agents to be deployed is first calculated based on the weighted and fused comprehensive configuration coefficient of the intelligent agents and the total number of intelligent agents. Then, according to the number of intelligent agents, a corresponding number of intelligent agents are randomly selected from the intelligent agent group to complete the final configuration of the target intelligent agent set. This realizes the dynamic and proportional precise allocation of system computing resources based on the complexity of the problem reflected in the current dialogue context, the urgency of the user's follow-up questions, and the defects in the quality of historical responses.

[0063] S30: Based on the target dialogue text, perform response simulation adversarial decision-making through the target intelligent agent set, generate target response text, and send the target response text to the user terminal.

[0064] Specifically, based on the target dialogue text, a response simulation adversarial decision is made through the target intelligent agent set to generate the target response text, including: The target intelligent agent set includes multiple selected intelligent agents, and the target dialogue text is distributed to each of the multiple selected intelligent agents in the target intelligent agent set, wherein each selected intelligent agent includes a response generator and a response evaluator; Each selected agent activates its corresponding response generator and response evaluator, and generates multiple candidate response texts through multi-round simulation adversarial optimization. The candidate response texts are scored and filtered to determine the candidate response text with the highest score, which is then used as the target response text.

[0065] First, the target dialogue text to be processed is simultaneously distributed to each selected agent in the target agent set. Each selected agent is an independent functional unit, which encapsulates two core components: a response generator responsible for generating candidate response texts, and a response evaluator responsible for quality assessment and guided optimization of the generated texts.

[0066] Each selected agent initiates its internal workflow in parallel. Each agent's response generator independently generates an initial candidate response text based on the received target dialogue text. Then, the agent's response evaluator performs a two-dimensional quality check on this candidate response text. Specifically, the evaluator's relevance evaluation branch calculates the semantic matching degree between the candidate response text and the target dialogue text, outputting a relevance score; the difference evaluation branch calculates the semantic difference between the candidate response text and the historical set of out-of-bounds response texts, outputting a difference score.

[0067] Specifically, when the relevance score falls below a preset relevance score threshold, the evaluator generates a relevance enhancement instruction to guide the response generator in improving the match between candidate responses and the target question. This includes supplementing key information points, adjusting semantic emphasis, and strengthening the expression of core viewpoints. The preset relevance score threshold is a critical value used to determine whether a response is sufficiently relevant to the topic. It is set according to the specific application scenario's requirements for accuracy and relevance of the answer. For example, in a medical consultation scenario requiring high precision, it can be set to 0.9; while in a general dialogue scenario, it can be set to 0.75.

[0068] When the difference score falls below a preset difference score threshold, the evaluator generates difference enhancement instructions to guide the response generator to avoid repetitive content with the excess response text set. This includes replacing frequently repeated words, adjusting expression methods, and adding novel perspectives. The preset difference score threshold is a critical value used to determine whether a response possesses sufficient novelty. It is set according to the scenario's requirements for answer diversity and innovation. For example, in educational scenarios that encourage creative solutions, it can be set to 0.6; while in technical support scenarios requiring stable and reliable information, it can be set to 0.3.

[0069] Upon receiving a two-dimensional optimization instruction, the response generator adjusts its generation strategy and internal parameters based on the instruction, specifically improving and reconstructing the current candidate response text. The improved candidate response is then submitted again to the response evaluator of the same agent for a new round of two-dimensional testing, thus forming a closed-loop iterative process of evaluation, guidance, optimization, and re-evaluation. This simulation adversarial loop continues within each agent until its generated candidate responses achieve satisfactory quality levels in both relevance and difference dimensions, or until the preset maximum number of adversarial rounds is reached. The maximum number of adversarial rounds is an upper limit for controlling the number of iterations in the optimization process within a single agent. It is determined based on the system's response time constraints and computational resource budget for a single request. For example, in online customer service scenarios with high real-time interaction requirements, the maximum number of adversarial rounds can be set to 3 rounds; while in automated report generation scenarios that allow for longer thinking times, this number can be set to 10 rounds.

[0070] In summary, through the above mechanisms, each agent can ultimately produce a candidate response text that has been optimized through multiple rounds of internal adversarial processes, generating multiple candidate response texts.

[0071] Finally, the generated candidate response texts are scored and filtered.

[0072] Specifically, the multiple candidate response texts are scored and filtered to determine the candidate response text with the highest score, which is then used as the target response text. Each selected agent's response evaluator evaluates the quality of multiple candidate response texts and generates corresponding quality scores. The quality scores of each selected agent for the same candidate response text are summed to obtain the cumulative score of each candidate response text; Sort all candidate response texts by their cumulative scores and determine the candidate response text with the highest cumulative score.

[0073] First, after all selected agents in the target agent set have completed their internal simulation adversarial optimization loops and generated multiple candidate response texts, a cross-agent centralized evaluation process is initiated. Specifically, for each selected agent's response evaluator, all generated candidate responses are evaluated for quality. Each selected agent's response evaluator independently assesses the quality of each candidate response text based on its internally preset two-dimensional evaluation criteria, namely relevance and difference, and generates a corresponding, quantified quality score.

[0074] Secondly, for each candidate response text, the quality scores generated by the response evaluators of all selected agents are summed to obtain a cumulative score for that candidate response text. The cumulative score reflects the overall quality acceptance of the candidate response text from the perspective of the entire agent evaluation.

[0075] Finally, the cumulative scores of all candidate response texts are compared and ranked. Through this ranking process, the candidate response text with the highest cumulative score is identified. This candidate response text is the one collectively recognized as the best-performing response under the mechanism of multi-agent parallel generation and cross-evaluation. Ultimately, this candidate response text with the highest cumulative score is determined as the target response text and sent to the user's terminal.

[0076] In summary, the scoring and screening mechanism based on collective wisdom ensures that the final output target response text is high-quality content that has been tested and optimized from multiple perspectives.

[0077] In summary, the embodiments of this application have at least the following technical effects: Compared to existing technologies, this method firstly automatically identifies users' repetitive questioning behavior through semantic analysis and quantifies the intensity of follow-up inquiries and the homogeneity of historical responses, thereby precisely triggering targeted optimization mechanisms. Secondly, by introducing a dual-dimensional evaluation and a multi-round simulation adversarial internal optimization loop, it ensures that each agent's generated candidate responses are highly relevant to the user's current question while effectively avoiding duplication with historical low-quality or homogeneous content, achieving a dual improvement in response quality in terms of relevance and novelty. Thirdly, the adaptive resource allocation mechanism based on the dynamic analysis results of the dialogue context can intelligently allocate computing resources proportionally according to the complexity of the question, the intensity of user follow-up inquiries, and historical performance deficiencies, ensuring high-quality output while avoiding resource waste. Finally, through a cross-agent parallel generation and collective voting screening mechanism, it gathers multi-faceted and differentiated generation and evaluation wisdom, resulting in the final selected target response text having higher accuracy, comprehensiveness, and user satisfaction, thus systematically improving the service capabilities and reliability of the intelligent question-answering system in various complex interaction scenarios.

[0078] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0079] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0080] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A multi-agent simulation adversarial decision generation method based on a large language model, characterized in that, The method comprises: receiving a target dialogue text input by a user through a user terminal, and obtaining a dialogue record in which the target dialogue text is located, the dialogue record comprising a plurality of historical dialogue groups, each of the historical dialogue groups comprising historical dialogue text input by the user and historical reply text received by the user; comparing the target dialogue text with a plurality of historical dialogue texts respectively to obtain a plurality of text semantic consistency degrees, and when there is a text semantic consistency degree greater than or equal to a consistency threshold, configuring a target agent set according to the target dialogue text, the plurality of historical dialogue texts and the plurality of historical reply texts; generating a target reply text through reply simulation counter-decision by the target agent set according to the target dialogue text, and sending the target reply text to the user terminal.

2. The method of claim 1, wherein the method is based on a large language model. comparing the target dialogue text with a plurality of historical dialogue texts respectively to obtain a plurality of text semantic consistency degrees, comprising: performing text feature extraction processing on the target dialogue text and the plurality of historical dialogue texts respectively to obtain a target text feature vector and a plurality of historical text feature vectors; extracting a first historical text feature vector from the plurality of historical text feature vectors, inputting the first historical text feature vector and the target text feature vector into a semantic analyzer to obtain a first text semantic consistency degree; obtaining text semantic consistency degrees corresponding to the remaining historical text feature vectors in the manner of obtaining the first text semantic consistency degree corresponding to the first historical text feature vector to obtain the plurality of text semantic consistency degrees.

3. The method of claim 2, wherein the method further comprises: The construction steps of the semantic analyzer comprise: collecting a historical dialogue record set, extracting all historical dialogue texts for each historical dialogue record in the historical dialogue record set, and pairing the extracted historical dialogue texts two by two to construct a sample feature vector pair set, wherein each sample feature vector pair comprises two text feature vectors; annotating each sample feature vector pair in the sample feature vector pair set for semantic consistency degree to obtain a sample semantic consistency degree set; performing supervised training until convergence with the sample feature vector pair set as input features and the sample semantic consistency degree set as supervised labels to obtain the semantic analyzer.

4. The method of claim 1, wherein, Configuring a target agent set according to a target dialogue text, a plurality of historical dialogue texts and a plurality of historical reply texts comprises: configuring a first agent configuration coefficient based on the target dialogue text and the plurality of historical dialogue texts, and configuring a second agent configuration coefficient based on the plurality of historical reply texts; performing weighted processing on the first agent configuration coefficient and the second agent configuration coefficient to obtain an agent comprehensive configuration coefficient; connecting an agent group to obtain a total number of agents of the agent group; determining an agent configuration number according to the agent comprehensive configuration coefficient and the total number of agents, and randomly selecting a plurality of selected agents from the agent group based on the agent configuration number to configure the target agent set.

5. The method of claim 4, wherein the method is based on a large language model. Performing weighted processing on the first agent configuration coefficient and the second agent configuration coefficient to obtain an agent comprehensive configuration coefficient comprises: obtaining a preset weight coefficient, the preset weight coefficient comprising a first weight coefficient and a second weight coefficient; The first agent configuration coefficient and the second agent configuration coefficient are weighted and summed based on the first weight coefficient and the second weight coefficient to obtain an agent comprehensive configuration coefficient.

6. The method of claim 4, wherein the method is based on a large language model. The agent group includes a plurality of agents, and the construction step of each agent includes: A lightweight language generation model is obtained as a reply generator by model compression and parameter reduction based on a pre-trained large language model; A reply evaluator is constructed, the reply evaluator including a relevance evaluation branch and a difference evaluation branch, the relevance evaluation branch being used to evaluate the matching degree of the candidate reply text and the target dialogue text, and the difference evaluation branch being used to evaluate the difference degree of the candidate reply text and the ultra-limit reply text set; The reply generator and the reply evaluator are integrated and packaged to form a single agent.

7. The method of claim 4, wherein the method is based on a large language model. The first agent configuration coefficient is configured based on the target dialogue text and the plurality of historical dialogue texts, including: The plurality of historical dialogue texts and the target dialogue text are sorted in chronological order to obtain a current dialogue text sequence, and the total number of text sequences is determined, wherein each dialogue text in the current dialogue text sequence has a position label; The plurality of text semantic consistencies are sorted based on the current dialogue text sequence to obtain a semantic consistency sequence; The first text semantic consistency greater than or equal to the consistency threshold is found from the semantic consistency sequence, and the position label of the corresponding dialogue text in the current dialogue text sequence is determined as the first ultra-limit position label; The position label of the target dialogue text in the current dialogue text sequence is obtained as the target position label, and the difference between the target position label and the first ultra-limit position label is calculated as the dialogue backoff number; The ratio of the dialogue backoff number to the total number of text sequences is calculated to obtain the first agent configuration coefficient.

8. The method of claim 7, wherein the method is based on a large language model. The second agent configuration coefficient is configured based on the plurality of historical reply texts, including: Based on the semantic consistency sequence, the historical dialogue texts corresponding to the text semantic consistencies exceeding the consistency threshold are filtered to determine a plurality of ultra-limit dialogue texts; The corresponding historical reply texts are obtained from the plurality of ultra-limit dialogue texts to form an ultra-limit reply text set; Key words are extracted from each reply text in the ultra-limit reply text set to obtain a plurality of reply key word sets; The intersection and union of all reply key words in the plurality of reply key word sets are calculated to obtain the number of intersection elements and the number of union elements; The ratio of the number of intersection elements to the number of union elements is obtained to obtain the second agent configuration coefficient.

9. The method of claim 1, wherein, The target dialogue text is simulated by the target agent set to generate a target reply text, including: The target agent set includes a plurality of selected agents, and the target dialogue text is distributed to each of the plurality of selected agents in the target agent set, wherein each selected agent includes a reply generator and a reply evaluator; Each selected agent starts the corresponding reply generator and reply evaluator to generate a plurality of candidate reply texts through multiple rounds of simulation and optimization; The plurality of candidate reply texts are scored and filtered to determine the candidate reply text with the highest score as the target reply text.

10. The method of claim 9, wherein the method is based on a large language model. The multiple candidate reply texts are scored and screened to determine the candidate reply text with the highest score as the target reply text, comprising: The reply evaluators of the selected agents respectively perform quality evaluation on the multiple candidate reply texts to generate corresponding quality scores; The quality scores of the same candidate reply text from the selected agents are accumulated to obtain the accumulated score of each candidate reply text; The accumulated scores of all candidate reply texts are sorted to determine the candidate reply text with the highest accumulated score.