Information reasoning method and device

By introducing a multi-agent debate framework in the large language model, task debate and collaboration between agents is solved, and the problem of limited capabilities of a single model in complex reasoning tasks is improved, and the accuracy and efficiency of reasoning are improved.

CN119940528AActive Publication Date: 2025-05-06SOUTH CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411881154.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-06
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

A single large language model is susceptible to model bias and thinking degradation in complex inference tasks, and the lack of external feedback makes it difficult to form multiple insights, resulting in limited capabilities when facing complex tasks.

Method used

The multi-agent debate framework is adopted to promote collaboration through task debate among agents, and the team joint actions of the learnt agent decision network are guided to the update of individual agent strategies, emphasizing behavioral differences between agents, so that each agent can contribute incrementally to task solutions according to the preset role identity.

Benefits of technology

It improves the reasoning accuracy and efficiency of the agent's decision-making network, so that each agent can collaborate more effectively in complex tasks, form diverse insights, and enhances the understanding and resolution of problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940528A_ABST
    Figure CN119940528A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large-scale language models, in particular to an information reasoning method and device, computer equipment and a storage medium, which utilize a multi-agent debate framework, promote cooperation through task debate among agents, learn team joint actions of an agent decision network to guide updating of a single agent strategy, and improve the strategy updating efficiency. The behavior difference between the agents is emphasized, so that each agent can incrementally make a contribution to a task solution according to the respective preset role identity when facing a complex task, and the reasoning accuracy and efficiency of the agent decision network are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large-scale language models, and in particular to an information reasoning method, device, computer equipment and storage medium. Background Art

[0002] Multi-agent debate has broad application prospects in the field of artificial intelligence, especially using the generative power of large language models to solve various tasks such as code programming, teamwork projects, bargaining, mathematical problems, etc.

[0003] Large language models have demonstrated superior capabilities in a wide range of tasks, but they often fail in complex reasoning. Currently, a single model is susceptible to model bias, thinking degradation and other issues during reasoning, and due to the lack of external feedback, it is difficult to form multiple insights into the problem, thus limiting its ability to face complex tasks. Summary of the invention

[0004] Based on this, the purpose of the present invention is to provide an information reasoning method, apparatus, computer equipment and storage medium, which utilize a multi-agent debate framework to promote collaboration through task debates between agents, learn the team joint actions of the agent decision network to guide the update of a single agent strategy, emphasize the behavioral differences between agents, and enable each agent to incrementally contribute to the task solution according to its preset role identity when facing complex tasks, thereby improving the accuracy and efficiency of the reasoning of the agent decision network.

[0005] In a first aspect, an embodiment of the present application provides an information reasoning method, wherein an agent decision network includes a plurality of agents; the method comprises the following steps:

[0006] Obtaining state space information of several agents in the current time slot of the agent decision network and preset role prompt information, wherein the state space information of the current time slot includes question text information and agent conversation history information;

[0007] According to the state space information and role prompt information of the several agents and the corresponding current time slot, the action space information of the several agents in the current time slot is obtained; according to the several agents and the corresponding action space information of the current time slot, the task debate is performed to obtain the state space information of the several agents in the next time slot, wherein the action space information is the answer text information generated by the agent based on the state space information;

[0008] Reward calculation is performed according to the action space information of the multiple agents in the current time slot to obtain the reward information of the multiple agents in the current time slot; the role prompt information of the multiple agents, the corresponding state space information of the current time slot, the reward information and the state space information of the next time slot are combined to construct the training information combination of the multiple agents in the current time slot;

[0009] According to the training information combination of several agents in the current time slot, the agent decision network is updated, and according to the several agents in the updated agent decision network and the corresponding role prompt information and the state space information of the next time slot, the training information combination of several agents in the next time slot is repeatedly constructed to update the agent decision network, and the agent decision network updated last time is used as the target agent decision network;

[0010] Obtain text information of the problem to be processed, input the text information of the problem to be processed into several agents in the target agent decision network respectively, iterate repeatedly according to a preset number of iterations, obtain answer text information output by several agents in the target agent decision network at the last iteration number, and use the answer text information with the highest frequency as the inference result of the text information of the problem to be processed.

[0011] In a second aspect, an embodiment of the present application provides an information reasoning device based on an agent decision network, comprising:

[0012] A data acquisition module, used to obtain state space information of several agents in the current time slot of the agent decision network and preset role prompt information, wherein the state space information of the current time slot includes question text information and agent conversation history information;

[0013] The task debate model is used to obtain the action space information of several agents in the current time slot according to the state space information and role prompt information of several agents and the corresponding current time slot; perform task debate according to the several agents and the corresponding action space information of the current time slot to obtain the state space information of several agents in the next time slot, wherein the action space information is the answer text information generated by the agent based on the state space information;

[0014] A training information combination construction module is used to calculate rewards according to the action space information of the multiple agents in the current time slot to obtain the reward information of the multiple agents in the current time slot; the role prompt information of the multiple agents, the corresponding state space information of the current time slot, the reward information and the state space information of the next time slot are combined to construct the training information combination of the multiple agents in the current time slot;

[0015] A decision network updating module is used to update the agent decision network according to the training information combination of the multiple agents in the current time slot, and repeatedly construct the training information combination of the multiple agents in the next time slot according to the multiple agents in the updated agent decision network and the corresponding role prompt information and the state space information of the next time slot, and update the agent decision network, and use the last updated agent decision network as the target agent decision network;

[0016] The information reasoning module is used to obtain the text information of the problem to be processed, input the text information of the problem to be processed into several agents in the target agent decision network respectively, iterate repeatedly according to a preset number of iterations, obtain the answer text information output by several agents in the target agent decision network at the last iteration number, and use the answer text information with the highest frequency as the reasoning result of the text information of the problem to be processed.

[0017] In a third aspect, an embodiment of the present application provides a computer device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor; when the computer program is executed by the processor, the steps of the information reasoning method described in the first aspect are implemented.

[0018] In a fourth aspect, an embodiment of the present application provides a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the information reasoning method described in the first aspect are implemented.

[0019] In an embodiment of the present application, an information reasoning method, apparatus, computer device and storage medium are provided, which utilize a multi-agent debate framework to promote collaboration through task debates between agents, learn the team joint actions of the agent decision network to guide the update of individual agent strategies, emphasize the behavioral differences between agents, and enable each agent to incrementally contribute to the task solution according to its preset role identity when faced with complex tasks, thereby improving the accuracy and efficiency of the reasoning of the agent decision network.

[0020] For better understanding and implementation, the present invention is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A flowchart of an information reasoning method provided for one embodiment of the present application;

[0022] Figure 2 A schematic diagram of the process of S2 in the information reasoning method provided in one embodiment of the present application;

[0023] Figure 3A schematic diagram of the process of S3 in the information reasoning method provided in one embodiment of the present application;

[0024] Figure 4 A schematic diagram of the process of S4 in the information reasoning method provided in one embodiment of the present application;

[0025] Figure 5 A schematic diagram of the process of S42 in the information reasoning method provided in one embodiment of the present application;

[0026] Figure 6 A schematic diagram of the flow of S4 in the information reasoning method provided in another embodiment of the present application;

[0027] Figure 7 A schematic diagram of the structure of an information reasoning device based on an agent decision network provided in one embodiment of the present application;

[0028] Figure 8 A schematic diagram of the structure of a computer device provided for one embodiment of the present application. DETAILED DESCRIPTION

[0029] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0030] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0031] It should be understood that, although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" / "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determination".

[0032] The executor of the information reasoning method is the reasoning device of the information reasoning method (hereinafter referred to as the reasoning device). The reasoning device can be implemented by software and / or hardware, and the information reasoning method can be implemented by software and / or hardware. The reasoning device can be composed of two or more physical entities, or it can be composed of one physical entity. The hardware pointed to by the reasoning device essentially refers to computer equipment. For example, the reasoning device can be a computer, a mobile phone, a tablet or an interactive tablet. In an optional embodiment, the reasoning device can specifically be a server, or a server cluster composed of multiple computer devices.

[0033] See also Figure 1 , Figure 1 A flowchart of an information reasoning method provided in one embodiment of the present application, the method comprising the following steps:

[0034] S1: Obtain the state space information of several agents in the current time slot of the agent decision network and the preset role prompt information.

[0035] The agent decision network includes a plurality of agents, in which each agent will perform multiple rounds of debate and collaboration on the same task, and each round is a time slot. The agent is a model built on the basis of a large language model (LLM) architecture.

[0036] In this embodiment, the reasoning device obtains the state space information of several agents in the current time slot of the agent decision network and the preset role prompt information, wherein the state space information of the current time slot includes the question text information and the agent conversation history information; the agent conversation history information is the text information record generated by the agent in the task debate with other agents, which is recorded as {u 0 ,u 1 ,...,u t-1},u 0 is the agent conversation history information of the initial time slot, u t-1 is the agent conversation history information of the t-1th time slot.

[0037] The role prompt information includes: "Take the opinions of other intelligent agents as supplementary suggestions", "Please stick to your own views in the debate" and "Focus on the answers of other intelligent agents as references" to activate the intelligent agents as diverse roles, which promotes effective collaboration within the intelligent agent decision-making network in the current environment.

[0038] S2: Based on the state space information and role prompt information of the several intelligent agents and the corresponding current time slot, the action space information of the several intelligent agents in the current time slot is obtained; based on the several intelligent agents and the corresponding action space information of the current time slot, task debate is performed to obtain the state space information of the several intelligent agents in the next time slot.

[0039] In this embodiment, the reasoning device obtains the action space information of several agents in the current time slot according to the state space information and role prompt information of several agents and the corresponding current time slot. Specifically, the action space information is the answer text information generated by the agent based on the state space information and the role prompt information, which is recorded as is the action space information of the ith agent in the tth time slot, w n The token indexed at the nth position indicates the character corresponding to the index at the nth position in the answer text information.

[0040] The reasoning device performs task debate based on several intelligent agents and the corresponding action space information of the current time slot to obtain the state space information of several intelligent agents in the next time slot.

[0041] The agent includes a role-aware network; see Figure 2 , Figure 2 The flowchart of S2 in the information reasoning method provided in one embodiment of the present application includes steps S21 to S22, which are as follows:

[0042] S21: performing embedding processing according to the plurality of intelligent agents and corresponding role prompt information to obtain role embedding information of the plurality of intelligent agents.

[0043] The role perception network is a RNN convolutional neural network with role guidance that can enhance role differentiation.

[0044] In this embodiment, the reasoning device performs embedding processing based on the plurality of the agents and the corresponding role prompt information to obtain the role embedding information of the plurality of the agents, so as to enhance the unique features associated with the agents and the preset roles.

[0045] S22: inputting the role embedding information of the several intelligent agents and the corresponding state space information of the current time slot into the role perception network in the corresponding intelligent agent respectively, and obtaining the action space information of the several intelligent agents in the current time slot according to the preset action space information generation algorithm.

[0046] The action space information generation algorithm is:

[0047]

[0048] In the formula, is the action space information of the ith agent in the tth time slot, is the state space information of the ith agent in the state space information of the tth time slot, e i Embedding information for the role of the i-th agent, RNN i (·) is the processing function of the role perception network of the ith agent, π i (·) is the policy function in the role perception network of the ith agent.

[0049] In this embodiment, the inference device inputs the role embedding information of several of the agents and the state space information of the corresponding current time slot into the role perception network in the corresponding agent, and obtains the action space information of several agents in the current time slot according to a preset action space information generation algorithm.

[0050] S3: Calculate rewards based on the action space information of several agents in the current time slot to obtain reward information of several agents in the current time slot; combine the role prompt information of several agents, the corresponding state space information of the current time slot, reward information and the state space information of the next time slot to construct a training information combination of several agents in the current time slot.

[0051] In this embodiment, the reasoning device calculates rewards based on the action space information of several agents in the current time slot to obtain reward information of several agents in the current time slot. The reward information can be used as feedback to improve the LLM strategy, thereby enhancing the consistency of its actions with expected results.

[0052] The reasoning device combines the role prompt information of several of the agents, the corresponding state space information of the current time slot, the reward information and the state space information of the next time slot to construct a training information combination of several agents in the current time slot.

[0053] See also Figure 3 , Figure 3 The flowchart of S3 in the information reasoning method provided in one embodiment of the present application includes steps S31 to S32, which are specifically as follows:

[0054] S31: Obtain standard answer information of several of the intelligent agents.

[0055] In this embodiment, the reasoning device obtains standard answer information of several of the intelligent agents, wherein the standard answer information is answer text information corresponding to the question text information in the corresponding state space information.

[0056] S32: Based on the action space information of several intelligent agents in the current time slot and the standard answer information, a cosine similarity calculation method is used to obtain the cosine similarity between the action space information of several intelligent agents in the current time slot and the standard answer information as the reward information.

[0057] In this embodiment, the reasoning device uses a cosine similarity calculation method to obtain the cosine similarity between the action space information of several intelligent agents in the current time slot and the standard answer information based on the action space information of several intelligent agents in the current time slot and the standard answer information, as the reward information. A higher cosine similarity indicates that the answer text information matches the standard answer information more closely, and thus a higher reward is obtained.

[0058] S4: According to the training information combination of several agents in the current time slot, the agent decision network is updated, and according to the several agents in the updated agent decision network and the corresponding role prompt information and the state space information of the next time slot, the training information combination of several agents in the next time slot is repeatedly constructed to update the agent decision network, and the last updated agent decision network is used as the target agent decision network;.

[0059] In this embodiment, the inference device updates the agent decision network according to the combination of training information of several agents in the current time slot, and uses the last updated agent decision network as the target agent decision network.

[0060] The reasoning device repeatedly constructs the training information combination of several agents in the next time slot based on the updated several agents in the agent decision network and the corresponding role prompt information and the state space information of the next time slot, updates the agent decision network, improves the action space information output by the role perception network, and conducts efficient and high-quality task debate.

[0061] See also Figure 4 , Figure 4 The flowchart of S4 in the information reasoning method provided in one embodiment of the present application includes steps S41 to S42, which are specifically as follows:

[0062] S41: Obtain the action space information of several agents in the next time slot according to the state space information of the next time slot in the training information combination of the agent and the role prompt information; input the reward information of the current time slot in the training information combination of several agents in the current time slot and the state space information and action space information of the next time slot into the target network, and obtain the individual value parameters of several agents in the current time slot according to a preset individual value calculation algorithm.

[0063] In this embodiment, the inference device obtains the action space information of several agents in the next time slot based on the state space information of the next time slot in the training information combination of the agent and the role prompt information. The specific implementation method can refer to steps S21 to S22 and will not be repeated here.

[0064] The inference device inputs the reward information of the current time slot in the training information combination of the multiple agents in the current time slot and the state space information and action space information of the next time slot into the target network, and obtains the individual value parameters of the multiple agents in the current time slot according to the preset individual value calculation algorithm, wherein the individual value calculation algorithm is:

[0065]

[0066] In the formula, is the individual value parameter of the ith agent in the tth time slot, r is the reward information of the tth time slot, γ is the discount factor, E(·) is the expected calculation function, To find the maximum function, Q(·) is the target network function, is the state space information of the ith agent in the state space information of the t+1th time slot, is the action space information of the i-th agent in the t+1-th time slot.

[0067] S42: Calculate the global value parameters of the agent decision network of the current time slot based on the individual value parameters of the agents in the current time slot and the hybrid network; construct a first loss value using a reinforcement learning method based on the reward information of the agents in the current time slot and the global value parameters of the agent decision network, and update the role perception networks of the agents in the agent decision network based on the first loss value.

[0068] In this embodiment, the inference device calculates the global value parameters based on the individual value parameters of several intelligent agents in the current time slot and the hybrid network to obtain the global value parameters of the intelligent agent decision network in the current time slot, wherein the global value parameters include a first global value parameter and a second global value parameter.

[0069] The inference device constructs a first loss value using a reinforcement learning method according to the reward information of several agents in the current time slot and the global value parameters of the agent decision network, and updates the role perception networks of several agents in the agent decision network according to the first loss value.

[0070] See also Figure 5 , Figure 5The flowchart of S42 in the information reasoning method provided in one embodiment of the present application includes steps S421 to S423, which are as follows:

[0071] S421: Combine the individual value parameters of several intelligent agents in the current time slot to construct the first individual value parameter set of the intelligent agent decision network in the current time slot; compare the individual value parameters of several intelligent agents in the current time slot to obtain the maximum individual value parameter of the current time slot, replace the individual value parameters of several intelligent agents in the current time slot with the maximum individual value parameter, combine the replaced individual value parameters of several intelligent agents in the current time slot to construct the second individual value parameter set of the intelligent agent decision network in the current time slot.

[0072] In this embodiment, the inference device combines the individual value parameters of several intelligent agents in the current time slot to construct a first individual value parameter set of the intelligent agent decision network of the current time slot; compares the individual value parameters of several intelligent agents in the current time slot to obtain the maximum individual value parameter of the current time slot, replaces the individual value parameters of several intelligent agents in the current time slot with the maximum individual value parameter, combines the replaced individual value parameters of several intelligent agents in the current time slot to construct a second individual value parameter set of the intelligent agent decision network of the current time slot.

[0073] S422: respectively use the first individual value parameter set and the second individual value parameter set of the agent decision network of the current time slot as the input parameter set of the hybrid network, calculate the global value parameters according to the input parameter sets, and obtain the first global value parameter and the second global value parameter of the agent decision network of the current time slot.

[0074] The hybrid network includes several fully connected layers whose parameters are determined by the hypernetwork. The input of the hypernetwork is a natural language text H describing the problem to be solved. In order to map to the same vector space, H is first input into the critic network to obtain its vector representation H. This vector and a linear layer are then used to calculate the parameters w and b of the hybrid network. The consistency between each agent and all agents in the agent decision network is ensured by the hybrid network, ensuring that each agent effectively contributes to the team task of the agent decision network.

[0075] In this embodiment, the inference device uses the first individual value parameter set and the second individual value parameter set of the agent decision network of the current time slot as the input parameter set of the hybrid network, performs global value parameter calculation based on the input parameter set, and obtains the first global value parameter and the second global value parameter of the agent decision network of the current time slot.

[0076] In an optional embodiment, the reasoning device constructs a consistency constraint condition, calculates the global value parameter according to the input parameter set and the consistency constraint condition, obtains the first global value parameter and the second global value parameter of the agent decision network of the current time slot, so as to ensure that the obtained first global value parameter is proportional to the individual value parameters of several agents of the current time slot, and the obtained second global value parameter is proportional to the maximum individual value parameter of the current time slot, so as to indicate that when the individual value parameter of each agent reaches its maximum individual value, the obtained first global value parameter and the second global value parameter also reach their maximum individual value, wherein the consistency constraint condition is:

[0077]

[0078] In the formula, Q tot is the global value parameter, Q i is the individual value parameter of the ith agent.

[0079] S423: Accumulate the reward information of several agents in the current time slot to obtain the cumulative reward information of the agent decision network in the current time slot; obtain the first loss value according to the cumulative reward information of the agent decision network in the current time slot, the first global value parameter and the second global value parameter and the preset first loss algorithm, and update the role perception network of several agents in the agent decision network according to the first loss value.

[0080] In this embodiment, the inference device accumulates the reward information of several agents in the current time slot to obtain the cumulative reward information of the agent decision network in the current time slot.

[0081] The reasoning device obtains a first loss value according to the accumulated reward information of the agent decision network at the current time slot, the first global value parameter, the second global value parameter, and a preset first loss algorithm, and updates the role perception networks of several agents in the agent decision network according to the first loss value, wherein the first loss algorithm is:

[0082]

[0083] Where, L mix is, T is the total number of time slots, R t is the cumulative reward information of the agent decision network at the tth time slot, γ is the discount factor, is the first global value parameter of the agent decision network at the tth time slot, is the second global value parameter of the agent decision network at the tth time slot.

[0084] By introducing the loss calculated by the hybrid network to update the agent's actor network, the coordination between multiple agents within the agent decision network is promoted to promote the global optimum, ensuring that the action space information generated by each agent contributes optimally and consistently to solving the problem, thereby ensuring that each agent contributes incrementally to the task solution and improving the accuracy and efficiency of the task debate of the agent decision network.

[0085] The agent decision network also includes an inference network; see Figure 6 , Figure 6 The flowchart of S4 in the information reasoning method provided in another embodiment of the present application includes steps S43 to S45, which are specifically as follows:

[0086] S43: Obtain observation values ​​of several intelligent agents in the current time slot, input the observation values ​​of several intelligent agents in the current time slot into the inference network, and obtain the Gaussian distribution of several intelligent agents in the current time slot; sample the Gaussian distribution of several intelligent agents in the current time slot respectively, and obtain the potential variable set of the current time slot.

[0087] In order to enhance the respective expertise of each agent, in this embodiment, the inference device obtains the observation values ​​of several agents in the current time slot, inputs the observation values ​​of several agents in the current time slot into the inference network for Gaussian distribution modeling, and obtains a set of latent variables in the current time slot, wherein the set of latent variables includes latent variables of several agents, and the latent variables represent the unique characteristics of the agents, which are used for model training, enhance the robustness and uncertainty of the reasoning task, thereby enhancing the diversity between agents and differentiating them from each other.

[0088] S44: Obtain the dissimilarity parameters between the several agents in the current time slot according to the Gaussian distribution of the several agents in the current time slot, the latent variable set and the preset dissimilarity calculation algorithm.

[0089] The dissimilarity calculation algorithm is:

[0090]

[0091] Where D φ (i,j) is the dissimilarity parameter between the ith agent and the jth agent, KL(·) is the KL divergence calculation function, are the Gaussian distributions of the i-th and j-th agents, respectively, and z i 、z j are the latent variables of the ith and jth agents respectively, b and c are the first and second balance coefficients respectively.

[0092] In this embodiment, the inference device obtains the dissimilarity parameters between the several agents in the current time slot based on the Gaussian distribution of the several agents in the current time slot, the set of latent variables and the preset dissimilarity calculation algorithm, by calculating the KL divergence between the corresponding Gaussian distributions and the mutual information between the latent variables, so as to reflect the mutual dependence and influence between the agents.

[0093] S45: Obtain a second loss value based on the observation values ​​of several agents in the current time slot, action space information, Gaussian distribution, the difference parameters between several agents, a set of latent variables and a preset second loss algorithm, and update the role perception networks of several agents in the agent decision network based on the first loss value and the second loss value.

[0094] The second loss algorithm is:

[0095]

[0096] Where, L dis is the second loss value, w MI is the first weight parameter, is the latent variable of the i-th agent, and a i is the action space information of the ith agent, w KL is the first weight parameter, MI(·) is the MI divergence calculation function, p(·|·) is the conditional probability distribution calculation function, o i is the observation value of the ith agent, o j is the observation value of the jth agent, w DI is the second weight parameter, w H is the fourth weight parameter, H(Z) is the entropy of the latent variable set, and Z is the latent variable set.

[0097] In this embodiment, the inference device predicts the potential variable z according to the observation values ​​of the multiple agents in the current time slot, the action space information, the Gaussian distribution, the dissimilarity parameters between the multiple agents, the potential variable set and the preset second loss algorithm. i and action space information a i The mutual information between them estimates the coexistence between them and predicts the latent variable z i Its conditional probability distribution p(z i |o i ), strengthen the association between the latent variable distribution and the observed value, and combine the calculated dissimilarity D between several agents φ (o i ,o j ) and information entropy to obtain the second loss value.

[0098] The reasoning device updates the role perception networks of several agents in the agent decision network according to the first loss value and the second loss value. Specifically, the reasoning device accumulates the first loss value and the second loss value to obtain a total loss value, and updates the role perception networks of several agents in the agent decision network according to the total loss value, ensuring that the update process not only focuses on maximizing the utility of each agent strategy and the acquisition of individual value parameters, but also emphasizes the behavioral differences between agents, so that each agent can contribute to the task solution incrementally according to its preset role identity when facing complex tasks, thereby improving the accuracy and efficiency of the reasoning of the agent decision network.

[0099] S5: Obtain text information of the problem to be processed, input the text information of the problem to be processed into several agents in the target agent decision network respectively, iterate repeatedly according to a preset number of iterations, obtain the answer text information output by several agents in the target agent decision network at the last iteration number, and use the answer text information with the highest frequency as the inference result of the text information of the problem to be processed.

[0100] In this embodiment, the reasoning device obtains text information of the problem to be processed, wherein the text information of the problem to be processed may be input by a user or obtained from a preset database.

[0101] The inference device inputs the text information of the problem to be processed into the updated several agents in the agent decision network respectively, and iterates repeatedly according to the preset number of iterations to obtain the answer text information output by the updated several agents for the last number of iterations, and takes the answer text information with the highest frequency as the inference result of the text information of the problem to be processed.

[0102] By utilizing the multi-agent debate framework, we promote collaboration through task debates between agents, learn the team joint actions of the agent decision network to guide the update of individual agent strategies, and emphasize the behavioral differences between agents. This allows each agent to incrementally contribute to the task solution according to its preset role identity when faced with complex tasks, thereby improving the accuracy and efficiency of the agent decision network's reasoning.

[0103] Please refer to Figure 7 , Figure 7 A schematic diagram of the structure of an information reasoning device based on an agent decision network provided by an embodiment of the present application. The device can implement all or part of the information reasoning device based on the agent decision network through software, hardware, or a combination of both. The device 7 includes:

[0104] A data acquisition module 71 is used to obtain state space information of several agents in the current time slot of the agent decision network and preset role prompt information, wherein the state space information of the current time slot includes question text information and agent conversation history information;

[0105] The task debate model 72 is used to obtain the action space information of several agents in the current time slot according to the state space information of several agents and the corresponding current time slot and the role prompt information; perform task debate according to the action space information of several agents and the corresponding current time slot to obtain the state space information of several agents in the next time slot;

[0106] The training information combination construction module 73 is used to calculate the reward according to the action space information of the multiple agents in the current time slot to obtain the reward information of the multiple agents in the current time slot; combine the role prompt information of the multiple agents, the corresponding state space information of the current time slot, the reward information and the state space information of the next time slot to construct the training information combination of the multiple agents in the current time slot;

[0107] A decision network updating module 74 is used to update the agent decision network according to the training information combination of the multiple agents in the current time slot, and repeatedly construct the training information combination of the multiple agents in the next time slot according to the multiple agents in the updated agent decision network and the corresponding role prompt information and the state space information of the next time slot, so as to update the agent decision network;

[0108] The information reasoning module 75 is used to obtain the text information of the problem to be processed, input the text information of the problem to be processed into several agents in the target agent decision network respectively, iterate repeatedly according to a preset number of iterations, obtain the answer text information output by several agents in the target agent decision network at the last iteration number, and use the answer text information with the highest frequency as the reasoning result of the text information of the problem to be processed.

[0109] In an embodiment of the present application, the state space information of several agents in the current time slot of the agent decision network and the preset role prompt information are obtained through a data acquisition module, wherein the state space information of the current time slot includes question text information and agent conversation history information; the action space information of several agents in the current time slot is obtained according to the state space information and role prompt information of several agents and the corresponding current time slot through a task debate model; task debate is performed according to several agents and the corresponding action space information of the current time slot to obtain the state space information of several agents in the next time slot; reward calculation is performed according to the action space information of several agents in the current time slot through a training information combination construction module to obtain the reward information of several agents in the current time slot; the role prompt information of several agents, the corresponding state space information of the current time slot, the reward information and the next time slot are combined into a task debate model; The state space information of the target agent decision network is combined to construct a training information combination of several agents in the current time slot; the agent decision network is updated according to the training information combination of several agents in the current time slot through a decision network update module, and the training information combination of several agents in the next time slot is repeatedly constructed according to the several agents in the updated agent decision network and the corresponding role prompt information and the state space information of the next time slot, and the agent decision network is updated; an information reasoning module is used to obtain text information of the problem to be processed, input the text information of the problem to be processed into several agents in the target agent decision network respectively, iterate repeatedly according to a preset number of iterations, obtain the answer text information output by several agents in the target agent decision network of the last iteration number, and use the answer text information with the highest frequency as the reasoning result of the text information of the problem to be processed. By utilizing the multi-agent debate framework, we promote collaboration through task debates between agents, learn the team joint actions of the agent decision network to guide the update of individual agent strategies, and emphasize the behavioral differences between agents. This allows each agent to incrementally contribute to the task solution according to its preset role identity when faced with complex tasks, thereby improving the accuracy and efficiency of the agent decision network's reasoning.

[0110] Please refer to Figure 8 , Figure 8 The computer device 8 is a schematic diagram of a structure of a computer device provided in an embodiment of the present application. The computer device 8 includes: a processor 81, a memory 82, and a computer program 83 stored in the memory 82 and executable on the processor 81; the computer device may store multiple instructions, which are suitable for being loaded and executed by the processor 81. Figures 1 to 6 The method steps shown in the figure can be found in the specific execution process. Figures 1 to 6 The specific description shown will not be repeated here.

[0111] Among them, the processor 81 may include one or more processing cores. The processor 81 uses various interfaces and lines to connect various parts in the server, and executes various functions and processes data of the information reasoning device 7 based on the intelligent agent decision network by running or executing instructions, programs, code sets or instruction sets stored in the memory 82, and calling the data in the memory 82. Optionally, the processor 81 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programble Logic Array, PLA). The processor 81 can integrate one or more combinations of a central processing unit 81 (Central Processing Unit, CPU), an image processor 81 (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the touch display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 81, and it can be implemented by a single chip.

[0112] Among them, the memory 82 may include a random access memory 82 (Random Access Memory, RAM), and may also include a read-only memory 82 (Read-Only Memory). Optionally, the memory 82 includes a non-transitory computer-readable storage medium. The memory 82 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 82 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch instructions, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 82 may also be optionally at least one storage device located away from the aforementioned processor 81.

[0113] The present application also provides a storage medium that can store multiple instructions, which are suitable for the processor to load and execute the above-mentioned Figures 1 to 6 The method steps shown in the figure can be found in the specific execution process. Figures 1 to 6 The specific description shown will not be repeated here.

[0114] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0115] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0116] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraint algorithm of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0117] In the embodiments provided by the present invention, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0118] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0119] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0120] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc.

[0121] The present invention is not limited to the above-mentioned embodiments. If various changes or modifications to the present invention do not depart from the spirit and scope of the present invention, and if these changes and modifications fall within the scope of the claims and equivalent technologies of the present invention, the present invention is also intended to include these changes and modifications.

Claims

1. An information reasoning method, wherein the agent decision network includes a plurality of agents; characterized in that: The method comprises the following steps: Obtaining state space information of several agents in the current time slot of the agent decision network and preset role prompt information, wherein the state space information of the current time slot includes question text information and agent conversation history information; According to the state space information and role prompt information of the several agents and the corresponding current time slot, the action space information of the several agents in the current time slot is obtained; according to the several agents and the corresponding action space information of the current time slot, the task debate is performed to obtain the state space information of the several agents in the next time slot, wherein the action space information is the answer text information generated by the agent based on the state space information; Reward calculation is performed according to the action space information of the multiple agents in the current time slot to obtain the reward information of the multiple agents in the current time slot; the role prompt information of the multiple agents, the corresponding state space information of the current time slot, the reward information and the state space information of the next time slot are combined to construct the training information combination of the multiple agents in the current time slot; According to the training information combination of several agents in the current time slot, the agent decision network is updated, and according to the several agents in the updated agent decision network and the corresponding role prompt information and the state space information of the next time slot, the training information combination of several agents in the next time slot is repeatedly constructed to update the agent decision network, and the agent decision network updated last time is used as the target agent decision network; Obtain text information of the problem to be processed, input the text information of the problem to be processed into several agents in the target agent decision network respectively, iterate repeatedly according to a preset number of iterations, obtain answer text information output by several agents in the target agent decision network at the last iteration number, and use the answer text information with the highest frequency as the inference result of the text information of the problem to be processed.

2. The information reasoning method according to claim 1, characterized in that: The intelligent agent includes a role perception network, and the role perception network is a RNN convolutional neural network; The step of obtaining the action space information of the plurality of agents in the current time slot according to the state space information of the plurality of agents and the corresponding current time slot and the role prompt information comprises the following steps: Performing embedding processing according to the plurality of agents and corresponding role prompt information to obtain role embedding information of the plurality of agents; The role embedding information of several agents and the state space information of the corresponding current time slot are respectively input into the role perception network in the corresponding agent, and the action space information of several agents in the current time slot is obtained according to the preset action space information generation algorithm, wherein the action space information generation algorithm is: In the formula, is the action space information of the ith agent in the tth time slot, is the state space information of the ith agent in the state space information of the tth time slot, e i Embedding information for the role of the i-th agent, RNN i (·) is the processing function of the role perception network of the ith agent, π i (·) is the policy function in the role perception network of the ith agent.

3. The information reasoning method according to claim 2, characterized in that: The step of performing reward calculation according to the action space information of the plurality of agents in the current time slot to obtain the reward information of the plurality of agents in the current time slot comprises the following steps: Obtaining standard answer information of a plurality of the intelligent agents, wherein the standard answer information is answer text information corresponding to the question text information in the corresponding state space information; According to the action space information of several intelligent agents in the current time slot and the standard answer information, a cosine similarity calculation method is adopted to obtain the cosine similarity between the action space information of several intelligent agents in the current time slot and the standard answer information as the reward information.

4. The information reasoning method according to claim 3, characterized in that: The agent also includes a target network; the agent decision network includes a hybrid network, and the hybrid network includes a recurrent neural network and a fully connected layer; The updating of the agent decision network according to the combination of training information of several agents in the current time slot comprises the steps of: Obtaining action space information of several agents in the next time slot according to the state space information of the next time slot and the role prompt information in the training information combination of the agent; The reward information of the current time slot in the training information combination of the multiple agents in the current time slot and the state space information and action space information of the next time slot are input into the target network, and the individual value parameters of the multiple agents in the current time slot are obtained according to the preset individual value calculation algorithm, wherein the individual value calculation algorithm is: In the formula, is the individual value parameter of the ith agent in the tth time slot, r is the reward information of the tth time slot, γ is the discount factor, E(·) is the expected calculation function, To find the maximum function, Q(·) is the target network function, is the state space information of the ith agent in the state space information of the t+1th time slot, is the action space information of the ith agent in the t+1th time slot; The global value parameters are calculated according to the individual value parameters of several agents in the current time slot and the hybrid network to obtain the global value parameters of the agent decision network in the current time slot; the first loss value is constructed by using the reinforcement learning method according to the reward information of several agents in the current time slot and the global value parameters of the agent decision network, and the role perception networks of several agents in the agent decision network are updated according to the first loss value.

5. The information reasoning method according to claim 4, characterized in that: The global value parameters include a first global value parameter and a second global value parameter; The step of calculating the global value parameters according to the individual value parameters of the multiple agents in the current time slot and the hybrid network to obtain the global value parameters of the agent decision network in the current time slot; and constructing the first loss value by using the reinforcement learning method according to the reward information of the multiple agents in the current time slot and the global value parameters of the agent decision network, comprises the following steps: The individual value parameters of several intelligent agents in the current time slot are combined to construct a first individual value parameter set of the intelligent agent decision network in the current time slot; the individual value parameters of several intelligent agents in the current time slot are compared to obtain the maximum individual value parameter of the current time slot, the individual value parameters of several intelligent agents in the current time slot are replaced with the maximum individual value parameter, and the individual value parameters of several intelligent agents in the current time slot after replacement are combined to construct a second individual value parameter set of the intelligent agent decision network in the current time slot; The first body value parameter set and the second body value parameter set of the agent decision network of the current time slot are respectively used as the input parameter set of the hybrid network, and the global value parameter is calculated according to the input parameter set to obtain the first global value parameter and the second global value parameter of the agent decision network of the current time slot; The reward information of several agents in the current time slot is accumulated to obtain the cumulative reward information of the agent decision network in the current time slot; the first loss value is obtained according to the cumulative reward information of the agent decision network in the current time slot, the first global value parameter and the second global value parameter and the preset first loss algorithm, and the role perception network of several agents in the agent decision network is updated according to the first loss value, wherein the first loss algorithm is: Where, L mix is, T is the total number of time slots, R t is the cumulative reward information of the agent decision network at the tth time slot, γ is the discount factor, is the first global value parameter of the agent decision network at the tth time slot, is the second global value parameter of the agent decision network at the tth time slot.

6. The information reasoning method according to claim 5, characterized in that: The agent decision network also includes a reasoning network; The updating of the agent decision network according to the combination of training information of several agents in the current time slot also includes the steps of: Obtaining observation values ​​of several agents in the current time slot, inputting the observation values ​​of several agents in the current time slot into the inference network, and obtaining Gaussian distribution of several agents in the current time slot; sampling the Gaussian distribution of several agents in the current time slot respectively, and obtaining a latent variable set of the current time slot, wherein the latent variable set includes latent variables of several agents; According to the Gaussian distribution of the several agents in the current time slot, the potential variable set and the preset dissimilarity calculation algorithm, the dissimilarity parameters between the several agents in the current time slot are obtained, wherein the dissimilarity calculation algorithm is: Where D φ (i,j) is the dissimilarity parameter between the ith agent and the jth agent, KL(·) is the KL divergence calculation function, are the Gaussian distributions of the i-th and j-th agents, respectively, and z i 、z j are the latent variables of the ith and jth agents, b and c are the first and second balance coefficients, respectively; According to the observation values ​​of several agents in the current time slot, action space information, Gaussian distribution, the dissimilarity parameters between several agents, the potential variable set and the preset second loss algorithm, a second loss value is obtained, and according to the first loss value and the second loss value, the role perception network of several agents in the agent decision network is updated, wherein the second loss algorithm is: Where, L dis is the second loss value, w MI is the first weight parameter, z i is the latent variable of the ith agent, a i is the action space information of the ith agent, w KL is the first weight parameter, MI(·) is the MI divergence calculation function, p(·|·) is the conditional probability distribution calculation function, o i is the observation value of the ith agent, o j is the observation value of the jth agent, w DI is the second weight parameter, w H is the fourth weight parameter, H(Z) is the entropy of the latent variable set, and Z is the latent variable set.

7. An information reasoning device based on an agent decision network, wherein the agent decision network includes a plurality of agents, characterized in that: include: A data acquisition module, used to obtain state space information of several agents in the current time slot of the agent decision network and preset role prompt information, wherein the state space information of the current time slot includes question text information and agent conversation history information; The task debate model is used to obtain the action space information of several agents in the current time slot according to the state space information and role prompt information of several agents and the corresponding current time slot; perform task debate according to the several agents and the corresponding action space information of the current time slot to obtain the state space information of several agents in the next time slot, wherein the action space information is the answer text information generated by the agent based on the state space information; A training information combination construction module is used to calculate rewards according to the action space information of the multiple agents in the current time slot to obtain the reward information of the multiple agents in the current time slot; the role prompt information of the multiple agents, the corresponding state space information of the current time slot, the reward information and the state space information of the next time slot are combined to construct the training information combination of the multiple agents in the current time slot; A decision network updating module is used to update the agent decision network according to the training information combination of the multiple agents in the current time slot, and repeatedly construct the training information combination of the multiple agents in the next time slot according to the multiple agents in the updated agent decision network and the corresponding role prompt information and the state space information of the next time slot, and update the agent decision network, and use the last updated agent decision network as the target agent decision network; The information reasoning module is used to obtain the text information of the problem to be processed, input the text information of the problem to be processed into several agents in the target agent decision network respectively, iterate repeatedly according to a preset number of iterations, obtain the answer text information output by several agents in the target agent decision network at the last iteration number, and use the answer text information with the highest frequency as the reasoning result of the text information of the problem to be processed.

8. A computer device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor; when the computer program is executed by the processor, the steps of the information reasoning method according to any one of claims 1 to 6 are implemented.

9. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the information reasoning method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Multi-agent collaborative confrontation decision-making method and device based on reinforcement learning

    CN117273057A

  • Clinical decision-making artificial intelligence object oriented system and method

    US20190333636A1