Large language model-based information processing method and device
By introducing multimodal information and retrieval answer constraints into the large language model, the problem of hallucination balance in the debate process of multiple agents is solved, the quality and consistency of the reply content is improved, and the debate delay is reduced.
Patent Information
- Application Number
- PCT/CN2025/078663
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-26
- Filing Date
- 2025-02-22
- Publication Date
- 2025-09-04
AI Technical Summary
The existing large language model has hallucinatory balance problems during the debate of multiple agents, resulting in the content of the reply not being consistent with the facts.
By introducing multimodal information and retrieval obtained answers as constraints, multiple agents debate and establish contact with the retrieval answers after reaching an agreement, correct the agent's answers in the form of self-feedback and optimize the model parameters.
It effectively alleviates the problem of hallucination balance in the debate process of multiple agents, improves the quality and consistency of the reply content, reduces debate delays, and improves user experience.
Smart Images

Figure CN2025078663_04092025_PF_FP_ABST
Abstract
Description
Information processing method and device based on large model
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on February 26, 2024, with application number 202410210769.X, and invention name “A method and device for information processing based on large models”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present invention relates to the field of artificial intelligence (AI), and in particular to a large model-based information processing method and device. Background Art
[0003] Large language models (LLMs) have demonstrated powerful end-to-end representation capabilities and excellent performance in numerous fields, garnering widespread attention. AI assistants such as ChatGPT, leveraging their powerful capabilities, have become indispensable tools in our daily lives. However, current LLMs can sometimes produce inaccurate responses, known as LLM hallucinations. This hallucination problem can cause LLMs to produce inconsistent responses. As shown in Figure 1, the LLM's response misinterprets "Lu Xun" and "Zhou Shuren" as two different individuals, contradicting objective reality. Current technologies for addressing LLM hallucinations primarily involve multi-agent debates. These techniques construct multiple agents based on LLMs, mimicking the human debate process. When agents disagree on their responses to user questions, they engage in debate until a consensus is reached. External knowledge base retrieval methods are currently limited to addressing LLMs in single-agent scenarios. Consequently, a growing number of researchers are focusing on multi-agent debates. However, multi-agent debates can lead to incorrect consensus, known as the hallucination balance problem, as shown in Figure 2.
[0004] How to solve the problem of illusory balance in the process of multi-agent debate is the current research focus of those skilled in the art. Summary of the Invention
[0005] The embodiments of the present application disclose a large-model-based information processing method and device, which can alleviate the model's illusory balance and improve the quality of the model's response content.
[0006] A first aspect provides a large model-based information processing method, applied to a first device, the method comprising:
[0007] obtaining a first answer to the first question from one or more second devices;
[0008] Having multiple intelligent agents debate the first question to obtain a second answer;
[0009] A reply content to the first question is generated based on the first answer and the second answer.
[0010] It should be understood that the intelligent agent is an object constructed through a model. Optionally, the intelligent agent is an object constructed based on a large model that is capable of understanding, planning, making decisions, and executing tasks.
[0011] In this method, after multiple agents reach a consensus after sufficient debate, the first answer obtained through retrieval is also introduced as a constraint to determine whether the second answer agreed upon by multiple agents is feasible. If it is feasible, the second answer will be used as the response to the first question. In this way, a connection is established between the first answer obtained through retrieval and the answer generated by the agent, and the agent's answer is corrected in the form of self-feedback, which can avoid the illusion of balance problem caused by the debate among multiple agents and improve the quality of the response content.
[0012] In combination with the first aspect, in a possible implementation of the first aspect, the first question includes multimodal information, and the multimodality includes at least two of text, pictures, videos, and sounds.
[0013] It can be understood that compared with the conventional single-modal pure text method, the introduction of multi-modal information can correct the understanding and cognition of the model (such as the intelligent agent), enhance the multi-perspective understanding ability of the same instruction or question during the debate of multiple intelligent agents, and thus improve the quality of the generated answers.
[0014] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in another possible implementation of the first aspect, the second device includes a database corresponding to multiple modalities, and obtaining a first answer to the first question from one or more second devices includes:
[0015] The first answer to the first question is obtained from a database corresponding to each modality in the multimodality.
[0016] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in another possible implementation of the first aspect, generating a response to the first question based on the first answer and the second answer includes:
[0017] If the second answer and the first answer meet a preset similarity condition, the second answer is determined as the response content of the first question, and the multiple intelligent agents stop debating.
[0018] In this approach, similarity is used to measure the degree of deviation between the second answer and the first answer, thereby determining whether the second answer deviates from common sense. Measuring the degree of deviation by similarity is easy to implement and has good results. In addition, when the answers generated by multiple agents meet the consistency judgment with the first retrieved answer, the debate is stopped early. Compared with the conventional implementation method in which all debates reach the maximum number of rounds, this embodiment of the application can significantly reduce the delay in the debate process and improve the user experience.
[0019] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation of the first aspect, further comprising:
[0020] If the second answer and the first answer do not meet the preset similarity condition, the process returns to executing the step of debating the first question through multiple agents to obtain the second answer.
[0021] It can be understood that according to conventional operations, if multiple intelligent agents fail to reach a consensus, they need to debate again. In the embodiment of the present application, when multiple intelligent agents reach a consensus, if the agreed answer is not consistent with the first answer obtained through retrieval, the multiple intelligent agents will also debate again, and after the debate reaches a consensus, they will again make a consistency judgment with the first answer until the two reach a consensus or the number of debates reaches a preset upper limit. In this way, it is highly likely that an answer consistent with the first answer can be generated in the end.
[0022] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in another possible implementation of the first aspect, the preset similarity condition includes that the similarity exceeds a first similarity threshold.
[0023] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in another possible implementation of the first aspect, the debating of the first question by multiple intelligent agents to obtain a second answer includes:
[0024] Generate answers to the first question by multiple intelligent agents respectively and conduct debate based on the generated answers;
[0025] If the multiple agents fail to reach a consensus during the debate, returning to the step of generating answers to the first question by the multiple agents and conducting a debate based on the generated answers;
[0026] If the multiple intelligent agents reach a consensus after debate, the agreed answer will be used as the second answer.
[0027] It can be understood that the debate in the embodiment of the present application is an iterative process. If the answers generated by multiple agents do not reach a consensus, the debate will continue until a consensus answer is generated or the number of debates reaches a preset upper limit.
[0028] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation of the first aspect, further comprising:
[0029] If the multiple agents fail to reach a consensus in the debate, it is determined based on the answers generated by the multiple agents whether the first answer needs to be used as the input for each agent to regenerate the answer to the first question; if it is determined that it is needed, the input when generating the answer to the first question includes the first answer; if it is determined that it is not needed, the input when generating the answer to the first question does not include the first answer.
[0030] It can be understood that the introduction of a self-feedback mechanism based on retrieval knowledge guidance can allow the intelligent agent to explore its own knowledge boundaries, introduce retrieval knowledge evidence (i.e., the first answer) at the appropriate time, improve the quality of the model's response, and further alleviate the illusion of balance in the debate process of multiple intelligent agents.
[0031] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in another possible implementation of the first aspect, before the debating of the first question by multiple agents to obtain a second answer further includes:
[0032] determining a first loss value based on a similarity between a third answer and a fourth answer, wherein the third answer is an answer to the second question obtained from the one or more second devices, and the fourth answer is an answer to the second question generated by the first agent;
[0033] Determining a loss function for updating the first agent according to the first loss value and the second loss value, wherein the second loss value is the loss of the fourth answer relative to the reference answer;
[0034] The first agent is updated using the loss function, wherein the first agent is any one of the multiple agents.
[0035] In this implementation, the third answer retrieved through retrieval is used as a guiding signal to evaluate the agent's model parameter loss, i.e., the first loss value. This first loss value is then used to optimize the conventional model parameter loss (i.e., the second loss value), resulting in a more accurate loss function for the model. Updating the agent's model parameters using this more accurate loss function can improve the quality of the answers generated by the model.
[0036] In combination with the first aspect, or any one of the above-mentioned possible implementations of the first aspect, in another possible implementation of the first aspect, it further includes: outputting the reply content of the first question.
[0037] A second aspect provides a device, such as a first device, comprising a unit or module for executing any method described in the first aspect.
[0038] The third aspect provides a device, which can be a single device or a distributed system composed of multiple nodes. The device includes one or more processors and one or more memories, wherein: the one or more memories are used to store one or more programs, and the one or more processors call the one or more programs to execute the method described in the first aspect or any possible implementation of the first aspect.
[0039] The fourth aspect provides a storage medium (also referred to as a computer-readable storage medium), which stores one or more programs. When the one or more programs are executed by a processor, they implement the method described in the first aspect or any possible implementation of the first aspect.
[0040] In a fifth aspect, a chip system is provided, comprising: the chip system comprises one or more processors and one or more memories, the one or more processors being used to read and execute one or more programs stored in the one or more memories to implement a method as described in any one of the first aspects.
[0041] According to a sixth aspect, a program product is provided. When the program product is run on a device, the device is caused to execute the method as described in any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The following is an introduction to the drawings used in the embodiments of this application.
[0043] FIG1 is a schematic diagram of a case of a large model illusion in the prior art;
[0044] FIG2 is a schematic diagram of another example of a large model illusion in the prior art;
[0045] FIG3 is a schematic diagram of the architecture of an information processing system based on a large model provided in an embodiment of the present application;
[0046] FIG4 is a logic diagram of a debate among multiple intelligent agents and the final output of a consistent response provided by an embodiment of the present application;
[0047] FIG5 is a flow chart of an information processing method based on a large model provided in an embodiment of the present application;
[0048] FIG6 is a schematic diagram of a multimodal retrieval scenario provided in an embodiment of the present application;
[0049] FIG7 is a schematic diagram of a consistency judgment principle provided by an embodiment of the present application;
[0050] FIG8 is a schematic diagram of a method of splicing multiple agent replies into a spliced answer according to an embodiment of the present application;
[0051] FIG9 is a schematic diagram showing a principle of a debate between two intelligent agents according to an embodiment of the present application;
[0052] FIG10 is a schematic diagram of a method for determining the compensation amount of model loss based on retrieval answers provided by an embodiment of the present application;
[0053] FIG11 is a schematic structural diagram of an information processing device based on a large model provided in an embodiment of the present application;
[0054] FIG12 is a schematic structural diagram of a first device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.
[0056] Artificial intelligence is used in many scenarios. Here are some examples:
[0057] Scenario 1: The user asks a question to the first device, and the first device generates an answer to the question through a large model. For example, the user inputs the question "Where is the best place to go this weekend" to the first device through voice, text, or other means, and the first device generates the answer "Go to scenic spot A. The weather at scenic spot A is sunny and the temperature is moderate this weekend, and the scenery is the best time of the year"; for another example, the user inputs the question "Is the animal in this picture a donkey or a horse" to the first device through voice, text, or other means, and the first device generates the answer, gives a reason, and outputs it; and so on.
[0058] Scenario 2: The user proposes a task to the first device, and the first device completes the task through a large model and outputs the result. For example, the user inputs the task "Generate a 1,000-word article describing the winter snow scene" to the first device through voice, text or other means, and the first device generates an article that meets the requirements and outputs it accordingly; for example, the user inputs two clothing styles to the first device, and the first device determines which of the two styles is more suitable for the user. Accordingly, the first device gives the recommendation result and reason and outputs it; for example, the user inputs the task "Find out the defects (or areas for improvement) in this article (or work or design)" to the first device through voice, text or other means, and the first device outputs the location of the defects, gives an optimization plan, and outputs it, etc.
[0059] It should be noted that the above scenarios are only examples. There are actually many other similar scenarios, which will not be described here one by one.
[0060] Please refer to Figure 3, which is a schematic diagram of an information processing architecture based on a large model provided by an embodiment of the present application. The architecture 30 includes a first device that performs tasks based on a large model. The first device can be a single node (or device) or a distributed system composed of multiple nodes (or devices). The nodes in the embodiment of the present application can be physical nodes or virtual nodes (such as nodes implemented by virtual machines). When there are multiple nodes, one or more agents can be deployed on each node. These multiple nodes can be deployed in the cloud, or all deployed locally, or partially deployed in the cloud and partially deployed locally, that is, end-cloud collaboration. Optionally, if an end-cloud collaboration scenario is adopted, a large-scale language model (including an agent) on the cloud side can debate with a small-scale language model (including an agent) on the end side (i.e., local), effectively improving the performance of the first device in performing tasks (or replies); in addition, the multiple agents in the embodiment of the present application can be two or more, and can be expanded to application scenarios for debates with more agents if computing resources permit.
[0061] Figure 3 illustrates the scenario of end-cloud collaboration as an example. The information processing architecture 30 based on the big model includes a node 301 deployed on the cloud, a node 302 deployed locally (i.e., on the end side), and a node 303. The nodes in the architecture 30 are connected by wired or wireless means, so that different nodes can communicate, wherein the intelligent agents deployed on the nodes can debate. In the information processing method based on the big model provided in the embodiment of the present application, operations other than debate (such as retrieval, outputting replies (or replies or tasks), consistency judgment, etc.) can be completed by node 301 or node 302, and can also be completed by a special node 303 where an intelligent agent is not deployed. It can be completed by one node or multiple nodes.
[0062] Optionally, the user can interact with the large-model-based information processing architecture 30 through the local node 302 or the local node 303, for example, by asking questions to the local node 302 or the local node 303. In this case, the local node 302 or the local node 303 can be a device to which the large-model is applied, such as a handheld device (e.g., a mobile phone, a tablet computer, a PDA, a laptop computer, etc.), an in-vehicle device (e.g., a car, a bicycle, an electric car, an airplane, a ship, etc.), a wearable device (e.g., a smart watch (e.g., iWatch, etc.), a smart bracelet, a pedometer, etc.), a smart home device (e.g., a refrigerator, a television, an air conditioner, an electric meter, etc.), an intelligent robot, workshop equipment, a computer, etc., and the specific details are not limited here.
[0063] Optionally, as shown in FIG4 , a logical diagram of a debate between agents and the final output of a consistent response is illustrated. The agent in FIG4 can be the agent in the architecture shown in FIG3 (called Agent in English), or it can be an agent in another architecture. The execution logic shown in FIG4 is as follows: Taking the debate between two agents (agent 1 and agent 2) as an example, the agent debate system can include a multimodal knowledge retrieval module, an agent response module, a knowledge self-feedback module, a consistency judgment module, and a retrieval timing module. For the query input by the user, a multimodal knowledge retrieval is first performed to obtain the retrieved knowledge fragment (i.e., the first answer mentioned later, which can be used as retrieval evidence) which will be used for subsequent knowledge self-feedback and agent response. Secondly, the two agent response modules generate responses (or answers) corresponding to the questions input by the user, and then input the responses into the consistency judgment module to determine whether the answers generated by the two agents are consistent. (1) If the responses of the agents are inconsistent, the retrieval timing module is entered. The agents judge whether the retrieval evidence (i.e., the first answer) is required as input for the next round of responses (or answers) based on their own responses, and then enter the next round of debate. In this process, if the retrieval evidence is required as input, the input includes the retrieved knowledge fragments and the answers generated by multiple agents in the past. If the retrieval evidence is not required as input, the input includes the answers generated by multiple agents in the past. (2) If the responses of the agents are consistent or the maximum number of debate rounds is reached, the debate ends. In addition, after each round of responses generated by the agents, the knowledge self-feedback mechanism can be used to correct the model responses to achieve the purpose of alleviating the illusion of balance in the debate between agents. The specific implementation details will be explained in detail in conjunction with the method flow shown in Figure 5.
[0064] Please refer to FIG5 , which shows an information processing method based on a large model provided by an embodiment of the present application. The method can be implemented based on the architecture shown in FIG3 or based on other architectures. The method includes but is not limited to the following steps:
[0065] Step S501: Obtain a first answer to a first question from one or more second devices.
[0066] Specifically, the second device can be a single device or a distributed system consisting of multiple nodes. The second device can be the first device described above. When the first device is a distributed system consisting of multiple nodes, the second device can be a portion of the nodes in the distributed system. Of course, the second device can also be a device other than the first device. When obtaining the first answer to the first question from the second device, it can be obtained by retrieval, reading, or other methods, which are not limited here.
[0067] The first question can be a question, such as "Where is the best place to go this weekend?" or "Is the animal in this picture a donkey or a horse?"; the first question can also be a task, such as "Generate a 1,000-word article describing the winter snow scene."
[0068] The first question may be single-modal. For example, if the first question is presented or input entirely in the form of text, then the first question is single-modal. For another example, if the first question is presented or input entirely in the form of audio, then the first question is single-modal.
[0069] In addition, the first question can also be multimodal. For example, if the first question consists of text and images, then the first question is multimodal. For example, if the first question consists of audio and images, then the first question is multimodal.
[0070] Optionally, the second device includes a database corresponding to multiple modalities. The database can be an internal database or an external database (such as a database provided by the Internet or other platforms). The subsequent examples are explained using the database as an external database. Optionally, there may be multiple databases. For example, as shown in Figure 6, if the multimodalities mentioned above include text, images, audio, etc., there may be a database (or knowledge base) corresponding to text, a database (or knowledge base) corresponding to images, a database (or knowledge base) corresponding to audio, and so on. When the first question is multimodal, answers can be retrieved from the databases corresponding to the multiple modalities. A multimodal retrieval module can be configured to implement multimodal retrieval (Retrieve Module). The answers retrieved separately are fused to obtain the first answer. Optionally, the information of different modalities in the first question can be encoded (or quantized) by the corresponding encoder (such as text encoder, audio encoder, image encoder, etc.) in the multimodal retrieval module, and then retrieved from the corresponding database in the multimodal retrieval module. For example, a similarity matching algorithm is used to retrieve knowledge fragments from the corresponding database. The knowledge fragments that meet the preset conditions can be used as the retrieved answers. Optionally, the first answer obtained by fusion may be a multimodal answer. As shown in FIG6 , the output first answer 601 includes images, text, and sound, and is therefore a multimodal answer.
[0071] When the first question is unimodal, the answer obtained (eg, retrieved) from a single database (or after processing) can be used as the first answer.
[0072] Step S502: Multiple intelligent agents debate the first question to obtain a second answer.
[0073] The intelligent agent is an object constructed based on a model, for example, the intelligent agent is an object constructed based on a large model that is capable of understanding, planning, making decisions, and executing tasks.
[0074] The principle of debate in the embodiment of this application is as follows:
[0075] Answers to the first question are generated by multiple agents respectively, and a debate is conducted based on the generated answers. For example, the answers (or replies) generated by different agents can be encoded using an encoder (such as TinyBert), and the similarity is calculated between them. If the similarity is greater than a preset second similarity threshold, it indicates that the answers (or replies) generated by the two agents are consistent, that is, they reach a consensus. Optionally, the second similarity threshold can be set according to actual needs, as shown in Figure 7, which takes the second similarity threshold of 0.5 as an example.
[0076] If any two of the preceding multiple intelligent agents reach a consensus, the consensus answer will be used as the second answer.
[0077] If any two of the multiple intelligent agents fail to reach a consensus, the multiple intelligent agents will regenerate answers to the first question and, based on the generated answers, perform consistency judgments between each other according to the previous principles. If no consensus is reached, the answer to the first question will be generated again. This cycle will continue until the multiple intelligent agents reach a consensus and use the agreed answer as the second answer. This process is the process of reaching consensus through intelligent agent debate.
[0078] In an optional solution, if the multiple agents fail to reach a consensus in the debate, the answers generated by the multiple agents determine whether the first answer needs to be used as input for each agent to regenerate the answer to the first question. Optionally, a model strategy is pre-trained that can concatenate (or fuse) the answers generated by the multiple agents to obtain a concatenated answer, which is then input to each agent. As shown in FIG8 , each agent thereby determines that the previous first answer needs to be used as one of the inputs when generating the answer (or reply) to the first question in the next round. From the perspective of a single agent, the agent's retrieval timing module concatenates the agent's answer (or reply) with the answers (or replies) of other agents into a concatenated answer, which is then used as the agent's input, allowing the agent to determine that the previous first answer needs to be used as one of the inputs when generating the answer (or reply) to the first question in the next round. If it is determined that it is required, the input for generating the answer to the first question includes the first answer. If it is determined that it is not required, the input for generating the answer to the first question does not include the first answer. This implementation scheme gives the intelligent agent the ability to explore the boundaries of knowledge. Specifically, it is based on a knowledge-guided feedback mechanism to guide the intelligent agent to explore the boundaries of knowledge, "know whether it knows" or not, and add the retrieved knowledge fragment (that is, the first answer obtained through the retrieval) at the right time to guide the intelligent agent to generate better quality responses.
[0079] 9 , a debate process is illustrated below using multiple agents (eg, agent 1 (Agent1) and agent 2 (Agent2)) as an example.
[0080] Among them, Agent1 is deployed on the cloud, and Agent2 is deployed on the client side (i.e., locally).
[0081] 1. A user inputs a first question to the large-scale model-based information processing system (i.e., the first device, which may be the architecture shown in FIG3 ). In this case, the first question includes “What is the animal in the picture?” and a picture.
[0082] 2. The multimodal retrieval module in the information processing system of the large model searches the corresponding database and obtains the first answer 1001. It can be seen that the first answer 1001 is a multimodal reply content, which includes text, images and sounds. The first answer can be used as multimodal retrieval evidence (Multi Modal Evidences) later.
[0083] 3. Enter the first round of debate process (round 1)
[0084] Agent 1 generates the answer to the first question, which is “the animal in the picture is a deer.”
[0085] Agent 2 generates the answer to the first question, which is "The animal in the picture is a horse."
[0086] 4. The consistency judgment module determines whether the answers of Agent 1 and Agent 2 are consistent. If no agreement is reached in round 1, the search timing modules of Agent 1 and Agent 2 are triggered to splice the answers of Agent 1 and Agent 2 to obtain a spliced answer (also called prompt in English).
[0087] 5. The retrieval timing modules of Agent 1 and Agent 2 each use the concatenated answers to determine whether to use the first answer (i.e., the answer previously retrieved) as one of the inputs when regenerating the answer to the first question in the next round. It can be seen that in the first round of the debate process (round 1), both Agent 1 and Agent 2 need to use the first answer as one of the inputs.
[0088] 6. Enter the second round of debate process (round 2)
[0089] Agent 1 uses the answers generated in Agent 1's history, the answers generated in Agent 2's history, and the first answer as input to regenerate the answer to the first question, which is "Horses generally do not have horns and are larger in size. According to the retrieved image, the animal in the image is more similar to a deer."
[0090] Agent 2 takes the answers generated in Agent 1's history, the answers generated in Agent 2's history, and the first answer as input, and regenerates the answer to the first question, which is "Horses and deer are herbivores and both have four legs. I feel that the animal in the picture looks more like a horse."
[0091] 7. The consistency judgment module determines whether the answers of Agent 1 and Agent 2 are consistent. If no agreement is reached in round 2, the search timing modules of Agent 1 and Agent 2 are triggered to splice the answers of Agent 1 and Agent 2 to obtain a spliced answer.
[0092] 8. The retrieval timing modules of Agent 1 and Agent 2 each use the concatenated answers to determine whether to use the first answer (i.e., the answer previously retrieved) as one of the inputs when regenerating the answer to the first question in the next round. It can be seen that in the second round of debate (round 2), Agent 1 does not need to use the first answer as one of the inputs, while Agent 2 does.
[0093] 9. Enter the third round of debate process (round 3)
[0094] Agent 1 uses the answers generated in Agent 1's history and Agent 2's history as input to regenerate the answer to the first question, which is "Although a horse also has four legs, it is tall, while the animal in the picture is smaller and has horns."
[0095] Agent 2 takes the answers generated in Agent 1's history, the answers generated in Agent 2's history, and the first answer as input, and regenerates the answer to the first question, which is "You are right, horses generally do not have horns but deer have horns, so the animal in the picture is a deer."
[0096] 10. The consistency judgment module determines whether Agent 1 and Agent 2's answers are consistent. If they reach a consensus in round 3, the debate between Agent 1 and Agent 2 is temporarily "over". If the answer agreed upon by Agent 1 and Agent 2 is consistent with the first answer mentioned above, the debate can be officially ended. Otherwise, the debate process will be restarted until the consistent answer formed by the debate is consistent with the first answer.
[0097] Step S503: Generate a reply to the first question based on the first answer and the second answer.
[0098] Specifically, the final response to the first question needs to comprehensively consider the first answer and the second answer. There are many specific implementation methods, and the following only uses some cases as examples to illustrate.
[0099] For example, the first answer and the second answer are combined to obtain the response content of the first question.
[0100] For another example, use the first answer as a reference for the second answer. If the second answer does not deviate too much from the first answer, then use the second answer as the response to the first question. For example, to determine whether the second answer and the first answer meet the preset similarity condition, the second answer and the first answer can be encoded (or quantized) by an encoder (such as TinyBert) and then compared for similarity. If the second answer and the first answer meet the preset similarity condition, the second answer is determined as the response to the first question, and the multiple agents stop debating. Optionally, the preset similarity condition includes a similarity exceeding a first similarity threshold. The first similarity threshold can be set according to actual needs, for example, set to 0.5. As shown in Figure 7, this process is the process of confirming whether the consensus answer (i.e., the second answer) of multiple agents is consistent with the retrieved "fact" (i.e., the first answer). Therefore, it can be called fact consistency judgment. If a consensus is reached in the fact consistency judgment process, the second answer can be used as the response to the first question and the above debate is stopped. If a consensus is not reached in the fact consistency judgment process, it is necessary to re-execute the debate process in step S502 to regenerate a consensus answer (i.e., the second answer), and then re-execute step S503 based on the regenerated consensus answer (i.e., the second answer) until a consensus is reached in the fact consistency judgment.
[0101] In the embodiment of the present application, before the above steps S501-S503, the large model may be trained first, including optimizing and updating the model parameters of the above multiple agents. An optional model parameter updating method is provided below:
[0102] A first loss value is determined based on the similarity between a third answer and a fourth answer, where the third answer is an answer to the second question obtained from the one or more second devices, such as an answer retrieved according to the principle of step S501, and the fourth answer is an answer to the second question generated by the first agent. By comparing the similarity between the third and fourth answers, the degree to which the third answer generated by the first agent deviates from the knowledge in the database can be determined. Optionally, the similarity between the third and fourth answers is negatively correlated with the first loss value, i.e., the smaller the similarity, the larger the first loss value. Figure 10 illustrates the process of generating the first loss value βsim(e, r). The module that performs this operation can be referred to as a self-supervised knowledge feedback module.
[0103] The loss function for updating the first agent is determined based on the first loss value and the second loss value, wherein the second loss value is the loss of the fourth answer relative to the reference answer, where the reference answer can be a pre-configured standard answer. Optionally, the first loss value, the second loss value, and the loss function Loss can directly satisfy the following relationship: Loss = CrossEntropyLoss - βsim(e, r)
[0104] Among them, CrossEntropyLoss is the second loss value, βsim(e, r) is the first loss value, sim(e, r) is the similarity between the third answer and the fourth answer, and β is the weight factor, whose size can be set according to actual needs.
[0105] The first agent is updated by the loss function, wherein the first agent is any one of the multiple agents, that is, each of the multiple agents can be updated by the first agent.
[0106] It can be seen that this model parameter update method mainly uses the third answer obtained through retrieval from the database as a guiding signal, restricting the answers generated by the agent from deviating from the retrieved knowledge, and to a certain extent alleviating the illusion balance problem in the debate of multiple agents.
[0107] It should be noted that, in addition to updating the model parameters of the intelligent agent according to the above-mentioned model parameter updating method when training the large model before the above-mentioned steps S501-S503, the model parameters of the intelligent agent can also be updated according to the principle of the above-mentioned model parameter updating method each time the intelligent agent generates an answer in steps S501-S503. The specific process will not be repeated here.
[0108] It should be noted that before the above steps S501-S503, when the large model is trained, the subject performing the model training can be the same as the subject performing the steps S501-S503, for example, it can be the architecture shown in Figure 3 or other architectures, which can be a single device or a distributed system composed of multiple nodes. Of course, the subject performing the model training can also be different from the subject performing the steps S501-S503, that is, the subject performing the model training trains the large model (including multiple agents) and sends it to the subject performing the steps S501-S503 for use.
[0109] Step S504: Output the reply content of the first question.
[0110] Specifically, the output method is not limited here and can be displayed through text, images, videos, etc., or output can be output through voice broadcast. It can be a single mode output or a mixed multi-modal output, such as the reply content including text and images.
[0111] In addition, the reply content to the first question may not be output, but some operations or tasks may be performed directly according to the reply content, or the reply content to the first question may be sent to other devices for use by other devices.
[0112] In the method shown in Figure 5, after multiple agents reach a consensus after sufficient debate, the first answer retrieved is also introduced as a constraint to determine whether the second answer agreed upon by multiple agents is feasible. If it is feasible, the second answer is used as the response to the first question. In this way, a connection is established between the first answer retrieved and the answer generated by the agent, and the answer of the agent is corrected in the form of self-feedback, which can avoid the illusion of balance problem caused by the debate among multiple agents and improve the quality of the response content.
[0113] The above describes in detail the method of the embodiment of the present application, and the following provides an apparatus of the embodiment of the present application.
[0114] Please refer to Figure 11, which is a structural diagram of a large model-based information processing device provided in an embodiment of the present application. The device can be the above-mentioned first device or a device or chip system in the above-mentioned first device. The device 110 may include a multimodal retrieval module 1101 and a consistency judgment module 1102, wherein each unit is described in detail as follows.
[0115] a multimodal retrieval module 1101 , configured to obtain a first answer to a first question from one or more second devices;
[0116] The consistency judgment module 1102 is configured to use multiple agents to debate the first question and obtain a second answer. It should be understood that the agent is an object constructed by the model. Optionally, the agent is an object constructed based on the large model that can understand, plan, make decisions, and execute tasks.
[0117] The consistency judgment module 1102 is further configured to generate a reply to the first question based on the first answer and the second answer.
[0118] In the method shown in Figure 5, after multiple agents reach a consensus after sufficient debate, the first answer obtained through retrieval is also introduced as a constraint to determine whether the second answer agreed upon by multiple agents is feasible. If feasible, the second answer is used as the response to the first question. In this way, a connection is established between the retrieved first answer and the answer generated by the agent, and the agent's answer is corrected in the form of self-feedback, which can avoid the illusion of balance problem caused by the debate among multiple agents and improve the quality of the response content.
[0119] In a possible implementation, the first question includes multimodal information, and the multimodality includes at least two of text, picture, video, and sound.
[0120] It can be understood that compared with the conventional single-modal pure text method, the introduction of multi-modal information can correct the understanding and cognition of the model (such as the intelligent agent), enhance the multi-perspective understanding ability of the same instruction or question during the debate of multiple intelligent agents, and thus improve the quality of the generated answers.
[0121] In another possible implementation, the second device includes databases corresponding to multiple modalities, and the first answer to the first question is obtained from one or more second devices. The multimodal retrieval module 1101 is specifically used to obtain the first answer to the first question from the database corresponding to each modality in the multimodality.
[0122] In another possible implementation, a reply to the first question is generated based on the first answer and the second answer, and the consistency determination module 1102 is specifically configured to:
[0123] If the second answer and the first answer meet a preset similarity condition, the second answer is determined as the response content of the first question, and the multiple intelligent agents stop debating.
[0124] In this approach, similarity is used to measure the degree of deviation between the second answer and the first answer, thereby determining whether the second answer deviates from common sense. Measuring the degree of deviation by similarity is easy to implement and has good results. In addition, when the answers generated by multiple agents meet the consistency judgment with the first retrieved answer, the debate is stopped early. Compared with the conventional implementation method in which all debates reach the maximum number of rounds, this embodiment of the application can significantly reduce the delay in the debate process and improve the user experience.
[0125] In another possible implementation, the consistency determination module 1102 is further configured to:
[0126] If the second answer and the first answer do not meet the preset similarity condition, the process returns to executing the step of debating the first question through multiple agents to obtain the second answer.
[0127] It can be understood that according to conventional operations, if there is no consensus between intelligent agents, a new debate is required. In the embodiment of the present application, when multiple intelligent agents reach a consensus, if the agreed answer is not consistent with the first answer obtained through retrieval, the multiple intelligent agents will also debate again, and after the debate reaches a consensus, the consistency judgment with the first answer will be made again until the two reach a consensus or the number of debates reaches a preset upper limit. In this way, it is highly likely that an answer consistent with the first answer can be generated in the end.
[0128] In yet another possible implementation, the preset similarity condition includes a similarity exceeding a first similarity threshold.
[0129] In another possible implementation, the first question is debated by multiple agents to obtain a second answer, and the consistency determination module 1102 is specifically configured to:
[0130] Generate answers to the first question by multiple intelligent agents respectively and conduct debate based on the generated answers;
[0131] If the multiple agents fail to reach a consensus during the debate, returning to the step of generating answers to the first question by the multiple agents and conducting a debate based on the generated answers;
[0132] If the multiple intelligent agents reach a consensus after debate, the agreed answer will be used as the second answer.
[0133] It can be understood that the debate in the embodiment of the present application is an iterative process. If the answers generated by multiple agents do not reach a consensus, the debate will continue until a consensus answer is generated or the number of debates reaches a preset upper limit.
[0134] In yet another possible implementation, the method further includes:
[0135] A retrieval timing module is used to determine whether the first answer needs to be used as input for each agent to regenerate the answer to the first question if the multiple agents fail to reach a consensus in the debate based on the answers generated by the multiple agents; if it is determined that it is needed, the input when generating the answer to the first question includes the first answer; if it is determined that it is not needed, the input when generating the answer to the first question does not include the first answer.
[0136] It can be understood that the introduction of a self-feedback mechanism based on retrieval knowledge guidance can allow the intelligent agent to explore its own knowledge boundaries, introduce retrieval knowledge evidence (i.e., the first answer) at the appropriate time, improve the quality of the model's response, and further alleviate the illusion of balance in the debate process of multiple intelligent agents.
[0137] In another possible implementation, a self-supervisory knowledge feedback module is further included, configured to:
[0138] determining a first loss value based on a similarity between a third answer and a fourth answer, wherein the third answer is an answer to the second question obtained from the one or more second devices, and the fourth answer is an answer to the second question generated by the first agent;
[0139] Determining a loss function for updating the first agent according to the first loss value and the second loss value, wherein the second loss value is the loss of the fourth answer relative to the reference answer;
[0140] The first agent is updated using the loss function, wherein the first agent is any one of the multiple agents.
[0141] In this implementation, the retrieved third answer is used as a guiding signal to evaluate the agent's model parameter loss, i.e., the first loss value. This first loss value is then used to optimize the conventional model parameter loss (i.e., the second loss value), resulting in a more accurate loss function for the model. Updating the agent's model parameters using this more accurate loss function can improve the quality of the answers generated by the model.
[0142] In another possible implementation, the method further includes: an output module, configured to output the reply content to the first question.
[0143] It should be noted that the implementation and beneficial effects of each unit may also correspond to the corresponding description of the method embodiment shown in FIG5 .
[0144] It is understood that the modules described above are functional modules divided according to their functions. In specific implementations, some functional blocks may be subdivided into more small functional modules, and some functional modules may be combined into a single functional module. However, regardless of whether these functional modules are subdivided or combined, the information processing flow based on the above large model is generally the same. Typically, each functional module corresponds to its own program code (or program instructions). When the program code corresponding to each functional module is executed on the processor, it causes the functional module to execute the corresponding process and thus realize the corresponding function.
[0145] Please refer to Figure 12, which is a schematic diagram of the structure of a device provided in an embodiment of the present application. For ease of distinction, this device can be referred to as a first device. The first device 120 includes one or more processors 1201 and one or more memories 1202. Optionally, it may also include a communication interface 1203 and a user interface 1204. The first device 120 can be a single device or a distributed system consisting of multiple nodes. Each module is described below.
[0146] Processor 1201 can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by one or more processors.
[0147] The memory 1202 may be one or more and is used to store program instructions and / or data. The memory 1202 is coupled to the processor 1201. The coupling in the embodiment of the present application is an indirect coupling or communication connection between devices, units or modules, which may be electrical, mechanical or other forms, and is used for information exchange between devices, units or modules. The processor 1201 may operate in conjunction with the memory 1202. The processor 1201 can execute program instructions stored in the memory 1202. Optionally, at least one of the one or more memories may be included in the processor. The memory may include but is not limited to non-volatile memories such as hard disks or solid-state drives, random access memories, erasable programmable read-only memories, read-only memories or portable read-only memories, etc. The memory is any storage medium that can be used to carry or store program code in the form of instructions or data structures and can be read and / or written by a device (such as the first device, etc.), but is not limited to this. The memory in the embodiment of the present application may also be a circuit or any other device that can implement a storage function, used to store program instructions and / or data.
[0148] The communication interface 1203 can communicate with other devices to receive and send data.
[0149] The user interface 1204 is mainly used for human-computer interaction, such as inputting and / or outputting information. Specifically, the user interface may include input and output devices, such as a touch screen, a display screen, a keyboard, etc., which are mainly used to receive data input by the user and output data to the user.
[0150] The specific connection medium between the user interface 1204, communication interface 1203, processor 1201, and memory 1202 is not limited in the embodiments of the present application. In Figure 12, the embodiment of the present application uses the bus connection between the user interface 1204, communication interface 1203, processor 1201, and memory 1202 as an example. The bus is represented by a bold line in Figure 12. The connection method between other components is only for schematic illustration and is not intended to be limiting. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one bold line is used in Figure 12, but this does not mean that there is only one bus or one type of bus.
[0151] The processor 1201 in the first device 120 is configured to read one or more programs stored in the memory 1202 and perform the following operations:
[0152] obtaining a first answer to the first question from one or more second devices;
[0153] Debating the first question by multiple agents to obtain a second answer, the agents being objects constructed by the model. Optionally, the agents are objects constructed based on the large model that can understand, plan decisions, and execute tasks;
[0154] A reply content to the first question is generated based on the first answer and the second answer.
[0155] In the method shown in Figure 5, after multiple agents reach a consensus after sufficient debate, the first answer retrieved is also introduced as a constraint to determine whether the second answer agreed upon by multiple agents is feasible. If it is feasible, the second answer is used as the response to the first question. In this way, a connection is established between the first answer retrieved and the answer generated by the agent, and the answer of the agent is corrected in the form of self-feedback, which can avoid the illusion of balance problem caused by the debate among multiple agents and improve the quality of the response content.
[0156] In a possible implementation, the first question includes multimodal information, and the multimodality includes at least two of text, picture, video, and sound.
[0157] It can be understood that compared with the conventional single-modal pure text method, the introduction of multi-modal information can correct the understanding and cognition of the model (such as the intelligent agent), enhance the multi-perspective understanding ability of the same instruction or question during the debate of multiple intelligent agents, and thus improve the quality of the generated answers.
[0158] In another possible implementation, the second device includes a database corresponding to multiple modalities, and the processor 1201 is specifically configured to obtain the first answer to the first question from one or more second devices:
[0159] The first answer to the first question is obtained from a database corresponding to each modality in the multimodality.
[0160] In yet another possible implementation, the generating of the reply content to the first question according to the first answer and the second answer, the processor 1201 is specifically configured to:
[0161] If the second answer and the first answer meet a preset similarity condition, the second answer is determined as the response content of the first question, and the multiple intelligent agents stop debating.
[0162] In this approach, similarity is used to measure the degree of deviation between the second answer and the first answer, thereby determining whether the second answer deviates from common sense. Measuring the degree of deviation by similarity is easy to implement and has good results. In addition, when the answers generated by multiple agents meet the consistency judgment with the first retrieved answer, the debate is stopped early. Compared with the conventional implementation method in which all debates reach the maximum number of rounds, this embodiment of the application can significantly reduce the delay in the debate process and improve the user experience.
[0163] In yet another possible implementation, the processor 1201 is further configured to:
[0164] If the second answer and the first answer do not meet the preset similarity condition, the process returns to executing the step of debating the first question through multiple agents to obtain the second answer.
[0165] It can be understood that according to conventional operations, if there is no consensus between intelligent agents, a new debate is required. In the embodiment of the present application, when multiple intelligent agents reach a consensus, if the agreed answer is not consistent with the first answer obtained through retrieval, the multiple intelligent agents will also debate again, and after the debate reaches a consensus, the consistency judgment with the first answer will be made again until the two reach a consensus or the number of debates reaches a preset upper limit. In this way, it is highly likely that an answer consistent with the first answer can be generated in the end.
[0166] In yet another possible implementation, the preset similarity condition includes a similarity exceeding a first similarity threshold.
[0167] In another possible implementation, the first question is debated by multiple agents to obtain a second answer, and the processor 1201 is specifically configured to:
[0168] Generate answers to the first question by multiple intelligent agents respectively and conduct debate based on the generated answers;
[0169] If the multiple agents fail to reach a consensus during the debate, returning to the step of generating answers to the first question by the multiple agents and conducting a debate based on the generated answers;
[0170] If the multiple intelligent agents reach a consensus after debate, the agreed answer will be used as the second answer.
[0171] It can be understood that the debate in the embodiment of the present application is an iterative process. If the answers generated by multiple agents do not reach a consensus, the debate will continue until a consensus answer is generated or the number of debates reaches a preset upper limit.
[0172] In yet another possible implementation, the processor 1201 is further configured to:
[0173] If the multiple agents fail to reach a consensus in the debate, it is determined based on the answers generated by the multiple agents whether the first answer needs to be used as the input for each agent to regenerate the answer to the first question; if it is determined that it is needed, the input when generating the answer to the first question includes the first answer; if it is determined that it is not needed, the input when generating the answer to the first question does not include the first answer.
[0174] It can be understood that the introduction of a self-feedback mechanism based on retrieval knowledge guidance can allow the intelligent agent to explore its own knowledge boundaries, introduce retrieval knowledge evidence (i.e., the first answer) at the appropriate time, improve the quality of the model's response, and further alleviate the illusion of balance in the debate process of multiple intelligent agents.
[0175] In yet another possible implementation, before the multiple agents debate the first question to obtain a second answer, the processor is further configured to:
[0176] determining a first loss value based on a similarity between a third answer and a fourth answer, wherein the third answer is an answer to the second question obtained from one or more second devices, and the fourth answer is an answer to the second question generated by the first agent;
[0177] Determining a loss function for updating the first agent according to the first loss value and the second loss value, wherein the second loss value is the loss of the fourth answer relative to the reference answer;
[0178] The first agent is updated using the loss function, wherein the first agent is any one of the multiple agents.
[0179] In this implementation, the retrieved third answer is used as a guiding signal to evaluate the agent's model parameter loss, i.e., the first loss value. This first loss value is then used to optimize the conventional model parameter loss (i.e., the second loss value), resulting in a more accurate loss function for the model. Updating the agent's model parameters using this more accurate loss function can improve the quality of the answers generated by the model.
[0180] In yet another possible implementation, the processor 1201 is further configured to: output a reply to the first question through the user interface 1204 .
[0181] It should be noted that the implementation and beneficial effects of each operation may also correspond to the corresponding description of the method embodiment shown in FIG5 .
[0182] The first device shown in the embodiment of the present application can implement the method provided in the embodiment of the present application in the form of hardware, or can implement the method provided in the embodiment of the present application in the form of software, etc., and the embodiment of the present application is not limited to this.
[0183] In addition, the present application also provides a program, which is used to be executed by a device (for example, a first device) to implement the method provided by the present application.
[0184] The present application also provides a storage medium, such as a computer-readable storage medium, which stores one or more programs (such as computer programs). When the one or more programs are called by a processor, the method provided in the application is executed by a device (for example, a first device).
[0185] The present application also provides a storage medium, such as a chip system including one or more processors and one or more memories, wherein the one or more processors are used to read and execute one or more programs stored in the one or more memories, so that the method provided in the application is executed by a device (for example, a first device).
[0186] The present application also provides a program product, which includes a code or a program. When the code or program is called by a processor, the method provided in the present application is executed by a device (for example, a first device).
[0187] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A large model-based information processing method, characterized in that: Applied to a first device, the method includes: obtaining a first answer to the first question from one or more second devices; Having multiple intelligent agents debate the first question to obtain a second answer; A reply content to the first question is generated based on the first answer and the second answer.
2. The method according to claim 1, characterized in that The first question includes multimodal information, and the multimodality includes at least two of text, picture, video, and sound.
3. The method according to claim 2, characterized in that The second device includes a database corresponding to multiple modalities, and obtaining a first answer to the first question from one or more second devices includes: The first answer to the first question is obtained from a database corresponding to each modality in the multimodality.
4. The method according to any one of claims 1 to 3, characterized in that Generating a reply to the first question based on the first answer and the second answer includes: If the second answer and the first answer meet a preset similarity condition, the second answer is determined as the response content of the first question, and the multiple intelligent agents stop debating.
5. The method according to claim 4, characterized in that Also includes: If the second answer and the first answer do not meet the preset similarity condition, the process returns to executing the step of debating the first question through multiple agents to obtain the second answer.
6. The method according to any one of claims 1 to 5, characterized in that The step of debating the first question by multiple intelligent agents to obtain a second answer includes: Generate answers to the first question respectively by the multiple intelligent agents and conduct debate based on the generated answers; If the multiple agents fail to reach a consensus during the debate, returning to the step of generating answers to the first question by the multiple agents and conducting a debate based on the generated answers; If the multiple intelligent agents reach a consensus after debate, the agreed answer will be used as the second answer.
7. The method according to claim 6, characterized in that Also includes: If the multiple agents fail to reach a consensus during the debate, determining, based on the answers generated by the multiple agents, whether the first answer needs to be used as input for each agent to regenerate an answer to the first question; If it is determined that it is required, the input when generating the answer to the first question includes the first answer; if it is determined that it is not required, the input when generating the answer to the first question does not include the first answer.
8. The method according to any one of claims 1 to 7, characterized in that Before the multiple intelligent agents debate the first question to obtain a second answer, the method further includes: determining a first loss value based on a similarity between a third answer and a fourth answer, wherein the third answer is an answer to the second question obtained from the one or more second devices, and the fourth answer is an answer to the second question generated by the first agent; Determining a loss function for updating the first agent according to the first loss value and the second loss value, wherein the second loss value is the loss of the fourth answer relative to the reference answer; The first agent is updated using the loss function, wherein the first agent is any one of the multiple agents.
9. The method according to any one of claims 1 to 8, characterized in that Also includes: Output the answer to the first question.
10. A device, characterized in that comprising one or more processors and one or more memories, wherein: The one or more memories are used to store one or more programs, and the one or more processors call the one or more programs to execute the method according to any one of claims 1 to 9.
11. A storage medium, characterized in that: The storage medium stores a program, and when the program is executed by the processor, the method according to any one of claims 1 to 9 is implemented.
12. A chip system comprising: The chip system includes one or more processors and one or more memories, and the one or more processors are used to read and execute one or more programs stored in the one or more memories to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Question and answer processing method and device, computer equipment and storage medium
CN115757725A
Construction method of multi-modal user mental perception question and answer model
CN117033602A
Information and data collaboration among multiple artificial intelligence (AI) systems
US20200372382A1