Intelligent Agent thinking method fusing longer memory and keeping content consistency

By adjusting user questions and using Mamba network model to select answers, the problem of inconsistent information in large models during long conversations is solved, and the consistency and user experience of answers are improved.

CN120067245APending Publication Date: 2025-05-30INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411959655.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Large models have weak consistency in long-term conversations, especially when multiple rounds of conversations, they are prone to problems with inconsistent answers before and after.

Method used

By adjusting user questions, obtaining the original and rewritten answers, and using the Mamba network model to select from them, we can finally obtain questions and answers that are more in line with user intentions.

Benefits of technology

Improves the information consistency of the big model in long conversations, slows down the shortcomings of the Transformer structure's inconsistent answers during long conversations, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067245A_ABST
    Figure CN120067245A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent Agent thinking method capable of fusing longer memory and keeping content consistency, which comprises the following steps of: asking a user question in a way instead of processing prompts and examples; respectively submitting the two questions to a large model of the intelligent Agent for answering, and carrying out binary quantization on the answers; an embedding algorithm is used for converting the question, the original answer and the adjusted question and answer into vectors, a difference vector of the original answer vector and the adjusted question and answer is calculated, the question vector and the difference vector are connected into a long vector, the long vector serves as input, an Agent answer binarization result serves as output, a set of input and output pairs are constructed, and a data set is arranged to train a Mama model; when a user puts forward a new question, two questions of long vector input and Agent are constructed according to the method, 0 and 1 prediction is performed by using a trained Mamba model, and an answer of a proper question is selected according to a prediction result and returned to the user. Compared with a large-model original answer, the method can obtain a question answer which is fused with longer memory and keeps content consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence large models, and specifically relates to an intelligent Agent thinking method for maintaining content consistency by integrating longer memories.

[0002] Adjust the user's question, obtain the original question answer and the adjusted question answer, and select from them using the Mamba network model, so as to obtain a question and answer that better conforms to the user's intention. Since the Mamba network has a better inference effect on longer memories, more consistent answers can be obtained before and after, alleviating the shortcoming that it is difficult to ensure the consistency of answers before and after during long conversations using the Transformer structure adopted by large models. Background Art

[0003] In the development of modern artificial intelligence, intelligent agents based on large models have gradually become an important technology for solving complex problems and enhancing user experience. Large models such as GPT and BERT, with their powerful natural language processing capabilities, are widely used in fields such as text generation, machine translation, and semantic analysis. These models can understand and generate natural language by processing massive amounts of data. However, although large models perform well in single tasks, their consistency in long conversations is weak. Especially in the face of multi-turn conversations, it is easy to have problems with inconsistent answers before and after. The literature [Siren’s Song in the AIOcean: A Survey on Hallucination in Large Language Models] points out that large models often have hallucinations when answering questions, resulting in errors. Currently, traditional large models rely on the Transformer architecture. Although this architecture performs excellently in processing short-term information, in long conversation scenarios, the model often forgets previous information or provides inconsistent answers. This phenomenon is mainly because the context window of large models is limited and they cannot remember the user's historical input and background information for a long time. When users have long conversations with intelligent agents, if the model fails to maintain consistency in information before and after, it may lead to a decline in user experience and reduce the trust in the answers generated by the model. To address this issue, researchers have proposed a solution to integrate long-term memory mechanisms into intelligent agents. By combining long-term memory, intelligent agents can retain and retrieve previous user information during the conversation process, thereby improving the model's context understanding ability and answer consistency. However, simply extending the memory range of large models is not sufficient to solve all problems. Users' questions are often vague and do not have a consistent structure, which increases the difficulty of generating accurate answers. Therefore, how to maintain the consistency of answers before and after in complex conversation scenarios has become an important challenge in the current application of large models and intelligent agents. The literature [Is Prompt All You NeedNo. A Comprehensive and Broader View of Instruction Learning] points out that by adjusting the way of asking questions or prompt words, the quality of large models' answers can be improved to a certain extent. Summary of the Invention

[0004] To solve the consistency problem of large models in long conversations, the purpose of the present invention is to provide an intelligent Agent thinking method for maintaining content consistency by integrating longer memories, which adjusts the user's question, obtains the original question answer and the adjusted question answer, and uses the Mamba network model to select from them, so as to obtain a question and answer that better conforms to the user's intention. Since the Mamba network has better reasoning effect on longer memories, more consistent answers can be obtained before and after, alleviating the shortcoming that it is difficult to ensure the consistency of answers before and after in the long conversation of the Transformer structure adopted by large models.

[0005] The purpose of the present invention is achieved by the following technical solutions:

[0006] An intelligent Agent thinking method for maintaining content consistency by integrating longer memories, characterized by including the following steps:

[0007] 1) Ask the user's question in another way without processing the prompts and examples;

[0008] 2) Submit the two questions to the large model of the intelligent Agent for answering respectively, and binary quantize the answers;

[0009] 3) Use the embedding algorithm to convert the question, the original answer and the adjusted question answer into vectors, calculate the difference vector between the original answer vector and the adjusted question answer, connect the question vector and the difference vector into a long vector, use the long vector as the input and the binary quantization result of the Agent answer as the output to construct a set of input-output pairs, and organize the data set to train the Mamba model;

[0010] 4) When the user asks a new question, construct the long vector input and the two questions of the Agent according to the foregoing method, use the trained Mamba model to make a 0, 1 prediction, and select the answer of the appropriate question according to the prediction result and return it to the user.

[0011] This method can obtain a question answer that maintains content consistency by integrating longer memories compared with the original answer of the large model.

[0012] In the present invention, the step of asking the user's question in another way without processing the prompts and examples is as follows:

[0013] Step 11: Set the current chat history record of the Agent as M, and back up M, that is, let Bak_M = M;

[0014] Step 12: Record the user's original question Q;

[0015] Step 13: Remove the prompt and example context in the original question, and ask the built-in large model integrated in the intelligent Agent: "What other question can this question be asked in? Try to make the words as dissimilar as possible";

[0016] Step 14. Obtain the answer Q' of the large model.

[0017] The steps of separately submitting the two questions to the large model of the intelligent Agent for answering and performing binary quantization on the answers are as follows:

[0018] Step 21. Restore the chat record history M of the Agent using the stored Bak_M, submit the original question Q to the large model for answering, and obtain the answer A;

[0019] Step 22. Restore the chat record history M of the Agent using the stored Bak_M, submit the rewritten question Q' to the large model of the intelligent agent for answering, and obtain the answer A';

[0020] Step 23. Manually judge the two answers A and A'. When the answer A is better than A', the binary variable B is 1; otherwise, the binary variable B is 0.

[0021] The steps of converting the question, the original answer, and the adjusted question answer into vectors using the embedding algorithm, calculating the difference vector between the original answer vector and the adjusted question answer, connecting the question vector and the difference vector into a long vector, using the long vector as the input and the binary quantization result of the Agent answer as the output, constructing a set of input-output pairs, and organizing the dataset to train the Mamba model are as follows:

[0022] Step 31. Use the embedding algorithm to map the question Q into a vector V, and use the embedding algorithm to map each sentence of the answer A into a vector where i is the i-th sentence, there are S sentences in total, assuming the vector has N dimensions, take V′=(V′ 1 , V′ 2 ,..., V′ j ,...V′ N ), 1 ≤ j ≤ N, where Use the embedding algorithm to map each sentence of the answer A' into a vector where i is the i-th sentence, there are S' sentences in total, assuming the vector has N dimensions, take V”=(V″ 1 , V″ 2 , …, V″ j , …V″ N ), 1 ≤ j ≤ N, where

[0023] Step 32. Normalize V to the range [-1, 1], let V - = V′ - V″, connect the vectors V and V - to obtain the connected long vector

[0024] Step 33: Quantize V. Quantize it into one level every 0.25. In this way, V is quantized into 8 levels, which is represented by 4 bits, V - is quantized. Quantize it into one level every 0.25. In this way, V is quantized into 16 levels, which is represented by 4 bits. Thus, each component in is represented by 8 bits. For perform compressed storage. Suppose after compression it is such as if each original component in is represented by 4 bytes, then after compression each component is represented by 1 byte, and every 4 bytes form a new component. Thus, the dimension of the input vector is compressed to 1 / 4 to 1 / 8 of the original. In this way, a set of input-output pairs, that is, samples

[0025] Step 34: Suppose is the sample sequence composed of sequences, and B = (B(1), B(2),..., B(t),...) is the output sequence composed of B sequences. Use V' and the sequence B as the training set of the Mamba model to train the Mamba model MambaModel.

[0026] When the user raises a new question, the steps of constructing the long vector input and the Agent's two types of questions according to the foregoing method, using the trained Mamba model to perform 0, 1 prediction, and selecting the appropriate answer to the question according to the prediction result and returning it to the user are as follows:

[0027] Step 41: Suppose the chat history record of the Agent at the current moment is M, and back up M, that is, let Bak_M = M;

[0028] Step 42: When the user raises a new question, suppose it is question QU, and call the large model integrated by the Agent to obtain the answer AU;

[0029] Step 43: Restore the chat history record M to Bak_M, remove the prompts and example contexts in the original question, and ask the internal large model integrated by the intelligent Agent: "What other kinds of questions can this question be asked in?" to obtain the answer QU' of the large model;

[0030] Step 44: Restore the chat history record M to Bak_M, and ask the large model of the Agent with the question QU' to obtain the answer AU';

[0031] Step 45: Use the embedding algorithm to map the question QU into a vector VU;

[0032] Step 46: Use the embedding algorithm to map each sentence of answer AU into a vector, and then take the maximum value of each component of all sentence vectors to form vector VU';

[0033] Step 47: Use the embedding algorithm to map each sentence of answer AU' into a vector, and then take the maximum value of each component of all sentence vectors to form vector VU";

[0034] Step 48: Calculate VU - = VU′ - VU", concatenate vector AU and VU - to obtain the concatenated long vector Quantize and compress the vector to obtain vector Use MambaModel for prediction to obtain the predicted value BU;

[0035] Step 49: When BU = 0, feedback AU' to the user as the answer for this time, otherwise feedback AU to the user as the answer for this time.

[0036] The beneficial effects of the present invention are as follows:

[0037] By adjusting the user's question, the present invention obtains the answers to the original question and the rewritten question, and uses the Mamba network to select the answer that best meets the user's intention. This method not only improves the memory ability of the large model, but also optimizes the question structure to ensure that the intelligent Agent can provide more coherent answers when dealing with complex conversations. The core of this method is to use the embedding algorithm to convert the user's question into a vector, construct a long vector input, and select the answers to multiple questions through the Mamba model. This method significantly improves the reasoning ability of the intelligent Agent in long-term conversations, solves the defect of inconsistent front and back information of the large model in dealing with long conversations, and improves the overall user experience at the same time. Description of the Drawings

[0038] Figure 1 It is the structural diagram of the training part of the present invention.

[0039] Figure 2 It is the normal working structural diagram of the present invention. Detailed Embodiment

[0040] An intelligent Agent thinking method for fusing longer memory and maintaining content consistency is realized through the following steps:

[0041] 1) Ask the user's question in another way without processing the prompts and examples; specifically as follows:

[0042] Step 11: Set the current chat history of the Agent as M, and back up M, i.e., let Bak_M = M;

[0043] Step 12: Record the user's original question Q;

[0044] Step 13: Remove the hints and example contexts from the original question, and ask the internal large model integrated in the intelligent Agent: "What other questions can this question be rephrased into? Try to make the words as dissimilar as possible";

[0045] Step 4: Obtain the answer Q' from the large model.

[0046] 2) Submit the two questions to the large model of the intelligent Agent for answering respectively, and binary quantize the answers; specifically as follows:

[0047] Step 21: Restore the chat record history M of the Agent with the stored Bak_M, submit the original question Q to the large model for answering, and obtain the answer A;

[0048] Step 22: Restore the chat record history M of the Agent with the stored Bak_M, submit the rewritten question Q' to the large model of the intelligent agent for answering, and obtain the answer A';

[0049] Step 23: Manually judge the two answers A and A'. When the answer A is better than A', the binary variable B is set to 1, otherwise the binary variable B is set to 0.

[0050] 3) Use the embedding algorithm to convert the question, the original answer, and the adjusted question answer into vectors, calculate the difference vector between the original answer vector and the adjusted question answer vector, concatenate the question vector and the difference vector into a long vector, use the long vector as the input and the binary quantization result of the Agent answer as the output to construct a set of input-output pairs, and organize the dataset to train the Mamba model; specifically as follows:

[0051] Step 31: Use the embedding algorithm to map the question Q into a vector V, and use the embedding algorithm to map each sentence of the answer A into a vector where i is the i-th sentence, there are S sentences in total, assuming the vector has N dimensions, take V'=(V′ 1 , V′ 2 ,..., V′ j ,...V′ N ), 1≤j≤N, where Use the embedding algorithm to map each sentence of the answer A' into a vector where i is the i-th sentence, there are S' sentences in total, assuming the vector has N dimensions, take V"=(V″ 1 , V″ 2, …, V″ j , …V″ N ), 1 ≤ j ≤ N, where

[0052] Step 32: Normalize V to the range [-1, 1], let V - = V′ - V″, concatenate the vectors V and V - to obtain the concatenated long vector

[0053] Step 33: Quantize V, quantize every 0.25 as one level, so V is quantized into 8 levels and represented by 4 bits, V - is quantized, quantize every 0.25 as one level, so V is quantized into 16 levels and represented by 4 bits, thus each component is represented by 8 bits, compress and store , assume it is compressed to As each original component in is represented by 4 bytes, then after compression each component is represented by 1 byte, and every 4 bytes form a new component, thus reducing the dimension of the input vector to 1 / 4 to 1 / 8 of the original, and obtaining a set of input-output pairs, that is, samples

[0054] Step 34: Assume is the sample sequence composed of sequences, B = (B(1), B(2),....B(t),...) is the output sequence composed of B sequences, use V′ and the sequence B as the training set of the Mamba model to train the Mamba model MambaModel.

[0055] 4) When the user asks a new question, construct the long vector input and the Agent's two types of questions according to the aforementioned method, use the trained Mamba model to perform 0, 1 prediction, and select the appropriate answer to the question according to the prediction result and return it to the user, specifically as follows:

[0056] Step 41: Assume the chat history of the Agent at the current moment is M, back up M, that is, let Bak_M = M;

[0057] Step 42: When the user asks a new question, assume it is the question QU, call the large model integrated by the Agent to obtain the answer AU;

[0058] Step 43: Restore the chat history M to Bak_M, remove the prompts and example contexts in the original question, and ask the internal large model integrated by the intelligent Agent: "What other kinds of questions can this question be asked in?" to obtain the answer QU' of the large model;

[0059] Step 44: Restore the chat history record M to Bak_M, ask the large model of the Agent with the question QU’, and obtain the answer AU’;

[0060] Step 45: Use the embedding algorithm to map the question QU into a vector VU;

[0061] Step 46: Use the embedding algorithm to map each sentence of the answer AU into a vector, and then take the maximum value of each component of all sentence vectors to form a vector VU’;

[0062] Step 47: Use the embedding algorithm to map each sentence of the answer AU’ into a vector, and then take the maximum value of each component of all sentence vectors to form a vector VU”;

[0063] Step 48: Calculate VU - = VU′ - VU", concatenate the vectors AU and VU - to obtain the concatenated long vector Quantize and compress the vector to obtain the vector Use MambaModel for prediction to obtain the predicted value BU;

[0064] Step 49: When BU = 0, feedback AU' as the answer to the user this time, otherwise feedback AU as the answer to the user this time.

[0065] This method is used to improve the question - answering effect of the intelligent Agent based on the large model, and can obtain question answers that fuse longer memories and maintain content consistency compared with the original answers of the large model.

Claims

1. A method for maintaining content consistency by integrating longer memory, characterized in that: The following steps are involved: 1) Ask the user's question in a different way, but do not process the prompts and examples; 2) Submit the two questions to the large model of the intelligent agent for answering, and quantize the answers into binary values; 3) Use the embedding algorithm to convert the question, the original answer, and the adjusted question answer into a vector, Calculate the difference vector between the original answer vector and the adjusted question answer, connect the question vector and the difference vector into a long vector, use the long vector as input, and use the binarized result of the Agent answer as output. Construct a set of input-output pairs, connect the question vector and the difference vector into a long vector, use the long vector as input, and the complement of the binary result of the Agent answer as output, and organize the data set for training Mamba model; 4) When the user asks a new question, two types of questions, long vector input and Agent, are constructed according to the above method. Use the trained Mamba model to make 0, 1 predictions, and select the appropriate answer to the question based on the prediction results and return it to the user.

2. According to claim 1, a method for maintaining content consistency by integrating longer memory, characterized in that: Step 1) is as follows: Step 11, let the current chat history record of Agent be M, and backup M, that is, let Bak_M = M; Step 12: Record the user's original question Q; Step 13: Remove the prompts and example context from the original question and integrate it into the intelligent agent. The inner model asks: "What other questions can be asked in this way? The texts should be as dissimilar as possible." Step 14: Get the answer Q' of the large model.

3. According to claim 2, a method for maintaining content consistency by integrating longer memory, characterized in that: Step 2) is as follows: Step 21: Use the stored Bak_M to restore the chat history M of the Agent, submit the original question Q to the big model to answer, and get the answer A; Step 22, use the stored Bak_M to restore the chat history M of the Agent, submit the rewritten question Q' to the big model of the intelligent agent to answer, and get the answer A'; Step 23: Manually judge the two answers A and A'. When answer A is better than A', binary variable B is converted to is 1, otherwise the binary variable B is 0.

4. According to claim 3, a method for maintaining content consistency by integrating longer memory, characterized in that: Step 3) is as follows: Step 31: Use the embedding algorithm to map the question Q into a vector V, and use the embedding algorithm to map each sentence of the answer A into a vector Where i is the i-th sentence, there are S sentences in total, Assume that the vector has N dimensions, and take V'=(V'1,V'2,...,V' j ,...V' N ),1≤j≤N, where Use the embedding algorithm to map each sentence of answer A' into a vector Where i is the i-th sentence, there are S' sentences in total, and the vector has N dimensions. V”=(V”1,V”2,...,V” j ,...V” N ),1≤j≤N, where Step 32: Normalize V to the range [-1,1], and let V - =V'-V", vector V, and V - Connect and get the connected long vector Step 33: quantize V, with each 0.25 being quantized into one level. Thus, V is quantized into 8 levels, represented by 4 bits. V - Quantize, every 0.25 quantization as a level, so V is quantized into 16 levels, represented by 4 bits, so each The components in are represented by 8 bits. Compress and store, set the compressed like Each component in the original is represented by 4 bytes, so after compression Each component is represented by 1 byte, and every 4 bytes form a new component, so that the input vector dimension is compressed to 1 / 4 to 1 / 8 of the original, thus obtaining a set of input-output pairs, i.e., samples Step 34: Set for is a sample sequence composed of sequences, B = (B(1), B(2),, ...B(t), ...) is an output sequence composed of sequences B, and V' and sequence B are used as the training sets of the Mamba model to train the Mamba model MambaModel.

5. According to claim 1, a method for maintaining content consistency by integrating longer memory, characterized in that: Step 4) is as follows: Step 41, assuming that the current chat history record of the Agent is M, and the backup is M, that is, Bak_M=M; Step 42: When the user raises a new question, set it as question QU and call the large model integrated by Agent. Type to get the answer AU; Step 43, restore the chat history record M to Bak_M, remove the prompt and example context in the original question, and ask the internal large model integrated by the intelligent agent: "What other questions can be asked in other ways?" Get the answer QU' of the large model; Step 44, restore the chat history record M to Bak_M, ask the Agent's large model a question with question QU', and obtain the answer AU'; Step 45: Use the embedding algorithm to map the question QU into a vector VU; Step 46: Use the embedding algorithm to map each sentence of the answer AU into a vector, and then take the maximum value of each component of all sentence vectors to form a vector VU'; Step 47, use the embedding algorithm to map each sentence of the answer AU' into a vector, and then take the maximum value of each component of all sentence vectors to form a vector VU'; Step 48: Calculate VU - =VU'-VU", vector AU and VU - Connect and get the connected long vector Pair Vector Quantize and compress to get vector Use MambaModel to make predictions and get the predicted value BU; Step 49: When BU=0, AU' is fed back to the user as the answer this time; otherwise, AU is fed back to the user as the answer this time.