Method, device, equipment for constructing large model memory based on message tree, and long-term memory generation method

Through the message tree-based large-model memory construction method, the problem of imbalance in energy consumption and accuracy of language models when answering questions is solved, efficient long-term and short-term memory construction and management are achieved, and system stability and user experience are improved.

CN120086265BActive Publication Date: 2025-06-24XINGFAN XINGQI (CHENGDU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510580707.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-24
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

In the prior art, the energy consumption and accuracy of language models are unbalanced when answering questions, resulting in low memory construction efficiency, loss or omission of information, and poor system complexity and stability.

Method used

The large-scale memory construction method based on the message tree is adopted, and the conversation message tree is obtained and constructed, and the timestamps and features in the message tree are used to sort and remember construction, so as to generate and manage short-term memory and long-term memory.

Benefits of technology

Without using additional vector embedding models and vector databases, long and short-term memory for large language model dialogue is realized, which improves memory construction efficiency and accuracy, and reduces system complexity and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086265B_ABST
    Figure CN120086265B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment for constructing large model memory based on a message tree, and a long-term memory generation method. The method includes: obtaining a plurality of messages and constructing the plurality of messages into a conversation; constructing the conversation into a message tree; tracing with a target message ID as a clue to obtain a linear message list corresponding to the target message ID; sorting the messages in the linear message list according to the timestamps and first features of the messages in the linear message list, or only according to the timestamps; obtaining short-term memory according to the sorting result; obtaining long-term memory after obtaining the short-term memory; adding the short-term memory and the long-term memory to a preset prompt template of a target model for user use. The present invention belongs to the field of artificial intelligence technology. The present invention can improve the long-term and short-term memory of large language model conversations without using additional vector embedding models and vector databases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, equipment for constructing a large model memory based on a message tree, and a long-term memory generation method. Background Art

[0002] In the use of large language models, the "memory" method is generally used to implement multi-turn conversations. Memory refers to combining the information of historical conversations in the form of a certain prompt word with the current question, so that the model can answer the current question according to the historical conversation information when reasoning. Existing memory construction schemes generally fall into two categories: directly filling historical information until the context length reaches the limit; or using a vector database to store memory, and retrieving the content with the highest semantic similarity as memory during the conversation.

[0003] The method of directly filling historical information is very likely to exceed the context length limit as the number of conversations increases. Therefore, information truncation is required, which actually loses the early historical information and only retains the recent historical information. In the method using a vector database, the relevant information retrieval will be relatively accurate, but it only contains a small part of the historical information with a high semantic similarity to the question sentence, which may lead to the loss and omission of key information in the remaining conversations. In addition, using a vector library also requires an additional vector embedding model, which not only takes extra time for vector embedding and retrieval in each conversation, but also requires an additional vector database system and a vector model service system, resulting in a delay in the request response speed, increasing the complexity and instability of the system, and consuming additional server resources. Therefore, the present invention provides a method for constructing a large model memory based on a message tree to solve the above problems. Summary of the Invention

[0004] By providing a method, device, equipment for constructing a large model memory based on a message tree, and a long-term memory generation method, the present invention solves the technical problem of the imbalance between energy consumption and accuracy in the question answering of language models in the prior art, and realizes the technical effect of being able to achieve long-term and short-term memory of large language model conversations without using an additional vector embedding model and a vector database.

[0005] In a first aspect, the present invention provides a method for constructing a large model memory based on a message tree, the method comprising:

[0006] Obtain a plurality of messages, and construct the plurality of messages into a conversation, where the messages at least include a message ID, a parent message ID, a user input, a model output, a timestamp, and a byte occupancy, and each message is obtained after fine-tuning the description of the user input or fine-tuning the conversation parameters;

[0007] Construct the conversation into a message tree, and determine a target message ID, the message tree including a plurality of branches, and each branch corresponding to a root node;

[0008] Trace back using the target message ID until reaching the root node of the target message ID in the message tree, so as to obtain the linear message list corresponding to the target message ID;

[0009] Sort the messages in the linear message list according to the timestamps and the first features of the messages in the linear message list, or only according to the timestamps, to obtain a sorting result, where the first features include byte occupancy and content relevance;

[0010] According to the sorting result, add the messages to the short-term memory list in sequence. When the preset length limit or the preset window size limit is reached, obtain the short-term memory;

[0011] After obtaining the short-term memory, fill the remaining messages in the linear message list into a preset long-term memory prompt template to obtain the long-term memory, where the long-term memory is summary-like information;

[0012] Add the short-term memory and the long-term memory to the preset prompt word template of the target model for the user to use.

[0013] Furthermore, sort the messages in the linear message list according to the timestamps and the first features of the messages in the linear message list to obtain a sorting result, including:

[0014] Determine the time difference between each message in the linear message list and the latest message in the linear message list according to the timestamps of the messages in the linear message list;

[0015] Determine the sorting feature of the message according to the time difference between the messages in the linear message list and the first feature;

[0016] Sort each message according to the sorting features of the messages in the linear message list to obtain a sorting result.

[0017] Furthermore, determine the sorting feature of the message according to the time difference between the messages in the linear message list and the first feature, including:

[0018]

[0019] Among them, is the sorting feature of the th message, is the content relevance between the th message and the latest message, is the time difference between the th message and the latest message, is the byte occupancy of the th message.

[0020] Further, sort the messages in the linear message list according to the timestamps of the messages in the linear message list to obtain a sorting result, including:

[0021] Determine the time difference between each message in the linear message list and the latest message in the linear message list according to the timestamps of the messages in the linear message list;

[0022] Sort each message according to the time difference between the messages in the linear message list to obtain a sorting result.

[0023] Further, obtain a number of messages, including:

[0024] Obtain the user's initial input to the model and obtain the message corresponding to the user's initial input;

[0025] Perform semantic replacement on the user's initial input, or perform synonym replacement on the keywords in the user's initial input, or add output restriction conditions for the model output, or replace the model to obtain a number of other messages.

[0026] Further, according to the sorting result, add the messages to the short-term memory list in sequence. When the preset length limit or preset window size limit is reached, obtain the short-term memory, including:

[0027] Reverse the sorting result;

[0028] Input the messages into the short-term memory list in a preset structured form according to the reversed sorting result;

[0029] When the preset length limit or preset window size limit of the short-term memory list is reached, stop the input.

[0030] Further, based on the message tree, determine the content relevance, including:

[0031] Determine the user input of the latest message;

[0032] Determine the branch length between the latest message and the target message in the linear message list, and the call frequency of the node where the target message is located;

[0033] Determine the call time consumption according to the branch length;

[0034] According to the call time consumption and the call frequency, determine the content relevance between the latest message and the target message in the linear message list, including:

[0035]

[0036] Among them, is the call frequency of the th message, is the The call time of a message is the call time between the message of the root node of the branch where the -th message is located and the latest message.

[0037] In a second aspect, the present invention provides a long-term memory generation method, the method comprising:

[0038] After the conversation ends, determine the target message ID;

[0039] According to the target message ID, extract a target message list from the message tree;

[0040] Continuously add the messages in the target message list to a preset prompt template in descending order of time until the preset length limit is reached;

[0041] For the messages not extracted in the target message list, based on a preset large language model, extract and compress the entities, relationships, and events of each message in the target message list to obtain long-term memory.

[0042] In a third aspect, the present invention provides a large model memory construction device based on a message tree, the device comprising:

[0043] A session module, configured to obtain a plurality of messages and construct the plurality of messages into a session, where the messages at least include a message ID, a parent message ID, user input, model output, timestamp, and byte occupancy, and each message is obtained after fine-tuning the description of the user input or fine-tuning the conversation parameters;

[0044] A message tree module, configured to construct the session into a message tree and determine the target message ID, the message tree including a plurality of branches, and each branch corresponding to a root node;

[0045] A message list module, configured to trace back using the target message ID as a clue until reaching the root node of the target message ID in the message tree, so as to obtain a linear message list corresponding to the target message ID;

[0046] A sorting module, configured to sort the messages in the linear message list according to the timestamp and the first feature of each message in the linear message list, or only the timestamp, to obtain a sorting result, where the first feature includes byte occupancy and content relevance;

[0047] A short-term memory module, configured to sequentially add the messages to a short-term memory list according to the sorting result, and obtain short-term memory when the preset length limit or the preset window size limit is reached;

[0048] A long-term memory module, which is used to fill the remaining messages in the linear message list into a preset long-term memory prompt template after obtaining the short-term memory, so as to obtain long-term memory, where the long-term memory is class summary information;

[0049] A prompt word module, which is used to add the short-term memory and the long-term memory to a preset prompt word template of a target model for user use.

[0050] In a fourth aspect, the present invention provides an electronic device, including:

[0051] A processor;

[0052] A memory for storing processor-executable instructions;

[0053] Wherein, the processor is configured to execute to implement the method for constructing a large model memory based on a message tree provided in the first aspect.

[0054] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:

[0055] The present invention can realize the long-term and short-term memory of large language model conversations without using an additional vector embedding model and a vector database. It can not only retain the global key information in the long-term history and the full amount of information in the recent history in memory, but also realize the millisecond-level construction of memory, improve the response speed of dialogue requests, reduce the consumption of additional database and model resources, improve system stability, and support the archiving and viewing of all dialogue branches in the session, thus enhancing the user experience.

[0056] By semantic replacement, or synonym replacement of keywords in the user's initial input, or adding output limit conditions to the model output, or replacing the model, the present invention can make the generated messages richer. The generated messages are not only relevant in content, but also can make the message tree in the subsequent text have more root nodes, enriching the tree structure of the message tree, facilitating the message tree to learn the generation rules of different models, and providing help for generating richer short-term and long-term memories.

[0057] The present invention can ensure the timing characteristics in the short-term memory through the time difference, ensuring the freshness and timeliness of the short-term memory. Since the short-term memory itself has the characteristic of small byte occupancy, the present invention can ensure the conciseness and high correlation of the subsequent messages incorporated into the short-term memory through the first feature, and can also ensure that more relevant messages can be accommodated in the short-term memory as much as possible due to the consideration of the byte occupancy of the message itself.

[0058] When determining the content relevance, the present invention combines the structural characteristics of the message tree. The call frequency usually indicates the status of the message in the message tree (indicating a higher degree of combination with the current message); while the call time-consuming reflects the transmission difficulty of the message (the call time-consuming includes the time spent by the message passing through the branches and the time for calling the node itself). The longer the time, the lower the efficiency of calling the node and the less likely it is to be called. The present invention determines the content relevance based on the call time-consuming and the call frequency, and can achieve the balance between efficiency and content. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0060] Figure 1 Flow chart of a method for constructing a large model memory based on a message tree provided by the present invention;

[0061] Figure 2 Schematic diagram of the message tree provided by the present invention;

[0062] Figure 3 Flow chart of another method for constructing a large model memory based on a message tree provided by the present invention;

[0063] Figure 4 Flow chart of the whole process of the method for constructing a large model memory based on a message tree provided by the present invention and a method for generating long-term memory. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The embodiments of the present invention provide a method for constructing a large model memory based on a message tree, which solves the technical problem of the imbalance between energy consumption and accuracy of language models when answering questions in the prior art.

[0065] The technical solution of the present invention for solving the above technical problem is generally as follows:

[0066] Method for constructing large model memory based on message tree, the method comprising: obtaining a plurality of messages and constructing the plurality of messages into a conversation, wherein the messages at least include a message ID, a parent message ID, user input, model output, a timestamp and byte occupancy, and each message is obtained after fine-tuning the description of the user input or fine-tuning the conversation parameters; constructing the conversation into a message tree and determining a target message ID, the message tree including a plurality of branches, each branch corresponding to a root node; tracing back with the target message ID as a clue until reaching the root node of the target message ID in the message tree, thereby obtaining a linear message list corresponding to the target message ID; sorting the messages in the linear message list according to the timestamps and a first feature of the messages in the linear message list, or only according to the timestamps, to obtain a sorting result, wherein the first feature includes byte occupancy and content relevance; adding the messages to a short-term memory list in sequence according to the sorting result, and when a preset length limit or a preset window size limit is reached, obtaining a short-term memory; after obtaining the short-term memory, filling the remaining messages in the linear message list into a preset long-term memory prompt template to obtain a long-term memory, wherein the long-term memory is summary-like information; adding the short-term memory and the long-term memory to a preset prompt word template of a target model for user use.

[0067] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0068] First, it should be noted that the term "and / or" appearing in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and rear associated objects.

[0069] The present invention provides a Figure 1 method for constructing large model memory based on message tree as shown, including steps S11-S17:

[0070] Step S11, obtaining a plurality of messages and constructing the plurality of messages into a conversation, wherein the messages at least include a message ID, a parent message ID, user input, model output, a timestamp and byte occupancy, and each message is obtained after fine-tuning the description of the user input or fine-tuning the conversation parameters.

[0071] The model in the present invention may refer to a large language model, such as ChatGPT, DeepSeek, and Doubao, etc. In the large language model question-answering and dialogue systems, the rule of one question and one answer is followed, that is, there is only one output for each user input. Therefore, a question-answer combination can be modeled as one message. The message ID refers to the identifier of each message, and the message ID ensures that each message in the system can be uniquely identified and referenced; the parent message ID refers to the ID of the previous message that the current message directly responds to or is associated with. By using the message ID and the parent message ID in pair, a message chain can be constructed.

[0072] For large language models, even if the user inputs are the same, the outputs of the large language models may be different, or, if the user needs to make minor modifications to some descriptions, dialogue parameters, etc. in the input, the outputs of the large language models are very likely to be different.

[0073] To expand the messages and make the model outputs related to each other, obtaining several messages may include: obtaining the user's initial input to the model and getting the message corresponding to the user's initial input; performing semantic substitution on the user's initial input, or performing synonym substitution on the keywords in the user's initial input, or adding output limiting conditions for the model output, or replacing the model, so as to obtain several other messages.

[0074] Both semantic substitution and synonym substitution of keywords can be implemented according to work experience or with the help of the language model, which is not limited here; adding output limiting conditions for the model output can be input in the generation box of the language model, such as: "Controlling the number of words in the answer within 300 words", "Highlighting a certain part", etc.

[0075] It should be emphasized that by performing semantic substitution on the user's initial input, or performing synonym substitution on the keywords in the user's initial input, or adding output limiting conditions for the model output, it is not only to increase the number of messages, but also to make the messages related to each other.

[0076] In addition, model replacement can also be carried out. Different models may generate different model outputs for the same question. It can be understood that the output logics of different models are different. Generating messages with different models can make the generated messages not only relevant in content, but also enable the message tree in the following text to have more root nodes, enriching the tree structure of the message tree, facilitating the message tree to learn the generation rules of different models, and providing help for generating richer short-term memories and long-term memories.

[0077] It can be understood that semantic substitution, or synonym substitution of keywords in the user's initial input, or adding output limiting conditions for the model output, or replacing the model can be used in combination, or used alone, or partially used.

[0078] The description fine-tuning can be keyword replacement, semantic replacement, etc.; the dialogue parameter fine-tuning can be adding output limit conditions, etc., which are not restricted here.

[0079] The message can also include a session ID, long-term memory, etc.

[0080] Step S12: Construct the conversation into a message tree and determine the target message ID. The message tree includes several branches, and each branch corresponds to a root node.

[0081] After obtaining the conversation, there are countless historical linear branches in the conversation. Therefore, the conversation is modeled as a message tree. There can be multiple root nodes in the message tree (as long as a node has no parent node, it is regarded as a root node).

[0082] Step S13: Trace back using the target message ID as a clue until reaching the root node of the target message ID in the message tree, so as to obtain the linear message list corresponding to the target message ID.

[0083] The message tree is a tree-like data structure used to represent the call relationship of message passing in a distributed system.

[0084] The target message ID is the unique identifier of a certain message in the message tree, used to locate a specific node in the message tree. Tracing back means starting from the node corresponding to the target message ID and backtracking upward along the parent nodes of the message tree until reaching the root node. The linear message list is the complete call path from the root node to the target message node, arranged in the call order. Through the linear message list, the call background and dependency relationship of the target message can be fully understood.

[0085] Step S14: Sort the messages in the linear message list according to the timestamps and the first feature of each message in the linear message list, or only the timestamps, to obtain a sorting result, where the first feature includes byte occupancy and content relevance.

[0086]

Method 1

[0087] Determine the time difference between each message in the linear message list and the latest message in the linear message list according to the timestamps of each message in the linear message list; determine the sorting feature of the message according to the time difference of the messages in the linear message list and the first feature; sort each message according to the sorting features of each message in the linear message list to obtain a sorting result.

[0088] Determine the sorting feature of the message according to the time difference of the messages in the linear message list and the first feature, including:

[0089]

[0090] Among them, is the sorting feature of the th message, is the content correlation degree between the th message and the latest message, is the time difference between the th message and the latest message, is the byte occupancy of the th message.

[0091] It should be noted that the latest message may not be included in the short-term memory or long-term memory.

[0092] The time difference can ensure the temporal characteristics in the short-term memory and ensure the freshness and timeliness of the short-term memory. Since the short-term memory itself has the characteristic of small byte occupancy, the first feature of the present invention can ensure the conciseness and high correlation degree of the subsequent messages included in the short-term memory. And because the byte occupancy of the message itself is considered, it is possible to ensure that more relevant messages are accommodated in the short-term memory as much as possible.

[0093] The content correlation degree can be realized by means of keyword matching, semantic similarity analysis, topic modeling, deep learning methods, etc. In addition, the present invention also provides a method for determining the content correlation degree based on the tree structure itself, including:

[0094] Determine the user input of the latest message; determine the branch length between the latest message and the target message in the linear message list, and the call frequency of the node where the target message is located; determine the call time according to the branch length; determine the content correlation degree between the latest message and the target message in the linear message list according to the call time and the call frequency, including:

[0095]

[0096] Among them, is the call frequency of the th message, is the call time of the th message, is the call time between the message at the root node of the branch where the th message is located and the latest message.

[0097] When determining the content relevance, the present invention combines the structural characteristics of the message tree. The call frequency usually indicates the status of the message in the message tree (indicating a higher degree of combination with the current message); while the call time consumption reflects the transmission difficulty of the message (the call time consumption includes the time taken for the message to pass through the branches and the time for calling the node itself). The longer the time, the lower the efficiency of calling the node and the less likely it is to be called. The present invention determines the content relevance based on the call time consumption and the call frequency, and can achieve a balance between efficiency and content.

[0098]

Method 2

[0099] According to the timestamps of each message in the linear message list, determine the time difference between each message in the linear message list and the latest message in the linear message list; according to the time differences of the messages in the linear message list, sort each message to obtain a sorting result.

[0100] The messages with smaller time differences are arranged in the front. Method 2 uses the time difference as the only basis for arrangement, which can make the messages entering the short-term memory completely conform to the timeline and the questioning logic of the timeline.

[0101] Step S15, according to the sorting result, add the messages to the short-term memory list in sequence. When the preset length limit or the preset window size limit is reached, obtain the short-term memory.

[0102] Specifically include: reverse the sorting result; according to the reversed sorting result, input the messages into the short-term memory list in a preset structured form; when the preset length limit or the preset window size limit of the short-term memory list is reached, stop the input.

[0103] It can be understood that in Method 1, the messages with larger sorting features need to enter the short-term memory list first, and in Method 2, the messages closer to the latest message need to enter the short-term memory list first. Therefore, by reversing the sorting result, it can be ensured that the messages with larger sorting features or the messages closer to the latest message enter the short-term memory list first.

[0104] The length limit refers to the context that the short-term memory in the large model can support. For example, if the large model supports a context of 32k, the preset length limit can be set to 16k.

[0105] The window refers to the number of dialogue turns. For example, if 10 rounds of conversations (10 messages) are carried out in the current branch, and the preset window size limit is 5 rounds, then at most 5 messages enter the short-term memory.

[0106] The preset structured form can be determined according to the structure of the large model used. The fields of the user input and the model output can be filled into the short-term memory prompt template in the preset structured form as the short-term memory.

[0107] Step S16, after obtaining the short-term memory, fill the remaining messages in the linear message list into a preset long-term memory prompt template to obtain long-term memory, where the long-term memory is class summary information.

[0108] Fill the remaining messages in the linear message list after short-term memory processing into the long-term memory prompt template in a preset structured form as long-term memory (the long-term memory is class summary information formed by the message fields and their historical information, and does not contain any information within the short-term memory window).

[0109] The preset structured form of the short-term memory can be different from that of the long-term memory, and the preset structured form of the long-term memory can be determined according to the structure of the large model used.

[0110] For example, the messages in the linear message list are m1 - m4 (the latest is m5). After screening, the conversation content of m3 and m4 is short-term memory. After extracting the information of the remaining m1 and m2, the summary information corresponding to m1 and m2 is obtained , then is used as the long-term memory.

[0111] Such as Figure 2 shown, for message h, perform branch extraction. a - b - d - h is the linear message list. When the window size is 2, a possible short-term memory is d and h, then it is the long-term memory ( is the class summary information of a and b).

[0112] Step S17, add the short-term memory and the long-term memory to the preset prompt word template of the target model for user use.

[0113] The preset prompt word template can include task instructions, user input, short-term memory, long-term memory, etc.

[0114] Figure 3 This is a schematic flowchart of another method for constructing the memory of a large model based on a message tree provided by the present invention.

[0115] Based on the same inventive concept, the present invention also provides a method for generating long-term memory. The method includes:

[0116] After the conversation ends, determine the target message ID;

[0117] According to the target message ID, extract the target message list from the message tree. The specific extraction method can refer to the above text;

[0118] In descending time order, the messages in the target message list are continuously added to the preset prompt template (i.e., the user input and model output fields of the message are iteratively filled into the prompt template for information extraction), and the message is removed from the target message list until the preset length limit is reached;

[0119] For the messages that are not extracted in the target message list, based on the preset large language model, the entities, relations and events of each message in the target message list are extracted and compressed to obtain long-term memory, that is, the long-term memory field of the latest message is extracted based on the preset large language model, as the long-term memory of the current message containing global key information, and filled into the prompt template of information extraction in a structured form. Each message node in the message tree contains compressed information of all its historical information.

[0120] like Figure 4 As shown, Figure 4 A schematic diagram of the entire process of a large model memory construction method based on a message tree and a long-term memory generation method.

[0121] In summary, the present invention can achieve long-term and short-term memory of large language model dialogues without using additional vector embedding models and vector databases. It can not only retain the global key information in the long-term history and the full amount of information in the recent history in the memory; it can also achieve millisecond-level construction of memory, improve the response speed of dialogue requests; and reduce the consumption of additional database and model resources, improve system stability; and support the archiving and viewing of all dialogue branches in the conversation, thereby improving the user experience.

[0122] The present invention can make the generated messages richer by semantic replacement, or replacing the keywords in the user's initial input with synonyms, or adding output restrictions to the model output, or replacing the model. The generated messages are not only relevant in content, but also can make the message tree in the subsequent text have more root nodes, enriching the tree structure of the message tree, making it easier for the message tree to learn the generation rules of different models, and providing assistance for generating richer short-term memory and long-term memory.

[0123] The present invention can ensure the timing characteristics in the short-term memory through the time difference, and ensure the freshness and timeliness of the short-term memory. Since the short-term memory itself has the characteristic of small byte occupancy, the present invention can ensure the simplicity and high relevance of the subsequent messages included in the short-term memory through the first feature, and since the byte occupancy of the message itself is taken into consideration, it can ensure that more highly relevant messages are accommodated in the short-term memory as much as possible.

[0124] When determining the content relevance, the present invention combines the structural characteristics of the message tree. The call frequency usually indicates the status of the message in the message tree (indicating a higher degree of combination with the current message); while the call time consumption reflects the difficulty of message transmission (the call time consumption includes the time spent by the message passing through the branches and the time for calling the node itself). The longer the time, the lower the efficiency of calling the node and the less likely it is to be called. The present invention determines the content relevance based on the call time consumption and the call frequency, and can achieve a balance between efficiency and content.

[0125] Based on the same inventive concept, the present invention provides a model memory construction device based on a message tree. The device includes:

[0126] A session module, configured to obtain a plurality of messages and construct the plurality of messages into a session, where the messages at least include a message ID, a parent message ID, user input, model output, a timestamp, and byte occupancy, and each message is obtained after fine-tuning the description of the user input or fine-tuning the dialogue parameters;

[0127] A message tree module, configured to construct the session into a message tree and determine a target message ID. The message tree includes a plurality of branches, and each branch corresponds to a root node;

[0128] A message list module, configured to trace back with the target message ID as a clue until reaching the root node of the target message ID in the message tree, so as to obtain a linear message list corresponding to the target message ID;

[0129] A sorting module, configured to sort the messages in the linear message list according to the timestamps and the first feature of the messages in the linear message list, or only the timestamps, to obtain a sorting result, where the first feature includes byte occupancy and content relevance;

[0130] A short-term memory module, configured to sequentially add the messages to a short-term memory list according to the sorting result, and obtain a short-term memory when reaching a preset length limit or a preset window size limit;

[0131] A long-term memory module, configured to, after obtaining the short-term memory, fill the remaining messages in the linear message list into a preset long-term memory prompt template to obtain a long-term memory, where the long-term memory is summary-like information;

[0132] A prompt word module, configured to add the short-term memory and the long-term memory to a preset prompt word template of a target model for user use.

[0133] Based on the same inventive concept, the present invention provides an electronic device, which includes:

[0134] A processor;

[0135] A memory for storing processor-executable instructions;

[0136] Among them, the processor is configured to execute to implement the message tree-based large model memory construction method provided above.

[0137] Since the electronic device introduced in this embodiment is the electronic device adopted for implementing the information processing method in the embodiments of the present invention, based on the information processing method introduced in the embodiments of the present invention, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the implementation of how this electronic device implements the method in the embodiments of the present invention will not be described in detail here. As long as the electronic device adopted by those skilled in the art to implement the information processing method in the embodiments of the present invention belongs to the scope protected by the present invention.

[0138] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0139] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0140] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps of the functions specified in one block or a plurality of blocks.

[0142] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0143] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A large model memory construction method based on a message tree, characterized in that: The method comprises: Obtain a number of messages, and construct the messages into a conversation, wherein the message at least includes a message ID, a parent message ID, a user input, a model output, a timestamp, and byte occupancy, and each message is obtained after fine-tuning the description of the user input or after fine-tuning the conversation parameters; Constructing a conversation as a message tree and determining a target message ID, wherein the message tree includes a plurality of branches, each branch corresponding to a root node; Tracing back with the target message ID as a clue until reaching the root node of the target message ID in the message tree, thereby obtaining a linear message list corresponding to the target message ID; Sorting the messages in the linear message list according to the timestamp and the first feature of each message in the linear message list, or only the timestamp, to obtain a sorting result, wherein the first feature includes byte occupancy and content relevance; According to the sorting results, the messages are added to the short-term memory list in sequence, and when the preset length limit or the preset window size limit is reached, the short-term memory is obtained; After the short-term memory is obtained, the remaining messages in the linear message list are filled into a preset long-term memory prompt template to obtain the long-term memory, wherein the long-term memory is summary-like information; The short-term memory and the long-term memory are added to the preset prompt word template of the target model for use by the user.

2. The large model memory construction method based on the message tree according to claim 1 is characterized in that: Sorting the messages in the linear message list according to the timestamp and the first feature of each message in the linear message list to obtain a sorting result, including: Determine the time difference between each message in the linear message list and the latest message in the linear message list according to the timestamp of each message in the linear message list; Determine a sorting feature of the message according to a time difference of the message in the linear message list and the first feature; According to the sorting characteristics of each message in the linear message list, each message is sorted to obtain a sorting result.

3. The large model memory construction method based on the message tree as claimed in claim 2, characterized in that: Determining the sorting feature of the message according to the time difference and the first feature of the message in the linear message list includes: in, For the The sorting characteristics of the messages, For the The content relevance of the message to the latest message, For the The time difference between the first message and the latest message, For the The bytes occupied by the message.

4. The large model memory construction method based on the message tree according to claim 1 is characterized in that: Sorting the messages in the linear message list according to the timestamps of the messages in the linear message list to obtain a sorting result, including: Determine the time difference between each message in the linear message list and the latest message in the linear message list according to the timestamp of each message in the linear message list; According to the time difference of the messages in the linear message list, the messages are sorted to obtain a sorting result.

5. The large model memory construction method based on the message tree according to claim 1, characterized in that: Get several messages, including: Get the user's initial input to the model and get the message corresponding to the user's initial input; Perform semantic replacement on the user's initial input, or replace the keywords in the user's initial input with synonyms, or increase output restrictions on the model output, or replace the model to obtain several other messages.

6. The large model memory construction method based on the message tree according to claim 1, characterized in that: According to the sorting results, the messages are added to the short-term memory list in sequence. When the preset length limit or the preset window size limit is reached, the short-term memory is obtained, including: Reverse the sorting results; According to the sorting result after flipping, the message is input into the short-term memory list in a preset structured form; When a preset length limit of the short-term memory list or a preset window size limit is reached, input stops.

7. The large model memory construction method based on the message tree as claimed in claim 3 is characterized in that: Based on the message tree, determine the content relevance, including: Determine the user input of the most recent message; Determine the branch length between the latest message and the target message in the linear message list, and the call frequency of the node where the target message is located; Determine the call time based on the length of the branch; Determine the content relevance between the latest message and the target message in the linear message list based on the call duration and call frequency, including: in, For the The call frequency of the message, For the The call time of the message is For the The call time between the root node message of the branch where the message is located and the latest message.

8. A method for generating long-term memory, characterized in that: A large model memory construction method based on a message tree applied to any one of claims 1-7, the method comprising: When the conversation is over, determine the target message ID; Extracting a target message list from the message tree according to the target message ID; Continuously adding messages in the target message list to the preset prompt template in descending time order until the preset length limit is reached; For the messages that are not extracted in the target message list, based on the preset large language model, the entities, relations and events of each message in the target message list are extracted and compressed to obtain long-term memory.

9. A large model memory construction device based on a message tree, characterized in that: The device comprises: A conversation module is used to obtain a number of messages and construct the messages into a conversation, wherein the message at least includes a message ID, a parent message ID, a user input, a model output, a timestamp and byte occupancy, and each message is obtained after fine-tuning the description of the user input or the conversation parameters; A message tree module, used to construct a conversation into a message tree and determine a target message ID, wherein the message tree includes a plurality of branches, each branch corresponding to a root node; A message list module, used to trace back the target message ID as a clue until the target message ID is reached at the root node of the message tree, thereby obtaining a linear message list corresponding to the target message ID; A sorting module, used for sorting the messages in the linear message list according to the timestamp and the first feature of each message in the linear message list, or only the timestamp, to obtain a sorting result, wherein the first feature includes byte occupancy and content relevance; A short-term memory module is used to add messages to the short-term memory list in sequence according to the sorting result, and obtain the short-term memory when a preset length limit or a preset window size limit is reached; A long-term memory module is used to fill the remaining messages in the linear message list into a preset long-term memory prompt template after obtaining the short-term memory to obtain the long-term memory, wherein the long-term memory is summary-like information; The prompt word module is used to add the short-term memory and the long-term memory to the preset prompt word template of the target model for user use.

10. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute to implement the large model memory construction method based on the message tree as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Session interaction method and device, electronic equipment and readable storage medium

    CN119185942A

  • Artificial intelligence framework combining a spiking neural network and a hyperdimensional computing block

    US20230071730A1