Context management method and device of large language model, medium, equipment and product
By using the first and second agents to manage contexts in a large language model, determining unique identifiers related to prompt words, assembling arrays and obtaining target contexts, the efficiency and correlation problems of context management in a multi-agent collaboration system are solved, and the performance of the large language model is improved.
Patent Information
- Application Number
- CN202510796909.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-18
AI Technical Summary
In multi-agent collaboration systems, the context management of large language models has problems of information loss and performance degradation, and how to effectively manage the context to improve efficiency and relevance.
The first proxy determines the message unique identifier related to the prompt word, assembles it into an array and sends it to the second proxy. The second proxy acquires the message and determines the target context, uses a large language model to output, filters the irrelevant context, and reduces irrelevant information interference.
The output with a highly relevant context and task is realized, reducing interference with irrelevant information, and improving the performance and efficiency of large language models.
Smart Images

Figure CN120337906A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of machine learning, and more particularly, to a method, apparatus, medium, device, and product for context management of a large language model. Background Art
[0002] The context of a large language model (LLM) refers to the relevant information that the large language model relies on when understanding and generating text. The context affects the large language model's understanding of the input and the large language model's output. In a multi-agent collaboration system, context management is a key technical issue. Since agents need to process a large amount of information in multiple rounds of execution, this information will quickly fill up the context window, resulting in information loss and performance degradation. Therefore, how to effectively manage the context is extremely important. Summary of the Invention
[0003] This Summary of the Invention section is provided to introduce concepts in a brief form that will be described in detail in the Detailed Implementation section later. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0004] In a first aspect, the present disclosure provides a method for context management of a large language model, including: Through a first agent, according to the received prompt, determine a unique identifier of a message related to the prompt, where the message is a message in the context of the large language model; Through the first agent, assemble the determined unique identifier into an array and send the array to a second agent; Through the second agent, according to the unique identifier included in the array, obtain the message corresponding to the unique identifier, and based on the obtained message, determine the target context of the large language model; Through the large language model included in the second agent, obtain the output result of the large language model according to the target context.
[0005] In a second aspect, the present disclosure provides a device for context management of a large language model, including: A determination module, configured to, through a first agent, according to the received prompt, determine a unique identifier of a message related to the prompt, where the message is a message in the context of the large language model; A sending module, configured to, through the first agent, assemble the determined unique identifier into an array and send the array to a second agent; An acquisition module, configured to obtain, through the second agent, a message corresponding to the unique identifier according to the unique identifier included in the array, and determine a target context of the large language model based on the obtained message; An output module, configured to obtain an output result of the large language model through the large language model included in the second agent according to the target context.
[0006] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processing device, the steps of the method described in the first aspect are implemented.
[0007] In a fourth aspect, the present disclosure provides an electronic device, including: A storage device having a computer program stored thereon; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.
[0008] In a fifth aspect, the present disclosure provides a computer program product including a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0009] Based on the above technical solutions, through the first agent, according to the received prompt, the unique identifier of the message related to the prompt is determined, and through the first agent, the determined unique identifier is assembled into an array and sent to the second agent. Then, through the second agent, according to the unique identifier included in the array, the message corresponding to the unique identifier is obtained, and based on the obtained message, the target context of the large language model is determined, and through the large language model included in the second agent, according to the target context, the output result of the large language model is obtained. This can not only select the context related to the prompt, filter out the irrelevant context, ensure that the context is highly relevant to the task, but also reduce the interference of irrelevant information on the large language model by transmitting the context through the unique identifier in multiple AI Agents (Artificial Intelligence Agents).
[0010] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Combined with the drawings and referring to the following specific implementation manners, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn to scale. In the drawings: Figure 1It is a flowchart of a context management method for a large language model shown according to some embodiments.
[0012] Figure 2 It is a schematic diagram of resources shown according to some embodiments.
[0013] Figure 3 It is a flowchart of a context management method for a large language model shown according to some other embodiments.
[0014] Figure 4 It is a flowchart of context compression shown according to some embodiments.
[0015] Figure 5 It is a flowchart of context compression shown according to some other embodiments.
[0016] Figure 6 It is a schematic structural diagram of a context management device for a large language model shown according to some embodiments.
[0017] Figure 7 It is a schematic structural diagram of an electronic device shown according to some embodiments. Detailed Embodiments
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0019] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0020] The term "including" and its variations used herein are open-ended, i.e., "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0021] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0022] It should be noted that the modifiers "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the target context, it should be understood as "one or more".
[0023] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0024] Figure 1 is a flowchart of a context management method for a large language model shown according to some embodiments. As Figure 1 shown, the embodiments of this disclosure provide a context management method for a large language model, which can specifically be executed by a context management device for a large language model, and this device can be implemented in software and / or hardware. As Figure 1 shown, the method may include the following steps.
[0025] In step 110, through a first agent, according to the received prompt, determine the unique identifier of the message related to the prompt.
[0026] Here, the first agent can be an artificial intelligence agent (AI Agent). The AI Agent takes the large language model as the core, provides natural language processing capabilities, and supports multi-language and multi-domain tasks. The received prompt can refer to the text instruction or question input by the user to the large language model received by the first agent, and the prompt is used to guide the large language model to generate a specific output.
[0027] Among them, the message related to the prompt is the message in the context of the large language model. The messages included in the context of the large language model can refer to messages from the user, messages from the large language model, and tool call results from the large language model, etc.
[0028] The first agent, through the received prompt, according to its own understanding of the task requirements corresponding to the prompt, searches for the message related to the prompt in the context of the large language model and obtains the unique identifier of the message.
[0029] It should be noted that corresponding unique identifiers (Identity document, ID) can be pre-assigned to each message in the context of the large language model, and the unique identifier is used to uniquely identify the message. Among them, a random string, a hash value, a timestamp, or a serial number can be used to generate the unique identifier corresponding to the message.
[0030] It should be noted that by filtering out the messages related to the prompt from the messages included in the context of the large language model, the large language model can perform reasoning based on the context composed of the messages related to the prompt, which not only reduces the transmission of irrelevant information and avoids filling up the context window of the large language model with excessive context information, but also ensures that the transmitted context is highly relevant to the current task corresponding to the prompt.
[0031] In step 120, through the first agent, the determined unique identifiers are assembled into an array and sent to the second agent.
[0032] Here, after the first agent obtains the unique identifiers of the messages related to the prompt, the first agent assembles the unique identifiers into an array. Exemplarily, the obtained unique identifiers can be assembled into an array in order, and the array contains the unique identifiers of all messages related to the prompt.
[0033] Then, the first agent sends the array composed of unique identifiers to the second agent to transmit the context to the second agent in the form of unique identifiers. The second agent can be an artificial intelligence agent. It should be noted that both the first agent and the second agent are artificial intelligence agents. In the embodiments of the present disclosure, the first agent is responsible for interacting with the user at the front end, and the second agent is responsible for the actual reasoning process.
[0034] It should be noted that in the scenario of multi-AI Agent collaboration, the first agent is used to receive the prompt, select the unique identifiers of the messages related to the prompt according to the prompt, assemble the selected unique identifiers into an array, and transmit the array to the second agent. The second agent is used to obtain the target context according to the received array and perform corresponding reasoning based on the target context to obtain the corresponding output result.
[0035] In step 130, through the second agent, according to the unique identifiers included in the array, the messages corresponding to the unique identifiers are obtained, and based on the obtained messages, the target context of the large language model is determined.
[0036] Here, the second agent receives the array sent by the first agent, obtains the messages corresponding to the unique identifiers through the unique identifiers included in the array, and uses all the obtained messages as the target context of the large language model. At this time, the target context for reasoning obtained by the second agent is the most useful and relevant context for the task corresponding to the prompt, filtering out the message content irrelevant to the task.
[0037] It should be understood that since the context is transmitted between the first agent and the second agent through unique identifiers, interference from irrelevant information can be avoided, greatly improving the efficiency of context transmission.
[0038] In step 140, the output result of the large language model is obtained according to the target context through the large language model included in the second agent.
[0039] Here, the core of the second agent is the large language model. After the second agent determines the target context, the second agent inputs the target context and the prompt words into the large language model, so that the large language model conducts reasoning based on the target context and the prompt words to obtain the output result of the large language model. It should be noted that the output result of the large language model can be displayed to the user through the first agent.
[0040] Thus, through the first agent, according to the received prompt words, the unique identifier of the message related to the prompt words is determined. Through the first agent, the determined unique identifiers are assembled into an array and sent to the second agent. Then, through the second agent, according to the unique identifiers included in the array, the messages corresponding to the unique identifiers are obtained, and based on the obtained messages, the target context of the large language model is determined. Moreover, through the large language model included in the second agent, according to the target context, the output result of the large language model is obtained. This can not only select the context related to the prompt words, filter out the irrelevant context, and ensure that the context is highly relevant to the task, but also reduce the interference of irrelevant information to the large language model by transmitting the context among multiple AI Agents through the unique identifier.
[0041] It is worth noting that the context management method of the large language model provided by the embodiments of the present disclosure can be used in scenarios such as large language model-driven dialogue systems, multi-Agent collaboration systems, conversational AI assistants, context management systems in long conversation scenarios, message passing in distributed AI systems, and so on. That is to say, the context management method of the large language model provided by the embodiments of the present disclosure does not require specific system architecture support, can be implemented in a conventional computer system, and can be quickly integrated into existing AI applications.
[0042] In some implementable embodiments, for each message in the context of the large language model, the message can be encapsulated as a resource, and a corresponding unique identifier is assigned to the resource.
[0043] Here, for each message in the context of the large language model, the message can be encapsulated as a resource by the first agent or the second agent, and a corresponding unique identifier (message ID) is assigned to each resource.
[0044] In some embodiments, the message is encapsulated as a resource through the tags of the Extensible Markup Language (XML).
[0045] XML tags are the markers used in XML to define and describe the structure and meaning of data. XML tags are pairs of strings that are used to mark data elements in an XML document, giving them a specific meaning. Each tag consists of a start tag and an end tag, with the data to be marked enclosed between the start and end tags.
[0046] Figure 2 is a schematic diagram of a resource shown according to some embodiments. As Figure 2 shown, the message content of each message is wrapped using XLM tags (), and moreover, each resource contains a corresponding unique identifier ( <resource id=""消息ID”">).
[0047] In specific implementation, the resource tagging system can be initialized so that the resource tagging system is ready to assign unique identifiers to messages. Then, traverse the context, process each message traversed, generate a random string with a length of 6 as the unique identifier of each message, and then wrap the content string of the message with XML tags, and replace the original message in the context with the encapsulated message.
[0048] It should be understood that by encapsulating messages with XML tags, based on basic string operations without complex algorithms, message encapsulation is simpler and the development cost is low. Moreover, XML tags are easy to be parsed and processed by various systems, have good scalability, and are easy to integrate with other systems.
[0049] Of course, in other embodiments, JSON (an open standard file format and data exchange format) format or specific delimiter tags for resources can also be used to encapsulate messages as resources.
[0050] Correspondingly, in step 110, the first agent can determine the unique identifier of the resource related to the received prompt word according to the received prompt word.
[0051] Here, since each message in the context is encapsulated as a resource, when the first agent receives a prompt word, the first agent selects the unique identifier corresponding to the resource related to the prompt word from each resource in the context according to the received prompt word.
[0052] It should be noted that through the prompt word, the first agent can be informed that when calling the second agent, it is necessary to select the unique identifier corresponding to the resource related to the prompt word in the context. The first agent makes its own judgment and selects the unique identifier corresponding to the relevant resource based on its own understanding, and forms an array with the selected unique identifier of the resource and passes it to the second agent.
[0053] Figure 3 is a flowchart of a context management method for a large language model shown according to some other embodiments. As Figure 3 shown, the first agent receives a prompt word, the first agent analyzes the corresponding task requirements according to the prompt word, the first agent selects the unique identifier of the message related to the task requirements according to the task requirements, the first agent assembles the unique identifier into an array, the first agent passes the array to the second agent, the second agent obtains the target context through the unique identifier included in the array, and the second agent executes the task according to the target context through the large language model to obtain the output result of the large language model.
[0054] It should be noted that in the prompt, there may be instructions for indicating the context screening, so that through this instruction, the first agent is informed that when calling the second agent, relevant messages to the prompt need to be selected as the target context.
[0055] Figure 4 It is a flowchart of context compression shown according to some embodiments. As Figure 4 shown, in some implementable embodiments, the following steps may be included.
[0056] In step 410, when the length of the target context is greater than the upper limit of the token window of the large language model, the messages of the target context are compressed through a preset compression strategy to obtain the compressed target context, and the length of the compressed target context is less than or equal to the token window upper limit.
[0057] Here, the upper limit of the token window of the large language model refers to the maximum value of the context window of the large language model. The context window refers to the amount of text that the large language model can receive when generating or understanding language, or the number of tokens that the large language model can process. In the large language model, a token can be a Chinese character / letter, a word, or a punctuation mark. Therefore, the context window represents the maximum number of characters or words that the large language model can process in one input. The size of the context window directly affects the context information that the large language model can utilize when processing information or the number of tokens when generating a response.
[0058] When the length of the target context is greater than the upper limit of the token window of the large language model, the messages of the target context can be compressed through a preset compression strategy to obtain the compressed target context. Among them, the length of the compressed target context is less than or equal to the token window upper limit.
[0059] It should be understood that by compressing the target context, it can be avoided that the target context fills up the context window of the large language model. Exemplarily, the first agent can compress the messages of the target context through a preset compression strategy.
[0060] In some implementable embodiments, in step 410, starting from the message at the preset position in the target context, one message can be selected in sequence, and at least part of the content in the selected message can be deleted until the length of the target context is less than or equal to the token window upper limit to obtain the compressed target context.
[0061] Here, when compressing the target context, one can start from the message at the preset position in the target context, select a message, and delete at least part of the content in the selected message. After deleting at least part of the content, determine again whether the length of the target context is less than or equal to the upper limit of the token window. If the length of the target context is greater than the upper limit of the token window, continue to select the next message and delete at least part of the content in the selected message until the length of the target context is less than or equal to the upper limit of the token window, then stop selecting the next message to obtain the compressed target context.
[0062] It should be understood that deleting at least part of the content in the selected message can be replacing the at least part of the content with an ellipsis.
[0063] Exemplarily, the message at the preset position in the target context can refer to the 20% - th message in the target context. For example, if there are 100 messages in the target context, the message at the preset position in the target context can refer to the 20 - th message in the target context.
[0064] Exemplarily, at least part of the content can be 50% of the message content in the middle of the message. Of course, at least part of the content can also be determined according to actual needs.
[0065] In some embodiments, for each message included in the target context, determine the message importance level of the message, and then determine the preset position according to the message importance levels of the messages included in the target context.
[0066] Here, the message importance level refers to the impact or value of each message on the large - language model's understanding and generation of accurate and relevant responses. For each message in the target context, the message importance level of the message can be determined.
[0067] As some examples, the message importance level can be determined by the chronological order of the messages. Exemplarily, the earlier the chronological order of the message, the lower the message importance level of the message.
[0068] As some other examples, the message importance level can be determined according to the length of the message. Exemplarily, the length of the message is proportional to the message importance level.
[0069] As some more examples, for each message in the target context, calculate the similarity between the semantics of the message and the overall semantics of the target context, and determine the message importance level corresponding to the message according to the similarity. Among them, the similarity is proportional to the message importance level, that is, the greater the similarity, the greater the message importance level.
[0070] After determining the importance level of each message corresponding to the target context, a preset position is determined according to the importance levels of the messages included in the target context.
[0071] Exemplarily, the position where the message with an importance level less than the preset threshold is located can be used as the preset position to preferentially compress the messages with relatively low importance levels.
[0072] It should be noted that in the above embodiment, in fact, the preset position is dynamically adjusted according to the message importance level to dynamically adjust the starting point of compression.
[0073] Thus, through the above embodiment, when the length of the context is not greater than the upper limit of the token window, sufficient context information can be retained, the retention of a longer conversation history can be achieved, and the performance metrics of the large language model are greatly improved.
[0074] In some implementable embodiments, in step 410, the importance level of the selected message can be determined, then according to the importance level, the compression ratio corresponding to the message is determined, and according to the compression ratio, part of the content in the message is deleted.
[0075] Here, how to determine the importance level of the message corresponding to the message can refer to the relevant description of the above embodiment, which will not be elaborated here. After determining the importance level of the selected message, the compression ratio corresponding to the selected message is determined according to the importance level.
[0076] It should be noted that the compression ratio refers to the proportion of the message content deleted in the message. Among them, the message importance level is inversely proportional to the compression ratio. That is to say, the higher the importance level of the message, the smaller the corresponding compression ratio, and the less message content of this message is deleted.
[0077] Thus, through the importance level of the message corresponding to the message, for each message, the compression ratio corresponding to the message can be dynamically determined, so that the message content with a greater importance level can be retained to a greater extent, and more important context information can be retained in the limited context window.
[0078] Figure 5 is a flowchart of context compression shown according to some other embodiments. As Figure 5 shown, it may include the following steps: S501, determine whether the length of the target context is greater than the upper limit of the token window; S502, in the case where the length of the target context is greater than the upper limit of the token window, select a message starting from the 20%th message in the target context; S503, delete 50% of the message content in the current message; S504, mark the current message; S505, recalculate the length of the target context; S506, determine whether the length of the target context is not greater than the upper limit of the token window; S507, if the length of the target context is greater than the upper limit of the token window, select the next message and return to execute step S503; S508, if the length of the target context is not greater than the upper limit of the token window, end the compression.
[0079] In some implementable embodiments, for a message from which at least part of the content has been deleted, the message can also be marked, and the mark is used to notify the large language model that the corresponding message has been compressed.
[0080] Here, for each message from which at least part of the content has been deleted, the message can be marked. Exemplarily, the message or resource can be marked as truncated.
[0081] By marking the message from which at least part of the content has been deleted, the large language model can be notified that the corresponding message has been compressed.
[0082] For the large language model, when processing the context, if the large language model needs to obtain the complete message content of the marked message, the large language model can read the complete message content through the corresponding unique identifier.
[0083] Thus, by marking the message from which at least part of the content has been deleted, the large language model can be notified that the corresponding message has been compressed, so that the large language model can know which message has been compressed.
[0084] In some implementable embodiments, in step 410, for each message included in the target context, determine the similarity between the message and the overall semantics corresponding to the target context, and then select the target message from the messages included in the target context according to the similarity, and delete the messages in the target context except the target message to obtain the compressed target context.
[0085] Here, the overall semantics corresponding to the target context refers to the overall meaning or theme conveyed by all the messages or texts in the target context. For each message included in the target context, calculate the similarity between the message and the overall semantics corresponding to the target context. Among them, the similarity can be vector similarity, such as Euclidean distance, cosine similarity, etc.
[0086] After calculating the similarities corresponding to each message, the target message is selected from each message included in the target context according to the similarity. Exemplarily, the target message is the message whose similarity is greater than a preset threshold. That is to say, the message in the target context whose similarity is greater than the preset threshold is used as the target message. It should be noted that the size of the preset threshold can be set according to the actual situation.
[0087] After determining the target message, the messages in the target context except the target message can be deleted to obtain the compressed target context. Therefore, in the compressed target context, only the target messages whose similarities are greater than the preset threshold are included.
[0088] Thus, by deleting the messages in the target context whose similarities are less than or equal to the preset threshold, the compressed target context can retain the semantically related messages.
[0089] In some implementable embodiments, in step 410, for each message included in the target context, according to the message type corresponding to the message, a compression strategy for the message is determined, and the message is compressed by the compression strategy to obtain the compressed target context.
[0090] Here, the message type corresponding to the message may refer to the source type of the message, and the sources corresponding to the messages of different message types are different. For example, the message from the user belongs to the first message type, the message from the large language model belongs to the second message type, and the message from the tool call result of the large language model belongs to the third message type.
[0091] When compressing the target context, for each message included in the target context, a compression strategy for the message can be determined according to the message type corresponding to the message, and the message is compressed by the compression strategy. That is to say, for the messages of different message types, different compression strategies can be adopted.
[0092] Continuing the above embodiments, a variety of compression strategies are provided. For example, the first compression strategy: deleting at least part of the content in the selected message. The second compression strategy: determining whether to delete the message according to the similarity between the message and the overall semantics corresponding to the target context.
[0093] For different message types, the target compression strategy can be selected from the first compression strategy and the second compression strategy, and the message is compressed by the target compression strategy.
[0094] Thus, for the messages of different message types in the context, different compression strategies can be used to compress the message.
[0095] In some implementable embodiments, in step 410, each message included in the target context can be input into a trained machine learning model to obtain the degree of relevance between each message and the target context, and the target context can be compressed according to the degree of relevance. Among them, the machine learning model can be a large language model trained to perform message classification.
[0096] Figure 6 It is a schematic structural diagram of a context management device of a large language model shown according to some embodiments. As Figure 6 shown, an embodiment of the present disclosure provides a context management device 600 of a large language model. The context management device 600 of the large language model includes: A determination module 601, configured to determine, through a first proxy, a unique identifier of a message related to the received prompt word according to the received prompt word, where the message is a message in the context of the large language model; A sending module 602, configured to assemble the determined unique identifier into an array through the first proxy and send the array to a second proxy; An acquisition module 603, configured to obtain, through the second proxy, a message corresponding to the unique identifier according to the unique identifier included in the array, and determine a target context of the large language model based on the obtained message; An output module 604, configured to obtain an output result of the large language model according to the target context through the large language model included in the second proxy.
[0097] Optionally, the context management device 600 of the large language model further includes: An encapsulation module, configured to encapsulate each message in the context of the large language model into a resource and assign a corresponding unique identifier to the resource; The determination module 601 is specifically configured to: Determine, through a first proxy, a unique identifier of a resource related to the received prompt word according to the received prompt word.
[0098] Optionally, the encapsulation module is specifically configured to: Encapsulate the message into a resource through a tag of the extensible markup language.
[0099] Optionally, the context management device 600 of the large language model further includes: A compression module, configured to, when the length of the target context is greater than the upper limit of the token window of the large language model, compress the message of the target context through a preset compression strategy to obtain a compressed target context, where the length of the compressed target context is less than or equal to the upper limit of the token window.
[0100] Optionally, the compression module is specifically configured to: Starting from the message at a preset position in the target context, sequentially select one message, and delete at least part of the content in the selected message until the length of the target context is less than or equal to the upper limit of the token window, to obtain the compressed target context.
[0101] Optionally, the compression module is specifically configured to: For each message included in the target context, determine the importance degree of the message; According to the importance degrees of the messages included in the target context, determine the preset position.
[0102] Optionally, the compression module is specifically configured to: Determine the importance degree corresponding to the selected message; According to the importance degree of the message, determine the compression ratio corresponding to the message, where the importance degree of the message is inversely proportional to the compression ratio; According to the compression ratio, delete part of the content in the message.
[0103] Optionally, the context management device 600 of the large language model further includes: A marking module, configured to mark the message for which at least part of the content has been deleted, and the marking is used to notify the large language model that the message corresponding to the marking has been compressed.
[0104] Optionally, the compression module is specifically configured to: For each message included in the target context, determine the similarity between the message and the overall semantics corresponding to the target context; According to the similarity, select target messages from the messages included in the target context, where the target messages are the messages whose similarity is greater than a preset threshold; Delete the messages in the target context except the target messages to obtain the compressed target context.
[0105] Optionally, the compression module is specifically configured to: For each message included in the target context, according to the message type corresponding to the message, determine a compression strategy for the message, and perform compression processing on the message through the compression strategy to obtain the compressed target context.
[0106] The functional logics executed by the respective functional modules in the context management device 600 of the large language model described above have been described in detail in the section on the method, and will not be elaborated here.
[0107] Figure 7 is a schematic structural diagram of an electronic device shown according to some embodiments. Referring below to Figure 7 , which shows a schematic structural diagram of an electronic device (such as a terminal device or a server) 700 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0108] As Figure 7 shown, the electronic device 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0109] Generally, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 shows the electronic device 700 having various devices, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0110] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0111] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0112] In some embodiments, the electronic device can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0113] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.
[0114] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: through a first agent, determine a unique identifier of a message related to the received prompt word according to the prompt word, where the message is a message in the context of a large language model; through the first agent, assemble the determined unique identifier into an array and send the array to a second agent; through the second agent, obtain the message corresponding to the unique identifier according to the unique identifier included in the array, and determine the target context of the large language model based on the obtained message; through the large language model included in the second agent, obtain the output result of the large language model according to the target context.
[0115] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0116] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0117] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the module itself.
[0118] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example, and not limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0119] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0120] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0121] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0122] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.< / resource>
Claims
1. A context management method for a large language model, characterized in that including: Through a first agent, according to the received prompt word, determine the unique identifier of the message related to the prompt word, where the message is a message in the context of a large language model; Through the first agent, assemble the determined unique identifier into an array and send the array to a second agent; Through the second agent, according to the unique identifier included in the array, obtain the message corresponding to the unique identifier, and based on the obtained message, determine the target context of the large language model; Through the large language model included in the second agent, obtain the output result of the large language model according to the target context.
2. The method according to claim 1, wherein The method further includes: For each message in the context of the large language model, encapsulate the message into a resource and assign a corresponding unique identifier to the resource; The step of through the first agent, according to the received prompt word, determine the unique identifier of the message related to the prompt word, includes: Through the first agent, according to the received prompt word, determine the unique identifier of the resource related to the prompt word.
3. The method according to claim 2, wherein The step of encapsulating the message into a resource includes: Encapsulate the message into a resource through tags of Extensible Markup Language.
4. The method according to any one of claims 1 to 3, characterized in that, After the step of based on the obtained message, determine the target context of the large language model, the method further includes: In the case where the length of the target context is greater than the upper limit of the token window of the large language model, perform compression processing on the messages in the target context through a preset compression strategy to obtain a compressed target context, where the length of the compressed target context is less than or equal to the upper limit of the token window.
5. The method according to claim 4, wherein The step of through a preset compression strategy, perform compression processing on the messages in the target context to obtain a compressed target context, includes: Starting from the message at a preset position in the target context, sequentially select a message and delete at least part of the content in the selected message until the length of the target context is less than or equal to the upper limit of the token window to obtain the compressed target context.
6. The method according to claim 5, wherein The preset position is determined through the following steps: For each message included in the target context, determine the message importance level of the message; According to the message importance levels of the messages included in the target context, determine the preset position.
7. The method according to claim 5, wherein The step of deleting at least part of the content in the selected message includes: Determine the message importance level corresponding to the selected message; According to the message importance level, determine the compression ratio corresponding to the message, where the message importance level is inversely proportional to the compression ratio; According to the compression ratio, delete part of the content in the message.
8. The method according to claim 5, characterized in that, The method further includes: For the message from which at least part of the content has been deleted, mark the message, and the mark is used to notify the large language model that the message corresponding to the mark has undergone compression processing.
9. The method according to claim 4, characterized in that, The step of through a preset compression strategy, perform compression processing on the messages in the target context to obtain a compressed target context, includes: For each message included in the target context, determine the similarity between the message and the overall semantics corresponding to the target context; According to the similarity, select a target message from each message included in the target context, where the target message is a message with a similarity greater than a preset threshold; Delete the messages in the target context except the target message to obtain the compressed target context.
10. The method according to claim 4, characterized in that, The compression processing of the messages in the target context through a preset compression strategy to obtain a compressed target context includes: For each message included in the target context, determine a compression strategy for the message according to the message type corresponding to the message, and perform compression processing on the message through the compression strategy to obtain a compressed target context.
11. A context management device for a large language model, characterized in that, Includes: A determination module configured to determine, through a first proxy, a unique identifier of a message related to the received prompt word according to the received prompt word, where the message is a message in the context of a large language model; A sending module configured to assemble the determined unique identifier into an array through the first proxy and send the array to a second proxy; An acquisition module configured to, through the second proxy, acquire a message corresponding to the unique identifier according to the unique identifier included in the array, and determine a target context of the large language model based on the acquired message; An output module configured to obtain an output result of the large language model through the large language model included in the second proxy according to the target context.
12. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processing device, the steps of the method according to any one of claims 1-10 are implemented.
13. An electronic device, characterized in that, Includes: A storage device having a computer program stored thereon; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-10.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1-10 are implemented.
Citation Information
Patent Citations
Large language model cue word injection attack detection method and device based on context learning
CN118734314A
Prompt word optimization method in combination with expert evaluation rule and large language model
CN118886427A
Large language model man-machine interaction method and system based on memory loop network enhancement
CN119002694A
Large language model RAG optimization method based on tree neighbor context
CN119293195A
Web component recommendation method, system, equipment and medium
CN119739383A
Cited By
Big language model dynamic dialogue history compression method and system based on double verification
CN121144461A
Dual-verification-based large language model dynamic dialogue history compression method and system
CN121144461B