A long text processing method and related device

By introducing external memory modules and multi-level memory structures into large language models, the problems of high computational complexity and context loss in long text processing are solved, efficiency and accuracy are improved, and the context understanding and adaptability of the model are enhanced.

CN119204234BActive Publication Date: 2025-07-11ZHONGHAO XINYING (HANGZHOU) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411747634.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-07-11
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

Existing large language models have high computational complexity and memory requirements when processing long text, resulting in inefficiency and loss of context information, which makes it difficult to meet user needs.

Method used

An external memory module and multi-level memory structure are introduced, including short-term memory areas and long-term memory areas, which store important context information in the current session and repeated context information in multiple sessions, and input them into a large language model for natural language processing.

Benefits of technology

It significantly improves the efficiency and accuracy of long text processing, enhances the model's context memory ability and flexibility and adaptability, and can better understand and generate coherent text content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119204234B_ABST
    Figure CN119204234B_ABST
Patent Text Reader

Abstract

This application belongs to the field of artificial intelligence, and particularly relates to a long text processing method and related devices, including: for the long text data to be processed in the current session, extracting the context information corresponding to each text segment from the long text data; storing each context information in different storage areas of an external memory module respectively; the external memory module includes a short-term memory area and a long-term memory area; the short-term memory area is used to store the first context information whose importance reaches a preset condition in the current session; the long-term memory area is used to store the second context information that appears repeatedly in multiple sessions; the multiple sessions include the current session and / or historical sessions; inputting the first context information and the second context information into a large language model, and realizing natural language processing of the long text data through the large language model. This method can improve the efficiency and accuracy of long text processing, and enhance the context memory ability and flexible adaptability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence, and particularly relates to a long text processing method and related devices. Background Art

[0002] Currently, large language models have made remarkable progress in the field of natural language processing (NLP) and are widely used in tasks such as machine translation, text generation, and dialogue systems.

[0003] With the explosive growth of Internet content, processing and understanding long texts have become increasingly important. Users hope to be able to process extremely long texts such as long articles, complete books, and cross-session conversations. In related technologies, large language models mainly rely on the self-attention mechanism to capture dependencies in the input sequence. However, as the input length increases, the computational complexity and memory requirements of the model grow exponentially, leading to technical problems such as low long text processing efficiency and loss of context information. It can be seen that existing large language models have obvious limitations in processing long texts and often struggle to meet the above-mentioned needs of users.

[0004] Therefore, there is an urgent need to design a brand-new technical solution to overcome at least one of the above technical problems. Summary of the Invention

[0005] This application provides a long text processing method and related devices to improve the efficiency and accuracy of long text processing and enhance the context memory ability and flexible adaptability of the model.

[0006] In a first aspect, this application provides a long text processing method, including:

[0007] For the long text data to be processed in the current session, extract the context information corresponding to each text segment from the long text data;

[0008] Store each context information in different storage areas of an external memory module respectively; the external memory module includes a short-term memory area and a long-term memory area; the short-term memory area is used to store first context information that reaches a preset condition in the current session; the long-term memory area is used to store second context information that appears repeatedly in multiple sessions; multiple sessions include the current session and / or historical sessions;

[0009] Input the first context information and the second context information into a large language model, and implement natural language processing of the long text data through the large language model.

[0010] In a second aspect, an embodiment of this application provides a long text processing device, including at least the following units:

[0011] An extraction unit, configured to extract context information corresponding to each text segment from the long text data to be processed in the current session from the long text data;

[0012] A storage unit, configured to store each context information into different storage areas of an external memory module respectively; the external memory module includes a short-term memory area and a long-term memory area; the short-term memory area is used to store first context information whose importance reaches a preset condition in the current session; the long-term memory area is used to store second context information that appears repeatedly in multiple sessions; multiple sessions include the current session and / or historical sessions;

[0013] An execution unit, configured to input the first context information and the second context information into a large language model, and implement natural language processing of the long text data through the large language model.

[0014] In a third aspect, an embodiment of the present application provides a chip for implementing the long text processing method described in the first aspect.

[0015] In a fourth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory, where the processor executes the computer program to implement the long text processing method described in the first aspect.

[0016] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed, the long text processing method described in the first aspect is implemented.

[0017] In the embodiment of the present application, first, for the long text data to be processed in the current session, context information corresponding to each text segment is extracted from the long text data. Furthermore, each context information is stored into different storage areas of an external memory module respectively; the external memory module includes a short-term memory area and a long-term memory area; the short-term memory area is used to store first context information whose importance reaches a preset condition in the current session; the long-term memory area is used to store second context information that appears repeatedly in multiple sessions; multiple sessions include the current session and / or historical sessions. Finally, the first context information and the second context information are input into a large language model, and natural language processing of the long text data is implemented through the large language model. In the embodiment of the present application, by introducing an external memory module and a multi-level memory structure, the efficiency and accuracy of long text processing are significantly improved, and at the same time, the context memory ability and flexible adaptability of the model are enhanced, providing more powerful and comprehensive technical support for natural language processing tasks. Description of the Drawings

[0018] The accompanying drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0019] Figure 1 is a flowchart of a long text processing method according to an embodiment of the present application;

[0020] Figure 2 is a structural block diagram of a long text processing device according to an embodiment of the present application;

[0021] Figure 3 is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0022] The embodiments of the present application provide a long text processing method and related devices. First, for the long text data to be processed in the current session, the context information corresponding to each text segment is extracted from the long text data. Furthermore, each context information is respectively stored in different storage areas of an external memory module; the external memory module includes a short-term memory area and a long-term memory area; the short-term memory area is used to store the first context information whose importance reaches a preset condition in the current session; the long-term memory area is used to store the second context information that appears repeatedly in multiple sessions; the multiple sessions include the current session and / or historical sessions. Finally, the first context information and the second context information are input into a large language model, and natural language processing of the long text data is realized through the large language model. In the embodiments of the present application, by introducing an external memory module and a multi-level memory structure, the efficiency and accuracy of long text processing are significantly improved, and at the same time, the context memory ability and flexible adaptability of the model are enhanced, providing more powerful and comprehensive technical support for natural language processing tasks.

[0023] Specifically, in the embodiments of the present application, by introducing an external memory module and a multi-level memory structure, the effect of long text processing is significantly optimized. When processing long text data, first, the context information of each text segment is extracted from the long text and stored in different storage areas of the external memory module respectively. The external memory module includes a short-term memory area and a long-term memory area. The short-term memory area stores the first context information that reaches a preset condition in the current session, while the long-term memory area stores the second context information that appears repeatedly in multiple sessions. This multi-level memory structure effectively solves the common context loss problem in traditional models, making the generated content more coherent and contextually logical. Secondly, by storing the context information in the external memory module, the computational overhead of the model when processing long text is reduced. In particular, the separation of the short-term memory area and the long-term memory area enables the model to quickly access and utilize relevant context information, thereby improving the inference speed. This design avoids the heavy task of recalculating or storing a large amount of context information during each inference process, significantly enhancing the overall computational efficiency. Thirdly, the dynamic adjustment and optimization mechanism of the external memory module can ensure that the model can obtain more accurate and relevant context information when generating and understanding long text. By inputting the first context information and the second context information into the large language model, the model can make full use of these context information for natural language processing, thereby generating more accurate and relevant results. This mechanism not only improves the accuracy of the model but also enhances the adaptability and generalization ability of the model. Finally, the design of the multi-level memory structure is not only applicable to long text processing but also can play a role in other tasks that require context memory. The separation of short-term memory and long-term memory enables the model to flexibly adapt to different types of task requirements. Whether it is the key information in the current session or the common information in historical sessions, they can be effectively managed and utilized in different storage areas.

[0024] The long text processing solution provided by the embodiments of the present application can also be executed by an electronic device, which can be a server, a server cluster, or a cloud server. The electronic device can also be a terminal device such as a mobile phone, a computer, a tablet computer, a wearable device, or a dedicated device (such as a dedicated terminal device with a long text processing system, etc.). These electronic devices can also be equipped with the chips introduced in the above embodiments. Alternatively, these electronic devices can also install a service program for executing the long text processing solution.

[0025] The long text processing method and related devices provided by the embodiments of the present application will be described below with reference to the accompanying drawings. Figure 1 A long text processing method provided by an embodiment of the present application is as Figure 1 shown, and the method includes:

[0026] S101. For the long text data to be processed in the current session, extract the context information corresponding to each text segment from the long text data;

[0027] S102. Store each context information in different storage areas of the external memory module respectively;

[0028] S103. Input the first context information and the second context information into the large language model, and implement natural language processing on the long text data through the large language model.

[0029] The above three steps provide significant technical optimizations and beneficial effects for long text data processing by introducing an external memory module and a multi-level memory structure. Specifically, first, by extracting the context information of each text segment, the model can more deeply understand the key content in the long text, avoiding the problem of context loss that traditional models are prone to when processing long texts. Decomposing the long text into multiple text segments and extracting their context information helps the model better capture the specific meaning and logical relationship of each segment, improving the refinement degree of processing. When generating text, it can more accurately refer to the context information of each segment, thereby generating more coherent and logical text content.

[0030] Furthermore, store each context information in different storage areas of the external memory module respectively. In this way, through the external memory module (including the short-term memory area and the long-term memory area), context information with different importance and timeliness can be reasonably stored and managed, avoiding information redundancy and chaos. By storing the important context information in the current session in the short-term memory area, the model can quickly access these key information, improving the reasoning and processing speed. The long-term memory area is used to store the context information that appears repeatedly in multiple sessions, which helps the model accumulate and utilize long-term knowledge, enhancing the model's knowledge base and adaptability.

[0031] Finally, input the first context information and the second context information into the large language model for natural language processing. By integrating the context information of the current session (short-term memory) and the historical session (long-term memory), the large language model can obtain more comprehensive and coherent context support when generating and understanding text, and the generated content is more natural and logical. Combining the context information of short-term and long-term memories, the model can more accurately understand the content of the long text and generate outputs highly relevant to the context, significantly improving the accuracy and relevance of the model. The design of the multi-level memory structure makes this technology not only applicable to long text processing but also able to play a role in other tasks that require context memory, enhancing the versatility and adaptability of the model.

[0032] In summary, these three steps significantly improve the context understanding ability, computational efficiency, and result accuracy of long text processing by introducing an external memory module and a multi-level memory structure. At the same time, the flexible adaptability of the model is enhanced, providing more comprehensive and efficient technical support for natural language processing tasks.

[0033] In the embodiments of the present application, the external memory module includes a short-term memory area and a long-term memory area. The division of these two areas enables the model to flexibly handle different types of context information, improving the overall processing efficiency and accuracy.

[0034] Among them, the short-term memory area is used to store the first context information that reaches a preset condition in the current session. That is to say, the information stored in the short-term memory area is important content in the recent period and can be frequently accessed in the current session. The short-term memory area is mainly used to store context information that frequently appears in the current session or is crucial for understanding the current content. These information may be mentioned or referenced multiple times in the current session, so it is necessary to be able to quickly access them in a short time. To avoid storing too much unimportant information, the short-term memory area only stores context information that reaches the preset condition. This screening mechanism ensures that the stored information is of high value, thus avoiding information redundancy and confusion. Since the short-term memory area stores the key information in the current session, the model can frequently access this information when processing the current session, thereby improving the inference speed and processing efficiency.

[0035] The long-term memory area is used to store the second context information that appears repeatedly in multiple sessions. The long-term memory area is mainly used to store context information that appears repeatedly in multiple sessions. These information may have appeared multiple times in historical sessions or are common background knowledge or themes across sessions. By storing these repeatedly appearing context information, the long-term memory area helps the model accumulate long-term knowledge. These knowledge can be referenced as background information in the current session to enhance the model's understanding and generation of the current content. The stored content in the long-term memory area can include context information in the current session and historical sessions, enabling the model to maintain context coherence and consistency when processing cross-session content.

[0036] It can be understood that multiple sessions include the current session and / or historical sessions. The current session refers to the current ongoing conversation or the text content being processed. The short-term memory area mainly stores the key information in the current session, and the long-term memory area may also include context information that appears repeatedly in the current session. Historical sessions refer to previous conversations or text content that has been processed. The context information stored in the long-term memory area may come from historical sessions, and these information can be referenced and utilized in the current session.

[0037] In summary, through the division of the short-term memory area and the long-term memory area, the external memory module realizes the efficient storage and management of context information. The short-term memory area ensures that key information in the current session can be quickly accessed and utilized, improving the real-time performance and efficiency of processing. The long-term memory area enhances the long-term knowledge accumulation and context coherence of the model by storing repeated information across sessions, making the model more accurate and efficient in processing complex and multi-session content. Thus, through the reasonable division of the short-term memory area and the long-term memory area, the context understanding ability and processing efficiency of the model in processing long text data are significantly improved, providing more comprehensive and powerful technical support for natural language processing tasks.

[0038] As an optional embodiment, in 102, storing each context information into different storage areas of the external memory module can be implemented as the following steps:

[0039] 201. Obtain the importance evaluation values corresponding to each context information in the current session through the self-attention mechanism; the importance evaluation value is used to indicate the correlation between the context information and the long text data; the higher the correlation between the context information and the long text data, the higher the importance evaluation value corresponding to the context information;

[0040] 202. Select the first context information in the current session whose importance evaluation value meets the preset conditions, and store the first context information in the short-term memory area;

[0041] 203. Match each context information in the current session with the second context information stored in the long-term memory area, and use the context information that appears repeatedly in the current session and the historical session as the newly added second context information to update the long-term memory area.

[0042] Specifically, in 201, through the self-attention mechanism, importance evaluation values corresponding to each context information in the current session are obtained. Among them, the self-attention mechanism is a technique used to calculate the correlation between elements in the input sequence. In this step, the self-attention mechanism is used to evaluate the correlation between each context information in the current session and the long text data. Specifically, through the self-attention mechanism, the model can calculate an importance evaluation value for each context information, which reflects the correlation between the context information and the long text data. The higher the correlation, the higher the importance evaluation value. In this way, the self-attention mechanism can accurately calculate the correlation between each context information and the long text data, providing an accurate basis for importance evaluation for subsequent steps. Through the importance evaluation value, the model can quickly screen out the context information highly relevant to the current long text data, reducing the interference of irrelevant information and improving the efficiency of storage and processing.

[0043] In 202, the first context information whose importance evaluation value meets the preset conditions in the current session is selected, and the first context information is stored in the short-term memory area. Specifically, according to the preset conditions (such as the threshold of the importance evaluation value), the first context information that meets the conditions is screened out. Thus, the context information whose importance evaluation value reaches or exceeds the preset threshold is stored in the short-term memory area for quick access and utilization of these key information in the current session. In this way, through the screening of the preset conditions, the short-term memory area only stores the context information that is crucial for understanding the current long text, avoiding information redundancy and chaos. The information in the short-term memory area can be quickly accessed, improving the reasoning and processing speed of the model in the current session.

[0044] It is understandable that the setting of preset conditions is the key to achieving efficient storage and management of the short-term memory area. The preset conditions can be set according to various factors to ensure that the context information crucial for understanding the current long text is stored. For example, multiple conditions such as importance evaluation value, access frequency, timeliness, semantic relevance, and information entropy can be combined according to actual needs to set multi-dimensional preset conditions. For instance, only the context information that simultaneously meets the thresholds of importance evaluation value, access frequency, and semantic similarity will be stored in the short-term memory area. In this way, according to the progress of the conversation and the changes in context information, the thresholds of each preset condition are dynamically adjusted to ensure that the stored information is always up-to-date and of high value. Here, the stored information has a high access frequency and timeliness, improving the access speed and processing efficiency of the short-term memory area. Through the combination of multiple conditions and dynamic adjustment, the performance and adaptability of the model are further enhanced. In summary, through the reasonable setting of various preset conditions, the efficient storage and management of context information can be achieved, providing more accurate and efficient technical support for the processing of long text data.

[0045] Exemplarily, the importance evaluation value of each context information is calculated using the Self-Attention Mechanism. The Self-Attention Mechanism evaluates the correlation between context information and the current long text data by calculating the attention scores between them. For each context information, the attention scores with all positions in the current long text data are calculated, and the attention scores are normalized through the Softmax function to obtain the importance evaluation value of each context information. A threshold for the importance evaluation value is preset. For example, the threshold is set to 0.8, indicating that only the context information whose importance evaluation value reaches or exceeds 0.8 will be stored in the short-term memory area. The Self-Attention Mechanism can accurately evaluate the correlation between context information and the current long text data, ensuring that the stored information is of high value, avoiding storing unimportant information, and improving the storage and processing efficiency. The threshold can be adjusted according to the actual application scenario and requirements to optimize the storage strategy and further enhance the performance and efficiency of the model. For example, when dealing with highly complex conversations, the threshold can be appropriately lowered to store more relevant information; when dealing with simple conversations, the threshold can be appropriately raised to reduce the storage of redundant information.

[0046] Exemplarily, in the current session, the number of times each context information is accessed is recorded in real time. The access frequency represents the frequency at which a certain context information is mentioned or referenced in the current session. Initialize the access frequency of each context information to 0, and update its access frequency whenever the context information is mentioned or referenced. A pre-set access frequency threshold is set. For example, set the threshold to 5, indicating that only those context information with an access frequency reaching or exceeding 5 times will be stored in the short-term memory area. Context information with a high access frequency is usually highly relevant to the content of the current session, ensuring that only important context information frequently mentioned in the current session is stored, and avoiding storing unimportant information. The access frequency threshold can be dynamically adjusted according to the progress of the session and changes in context information to ensure that the stored information is always up-to-date and of high value. For example, a lower threshold can be set at the beginning of the session to quickly accumulate important information; the threshold can be appropriately increased in the later stage of the session to screen out truly important information.

[0047] Exemplarily, calculate the information entropy value of each context information. The information entropy value reflects the amount of information contained in the context information. The higher the information entropy value, the greater the amount of information contained in the context information. Calculate the information entropy value for each context information. The information entropy value is calculated by the formula: . Among them, represents the probability of each possible state appearing in the context information, represents taking the logarithm of the probability , represents the sum of the self-information values of all probability terms from i = 0 to n for the probability . The uncertainty of the entire data set is quantified by calculating the expected value of the self-information of all possible results. The greater the information entropy, the higher the uncertainty of the data set; conversely, the smaller the information entropy, the lower the uncertainty of the data set. Calculate the information entropy value for all context information and sort them. Further optionally, a pre-set information entropy threshold is set. For example, set the threshold to 0.7, indicating that only those context information with an information entropy value reaching or exceeding 0.7 will be stored in the short-term memory area. In this way, it is ensured that the stored information has a high amount of information, and redundant and low-value information is avoided from being stored. Context information with a high information entropy value usually contains more key information, which helps to improve the model's understanding and reasoning ability.

[0048] In practical applications, further, the information entropy threshold can be adjusted according to the actual application scenario and requirements to optimize the storage strategy and further improve the performance and efficiency of the model. For example, when dealing with highly complex sessions, the threshold can be appropriately lowered to store more information-rich information; when dealing with simple sessions, the threshold can be appropriately increased to reduce the storage of redundant information.

[0049] Exemplarily, a time window is set to store only the context information that appears within the time window. For example, only the context information that appears within the past 10 minutes is stored. A time decay function is used to evaluate the timeliness of the context information. The context information closer to the current time has higher timeliness and is preferentially stored. Thus, it is ensured that the stored information has high timeliness and obsolete information is avoided. The time window can be adjusted according to the duration of the session to ensure that the stored information is always up-to-date and relevant.

[0050] Exemplarily, a semantic similarity algorithm (such as cosine similarity) is used to calculate the semantic relevance between the context information and the current long text data. A semantic similarity threshold is preset in advance. Only the context information whose semantic similarity reaches or exceeds the threshold will be stored in the short-term memory area. Thus, it is ensured that the stored information has high semantic relevance to the current long text data and the understanding ability is enhanced. The threshold can be adjusted according to the actual application scenario and requirements to optimize the storage strategy and further improve the performance and efficiency of the model.

[0051] In 203, each context information in the current session is matched with the second context information stored in the long-term memory area, and the context information that repeatedly appears in the current session and the historical sessions is used as the newly added second context information and updated to the long-term memory area. Specifically, each context information in the current session is matched with the second context information stored in the long-term memory area to find the context information that repeatedly appears in multiple sessions. Furthermore, the matched and repeatedly appearing context information is used as the newly added second context information and updated to the long-term memory area. This information may come from the current session and / or historical sessions. Thus, through matching and updating, the long-term memory area can accumulate the repeated context information across sessions, enhancing the long-term knowledge base and adaptability of the model. The updated long-term memory area can provide more comprehensive background information for the current session, enhancing the model's understanding and generation ability of the current content and maintaining the coherence and consistency of the context. The information in the long-term memory area is not only applicable to the current session but can also be referenced and utilized in future sessions, improving the generalization ability and application scope of the model.

[0052] Similar to step 202, in combination with the storage requirements of the long-term memory area for context information in the actual application scenario, a method similar to that exemplified in step 202 can be adopted to implement step 203, which will not be elaborated here for the time being.

[0053] Through the implementation of steps 201 to 203, the efficient storage and management capabilities of context information are significantly improved. By means of the self-attention mechanism, the importance of context information is accurately evaluated, providing an accurate evaluation basis for subsequent steps. By screening and storing the key information in the current session according to preset conditions, the access speed and processing efficiency of the short-term memory area are improved. By matching and updating the long-term memory area, the long-term knowledge accumulation and context coherence of the model are enhanced, improving the generalization ability and application scope of the model. In summary, optional embodiment 102 significantly optimizes the storage and management of context information through the self-attention mechanism and a reasonable storage strategy, providing more efficient and accurate technical support for the processing of long text data.

[0054] In another optional embodiment 102, each piece of context information is separately stored in different storage areas of the external memory module, and the access frequency corresponding to each piece of context information in the current session can be further obtained. Furthermore, the first context information with an access frequency higher than the set access frequency threshold in the current session is selected, and the first context information is stored in the short-term memory area.

[0055] For example, in the current session, the number of times each piece of context information is accessed is recorded. The access frequency represents the frequency at which a piece of context information is mentioned or referred to in the current session. By monitoring the access behavior in the current session in real time, the model can dynamically calculate the access frequency of each piece of context information. A preset access frequency threshold is used as a condition for screening context information. Only those context information with an access frequency higher than or equal to this threshold will be considered important. According to the access frequency threshold, the context information with a higher access frequency in the current session is screened out and marked as the first context information. The screened first context information is stored in the short-term memory area. These information may be frequently mentioned or referred to in the current session, so they need to be quickly accessible in a short time.

[0056] It can be understood that the context information in the short-term memory area has a high access frequency, which means that these information have high importance in the current session. The model can quickly access these key information, improving the inference and processing speed. The information in the short-term memory area can be dynamically updated according to the real-time access situation of the current session to ensure that the stored information is always up-to-date and of high value. The context information that is frequently accessed is usually highly relevant to the content of the current session. Storing this information in the short-term memory area helps the model maintain context coherence and consistency when processing the current session. When generating text, the model can more accurately refer to the context information in the short-term memory area, thus generating more coherent and logical text content.

[0057] Further optionally, different access methods corresponding to different levels can also be set based on the access frequency to improve the response speed of access requests.

[0058] As an alternative embodiment, in 103, inputting the first context information and the second context information into the large language model can be implemented as the following steps:

[0059] 301. Obtain the target text information to be processed currently from the current session;

[0060] 302. Generate a corresponding retrieval request based on the target text information; at least one or more of the text content information, text type, and text theme in the target text information are included in the retrieval request;

[0061] 303. Retrieve the first context information and the second context information that match the target text information from the external memory module based on the retrieval request;

[0062] 304. Perform a fusion process on the retrieved first context information and the second context information, and input the result of the fusion process into the large language model.

[0063] Exemplarily, in this alternative embodiment, the first context information and the second context information are input into the large language model through multiple steps to achieve efficient processing of the target text information. Assume an application scenario: a customer service robot provides intelligent answers to questions raised by users. Assume the target text information is the question raised by the user, "How can I upgrade my account membership?"

[0064] In 301, the customer service robot extracts the question raised by the user, "How can I upgrade my account membership?" from the current user's session record. Here, the target text information = "How can I upgrade my account membership?"

[0065] In 302, a retrieval request is generated according to the target text information. This request contains the characteristics of the target text, such as text content, type, and theme. For example, the text content information is "How can I upgrade my account membership?", the text type is a question, and the text theme is account upgrade. Thus, a corresponding retrieval request is generated, i.e., {"text": "How can I upgrade my account membership?", "type": "question", "topic": "account upgrade"}.

[0066] In 303, according to the generated retrieval request, context information related to the target text information is retrieved from an external memory module (such as a knowledge base or database). For example, the first context information is the step description for account upgrade. The second context information is the part about account upgrade in the Frequently Asked Questions (FAQ). Then, the first context information = "You can upgrade your account membership through the following steps: Log in to the account -> Select the membership package -> Complete the payment"; the second context information = "Having problems when upgrading the account? Please refer to Article 4 in the Frequently Asked Questions".

[0067] In 304, the retrieved first and second context information are fused to form a more comprehensive and detailed information set. For example, the fusion processing result = "You can upgrade your account membership through the following steps: Log in to the account -> Select the membership package -> Complete the payment. If you have problems during the upgrade process, please refer to Article 4 in the Frequently Asked Questions". In this way, the fusion processing result is input into the large language model, and the model will generate an intelligent answer. Here, the answer generated by the large language model can be, "To upgrade your account membership, please log in to your account, select your preferred membership package, and then complete the payment. If you encounter any problems during the process, please feel free to refer to Article 4 in our Frequently Asked Questions".

[0068] Through the above example steps, the customer service robot can extract relevant information from the external memory module and generate an accurate and comprehensive answer by combining this information. This method not only improves the accuracy of the answer but also enhances the user experience.

[0069] As an optional embodiment, in 303, retrieving the first context information and the second context information that match the target text information from the external memory module based on the retrieval request can be implemented as the following steps:

[0070] 401. Extract the target text fragment from the target text information;

[0071] 402. Obtain the cosine similarity between the first context information and the second context information in the external memory module and the target text fragment respectively;

[0072] 403. Select the first context information and the second context information whose cosine similarity meets the preset conditions as the first context information and the second context information that match the target text fragment.

[0073] Among them, cosine similarity is a measurement method used to measure the angle between two vectors. In text processing and information retrieval, it can be used to evaluate the similarity between two text vectors. The value range of cosine similarity is [-1, 1]. The larger the value, the more similar the two vectors are, and the smaller the value, the less similar the two vectors are. In text processing, the bag-of-words model or word embedding is used to convert text into vector form, and then the cosine similarity between these vectors is calculated.

[0074] Exemplarily, an external memory module is set up, which stores context information related to account upgrade. Suppose cosine similarity is used to retrieve context information related to the user's question "How can I upgrade my account membership?". In 401, key fragments are extracted from the target text information "How can I upgrade my account membership?". Suppose the target text fragment is "upgrade my account membership".

[0075] In 402, the first context information and the second context information are obtained from the external memory module. For example, the first context information is "You can upgrade your account membership through the following steps: log in to the account -> select a membership package -> complete the payment.". The second context information is "Having problems when upgrading the account? Please refer to item 4 in the Frequently Asked Questions.". Furthermore, the target text fragment "upgrade my account membership" and the context information are converted into vector form. The cosine similarity calculation formula is used to calculate the cosine similarity between the target text fragment and each context information. The target text fragment vector is [0.1, 0.2, 0.3, 0.4]. The first context information vector is [0.15, 0.25, 0.35, 0.45]. The second context information vector is [0.2, 0.3, 0.4, 0.5]. The calculation result is that the cosine similarity between the target text fragment and the first context information is 0.98, and the cosine similarity between the target text fragment and the second context information is 0.95.

[0076] In 403, a threshold of cosine similarity is set, for example, 0.9. The context information with a cosine similarity greater than or equal to 0.9 is selected. The cosine similarity of the first context information is 0.98 (meeting the condition), and the cosine similarity of the second context information is 0.95 (meeting the condition). The first context information and the second context information are selected as the context information matching the target text fragment.

[0077] In this way, by using the cosine similarity algorithm, it is possible to efficiently retrieve context information highly relevant to the target text information from the external memory module. This method not only improves the accuracy of retrieval but also reduces the interference of irrelevant information, thereby enhancing the overall system performance and user experience. The cosine similarity can accurately measure the similarity between texts, ensuring that the retrieved context information is highly relevant to the target text. The calculation of the cosine similarity is relatively simple and suitable for large-scale text retrieval tasks. The threshold of the cosine similarity can be adjusted according to actual needs to optimize the retrieval strategy. If there is a large amount of noise in the text vector, it may affect the accuracy of the cosine similarity. In a high-dimensional space, the calculation of the cosine similarity may be affected by the sparsity problem.

[0078] In summary, the cosine similarity algorithm has a wide range of applications in text processing and information retrieval. It can efficiently measure the similarity between texts and filter out the most relevant context information according to the set conditions. By reasonably setting the threshold and optimizing the vector representation method, the retrieval effect can be further improved.

[0079] Further optionally, assume that the first context information and the second context information in the external memory module are respectively associated with corresponding dynamic indexes according to their respective text content characteristics and / or text format types. Based on this, before 403, it is also possible to identify the text attribute features of the target text segment. The text attribute features are used to represent the text content characteristics and / or text format types of the target text segment. Furthermore, based on the text content characteristics and / or text format types, construct the text retrieval preconditions for the target text segment. Then, based on the text retrieval preconditions, determine the target dynamic index in the external memory module that matches the target text segment, and obtain the first context information and the second context information associated with the target dynamic index as the first context information and the second context information to be filtered.

[0080] In this embodiment, the context information in the external memory module is associated with corresponding dynamic indexes according to their respective text content characteristics and / or text format types. By identifying the text attribute features of the target text segment, the text retrieval preconditions for the target text segment are constructed, so as to more accurately determine the target dynamic index in the external memory module that matches the target text segment. This step can significantly improve the efficiency and accuracy of retrieval.

[0081] Suppose there is an external memory module that stores a large amount of context information related to account upgrades. The information in the external memory module is associated with different dynamic indexes according to content characteristics and format types. The target text information is the user's question "How can I upgrade my account membership?", and the keyword is "upgrade account membership". Based on this, the keyword vocabulary and theme in the target text segment can be analyzed to determine the type of the target text segment (such as question, statement, command, etc.). The text content characteristics are that the keywords = "upgrade", "account", "membership"; the text format type is a question.

[0082] According to the text attribute characteristics of the target text segment, a retrieval condition is constructed to filter relevant information from the external memory module. The keywords include "upgrade", "account", "membership", and the text format is a question. According to the preconditions for text retrieval, the dynamic index in the external memory module that matches the target text segment is determined. Suppose dynamic index 1: keywords = "upgrade", "account", "membership", text format = question; dynamic index 2: keywords = "membership", "account", text format = statement; dynamic index 3: keywords = "upgrade", "account", text format = command. Based on the above assumptions and implementation principles, it can be determined that the dynamic index that the target text segment matches is dynamic index 1.

[0083] Furthermore, relevant context information is obtained from the determined target dynamic index. The first context information associated with the target dynamic index 1 is, "You can upgrade your account membership through the following steps: log in to the account -> select a membership package -> complete the payment." The second context information is, "Having problems when upgrading the account? Please refer to item 4 in the Frequently Asked Questions."

[0084] In this way, through the dynamic index, relevant context information can be quickly located according to the text attribute characteristics of the target text segment, reducing unnecessary full-text retrieval and improving the retrieval speed. Identifying the text attribute characteristics of the target text segment and constructing preconditions for retrieval can reduce interference from irrelevant information and ensure that the retrieved context information is highly relevant to the target text. More accurate and efficient retrieval of context information can improve user satisfaction, especially in application scenarios such as customer service robots, where the information needed by users can be provided in a timely manner. The dynamic index can be dynamically adjusted according to the characteristics and format types of the text content to adapt to different types of text information, improving the flexibility and adaptability of the system. In short, by identifying the text attribute characteristics of the target text segment and constructing preconditions for retrieval, the dynamic index in the external memory module that matches the target text segment can be determined more precisely, thereby improving the efficiency and accuracy of context information retrieval. This method not only optimizes the performance of the system but also significantly enhances the user experience.

[0085] Furthermore, in 304, the retrieved first context information and second context information are fused. Specifically, the retrieved first context information and second context information are sorted in descending order of cosine similarity to select the first context information and second context information whose similarity is at a preset position or reaches a preset similarity threshold. Furthermore, the selected first context information and second context information are subjected to multi-level memory fusion processing to obtain the fusion result.

[0086] In step 304, the retrieved first context information and second context information are sorted in descending order of cosine similarity, and the context information whose similarity is at a preset position or reaches a preset similarity threshold is selected. Then, the selected context information is subjected to multi-level memory fusion processing to obtain the final fusion result.

[0087] For example, the hybrid splicing method is adopted. Specifically, the retrieved context information is sorted according to the cosine similarity. The context information whose similarity reaches the preset threshold or is at the preset position is selected. The screened context information is textually spliced to form a complete answer.

[0088] Exemplarily, the first context information (similarity 0.98): You can upgrade your account membership through the following steps: Log in to the account -> Select the membership package -> Complete the payment. The second context information (similarity 0.95): Having problems when upgrading the account? Please refer to Article 4 in the FAQ. The third context information (similarity 0.88): After the account upgrade is successful, you can enjoy more membership benefits. The preset similarity threshold is 0.9, and the first context information and the second context information are screened out. The screened context information is spliced to form a complete answer. Thus, the final fusion result is, "You can upgrade your account membership through the following steps: Log in to the account -> Select the membership package -> Complete the payment. If you have problems during the process, please refer to Article 4 in the FAQ."

[0089] Alternatively, the information superposition method can also be adopted. Specifically, the retrieved context information is sorted according to the cosine similarity. The context information whose similarity reaches the preset threshold or is at the preset position is selected. The information contents of the screened context information are superposed to form a more detailed and comprehensive answer.

[0090] Exemplarily, the first context information (similarity 0.98): You can upgrade your account membership by following these steps: Log in to the account -> Select the membership package -> Complete the payment. The second context information (similarity 0.95): Having problems during account upgrade? Please refer to Article 4 in the FAQ. The third context information (similarity 0.88): After the account upgrade is successful, you can enjoy more membership benefits. The preset similarity threshold is 0.9. Filter out the first context information and the second context information. Overlay the information content of the filtered context information to form a detailed answer. Thus, the final fusion result is, "You can upgrade your account membership by following these steps: Log in to the account -> Select the membership package -> Complete the payment. If you encounter problems during the process, please refer to Article 4 in the FAQ. After the account upgrade is successful, you can enjoy more membership benefits."

[0091] Alternatively, the information synthesis method can also be adopted. Specifically, sort the retrieved context information according to the cosine similarity, select the context information whose similarity reaches the preset threshold or is in the preset position, and use natural language processing techniques (such as sentence generation models) to synthesize a natural and fluent answer.

[0092] Exemplarily, the first context information (similarity 0.98): You can upgrade your account membership by following these steps: Log in to the account -> Select the membership package -> Complete the payment. The second context information (similarity 0.95): Having problems during account upgrade? Please refer to Article 4 in the FAQ. The third context information (similarity 0.88): After the account upgrade is successful, you can enjoy more membership benefits. The preset similarity threshold is 0.9. Filter out the first context information and the second context information. Use the sentence generation model to synthesize the filtered context information into a natural and fluent answer. Thus, the final fusion result is, "To upgrade your account membership, you need to log in to the account, select a suitable membership package, and complete the payment. If you encounter any problems during the upgrade process, you can refer to Article 4 in our FAQ."

[0093] Alternatively, an information screening method can also be adopted. Specifically, the retrieved context information is sorted according to the cosine similarity, and the context information with a similarity reaching a preset threshold or being in a preset position is selected. Then, based on the content of the context information and the user's intention, the most relevant information is further screened out. Finally, the screened-out information is concatenated or synthesized to form a precise answer. For example, the first context information (similarity 0.98): You can upgrade your account membership through the following steps: Log in to the account -> Select the membership package -> Complete the payment. The second context information (similarity 0.95): Having problems when upgrading the account? Please refer to Article 4 in the Frequently Asked Questions. The third context information (similarity 0.88): After the account upgrade is successful, you can enjoy more membership privileges. The preset similarity threshold is 0.9, and the first context information and the second context information are screened out. Further optionally, according to the user's intention (i.e., the user mainly cares about how to upgrade the account), the most relevant information is further screened out. The screened-out information is concatenated or synthesized. Thus, the final integration result is, "To upgrade your account membership, please log in to the account, select the membership package, and complete the payment. If you have problems during the upgrade process, please refer to Article 4 in the Frequently Asked Questions."

[0094] In summary, through multi-level memory fusion processing, multiple relevant context information can be integrated together to provide a more accurate and comprehensive answer. The information synthesis method can generate natural and fluent text, enhancing the user experience. The information screening method can further eliminate irrelevant information to ensure the precision of the answer. Through cosine similarity sorting and preset threshold screening, the amount of context information to be processed can be significantly reduced, improving the processing speed and performance of the system.

[0095] In practical applications, different fusion methods can be selected according to specific requirements in the multi-level memory fusion processing, such as the hybrid splicing method, the information superposition method, the information synthesis method, and the information screening method. These methods not only improve the accuracy and comprehensiveness of the answer but also enhance the naturalness and fluency, thus significantly improving the user experience and system performance.

[0096] As an alternative embodiment, after storing each context information in different storage areas of the external memory module in 103, corresponding memory decay coefficients can also be set for the first context information and the second context information in the external memory module; wherein, the lower the access frequency of the context information, the larger the corresponding memory decay coefficient. Then, the Least Recently Used (LRU) strategy is adopted to perform memory management on the first context information and the second context information stored in the external memory module based on the memory decay coefficient.

[0097] Among them, the LRU policy is a common cache eviction algorithm used to manage data in the cache. In this policy, the least recently used data is preferentially deleted to make room for new data. The LRU policy is widely applied in various fields of computer science, including operating systems, databases, and various cache systems. The LRU policy maintains a data structure that records the data access time or access frequency.

[0098] When the cache is full, the least recently used data is preferentially deleted. The LRU policy usually uses a combined data structure called a doubly linked list and a hash table to implement. The doubly linked list is used to record the access order of the data. The newly accessed data is added to the head of the list, and the tail of the list is the least recently used data. The hash table is used to quickly find the data. The key of the hash table is the unique identifier of the data, and the value is a pointer to the corresponding node in the doubly linked list.

[0099] Specifically, create a doubly linked list and a hash table. If the cache is not full, insert the new data into the head of the list and add the corresponding key-value pair to the hash table. If the cache is full, delete the node at the tail of the list (i.e., the least recently used data), then insert the new data into the head of the list and add the corresponding key-value pair to the hash table. When accessing a piece of data, quickly find the node of the data through the hash table. Remove the node from the list and re-insert it into the head of the list. When data needs to be deleted, delete the node from the tail of the list and delete the corresponding key-value pair from the hash table.

[0100] In the external memory module, the LRU policy can be combined with a memory decay coefficient to manage context information. Set corresponding memory decay coefficients for the first context information and the second context information in the external memory module respectively. That is, set the memory decay coefficient according to the access frequency of the context information. The lower the access frequency, the larger the corresponding memory decay coefficient. For example, if a certain context information has not been accessed in the past month, its memory decay coefficient can be set to 1.5, while the context information that is frequently accessed can be set to 1.0.

[0101] Adopt the least recently used policy LRU to perform memory management on the first context information and the second context information stored in the external memory module based on the memory decay coefficient. The access record is to maintain a doubly linked list and a hash table to record the access time and frequency of each context information. Each time a context information is accessed, update its access frequency and adjust the memory decay coefficient. When the cache is full, calculate the comprehensive score of each context information. Comprehensive score = access time + memory decay coefficient. Select the context information with the lowest comprehensive score for eviction.

[0102] Suppose the following context information is stored in the external memory module. The content of the first context information is, "You can upgrade your account membership through the following steps: Log in to the account -> Select a membership package -> Complete the payment". The access frequency is 10 times, and the memory decay coefficient is 1.0. The content of the second context information is, "Having problems when upgrading the account? Please refer to Article 4 in the FAQ". The access frequency is 5 times, and the memory decay coefficient is 1.2. The content of the third context information is, "After the account upgrade is successful, you can enjoy more membership privileges". The access frequency is 2 times, and the memory decay coefficient is 1.5. When the cache is full, one context information needs to be eliminated. Calculate the comprehensive score of each context information. For example, for the first context information: Comprehensive score = recent access time + memory decay coefficient (assuming the recent access times are the same, all 0); for the second context information: Comprehensive score = recent access time + memory decay coefficient; for the third context information: Comprehensive score = recent access time + memory decay coefficient. Since the third context information has the lowest access frequency and the largest memory decay coefficient, its comprehensive score is the lowest and it will be eliminated first.

[0103] In this way, through the LRU strategy, it can be ensured that the context information that is frequently accessed recently is retained in the cache, improving the cache hit rate. Adopting the memory decay coefficient can better manage the storage space, and preferentially retain the context information with high access frequency and high importance. The improvement of the cache hit rate can reduce the number of accesses to the external memory module, thereby enhancing the overall performance of the system. The memory decay coefficient can be dynamically adjusted according to the actual situation to adapt to different usage scenarios and requirements. Thus, by setting the memory decay coefficient for the context information in the external memory module and adopting the LRU strategy for management, the efficiency of the cache and the utilization of the storage space can be effectively improved, thereby enhancing the performance of the system and the user experience. This strategy is not only applicable to text information management, but also can be widely applied to other scenarios that require cache management.

[0104] Here, the update management of context information can be divided into two methods: batch management and scattered management. Batch management is applicable to a large amount of information stored in the same period. By setting a unified update management cycle, these information are uniformly updated or eliminated at the end of each cycle, simplifying the management process. Scattered management, based on the characteristics and importance of context information, formulates different update strategies for each label, and regularly checks and updates the information under each label to achieve more refined and dynamic management.

[0105] Exemplarily, assume that the following context information was stored between 10:00 and 11:00 on November 19, 2024. Further assume that for context information 1, the storage time = November 19, 2024, 10:05, the content = "How to add a bank card", the access frequency = 5 times, and the memory decay coefficient = 1.0. For context information 2, the storage time = November 19, 2024, 10:15, the content = "What to do if you forget your password", the access frequency = 3 times, and the memory decay coefficient = 1.2. For context information 3: the storage time = November 19, 2024, 10:30, the content = "Introduction to membership benefits", the access frequency = 8 times, and the memory decay coefficient = 1.5. The update management period is 1 month. Based on the above assumptions, the update operation could be to check this batch of context information at 11:00 on December 19, 2024. The access frequency of context information 2 is the lowest (3 times), and the comprehensive score is: storage time + memory decay coefficient = 10:15 + 1.2 = 11:15.2. The comprehensive score of context information 1 is: 10:05 + 1.0 = 10:06. The comprehensive score of context information 3 is: 10:30 + 1.5 = 10:31.5. Select context information 2 for elimination.

[0106] Suppose there is the following context information with different tags in the external memory module, such as:

[0107] Context information 1: Tag = account upgrade, storage time = November 19, 2024, 10:00, content = "How to upgrade the account membership", access frequency = 10 times, memory decay coefficient = 1.0.

[0108] Context information 2: Tag = frequently asked questions, storage time = November 19, 2024, 11:00, content = "What to do if you forget your password", access frequency = 5 times, memory decay coefficient = 1.2.

[0109] Context information 3: Tag = membership benefits, storage time = November 19, 2024, 12:00, content = "Introduction to membership benefits", access frequency = 12 times, memory decay coefficient = 1.5.

[0110] Based on the above settings, for tag 1 (account upgrade): the update period is 1 week, the access frequency threshold is 5 times, and the memory decay coefficient is 1.0. For tag 2 (frequently asked questions): the update period is 1 month, the access frequency threshold is 10 times, and the memory decay coefficient is 1.2. For tag 3 (membership benefits): the update period is 3 months, the access frequency threshold is 15 times, and the memory decay coefficient is 1.5.

[0111] The update strategy based on tags could be that for tag 1 (account upgrade): check at 10:00 on the 19th of each week. The access frequency of context information 1 is 10 times, which is higher than the threshold of 5 times, so it is retained.

[0112] Label 2 (Frequently Asked Questions): Check at 11:00 on the 19th of each month. The access frequency of Context Information 2 is 5 times, which is lower than the threshold of 10 times. The comprehensive score is: 11:00 + 1.2 = 11:01.2. Select for elimination.

[0113] Label 3 (Membership Benefits): Check at 12:00 on the 19th of every 3 months. The access frequency of Context Information 3 is 12 times, which is higher than the threshold of 15 times. Retain.

[0114] In this way, batch management simplifies the management process and is applicable to a large amount of information stored in the same time period. Scattered management is more flexible and precise, and different update strategies can be formulated according to the characteristics and importance of context information. Scattered management based on labels can perform refined management for different types of context information to ensure that important information is retained. By setting reasonable update cycles and access frequency thresholds, storage resources can be optimized and unnecessary storage occupancy can be reduced. Both batch management and scattered management can be dynamically adjusted according to the system operation status and user needs to ensure that the system always operates efficiently. By managing context information through batch management and scattered management, the performance of the system and the user experience can be effectively improved, ensuring that the information in the external memory module always remains up-to-date and most relevant.

[0115] Further optionally, in the above steps, the LRU strategy is adopted to perform memory management on the first context information and the second context information stored in the external memory module based on the memory decay coefficient. Specifically, if the external memory module meets the management conditions, the information replacement strategy of the external memory module is determined based on the model application scenario deployed by the external memory module. Among them, the management conditions include at least one of the following: the storage capacity of the external memory module reaches the upper limit, the preset update cycle of the external memory module is reached, and the preset update event of the external memory module is triggered. Furthermore, based on the information replacement strategy, any one or more operations of update, deletion, modification, and merging on the first context information and the second context information stored in the external memory module are performed. Similar to the specific embodiments of the above LRU strategy, it will not be elaborated here.

[0116] Further optionally, in the external memory module, in order to better adapt to different application scenarios, the application scenario of the model can be identified before determining the information replacement strategy, and the information replacement strategy can be dynamically configured and updated according to the model running parameters and scenario environment parameters. Specifically, before determining the information replacement strategy of the external memory module based on the model application scenario deployed by the external memory module, the application scenario of the model deployed by the external memory module can also be identified; the model running parameters and / or scenario environment parameters in the model application scenario can be obtained; based on the model running parameters and / or scenario environment parameters, parameter configuration and parameter update management are performed on the information replacement strategy of the external memory module to dynamically maintain any one or more of the context information types, storage capacities, and storage space attributes to be managed in the external memory module.

[0117] For example, the application scenario of the model deployed by the external memory module is identified through the configuration information, log records, or user-defined methods of the model. The model running parameters and scenario environment parameters in this model application scenario are obtained. Suppose the model application scenario is a customer service chatbot, and the model running parameters are request processing rate, number of dialogue turns, user activity level, etc. Suppose the scenario environment parameters are user access volume, peak time period, system resource utilization rate, etc. Exemplarily, the request processing rate is the number of requests processed per second. The number of dialogue turns is the average number of dialogue turns between the user and the chatbot. The user activity level is the active time period and active frequency of the user. The user access volume is the number of user accesses per day. The peak time period is the peak time period of user access. The system resource utilization rate is the utilization rate of system resources such as CPU and memory.

[0118] Exemplarily, the request processing rate is 100 requests processed per second. The number of dialogue turns is 5 turns per average dialogue. The user activity level is that the user is most active from 2 pm to 6 pm. The user access volume is 1000 users per day. The peak time period is from 4 pm to 6 pm every day. The system resource utilization rate is that the CPU utilization rate reaches 80% during the peak period.

[0119] The context information type to be managed is to determine the information type that needs to be prioritized according to the user activity level and peak time period, such as frequently asked questions (FAQs). The storage capacity is to determine an appropriate storage capacity according to the user access volume and request processing rate, such as storing 10,000 pieces of context information. The storage space attribute is to optimize the allocation attribute of the storage space according to the system resource utilization rate, such as reducing the elimination frequency of expired information during the peak period.

[0120] Periodic updates are for regularly checking the model operation parameters and scenario environment parameters, and adjusting the information replacement strategy according to the new parameters. Event-triggered updates are when the system detects a specific event (such as a sudden increase in user traffic), the information replacement strategy is immediately updated.

[0121] For example, the application scenario is a customer service chatbot. The initial configuration is that the types of context information to be managed are FAQs, account management, and membership benefits. The storage capacity is 10,000 pieces of context information. The storage space attribute is to reduce the elimination frequency of expired information during peak hours.

[0122] Based on this, the obtained parameters are: the request processing rate is 100 requests per second, the number of dialogue turns is 5 turns per average dialogue, the user activity is that users are most active from 2 pm to 6 pm, the user traffic is 1,000 users per day, the peak time period is from 4 pm to 6 pm every day, and the system resource utilization rate is that the CPU utilization rate reaches 80% during peak hours.

[0123] In the dynamic adjustment strategy, for the types of context information to be managed, it can be analyzed that users are most active from 2 pm to 6 pm, and the peak time period is from 4 pm to 6 pm, and the access frequencies of common questions (FAQs) and account management information are relatively high. Or, prioritize the management of FAQs and account management information and reduce the elimination frequency of these types of information.

[0124] The storage capacity is 1,000 user accesses per day, 100 requests are processed per second, and the average number of dialogue turns per dialogue is 5. It is adjusted to adjust the storage capacity to 15,000 pieces of context information to cope with the high traffic during peak hours.

[0125] For the storage space attribute, it is analyzed that the system resource utilization rate reaches 80% during peak hours, and it is necessary to ensure the stable operation of the system during peak hours. It is adjusted to reduce the elimination frequency of expired information during peak hours (from 4 pm to 6 pm) and keep more context information in the cache. During non-peak hours, increase the elimination frequency of expired information to release storage space.

[0126] The periodic update is to set the update period to update the model operation parameters and scenario environment parameters once a week. The update operation is that at 10:00 on the 19th of each week, the system automatically obtains the latest model operation parameters and scenario environment parameters. According to the new parameters, the information replacement strategy is readjusted. For example, if the user traffic increases to 1,500, the storage capacity can be further increased. If the user active time period is adjusted to from 3 pm to 7 pm, the management strategy during peak hours can be adjusted accordingly.

[0127] Event-triggered update: Set the trigger condition to update the information replacement strategy immediately when the user access volume exceeds 1,200 or the system resource utilization rate exceeds 90%. The update operations are as follows: When the user access volume is triggered, that is, when the system detects that the user access volume exceeds 1,200, immediately increase the storage capacity and reduce the information elimination frequency. When the system resource utilization rate is triggered, that is, when the system detects that the CPU utilization rate exceeds 90%, immediately lower the information elimination frequency to reduce the system burden.

[0128] Suppose on November 19, 2024, the context information in the external memory module is as follows: Context information 1 has a label = Account upgrade, a storage time = 10:00 on November 19, 2024, content = "How to upgrade the account membership", an access frequency = 100 times, and a memory decay coefficient = 1.0. Context information 2 has a label = Frequently Asked Questions, a storage time = 11:00 on November 19, 2024, content = "What to do if you forget your password", an access frequency = 80 times, and a memory decay coefficient = 1.2. Context information 3 has a label = Membership benefits, a storage time = 12:00 on November 19, 2024, content = "Introduction to membership benefits", an access frequency = 20 times, and a memory decay coefficient = 1.5.

[0129] The initial information replacement strategy is that the types of context information to be managed are FAQs, Account management, and Membership benefits, the storage capacity is 10,000 pieces of context information, and the storage space attribute is to reduce the elimination frequency of expired information during peak hours.

[0130] The newly obtained model operation parameters and scenario environment parameters are as follows: The user activity level is that users are most active from 3 pm to 7 pm, the user access volume is 1,200 users per day, the peak time period is from 5 pm to 7 pm every day, and the system resource utilization rate is that the CPU utilization rate reaches 95% during peak hours.

[0131] The dynamic adjustment operations are as follows: The types of context information to be managed are to prioritize the management of FAQs and Account management information to maintain a high access frequency for this information. The storage capacity is adjusted to 12,000 pieces of context information to cope with a higher user access volume. The storage space attribute is to reduce the elimination frequency of expired information during peak hours (from 5 pm to 7 pm) to keep more context information in the cache. During non-peak hours, increase the elimination frequency of expired information to free up storage space.

[0132] In summary, by identifying the application scenarios of the recognition model and obtaining the model operation parameters and scenario environment parameters, the information replacement strategy of the external memory module can be dynamically configured and updated. This dynamic management method not only improves the flexibility and adaptability of the system, but also optimizes the utilization of storage resources, ensuring that the information in the external memory module remains up-to-date and most relevant under different application scenarios, and enhancing the system performance and user experience.

[0133] As an optional embodiment, after each context information is separately stored in different storage areas of the external memory module in 103, the storage levels of the context information in different storage areas are also dynamically adjusted. For example, based on different categories of sessions, the management level index is adjusted, and different indexes correspond to the context information of different types of sessions. Specifically, by dynamically adjusting the storage levels of different storage areas in the external memory module, the access efficiency can be significantly improved, the storage resource utilization can be optimized, different user requirements can be adapted, and the flexibility of the system can be enhanced. Specifically, by classifying and storing the context information and setting different level indexes, the system can dynamically adjust the priority of the storage levels according to the activity, importance, user requirements, etc. of the session. For example, assuming that the access frequency of FAQs continues to increase, the system dynamically raises the level of storage area 1 to level 0 (the highest priority) to ensure that the information of FAQs can be accessed quickly, thereby improving the user experience. At the same time, for the context information with higher importance, the system can dynamically allocate more storage space. For example, when the access frequency of account management information increases, the level is raised and more storage space is allocated to avoid performance degradation. In addition, the system can also dynamically adjust the storage level according to the personalized needs of the user. For example, when a certain user frequently queries personal preference information, the level of storage area 4 is raised and the personalized preference information of this user is preferentially stored to provide more accurate services. This dynamic management method not only enhances the system performance and user experience, but also ensures that the external memory module can efficiently and intelligently manage the context information under different application scenarios, quickly respond to changes in user requirements, and provide the best services.

[0134] As an alternative embodiment, after each piece of context information is stored in different storage areas of the external memory module in 103, the storage levels of different categories of context information can also be adjusted based on the predicted value of the business development trend. For example, the context information corresponding to the information content with a higher search frequency has a higher storage level, so that it can be quickly called during the session processing. Specifically, by dynamically adjusting the storage levels of different categories of context information in the external memory module based on the predicted value of the business development trend, the user's search experience can be significantly improved, the utilization of storage resources can be optimized, the intelligence of the system can be enhanced, and marketing activities can be effectively supported. Specifically, when "healthy diet" is searched more times than the preset number of times, the system will dynamically raise the storage level of the context information related to healthy diet (such as healthy diet suggestions, healthy recipes, etc.) to a high priority level to ensure that users can quickly obtain relevant information during the session and improve the user experience. This dynamic management method not only improves the performance of the system and the user experience, but also enables the external memory module to manage context information more intelligently and efficiently, respond to different application scenarios and user needs, quickly respond to changes in user needs, and provide the best service.

[0135] As an alternative embodiment, after each piece of context information is stored in different storage areas of the external memory module in 103, the storage levels of different categories of context information can also be adjusted based on the user's historical behavior.

[0136] For example, in the external memory module, after the context information is stored in different storage areas respectively, the storage levels of different categories of context information can be dynamically adjusted based on the user's historical behavior. This adjustment can raise the priority of the context information of the user's preferences or frequently accessed context information, so as to achieve quick call during the session processing, bringing beneficial effects in many aspects. User behavior analysis is to collect and analyze data such as the user's search history, click records, purchase behavior, etc. to identify the user's preferences and common context information. Dynamically adjusting the storage level is to adjust the storage levels of different categories of context information according to the user's historical behavior and raise the priority of the context information of the user's preferences.

[0137] Exemplarily, the user's historical behavior is that user A frequently searches for information such as "healthy diet" and "fitness plan". The context information categories are healthy diet suggestions, fitness plans, and nutritional knowledge. The effect is that by raising the priority of the context information of the user's preferences, it is ensured that users can quickly obtain personalized information during the session and improve the user experience.

[0138] For example, dynamic adjustment means that based on the search history of User A, the system dynamically elevates the storage levels of "healthy diet suggestions" and "fitness plans" to high priority. Quick response means that when User A asks relevant questions, the system can quickly call the high-priority context information to provide immediate and accurate answers. The effect is that by adjusting the storage levels, the utilization of storage resources can be optimized to ensure that the context information preferred by the user occupies more storage space and computing resources.

[0139] For example, dynamic adjustment means that assuming User A frequently accesses information related to "fitness plans", the system dynamically adjusts the storage space allocation to ensure that there is sufficient storage space for this information. Resource optimization means ensuring that the information related to "fitness plans" preferred by User A has sufficient storage space during peak periods to avoid performance degradation due to insufficient storage space. The effect is that by dynamically adjusting the storage levels, the system can intelligently adjust the storage strategy according to user behavior, improving the intelligence level of the system.

[0140] For instance, dynamic adjustment means that the system can automatically adjust the storage levels based on the behavior data of User A without manual intervention. Intelligent management means that the system can intelligently identify and respond to the preferences of User A, improving the intelligence level of the system. The effect is that by providing personalized services and elevating the priority of the context information preferred by the user, the satisfaction and stickiness of the user can be effectively improved. Dynamic adjustment means that assuming User A frequently queries "healthy diet suggestions", the system dynamically elevates the storage level of this context information to high priority. User stickiness means that when User A visits again, the system can quickly provide detailed healthy diet suggestions, enhancing user satisfaction and stickiness. Suppose on November 19, 2024, the context information in the external memory module is as follows: Context Information 1: Category = healthy diet suggestions, Storage Area = Storage Area 1, Level = Level 1, Content = "How to formulate a healthy diet plan", Access Frequency = 50 times. Context Information 2: Category = fitness plan, Storage Area = Storage Area 2, Level = Level 2, Content = "Personal fitness plan recommendations", Access Frequency = 30 times. Context Information 3: Category = nutritional knowledge, Storage Area = Storage Area 3, Level = Level 3, Content = "Popularization of nutritional knowledge", Access Frequency = 20 times. Analysis of User A's historical behavior: Search history shows that User A frequently searches for content related to "healthy diet" and "fitness plan". Click record shows that User A has clicked on links related to "healthy diet suggestions" and "fitness plan" multiple times. Purchase behavior shows that User A has purchased multiple fitness-related products.

[0141] Accordingly, the dynamic adjustment operation is as follows: Analyze the user behavior: The healthy diet advice is frequently searched and clicked by User A, and the access frequency increases to 80 times. The fitness plan is frequently searched and purchased by User A, and the access frequency increases to 60 times. The nutrition knowledge has a stable access frequency of 20 times for User A. The dynamic adjustment of the storage level is as follows: Storage area 1 (healthy diet advice): The level is upgraded to level 0 (the highest priority). Storage area 2 (fitness plan): The level is upgraded to level 1. Storage area 3 (nutrition knowledge): The level remains at level 3.

[0142] In the above example, enhancing the personalized user experience means that when User A asks for healthy diet advice or a fitness plan, the system can quickly provide detailed answers, improving the user experience. Optimizing resource utilization means that there is sufficient storage space for information related to healthy diet advice and fitness plans, avoiding performance degradation. Enhancing system intelligence means that the system can automatically identify and adjust the storage level of context information preferred by User A, improving the system's intelligence level. Increasing user stickiness means that by providing personalized services and elevating the priority of context information preferred by User A, the user's satisfaction and stickiness are effectively increased.

[0143] In summary, by dynamically adjusting the storage levels of different categories of context information in the external memory module based on the user's historical behavior, the personalized user experience can be significantly improved, the utilization of storage resources can be optimized, the system's intelligence can be enhanced, and the user's satisfaction and stickiness can be effectively increased. This dynamic management method not only improves the system's performance and user experience but also enables the external memory module to manage context information more intelligently and efficiently, meeting different application scenarios and user requirements.

[0144] As an optional embodiment, after storing each context information into different storage areas of the external memory module in 103, the space capacity of different storage areas in the external memory module can also be dynamically adjusted. Specifically, by dynamically adjusting the space capacity of different storage areas in the external memory module, the allocation of storage resources can be significantly optimized, the performance of the system and the user experience can be improved, and the adaptability and flexibility of the system can be enhanced. This dynamic management method not only improves the overall performance of the system and the user experience, but also enables the external memory module to manage storage resources more intelligently and efficiently to cope with different application scenarios and user requirements. Specifically, by analyzing data to identify the access frequency and storage requirements of each storage area, the system can dynamically adjust the space capacity to ensure that important or frequently accessed context information has sufficient storage space, thereby avoiding performance degradation and improving access efficiency. For example, assuming that the access frequency of healthy diet suggestions and fitness plans increases significantly, the system can automatically increase the space capacity of the corresponding storage area to ensure that these information can respond quickly during peak periods. This dynamic adjustment not only optimizes the allocation of storage resources, improves the performance of the system, but also enhances the adaptability of the system, enabling it to automatically adjust the storage strategy according to actual needs without manual intervention. Ultimately, this intelligent management method significantly improves the user experience. Users can obtain faster and more accurate responses when searching for important information, enhancing satisfaction and the overall usage experience.

[0145] In the embodiments of the present application, the above three steps provide significant technical optimizations and beneficial effects for long text data processing by introducing an external memory module and a multi-level memory structure. Specifically, first, by extracting the context information of each text segment, the model can more deeply understand the key content in the long text, avoiding the problem of context loss that easily occurs in traditional models when processing long texts. Decomposing the long text into multiple text segments and extracting their context information helps the model better capture the specific meaning and logical relationship of each segment, improving the refinement of processing. When generating text, it can more accurately refer to the context information of each segment, thereby generating more coherent and logical text content. Furthermore, the respective context information is stored in different storage areas of the external memory module. In this way, through the external memory module (including the short-term memory area and the long-term memory area), context information of different importance and timeliness can be reasonably stored and managed, avoiding information redundancy and chaos. By storing the important context information in the current session in the short-term memory area, the model can quickly access these key information, improving the reasoning and processing speed. The long-term memory area is used to store the context information that appears repeatedly in multiple sessions, helping the model accumulate and utilize long-term knowledge, enhancing the knowledge base and adaptability of the model. Finally, the first context information and the second context information are input into a large language model for natural language processing. By integrating the context information of the current session (short-term memory) and the historical session (long-term memory), the large language model can obtain more comprehensive and coherent context support when generating and understanding text, and the generated content is more natural and logical. Combining the context information of short-term and long-term memories, the model can more accurately understand the content of long texts and generate outputs highly relevant to the context, significantly improving the accuracy and relevance of the model. The design of the multi-level memory structure enables this technology to be applicable not only to long text processing but also to play a role in other tasks that require context memory, enhancing the versatility and adaptability of the model. In summary, these three steps significantly improve the context understanding ability, computational efficiency, and result accuracy of long text processing by introducing an external memory module and a multi-level memory structure, while enhancing the flexible adaptability of the model, providing more comprehensive and efficient technical support for natural language processing tasks.

[0146] Based on the same implementation principle, the embodiments of the present application also provide a long text processing device for implementing the long text processing method in the above embodiments. Figure 2 The structural block diagram of the long text processing device provided by the embodiments of the present application is as follows Figure 2 As shown, the device at least includes the following units:

[0147] An extraction unit, configured to extract the context information corresponding to each text segment from the long text data to be processed in the current session;

[0148] A storage unit configured to separately store respective context information in different storage areas of an external memory module; the external memory module includes a short-term memory area and a long-term memory area; the short-term memory area is used to store first context information whose importance reaches a preset condition in the current session; the long-term memory area is used to store second context information that appears repeatedly in multiple sessions; the multiple sessions include the current session and / or historical sessions;

[0149] An execution unit configured to input the first context information and the second context information into a large language model, and implement natural language processing of the long text data through the large language model.

[0150] In an optional embodiment, when the storage unit separately stores respective context information in different storage areas of the external memory module, it is configured to:

[0151] Through a self-attention mechanism, obtain importance evaluation values corresponding to respective context information in the current session; the importance evaluation values are used to indicate the correlation between the context information and the long text data; the higher the correlation between the context information and the long text data, the higher the importance evaluation value corresponding to the context information;

[0152] Select first context information in the current session whose importance evaluation values meet the preset conditions, and store the first context information in the short-term memory area;

[0153] Match respective context information in the current session with the second context information stored in the long-term memory area, and use the context information that appears repeatedly in the current session and historical sessions as new second context information to update the long-term memory area.

[0154] In an optional embodiment, when the storage unit separately stores respective context information in different storage areas of the external memory module, it is further configured to:

[0155] Obtain access frequencies corresponding to respective context information in the current session;

[0156] Select first context information in the current session whose access frequencies are higher than a set access frequency threshold, and store the first context information in the short-term memory area.

[0157] In an optional embodiment, when the execution unit inputs the first context information and the second context information into the large language model, it is configured to:

[0158] Obtain target text information to be processed currently from the current session;

[0159] Generate a corresponding retrieval request based on the target text information; at least one or more of the text content information, text type, and text theme in the target text information are included in the retrieval request;

[0160] Retrieve first context information and second context information that match the target text information from the external memory module based on the retrieval request;

[0161] Fuse the retrieved first context information and second context information, and input the result of the fusion processing into the large language model.

[0162] In an alternative embodiment, when the execution unit retrieves the first context information and the second context information that match the target text information from the external memory module based on the retrieval request, it is configured to:

[0163] Extract the target text segment from the target text information;

[0164] Obtain the cosine similarity between each of the first context information and the second context information in the external memory module and the target text segment;

[0165] Select the first context information and the second context information whose cosine similarity meets the preset conditions as the first context information and the second context information that match the target text segment.

[0166] In an alternative embodiment, when the execution unit fuses the retrieved first context information and second context information, it is configured to:

[0167] Sort the retrieved first context information and second context information in descending order according to the cosine similarity to select the first context information and the second context information whose similarity is at a preset position or reaches a preset similarity threshold;

[0168] Perform multi-level memory fusion processing on the selected first context information and second context information to obtain the fusion result.

[0169] In an alternative embodiment, the first context information and the second context information in the external memory module are respectively associated with corresponding dynamic indexes according to their respective text content characteristics and / or text format types; before the execution unit obtains the cosine similarity between each of the first context information and the second context information in the external memory module and the target text segment, it is also configured to:

[0170] Identify the text attribute features of the target text segment; the text attribute features are used to represent the text content characteristics and / or text format types of the target text segment;

[0171] Construct the text retrieval preconditions of the target text segment based on the text content characteristics and / or text format types;

[0172] Based on the text retrieval preconditions, determine the target dynamic index in the external memory module that matches the target text segment, and obtain the first context information and the second context information associated with the target dynamic index as the first context information and the second context information to be filtered. In an optional embodiment, after the storage unit stores each context information in different storage areas of the external memory module respectively, it is further configured to:

[0173] Set corresponding memory decay coefficients for the first context information and the second context information in the external memory module respectively; among them, the lower the access frequency of the context information, the larger the corresponding memory decay coefficient;

[0174] Adopt the least recently used (LRU) strategy to perform memory management on the first context information and the second context information stored in the external memory module based on the memory decay coefficient.

[0175] In an optional embodiment, when the storage unit adopts the least recently used (LRU) strategy to perform memory management on the first context information and the second context information stored in the external memory module based on the memory decay coefficient, it is configured to:

[0176] If the external memory module meets the management conditions, determine the information replacement strategy of the external memory module based on the model application scenario deployed by the external memory module; where the management conditions include at least one of the following: the storage capacity of the external memory module reaches the upper limit, reaches the preset update period of the external memory module, triggers the preset update event of the external memory module;

[0177] Based on the information replacement strategy, perform any one or more operations of updating, deleting, modifying, and merging on the first context information and the second context information stored in the external memory module.

[0178] In an optional embodiment, before the storage unit determines the information replacement strategy of the external memory module based on the model application scenario deployed by the external memory module, it is further configured to:

[0179] Identify the model application scenario deployed by the external memory module;

[0180] Obtain the model operation parameters and / or scenario environment parameters in the model application scenario;

[0181] Based on the model operation parameters and / or scenario environment parameters, perform parameter configuration and parameter update management on the information replacement strategy of the external memory module, so as to dynamically maintain any one or more of the context information types, storage capacities, and storage space attributes to be managed in the external memory module.

[0182] It should be noted that regarding the long text processing device provided in the implementation of this application, the specific functions and implementation details of each module have the same implementation principle as the implementation process of the corresponding steps in the foregoing method embodiments. Specifically, reference can be made to the description of the corresponding parts in the above method embodiments, and details will not be repeated here.

[0183] Based on the same implementation principle, an embodiment of this application also provides an electronic device, Figure 3 which is the structural block diagram of the electronic device 300, as Figure 3 shown. The electronic device 300 includes: a processor 301, a memory 302, a communication interface 303, a communication bus 304, and a controller 305; wherein, the processor 301, the memory 302, and the communication interface 303 complete mutual communication through the communication bus 304; the memory 302 is used to store computer programs; the processor 301 is used to execute the programs stored in the memory 302 to implement corresponding processing functions; the communication bus 304 is used for communication between the electronic device 300 and other devices; the controller 305 is used to implement the long text processing method described in the foregoing method embodiments.

[0184] In the embodiment of this application, the communication bus 304 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 304 may be divided into an address bus, a data bus, a control bus, etc., and the specific form is not limited. For the sake of representation, Figure 3 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0185] The memory 302 may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the foregoing processor 301.

[0186] The processor 301 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the specific form may be determined according to actual requirements.

[0187] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement each step executable by the electronic device in the above method embodiment.

[0188] The above describes one or more embodiments of the present application in combination with optional embodiments. However, these embodiments are merely exemplary and only serve an illustrative purpose. On this basis, various replacements and improvements can be made to one or more embodiments of the present application, and these all fall within the protection scope of one or more embodiments of the present application.

Claims

1. A long text processing method, characterized in that, The method includes: For the long text data to be processed in the current session, extracting the context information corresponding to each text segment from the long text data; Storing each context information into different storage areas of an external memory module respectively; the external memory module includes a short-term memory area and a long-term memory area; the short-term memory area is used to store the first context information whose importance reaches a preset condition in the current session; the long-term memory area is used to store the second context information that repeatedly appears in multiple sessions; multiple sessions include the current session and / or historical sessions; After storing each context information into different storage areas of the external memory module respectively, dynamically adjusting the storage levels of the context information in different storage areas; among them, based on different categories of sessions, managing the level index, and different indexes correspond to the context information of different types of sessions; dynamically adjusting the priority of the storage levels according to the activity, importance, and user requirements of different categories of sessions; based on the predicted value of the business development trend, adjusting the storage levels of different categories of context information; Inputting the first context information and the second context information into a large language model, and implementing natural language processing on the long text data through the large language model; in the large language model, performing multi-level memory fusion processing on the first context information and the second context information; selecting different fusion methods according to specific requirements in the multi-level memory fusion processing; the fusion methods are at least one of the hybrid splicing method, the information superposition method, the information synthesis method, and the information screening method.

2. The method according to claim 1, characterized in that, The storing each context information into different storage areas of the external memory module respectively includes: Obtaining the importance evaluation value corresponding to each context information in the current session through a self-attention mechanism; the importance evaluation value is used to indicate the correlation between the context information and the long text data; the higher the correlation between the context information and the long text data, the higher the importance evaluation value corresponding to the context information; Selecting the first context information in the current session whose importance evaluation value meets the preset condition, and storing the first context information into the short-term memory area; Matching each context information in the current session with the second context information stored in the long-term memory area, and taking the context information that repeatedly appears in the current session and historical sessions as the newly added second context information and updating it into the long-term memory area.

3. The method according to claim 2, characterized in that, The storing each context information into different storage areas of the external memory module respectively further includes: Obtaining the access frequency corresponding to each context information in the current session; Selecting the first context information in the current session whose access frequency is higher than the set access frequency threshold, and storing the first context information into the short-term memory area.

4. The method according to claim 1, wherein The inputting the first context information and the second context information into the large language model includes: Obtaining the target text information to be processed currently from the current session; Generate a corresponding retrieval request based on the target text information; at least one or more of the text content information, text type, and text theme in the target text information are included in the retrieval request; Retrieve first context information and second context information that match the target text information from the external memory module based on the retrieval request; Fuse the retrieved first context information and second context information, and input the result of the fusion processing into the large language model.

5. The method according to claim 4, wherein The retrieving first context information and second context information that match the target text information from the external memory module based on the retrieval request includes: Extract the target text fragment from the target text information; Obtain the cosine similarity between the first context information and the second context information in the external memory module and the target text fragment respectively; Select the first context information and the second context information whose cosine similarity meets the preset conditions as the first context information and the second context information that match the target text fragment.

6. The method according to claim 5, wherein The fusing the retrieved first context information and second context information includes: Sort the retrieved first context information and second context information in descending order of cosine similarity to select the first context information and the second context information whose similarity is at a preset position or reaches a preset similarity threshold; Perform multi-level memory fusion processing on the selected first context information and second context information to obtain a fusion result.

7. The method according to claim 5, wherein The first context information and the second context information in the external memory module are respectively associated with corresponding dynamic indexes according to their respective text content characteristics and / or text format types; Before obtaining the cosine similarity between the first context information and the second context information in the external memory module and the target text fragment respectively, it further includes: Identify the text attribute features of the target text fragment; the text attribute features are used to represent the text content characteristics and / or text format types of the target text fragment; Construct a text retrieval precondition for the target text fragment based on the text content characteristics and / or text format types; Based on the text retrieval precondition, determine the target dynamic index in the external memory module that matches the target text fragment, and obtain the first context information and the second context information associated with the target dynamic index as the first context information and the second context information to be screened.

8. The method according to claim 1, wherein After storing each context information in different storage areas of the external memory module respectively, it further includes: Set corresponding memory decay coefficients for the first context information and the second context information in the external memory module respectively; among them, the lower the access frequency of the context information, the larger the corresponding memory decay coefficient; Adopt the least recently used strategy LRU to manage the memory of the first context information and the second context information stored in the external memory module based on the memory decay coefficient.

9. The method according to claim 8, characterized in that The Least Recently Used (LRU) strategy is adopted, and memory management is performed on the first context information and the second context information stored in the external memory module based on a memory decay coefficient, including: If the external memory module meets the management conditions, an information replacement strategy for the external memory module is determined based on the model application scenario in which the external memory module is deployed; wherein, the management conditions include at least one of the following: the storage capacity of the external memory module reaches the upper limit, the preset update period of the external memory module is reached, or a preset update event of the external memory module is triggered; Based on the information replacement strategy, any one or more operations of updating, deleting, modifying, and merging the first context information and the second context information stored in the external memory module are performed.

10. The method according to claim 9, wherein Before determining the information replacement strategy for the external memory module based on the model application scenario in which the external memory module is deployed, it further includes: Identifying the model application scenario in which the external memory module is deployed; Obtaining the model operation parameters and / or scenario environment parameters in the model application scenario; Based on the model operation parameters and / or scenario environment parameters, parameter configuration and parameter update management are performed on the information replacement strategy of the external memory module to dynamically maintain any one or more of the context information types, storage capacity, and storage space attributes to be managed in the external memory module.

11. A long text processing device, characterized in that, The device at least includes the following units: An extraction unit configured to extract the context information corresponding to each text segment from the long text data to be processed in the current session; A storage unit configured to store each context information in different storage areas of an external memory module; the external memory module includes a short-term memory area and a long-term memory area; the short-term memory area is used to store the first context information whose importance reaches a preset condition in the current session; the long-term memory area is used to store the second context information that appears repeatedly in multiple sessions; the multiple sessions include the current session and / or historical sessions; after storing each context information in different storage areas of the external memory module, the storage levels of the context information in different storage areas are dynamically adjusted; wherein, based on different categories of sessions, the management level index is used, and different indexes correspond to the context information of different types of sessions; the priority of the storage level is dynamically adjusted according to the activity, importance, and user needs of different categories of sessions; based on the predicted value of the business development trend, the storage levels of different categories of context information are adjusted; An execution unit configured to input the first context information and the second context information into a large language model, and implement natural language processing of the long text data through the large language model; in the large language model, multi-level memory fusion processing is performed on the first context information and the second context information; different fusion methods are selected according to specific requirements in the multi-level memory fusion processing; the fusion methods include at least one of the hybrid splicing method, the information superposition method, the information synthesis method, and the information screening method.

12. A chip, characterized in that, The chip includes a processor coupled to a transceiver for performing the long text processing method according to any one of claims 1-10.

13. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the long text processing method according to any one of claims 1-10.

14. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed, the long text processing method according to any one of the above claims 1-10 is implemented.

Citation Information

Patent Citations

  • Human-computer interaction method, device and equipment and storage medium

    CN117421398A