Retrieval-Enhanced Generation Method, System, Electronic Device and Storage Medium Based on Long-Term Memory
By compressing the database content into global memory and combining in-depth analysis to generate answers, the traditional system's insufficient answer generation in long text and complex queries is solved, and efficient and accurate answer provision is achieved.
Patent Information
- Application Number
- CN202510026857.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-01-08
AI Technical Summary
Traditional retrieval enhancement generation systems are difficult to effectively handle long text and implicit information needs, and cannot generate accurate and comprehensive answers when faced with cross-domain or multi-step complex queries.
The database content is compressed into global memory by using a large language model, and the preliminary answers are generated as search clues, and relevant information is retrieved from global memory, and the final answer is generated in combination with in-depth analysis. A refined clue generation and search strategy is designed to optimize the answer quality.
It realizes efficient, accurate and flexible answer generation in processing large amounts of text data and complex query scenarios, significantly improving the response quality of complex query.
Smart Images

Figure CN119415664B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a retrieval-enhanced generation method, system, electronic device and storage medium based on long-term memory. Background Art
[0002] Large language models are powerful artificial intelligence models with excellent text generation and understanding capabilities. By training on a large amount of text data, they can learn the patterns and rules of language and generate coherent and reasonable answers or texts based on the input prompts or questions. Large language models have been widely used in fields such as machine translation, automatic summarization, dialogue systems, and text generation.
[0003] Traditional retrieval-enhanced generation systems often rely on short-term, query-based direct relevance matching. This method has obvious deficiencies in dealing with tasks that require in-depth context understanding and complex reasoning, and cannot effectively handle long texts and implicit information requirements. At the same time, in the face of cross-domain or multi-step complex queries, traditional methods are also difficult to perform effective information integration and reasoning, resulting in the inability to generate accurate and comprehensive answers. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a retrieval-enhanced generation method, system, electronic device and storage medium based on long-term memory, which is used to fully or at least partially solve the technical problems existing in the above-mentioned prior art, such as the inability to effectively process long texts and implicit information requirements. At the same time, in the face of cross-domain or multi-step complex queries, it is difficult to perform effective information integration and reasoning, resulting in the inability to generate accurate and comprehensive answers.
[0005] In a first aspect, an embodiment of the present application provides a retrieval-enhanced generation method based on long-term memory, including:
[0006] Using a large language model, compress and semantically encode the content of the entire database to form a global memory;
[0007] Based on the query proposed by the user, generate a preliminary answer, and use the preliminary answer as retrieval information to retrieve relevant information from the database;
[0008] Based on the retrieved relevant information and the global memory, generate a final answer and provide it to the user.
[0009] Optionally, using a large language model, compress and semantically encode the content of the entire database to form a global memory, including:
[0010] Read and aggregate the data in the database, and use a large language model based on Transformer to compress long texts into smaller memory units through a multi-layer attention mechanism, while retaining key semantic information to form global memory. Among them, the database includes a relational database and a non-relational database.
[0011] Optionally, reading and aggregating the data in the database, and using a large language model based on Transformer to compress long texts into smaller memory units through a multi-layer attention mechanism, while retaining key semantic information to form global memory, includes:
[0012] Introduce a global memory construction algorithm for constructing a global memory representation containing semantic information from the original long text data:
[0013] Representation step of the original input token: The original input token is processed by the Transformer model and converted into query, key, and value matrices for calculating the attention weights of the query, key, and value matrices;
[0014] Calculation step of the attention mechanism: Use the softmax function to calculate the attention weights and obtain the weighted values using the attention weights;
[0015] Initialization step of the memory token: After each context window, append a preset number of memory tokens and initialize another set of weight matrices for the query, key, and value of the memory tokens;
[0016] Update step of the memory token: Use the same attention mechanism as the original input token and use the weight matrix of the memory token to calculate the updated representation of the memory token;
[0017] Memory formation step: After being processed by multiple layers of Transformer, the original input token is encoded into a hidden state, and after the memory is formed, the key-value cache of the original token is discarded, where the hidden state includes the hidden state of the original token and the hidden state of the memory token;
[0018] Semantic compression step: Compress the long text into global memory;
[0019] Training step of the global memory construction algorithm: Includes a pre-training and a supervised fine-tuning phase. Among them, in the pre-training phase, the global memory construction algorithm uses a long context dataset to learn to form memories from the original context. In the supervised fine-tuning phase, the global memory construction algorithm uses the data of a preset task to generate task cues based on the formed memories.
[0020] Optionally, the global memory representation is:
[0021] ;
[0022] Wherein, is the i-th vector in the hidden state representation, is the attention weight corresponding to the i-th word in the attention weight matrix output by the attention mechanism.
[0023] Optionally, based on the query proposed by the user, a preliminary answer is generated, and the preliminary answer is used as retrieval information to retrieve relevant information from the database, including:
[0024] After receiving the query proposed by the user, the query is parsed, and a preliminary answer draft is generated according to the global memory, and the preliminary answer draft is used as a retrieval clue to represent the information requirement behind the user's query;
[0025] Use a large language model and prompt engineering to evaluate the generated retrieval clue. If it does not meet the requirements, it is regenerated to ensure the accuracy of the retrieval clue;
[0026] Using the retrieval clue, retrieve the information fragment most relevant to the user's query from the database, and screen and integrate the retrieved information fragments to remove redundant and irrelevant content, forming an evidence set for subsequent answer generation.
[0027] Optionally, based on the retrieved relevant information and global memory, a final answer is generated and provided to the user, including:
[0028] Using a large language model, fuse the query proposed by the user, the retrieved evidence set, and the global memory to generate an initial answer, and post-process the generated initial answer to generate a final answer.
[0029] Optionally, before generating a final answer based on the retrieved relevant information and global memory and providing it to the user, the retrieval enhanced generation method based on long-term memory further includes:
[0030] Fine-tune the Transformer-based large language model based on the collected data so that the Transformer-based large language model can better understand the user's query intention.
[0031] In a second aspect, an embodiment of the present application further provides a retrieval enhanced generation system based on long-term memory, including:
[0032] A global memory construction module, configured to use a large language model to compress and semantically encode the content of the entire database to form a global memory;
[0033] A retrieval clue generation and information retrieval module, configured to generate a preliminary answer based on a query proposed by a user, and retrieve relevant information from a database using the preliminary answer as retrieval information;
[0034] An answer generation module, configured to generate a final answer based on the retrieved relevant information and global memory, and provide it to the user.
[0035] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned retrieval enhanced generation method based on long-term memory are implemented.
[0036] In a fourth aspect, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned retrieval enhanced generation method based on long-term memory are implemented.
[0037] As can be seen from the above technical solutions, the present invention has the following advantages:
[0038] In the retrieval enhanced generation method, system, electronic device, and storage medium based on long-term memory provided by the present application, first, a lightweight large language model is used to compress and semantically encode the content of the entire database to form global memory. Subsequently, based on a query proposed by a user, the system generates preliminary answers, and these answers are used as retrieval clues to retrieve relevant information from the global memory. At the same time, in combination with the retrieved information, a heavyweight large language model is used for in-depth analysis and answer generation, leveraging its powerful expressive ability to optimize the quality and accuracy of the answers. Moreover, by designing a set of refined clue generation and retrieval strategies, by adjusting the detail level and relevance of the retrieval clues, it is ensured that the most valuable information can be quickly and accurately located. Finally, in combination with the retrieved detailed evidence and preliminary answers, a final detailed answer is generated and provided to the user. It has the advantages of high efficiency, accuracy, and flexibility, and provides an effective solution for application scenarios that require processing a large amount of text data and complex queries. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0040] Figure 1 FIG. is an implementation flowchart of a retrieval enhanced generation method based on long-term memory provided by an embodiment of the present invention;
[0041] Figure 2 A flowchart for retrieving relevant information provided by an embodiment of the present invention;
[0042] Figure 3 A detailed implementation flowchart of a retrieval enhanced generation method based on long-term memory provided by an embodiment of the present invention;
[0043] Figure 4 A schematic structural diagram of a retrieval enhanced generation system based on long-term memory provided by an embodiment of the present invention;
[0044] Figure 5 A schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0045] In the following detailed description, various embodiments of the present disclosure will be more fully described. The present disclosure may have various embodiments and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents and / or alternative solutions falling within the spirit and scope of the various embodiments of the present disclosure.
[0046] In the following, the term "comprising" or "may comprise" that may be used in various embodiments of the present disclosure indicates the presence of the disclosed function or operation and does not limit the addition of one or more functions or operations. Further, as used in various embodiments of the present disclosure, the terms "comprising", "having" and their cognates are only intended to indicate a specific feature, number, step, operation or combination of the foregoing items, and should not be construed as first excluding the existence or addition of one or more other features, numbers, steps, operations or combinations of the foregoing items.
[0047] In various embodiments of the present disclosure, the expression "or" or "at least one of A or / and B" includes any combination or all combinations of the listed words. For example, the expression "A or B" or "at least one of A or / and B" may include A, may include B, or may include both A and B.
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] Refer to Figure 1The following is a flowchart of a retrieval-augmented generation method based on long-term memory in a specific embodiment, including the following execution steps:
[0050] Step 100: Use a large language model to compress and semantically encode the content of the entire database to form a global memory.
[0051] Specifically, using a large language model to compress and semantically encode the content of the entire database to form a global memory includes: reading and aggregating the data in the database, and using a large language model based on Transformer to compress long texts into smaller memory units through a multi-layer attention mechanism while retaining key semantic information to form a global memory, where the database includes a relational database and a non-relational database.
[0052] It should be understood that this global memory can not only effectively store the key content of the database, but also quickly respond to various query requirements in subsequent steps.
[0053] In a specific embodiment, a relational database such as MySQL or PostgreSQL is connected using a JDBC (Java Database Connectivity) or ORM (Object-Relational Mapping) framework (such as Hibernate or MyBatis). Database connection configuration information is written, including the database URL, username, password, etc. According to business requirements, SQL query statements are written to extract the required data from the table. The database-provided query language (such as the MongoDB QueryLanguage for MongoDB) or API is used to extract the data. For complex queries, an aggregation pipeline (such as the AggregationPipeline in MongoDB) is considered. Invalid or redundant data fields are removed. The data is converted into a unified format, such as JSON or CSV, and data cleaning tools (such as Apache Spark, Pandas, etc.) are used for data cleaning and conversion. According to business requirements, aggregation operations such as grouping, summing, and averaging are performed on the data, using SQL aggregation functions (such as SUM, AVG, GROUP BY) or the aggregation capabilities provided by the database. The aggregated data is stored in a temporary table, an in-memory database (such as Redis), or a file for subsequent processing. Natural language processing libraries (such as NLTK, spaCy) are used for preprocessing operations such as tokenization, stop-word removal, and stemming. When tokenizing the text, a pre-trained word embedding model (such as Word2Vec, GloVe) is used to improve the tokenization effect. A pre-trained large language model based on Transformer is selected. The encoder of the Transformer model is used to encode the preprocessed text. The encoder converts the text into a series of feature vectors that contain the semantic information of the text. In the Transformer model, the multi-head self-attention mechanism is one of the core components. The API or library provided by the model is used to calculate the attention weights at each position in the text. Based on the attention weights, important information in the text can be extracted and irrelevant information can be ignored. A pooling layer (such as mean pooling, max pooling, etc.) is used to compress the output matrix of the encoder into a smaller memory unit. The compressed memory unit is stored in a database (such as MySQL, MongoDB, etc.) or in memory (such as Redis). According to business requirements, query statements or API calls are written to retrieve relevant information from the global memory. When the data in the database changes, the global memory is updated in a timely manner.
[0054] Step 101: Based on the query proposed by the user, a preliminary answer is generated, and the preliminary answer is used as retrieval information to retrieve relevant information from the database.
[0055] Exemplarily, a query proposed by the user is, for example, "Please tell me the relationships between the main characters in the Harry Potter series."
[0056] Step 102: Generate a final answer based on the retrieved relevant information and the global memory, and provide it to the user.
[0057] In this embodiment, first, a lightweight large language model (LLM) is used to compress and semantically encode the content of the entire database to form a global memory. Subsequently, based on the query proposed by the user, the system generates preliminary answers, which serve as retrieval clues to guide the system to retrieve relevant information from the global memory. At the same time, in combination with the retrieved information, a heavyweight LLM is used for in-depth analysis and answer generation, leveraging its powerful expressive ability to optimize the quality and accuracy of the answers. To optimize the retrieval performance, a refined clue generation and retrieval strategy is also designed to ensure that the most valuable information can be quickly and accurately located by adjusting the detail and relevance of the retrieval clues. Finally, a final detailed answer is generated by combining the retrieved detailed evidence and the preliminary answers and provided to the user. It has the advantages of high efficiency, accuracy, and flexibility, providing an effective solution for application scenarios that require processing a large amount of text data and complex queries.
[0058] In one embodiment of the present invention, based on step 100, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.
[0059] Introduce a global memory construction algorithm for constructing a global memory representation containing semantic information from the original long text data, which specifically includes the following steps:
[0060] S1: Representation step of the original input token: The original input token is processed by the Transformer model and converted into query, key, and value matrices for calculating the attention weights of the query, key, and value matrices.
[0061] , , ;
[0062] wherein, is the embedding representation of the input token, , and are the weight matrices of the query (Query, Q), key (Key, K), and value (Value, V) respectively.
[0063] S2: Calculation step of the attention mechanism: Use the softmax function to calculate the attention weights and obtain the weighted values using the attention weights.
[0064] ;
[0065] wherein, is the dimension of the key vector.
[0066] S3: Initialization step of memory tokens: After each context window, append a preset number of memory tokens and initialize another set of weight matrices for queries, keys, and values for the memory tokens.
[0067] Exemplarily, another set of weight matrices for queries, keys, and values is denoted as , and .
[0068] S4: Update step of memory tokens: Calculate the updated representation of the memory tokens using the same attention mechanism as the original input tokens and the weight matrices of the memory tokens.
[0069] , , ;
[0070] ;
[0071] S5: Memory formation step: After being processed by multiple layers of Transformer, the original input tokens are encoded into hidden states, and after memory formation, the key-value caches of the original tokens are discarded, wherein the hidden states include the hidden states of the original tokens and the hidden states of the memory tokens.
[0072] It should be understood that after being processed by multiple layers of Transformer, the original input tokens are encoded into hidden states, and these hidden states include the hidden states of the original tokens and the hidden states of the memory tokens. After memory formation, the model discards the key-value (KV) caches of the original tokens, similar to the forgetting process of human memory.
[0073] S6: Semantic compression step: Compress the long text into global memory.
[0074] S7: Training step of the global memory construction algorithm: It includes a pre-training and a supervised fine-tuning phase. In the pre-training phase, the global memory construction algorithm uses a long context dataset to learn to form memories from the original context. In the supervised fine-tuning phase, the global memory construction algorithm uses the data of a preset task to generate task cues based on the formed memories.
[0075] Exemplarily, the global memory is represented as:
[0076] ;
[0077] In the formula, is the i-th vector in the hidden state representation, is the attention weight corresponding to the i-th word in the attention weight matrix output by the attention mechanism.
[0078] In an embodiment of the present invention, referring to Figure 2 shown, based on step 101, the following will give a possible embodiment to non-restrictively elaborate on its specific implementation.
[0079] S1010: After receiving the query proposed by the user, parse the query, generate a preliminary answer draft according to the global memory, and use the preliminary answer draft as a retrieval clue to represent the information requirement behind the user's query.
[0080] Exemplarily, query parsing includes preprocessing the query input by the user, including operations such as removing special symbols, coreference resolution, and normalization of time and location.
[0081] S1011: Use a large language model and prompt engineering to evaluate the generated retrieval clue. If it does not meet the requirements, regenerate it to ensure the accuracy of the retrieval clue.
[0082] S1012: Use the retrieval clue to retrieve the information fragment most relevant to the user's query in the database, and screen and integrate the retrieved information fragments to remove redundant and irrelevant content, forming an evidence set for subsequent answer generation.
[0083] Exemplarily, this step uses a variety of retrieval methods, including sparse retrieval, dense retrieval, and re-ranking, and selects a suitable retrieval method according to the complexity of the task and the characteristics of the database. For example, when dealing with a large amount of unstructured text, dense retrieval can be selected.
[0084] Exemplarily, post-process the generated answer, such as removing redundant information, adjusting word order, and correcting grammar errors, etc., to improve the accuracy and fluency of the answer.
[0085] In some embodiments, when performing step 102, the steps can be executed: use a large language model to fuse the query proposed by the user, the retrieved evidence set, and the global memory to generate an initial answer, and post-process the generated initial answer to generate a final answer.
[0086] In some embodiments, before generating the final answer based on the retrieved relevant information and global memory and providing it to the user, the method flow of retrieval-enhanced generation based on long-term memory further performs the step of fine-tuning the Transformer-based large language model based on the collected data so that the Transformer-based large language model can better understand the user's query intention.
[0087] In one embodiment, Figure 3 FIG. 5 is a detailed flowchart of a retrieval-enhanced generation method based on long-term memory according to an embodiment of the present invention. This embodiment is further optimized and extended on the basis of the above embodiments.
[0088] S300: Use the large language model to compress and semantically encode the content of the entire database to form global memory.
[0089] S301: After receiving the query proposed by the user, parse the query, generate a preliminary answer draft based on the global memory, and use the preliminary answer draft as a retrieval clue to represent the information requirement behind the user's query.
[0090] S302: Use the large language model and prompt engineering to evaluate the generated retrieval clue. If the requirement is not met, regenerate it to ensure the accuracy of the retrieval clue.
[0091] S303: Use the retrieval clue to retrieve the information fragments most relevant to the user's query in the database, and screen and integrate the retrieved information fragments to remove redundant and irrelevant content to form an evidence set for subsequent answer generation.
[0092] S304: Use the large language model to fuse the query proposed by the user, the retrieved evidence set, and the global memory to generate an initial answer, and post-process the generated initial answer to generate the final answer.
[0093] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0094] As Figure 4 shown, the following is an embodiment of a retrieval-enhanced generation system based on long-term memory provided by an embodiment of the present disclosure. It belongs to the same inventive concept as the retrieval-enhanced generation method based on long-term memory in the above embodiments. For the details not described in detail in the embodiment of the retrieval-enhanced generation system based on long-term memory, reference may be made to the embodiment of the retrieval-enhanced generation method based on long-term memory.
[0095] A global memory construction module, which uses a large language model to compress and semantically encode the content of the entire database to form global memory;
[0096] A retrieval clue generation and information retrieval module, which is used to generate a preliminary answer based on the query proposed by the user, and retrieve relevant information from the database using the preliminary answer as the retrieval information;
[0097] An answer generation module, which is used to generate a final answer based on the retrieved relevant information and global memory and provide it to the user.
[0098] The retrieval augmented generation system based on long-term memory of this application refers to Figure 4 As shown, it includes multiple core modules: a global memory construction module, a retrieval clue generation module, an information retrieval module, an answer generation module, and a model fine-tuning module. Among them, the global memory construction module is responsible for processing and encoding the content of the entire database, providing a comprehensive knowledge base for the system. The retrieval clue generation module uses the global memory to extract key information from the user query to ensure that the generated clues can accurately guide the subsequent retrieval process. The information retrieval module retrieves the most relevant information fragments from the database for these generated clues through an efficient search algorithm. The answer generation module constructs accurate and comprehensive answers based on the retrieved information and the original query through natural language processing technology. The model fine-tuning module improves the performance and adaptability of the system for specific application scenarios and user requirements. These modules work together to form the retrieval augmented generation system based on long-term memory of this application, significantly improving the processing ability for complex queries and long text data.
[0099] To facilitate the understanding of this application, an implementation example is given below. Taking the user input "Please tell me the relationships between the main characters in the Harry Potter series." as an example:
[0100] 1) Global memory construction module: The system first extracts all the text content of the Harry Potter series from the database, including books, character descriptions, and plot summaries. Then, it uses a Transformer-based LLM to process and encode these texts to generate a set of memory tokens.
[0101] 2) Retrieval clue generation module: When the user enters a query, the system uses the global memory to extract key information from it. The system analyzes the user's query, identifies "main characters" and "relationships" as key elements, and generates relevant retrieval clues, such as "the friendship between Harry and Hermione", "the relationship between Ron and Harry", etc.
[0102] 3) Information Retrieval Module: The information retrieval module receives the generated retrieval clues and uses an efficient search algorithm (such as vector-based similarity search) to find information fragments related to these clues in the database. The content retrieved by the system includes interactions, conversations, and plot descriptions between characters.
[0103] 4) Answer Generation Module: The answer generation module uses natural language processing techniques to construct an accurate and comprehensive answer based on the retrieved information fragments and the user's original query. The system integrates the retrieved information to generate the final response.
[0104] 5) Model Fine-tuning Module: The model fine-tuning module fine-tunes the large model based on the collected data to continuously optimize the performance of the system for specific application scenarios and user needs. Through fine-tuning, the system can better understand the user's query intention and improve the response quality to complex queries.
[0105] Figure 5 It is a schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.
[0106] The retrieval-enhanced generation method based on long-term memory provided in the embodiments of the present application can be applied to an electronic device. Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of the present invention does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.
[0107] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, keys, a camera, a display screen, and a SIM card interface, etc.
[0108] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.
[0109] The processor may include one or more processing units. For example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices or integrated in one or more processors.
[0110] Among them, the processor can be the nerve center and command center of the electronic device. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0111] A memory can also be set in the processor to store instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can save the instructions or data that the processor has just used or recycled. If the processor needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor, and thus improves the system efficiency.
[0112] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor through the external memory interface to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.
[0113] The internal memory can be used to store computer-executable program code, and the computer-executable program code includes instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory can include a program storage area and a data storage area. The internal memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0114] The wireless communication function of the electronic device can be implemented by an antenna, a wireless communication module, a modulation and demodulation processor, a baseband processor, etc.
[0115] The wireless communication module can provide wireless communication solutions applied to the electronic device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0116] The electronic device can implement audio functions, etc. through an audio module, a speaker, a receiver, a microphone, a headphone jack, an application processor, etc.
[0117] The electronic device can implement a shooting function through an ISP, a camera, a video codec, a GPU, a display screen, an application processor, etc.
[0118] The electronic device can implement a display function through a GPU, a display screen, an application processor, etc.
[0119] The GPU is a microprocessor for image processing, connecting the display screen and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor can include one or more GPUs, which execute program instructions to generate or change display information.
[0120] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0121] The above electronic device implements the technical solution of the method for retrieval-augmented generation based on long-term memory in the theme name of this application, which uses a large language model to compress and semantically encode the content of the entire database to form a global memory; based on a query proposed by a user, generate a preliminary answer, and use the preliminary answer as retrieval information to retrieve relevant information from the database; based on the retrieved relevant information and the global memory, generate a final answer and provide it to the user, achieving the beneficial effects of improving the response quality and accuracy for complex queries and significantly enhancing the ability to process long texts and implicit information requirements.
[0122] In the storage medium provided by this application, there is a program product capable of implementing the method for retrieval-augmented generation based on long-term memory.
[0123] The method for retrieval-augmented generation based on long-term memory includes: using a large language model to compress and semantically encode the content of the entire database to form a global memory; based on a query proposed by a user, generate a preliminary answer, and use the preliminary answer as retrieval information to retrieve relevant information from the database; based on the retrieved relevant information and the global memory, generate a final answer and provide it to the user.
[0124] In some possible implementation manners, the theme name of this disclosure, the method and system for retrieval-augmented generation based on long-term memory, can be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of this disclosure described in the "Exemplary Method" section above in this specification.
[0125] The storage medium of this disclosure can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0126] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A retrieval-augmented generation method based on long-term memory, characterized in that, Including: Using a large language model, compress and semantically encode the content of the entire database to form a global memory; Based on the query proposed by the user, generate a preliminary answer, and use the preliminary answer as retrieval information to retrieve relevant information from the database; Based on the retrieved relevant information and the global memory, generate a final answer and provide it to the user; Among them, using a large language model, compress and semantically encode the content of the entire database to form a global memory, including: Read and aggregate the data in the database, and use a large language model based on Transformer to compress long texts into memory units through a multi-layer attention mechanism while retaining key semantic information to form a global memory, where the database includes a relational database and a non-relational database; Among them, process the original input tokens through a Transformer model to generate query, key, and value matrices, and calculate weights using the attention mechanism; After each context window, append memory tokens, and update the memory representation using the weight matrix of the memory tokens to form a global memory; Among them, the global memory is characterized as: ; wherein, is the i-th vector in the hidden state representation, is the attention weight corresponding to the i-th word in the attention weight matrix output by the attention mechanism; The method also reads and aggregates the data in the database, and uses a large language model based on Transformer to compress long texts into memory units through a multi-layer attention mechanism while retaining key semantic information to form a global memory, including: Introduce a global memory construction algorithm for constructing a global memory representation containing semantic information from the original long text data: Representation step of the original input tokens: The original input tokens are converted into query, key, and value matrices through the processing of the Transformer model for calculating the attention weights of the query, key, and value matrices; Calculation step of the attention mechanism: Use the softmax function to calculate the attention weights and obtain the weighted values using the attention weights; ; In the formula, is the dimension of the key vector; Initialization step of the memory tokens: After each context window, append a preset number of memory tokens and initialize another set of query, key, and value weight matrices for the memory tokens; Another weight matrix for a set of queries, keys, and values is represented as , and ; Update step of the memory tokens: Use the same attention mechanism as the original input tokens and use the weight matrix of the memory tokens to calculate the updated representation of the memory tokens; , , ; ; Memory formation step: After being processed by multiple layers of Transformer, the original input tokens are encoded into hidden states, and after the memory is formed, the key-value cache of the original tokens is discarded, where the hidden states include the hidden states of the original tokens and the hidden states of the memory tokens; Semantic compression step: Compress the long text into a global memory; Training step of the global memory construction algorithm: including a pre-training and a supervised fine-tuning phase. Among them, in the pre-training phase, the global memory construction algorithm uses a long context dataset to learn to form memories from the original context, and in the supervised fine-tuning phase, the global memory construction algorithm uses the data of a preset task to generate task cues based on the formed memories.
2. The retrieval-augmented generation method based on long-term memory according to claim 1, wherein, Based on the query proposed by the user, generate a preliminary answer, and use the preliminary answer as retrieval information to retrieve relevant information from the database, including: After receiving the query proposed by the user, parse the query, generate a draft of the preliminary answer according to the global memory, and use the draft of the preliminary answer as a retrieval clue to represent the information needs behind the user's query; Use a large language model and prompt engineering to evaluate the generated retrieval clue, and regenerate it if it does not meet the requirements to ensure the accuracy of the retrieval clue; Use the retrieval clue to retrieve the information fragment most relevant to the user's query in the database, and screen and integrate the retrieved information fragments to remove redundant and irrelevant content, forming an evidence set for subsequent answer generation.
3. The retrieval augmented generation method based on long-term memory according to claim 2, wherein Based on the retrieved relevant information and global memory, generate a final answer and provide it to the user, including: Use a large language model to fuse the query proposed by the user, the retrieved evidence set, and the global memory to generate an initial answer, and post-process the generated initial answer to generate a final answer.
4. The retrieval augmented generation method based on long-term memory according to claim 1, wherein Before generating a final answer based on the retrieved relevant information and global memory and providing it to the user, the retrieval enhanced generation method based on long-term memory further includes: Fine-tune the Transformer-based large language model based on the collected data to enable the Transformer-based large language model to better understand the user's query intent.
5. A retrieval-augmented generation system based on long-term memory for use in the retrieval-augmented generation method based on long-term memory according to any one of claims 1-4, characterized in that, Including: A global memory construction module for using a large language model to compress and semantically encode the content of the entire database to form global memory; A retrieval clue generation and information retrieval module for generating a preliminary answer based on the query proposed by the user, and using the preliminary answer as retrieval information to retrieve relevant information from the database; An answer generation module for generating a final answer based on the retrieved relevant information and global memory and providing it to the user.
6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the retrieval enhanced generation method based on long-term memory according to any one of claims 1 to 4.
7. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the retrieval enhanced generation method based on long-term memory according to any one of claims 1 to 4.