Historical record backtracking method based on LLM
By recording screen and audio information on the user's local device, and using large language models and RAG technology, users can trace back the history through natural language questions, solving massive information management and search problems, improving user experience and work efficiency.
Patent Information
- Application Number
- CN202510169056.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
Smart Images

Figure CN120104660A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of data retrieval and natural language processing, and in particular to a historical record backtracking method based on LLM. Background Art
[0002] LLM, the full name of Large Language Model, is an artificial intelligence algorithm based on deep learning. It can model natural language text by training a large amount of text data to learn the grammar, semantics and context information of the language. This model has a wide range of applications in the field of natural language processing, including but not limited to text generation, text classification, machine translation, sentiment analysis, etc.
[0003] RAG, the full name of Retrieval-Augmented Generation, is retrieval-augmented generation. RAG technology enhances the answering ability of the generation model by introducing external knowledge bases or document bases, enabling it to generate more detailed and accurate answers.
[0004] With the rapid development of information technology, we have entered the information age. Users perform a large number of operations on computers every day, including but not limited to browsing web pages, editing documents, watching videos, attending meetings, and chatting with instant messengers. However, due to the huge amount of information, users often encounter the following problems: 1. Information loss: After browsing a large amount of information, users often find it difficult to find specific content they have seen before. 2. Emergencies: Accidentally closed programs, withdrawn messages, or temporarily lost important content are often difficult to recover. 3. Memory confusion: Users are unable to efficiently organize and review past experiences and information, resulting in the forgetting of important information and difficulty in recalling specific details. Therefore, for users who need to frequently recall and find information, as time goes by, the user accumulates more and more operation records. How to efficiently manage and find these historical records becomes an important issue. Therefore, it is necessary to design a historical record backtracking method based on LLM. Summary of the invention
[0005] The present invention aims to solve the problems of memory management and information retrieval. With the rapid growth of digital information, users often find it difficult to accurately recall or find important content they have seen when faced with massive amounts of information. In order to solve this problem, the present invention proposes a historical record backtracking method based on a large language model (LLM). The system can fully record screen content and audio information on the user's local device, and provide basic retrieval functions to help users quickly trace back past memory fragments. Furthermore, the present invention allows users to interact with their own timelines by introducing LLM and RAG technologies, allowing users to trace back historical records through natural language questions and dialogues, and to search, summarize and review past operations or information, further enhancing the intelligence and ease of use of the system. This innovative interactive mode allows users to access and manage their digital memories more intuitively, improving the availability of information and user experience.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] A historical record backtracking method based on LLM includes the following steps:
[0008] 1. Data Collection and Storage
[0009] First, various data are obtained in the local computer, and then the text information is obtained by combining OCR technology and ASR technology. Finally, the required information is stored in the database, which specifically includes the following aspects:
[0010] 1. Screenshots. Regularly take screenshots of the user's computer screen to capture the user's operations when using the computer, such as browsing the web, editing documents, and using software.
[0011] 2. Video recording: Directly record the video of the user's operation and process it into a temporary video with a specified number of frames, while extracting key video picture frames.
[0012] 3. Optical Character Recognition (OCR). The complete OCR task can be viewed as two consecutive tasks: text detection and text recognition. Whether performed simultaneously or separately, each task corresponds to a deep learning architecture. Using OCR technology, you can obtain the OCR results, i.e., text information, in screenshots and image frames, and finally save the required information into a temporary data file.
[0013] 4. Automatic Speech Recognition (ASR): Start audio recording to capture voice signals, combine with the locally deployed automatic speech recognition model, transcribe the voice into text, and then save the required data to a temporary data file.
[0014] 5. After the task is completed, the generated temporary data files will be read and converted into a unified format and stored in the database.
[0015] 2. Data Processing
[0016] This part is responsible for processing the collected multimodal data information. This part will first deduplicate and clean the collected data, then convert the collected screenshots into videos, and finally embed and index the video image frames, OCR results, and speech transcription text. It specifically includes the following parts:
[0017] 1. Text merging and deduplication. This is performed synchronously during the data collection process. By setting a similarity threshold, the currently acquired data can be compared with the stored data to avoid repeated storage of similar content and reduce data redundancy.
[0018] 2. Text cleaning and optimization. Clean and optimize the collected text data, delete unreadable characters, garbled characters, and redundant content caused by hallucinations, and remove unnecessary information such as spoken language.
[0019] 3. Screenshot conversion: Regularly convert the collected screenshots into video files.
[0020] 4. Embedding index. Convert the video image frame and the OCR results and speech transcription text in the database into vectors and add them to the vector database. This operation is also the core of RAG technology. It includes the following steps:
[0021] (1) Image embedding. Read the video file and extract the corresponding frame images based on the frame information recorded in the database. These frame images are embedded into vectors through a specific Embedding model, and the vectors, the unique identifier of the corresponding data collection phase database, and related metadata are stored in the image vector database.
[0022] (2) Text embedding. Read the OCR text and speech transcription text in the database, embed these texts into vectors through a specific Embedding model, and store the vectors, the unique identifier of the corresponding database in the data collection phase, and related metadata in the text vector database.
[0023] (3) The vector is subjected to L2 normalization. L2 normalization is to normalize the length (or norm) of the vector to 1, so that the vector information is retained while making the similarity calculation and retrieval more stable and effective.
[0024] 3. LLM fine-tuning training
[0025] LLM fine-tuning refers to further training the pre-trained LLM using a smaller and more specific data set to enhance its performance in specific domain tasks (such as understanding computer industry terminology). In order to make the subsequent timeline dialogues better adapt to specific application scenarios according to user needs, the pre-trained LLM is optimized using fine-tuning technology. The specific process includes the following parts:
[0026] 1. Load the pre-trained LLM model and obtain it from the target domain (such as user interaction records, industry documents, domain knowledge base)
[0027] The model can be based on a mainstream large language model (such as a GPT-like model or other equivalent models).
[0028] 2. Set fine-tuning parameters, including learning rate, batch size, number of gradient accumulation steps, number of fine-tuning layers, optimizer, etc.
[0029] 3. Create a fine-tuner, select the corresponding fine-tuning strategy based on the target task, and initialize its weights, embedding layer, and objective function.
[0030] 4. Create a data loader and define the data loader to load corpus data in batches.
[0031] 5. Integrate the model with the fine-tuner, use the data loader to feed the training data in batches, and perform multiple rounds of iterative training. During the training process, the learning rate can be dynamically adjusted to prevent the model from overfitting.
[0032] 4. Timeline Dialogue
[0033] The chatbot can realize the dialogue with the historical timeline, answer the user's general questions, and answer personalized questions based on specific data and personal data and specific time. In addition, RAG technology and memory function are introduced to enhance the personalized interactive experience. In order to realize these functions, the present invention improves the conventional chatbot, which specifically includes the following parts:
[0034] 1. Event processing. Capture and respond to user operations, process keyboard key presses, mouse clicks and other events, such as activating and deactivating the chatbot, pressing Enter to initiate a query, etc., and store the input obtained from the text box in the query queue.
[0035] 2. Dialogue processing. Responsible for interacting with the pre-fine-tuned LLM, using a technical framework design based on RAG and memory module (optional) for optimization, and providing users with more accurate, personalized and context-related dialogue responses by combining LLM, external knowledge base retrieval function and memory function. Its functional steps include but are not limited to the following:
[0036] (1) Build a memory module. The memory module (such as Mem0 or other implementations) stores and manages user preferences and historical interaction data through an intelligent memory layer, supporting dynamic learning and continuous optimization. This module can be based on memory or persistent storage to ensure that user session data is not lost during restart and provide continuous context support.
[0037] (2) Create a historical conversation list to store chat history records, including questions entered by the user, corresponding answers, and query-related image frame information.
[0038] (3) Receive user input. The system receives the query content entered by the user, parses it and stores it in the historical conversation record.
[0039] (4) Dynamic memory update. After receiving user input, the memory module automatically extracts relevant information, such as the user's points of interest and contextual preferences, and dynamically updates the stored data when necessary.
[0040] (5) Vector retrieval. The core step of RAG is to embed the input question into a vector and search it in the vector database, sort the record vectors from high to low according to similarity, and return the k records that are most similar to the input. The retrieval results include not only the text content but also the relevant metadata.
[0041] (6) Parsing and processing time information in questions. When dealing with questions involving time ranges or time limits, the system needs to accurately parse the time information in the user's question and make precise adjustments to the answer accordingly, including extracting the clear time range and corresponding time point from the question.
[0042] (7) Generate dialogue prompts. Prompt engineering is a key step in generating high-quality dialogues in this system. Highly personalized dialogue prompts can be generated by combining the user's current input, metadata of related text, historical dialogue records, and previously stored memory data. The main steps are as follows:
[0043] (7.1) Integrate search results. Organize the external knowledge fragments related to the user input extracted by the search and sort them by priority.
[0044] (7.2) Obtain memory data. Retrieve user preference information through the memory module to ensure that the generated prompts are consistent with the user context.
[0045] (7.3) Extract conversation history information. Extract the most recent conversation information from the historical conversation list, giving priority to a few recent conversation messages, which can be adjusted according to needs.
[0046] (7.4) Build prompt templates. Customize prompt templates based on target tasks and user needs, comprehensively consider user input, retrieved text and its metadata, historical records, memory data, etc., and generate personalized prompts to guide the model to generate responses.
[0047] (8) Model response generation. The generated prompt is passed to the LLM for processing, and the model generates a response based on the context, memory information, and retrieval results. Finally, the system adds the generated response content, related image frame ID, and audio ID to the historical conversation list.
[0048] 3. Conversation display: Display each question and answer in the chat history, and then display the relevant pictures or audio links below the answer.
[0049] The beneficial effects of the present invention are:
[0050] ① The present invention proposes a historical record backtracking method based on LLM. The system can record and manage the user's digital memory, and help the user retrieve and organize past memory fragments through retrieval function and timeline dialogue function.
[0051] ② The present invention introduces LLM and RAG technologies to achieve intelligent interaction between users and timelines. Users can retrieve historical records and obtain specific information through natural language dialogues. It can also help users summarize past events and meetings attended at a certain time. This intelligent interaction enhances the usability and flexibility of the system and significantly improves user experience and work efficiency.
[0052] ③ In the prior art, AI dialogue systems usually rely on predefined rules or standardized models for dialogue responses. However, these systems have limitations when dealing with long-term dialogues and personalized user interactions. To solve this problem, the present invention additionally introduces a memory module, which enables the system to dynamically adjust the dialogue style and corresponding content, and generate more natural dialogue content that meets user needs based on the user's personalized historical data and memory.
[0053] ④ The present invention runs on the user's local device, effectively protecting the privacy and security of data and avoiding the potential safety hazards brought by external storage. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a system structure diagram of the present invention;
[0055] Figure 2 is a flow chart of the method of the present invention;
[0056] Figure 3 This is a schematic diagram of the structure of the timeline dialogue robot of the present invention. DETAILED DESCRIPTION
[0057] The present invention will be further described below in conjunction with the accompanying drawings.
[0058] like Figure 1 and Figure 2 As shown, a historical record backtracking method based on LLM includes the following steps:
[0059] Step 1: Collect and store local personal data. First, obtain various data in the local computer, and then combine OCR technology and ASR technology to obtain text information. Specifically, it includes the following aspects:
[0060] Step 1.1, screenshot. You can use MSS to regularly take screenshots of the user's computer screen to capture the user's operations. These screenshots include the user's operations in scenarios such as browsing web pages, editing documents, and using software. MSS is a lightweight cross-platform screen capture library that supports efficient acquisition of screen content. It has the characteristics of fast speed and low resource usage. It supports multiple modes such as custom foreground window screenshots, single monitor screenshots, and full-range monitor screenshots.
[0061] Step 1.2, video recording. You can directly record the video through FFmpeg, then convert the input video into a 2fps temporary video and extract the image frames of the video. FFmpeg is a set of open source tools that can be used to record, convert digital audio and video, and convert them into streams.
[0062] Step 1.3, Optical Character Recognition (OCR). The complete optical character recognition task can be regarded as two consecutive tasks: text detection and text recognition. Whether performed simultaneously or separately, each task corresponds to a deep learning architecture. Using OCR technology, we can obtain the OCR results in screenshots and image frames, that is, the text information of screenshots and image frames. Finally, the file type, file path, OCR result, window title, timestamp and other information are saved to a temporary JSON database file.
[0063] Step 1.4, Automatic Speech Recognition (ASR). Start audio recording to capture speech signals, combine with the locally deployed automatic speech recognition model, transcribe the speech into text, and optimize it with the Voice Activity Detection (VAD) algorithm, whose main function is to distinguish the voice signal from various background noise signals, thereby reducing the amount of calculation and improving the accuracy of transcription. Finally, save the file type, file path, speech transcription results, timestamp and other information into a temporary JSON database file.
[0064] Step 1.5, at the end of recording, read the indexed temporary JSON file and convert it into Pandas DataFrame format and submit it to the SQLite database.
[0065] Step 2: Process the collected multimodal data information. First, remove duplicates and clean the collected data, then convert the collected screenshots into videos, and then embed and index the video image frames, OCR results, and audio text. Specifically, it includes the following parts:
[0066] Step 2.1, text merging and deduplication. This is done simultaneously during the data collection process. By setting a similarity threshold, the currently acquired screenshots and OCR results are compared with the previous screenshots and OCR results to avoid repeated storage of similar content and reduce data redundancy.
[0067] Step 2.2, text cleaning and optimization. Clean and optimize the collected text data, delete unreadable characters, garbled characters, and redundant content caused by hallucinations, and remove unnecessary information such as spoken language.
[0068] Step 2.3, screenshot conversion. Convert screenshots to video files regularly. Please refer to the following steps for details:
[0069] (1) Parse the file name of the screenshot file, extract the date and time string and convert it into a Unix timestamp;
[0070] (2) using the screenshot information in the temporary database to calculate the relative time of each screenshot in the video, specifically by calculating the time difference of each screenshot relative to the first screenshot, and creating a mapping list of timestamps and file paths;
[0071] (3) Read the first screenshot to obtain the resolution of the video to unify the size of the screenshots and ensure that all screenshots have the same resolution;
[0072] (4) Write the corresponding screenshots into the video file according to the timestamp of each screenshot and the set duration.
[0073] Step 2.4, embedding index. Convert the video image frames and OCR results in the SQLite database and the speech transcription text into vectors and add them to the vector database. This can be used for vector image retrieval and semantic search. This step is also the core of RAG technology. For details, please refer to the following steps:
[0074] (1) Load the pre-trained Embedding model and the corresponding processor.
[0075] (2) Read the video file and extract the corresponding frame images according to the frame information recorded in the database. These frame images are embedded into vectors through a specific Embedding model, and the vectors, the unique identifier of the corresponding data collection phase database, and the relevant metadata are stored in the image vector database.
[0076] (3) Text embedding. Read the OCR text and speech transcription text in the database, embed these texts into vectors through a specific Embedding model, and store the vectors, the unique identifier of the corresponding database in the data collection phase, and related metadata in the text vector database.
[0077] (4) Perform L2 normalization on the vector. L2 normalization of the vector is to convert the input vector into a unit vector by calculating the L2 norm of the vector and dividing each component by the norm while keeping its direction unchanged.
[0078] Step 3: Fine-tune the pre-trained LLM. This allows the timeline dialogue part of the system to better adapt to specific application scenarios according to user needs. The specific process of using fine-tuning technology to optimize the pre-trained LLM includes the following steps:
[0079] Step 3.1, load the pre-trained model and the high-quality corpus dataset obtained from the target domain (such as user interaction records, industry documents, domain knowledge base).
[0080] Step 3.2, set fine-tuning parameters, where fine-tuning parameters include learning rate, batch size, number of gradient accumulation steps, number of fine-tuning layers, optimizer, etc.
[0081] Step 3.3, create a fine-tuner, select the corresponding fine-tuning strategy based on the target task, and initialize its weights, embedding layer, and objective function.
[0082] Step 3.4, create a data loader and define the data loader to load corpus data in batches.
[0083] In step 3.5, integrate the model with the fine-tuner, use the data loader to feed the training data in batches, and perform multiple rounds of iterative training. During the training process, the learning rate can be dynamically adjusted to prevent the model from overfitting.
[0084] Step 4, realize the dialogue with the historical timeline, which is the core part of this system. It is like a digital assistant that can ask you anything you see, say or hear. It can answer general questions like ChatGPT, and can also answer questions about specific data. In addition, RAG technology and memory function are introduced to increase personalized interaction. In order to realize the above functions, the present invention improves the conventional chat robot, see Figure 3, which specifically includes the following parts:
[0085] Step 4.1, event processing. Responsible for capturing and responding to user operations, processing keyboard key presses, mouse clicks and other events, such as activating and deactivating the chatbot, pressing Enter to initiate a query, etc., and storing the input obtained from the text box in the query queue.
[0086] Step 4.2, dialogue processing. Responsible for interacting with LLM, using a technical framework design based on RAG and memory module (optional) for optimization, by combining LLM, external knowledge base retrieval function and memory function, to provide users with more accurate, personalized and context-related dialogue responses. Its functional steps include but are not limited to the following:
[0087] (1) Load a pre-fine-tuned LLM model, which can be based on a mainstream large language model architecture (such as a GPT-like model or other equivalent models).
[0088] (2) Build a memory module. The memory module (such as Mem0 or other implementations) stores and manages user preferences and historical interaction data through an intelligent memory layer, supporting dynamic learning and continuous optimization. This module can be based on memory or persistent storage to ensure that user session data is not lost during restart and provide continuous context support.
[0089] (3) Create a historical conversation list to store chat history records, including questions entered by the user, corresponding answers, and query-related image frame information.
[0090] (4) Receive user input. The system receives the query content entered by the user, parses it and stores it in the historical conversation record.
[0091] (5) Dynamic memory update. After receiving user input, the memory module automatically extracts relevant information, such as the user's points of interest and contextual preferences, and dynamically updates the stored data when necessary.
[0092] (6) Vector retrieval. The core step of RAG is to embed the input question into a vector and search it in the vector database, sort the record vectors from high to low according to similarity, and return the k records that are most similar to the input. The retrieval results include not only the text content but also the relevant metadata.
[0093] (7) Analysis and processing of time information in questions. When dealing with questions involving time ranges or time limits, the system needs to accurately analyze the time information in the user's question and make precise adjustments to the answer accordingly, including extracting a clear time range and corresponding time point from the question. The main ideas are as follows:
[0094] (7.1) Use dateparser or related NLP libraries to parse the time information in the question and convert it into Unix timestamp format. dateparser is a powerful date parsing library that supports multi-language and multi-format time expression parsing and can automatically identify relative and absolute time.
[0095] (7.2) For a single time, a reasonable time range can be set. For a specific time, for example, 9:00 is the center, and it can be extended by 1-2 minutes. For fuzzy time, such as a year, month, week, or day, the time range can be set according to the corresponding situation.
[0096] (7.3) For multiple times, regular expressions can be used to assist in parsing. For a time range, such as August-September, the unix timestamps of the two times before and after can be used as the start time and end time respectively. For multiple independent times, they can be processed as single times, and the start time and end time can be stored separately.
[0097] (7.4) Searching in vector databases, such as the Chroma vector database, can use the where parameter to filter documents by metadata, specifying $gte and $lte to limit the time range, so that only records that match the time period are obtained. For multiple time ranges, multiple filters can be combined using $or.
[0098] (8) Generate dialogue prompts. Prompt engineering is a key step in generating high-quality dialogues in this system. Highly personalized dialogue prompts can be generated by combining the user's current input, metadata of related text, historical dialogue records, and previously stored memory data. The main steps are as follows:
[0099] (8.1) Integrate search results. Organize the external knowledge fragments related to the user input extracted by the search and sort them by priority.
[0100] (8.2) Obtain memory data. Retrieve user preference information through the memory module to ensure that the generated prompts are consistent with the user context.
[0101] (8.3) Extract conversation history information. Extract the most recent conversation information from the historical conversation list, giving priority to a few recent conversation messages, which can be adjusted according to needs.
[0102] (8.4) Construct Prompt template. Customize prompt template according to target tasks and user needs, comprehensively consider user input, retrieved text and its metadata, historical records, memory data, etc., and generate personalized prompts to guide the model to generate responses.
[0103] (9) Model response generation. The generated prompt is passed to the LLM for processing, and the model generates a response based on the context, memory information, and retrieval results. Finally, the system adds the generated response content, related image frame ID, and audio ID to the historical conversation list.
[0104] Step 4.3, dialogue display. Display each question and answer in the chat history, and then display the history-related pictures or audio links below the answers.
[0105] The basic principles, main features and advantages of the present invention are shown and described above. For those skilled in the art, it should be understood that the present invention is not limited to the above embodiments, and these embodiments and descriptions are only used to explain the principles of the present invention. Various changes and improvements can be made to the present invention without departing from the spirit and scope of the present invention, and all these changes and improvements belong to the protection scope of the present invention. The scope of protection claimed in the present invention is defined by the attached claims and their equivalents.
Claims
1. A historical record backtracking method based on LLM, characterized in that The implementation steps are: (1) Collect and obtain various types of data on the local computer, including screenshots, videos, voice and other user operation records, and then combine OCR technology and ASR technology to extract text information from pictures and voices, and store the generated results and required information in a unified format in the database; (2) Processing the collected multimodal data, including deduplication and cleaning of text data, converting the collected screenshots into video files, and finally embedding the text data and video image frames into vectors and storing them in a vector database; (3) Fine-tune the pre-trained LLM using domain-specific datasets; (4) First, user input is captured and stored in the query queue. Then, personalized responses are generated by interacting with the LLM, combining the memory module and RAG technology, and displaying the answers and related images or audio links.
2. The historical record backtracking method based on LLM according to claim 1 is characterized in that This method obtains various types of data in the local computer, and then combines OCR technology and ASR technology to obtain text information: (1) Regularly take screenshots of the user's computer screen to capture the user's operations during computer use, such as browsing web pages, editing documents, and using software; (2) Directly record the video of the user's operation and process it into a temporary video with a specified number of frames, while extracting key video picture frames; (3) Obtain text information from screenshots and image frames through OCR technology, and then save the required information into a temporary data file; (4) Start audio recording to capture voice signals, combine with the locally deployed automatic speech recognition model, transcribe the voice into text, and then save the required data into a temporary data file; (5) After the task is completed, the generated temporary data files will be read and converted into a unified format and stored in the database.
3. The historical record backtracking method based on LLM according to claim 1 is characterized in that This method processes the collected multimodal data: (1) Setting a similarity threshold, comparing the currently acquired data with the stored data, deleting duplicates and saving similar content; (2) Delete unreadable characters, garbled characters, and redundant content caused by hallucination problems, and remove unnecessary information such as spoken language; (3) Regularly convert the collected screenshots into video files; (4) Convert the video image frames, OCR results in the database, and speech transcription text into vectors and add them to the vector database.
4. A historical record backtracking method based on LLM according to claim 1 or 3, characterized in that Convert image video frames, OCR text and speech text into vectors and store them in the corresponding vector database: (1) Read the video file and extract the corresponding frame images according to the frame information recorded in the database. These frame images are embedded into vectors through a specific Embedding model, and the vectors, the unique identifier of the corresponding data collection phase database, and the relevant metadata are stored in the image vector database; (2) Read the OCR text and speech transcription text in the database, embed these texts into vectors through a specific Embedding model, and store the vectors, the unique identifier of the corresponding database in the data collection phase, and related metadata in the text vector database; (3) Ensure that the vector is L2 normalized.
5. The historical record backtracking method based on LLM according to claim 1 is characterized in that This method fine-tunes the pre-trained LLM: (1) Loading pre-trained models and high-quality corpus datasets obtained from the target domain (such as user interaction records, industry documents, domain knowledge bases, etc.); (2) Setting fine-tuning parameters, including learning rate, batch size, number of gradient accumulation steps, number of fine-tuning layers, optimizer, etc. (3) Create a fine-tuner, select the corresponding fine-tuning strategy based on the target task, and initialize its weights, embedding layer, and objective function; (4) Create a data loader and define the data loader to load corpus data in batches; (5) Integrate the model with the fine-tuner, use the data loader to feed the training data in batches, and perform multiple rounds of iterative training.
6. The historical record backtracking method based on LLM according to claim 1 is characterized in that This method enables a dialogue with the historical timeline: (1) Capture and respond to user operations (such as keyboard and mouse events), handle the start and stop of the chatbot, obtain user input and store it in the query queue; (2) Retrieve relevant historical data based on user input, generate personalized dialogue prompts, and generate answers corresponding to user questions through the LLM model; (3) Display chat history, including questions, answers, and related images or audio links.
7. The LLM-based historical record backtracking method according to claim 1 or 6, characterized in that This method retrieves relevant historical data based on user input, generates personalized dialogue prompts, and generates answers corresponding to user questions through the LLM model: (1) Build a memory module to store and manage user historical personalized data; (2) Create a conversation history list to store the user's historical input and answers; (3) Dynamically update user preferences, points of interest, and other information in the memory module based on the user's current input; (4) Embed the input question into a vector and search it in the vector database, sort the record vectors from high to low according to similarity, and return the k records that are most similar to the input; (5) Parse the time information in the input question and filter out the records related to the time range in the vector database; (6) Generate highly personalized conversation prompts by combining the user’s current input, metadata of related text, historical conversation records, and previously stored memory data; (7) Submit the generated dialogue prompt to the LLM model, obtain the answer, polish it and generate the final answer, and then add the answer generated by the model to the dialogue history list.
Citation Information
Cited By
Digital media resource intelligent retrieval and management system based on multi-modal interaction
CN121278155A