Historical dialogue screening method and device for dialogue system
By calculating the multi-dimensional similarity between the current question and historical dialogue data and filtering related content, the traditional method omits cross-time, cross-topic content and difficult to deal with dialogue without obvious keywords, achieving a more accurate and personalized historical dialogue screening effect.
Patent Information
- Application Number
- CN202510163965.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional historical dialogue screening methods may miss content across time and topics, making it difficult to deal with dialogues without obvious keywords, resulting in the screening results deviating from user expectations and poor screening of dialogues without obvious intentions.
By calculating the multi-dimensional similarity between the current question and the historical conversation data, including the weighted fusion of semantic similarity, keyword similarity, topic similarity, and context similarity, and filtering out the historical conversation data that is most relevant to the current question based on the preset dialogue filtering threshold.
It realizes the more accurate screening of historical dialogue records that are most relevant to the current question, adapting to different conversation scenarios, ensuring the timeliness and relevance of the screening results, and automatically optimizing the filter parameters through user feedback to enhance personalization and accuracy.
Smart Images

Figure CN120104735A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a method and device for screening historical dialogues of a dialogue system. Background Art
[0002] Historical dialogue screening is of great significance to multi-round dialogue systems: it can improve dialogue efficiency, avoid processing lengthy irrelevant information, enable the system to quickly focus on current issues in complex multi-round dialogue scenarios, and improve response speed and experience. It can also enhance the accuracy of answers, accurately grasp user intentions and needs, and avoid repeated or off-point answers in scenarios such as knowledge questions and answers and recommendation systems. At the same time, it helps to optimize dialogue strategies. Developers can use this to understand user behavior and problems, optimize speech and guidance strategies, and achieve personalized customization.
[0003] Traditional historical conversation screening methods mainly include screening based on time windows, screening based on keyword matching, and screening based on intent matching, as follows:
[0004] The time window-based filtering method sets a fixed time range, such as the past 5 minutes or the last 10 rounds of conversation, and only retains the conversation history within this time window. This method is suitable for scenarios where the conversation topic changes quickly or real-time requirements are high. For example, in online customer service, users may ask questions about different products in a short period of time. Only focusing on recent conversations can quickly understand the key points of the current problem. However, for long-term cross-topic conversation scenarios, this method may ignore content that is still important despite the long time.
[0005] The keyword matching-based screening method extracts keywords from the current conversation and searches for content that contains these keywords or is highly related to them in historical conversations. For example, in a travel consultation conversation, if the user currently mentions "Paris attractions", the system will filter out conversations that mention related words such as "Eiffel Tower" and "Louvre" from the historical records. This method can better focus on the current topic, but its effect depends on the accuracy of keyword extraction. If there are no significant keywords in the current conversation, the screening effect may not be ideal.
[0006] The filtering method based on intent matching uses intent recognition technology to determine the user's intent in each round of conversation, and selects those conversation contents that are the same or similar to the current intent when filtering the historical records. This method performs well in topic-focused conversations, but its performance depends on the accuracy of the intent recognition model, and the filtering effect is poor for conversations that lack obvious intent. Summary of the invention
[0007] The technical problem to be solved by the present invention is: to propose a method and device for screening historical conversations in a dialogue system, so as to solve the problems that the historical conversation screening schemes in traditional technologies may miss content across time and topics, have difficulty in processing conversations without obvious keywords, resulting in screening results deviating from user expectations, and have poor screening effects on conversations lacking obvious intentions.
[0008] The technical solution adopted by the present invention to solve the above technical problems is:
[0009] In one aspect, the present invention provides a method for screening historical dialogues in a dialogue system, comprising the following steps:
[0010] Store historical conversation data between the user and the conversation system based on the time and topic of the conversation;
[0011] When screening historical conversations, first calculate the similarity between the current question and the stored historical conversation data, which includes a weighted fusion of semantic similarity, keyword similarity, topic similarity, and context similarity;
[0012] Then, based on a preset conversation screening threshold, combined with the similarity calculation result between the current question and the stored historical conversation data, the historical conversation data most relevant to the current question is screened out from the stored historical conversation data; the conversation screening threshold includes an adjustable similarity threshold and time window.
[0013] Furthermore, storing historical conversation data between the user and the conversation system based on the time and topic of the conversation includes:
[0014] Calculate the topic of each round of dialogue between the user and the dialogue system;
[0015] Based on the time and topic of the conversation, the conversation content is stored in multi-level classification.
[0016] Furthermore, the method for calculating the similarity between the current question and the stored historical dialogue data includes:
[0017] Calculate the semantic similarity between the current question and the historical dialogue through the semantic similarity model;
[0018] Calculate the keyword similarity between the current question and the historical dialogue through the keyword similarity algorithm;
[0019] Calculate the topic similarity between the current question and the historical dialogue through the topic similarity algorithm;
[0020] Calculate the context similarity between the current question and the historical dialogue through the context similarity algorithm;
[0021] The semantic similarity, keyword similarity, topic similarity, and context similarity between the current question and the historical conversation are weightedly fused to obtain the similarity between the current question and the historical conversation.
[0022] Furthermore, the keyword similarity between the current question and the historical dialogue is calculated by using a keyword similarity algorithm, including:
[0023] Extract keywords from the current question and historical conversations respectively;
[0024] The extracted keywords are converted into vector representations using a vectorization model;
[0025] The similarity calculation method is used to calculate the similarity between the vector representation of the keywords in the current question and the vector representation of the keywords in the historical conversation.
[0026] Furthermore, the topic similarity between the current question and the historical dialogue is calculated by using a topic similarity algorithm, including:
[0027] Use the topic model to calculate the topic words of the current question;
[0028] Using a vector model to convert the subject word of the current question into a vector representation;
[0029] The similarity calculation method is used to calculate the similarity between the topic word vector of the current question and the topic vector of the historical dialogue.
[0030] Furthermore, the calculating of the context similarity between the current question and the historical dialogue by using a context similarity algorithm includes:
[0031] constructing prompt templates for assessing conversational continuity and cohesion;
[0032] Fill the current question and historical conversation content as the input of the large language model;
[0033] A large language model is used to generate contextual similarity between the current question and historical conversations.
[0034] Furthermore, the method for setting the conversation screening threshold includes:
[0035] Set an adjustable similarity threshold;
[0036] Set adjustable time windows;
[0037] Automatically optimize the similarity threshold and time window based on user feedback data.
[0038] Furthermore, the method for automatically optimizing the similarity threshold and time window according to user feedback data includes:
[0039] Construct a prompt word template for detecting the relevance of the current question and the historical dialogue;
[0040] Based on the current question and historical conversations, the relevance detection prompt word template is used to guide the large language model to output the relevance detection results;
[0041] According to the relevance test results, determine the proportion of historical conversations that are relevant and irrelevant to the current question;
[0042] If the relevant ratio exceeds the set threshold, the similarity threshold is increased and the time window is reduced;
[0043] If the proportion of irrelevant data exceeds the set threshold, the similarity threshold is reduced and the time window is expanded.
[0044] Furthermore, the method of selecting the historical conversation data most relevant to the current question from the stored historical conversation data includes:
[0045] According to the time window, select the historical dialogue within the time window before the current question;
[0046] Based on the similarity threshold, select the top n historical conversations with the highest similarity from the historical conversations outside the time window;
[0047] The historical dialogues within the time window are merged with the selected Top n rounds of historical dialogues to obtain the historical dialogues most relevant to the current question.
[0048] On the other hand, the present invention also provides a device for screening historical dialogues in a dialogue system, comprising:
[0049] A storage module, used to store historical conversation data between the user and the conversation system based on the time and topic of the conversation;
[0050] A configuration module, used to configure a conversation screening threshold and automatically optimize the conversation screening threshold based on user feedback data; the conversation screening threshold includes an adjustable similarity threshold and a time window;
[0051] A similarity calculation module is used to calculate the similarity between the current question and the stored historical dialogue data, wherein the similarity includes a weighted fusion of semantic similarity, keyword similarity, topic similarity and context similarity;
[0052] The screening module is used to screen out the historical dialogue data most relevant to the current question from the stored historical dialogue data based on a preset dialogue screening threshold and the similarity calculation result between the current question and the stored historical dialogue data.
[0053] The beneficial effects of the present invention are:
[0054] (1) By combining multiple similarity calculation methods such as semantic similarity, keyword similarity, topic similarity, and context similarity, we can more comprehensively evaluate the relevance of the current question and historical conversations, thereby more accurately screening out the historical conversation records that are most relevant to the current question.
[0055] (2) Dynamic dialogue screening thresholds, including adjustable similarity thresholds and time windows, can be used to flexibly adjust screening parameters according to different dialogue scenarios to ensure the timeliness and relevance of screening results.
[0056] (3) Automatically optimizing the similarity threshold and time window based on user feedback data can enhance the personalization and accuracy of screening results.
[0057] (4) Multi-level classification and storage of historical conversations based on time and topic can effectively improve the retrieval efficiency of historical conversations. This enables the system to quickly locate and retrieve relevant content when processing large-scale historical conversation data. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 This is a flow chart of a method for screening historical conversations in a conversation system in Embodiment 1 of the present invention;
[0059] Figure 2 This is a structural diagram of the historical dialogue screening device of the dialogue system in Example 2 of the present invention. DETAILED DESCRIPTION
[0060] The present invention aims to provide a method and device for filtering historical conversations in a dialogue system, so as to solve the problems that the historical conversation filtering schemes in the traditional technology may miss the content across time and across topics, have difficulty in processing conversations without obvious keywords, resulting in the screening results deviating from the user's expectations, and have poor filtering effects on conversations lacking obvious intentions. The core idea is: in view of the problem that the time window-based filtering schemes in the traditional technology may miss the content across time and across topics, when filtering historical information, the present invention not only considers the historical conversations within the time window, but also filters out the top Top Conversations with the highest similarity from the historical conversations outside the time window. n rounds of historical dialogues are collected and merged with the dialogues in the time window as the final output, thereby avoiding missing important dialogue content across time and topics; in view of the problem that the screening method based on keyword matching in traditional technology relies on the accuracy of keyword extraction and is difficult to effectively handle dialogues without obvious keywords, the present invention not only considers keyword similarity when calculating the similarity between the current question and the historical dialogue, but also considers semantic similarity, topic similarity and context similarity, and then fuses them. The use of this multi-dimensional similarity fusion calculation method reduces the dependence on keyword extraction, can effectively handle dialogues without obvious keywords, and improves the accuracy and adaptability of the screening results; in view of the problem that the screening method based on intent matching in traditional technology is highly dependent on the accuracy of the intent recognition model and has poor effect on dialogues without clear intent, the present invention adopts a multi-dimensional similarity fusion calculation method, so that the screening of historical dialogues does not completely rely on intent recognition, and can also achieve better screening effects in dialogue scenarios without clear intent, thereby improving the system's adaptability to complex dialogue scenarios.
[0061] In addition, when setting the conversation screening threshold, the present invention adopts a time-dynamic threshold rather than a fixed threshold, which includes an adjustable similarity threshold and time window, and can flexibly adjust the screening parameters according to different conversation scenarios to ensure the timeliness and relevance of the screening results; and, it can also automatically optimize the similarity threshold and time window based on user feedback data on a regular basis, which can enhance the personalization and accuracy of the screening results.
[0062] The scheme of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0063] Example 1
[0064] This embodiment provides a method for filtering historical dialogues in a dialogue system. Figure 1 , including the following steps:
[0065] 1. Storing historical conversations
[0066] In this step, the historical conversation data between the user and the conversation system is stored based on the time and topic of the conversation.
[0067] In an exemplary embodiment, first, traditional topic models including but not limited to LDA, LSA, and large language models such as Llama, Qwen, and GPT-4 are used to calculate the potential topics in each round of historical conversations to ensure accurate identification of topics in conversations and track the evolution of topics; then, based on the time information and topic information of the conversation, the conversation content is classified and stored in a multi-level manner.
[0068] In the above process, the topic mining of the traditional topic model and the large language model can not only efficiently generate highly interpretable topics, but also deeply understand complex semantics and achieve more accurate topic calculation. Multi-level classification and storage of conversation content can effectively improve the retrieval efficiency of historical conversations.
[0069] Specifically, in each round of dialogue, the traditional topic model and the large language model are used to calculate the topic of the current question and the system response. For example, when a user asks "How many degrees is it today", the topic is "weather". Then, the calculation results of the traditional model and the large language model are merged into the final topic. Finally, it is determined whether the topic of the user's question and the system response are consistent. If they are consistent, the current question, the topic of the current question, the system response, the topic of the current question, and the current dialogue time are stored; if they are inconsistent, they are not saved.
[0070] 2. Screening Threshold Configuration
[0071] In this step, a dynamic conversation screening threshold is configured, and the conversation screening threshold is automatically optimized regularly based on user feedback, wherein the conversation screening threshold includes an adjustable similarity threshold and time window.
[0072] In an exemplary embodiment, the method for configuring the screening threshold includes:
[0073] Set an adjustable similarity threshold to ensure that the filtered historical conversations are sufficiently relevant to the current question;
[0074] Set the time window to select the historical conversations within the time window before the current question;
[0075] Based on user feedback data, the similarity threshold and time window are automatically optimized regularly to enhance the personalization and accuracy of the screening results, thereby improving the system's response to user questions.
[0076] By dynamically adjusting the similarity threshold and time window, it can flexibly adapt to different dialogue scenarios and ensure the relevance and real-time nature of the screening results; based on regular optimization of user feedback data, it can continuously improve the screening parameters and enhance the ability to adapt to users' personalized needs.
[0077] Among them, the similarity threshold and time window are regularly and automatically optimized based on user feedback, including:
[0078] Construct a prompt word template for detecting the relevance of the current question and the historical dialogue;
[0079] Based on the question to be identified and its historical dialogue, the relevance detection prompt word template is used to guide large language models including but not limited to Llama, Qwen, and GPT-4 to output the relevance detection results;
[0080] Determine the proportion of related and irrelevant items based on the correlation test results;
[0081] If the relevant ratio exceeds the set threshold, the similarity threshold is increased and the time window is shortened to enhance the relevance of the screening;
[0082] If the proportion of irrelevant content exceeds the set threshold, the similarity threshold is reduced and the time window is expanded to increase the coverage of conversation screening.
[0083] In the above process, by constructing a relevance detection prompt word template, the evaluation criteria for dialogue relevance can be clarified, and the accuracy and consistency of relevance judgment can be improved; by using the powerful semantic analysis capabilities of the large language model, the potential relevance between the current question and the historical dialogue can be more accurately identified; by dynamically adjusting the similarity threshold and time window, the system can adapt to the changes in different user needs and enhance the flexibility and personalization of the screening results. At the same time, this optimization method can effectively reduce redundant dialogues that are irrelevant to the user's current intentions, significantly improve the response efficiency and user experience of the dialogue system, and meet the dual requirements of dialogue screening accuracy and breadth in multiple scenarios.
[0084] For example, the template for reply quality sentiment detection is: "Based on the following conversation content, determine the relevance of the current question with the historical conversation, and directly output 'relevant' or 'irrelevant'." The template for relevance detection after filling in the i-th round of questions to be identified and their historical conversations is: "Based on the following conversation content, determine the relevance of the current question with the historical conversation, and directly output 'relevant' or 'irrelevant'.##Historical conversation content: User question 1: How to return a product? System reply 1: Please click 'Apply for Return' on the order page and fill in the reason for the return. User question 2: How long does it take to process a return? System reply 2: It usually takes 5 working days to complete the processing. Current question: My return application has not been processed, what's the matter?##"
[0085] 3. Calculation of Similarity between Current Question and Historical Dialogue
[0086] In this step, when screening historical conversations, the similarity between the current question and the stored historical conversation data is first calculated, and the similarity includes a weighted fusion of semantic similarity, keyword similarity, topic similarity and context similarity.
[0087] In an exemplary embodiment, the similarity calculation process includes:
[0088] The semantic similarity between the question and the historical conversation is calculated through the semantic similarity model to identify the deep correlation between the question and the historical conversation at the semantic level;
[0089] The keyword similarity algorithm is used to calculate the keyword similarity between the question and the historical dialogue to capture the superficial keyword matching degree;
[0090] The similarity between the question and the historical conversation is calculated through the topic similarity algorithm to evaluate the topic consistency between the question and the historical conversation content;
[0091] The context similarity calculation method is used to evaluate the continuity and cohesion of historical conversations, generate context similarity, and ensure the context logic and coherence of the conversation;
[0092] At least two of the similarities among semantic similarity, keyword similarity, topic similarity and context similarity are weighted and fused to obtain the final similarity between the current question and the historical dialogue.
[0093] In the above process, the semantic similarity, keyword similarity, topic similarity and context similarity between the current question and the historical dialogue are comprehensively considered to ensure the consistency between the two in terms of semantic level, text surface, topic and context logic and coherence.
[0094] In specific implementation, in the similarity calculation process, two or more of semantic similarity, keyword similarity, topic similarity and context similarity can be used for calculation and fusion. For example, suppose that only semantic similarity, keyword similarity and topic similarity are calculated in the configuration file. Then traverse the historical dialogues and calculate the semantic similarity s between the current question and each round of historical dialogues. 1 , keyword similarity 2 , topic similarity s 3 , according to the weights w configured separately 1 、w 2 and w 3 Weighted summation to get the final similarity: s 1 *w 1 +s 2 *w 2 +s 3 *w 3 .
[0095] In the above similarity calculation process, for the calculation of semantic similarity, the semantic similarity model used includes but is not limited to a model based on the SimCSE algorithm, a model based on BERT, and a semantic similarity calculation method based on large language models such as Llama, Qwen, and GPT-4. It is understandable that in order to improve efficiency and reduce the semantic similarity calculation time, only the semantic similarity between the current question and each round of user questions in the historical dialogue can be calculated.
[0096] In the above similarity calculation process, the calculation process of keyword similarity is as follows: extract keywords from questions and historical conversations using algorithms including but not limited to TF-IDF, TextRank, RAKE, and WordNet; convert the extracted keywords into vector representations using vectorization models including but not limited to bag-of-words (BoW) and Word2Vec; and measure the keyword similarity between questions and historical conversations using similarity calculation methods including but not limited to cosine similarity, Jaccard similarity, and Euclidean distance. It is understandable that in order to improve efficiency and reduce the time for keyword similarity calculation, only the keyword similarity between the current question and each round of user questions in the historical conversation can be calculated. At the same time, the keywords of each round of user questions are saved in advance so that the calculated keywords can be directly reused in subsequent conversations to improve the efficiency of similarity calculation.
[0097] In the above similarity calculation process, the calculation process for topic similarity is as follows: use the topic model to calculate the topic words of the current question; use the vector model to convert the topic words into vector representation; use the similarity calculation method to calculate the similarity between the question topic vector and the historical conversation topic vector as the topic similarity indicator. In the above process, for consistency considerations, the topic word calculation method of the current question is consistent with that of the historical conversation; the vector model is consistent with the vector model in the keyword similarity calculation process; the similarity calculation method is consistent with the similarity calculation method in the keyword similarity algorithm. It can be understood that in order to improve efficiency and reduce the topic similarity calculation time, only the topic similarity of the current question and each round of user questions in the historical conversation can be calculated.
[0098] In the above similarity calculation process, the calculation process of context similarity is as follows: construct a prompt template for evaluating the continuity and coherence of the conversation; fill in the current question and the historical conversation content as the input of the large language model; use the large language model to generate the conversation context similarity, and use this method to evaluate the continuity and coherence of the current question and the historical conversation. In the above process, by constructing the prompt template, the evaluation criteria for the continuity and coherence of the conversation can be clarified, and the pertinence of the similarity calculation can be improved; combined with the powerful semantic understanding and context association capabilities of the large language model, the logical coherence between the current question and the historical conversation can be more accurately evaluated; at the same time, filling in the question and the historical conversation content as the model input helps to make full use of the context information, thereby generating similarity results that are more in line with the actual context, significantly improving the response quality and user experience of the dialogue system.
[0099] For example, suppose the user's current question is: "Will it be sunny tomorrow?", and the historical dialogue of the i-th round is: "User asked: What is the weather in Shanghai today? System reply: It is sunny in Shanghai today, and the temperature is about 25 degrees."
[0100] Then, the context similarity calculation process is as follows:
[0101] (1) Construct the following example prompt template: "Based on the following dialogue content, evaluate the logical continuity and cohesion between the current question and the historical dialogue, and output the similarity score."
[0102] (2) Fill in the question and the historical conversation content into the prompt template as the input of the large language model: "Historical conversation: User question: What is the weather in Shanghai today? System reply: It is sunny in Shanghai today, and the temperature is about 25 degrees. Current question: Will it be sunny tomorrow? Based on the above content, evaluate the logical continuity and cohesion of the current question and the historical conversation, and output the similarity score."
[0103] (3) The large language model generates context similarity: "The context similarity score between the current question and the historical dialogue is 0.85".
[0104] 4. Historical dialogue screening
[0105] In this step, based on a preset dialogue screening threshold and in combination with the similarity calculation result between the current question and the stored historical dialogue data, the historical dialogue data most relevant to the current question is screened out from the stored historical dialogue data.
[0106] In an exemplary embodiment, the method for filtering historical conversations includes:
[0107] According to the time window, select the historical dialogue within the time window before the current question;
[0108] Based on the similarity threshold, select the top n historical conversations with the highest similarity from the historical conversations outside the time window;
[0109] The historical conversations within the time window are merged with the selected Top n rounds of historical conversations to generate the final filtered historical conversations.
[0110] In the above process, by combining the historical conversations within the time window and the historical conversations with the highest similarity outside the time window, the timeliness and semantic relevance of the conversation are fully taken into account; the setting of the time window ensures the system's rapid response to recent conversation content, and filtering high-similarity conversations from outside the time window avoids missing important content across time; finally, the two parts of the conversation are fused to generate the screening results, which not only enhances the comprehensiveness of the screening results, but also ensures a high relevance to the current question, thereby effectively improving the response quality of the system.
[0111] According to the method provided in this embodiment, the problem of possible omission of important content across time and topics in the prior art is solved, the dependence on keyword extraction is reduced, and conversations without obvious keywords can be effectively processed. At the same time, the dependence on the accuracy of the intent recognition model is weakened, and better screening effects can be achieved in conversation scenarios that lack clear intent, thereby improving the adaptability to complex conversation scenarios and meeting user needs.
[0112] Example 2
[0113] This embodiment provides a device for filtering historical dialogues in a dialogue system. Figure 2 , which includes:
[0114] A storage module, used to store historical conversation data between the user and the conversation system based on the time and topic of the conversation;
[0115] A configuration module, used to configure a conversation screening threshold and automatically optimize the conversation screening threshold based on user feedback data; the conversation screening threshold includes an adjustable similarity threshold and a time window;
[0116] A similarity calculation module is used to calculate the similarity between the current question and the stored historical dialogue data, wherein the similarity includes a weighted fusion of semantic similarity, keyword similarity, topic similarity and context similarity;
[0117] The screening module is used to screen out the historical dialogue data most relevant to the current question from the stored historical dialogue data based on a preset dialogue screening threshold and the similarity calculation result between the current question and the stored historical dialogue data.
[0118] Since the functional modules of the screening device in this embodiment correspond to the step description of the screening method in Embodiment 1, and the specific implementation of the method steps has been described in Embodiment 1, the specific implementation of each functional module will not be repeated here.
[0119] The screening device provided in this embodiment solves the problem that important content across time and topics may be missed in existing devices, reduces the reliance on keyword extraction, and can effectively handle conversations without obvious keywords. At the same time, it weakens the reliance on the accuracy of the intent recognition model, and can achieve better screening effects in conversation scenarios that lack clear intent, thereby improving the ability to adapt to complex conversation scenarios and meeting user needs.
[0120] Finally, it should be noted that the above embodiments are only preferred implementations and are not intended to limit the present invention. It should be pointed out that for those skilled in the art, several modifications, equivalent replacements, improvements, etc. can be made without departing from the scope of the present invention and the scope of protection of the claims, and all of these should be included in the protection scope of the present invention.
Claims
1. A method for screening historical dialogues in a dialogue system, characterized in that: The following steps are involved: Store historical conversation data between the user and the conversation system based on the time and topic of the conversation; When screening historical conversations, first calculate the similarity between the current question and the stored historical conversation data, which includes a weighted fusion of semantic similarity, keyword similarity, topic similarity, and context similarity; Then, based on a preset dialogue screening threshold and in combination with the similarity calculation result between the current question and the stored historical dialogue data, the historical dialogue data most relevant to the current question is screened out from the stored historical dialogue data; The conversation screening threshold includes an adjustable similarity threshold and a time window.
2. A method for screening historical dialogues in a dialogue system as claimed in claim 1, characterized in that: The storing of historical conversation data between the user and the conversation system based on the time and topic of the conversation includes: Calculate the topic of each round of dialogue between the user and the dialogue system; Based on the time and topic of the conversation, the conversation content is stored in multi-level classification.
3. A method for screening historical dialogues in a dialogue system as claimed in claim 1, characterized in that: The method for calculating the similarity between the current question and the stored historical dialogue data includes: Calculate the semantic similarity between the current question and the historical dialogue through the semantic similarity model; Calculate the keyword similarity between the current question and the historical dialogue through the keyword similarity algorithm; Calculate the topic similarity between the current question and the historical dialogue through the topic similarity algorithm; Calculate the context similarity between the current question and the historical dialogue through the context similarity algorithm; The semantic similarity, keyword similarity, topic similarity, and context similarity between the current question and the historical conversation are weightedly fused to obtain the similarity between the current question and the historical conversation.
4. A method for screening historical dialogues in a dialogue system as claimed in claim 3, characterized in that: The method of calculating the keyword similarity between the current question and the historical dialogue by using a keyword similarity algorithm includes: Extract keywords from the current question and historical conversations respectively; The extracted keywords are converted into vector representations using a vectorization model; The similarity calculation method is used to calculate the similarity between the vector representation of the keywords in the current question and the vector representation of the keywords in the historical conversation.
5. A method for screening historical dialogues in a dialogue system as claimed in claim 3, characterized in that: The topic similarity between the current question and the historical dialogue is calculated by using a topic similarity algorithm, including: Use the topic model to calculate the topic words of the current question; Using a vector model to convert the subject word of the current question into a vector representation; The similarity calculation method is used to calculate the similarity between the topic word vector of the current question and the topic vector of the historical dialogue.
6. A method for screening historical dialogues in a dialogue system as claimed in claim 3, characterized in that: The calculating the context similarity between the current question and the historical dialogue by using the context similarity algorithm includes: constructing prompt templates for assessing conversational continuity and cohesion; Fill the current question and historical conversation content as the input of the large language model; A large language model is used to generate contextual similarity between the current question and historical conversations.
7. A method for screening historical dialogues in a dialogue system as claimed in claim 1, characterized in that: The method for setting the conversation screening threshold includes: Set an adjustable similarity threshold; Set adjustable time windows; Automatically optimize the similarity threshold and time window based on user feedback data.
8. A method for screening historical dialogues in a dialogue system as claimed in claim 7, characterized in that: The method for automatically optimizing the similarity threshold and time window according to user feedback data includes: Construct a prompt word template for detecting the relevance of the current question and the historical dialogue; Based on the current question and historical conversations, the relevance detection prompt word template is used to guide the large language model to output the relevance detection results; According to the relevance test results, determine the proportion of historical conversations that are relevant and irrelevant to the current question; If the relevant ratio exceeds the set threshold, the similarity threshold is increased and the time window is reduced; If the proportion of irrelevant data exceeds the set threshold, the similarity threshold is reduced and the time window is expanded.
9. A method for screening historical dialogues in a dialogue system as claimed in claim 8, characterized in that: The method of selecting the historical dialogue data most relevant to the current question from the stored historical dialogue data includes: According to the time window, select the historical dialogue within the time window before the current question; Based on the similarity threshold, select the top n historical conversations with the highest similarity from the historical conversations outside the time window; The historical dialogues within the time window are merged with the selected Top n rounds of historical dialogues to obtain the historical dialogues most relevant to the current question.
10. A device for screening historical dialogues in a dialogue system, characterized in that: include: A storage module, used to store historical conversation data between the user and the conversation system based on the time and topic of the conversation; A configuration module, used to configure a conversation screening threshold and automatically optimize the conversation screening threshold based on user feedback data; the conversation screening threshold includes an adjustable similarity threshold and a time window; A similarity calculation module is used to calculate the similarity between the current question and the stored historical dialogue data, wherein the similarity includes a weighted fusion of semantic similarity, keyword similarity, topic similarity and context similarity; The screening module is used to screen out the historical dialogue data most relevant to the current question from the stored historical dialogue data based on a preset dialogue screening threshold and the similarity calculation result between the current question and the stored historical dialogue data.
Citation Information
Cited By
Short message content AI iteration method and system based on user feedback
CN120509418A
A method and system for AI-based iterative analysis of SMS content based on user feedback
CN120509418B
Information retrieval method of smart city system
CN121562603A