Multi-turn Dialogue Interaction Method and System Based on Context Reconstruction and Multi-database Retrieval
Through the methods based on context reconstruction and multi-store retrieval, the dialogue incoherence and response deviation of multiple rounds of dialogue systems in complex scenarios are solved, and the self-learning and optimization of the dialogue system is realized, and the accuracy and user experience of the dialogue system are improved.
Patent Information
- Application Number
- CN202510585691.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing multi-round dialogue system lacks in-depth understanding and effective reconstruction of the dialogue context when dealing with complex dialogue scenarios, resulting in incoherence of responses, loss of information or misunderstanding, lack of flexibility and diversity in retrieval strategies, deviation from response statements from actual needs, lack of self-learning and optimization capabilities, affecting user experience and system performance.
Through the methods based on context reconstruction and multi-store retrieval, multiple rounds of dialogue data are obtained, context reconstruction process is performed, semantic focus offset and semantic dependency intensity characteristics are generated, search range and strategy are dynamically adjusted, multi-strategy matching and optimization are performed, and feedback is carried out to the system to update database parameters to achieve self-learning and optimization.
It improves the accuracy and fluency of multiple rounds of dialogue interaction, enhances the comprehensiveness and adaptability of retrieval, ensures that the response conforms to context logic and has high confidence, and improves the performance and user experience of the dialogue system.
Smart Images

Figure CN120123485B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a multi-turn dialogue interaction method and system based on context reconstruction and multi-library retrieval. Background Art
[0002] In the current development field of intelligent dialogue systems, with the continuous progress of artificial intelligence technology and the increasing complexity of application scenarios, multi-turn dialogue interaction has become one of the key indicators for measuring the performance of dialogue systems. However, existing multi-turn dialogue interaction technologies still face many challenges when dealing with complex dialogue scenarios, which limit the intelligence level and user experience of dialogue systems.
[0003] On the one hand, when existing dialogue systems process multi-turn dialogues, they often lack in-depth understanding and effective reconstruction of dialogue context. Most systems generate responses only based on the user input of the current turn and preset response templates, while ignoring the semantic information and dependency relationships in historical dialogue turns. This processing method causes dialogue systems to easily have problems such as discontinuous responses, information loss, or misunderstandings when facing dialogue scenarios with closely related context and complex semantics, seriously affecting the fluency and accuracy of the dialogue.
[0004] On the other hand, when existing dialogue systems retrieve response statements, they usually rely on a single database or retrieval strategy, lacking flexibility and diversity. This single retrieval method not only limits the sources and scopes of response statements but also is difficult to adapt to the diverse needs of different users and different scenarios. In addition, existing retrieval strategies often lack dynamic adjustment and optimization of context information, resulting in a deviation between retrieval results and the actual needs of the current dialogue, reducing the pertinence and effectiveness of responses.
[0005] Furthermore, when existing dialogue systems generate response strategies, most of them adopt static or simple dynamic adjustment methods and cannot perform multi-strategy matching and dynamic optimization according to the context characteristics of the dialogue and retrieval results. This processing method makes it difficult for dialogue systems to generate response statements that not only conform to the context logic but also have a high confidence level when facing complex and changeable dialogue scenarios, affecting the quality of the dialogue and user experience.
[0006] Finally, there are obvious deficiencies in the response optimization of existing dialogue systems. Most systems lack effective feedback mechanisms and learning capabilities and cannot dynamically adjust and optimize parameters such as retrieval weights and context windows according to the actual feedback of users and the operation data of dialogue systems. Such dialogue systems lacking self-learning and optimization capabilities are difficult to adapt to the changing user needs and application scenarios, limiting their long-term development and application potential. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide a multi-round dialogue interaction method and system based on context reconstruction and multi-database retrieval.
[0008] According to the first aspect of the present invention, there is provided a multi-round dialogue interaction method based on context reconstruction and multi-database retrieval, the method comprising:
[0009] Obtain a multi-round dialogue data set of a target user, the multi-round dialogue data set including a plurality of dialogue turn sequences, each dialogue turn sequence being composed of a user input statement and a corresponding system response statement;
[0010] Perform context reconstruction processing on the multi-round dialogue data set to generate context reconstruction features for each dialogue turn sequence, the context reconstruction features including a semantic focus offset of the current dialogue turn and a semantic dependence strength of historical dialogue turns;
[0011] Retrieve multi-database retrieval results matching the context reconstruction features from a preset heterogeneous database based on a dynamic context window, the multi-database retrieval results including a candidate response statement set and corresponding response confidence levels;
[0012] Perform multi-strategy matching processing on the context reconstruction features and the multi-database retrieval results to generate a response generation strategy for the current dialogue turn, the response generation strategy being used to dynamically adjust the semantic priority and confidence threshold of candidate response statements;
[0013] Feed the response generation strategy back to the dialogue service system to trigger a response optimization operation, the response optimization operation being used to update the retrieval weights of the heterogeneous database and the dynamic adjustment parameters of the context window.
[0014] In a possible implementation manner of the first aspect, the performing context reconstruction processing on the multi-round dialogue data set to generate context reconstruction features for each dialogue turn sequence includes:
[0015] Extract local semantic features of the user input statement of the current dialogue turn and feedback semantic features of the corresponding system response statement;
[0016] Traverse the historical dialogue turns of the dialogue turn sequence, perform long-range dependence analysis on the user input statements of each historical dialogue turn, and generate historical semantic chain features, the historical semantic chain features including the decay weight of the historical semantic focus in time series and the jump association strength;
[0017] Perform consistency verification processing on the local semantic features, feedback semantic features, and historical semantic chain features to obtain the semantic focus offset of the current dialogue turn, the semantic focus offset being used to quantify the semantic coherence difference between the current dialogue turn and historical dialogue turns;
[0018] Dynamically intercept the historical semantic chain features based on a preset context window threshold to generate the semantic dependence strength of the historical dialogue turn, where the semantic dependence strength is used to characterize the semantic contribution degree of the intercepted historical semantic chain to the current dialogue turn;
[0019] Normalize and splice the semantic focus offset and the semantic dependence strength to generate the context reconstruction feature of the dialogue turn sequence.
[0020] In a possible implementation manner of the first aspect, the performing consistency verification processing on the local semantic feature, the feedback semantic feature, and the historical semantic chain feature to obtain the semantic focus offset of the current dialogue turn includes:
[0021] Calculate the semantic alignment degree between the local semantic feature and the feedback semantic feature, where the semantic alignment degree is used to quantify the coverage range of the system response statement to the user input statement;
[0022] Traverse the historical semantic chain features of each historical dialogue turn, and calculate the jump correlation degree between the local semantic feature and the historical semantic chain feature, where the jump correlation degree is used to characterize the implicit semantic connection between the current dialogue turn and the historical dialogue turn;
[0023] Construct a consistency scoring function based on the semantic alignment degree and the jump correlation degree, where the consistency scoring function is used to measure whether the current dialogue turn deviates from the core semantic path of the historical dialogue turn;
[0024] Determine the semantic focus offset according to the output value of the consistency scoring function. If the output value is lower than the fifth threshold, it is determined that the current dialogue turn has a semantic focus offset, and a corresponding semantic focus offset is generated based on the offset direction.
[0025] In a possible implementation manner of the first aspect, the constructing a consistency scoring function based on the semantic alignment degree and the jump correlation degree includes:
[0026] Perform normalization processing on the semantic alignment degree to obtain a first sub-score;
[0027] Perform time decay weighted processing on the jump correlation degree to obtain a second sub-score;
[0028] Perform a linear combination of the first sub-score and the second sub-score according to the position weight of the current dialogue turn in the dialogue turn sequence to generate a consistency score;
[0029] Wherein, the weight coefficient of the time decay weighted processing is inversely proportional to the number of intervening turns between the historical dialogue turn and the current dialogue turn, and the position weight is proportional to the position serial number of the current dialogue turn in the sequence.
[0030] In a possible implementation manner of the first aspect, the retrieving a multi-database search result matching the context reconstruction feature from preset heterogeneous databases based on the dynamic context window includes:
[0031] adjusting the interception range of the dynamic context window according to the semantic focus offset to determine the number of valid historical rounds of the current dialogue round;
[0032] Based on the semantic dependency strength, weighted fusion is performed on the historical semantic chain features within the valid historical round number to generate a context retrieval vector;
[0033] Invoke the distributed index service of the heterogeneous database to map the context search vector to the search space of multiple sub-databases, wherein the search space of each sub-database corresponds to a data partition of a response type;
[0034] In the retrieval space of each sub-database, the cosine similarity between the context retrieval vector and the semantic embedding vector of the candidate response sentence is calculated, and a set of candidate response sentences having a similarity higher than a first threshold is screened out according to the cosine similarity;
[0035] The confidence calibration process is performed on the screened candidate response sentence set to generate the response confidence of the candidate response sentence, and the confidence calibration process includes: performing weighted calculation based on the occurrence frequency, semantic conflict degree and user feedback score of the candidate response sentence in the historical dialogue rounds.
[0036] In a possible implementation manner of the first aspect, calling the distributed index service of the heterogeneous database to map the context search vector to the search space of the plurality of sub-databases includes:
[0037] Determining, according to the semantic distribution feature of the context retrieval vector, a matching degree between the semantic distribution feature and the response type of each sub-database;
[0038] Selecting a sub-database with a matching degree higher than a sixth threshold as a target search space, and allocating dynamic search resources to the target search space;
[0039] In the target retrieval space, clustering the candidate response sentences based on the dimensional distribution of the context retrieval vector to generate a plurality of semantic clusters;
[0040] Selecting a central candidate response sentence from each of the semantic clusters as a retrieval anchor point, and calculating a similarity difference between the context retrieval vector and the retrieval anchor point;
[0041] If the similarity difference is less than the seventh threshold, all candidate response statements in the corresponding semantic cluster are added to the candidate response statement set.
[0042] In a possible implementation of the first aspect, the multi-strategy matching process for the context reconstruction feature and the multi-library retrieval results to generate a response generation strategy for the current dialogue turn includes:
[0043] Determine the semantic stability level of the current dialogue turn according to the semantic focus offset, and the semantic stability level is used to divide the semantic correction intensity of candidate response statements;
[0044] Based on the response confidence, perform priority sorting on the candidate response statement set to generate an initial response priority sequence;
[0045] Dynamically adjust the initial response priority sequence according to the semantic stability level to generate an adjusted response priority sequence. Among them, if the semantic stability level is lower than the preset level threshold, the priority of the candidate response statement with the highest matching degree with the historical semantic chain feature is increased, and at the same time, the priority of the candidate response statement with a semantic conflict degree higher than the second threshold is decreased;
[0046] Generate a response generation strategy based on the adjusted response priority sequence and the confidence threshold. The response generation strategy includes: semantic correction rules, priority update rules, and confidence filtering rules for response statements.
[0047] In a possible implementation of the first aspect, the feedback of the response generation strategy to the dialogue service system to trigger a response optimization operation includes:
[0048] Update the semantic mapping relationship of the corresponding data partition in the heterogeneous database according to the semantic correction rules;
[0049] Based on the priority update rules, adjust the retrieval weight distribution parameters of the distributed index service so that the data partitions of higher-priority response types obtain larger weight coefficients in subsequent retrievals;
[0050] Re-calibrate the initial confidence of each candidate response statement in the heterogeneous database according to the confidence filtering rules, including: degrading the confidence of the candidate response statement with a user feedback score lower than the third threshold, and removing the candidate response statement with a semantic conflict degree exceeding the fourth threshold from the retrieval results;
[0051] Among them, the degradation of the confidence of the candidate response statement with a user feedback score lower than the third threshold includes:
[0052] Obtain the explicit feedback data and implicit feedback data of the candidate response statement in the historical dialogue turn. The explicit feedback data includes the user score or correction instruction, and the implicit feedback data includes whether the user repeats the same semantic input in the subsequent dialogue turn;
[0053] Generate a comprehensive feedback score based on the weighted sum result of explicit feedback data and implicit feedback data;
[0054] If the comprehensive feedback score is lower than the third threshold, downgrade the confidence level of the candidate response statement to the preset lowest level and move it to the response pool to be verified;
[0055] Regularly perform semantic conflict analysis and context adaptability testing on the candidate response statements in the response pool to be verified, and restore their original confidence levels if the tests pass.
[0056] In a possible implementation manner of the first aspect, the determining the matching degree between the semantic distribution feature and the response types of each sub-database includes:
[0057] Extract the projection components of the context retrieval vector in the preset response type dimensions, where the response type dimensions include task-based, question-and-answer-based, and casual chat-based;
[0058] Calculate the Euclidean distance between the response type label vector of each sub-database and the projection component, where the response type label vector is composed of the statistical values of the historical response statement types of the sub-database;
[0059] Normalize the Euclidean distance to obtain the initial matching degree of each sub-database;
[0060] Dynamically compensate the initial matching degree based on the semantic focus offset of the current conversation turn. If the semantic focus offset exceeds the eighth threshold, increase the matching degree compensation coefficient of the task-based sub-database and decrease the matching degree compensation coefficient of the casual chat-based sub-database at the same time;
[0061] Determine the matching degree between the semantic distribution feature and the response types of the sub-database according to the compensated matching degree, and sort the matching degrees in descending order to select the target retrieval space.
[0062] According to the second aspect of the present invention, there is provided a multi-turn dialogue interaction system, where the multi-turn dialogue interaction system includes a machine-readable storage medium and a processor. The machine-readable storage medium stores machine-executable instructions, and when the processor executes the machine-executable instructions, the multi-turn dialogue interaction system implements the foregoing multi-turn dialogue interaction method based on context reconstruction and multi-database retrieval.
[0063] According to the third aspect of the present invention, there is provided a computer-readable storage medium, where computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed, the foregoing multi-turn dialogue interaction method based on context reconstruction and multi-database retrieval is implemented.
[0064] According to any of the above aspects, the technical effects of the present invention are as follows:
[0065] In the embodiment of the present application, by constructing a multi-round dialogue interaction method based on context reconstruction and multi-database retrieval, the full-process intelligent processing from dialogue data parsing to response strategy generation is realized, significantly improving the accuracy and fluency of multi-round dialogue interaction, and providing an innovative technical solution for the optimization and upgrade of the dialogue system. Specifically, by obtaining the multi-round dialogue data set of the target user and performing context reconstruction processing on it, the semantic focus offset of each dialogue turn and the semantic dependence strength of the historical dialogue turns can be accurately captured, providing rich and accurate context features for subsequent retrieval and matching, and effectively solving the problem of context information loss or misunderstanding in traditional dialogue systems when dealing with complex dialogue scenarios. Retrieving multi-database retrieval results that match the context reconstruction features from a preset heterogeneous database based on a dynamic context window not only broadens the retrieval scope and improves the comprehensiveness of the retrieval, but also enhances the flexibility and adaptability of the retrieval by dynamically adjusting the size of the context window, making the retrieval results more in line with the actual needs of the current dialogue. Further, performing multi-strategy matching processing on the context reconstruction features and the multi-database retrieval results to generate a response generation strategy for the current dialogue turn, which can dynamically adjust the semantic priority and confidence threshold of candidate response sentences, ensuring that the generated response not only conforms to the context logic of the dialogue, but also has a high confidence and user satisfaction. Finally, feedback the response generation strategy to the dialogue service system to trigger response optimization operations, and realize the self-learning and continuous optimization of the dialogue system by continuously updating the retrieval weights of the heterogeneous database and the dynamic adjustment parameters of the context window, effectively improving the performance and user experience of the dialogue system. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0067] Figure 1 The flowchart showing the multi-round dialogue interaction method based on context reconstruction and multi-database retrieval provided by the embodiment of the present invention;
[0068] Figure 2 The component structure diagram showing the multi-round dialogue interaction system for implementing the above multi-round dialogue interaction method based on context reconstruction and multi-database retrieval provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] Embodiments of the present invention will be described below with reference to the accompanying drawings in the present invention. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present invention, and do not limit the technical solutions of the embodiments of the present invention.
[0070] Those skilled in the art of the present technology can understand that, unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present invention mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations, etc. supported by the art of the present technology. It should be understood that when an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include a wireless connection or a wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term. For example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0071] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Through the description of several exemplary embodiments below, the technical solutions of the embodiments of the present invention and the technical effects produced by the technical solutions of the present invention will be illustrated. It should be noted that the following embodiments can be referred to, learned from or combined with each other. For the same terms, similar features and similar implementation steps in different embodiments, they will not be described repeatedly.
[0072] Figure 1 The flowchart of the multi-round dialogue interaction method and system based on context reconstruction and multi-library retrieval provided by the embodiments of the present invention is shown. It should be understood that in other embodiments, the order of some steps of the multi-round dialogue interaction method based on context reconstruction and multi-library retrieval in this embodiment can be shared according to actual needs, or some of the steps can also be omitted or maintained. The detailed steps of the multi-round dialogue interaction method based on context reconstruction and multi-library retrieval include:
[0073] Step S110: Obtain a multi-round dialogue data set of a target user, where the multi-round dialogue data set includes multiple dialogue round sequences, and each dialogue round sequence is composed of a user input statement and a corresponding system response statement.
[0074] In the field of medical data retrieval, the target users mainly refer to individuals who interact with the dialogue service system in a medical scenario, which could be either patients or medical staff. Taking patients as an example, when seeking medical help, patients will conduct multiple rounds of conversations with the medical dialogue service system to consult information related to diseases. The multi-round dialogue data set is actually a detailed record of a series of conversations between the patient and the system.
[0075] Specifically, these sequences of dialogue turns are arranged in chronological order. For example, in the first round of conversation between the patient and the system, the patient's input statement is "I've been coughing a lot lately. What could be the reason?" This indicates that the patient is currently experiencing the symptom of coughing and hopes to understand the possible causes. Based on the preset rules and knowledge reserves, the system gives the response statement "Coughing can be caused by various reasons, such as colds, allergies, etc.", providing the patient with some possible directions for the cause of the illness.
[0076] In the second round of conversation, the patient's input statement is "I don't have any other symptoms of a cold. Could it be an allergy?" This shows that after referring to the system's response in the first round and considering their own actual situation, the patient further asks whether it is due to an allergic reaction. The system then gives another response: "It's possible. Coughing caused by allergies usually also comes with symptoms such as sneezing and runny nose", further guiding the patient to judge the allergic situation.
[0077] In the third round of conversation, the patient's input statement is "I also don't have the symptoms of sneezing and runny nose. What could be other possible reasons?" This indicates that after self-examination, the patient has ruled out the possibility of an allergy and continues to seek other possible causes from the system. The system responds with "In addition to colds and allergies, some respiratory diseases may also cause coughing, such as bronchitis", providing the patient with new possible causes of the illness.
[0078] These sequentially arranged dialogue turns together constitute the multi-round dialogue data set, recording the gradually in-depth communication process between the patient and the system.
[0079] Step S120: Perform context reconstruction processing on the multi-round dialogue data set to generate context reconstruction features for each sequence of dialogue turns. The context reconstruction features include the semantic focus offset of the current dialogue turn and the semantic dependence strength of the historical dialogue turns.
[0080] After obtaining the multi-round dialogue data set between the patient and the system, it is necessary to perform context reconstruction processing on it. The purpose of this processing is to generate context reconstruction features for each sequence of dialogue turns. The semantic focus offset and semantic dependence strength included in the context reconstruction features can help the system more accurately understand the relationship between the current dialogue turn and the historical dialogue, thereby providing a more accurate basis for subsequent retrieval and response.
[0081] Step S121: Extract the local semantic features of the user input statement and the feedback semantic features of the corresponding system response statement in the current conversation turn.
[0082] Taking the third-round conversation as an example, the user input statement in the current conversation turn is "I don't have the symptoms of sneezing or running nose either. Then what could be the other reasons". To extract the local semantic features of this statement, a series of natural language processing techniques need to be applied. First is lexical analysis. The statement is tokenized, splitting it into individual words, obtaining "I", "don't have", "sneezing", "running nose", "symptoms", "then", "could be", "what", "reasons", etc. These words are the basic units of the statement, and subsequent processing will be based on these words.
[0083] Then perform syntactic analysis to analyze the grammatical relationships between these words. For example, "I" is the subject, "don't have" is the predicate, and "the symptoms of sneezing or running nose" is the object, etc. Through syntactic analysis, the structure of the statement can be clarified, which helps to better understand the semantics of the statement.
[0084] Next is semantic understanding. The tokenized words are mapped into a predefined semantic space. Here, a pre-trained word vector model is used, which can map each word into a 300-dimensional vector space. For example, the word "cough" corresponds to a 300-dimensional vector in the word vector model, and each dimension of this vector represents a feature of the word in the semantic space.
[0085] Perform weighted average processing on the vectors of these words. Assume that the weight of each word is determined according to its importance in the statement. For some key words, such as "symptoms", "reasons", higher weights are assigned, while for some function words or auxiliary words, such as "then", "don't have", lower weights are assigned. Through weighted average, the local semantic feature vector of the user input statement is obtained, which is a 300-dimensional vector and contains the semantic information of the current statement.
[0086] The corresponding system response statement is "Besides cold and allergy, some respiratory diseases may also cause cough, such as bronchitis". Similarly, perform tokenization, word vector mapping, and weighted average processing on this statement. First, tokenize the statement into words such as "besides", "cold", "and", "allergy", "some", "respiratory diseases", "also", "may", "cause", "cough", "such as", "bronchitis", etc. Then map these words into a 300-dimensional vector space, and then perform weighted average according to the importance of the words to obtain the feedback semantic feature vector of the system response statement, which is also a 300-dimensional vector.
[0087] Step S122: Traverse the historical dialogue turns in the dialogue turn sequence, perform long-range dependence analysis on the user input statements in each historical dialogue turn, and generate historical semantic chain features, where the historical semantic chain features include the decay weight of the historical semantic focus in time series and the jump association strength.
[0088] Continuing with the third round of dialogue as an example, the historical dialogue turns include the first round and the second round. For the first round of dialogue, the user input statement is "I always cough recently. What's going on?", and long-range dependence analysis is performed on it. Here, a long short-term memory network (LSTM) or gated recurrent unit (GRU) based on deep learning is used to capture the long-range dependence relationships in the statement.
[0089] Input the word vector sequence of this statement into the LSTM model. The LSTM model has memory cells that can remember the dependence relationships between words at different positions in the statement. For example, there is a temporal dependence relationship between the words "recently" and "always", and the LSTM model can capture this relationship. The LSTM model outputs a hidden state vector, which contains the semantic information of the statement and is a 300-dimensional vector.
[0090] Then, according to the time interval between this historical dialogue turn and the current dialogue turn, determine the decay weight of the historical semantic focus in time series. Assume that the preset decay function is an exponential decay function, and its formula is decay weight = base ^ interval turn number, where base is a constant less than 1. Assume base is 0.9. The first round of dialogue and the third round of dialogue are separated by 2 turns, so the decay weight for this turn is 0.9 ^ 2 = 0.81.
[0091] At the same time, calculate the jump association strength between the current dialogue turn and the first round of dialogue. This can be achieved by calculating the similarity between the semantic vectors of the two statements. Here, cosine similarity is used. Calculate the cosine similarity between the local semantic feature vector of the third round of dialogue and the semantic vector of the first round of dialogue. Assume that the cosine similarity between the semantic vector of the first round of dialogue and the local semantic feature vector of the third round of dialogue is 0.6, and use this similarity as the jump association strength.
[0092] For the second round of conversation, long-term dependence analysis, decay weight calculation, and jump association strength calculation are also performed. The word vector sequence of the user input statement "I don't have other symptoms of a cold. Could it be an allergy?" in the second round of conversation is input into the LSTM model to obtain a 300-dimensional hidden state vector. Since there is a one-round interval between the second round of conversation and the third round of conversation, according to the exponential decay function, the decay weight is 0.9^1 = 0.9. Calculate the cosine similarity between the semantic vector of the second round of conversation and the local semantic feature vector of the third round of conversation. Suppose it is 0.7. This similarity is the jump association strength between the current conversation turn and the second round of conversation.
[0093] Concatenate the decay weight, jump association strength, and semantic vector of each round of historical conversation to obtain the historical semantic chain feature. Suppose the semantic vector of each round is 300-dimensional, and the decay weight and jump association strength are each 1-dimensional. Then the historical semantic chain feature of the first round is a 302-dimensional vector, and the historical semantic chain feature of the second round is also a 302-dimensional vector.
[0094] Step S123: Perform a consistency verification process on the local semantic feature, feedback semantic feature, and historical semantic chain feature to obtain the semantic focus offset of the current conversation turn. The semantic focus offset is used to quantify the semantic coherence difference between the current conversation turn and the historical conversation turns.
[0095] Step S1231: Calculate the semantic alignment degree between the local semantic feature and the feedback semantic feature. The semantic alignment degree is used to quantify the coverage range of the system response statement for the user input statement.
[0096] Taking the third round of conversation as an example, calculate the semantic alignment degree between the local semantic feature vector of the user input statement in the third round and the feedback semantic feature vector of the system response statement. Use cosine similarity to calculate the similarity between the two vectors and take it as the semantic alignment degree. The calculation of cosine similarity is obtained by dividing the dot product of the two vectors by the product of their magnitudes. Suppose the local semantic feature vector in the third round is V1, the feedback semantic feature vector is V2, their dot product is V1·V2, the magnitude of V1 is |V1|, and the magnitude of V2 is |V2|. Then the semantic alignment degree = V1·V2 / (|V1|*|V2|). Suppose the calculated cosine similarity is 0.7, that is, the semantic alignment degree is 0.7. This means that the coverage range of the system response statement for the user input statement is 70%.
[0097] Step S1232: Traverse the historical semantic chain features of each historical conversation turn, and calculate the jump association degree between the local semantic feature and the historical semantic chain feature. The jump association degree is used to characterize the implicit semantic connection between the current conversation turn and the historical conversation turns.
[0098] For the historical semantic chain features of the first-round historical dialogue, extract the semantic vector part (300 dimensions) and calculate its cosine similarity with the local semantic feature vector of the third round. Also use the calculation formula of cosine similarity. Suppose the obtained similarity is 0.6, and this similarity is the jump correlation degree between the current dialogue turn and the first-round historical dialogue. For the historical semantic chain features of the second-round historical dialogue, similarly extract its semantic vector part and calculate its cosine similarity with the local semantic feature vector of the third round. Suppose it is 0.7, and this similarity is the jump correlation degree between the current dialogue turn and the second-round historical dialogue.
[0099] Step S1233: Construct a consistency scoring function based on the semantic alignment degree and the jump correlation degree. The consistency scoring function is used to measure whether the current dialogue turn deviates from the core semantic path of the historical dialogue turns.
[0100] Step S12331: Normalize the semantic alignment degree to obtain a first sub-score.
[0101] The value range of the semantic alignment degree is usually between 0 and 1. To incorporate it into the consistency scoring function, normalization is required. Here, the linear normalization method is used. Since the semantic alignment degree itself is between 0 and 1, it is directly used as the first sub-score. In the third-round dialogue, the semantic alignment degree is 0.7, so the first sub-score is also 0.7.
[0102] Step S12332: Perform time decay weighting on the jump correlation degree to obtain a second sub-score.
[0103] For the jump correlation degree of 0.6 of the first-round historical dialogue, since it is 2 turns away from the third-round dialogue, according to the preset time decay function, such as the exponential decay function, assuming the decay coefficient is 0.9, the weighted jump correlation degree is calculated as 0.6 * 0.9^2 = 0.486. For the jump correlation degree of 0.7 of the second-round historical dialogue, since it is 1 turn away from the third-round dialogue and the decay coefficient is 0.9, the weighted jump correlation degree is calculated as 0.7 * 0.9^1 = 0.63.
[0104] Sum up the weighted jump correlation degrees of all historical dialogue turns. Suppose there are only two rounds of historical dialogue, and the sum is 0.486 + 0.63 = 1.116. Then perform normalization. Suppose the linear normalization method is used, and divide the sum by the number of historical dialogue turns (2) to obtain the second sub-score of approximately 0.558.
[0105] Step S12333: Linearly combine the first sub-score and the second sub-score according to the position weight of the current dialogue turn in the dialogue turn sequence to generate a consistency score.
[0106] Assume that the current conversation turn is the third turn, and its position serial number in the conversation turn sequence is 3. The position weight is proportional to the position serial number of the current conversation turn in the sequence. Assume that the calculation formula for the position weight is position weight = 0.1 * position serial number. Then the position weight of the third turn is 0.1 * 3 = 0.3.
[0107] Consistency score = first sub-score * (1 - position weight) + second sub-score * position weight = 0.7 * (1 - 0.3) + 0.558 * 0.3 = 0.49 + 0.1674 = 0.6574.
[0108] Step S1234: Determine the semantic focus offset amount according to the output value of the consistency scoring function. If the output value is lower than the fifth threshold, it is determined that the current conversation turn has a semantic focus shift, and a corresponding semantic focus offset amount is generated based on the offset direction.
[0109] Assume that the fifth threshold is 0.7, and the consistency score of the third conversation turn is 0.6574, which is lower than the fifth threshold. It is determined that the current conversation turn has a semantic focus shift. To determine the offset direction, the difference between the local semantic features and the historical semantic chain features can be analyzed. By comparing the values of each dimension of the semantic vector, the dimension with a larger difference is found. For example, in the 100th dimension, the value of the local semantic feature vector of the third turn is 0.8, while the value of the historical semantic chain feature vector of the first turn in this dimension is 0.3, and the value of the historical semantic chain feature vector of the second turn in this dimension is 0.4. The difference value of this dimension is 0.8 - ((0.3 + 0.4) / 2) = 0.45. This difference value of this dimension is used as an index of the offset direction.
[0110] Then, according to the preset rule, the offset direction is converted into a numerical value. Assume that the rule is to multiply the index of the offset direction by a coefficient of 0.2, resulting in 0.45 * 0.2 = 0.09. Combining with the difference between the consistency score and the fifth threshold, the difference is 0.7 - 0.6574 = 0.0426. The semantic focus offset amount = 0.09 + 0.0426 = 0.1326.
[0111] Step S124: Dynamically intercept the historical semantic chain features based on a preset context window threshold to generate the semantic dependence strength of the historical conversation turn. The semantic dependence strength is used to represent the semantic contribution degree of the intercepted historical semantic chain to the current conversation turn.
[0112] Assume that the preset context window threshold is 2, that is, only the historical semantic chain features of the last 2 historical conversation turns are considered. For the third conversation turn, the historical semantic chain features of the second turn and the first turn are intercepted.
[0113] In order to calculate the semantic dependency strength, the intercepted historical semantic chain features can be weighted and summed. According to the decay weight of each round of historical dialogue, the semantic vector part of its historical semantic chain features is weighted. Assuming that the decay weight of the second round is 0.9 and the decay weight of the first round is 0.81, the semantic vector of the second round of historical semantic chain features is V2, and the semantic vector of the first round of historical semantic chain features is V1. Then the weighted semantic vector is 0.9*V2+0.81*V1.
[0114] Then, the cosine similarity between the weighted semantic vector and the local semantic feature vector of the third round is calculated. Assuming that the calculated cosine similarity is 0.76, the similarity is used as the semantic dependency strength of the historical dialogue rounds.
[0115] Step S125: normalizing and concatenating the semantic focus offset and the semantic dependency strength to generate context reconstruction features of the dialogue turn sequence.
[0116] The value ranges of semantic focus offset and semantic dependency strength may be different, and they need to be normalized. Both semantic focus offset and semantic dependency strength are between 0 and 1, so their values are used directly.
[0117] The semantic focus offset and semantic dependency strength are concatenated to obtain a 2D vector, which is used as the context reconstruction feature of the dialogue turn sequence. In the third round of dialogue, the semantic focus offset is 0.1326, the semantic dependency strength is 0.76, and the context reconstruction feature vector is [0.1326, 0.76].
[0118] Step S130: retrieving multi-database search results matching the context reconstruction feature from preset heterogeneous databases based on the dynamic context window, wherein the multi-database search results include a set of candidate response statements and corresponding response confidences.
[0119] After obtaining the context reconstruction features of each dialogue turn sequence, it is necessary to retrieve matching multi-database search results from the preset heterogeneous database based on the dynamic context window. The heterogeneous database may contain multiple sub-databases of different types, such as medical knowledge base, case database, expert experience database, etc.
[0120] Step S131: adjusting the capture range of the dynamic context window according to the semantic focus offset, and determining the number of valid historical rounds of the current dialogue round.
[0121] Taking the third round of conversation as an example, its semantic focus offset is 0.1326. Suppose the mapping relationship between the preset semantic focus offset and the number of effective historical turns is as follows: when the semantic focus offset is less than 0.2, the number of effective historical turns is 2; when the semantic focus offset is between 0.2 and 0.5, the number of effective historical turns is 1; when the semantic focus offset is greater than 0.5, the number of effective historical turns is 0. Since the semantic focus offset of the third round of conversation is 0.1326, which is less than 0.2, the number of effective historical turns of the current conversation turn is determined to be 2, that is, the last 2 historical conversation turns are considered.
[0122] Step S132: Based on the semantic dependency strength, perform weighted fusion on the historical semantic chain features within the number of effective historical turns to generate a context retrieval vector.
[0123] In the third round of conversation, the number of effective historical turns is 2, that is, the historical semantic chain features of the second round and the first round. The semantic dependency strength is 0.76.
[0124] For the semantic vector V2 of the historical semantic chain features of the second round and the semantic vector V1 of the historical semantic chain features of the first round, perform weighted fusion according to the semantic dependency strength. Suppose the weight of the second round is 0.76 * 0.9 (the decay weight of the second round), and the weight of the first round is 0.76 * 0.81 (the decay weight of the first round). Then the context retrieval vector = 0.76 * 0.9 * V2 + 0.76 * 0.81 * V1.
[0125] Step S133: Invoke the distributed index service of the heterogeneous database to map the context retrieval vector to the retrieval spaces of multiple sub-databases, where the retrieval space of each sub-database corresponds to a data partition of a response type.
[0126] Step S1331: Determine the matching degree between the semantic distribution feature and each sub-database's response type according to the semantic distribution feature of the context retrieval vector.
[0127] Step S13311: Extract the projection components of the context retrieval vector in the preset response type dimensions, and the response type dimensions include task type, question-and-answer type, and casual chat type.
[0128] Assume that the context retrieval vector is a 300-dimensional vector, and the preset response type dimension is 3-dimensional, corresponding to task type, Q&A type, and casual chat type respectively. To extract the projection components, three basis vectors corresponding to the response type dimensions need to be determined first. Assume that the basis vector for the task type dimension is T, the basis vector for the Q&A type dimension is Q, and the basis vector for the casual chat type dimension is C. These three basis vectors are all 300-dimensional vectors. By calculating the inner product of the context retrieval vector and the basis vectors of each response type dimension, the projection components in each response type dimension are obtained. The calculation method of the inner product is to multiply the values of the corresponding dimensions of the context retrieval vector and the basis vector and then sum them up. For example, for the task type dimension, assume the context retrieval vector is R and the task type basis vector is T. The projection component of the task type dimension is equal to the first dimension value of R multiplied by the first dimension value of T, plus the second dimension value of R multiplied by the second dimension value of T, and so on until all 300 dimensions are calculated and summed. Assume that after calculation, the projection component in the task type dimension is 0.3, the projection component in the Q&A type dimension is 0.6, and the projection component in the casual chat type dimension is 0.1.
[0129] Step S13312: Calculate the Euclidean distance between the response type label vector of each sub-database and the projection component, where the response type label vector is composed of the statistical values of the historical response statement type distribution of this sub-database.
[0130] Assume that there are three sub-databases in the heterogeneous database, namely Sub-database A, Sub-database B, and Sub-database C. The statistical distribution of the historical response statement types of Sub-database A shows that the task type response accounts for 20%, the Q&A type response accounts for 70%, and the casual chat type response accounts for 10%. Then the response type label vector of Sub-database A is [0.2, 0.7, 0.1]. The statistical distribution of the historical response statement types of Sub-database B shows that the task type response accounts for 40%, the Q&A type response accounts for 50%, and the casual chat type response accounts for 10%. Its response type label vector is [0.4, 0.5, 0.1]. The statistical distribution of the historical response statement types of Sub-database C shows that the task type response accounts for 10%, the Q&A type response accounts for 30%, and the casual chat type response accounts for 60%. Its response type label vector is [0.1, 0.3, 0.6].
[0131] Calculate the Euclidean distance between the response type label vector of sub-database A and the projection component [0.3, 0.6, 0.1]. First, calculate the square of the difference for each dimension. That is, for the task type dimension, the difference is 0.3 - 0.2 = 0.1, and its square is 0.01; for the Q&A type dimension, the difference is 0.6 - 0.7 = -0.1, and its square is 0.01; for the casual chat type dimension, the difference is 0.1 - 0.1 = 0, and its square is 0. Then add these squared values together to get 0.01 + 0.01 + 0 = 0.02. Finally, take the square root of this sum (although a formula editor is not allowed here, the logic is like this), and we get approximately 0.1414.
[0132] Similarly, calculate the Euclidean distance between the response type label vector of sub-database B and the projection component. For the task type dimension, the difference is 0.3 - 0.4 = -0.1, and its square is 0.01; for the Q&A type dimension, the difference is 0.6 - 0.5 = 0.1, and its square is 0.01; for the casual chat type dimension, the difference is 0.1 - 0.1 = 0, and its square is 0. Add them up to get 0.02, and after taking the square root, it is approximately 0.1414.
[0133] Calculate the Euclidean distance between the response type label vector of sub-database C and the projection component. For the task type dimension, the difference is 0.3 - 0.1 = 0.2, and its square is 0.04; for the Q&A type dimension, the difference is 0.6 - 0.3 = 0.3, and its square is 0.09; for the casual chat type dimension, the difference is 0.1 - 0.6 = -0.5, and its square is 0.25. Add them up to get 0.04 + 0.09 + 0.25 = 0.38, and after taking the square root, it is approximately 0.6164.
[0134] Step S13313: Normalize the Euclidean distance to obtain the initial matching degree of each sub-database.
[0135] Using the linear normalization method, set the maximum value of the Euclidean distance to 1 and the minimum value to 0. The Euclidean distances of sub-database A and sub-database B are the smallest, both being 0.1414, and the Euclidean distance of sub-database C is the largest, being 0.6164.
[0136] The calculation method for the initial matching degree of sub-database A is: subtract the difference between the Euclidean distance of sub-database A and the minimum Euclidean distance divided by the difference between the maximum Euclidean distance and the minimum Euclidean distance from 1, that is, 1 - (0.1414 - 0.1414) / (0.6164 - 0.1414) = 1.
[0137] The initial matching degree of sub-database B is calculated as: 1 - (0.1414 - 0.1414) / (0.6164 - 0.1414) = 1.
[0138] The initial matching degree of sub-database C is calculated as: 1 - (0.6164 - 0.1414) / (0.6164 - 0.1414) = 0.
[0139] Step S13314: Dynamically compensate the initial matching degree based on the semantic focus offset of the current conversation turn. If the semantic focus offset exceeds the eighth threshold, increase the matching degree compensation coefficient of the task-based sub-database and decrease the matching degree compensation coefficient of the chit-chat sub-database simultaneously.
[0140] Assume that the eighth threshold is 0.1, and the semantic focus offset of the third conversation turn is 0.1326, which exceeds the eighth threshold. Set the matching degree compensation coefficient of the task-based sub-database to 0.2, and the matching degree compensation coefficient of the chit-chat sub-database to -0.2.
[0141] Sub-database A contains task-based responses, and its initial matching degree is 1. The compensated matching degree is 1 + 0.2 = 1.2, but the matching degree cannot exceed 1, so the compensated matching degree of sub-database A is still 1.
[0142] Sub-database C contains more chit-chat responses, and its initial matching degree is 0. The compensated matching degree is 0 - 0.2 = -0.2, but the matching degree cannot be less than 0, so the compensated matching degree of sub-database C is 0.
[0143] Sub-database B contains a certain proportion of task-based and Q&A responses. Since there is compensation for the task-based part, assuming the task-based proportion is calculated for compensation at 40%, the compensation amount is 0.2 * 0.4 = 0.08, and the compensated matching degree is 1 + 0.08 = 1.08. Similarly, the matching degree cannot exceed 1, so the compensated matching degree of sub-database B is still 1.
[0144] Determine the matching degree between the semantic distribution feature and the response type of the sub-database according to the compensated matching degree, and sort the matching degrees in descending order to select the target retrieval space. Here, the matching degrees of sub-database A and sub-database B are 1, and the matching degree of sub-database C is 0, so select sub-database A and sub-database B as the target retrieval space.
[0145] Step S1332: Select the sub-databases with matching degrees higher than the sixth threshold as the target retrieval space and allocate dynamic retrieval resources to the target retrieval space.
[0146] Assume that the sixth threshold is 0.8, and the matching degrees of sub-database A and sub-database B are 1, which are higher than the sixth threshold. So, take sub-database A and sub-database B as the target retrieval space. To improve the retrieval efficiency, allocate dynamic retrieval resources according to the matching degree for the target retrieval space. Assume the total retrieval resources are 100 units, and the matching degrees of sub-database A and sub-database B are the same, then each is allocated 50 units of retrieval resources.
[0147] Step S1333: In the target retrieval space, perform clustering on the candidate response statements based on the dimensionality distribution of the context retrieval vector to generate multiple semantic clusters.
[0148] In sub-database A and sub-database B, there are a large number of candidate response statements. First, each candidate response statement is transformed into a 300-dimensional semantic embedding vector through a word vector model. Then, based on the dimensionality distribution of the context retrieval vector, a clustering algorithm (such as the K-Means clustering algorithm) is used to cluster the semantic embedding vectors of these candidate response statements. Suppose they are clustered into 5 semantic clusters, and the candidate response statements in each semantic cluster are semantically similar.
[0149] Step S1334: Select the central candidate response statement from each semantic cluster as the retrieval anchor point, and calculate the similarity difference between the context retrieval vector and the retrieval anchor point.
[0150] In each semantic cluster, the central vector of the semantic cluster is obtained by calculating the average of the semantic embedding vectors of all candidate response statements within the cluster. The candidate response statement closest to this central vector is selected as the retrieval anchor point. For each retrieval anchor point, calculate the cosine similarity between its semantic embedding vector and the context retrieval vector. Suppose the context retrieval vector is R and the semantic embedding vector of the retrieval anchor point is S. The cosine similarity calculation formula is the dot product of R and S divided by the product of the norm of R and the norm of S. Suppose the cosine similarities between 5 retrieval anchor points and the context retrieval vector are calculated to be 0.7, 0.75, 0.8, 0.85, and 0.9 respectively. Suppose the cosine similarity between the context retrieval vector and itself is 1. Then the similarity differences between these 5 retrieval anchor points and the context retrieval vector are 1 - 0.7 = 0.3, 1 - 0.75 = 0.25, 1 - 0.8 = 0.2, 1 - 0.85 = 0.15, and 1 - 0.9 = 0.1 respectively.
[0151] Step S1335: If the similarity difference is less than the seventh threshold, add all candidate response statements within the corresponding semantic cluster to the candidate response statement set.
[0152] Suppose the seventh threshold is 0.2. Then all candidate response statements within the semantic clusters corresponding to the two retrieval anchor points with similarity differences of 0.15 and 0.1 will be added to the candidate response statement set. Suppose these two semantic clusters have 10 and 15 candidate response statements respectively. Then 25 candidate response statements are added to the candidate response statement set.
[0153] Step S134: In the retrieval space of each sub-database, calculate the cosine similarity between the context retrieval vector and the semantic embedding vector of the candidate response statement, and filter out the candidate response statement set with a similarity higher than the first threshold according to the cosine similarity.
[0154] In the retrieval spaces of sub-database A and sub-database B, for each candidate response statement in the set of candidate response statements, calculate the cosine similarity between its semantic embedding vector and the context retrieval vector. Assuming the first threshold is 0.7, after calculation, filter out the candidate response statements with a cosine similarity higher than 0.7. Suppose there are 25 statements in the set of candidate response statements. After filtering, 18 statements have a cosine similarity higher than 0.7, and these 18 statements form the new set of candidate response statements.
[0155] Step S135: Perform confidence calibration processing on the filtered set of candidate response statements to generate the response confidence of the candidate response statements. The confidence calibration processing includes: performing weighted calculation based on the occurrence frequency, semantic conflict degree, and user feedback score of the candidate response statement in the historical dialogue turns.
[0156] For the 18 filtered candidate response statements, calculate their occurrence frequency, semantic conflict degree, and user feedback score in the historical dialogue turns respectively. Suppose the weight of the occurrence frequency is 0.3, the weight of the semantic conflict degree is 0.2, and the weight of the user feedback score is 0.5.
[0157] For a certain candidate response statement, it appears 3 times in the historical dialogue turns. Suppose according to the preset rule, the score calculation of the occurrence frequency is the number of occurrences divided by the total number of historical dialogue turns (assumed to be 10 turns), that is, 3 / 10 = 0.3. The calculation of the semantic conflict degree is to compare the semantic embedding vector of this candidate response statement with the semantic vectors of each turn in the historical dialogue, find the dimensions with larger semantic differences, and score according to the degree of difference. Suppose the semantic conflict degree score of this statement is 0.1. The user feedback score is obtained by normalizing the score given by the user to this statement in the historical dialogue (assumed to be a 0 - 10 score system). Suppose the user feedback score of this statement is 0.8.
[0158] Then the response confidence of this candidate response statement is 0.3 * 0.3 + 0.2 * 0.1 + 0.5 * 0.8 = 0.09 + 0.02 + 0.4 = 0.51. Perform such calculations on the 18 candidate response statements to obtain the response confidence of each statement.
[0159] Step S140: Perform multi-strategy matching processing on the context reconstruction feature and the multi-database retrieval result to generate the response generation strategy for the current dialogue turn. The response generation strategy is used to dynamically adjust the semantic priority and confidence threshold of the candidate response statements.
[0160] Step S141: Determine the semantic stability level of the current dialogue turn according to the semantic focus offset amount. The semantic stability level is used to divide the semantic correction intensity of the candidate response statements.
[0161] Suppose the mapping relationship between the semantic focus shift and the semantic stability level is as follows: when the semantic focus shift is less than 0.1, the semantic stability level is high; when the semantic focus shift is between 0.1 and 0.3, the semantic stability level is medium; when the semantic focus shift is greater than 0.3, the semantic stability level is low. The semantic focus shift of the third-round conversation is 0.1326, so the semantic stability level is medium.
[0162] For the case where the semantic stability level is medium, the semantic correction intensity for dividing the candidate response sentences is medium. This means that for the candidate response sentences, a certain degree of semantic correction is required to better adapt to the semantics of the current conversation.
[0163] Step S142: Based on the response confidence, perform priority sorting on the set of candidate response sentences to generate an initial response priority sequence.
[0164] For the 18 selected candidate response sentences and their corresponding response confidences, sort them from high to low according to the response confidence. Suppose the initial response priority sequence obtained after sorting is: candidate response sentence 1 (confidence 0.8), candidate response sentence 2 (confidence 0.75), candidate response sentence 3 (confidence 0.72)... candidate response sentence 18 (confidence 0.5).
[0165] Step S143: Dynamically adjust the initial response priority sequence according to the semantic stability level to generate an adjusted response priority sequence. Among them, if the semantic stability level is lower than the preset level threshold, the priority of the candidate response sentence with the highest matching degree with the historical semantic chain feature is increased, and at the same time, the priority of the candidate response sentence with a semantic conflict degree higher than the second threshold is decreased.
[0166] Suppose the preset level threshold is high, and the semantic stability level of the third-round conversation is medium, lower than the preset level threshold. First, calculate the cosine similarity between the semantic embedding vector of each candidate response sentence and the semantic vector of the historical semantic chain feature, and find the candidate response sentence with the highest matching degree with the historical semantic chain feature. Suppose the candidate response sentence 5 has the highest matching degree with the historical semantic chain feature, and its cosine similarity is 0.85.
[0167] At the same time, suppose the second threshold is 0.2. For each candidate response sentence, calculate its semantic conflict degree, and find the candidate response sentence with a semantic conflict degree higher than 0.2. Suppose the semantic conflict degree of candidate response sentence 10 is 0.25, higher than the second threshold.
[0168] Promote the priority of candidate response statement 5 to the front of the sequence and lower the priority of candidate response statement 10 to the back of the sequence. The adjusted response priority sequence is: candidate response statement 5 (confidence 0.7), candidate response statement 1 (confidence 0.8), candidate response statement 2 (confidence 0.75) …… candidate response statement 10 (confidence 0.6), candidate response statement 18 (confidence 0.5).
[0169] Step S144: Generate a response generation strategy based on the adjusted response priority sequence and the confidence threshold. The response generation strategy includes: semantic correction rules for response statements, priority update rules, and confidence filtering rules.
[0170] According to the adjusted response priority sequence, formulate semantic correction rules for response statements. Since the semantic stability level is medium, for candidate response statements with higher priorities, perform moderate semantic corrections. For example, for candidate response statement 5, according to the semantic focus of the current conversation, perform some lexical substitutions or sentence structure adjustments on it, but without changing its core semantics.
[0171] The priority update rule is to dynamically adjust the priorities of candidate response statements according to the progress of the conversation. If in subsequent conversations, a certain candidate response statement receives positive feedback from the user, increase its priority; if it receives negative feedback, lower its priority.
[0172] The confidence filtering rule is to set a confidence threshold, assumed to be 0.6. For candidate response statements in the adjusted response priority sequence with a confidence lower than 0.6, remove them from the candidate set. In this way, after filtering, the remaining candidate response statements will be used for subsequent conversation responses.
[0173] Step S150: Feed the response generation strategy back to the dialogue service system to trigger a response optimization operation, which is used to update the retrieval weights of the heterogeneous database and the dynamic adjustment parameters of the context window.
[0174] Step S151: Update the semantic mapping relationship of the corresponding data partition in the heterogeneous database according to the semantic correction rule.
[0175] According to the semantic correction rule in the response generation strategy, update the semantic mapping relationship of the corresponding data partition in the heterogeneous database. For example, for candidate response statement 5, semantic corrections are made, and the corrected semantic information is updated to the corresponding position in the database. If candidate response statement 5 comes from a certain data partition of sub-database A, then in this data partition, replace the original semantic mapping relationship with the corrected semantic mapping relationship to ensure that the subsequent retrieved response statements better meet the semantic requirements of the current conversation.
[0176] Step S152: Adjust the retrieval weight allocation parameters of the distributed index service based on the priority update rule, so that data partitions of higher-priority response types obtain larger weight coefficients in subsequent retrievals.
[0177] According to the priority update rule, adjust the retrieval weight allocation parameters of the distributed index service. Suppose in the adjusted response priority sequence, the candidate response statements of the task-based response type have a higher priority. Then in the distributed index service, increase the retrieval weight coefficient of the data partition corresponding to the task-based response type. For example, the original retrieval weight coefficient of the task-based data partition is 0.3, and now it is increased to 0.5. In this way, in subsequent retrievals, the task-based data partition will obtain more retrieval resources and it is easier to retrieve relevant response statements.
[0178] Step S153: Recalibrate the initial confidence levels of each candidate response statement in the heterogeneous database according to the confidence filtering rule, including: degrading the confidence levels of candidate response statements with user feedback scores lower than the third threshold, and removing candidate response statements with semantic conflict degrees exceeding the fourth threshold from the retrieval results.
[0179] Step S1531: Degrade the confidence levels of candidate response statements with user feedback scores lower than the third threshold.
[0180] Suppose the third threshold is 0.6, and obtain the explicit feedback data and implicit feedback data of the candidate response statements in historical conversation turns. The explicit feedback data includes user scores or correction instructions, and the implicit feedback data includes whether the user repeats the same semantic input in subsequent conversation turns.
[0181] For a certain candidate response statement, its explicit feedback data is that the user score is 0.5, and the implicit feedback data is that the user does not repeat the same semantic input in subsequent conversations. Suppose the weight of the explicit feedback data is 0.7 and the weight of the implicit feedback data is 0.3. First, convert the implicit feedback data into a score. Suppose the score corresponding to the user not repeating the same semantic input in subsequent conversations is 0.4. Generate a comprehensive feedback score according to the weighted sum of the explicit feedback data and the implicit feedback data, that is, comprehensive feedback score = explicit feedback score × explicit feedback data weight + implicit feedback score × implicit feedback data weight, which is 0.5×0.7 + 0.4×0.3 = 0.35 + 0.12 = 0.47. Since this comprehensive feedback score of 0.47 is lower than the third threshold of 0.6, the confidence level of this candidate response statement is degraded to the preset lowest level. Suppose the preset lowest level is 0.2, then the confidence level of this candidate response statement is reduced from the original value (suppose it is 0.65) to 0.2, and it is moved to the pool of response statements to be verified.
[0182] The response pool to be verified is used to store candidate response sentences whose confidence levels have been downgraded and need to be further verified. The candidate response sentences in this response pool to be verified are regularly subjected to semantic conflict analysis and context adaptability testing. Semantic conflict analysis involves comparing the semantic embedding vectors of the candidate response sentences with the context retrieval vectors of the current conversation and the semantic vectors of the historical conversation turns to check for obvious semantic conflicts. Context adaptability testing, on the other hand, simulates different conversation scenarios to check whether the candidate response sentences can respond reasonably in these scenarios.
[0183] For example, for this candidate response sentence moved into the response pool to be verified, when performing semantic conflict analysis, calculate the cosine similarity between its semantic embedding vector and the context retrieval vector of the current third-round conversation. Suppose the similarity obtained is 0.3, indicating a certain semantic difference. Then calculate the cosine similarity between it and the semantic vectors of the historical semantic chain features of the first and second rounds of conversations respectively. Suppose they are 0.2 and 0.25 respectively, which also shows a large difference from the historical conversation semantics and there is a semantic conflict.
[0184] In the context adaptability testing, simulate different patient consultation scenarios, such as different symptom combinations, different questioning methods, etc. Put this candidate response sentence into these scenarios to see if it can respond reasonably. If in multiple simulated scenarios, this candidate response sentence cannot be well adapted, then it needs to be further corrected or eliminated. If after a series of tests, it is found that this candidate response sentence performs well in both semantic conflict analysis and context adaptability testing, for example, the semantic similarity reaches above 0.7 and it can respond reasonably in the simulated scenarios, then restore its original confidence level (here it is 0.65), and remove it from the response pool to be verified and put it back into the normal set of candidate response sentences.
[0185] Step S1532: Remove the candidate response sentences whose semantic conflict degree exceeds the fourth threshold from the retrieval results.
[0186] Suppose the fourth threshold is 0.3. For each candidate response sentence in the heterogeneous database, calculate its semantic conflict degree. The calculation method of the semantic conflict degree has been mentioned before. It is to compare the semantic embedding vector of the candidate response sentence with the semantic vectors of each round in the historical conversation, find the dimensions with larger semantic differences, and score according to the degree of difference.
[0187] For example, there is a candidate response statement. The difference value between its semantic embedding vector and the semantic vector of the historical semantic chain feature of the first-round conversation is 0.2 in a certain key dimension, the difference value between its semantic embedding vector and the semantic vector of the historical semantic chain feature of the second-round conversation is 0.15 in another key dimension, and the difference value between its semantic embedding vector and the context retrieval vector of the third-round conversation is 0.1 in a certain dimension. These difference values are combined according to the preset rules to calculate the semantic conflict degree. Suppose the obtained semantic conflict degree is 0.25, which is less than the fourth threshold 0.3. Then this candidate response statement can be retained in the retrieval results.
[0188] For another candidate response statement, after calculation, its semantic conflict degree is 0.35, which exceeds the fourth threshold 0.3. Then this candidate response statement is removed from the retrieval results. This can ensure that the candidate response statements retrieved subsequently have higher semantic consistency with the historical conversation and the current conversation, improving the quality and accuracy of the conversation service.
[0189] Figure 2 Fig. shows a multi-turn dialogue interaction system 100 provided in an embodiment of the present invention, including a processor 1001, a memory 1003, and program code stored on the memory 1003. The processor 1001 executes the above program code to implement the steps of the multi-turn dialogue interaction method based on context reconstruction and multi-library retrieval.
[0190] Figure 2 The multi-turn dialogue interaction system 100 shown in the figure includes: a processor 1001 and a memory 1003. Among them, the processor 1001 and the memory 1003 are connected, such as connected through a bus 1002. Optionally, the multi-turn dialogue interaction system 100 may further include a transceiver 1004. The transceiver 1004 can be used for data interaction between this multi-turn dialogue interaction system and other multi-turn dialogue interaction systems, such as sending and / or receiving data, etc. It should be noted that in actual scheduling, the transceiver 1004 is not limited to one, and the structure of this multi-turn dialogue interaction system 100 does not constitute a limitation to the embodiment of the present invention.
[0191] The processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present invention. The processor 1001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0192] The bus 1002 may include a path for transmitting information between the above components. The bus 1002 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 1002 may be divided into an address bus, a data bus, a control bus, etc.
[0193] The memory 1003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to have or store program code and can be read by a computer, which is not limited herein.
[0194] The memory 1003 is used to store the program code for implementing the embodiments of the present invention and is controlled by the processor 1001 for execution. The processor 1001 is used to execute the program code stored in the memory 1003 to implement the steps shown in the foregoing method embodiments.
[0195] An embodiment of the present invention provides a computer-readable storage medium, on which program code is stored. When the program code is executed by a processor, the steps and corresponding content of the foregoing method embodiment can be implemented.
[0196] It should be understood that although the flowcharts of the embodiments of the present invention indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present invention, the implementation steps in each flowchart can be executed in other orders based on requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages according to the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present invention do not limit this.
[0197] The above are only optional implementation manners of some implementation scenarios of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present invention, adopting other similar implementation means based on the technical idea of the present invention also belongs to the protection scope of the embodiments of the present invention.
Claims
1. A multi-round dialogue interaction method based on context reconstruction and multi-library retrieval, characterized in that The method includes: Obtaining a multi-turn conversation data set of a target user, where the multi-turn conversation data set includes multiple conversation turn sequences, and each conversation turn sequence is composed of a user input statement and a corresponding system response statement; Performing context reconstruction processing on the multi-turn conversation data set to generate context reconstruction features for each conversation turn sequence, where the context reconstruction features include the semantic focus offset of the current conversation turn and the semantic dependence strength of historical conversation turns; Retrieving multi-database retrieval results matching the context reconstruction features from a preset heterogeneous database based on a dynamic context window, where the multi-database retrieval results include a set of candidate response statements and corresponding response confidence levels; Performing multi-strategy matching processing on the context reconstruction features and the multi-database retrieval results to generate a response generation strategy for the current conversation turn, where the response generation strategy is used to dynamically adjust the semantic priority and confidence threshold of candidate response statements; Feeding back the response generation strategy to the conversation service system to trigger a response optimization operation, where the response optimization operation is used to update the retrieval weights of the heterogeneous database and the dynamic adjustment parameters of the context window.
2. The multi-turn dialogue interaction method based on context reconstruction and multi-library retrieval according to claim 1, wherein, The performing context reconstruction processing on the multi-turn conversation data set to generate context reconstruction features for each conversation turn sequence includes: Extracting local semantic features of the user input statement of the current conversation turn and feedback semantic features of the corresponding system response statement; Traversing the historical conversation turns of the conversation turn sequence, performing long-range dependence analysis on the user input statements of each historical conversation turn to generate historical semantic chain features, where the historical semantic chain features include the decay weight of the historical semantic focus in time series and the jump association strength; Performing consistency verification processing on the local semantic features, feedback semantic features, and historical semantic chain features to obtain the semantic focus offset of the current conversation turn, where the semantic focus offset is used to quantify the semantic coherence difference between the current conversation turn and historical conversation turns; Dynamically intercepting the historical semantic chain features based on a preset context window threshold to generate the semantic dependence strength of historical conversation turns, where the semantic dependence strength is used to represent the semantic contribution degree of the intercepted historical semantic chain to the current conversation turn; Normalizing and splicing the semantic focus offset and the semantic dependence strength to generate the context reconstruction features of the conversation turn sequence.
3. The multi-turn dialogue interaction method based on context reconstruction and multi-library retrieval according to claim 2, wherein The performing consistency verification processing on the local semantic features, feedback semantic features, and historical semantic chain features to obtain the semantic focus offset of the current conversation turn includes: Calculating the semantic alignment degree between the local semantic features and the feedback semantic features, where the semantic alignment degree is used to quantify the coverage range of the system response statement for the user input statement; Traversing the historical semantic chain features of each historical conversation turn and calculating the jump association degree between the local semantic features and the historical semantic chain features, where the jump association degree is used to represent the implicit semantic connection between the current conversation turn and historical conversation turns; Building a consistency scoring function based on the semantic alignment and the jump correlation, wherein the consistency scoring function is used to measure whether the current dialogue turn deviates from the core semantic path of the historical dialogue turn; The semantic focus offset is determined according to the output value of the consistency scoring function. If the output value is lower than the fifth threshold, it is determined that a semantic focus offset occurs in the current dialogue turn, and a corresponding semantic focus offset is generated based on the offset direction.
4. The multi-round dialogue interaction method based on context reconstruction and multi-library retrieval according to claim 3, characterized in that The constructing a consistency scoring function based on the semantic alignment and the jump association includes: Normalizing the semantic alignment to obtain a first sub-score; Performing time-attenuated weighted processing on the jump correlation to obtain a second sub-score; Linearly combine the first sub-score and the second sub-score according to the position weight of the current dialogue turn in the dialogue turn sequence to generate a consistency score; The weight coefficient of the time decay weighted processing is inversely proportional to the number of intervals between the historical dialogue rounds and the current dialogue round, and the position weight is proportional to the position number of the current dialogue round in the sequence.
5. The multi-round dialogue interaction method based on context reconstruction and multi-library retrieval according to claim 1, characterized in that The retrieving multi-database search results matching the context reconstruction feature from preset heterogeneous databases based on the dynamic context window includes: adjusting the interception range of the dynamic context window according to the semantic focus offset to determine the number of valid historical rounds of the current dialogue round; Based on the semantic dependency strength, weighted fusion is performed on the historical semantic chain features within the valid historical round number to generate a context retrieval vector; Invoke the distributed index service of the heterogeneous database to map the context search vector to the search space of multiple sub-databases, wherein the search space of each sub-database corresponds to a data partition of a response type; In the retrieval space of each sub-database, the cosine similarity between the context retrieval vector and the semantic embedding vector of the candidate response sentence is calculated, and a set of candidate response sentences having a similarity higher than a first threshold is screened out according to the cosine similarity; The confidence calibration process is performed on the screened candidate response sentence set to generate the response confidence of the candidate response sentence, and the confidence calibration process includes: performing weighted calculation based on the occurrence frequency, semantic conflict degree and user feedback score of the candidate response sentence in the historical dialogue rounds.
6. The multi-turn dialogue interaction method based on context reconstruction and multi-library retrieval according to claim 5, wherein, The calling of the distributed index service of the heterogeneous database to map the context search vector to the search space of multiple sub-databases includes: Determining, according to the semantic distribution feature of the context retrieval vector, a matching degree between the semantic distribution feature and the response type of each sub-database; Selecting a sub-database with a matching degree higher than a sixth threshold as a target search space, and allocating dynamic search resources to the target search space; In the target retrieval space, clustering the candidate response sentences based on the dimensional distribution of the context retrieval vector to generate a plurality of semantic clusters; Selecting a central candidate response sentence from each of the semantic clusters as a retrieval anchor point, and calculating a similarity difference between the context retrieval vector and the retrieval anchor point; If the similarity difference is less than the seventh threshold, all candidate response statements in the corresponding semantic cluster are added to the candidate response statement set.
7. The multi-turn dialogue interaction method based on context reconstruction and multi-library retrieval according to claim 1, characterized in that Performing multi-strategy matching processing on the context reconstruction feature and the multi-database retrieval results to generate a response generation strategy for the current dialogue turn, including: Determining the semantic stability level of the current dialogue turn according to the semantic focus offset, where the semantic stability level is used to divide the semantic correction intensity of candidate response sentences; Performing priority sorting on the set of candidate response sentences based on the response confidence to generate an initial response priority sequence; Dynamically adjusting the initial response priority sequence according to the semantic stability level to generate an adjusted response priority sequence, where if the semantic stability level is lower than a preset level threshold, the priority of the candidate response sentence with the highest matching degree with the historical semantic chain feature is increased, and at the same time, the priority of the candidate response sentence with a semantic conflict degree higher than the second threshold is decreased; Generating a response generation strategy based on the adjusted response priority sequence and the confidence threshold, where the response generation strategy includes: semantic correction rules, priority update rules, and confidence filtering rules of response sentences; 8. The multi-turn dialogue interaction method based on context reconstruction and multi-library retrieval according to claim 7, characterized in that Feeding back the response generation strategy to the dialogue service system to trigger a response optimization operation, including: Updating the semantic mapping relationship of the corresponding data partition in the heterogeneous database according to the semantic correction rules; Adjusting the retrieval weight distribution parameters of the distributed index service based on the priority update rules, so that the data partition of the higher-priority response type obtains a larger weight coefficient in subsequent retrievals; Recalibrating the initial confidence of each candidate response sentence in the heterogeneous database according to the confidence filtering rules, including: downgrading the confidence of the candidate response sentence with a user feedback score lower than the third threshold, and removing the candidate response sentence with a semantic conflict degree exceeding the fourth threshold from the retrieval results; Among them, the downgrading the confidence of the candidate response sentence with a user feedback score lower than the third threshold includes: Obtaining the explicit feedback data and implicit feedback data of the candidate response sentence in the historical dialogue turn, where the explicit feedback data includes user scores or correction instructions, and the implicit feedback data includes whether the user repeats the same semantic input in subsequent dialogue turns; Generating a comprehensive feedback score according to the weighted sum result of the explicit feedback data and the implicit feedback data; If the comprehensive feedback score is lower than the third threshold, downgrading the confidence of the candidate response sentence to the preset lowest level and moving it to the pool of candidate responses to be verified; Regularly performing semantic conflict analysis and context adaptability tests on the candidate response sentences in the pool of candidate responses to be verified, and restoring their original confidence levels if the tests pass; 9. The multi-turn dialogue interaction method based on context reconstruction and multi-library retrieval according to claim 6, characterized in that, Determining the matching degree between the semantic distribution feature and the response type of each sub-database, including: Extracting the projection components of the context retrieval vector in the preset response type dimension, where the response type dimension includes task type, question and answer type, and chatting type; Calculating the Euclidean distance between the response type label vector of each sub-database and the projection component, where the response type label vector is composed of the statistical values of the historical response sentence types of the sub-database; Performing normalization processing on the Euclidean distance to obtain the initial matching degree of each sub-database; Dynamically compensate the initial matching degree based on the semantic focus offset of the current conversation turn. If the semantic focus offset exceeds the eighth threshold, increase the matching degree compensation coefficient of the task-based sub-database, and at the same time decrease the matching degree compensation coefficient of the chit-chat sub-database; Determine the matching degree between the semantic distribution feature and the response type of the sub-database according to the compensated matching degree, and sort the matching degrees in descending order to select the target retrieval space.
10. A multi-round dialogue interaction system, characterized in that It includes a processor and a computer-readable storage medium. The computer-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by the processor, the multi-turn dialogue interaction method based on context reconstruction and multi-database retrieval described in any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Multi-round dialogue method and device, equipment and storage medium
CN118093796A
Dynamic question and answer matching method and system based on context awareness
CN119691133A