Multi-round dialogue interaction method and system based on context reconstruction and multi-library retrieval

By conducting context reconstruction and multi-store search on multiple rounds of dialogue data, a response generation strategy is generated, which solves the problem of insufficient context information processing in the dialogue system in the existing technology in complex scenarios, and achieves higher dialogue accuracy and fluency.

CN120123485AActive Publication Date: 2025-06-10JIEHELIX (SHANGHAI) MEDICAL TECH CO LTD

Patent Information

Application Number
CN202510585691.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-10
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing multi-round dialogue interaction technology lacks in-depth understanding and effective reconstruction of the dialogue context when dealing with complex dialogue scenarios, resulting in incoherence of responses, loss of information or misunderstandings, affecting the fluency and accuracy of dialogue.

Method used

By obtaining the multiple rounds of conversation data set of the target user, the context reconstruction process is performed, the context reconstruction features of each conversation round are generated, and the matching multi-store search results are retrieved from the heterogeneous database based on the dynamic context window, and the multi-strategy matching process is performed to generate a response generation strategy.

Benefits of technology

It significantly improves the accuracy and fluency of multiple rounds of dialogue interaction, solves the problem of context information loss or misunderstanding of dialogue systems in complex scenarios, and improves the targeted response and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123485A_ABST
    Figure CN120123485A_ABST
Patent Text Reader

Abstract

The invention provides a multi-round dialogue interaction method and system based on context reconstruction and multi-library retrieval, and the method comprises the steps: firstly obtaining a multi-round dialogue data set of a target user, which comprises a plurality of dialogue round sequences composed of user input statements and system response statements; performing context reconstruction processing on the multi-round dialogue data set to generate context reconstruction characteristics of each dialogue round sequence, covering current round semantic focus offset and historical round semantic dependence intensity, and retrieving a matched multi-database retrieval result from a preset heterogeneous database based on a dynamic context window, a candidate response statement set and confidence are included. Performing multi-strategy matching on context reconstruction features and retrieval results, generating a response generation strategy to adjust candidate response semantic priorities and confidence thresholds, and finally feeding back the response generation strategy to a dialogue service system to trigger response optimization operation, and updating heterogeneous database retrieval weights and context window dynamic adjustment parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to a multi-round dialogue interaction method and system based on context reconstruction and multi-library retrieval. Background Art

[0002] In the current development field of intelligent dialogue systems, with the continuous progress of artificial intelligence technology and the increasing complexity of application scenarios, multi-round dialogue interaction has become one of the key indicators for measuring the performance of dialogue systems. However, existing multi-round dialogue interaction technologies still face many challenges when dealing with complex dialogue scenarios, which limit the intelligence level and user experience of dialogue systems.

[0003] On the one hand, when existing dialogue systems handle multi-round dialogues, they often lack in-depth understanding and effective reconstruction of dialogue context. Most systems generate responses only based on the current round of user input and preset response templates, while ignoring the semantic information and dependency relationships in historical dialogue rounds. This processing method causes dialogue systems to easily encounter problems such as incoherent responses, information loss, or misunderstandings when facing dialogue scenarios with closely related context and complex semantics, seriously affecting the fluency and accuracy of the dialogue.

[0004] On the other hand, when existing dialogue systems retrieve response statements, they usually rely on a single database or retrieval strategy, lacking flexibility and diversity. This single retrieval method not only limits the sources and scope of response statements but also makes it difficult to adapt to the diverse needs of different users and different scenarios. In addition, existing retrieval strategies often lack dynamic adjustment and optimization of context information, resulting in a deviation between the retrieval results and the actual needs of the current dialogue, reducing the pertinence and effectiveness of the response.

[0005] Furthermore, when existing dialogue systems generate response strategies, most of them adopt static or simple dynamic adjustment methods and cannot perform multi-strategy matching and dynamic optimization according to the context characteristics of the dialogue and the retrieval results. This processing method makes it difficult for dialogue systems to generate response statements that not only conform to the context logic but also have a high confidence level when facing complex and changeable dialogue scenarios, affecting the quality of the dialogue and the user experience.

[0006] Finally, there are obvious deficiencies in the response optimization of existing dialogue systems. Most systems lack effective feedback mechanisms and learning capabilities and cannot dynamically adjust and optimize parameters such as retrieval weights and context windows based on the actual feedback of users and the operation data of the dialogue system. Such dialogue systems lacking self-learning and optimization capabilities are difficult to adapt to the changing user needs and application scenarios, limiting their long-term development and application potential. Summary of the Invention

[0007] In view of this, the purpose of the present invention is to provide a multi-turn dialogue interaction method and system based on context reconstruction and multi-database retrieval.

[0008] According to the first aspect of the present invention, there is provided a multi-turn dialogue interaction method based on context reconstruction and multi-database retrieval, the method comprising: Obtain a multi-turn dialogue data set of a target user, the multi-turn dialogue data set including a plurality of dialogue turn sequences, each dialogue turn sequence consisting of a user input statement and a corresponding system response statement; Perform context reconstruction processing on the multi-turn dialogue data set to generate context reconstruction features for each dialogue turn sequence, the context reconstruction features including a semantic focus offset of the current dialogue turn and a semantic dependence strength of historical dialogue turns; Retrieve multi-database retrieval results matching the context reconstruction features from a preset heterogeneous database based on a dynamic context window, the multi-database retrieval results including a set of candidate response statements and corresponding response confidence levels; Perform multi-strategy matching processing on the context reconstruction features and the multi-database retrieval results to generate a response generation strategy for the current dialogue turn, the response generation strategy being used to dynamically adjust the semantic priority and confidence threshold of the candidate response statements; Feed the response generation strategy back to the dialogue service system to trigger a response optimization operation, the response optimization operation being used to update the retrieval weights of the heterogeneous database and the dynamic adjustment parameters of the context window.

[0009] In a possible implementation manner of the first aspect, the performing context reconstruction processing on the multi-turn dialogue data set to generate context reconstruction features for each dialogue turn sequence includes: Extract local semantic features of the user input statement of the current dialogue turn and feedback semantic features of the corresponding system response statement; Traverse the historical dialogue turns of the dialogue turn sequence, perform long-range dependence analysis on the user input statements of each historical dialogue turn to generate historical semantic chain features, the historical semantic chain features including a decay weight of the historical semantic focus in time series and a jump association strength; Perform consistency verification processing on the local semantic features, feedback semantic features and historical semantic chain features to obtain a semantic focus offset of the current dialogue turn, the semantic focus offset being used to quantify the semantic coherence difference between the current dialogue turn and the historical dialogue turns; Dynamically intercept the historical semantic chain features based on a preset context window threshold to generate a semantic dependence strength of the historical dialogue turns, the semantic dependence strength being used to characterize the semantic contribution degree of the intercepted historical semantic chain to the current dialogue turn; The semantic focus offset and the semantic dependency strength are normalized and concatenated to generate context reconstruction features of the dialogue turn sequence.

[0010] In a possible implementation manner of the first aspect, performing consistency verification processing on the local semantic features, the feedback semantic features, and the historical semantic chain features to obtain the semantic focus offset of the current dialogue turn includes: Calculating the semantic alignment between the local semantic feature and the feedback semantic feature, wherein the semantic alignment is used to quantify the coverage of the system response sentence to the user input sentence; Traversing the historical semantic chain features of each historical dialogue turn, calculating the jump correlation between the local semantic features and the historical semantic chain features, wherein the jump correlation is used to characterize the implicit semantic connection between the current dialogue turn and the historical dialogue turn; Building a consistency scoring function based on the semantic alignment and the jump correlation, wherein the consistency scoring function is used to measure whether the current dialogue turn deviates from the core semantic path of the historical dialogue turn; The semantic focus offset is determined according to the output value of the consistency scoring function. If the output value is lower than the fifth threshold, it is determined that a semantic focus offset occurs in the current dialogue turn, and a corresponding semantic focus offset is generated based on the offset direction.

[0011] In a possible implementation of the first aspect, constructing a consistency scoring function based on the semantic alignment and the jump relevance includes: Normalizing the semantic alignment to obtain a first sub-score; Performing time-attenuated weighted processing on the jump correlation to obtain a second sub-score; Linearly combine the first sub-score and the second sub-score according to the position weight of the current dialogue turn in the dialogue turn sequence to generate a consistency score; The weight coefficient of the time decay weighted processing is inversely proportional to the number of intervals between the historical dialogue rounds and the current dialogue round, and the position weight is proportional to the position number of the current dialogue round in the sequence.

[0012] In a possible implementation manner of the first aspect, the retrieving a multi-database search result matching the context reconstruction feature from preset heterogeneous databases based on the dynamic context window includes: adjusting the interception range of the dynamic context window according to the semantic focus offset to determine the number of valid historical rounds of the current dialogue round; Based on the semantic dependency strength, weighted fusion is performed on the historical semantic chain features within the valid historical round number to generate a context retrieval vector; Call the distributed index service of the heterogeneous database to map the context retrieval vector to the retrieval spaces of multiple sub-databases, where the retrieval space of each sub-database corresponds to a data partition of a response type; In the retrieval space of each sub-database, calculate the cosine similarity between the context retrieval vector and the semantic embedding vector of the candidate response statement, and filter out a set of candidate response statements with a similarity higher than the first threshold according to the cosine similarity; Perform confidence calibration processing on the filtered set of candidate response statements to generate the response confidence of the candidate response statements. The confidence calibration processing includes: performing weighted calculation based on the occurrence frequency, semantic conflict degree, and user feedback score of the candidate response statement in the historical conversation rounds.

[0013] In a possible implementation manner of the first aspect, the calling the distributed index service of the heterogeneous database to map the context retrieval vector to the retrieval spaces of multiple sub-databases includes: According to the semantic distribution characteristics of the context retrieval vector, determine the matching degree between the semantic distribution characteristics and the response type of each sub-database; Select the sub-databases with a matching degree higher than the sixth threshold as the target retrieval spaces, and allocate dynamic retrieval resources to the target retrieval spaces; In the target retrieval space, perform clustering processing on the candidate response statements based on the dimension distribution of the context retrieval vector to generate multiple semantic clusters; Select the central candidate response statement from each semantic cluster as the retrieval anchor point, and calculate the similarity difference between the context retrieval vector and the retrieval anchor point; If the similarity difference is less than the seventh threshold, add all the candidate response statements within the corresponding semantic cluster to the set of candidate response statements.

[0014] In a possible implementation manner of the first aspect, the performing multi-strategy matching processing on the context reconstruction feature and the multi-database retrieval result to generate the response generation strategy for the current conversation round includes: Determine the semantic stability level of the current conversation round according to the semantic focus offset amount. The semantic stability level is used to divide the semantic correction intensity of the candidate response statements; Perform priority sorting on the set of candidate response statements based on the response confidence to generate an initial response priority sequence; Dynamically adjust the initial response priority sequence according to the semantic stability level to generate an adjusted response priority sequence. Among them, if the semantic stability level is lower than the preset level threshold, increase the priority of the candidate response statement with the highest matching degree with the historical semantic chain features, and at the same time decrease the priority of the candidate response statement with a semantic conflict degree higher than the second threshold; Generate a response generation strategy based on the adjusted response priority sequence and the confidence threshold. The response generation strategy includes: semantic correction rules, priority update rules, and confidence filtering rules for response statements.

[0015] In a possible implementation manner of the first aspect, the feedback of the response generation strategy to the dialogue service system to trigger a response optimization operation includes: Update the semantic mapping relationship of the corresponding data partition in the heterogeneous database according to the semantic correction rules; Adjust the retrieval weight distribution parameters of the distributed index service based on the priority update rules, so that the data partition of the higher-priority response type obtains a larger weight coefficient in subsequent retrievals; Recalibrate the initial confidence of each candidate response statement in the heterogeneous database according to the confidence filtering rules, including: downgrading the confidence of the candidate response statement with a user feedback score lower than the third threshold, and removing the candidate response statement with a semantic conflict degree exceeding the fourth threshold from the retrieval results; Among them, the downgrading of the confidence of the candidate response statement with a user feedback score lower than the third threshold includes: Obtain the explicit feedback data and implicit feedback data of the candidate response statement in the historical dialogue rounds. The explicit feedback data includes user scores or correction instructions, and the implicit feedback data includes whether the user repeats the same semantic input in subsequent dialogue rounds; Generate a comprehensive feedback score according to the weighted sum result of the explicit feedback data and the implicit feedback data; If the comprehensive feedback score is lower than the third threshold, downgrade the confidence of the candidate response statement to the preset lowest level and move it to the pool of responses to be verified; Regularly perform semantic conflict analysis and context adaptability tests on the candidate response statements in the pool of responses to be verified, and restore their original confidence levels if the tests pass.

[0016] In a possible implementation manner of the first aspect, the determination of the matching degree between the semantic distribution feature and the response type of each sub-database includes: Extract the projection components of the context retrieval vector in the preset response type dimension. The response type dimension includes task type, question-and-answer type, and chat type; Calculate the Euclidean distance between the response type label vector of each sub-database and the projection component, where the response type label vector is composed of the statistical values of the historical response statement types of the sub-database; Normalize the Euclidean distance to obtain the initial matching degree of each sub-database; Dynamically compensate the initial matching degree based on the semantic focus offset of the current conversation turn. If the semantic focus offset exceeds the eighth threshold, increase the matching degree compensation coefficient of the task-based sub-database and decrease the matching degree compensation coefficient of the chatty sub-database; Determine the matching degree between the semantic distribution feature and the response type of the sub-database according to the compensated matching degree, and sort the matching degrees in descending order to select the target retrieval space.

[0017] According to the second aspect of the present invention, there is provided a multi-turn dialogue interaction system, which includes a machine-readable storage medium and a processor. The machine-readable storage medium stores machine-executable instructions. When the processor executes the machine-executable instructions, the multi-turn dialogue interaction system implements the foregoing multi-turn dialogue interaction method based on context reconstruction and multi-database retrieval.

[0018] According to the third aspect of the present invention, there is provided a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed, the foregoing multi-turn dialogue interaction method based on context reconstruction and multi-database retrieval is implemented.

[0019] According to any one of the above aspects, the technical effect of the present invention is as follows: Embodiments of the present application implement intelligent processing of the entire process from dialogue data parsing to response strategy generation by constructing a multi-turn dialogue interaction method based on context reconstruction and multi-database retrieval, significantly improving the accuracy and fluency of multi-turn dialogue interaction, and providing an innovative technical solution for the optimization and upgrade of dialogue systems. Specifically, by obtaining the multi-turn dialogue data set of the target user and performing context reconstruction processing on it, the semantic focus offset of each dialogue turn and the semantic dependence strength of historical dialogue turns can be accurately captured, providing rich and accurate context features for subsequent retrieval and matching, and effectively solving the problem of context information loss or misunderstanding in traditional dialogue systems when dealing with complex dialogue scenarios. Retrieving multi-database retrieval results that match the context reconstruction features from a preset heterogeneous database based on a dynamic context window not only broadens the retrieval scope and improves the comprehensiveness of retrieval, but also enhances the flexibility and adaptability of retrieval by dynamically adjusting the size of the context window, making the retrieval results more in line with the actual needs of the current dialogue. Further, multi-strategy matching processing is performed on the context reconstruction features and multi-database retrieval results to generate a response generation strategy for the current dialogue turn. This strategy can dynamically adjust the semantic priority and confidence threshold of candidate response statements to ensure that the generated response not only conforms to the context logic of the dialogue but also has a high confidence level and user satisfaction. Finally, the response generation strategy is fed back to the dialogue service system to trigger response optimization operations. By continuously updating the retrieval weights of the heterogeneous database and the dynamic adjustment parameters of the context window, self-learning and continuous optimization of the dialogue system are achieved, effectively improving the performance and user experience of the dialogue system. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 FIG. shows a schematic flowchart of the multi-turn dialogue interaction method based on context reconstruction and multi-database retrieval provided by the embodiments of the present invention; Figure 2 FIG. shows a schematic diagram of the component structure of the multi-turn dialogue interaction system for implementing the above multi-turn dialogue interaction method based on context reconstruction and multi-database retrieval provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] Embodiments of the present invention will be described below with reference to the accompanying drawings in the present invention. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present invention, and do not limit the technical solutions of the embodiments of the present invention.

[0023] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present invention mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the art of the present technology. It should be understood that when an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include a wireless connection or a wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. The technical solutions of the embodiments of the present invention and the technical effects produced by the technical solutions of the present invention will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be referenced, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0025] Figure 1 A flowchart showing a multi-turn dialogue interaction method and system based on context reconstruction and multi-library retrieval provided by an embodiment of the present invention is shown. It should be understood that in other embodiments, the order of some steps of the multi-turn dialogue interaction method based on context reconstruction and multi-library retrieval in this embodiment can be shared according to actual needs, or some of the steps can also be omitted or maintained. The detailed steps of the multi-turn dialogue interaction method based on context reconstruction and multi-library retrieval include: Step S110: Obtain a multi-turn dialogue data set of a target user, where the multi-turn dialogue data set includes a plurality of dialogue turn sequences, and each dialogue turn sequence is composed of a user input statement and a corresponding system response statement.

[0026] In the field of medical data retrieval, the target user mainly refers to an individual who interacts with the dialogue service system in a medical scenario, which could be either a patient or a medical staff member. Taking the patient as an example, when seeking medical help, the patient will conduct multiple rounds of conversations with the medical dialogue service system to consult information related to the disease. The multi-round dialogue data set is actually a detailed record of a series of conversations between the patient and the system.

[0027] Specifically, these sequences of dialogue turns are arranged in chronological order. For example, in the first round of conversation between the patient and the system, the patient's input statement is "I've been coughing a lot lately. What's going on?", which indicates that the patient is currently experiencing coughing symptoms and hopes to understand the possible causes. Based on the preset rules and knowledge reserves, the system gives the response statement "Coughing can be caused by various reasons, such as colds, allergies, etc.", providing the patient with some possible directions for the cause of the disease.

[0028] In the second round of conversation, the patient's input statement is "I don't have any other symptoms of a cold. Could it be an allergy?", which shows that after referring to the system's first-round response and considering their own actual situation, the patient further asks whether it is due to an allergic reaction. The system then gives another response: "It's possible. Coughing caused by allergies usually also comes with symptoms such as sneezing and runny nose", further guiding the patient to judge the allergic situation.

[0029] In the third round of conversation, the patient's input statement is "I also don't have the symptoms of sneezing and runny nose. What else could be the reason?", indicating that after self-examination, the patient has ruled out the possibility of an allergy and continues to seek other possible causes from the system. The system responds with "In addition to colds and allergies, some respiratory diseases may also cause coughing, such as bronchitis", providing the patient with new possible causes of the disease.

[0030] These sequentially arranged dialogue turns together constitute the multi-round dialogue data set, recording the gradually deepening communication process between the patient and the system.

[0031] Step S120: Perform context reconstruction processing on the multi-round dialogue data set to generate context reconstruction features for each sequence of dialogue turns. The context reconstruction features include the semantic focus offset of the current dialogue turn and the semantic dependency strength of the historical dialogue turns.

[0032] After obtaining the multi-round dialogue data set between the patient and the system, it is necessary to perform context reconstruction processing on it. The purpose of this processing is to generate context reconstruction features for each sequence of dialogue turns. The semantic focus offset and semantic dependency strength included in the context reconstruction features can help the system more accurately understand the relationship between the current dialogue turn and the historical dialogue, thus providing a more accurate basis for subsequent retrieval and response.

[0033] Step S121: Extract the local semantic features of the user input statement in the current dialogue turn and the feedback semantic features of the corresponding system response statement.

[0034] Taking the third-round dialogue as an example, the user input statement in the current dialogue turn is "I don't have symptoms of sneezing or runny nose either. Then what could be the reason?". To extract the local semantic features of this statement, a series of natural language processing techniques need to be applied. First is lexical analysis, which tokenizes the statement, splitting it into individual words, resulting in words such as "I", "don't have", "sneezing", "runny nose", "symptoms", "then", "could be", "what", "reason", etc. These words are the basic units of the statement, and subsequent processing will be based on these words.

[0035] Next is syntactic analysis, which analyzes the grammatical relationships between these words. For example, "I" is the subject, "don't have" is the predicate, and "symptoms of sneezing or runny nose" is the object, etc. Through syntactic analysis, the structure of the statement can be clarified, which helps to better understand the semantics of the statement.

[0036] Then is semantic understanding, which maps the tokenized words into a predefined semantic space. Here, a pre-trained word vector model is used, which can map each word into a 300-dimensional vector space. For example, the word "cough" corresponds to a 300-dimensional vector in the word vector model, and each dimension of this vector represents a feature of the word in the semantic space.

[0037] Perform weighted average processing on the vectors of these words. Assume that the weight of each word is determined according to its importance in the statement. For some key words, such as "symptoms" and "reason", higher weights are assigned, while for some function words or auxiliary words, such as "then" and "don't have", lower weights are assigned. Through weighted average, the local semantic feature vector of the user input statement is obtained, which is a 300-dimensional vector and contains the semantic information of the current statement.

[0038] The corresponding system response statement is "In addition to colds and allergies, some respiratory diseases may also cause cough, such as bronchitis". Similarly, perform tokenization, word vector mapping, and weighted average processing on this statement. First, tokenize the statement into words such as "In addition to", "colds", "and", "allergies", "some", "respiratory diseases", "also", "may", "cause", "cough", "such as", "bronchitis", etc., then map these words into a 300-dimensional vector space, and then perform weighted average according to the importance of the words to obtain the feedback semantic feature vector of the system response statement, which is also a 300-dimensional vector.

[0039] Step S122: Traverse the historical dialogue turns in the dialogue turn sequence, perform long-range dependency analysis on the user input statements in each historical dialogue turn, and generate historical semantic chain features, where the historical semantic chain features include the decay weight of the historical semantic focus in time series and the jump association strength.

[0040] Continuing with the third-round dialogue as an example, the historical dialogue turns include the first round and the second round. For the first-round dialogue, the user input statement is "I always cough recently. What's going on?", and long-range dependency analysis is performed on it. Here, a deep learning-based long short-term memory network (LSTM) or gated recurrent unit (GRU) is used to capture the long-range dependency relationships in the statement.

[0041] Input the word vector sequence of this statement into the LSTM model. The LSTM model has memory cells that can remember the dependency relationships between words at different positions in the statement. For example, there is a temporal dependency relationship between the words "recently" and "always", and the LSTM model can capture this relationship. The LSTM model outputs a hidden state vector, which contains the semantic information of the statement and is a 300-dimensional vector.

[0042] Then, according to the time interval between this historical dialogue turn and the current dialogue turn, determine the decay weight of the historical semantic focus in time series. Assume that the preset decay function is an exponential decay function, and its formula is decay weight = base ^ interval turn number, where base is a constant less than 1. Assume base is 0.9. The first-round dialogue and the third-round dialogue are separated by 2 turns, so the decay weight for this turn is 0.9 ^ 2 = 0.81.

[0043] At the same time, calculate the jump association strength between the current dialogue turn and the first-round dialogue. This can be achieved by calculating the similarity between the semantic vectors of the two statements. Here, cosine similarity is used. Calculate the cosine similarity between the local semantic feature vector of the third-round dialogue and the semantic vector of the first-round dialogue. Assume that the cosine similarity between the semantic vector of the first-round dialogue and the local semantic feature vector of the third-round dialogue is 0.6, and use this similarity as the jump association strength.

[0044] For the second round of conversation, long-term dependence analysis, decay weight calculation, and jump correlation strength calculation are also performed. The word vector sequence of the user input statement "I don't have other symptoms of a cold. Could it be an allergy?" in the second round of conversation is input into the LSTM model to obtain a 300-dimensional hidden state vector. Since there is an interval of 1 round between the second round of conversation and the third round of conversation, according to the exponential decay function, the decay weight is 0.9^1 = 0.9. Calculate the cosine similarity between the semantic vector of the second round of conversation and the local semantic feature vector of the third round of conversation. Suppose it is 0.7. This similarity is the jump correlation strength between the current conversation round and the second round of conversation.

[0045] Concatenate the decay weight, jump correlation strength, and semantic vector of each round of historical conversation to obtain the historical semantic chain feature. Suppose the semantic vector of each round is 300-dimensional, and the decay weight and jump correlation strength are each 1-dimensional. Then the historical semantic chain feature of the first round is a 302-dimensional vector, and the historical semantic chain feature of the second round is also a 302-dimensional vector.

[0046] Step S123: Perform consistency verification processing on the local semantic feature, feedback semantic feature, and historical semantic chain feature to obtain the semantic focus offset of the current conversation round. The semantic focus offset is used to quantify the semantic coherence difference between the current conversation round and the historical conversation rounds.

[0047] Step S1231: Calculate the semantic alignment degree between the local semantic feature and the feedback semantic feature. The semantic alignment degree is used to quantify the coverage range of the system response statement for the user input statement.

[0048] Taking the third round of conversation as an example, calculate the semantic alignment degree between the local semantic feature vector of the user input statement in the third round and the feedback semantic feature vector of the system response statement. Use cosine similarity to calculate the similarity between the two vectors and take it as the semantic alignment degree. The calculation of cosine similarity is obtained by dividing the dot product of the two vectors by the product of their magnitudes. Suppose the local semantic feature vector in the third round is V1, the feedback semantic feature vector is V2, their dot product is V1·V2, the magnitude of V1 is |V1|, and the magnitude of V2 is |V2|. Then the semantic alignment degree = V1·V2 / (|V1|*|V2|). Suppose the calculated cosine similarity is 0.7, that is, the semantic alignment degree is 0.7. This means that the coverage range of the system response statement for the user input statement is 70%.

[0049] Step S1232: Traverse the historical semantic chain features of each historical conversation round, and calculate the jump correlation degree between the local semantic feature and the historical semantic chain feature. The jump correlation degree is used to characterize the implicit semantic connection between the current conversation round and the historical conversation rounds.

[0050] For the historical semantic chain features of the first-round historical dialogue, extract the semantic vector part (300 dimensions) and calculate its cosine similarity with the local semantic feature vector of the third round. Also use the calculation formula of cosine similarity. Suppose the obtained similarity is 0.6, and this similarity is the jump correlation degree between the current dialogue turn and the first-round historical dialogue. For the historical semantic chain features of the second-round historical dialogue, similarly extract its semantic vector part and calculate its cosine similarity with the local semantic feature vector of the third round. Suppose it is 0.7, and this similarity is the jump correlation degree between the current dialogue turn and the second-round historical dialogue.

[0051] Step S1233: Construct a consistency scoring function based on the semantic alignment degree and the jump correlation degree. The consistency scoring function is used to measure whether the current dialogue turn deviates from the core semantic path of the historical dialogue turns.

[0052] Step S12331: Normalize the semantic alignment degree to obtain the first sub-score.

[0053] The value range of the semantic alignment degree is usually between 0 and 1. To incorporate it into the consistency scoring function, normalization is required. Here, the linear normalization method is used. Since the semantic alignment degree itself is between 0 and 1, it is directly used as the first sub-score. In the third-round dialogue, the semantic alignment degree is 0.7, so the first sub-score is also 0.7.

[0054] Step S12332: Perform time decay weighting on the jump correlation degree to obtain the second sub-score.

[0055] For the jump correlation degree of 0.6 of the first-round historical dialogue, since there are 2 turns between it and the third-round dialogue, according to the preset time decay function, such as the exponential decay function, assuming the decay coefficient is 0.9, the weighted jump correlation degree is calculated as 0.6 * 0.9^2 = 0.486. For the jump correlation degree of 0.7 of the second-round historical dialogue, since there is 1 turn between it and the third-round dialogue and the decay coefficient is 0.9, the weighted jump correlation degree is calculated as 0.7 * 0.9^1 = 0.63.

[0056] Sum up the weighted jump correlation degrees of all historical dialogue turns. Suppose there are only two rounds of historical dialogue, and the sum is 0.486 + 0.63 = 1.116. Then perform normalization. Suppose the linear normalization method is used, and divide the sum by the number of historical dialogue turns (2) to obtain the second sub-score of approximately 0.558.

[0057] Step S12333: Linearly combine the first sub-score and the second sub-score according to the position weight of the current dialogue turn in the dialogue turn sequence to generate the consistency score.

[0058] Assume that the current conversation turn is the third turn, and its position serial number in the conversation turn sequence is 3. The position weight is proportional to the position serial number of the current conversation turn in the sequence. Assume that the calculation formula of the position weight is position weight = 0.1 * position serial number. Then the position weight of the third turn is 0.1 * 3 = 0.3.

[0059] Consistency score = first sub-score * (1 - position weight) + second sub-score * position weight = 0.7 * (1 - 0.3) + 0.558 * 0.3 = 0.49 + 0.1674 = 0.6574.

[0060] Step S1234: Determine the semantic focus offset amount according to the output value of the consistency scoring function. If the output value is lower than the fifth threshold, it is determined that the current conversation turn has a semantic focus shift, and a corresponding semantic focus offset amount is generated based on the shift direction.

[0061] Assume that the fifth threshold is 0.7, and the consistency score of the third conversation turn is 0.6574, which is lower than the fifth threshold. It is determined that the current conversation turn has a semantic focus shift. To determine the shift direction, the difference between the local semantic features and the historical semantic chain features can be analyzed. By comparing the values of each dimension of the semantic vector, the dimension with a large difference is found. For example, in the 100th dimension, the value of the local semantic feature vector of the third turn is 0.8, while the value of the historical semantic chain feature vector of the first turn in this dimension is 0.3, and the value of the historical semantic chain feature vector of the second turn in this dimension is 0.4. The difference value of this dimension is 0.8 - ((0.3 + 0.4) / 2) = 0.45. This difference value of this dimension is used as an index of the shift direction.

[0062] Then, according to the preset rule, the shift direction is converted into a numerical value. Assume that the rule is to multiply the index of the shift direction by a coefficient of 0.2, resulting in 0.45 * 0.2 = 0.09. Combining with the difference between the consistency score and the fifth threshold, the difference is 0.7 - 0.6574 = 0.0426. The semantic focus offset amount = 0.09 + 0.0426 = 0.1326.

[0063] Step S124: Dynamically intercept the historical semantic chain features based on the preset context window threshold to generate the semantic dependence strength of the historical conversation turn. The semantic dependence strength is used to represent the semantic contribution degree of the intercepted historical semantic chain to the current conversation turn.

[0064] Assume that the preset context window threshold is 2, that is, only the historical semantic chain features of the last 2 historical conversation turns are considered. For the third conversation turn, the historical semantic chain features of the second turn and the first turn are intercepted.

[0065] To calculate the semantic dependency strength, the weighted sum of the historical semantic chain features after truncation can be performed. According to the decay weights of each round of historical conversations, the semantic vector part of its historical semantic chain features is weighted. Suppose the decay weight of the second round is 0.9, the decay weight of the first round is 0.81, the semantic vector of the historical semantic chain features in the second round is V2, and the semantic vector of the historical semantic chain features in the first round is V1. Then the weighted semantic vector is 0.9*V2 + 0.81*V1.

[0066] Then, calculate the cosine similarity between the weighted semantic vector and the local semantic feature vector in the third round. Suppose the calculated cosine similarity is 0.76, and this similarity is used as the semantic dependency strength of the historical conversation round.

[0067] Step S125: Normalize and splice the semantic focus offset and the semantic dependency strength to generate the context reconstruction feature of the conversation round sequence.

[0068] The value ranges of the semantic focus offset and the semantic dependency strength may be different, and normalization processing is required. The semantic focus offset and the semantic dependency strength are both between 0 and 1 themselves, so their values are directly used.

[0069] Splice the semantic focus offset and the semantic dependency strength to obtain a 2D vector as the context reconstruction feature of the conversation round sequence. In the third round of conversation, the semantic focus offset is 0.1326, the semantic dependency strength is 0.76, and the context reconstruction feature vector is [0.1326, 0.76].

[0070] Step S130: Retrieve the multi-database retrieval results matching the context reconstruction feature from a preset heterogeneous database based on a dynamic context window. The multi-database retrieval results include a set of candidate response statements and their corresponding response confidence levels.

[0071] After obtaining the context reconstruction feature of each conversation round sequence, it is necessary to retrieve the matching multi-database retrieval results from a preset heterogeneous database based on a dynamic context window. The heterogeneous database may contain multiple different types of sub-databases, such as a medical knowledge base, a case database, an expert experience database, etc.

[0072] Step S131: Adjust the truncation range of the dynamic context window according to the semantic focus offset to determine the number of valid historical rounds in the current conversation round.

[0073] Taking the third-round conversation as an example, its semantic focus offset is 0.1326. Suppose the mapping relationship between the preset semantic focus offset and the number of effective historical turns is as follows: when the semantic focus offset is less than 0.2, the number of effective historical turns is 2; when the semantic focus offset is between 0.2 and 0.5, the number of effective historical turns is 1; when the semantic focus offset is greater than 0.5, the number of effective historical turns is 0. Since the semantic focus offset of the third-round conversation is 0.1326, which is less than 0.2, it is determined that the number of effective historical turns of the current conversation turn is 2, that is, the last 2 historical conversation turns are considered.

[0074] Step S132: Based on the semantic dependence strength, perform weighted fusion on the historical semantic chain features within the number of effective historical turns to generate a context retrieval vector.

[0075] In the third-round conversation, the number of effective historical turns is 2, that is, the historical semantic chain features of the second round and the first round. The semantic dependence strength is 0.76.

[0076] For the semantic vector V2 of the historical semantic chain features of the second round and the semantic vector V1 of the historical semantic chain features of the first round, perform weighted fusion according to the semantic dependence strength. Suppose the weight of the second round is 0.76 * 0.9 (the decay weight of the second round), and the weight of the first round is 0.76 * 0.81 (the decay weight of the first round). Then the context retrieval vector = 0.76 * 0.9 * V2 + 0.76 * 0.81 * V1.

[0077] Step S133: Invoke the distributed index service of the heterogeneous database to map the context retrieval vector to the retrieval spaces of multiple sub-databases, where the retrieval space of each sub-database corresponds to a data partition of a response type.

[0078] Step S1331: Determine the matching degree between the semantic distribution feature and each sub-database's response type according to the semantic distribution feature of the context retrieval vector.

[0079] Step S13311: Extract the projection component of the context retrieval vector in the preset response type dimension, and the response type dimension includes task type, Q&A type, and chat type.

[0080] Suppose the context retrieval vector is a 300-dimensional vector, and the preset response type dimension is 3-dimensional, corresponding to task type, Q&A type, and casual chat type respectively. To extract the projection components, three basis vectors corresponding to the response type dimensions need to be determined first. Suppose the basis vector for the task type dimension is T, the basis vector for the Q&A type dimension is Q, and the basis vector for the casual chat type dimension is C. These three basis vectors are all 300-dimensional vectors. By calculating the inner product of the context retrieval vector and the basis vectors of each response type dimension, the projection components in each response type dimension are obtained. The calculation method of the inner product is to multiply the values of the corresponding dimensions of the context retrieval vector and the basis vector and then sum them up. For example, for the task type dimension, suppose the context retrieval vector is R and the task type basis vector is T. The projection component of the task type dimension is equal to the first dimension value of R multiplied by the first dimension value of T, plus the second dimension value of R multiplied by the second dimension value of T, and so on until all 300 dimensions are calculated and summed. Suppose after calculation, the projection component in the task type dimension is 0.3, the projection component in the Q&A type dimension is 0.6, and the projection component in the casual chat type dimension is 0.1.

[0081] Step S13312: Calculate the Euclidean distance between the response type label vector of each sub-database and the projection component, where the response type label vector is composed of the statistical values of the historical response statement type distribution of this sub-database.

[0082] Suppose there are three sub-databases in the heterogeneous database, namely sub-database A, sub-database B, and sub-database C. The statistical distribution of the historical response statement types of sub-database A shows that the task type response accounts for 20%, the Q&A type response accounts for 70%, and the casual chat type response accounts for 10%. Then the response type label vector of sub-database A is [0.2, 0.7, 0.1]. The statistical distribution of the historical response statement types of sub-database B shows that the task type response accounts for 40%, the Q&A type response accounts for 50%, and the casual chat type response accounts for 10%. Its response type label vector is [0.4, 0.5, 0.1]. The statistical distribution of the historical response statement types of sub-database C shows that the task type response accounts for 10%, the Q&A type response accounts for 30%, and the casual chat type response accounts for 60%. Its response type label vector is [0.1, 0.3, 0.6].

[0083] Calculate the Euclidean distance between the response type label vector of sub-database A and the projection component [0.3, 0.6, 0.1]. First, calculate the square of the difference in each dimension. That is, the difference in the task dimension is 0.3 - 0.2 = 0.1, and its square is 0.01; the difference in the Q&A dimension is 0.6 - 0.7 = -0.1, and its square is 0.01; the difference in the chat dimension is 0.1 - 0.1 = 0, and its square is 0. Then add these squared values together to get 0.01 + 0.01 + 0 = 0.02. Finally, take the square root of this sum (although the formula editor is not allowed here, the logic is like this) to get approximately 0.1414.

[0084] Similarly, calculate the Euclidean distance between the response type label vector of sub-database B and the projection component. The difference in the task dimension is 0.3 - 0.4 = -0.1, and its square is 0.01; the difference in the Q&A dimension is 0.6 - 0.5 = 0.1, and its square is 0.01; the difference in the chat dimension is 0.1 - 0.1 = 0, and its square is 0. Add them up to get 0.02, and after taking the square root, it is approximately 0.1414.

[0085] Calculate the Euclidean distance between the response type label vector of sub-database C and the projection component. The difference in the task dimension is 0.3 - 0.1 = 0.2, and its square is 0.04; the difference in the Q&A dimension is 0.6 - 0.3 = 0.3, and its square is 0.09; the difference in the chat dimension is 0.1 - 0.6 = -0.5, and its square is 0.25. Add them up to get 0.04 + 0.09 + 0.25 = 0.38, and after taking the square root, it is approximately 0.6164.

[0086] Step S13313: Normalize the Euclidean distance to obtain the initial matching degree of each sub-database.

[0087] Using the linear normalization method, set the maximum value of the Euclidean distance to 1 and the minimum value to 0. The Euclidean distances of sub-database A and sub-database B are the smallest, both being 0.1414, and the Euclidean distance of sub-database C is the largest, being 0.6164.

[0088] The calculation method for the initial matching degree of sub-database A is: subtract the difference between the Euclidean distance of sub-database A and the minimum Euclidean distance divided by the difference between the maximum Euclidean distance and the minimum Euclidean distance from 1, that is, 1 - (0.1414 - 0.1414) / (0.6164 - 0.1414) = 1.

[0089] The initial matching degree of sub-database B is calculated as: 1 - (0.1414 - 0.1414) / (0.6164 - 0.1414) = 1.

[0090] The initial matching degree of sub-database C is calculated as: 1 - (0.6164 - 0.1414) / (0.6164 - 0.1414) = 0.

[0091] Step S13314: Dynamically compensate the initial matching degree based on the semantic focus offset of the current conversation turn. If the semantic focus offset exceeds the eighth threshold, increase the matching degree compensation coefficient of the task-based sub-database and decrease the matching degree compensation coefficient of the chit-chat sub-database simultaneously.

[0092] Suppose the eighth threshold is 0.1, and the semantic focus offset of the third-round conversation is 0.1326, which exceeds the eighth threshold. Set the matching degree compensation coefficient of the task-based sub-database to 0.2, and the matching degree compensation coefficient of the chit-chat sub-database to -0.2.

[0093] Sub-database A contains task-based responses, with an initial matching degree of 1. After compensation, the matching degree is 1 + 0.2 = 1.2, but the matching degree cannot exceed 1, so the matching degree of sub-database A after compensation remains 1.

[0094] Sub-database C contains more chit-chat responses, with an initial matching degree of 0. After compensation, the matching degree is 0 - 0.2 = -0.2, but the matching degree cannot be less than 0, so the matching degree of sub-database C after compensation is 0.

[0095] Sub-database B contains a certain proportion of task-based and Q&A responses. Since there is compensation for the task-based part, assuming the task-based proportion is calculated for compensation at 40%, the compensation amount is 0.2 * 0.4 = 0.08. The matching degree after compensation is 1 + 0.08 = 1.08. Similarly, the matching degree cannot exceed 1, so the matching degree of sub-database B after compensation remains 1.

[0096] Determine the matching degree between the semantic distribution feature and the response type of the sub-database according to the compensated matching degree, and sort the matching degrees in descending order to select the target retrieval space. Here, the matching degrees of sub-database A and sub-database B are 1, and the matching degree of sub-database C is 0, so select sub-database A and sub-database B as the target retrieval space.

[0097] Step S1332: Select the sub-databases with a matching degree higher than the sixth threshold as the target retrieval space, and allocate dynamic retrieval resources to the target retrieval space.

[0098] Suppose the sixth threshold is 0.8, and the matching degrees of sub-database A and sub-database B are 1, which are higher than the sixth threshold, so take sub-database A and sub-database B as the target retrieval space. To improve the retrieval efficiency, allocate dynamic retrieval resources to the target retrieval space according to the matching degree. Suppose the total retrieval resources are 100 units, and the matching degrees of sub-database A and sub-database B are the same, then each is allocated 50 units of retrieval resources.

[0099] Step S1333: In the target retrieval space, perform clustering processing on the candidate response statements based on the dimensionality distribution of the context retrieval vector to generate multiple semantic clusters.

[0100] In sub-database A and sub-database B, there are a large number of candidate response statements. First, each candidate response statement is transformed into a 300-dimensional semantic embedding vector through a word vector model. Then, based on the dimensionality distribution of the context retrieval vector, a clustering algorithm (such as the K-Means clustering algorithm) is used to perform clustering processing on the semantic embedding vectors of these candidate response statements. Suppose they are clustered into 5 semantic clusters, and the candidate response statements in each semantic cluster are semantically similar.

[0101] Step S1334: Select the central candidate response statement from each of the semantic clusters as the retrieval anchor, and calculate the similarity difference between the context retrieval vector and the retrieval anchor.

[0102] In each semantic cluster, the central vector of the semantic cluster is obtained by calculating the average value of the semantic embedding vectors of all candidate response statements within the cluster. The candidate response statement closest to this central vector is selected as the retrieval anchor. For each retrieval anchor, calculate the cosine similarity between its semantic embedding vector and the context retrieval vector. Suppose the context retrieval vector is R, and the semantic embedding vector of the retrieval anchor is S. The cosine similarity calculation formula is the dot product of R and S divided by the product of the norm of R and the norm of S. Suppose the cosine similarities between 5 retrieval anchors and the context retrieval vector are calculated as 0.7, 0.75, 0.8, 0.85, and 0.9 respectively. Suppose the cosine similarity between the context retrieval vector and itself is 1, then the similarity differences between these 5 retrieval anchors and the context retrieval vector are 1 - 0.7 = 0.3, 1 - 0.75 = 0.25, 1 - 0.8 = 0.2, 1 - 0.85 = 0.15, and 1 - 0.9 = 0.1 respectively.

[0103] Step S1335: If the similarity difference is less than the seventh threshold, add all the candidate response statements within the corresponding semantic cluster to the candidate response statement set.

[0104] Suppose the seventh threshold is 0.2, then all the candidate response statements within the semantic clusters corresponding to the two retrieval anchors with similarity differences of 0.15 and 0.1 will be added to the candidate response statement set. Suppose these two semantic clusters have 10 and 15 candidate response statements respectively, then 25 candidate response statements are added to the candidate response statement set.

[0105] Step S134: In the retrieval space of each sub-database, calculate the cosine similarity between the context retrieval vector and the semantic embedding vector of the candidate response statement, and filter out the candidate response statement set with a similarity higher than the first threshold according to the cosine similarity.

[0106] In the retrieval space of sub-database A and sub-database B, for each candidate response statement in the candidate response statement set, the cosine similarity between its semantic embedding vector and the context retrieval vector is calculated. Assuming the first threshold is 0.7, after calculation, the candidate response statements with a cosine similarity higher than 0.7 are screened out. Assuming there are 25 statements in the candidate response statement set, after screening, 18 statements have a cosine similarity higher than 0.7, and these 18 statements constitute a new candidate response statement set.

[0107] Step S135: calibrate the confidence of the selected candidate response statement set to generate the response confidence of the candidate response statement. The confidence calibration process includes: performing weighted calculation based on the frequency of occurrence of the candidate response statement in historical dialogue rounds, semantic conflict degree and user feedback score.

[0108] For the 18 candidate response statements selected, their occurrence frequency, semantic conflict degree, and user feedback score in the historical dialogue rounds are calculated respectively. Assume that the weight of the occurrence frequency is 0.3, the weight of the semantic conflict degree is 0.2, and the weight of the user feedback score is 0.5.

[0109] For a candidate response statement, it appeared 3 times in the historical dialogue rounds. Assuming that according to the preset rules, the frequency score is calculated as the number of appearances divided by the total number of historical dialogue rounds (assuming 10 times), that is, 3 / 10=0.3. The semantic conflict degree is calculated by comparing the semantic embedding vector of the candidate response statement with the semantic vectors of each round in the historical dialogue, finding the dimension with the largest semantic difference, and scoring it according to the degree of difference. Assume that the semantic conflict degree score of the statement is 0.1. The user feedback score is obtained by normalizing the user's score on the statement in the historical dialogue (assuming a 0-10 point system). Assume that the user feedback score of the statement is 0.8.

[0110] Then the response confidence of the candidate response statement is 0.3*0.3+0.2*0.1+0.5*0.8=0.09+0.02+0.4=0.51. This calculation is performed on all 18 candidate response statements to obtain the response confidence of each statement.

[0111] Step S140: performing multi-strategy matching processing on the context reconstruction feature and the multi-database search result to generate a response generation strategy for the current dialogue round, wherein the response generation strategy is used to dynamically adjust the semantic priority and confidence threshold of the candidate response statement.

[0112] Step S141: determining the semantic stability level of the current dialogue turn according to the semantic focus offset, wherein the semantic stability level is used to classify the semantic correction strength of the candidate response sentences.

[0113] Suppose the mapping relationship between the semantic focus shift amount and the semantic stability level is as follows: when the semantic focus shift amount is less than 0.1, the semantic stability level is high; when the semantic focus shift amount is between 0.1 and 0.3, the semantic stability level is medium; when the semantic focus shift amount is greater than 0.3, the semantic stability level is low. The semantic focus shift amount of the third-round conversation is 0.1326, so the semantic stability level is medium.

[0114] For the case where the semantic stability level is medium, the semantic correction intensity for dividing the candidate response statements is medium. This means that for the candidate response statements, a certain degree of semantic correction is required to better adapt to the semantics of the current conversation.

[0115] Step S142: Based on the response confidence, perform priority sorting on the candidate response statement set to generate an initial response priority sequence.

[0116] For the 18 candidate response statements and their corresponding response confidences selected, sort them in descending order of response confidence. Suppose the initial response priority sequence obtained after sorting is: candidate response statement 1 (confidence 0.8), candidate response statement 2 (confidence 0.75), candidate response statement 3 (confidence 0.72)... candidate response statement 18 (confidence 0.5).

[0117] Step S143: Dynamically adjust the initial response priority sequence according to the semantic stability level to generate an adjusted response priority sequence, where if the semantic stability level is lower than the preset level threshold, then increase the priority of the candidate response statement with the highest matching degree with the historical semantic chain features, and at the same time decrease the priority of the candidate response statement with a semantic conflict degree higher than the second threshold.

[0118] Suppose the preset level threshold is high, and the semantic stability level of the third-round conversation is medium, lower than the preset level threshold. First, calculate the cosine similarity between the semantic embedding vector of each candidate response statement and the semantic vector of the historical semantic chain features, and find the candidate response statement with the highest matching degree with the historical semantic chain features. Suppose the candidate response statement 5 has the highest matching degree with the historical semantic chain features, and its cosine similarity is 0.85.

[0119] At the same time, suppose the second threshold is 0.2. For each candidate response statement, calculate its semantic conflict degree, and find the candidate response statement with a semantic conflict degree higher than 0.2. Suppose the semantic conflict degree of candidate response statement 10 is 0.25, higher than the second threshold.

[0120] Promote the priority of candidate response statement 5 to the front of the sequence and lower the priority of candidate response statement 10 to the back of the sequence. The adjusted response priority sequence is: candidate response statement 5 (confidence 0.7), candidate response statement 1 (confidence 0.8), candidate response statement 2 (confidence 0.75) …… candidate response statement 10 (confidence 0.6), candidate response statement 18 (confidence 0.5).

[0121] Step S144: Generate a response generation strategy based on the adjusted response priority sequence and the confidence threshold. The response generation strategy includes: semantic correction rules for response statements, priority update rules, and confidence filtering rules.

[0122] According to the adjusted response priority sequence, formulate semantic correction rules for response statements. Since the semantic stability level is medium, moderately correct the semantics of candidate response statements with higher priorities. For example, for candidate response statement 5, according to the semantic focus of the current conversation, perform some lexical substitutions or sentence structure adjustments on it, but do not change its core semantics.

[0123] The priority update rule is to dynamically adjust the priorities of candidate response statements according to the progress of the conversation. If a certain candidate response statement receives positive feedback from the user in subsequent conversations, increase its priority; if it receives negative feedback, lower its priority.

[0124] The confidence filtering rule is to set a confidence threshold, assumed to be 0.6. For candidate response statements in the adjusted response priority sequence with a confidence lower than 0.6, remove them from the candidate set. In this way, after filtering, the remaining candidate response statements will be used for subsequent conversation responses.

[0125] Step S150: Feed the response generation strategy back to the dialogue service system to trigger response optimization operations, which are used to update the retrieval weights of the heterogeneous database and the dynamic adjustment parameters of the context window.

[0126] Step S151: Update the semantic mapping relationship of the corresponding data partition in the heterogeneous database according to the semantic correction rules.

[0127] According to the semantic correction rules in the response generation strategy, update the semantic mapping relationship of the corresponding data partition in the heterogeneous database. For example, for candidate response statement 5, semantic correction has been performed, and the corrected semantic information is updated to the corresponding position in the database. If candidate response statement 5 comes from a certain data partition of sub-database A, then in this data partition, replace the original semantic mapping relationship with the corrected semantic mapping relationship to ensure that the subsequent retrieved response statements better meet the semantic requirements of the current conversation.

[0128] Step S152: Adjust the retrieval weight allocation parameters of the distributed index service based on the priority update rule, so that data partitions of higher-priority response types obtain larger weight coefficients in subsequent retrievals.

[0129] According to the priority update rule, adjust the retrieval weight allocation parameters of the distributed index service. Suppose in the adjusted response priority sequence, the candidate response statements of the task-based response type have a higher priority. Then in the distributed index service, increase the retrieval weight coefficient of the data partition corresponding to the task-based response type. For example, the original retrieval weight coefficient of the task-based data partition was 0.3, and now it is increased to 0.5. In subsequent retrievals, the task-based data partition will obtain more retrieval resources and it will be easier to retrieve relevant response statements.

[0130] Step S153: Recalibrate the initial confidence levels of each candidate response statement in the heterogeneous database according to the confidence filtering rule, including: degrading the confidence levels of candidate response statements with user feedback scores lower than the third threshold, and removing candidate response statements with semantic conflict degrees exceeding the fourth threshold from the retrieval results.

[0131] Step S1531: Degrade the confidence levels of candidate response statements with user feedback scores lower than the third threshold.

[0132] Suppose the third threshold is 0.6. Obtain the explicit feedback data and implicit feedback data of the candidate response statements in the historical conversation rounds. The explicit feedback data includes user scores or correction instructions, and the implicit feedback data includes whether the user repeats the same semantic input in subsequent conversation rounds.

[0133] For a certain candidate response statement, its explicit feedback data is that the user score is 0.5, and the implicit feedback data is that the user does not repeat the same semantic input in the subsequent conversation. Suppose the weight of the explicit feedback data is 0.7 and the weight of the implicit feedback data is 0.3. First, convert the implicit feedback data into a score. Suppose the score corresponding to the user not repeating the same semantic input in the subsequent conversation is 0.4. Generate a comprehensive feedback score based on the weighted sum of the explicit feedback data and the implicit feedback data, that is, comprehensive feedback score = explicit feedback score × explicit feedback data weight + implicit feedback score × implicit feedback data weight, which is 0.5×0.7 + 0.4×0.3 = 0.35 + 0.12 = 0.47. Since this comprehensive feedback score of 0.47 is lower than the third threshold of 0.6, degrade the confidence level of this candidate response statement to the preset lowest level. Suppose the preset lowest level is 0.2, then the confidence level of this candidate response statement is reduced from the original value (suppose it is 0.65) to 0.2, and it is moved to the pool of response statements to be verified.

[0134] The response pool to be verified is used to store candidate response statements whose confidence has been downgraded and need further verification. Semantic conflict analysis and context adaptability test are performed regularly on the candidate response statements in the response pool to be verified. Semantic conflict analysis is to compare the semantic embedding vector of the candidate response statement with the context retrieval vector of the current conversation and the semantic vector of the historical conversation round to see if there is an obvious semantic conflict. The context adaptability test simulates different conversation scenarios to check whether the candidate response statement can respond reasonably in these scenarios.

[0135] For example, when performing semantic conflict analysis on the candidate response sentence moved into the response pool to be verified, the cosine similarity between its semantic embedding vector and the context retrieval vector of the current third round of dialogue is calculated. Assuming the similarity is 0.3, it indicates that there is a certain semantic difference. Then the cosine similarity is calculated with the semantic vectors of the historical semantic chain features of the first and second rounds of dialogue, assuming they are 0.2 and 0.25 respectively, which also shows that there is a large difference with the semantics of the historical dialogue and there is a semantic conflict.

[0136] In the context adaptability test, different patient consultation scenarios are simulated, such as different symptom combinations, different questioning methods, etc., and the candidate response statement is placed in these scenarios to see if it can respond reasonably. If the candidate response statement cannot adapt well in multiple simulated scenarios, it needs to be further revised or eliminated. If, after a series of tests, it is found that the candidate response statement performs well in both semantic conflict analysis and context adaptability tests, for example, the semantic similarity reaches above 0.7, and it can respond reasonably in the simulated scenario, then restore its original confidence level (here is 0.65), remove it from the response pool to be verified, and put it back into the normal candidate response statement set.

[0137] Step S1532: Eliminate candidate response sentences whose semantic conflict degree exceeds a fourth threshold from the search results.

[0138] Assuming that the fourth threshold is 0.3, for each candidate response statement in the heterogeneous database, calculate its semantic conflict degree. As mentioned above, the calculation method of the semantic conflict degree is to compare the semantic embedding vector of the candidate response statement with the semantic vector of each round in the historical dialogue, find the dimension with the largest semantic difference, and score it according to the degree of difference.

[0139] For example, there is a candidate response statement. The difference value between its semantic embedding vector and the semantic vector of the historical semantic chain feature of the first-round conversation in a certain key dimension is 0.2, the difference value between its semantic embedding vector and the semantic vector of the historical semantic chain feature of the second-round conversation in another key dimension is 0.15, and the difference value between its semantic embedding vector and the context retrieval vector of the third-round conversation in a certain dimension is 0.1. These difference values are combined according to the preset rules to calculate the semantic conflict degree. Assuming that the obtained semantic conflict degree is 0.25, which is less than the fourth threshold of 0.3, then this candidate response statement can be retained in the retrieval results.

[0140] For another candidate response statement, after calculation, its semantic conflict degree is 0.35, which exceeds the fourth threshold of 0.3. Then this candidate response statement is excluded from the retrieval results. This can ensure that the subsequent retrieved candidate response statements have higher semantic consistency with the historical conversation and the current conversation, improving the quality and accuracy of the conversation service.

[0141] Figure 2 Fig. shows a multi-round conversation interaction system 100 provided in an embodiment of the present invention, including a processor 1001, a memory 1003, and program code stored on the memory 1003. The processor 1001 executes the above program code to implement the steps of the multi-round conversation interaction method based on context reconstruction and multi-library retrieval.

[0142] Figure 2 The multi-round conversation interaction system 100 shown includes: a processor 1001 and a memory 1003. Among them, the processor 1001 and the memory 1003 are connected, such as through a bus 1002. Optionally, the multi-round conversation interaction system 100 may further include a transceiver 1004. The transceiver 1004 can be used for data interaction between this multi-round conversation interaction system and other multi-round conversation interaction systems, such as data sending and / or data receiving, etc. It should be noted that in actual scheduling, the transceiver 1004 is not limited to one, and the structure of this multi-round conversation interaction system 100 does not constitute a limitation to the embodiment of the present invention.

[0143] The processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present invention. The processor 1001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0144] The bus 1002 may include a path for transmitting information between the above components. The bus 1002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 1002 may be divided into an address bus, a data bus, a control bus, etc.

[0145] The memory 1003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to have or store program code and can be read by a computer, which is not limited herein.

[0146] The memory 1003 is used to store the program code for implementing the embodiments of the present invention and is controlled by the processor 1001 for execution. The processor 1001 is used to execute the program code stored in the memory 1003 to implement the steps shown in the foregoing method embodiments.

[0147] An embodiment of the present invention provides a computer-readable storage medium, on which program code is stored. When the program code is executed by a processor, the steps and corresponding content of the foregoing method embodiment can be implemented.

[0148] It should be understood that although the flowchart of the embodiment of the present invention indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless there is a clear description in this article, in some implementation scenarios of the embodiment of the present invention, the implementation steps in each flowchart can be executed in other orders based on requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages according to the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiment of the present invention does not limit this.

[0149] The above are only optional implementation manners of some implementation scenarios of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present invention, using other similar implementation means based on the technical idea of the present invention also belongs to the protection scope of the embodiment of the present invention.

Claims

1. A multi-round dialogue interaction method based on context reconstruction and multi-database retrieval, characterized in that: The method comprises: Acquire a multi-round dialogue data set of a target user, wherein the multi-round dialogue data set includes a plurality of dialogue turn sequences, each dialogue turn sequence consisting of a user input statement and a corresponding system response statement; Performing context reconstruction processing on the multi-round dialogue data set to generate context reconstruction features of each dialogue turn sequence, wherein the context reconstruction features include a semantic focus offset of a current dialogue turn and a semantic dependency strength of historical dialogue turns; Retrieving a multi-database search result matching the context reconstruction feature from a preset heterogeneous database based on a dynamic context window, wherein the multi-database search result includes a set of candidate response statements and corresponding response confidences; Performing multi-strategy matching processing on the context reconstruction feature and the multi-database search result to generate a response generation strategy for the current dialogue round, wherein the response generation strategy is used to dynamically adjust the semantic priority and confidence threshold of the candidate response statement; The response generation strategy is fed back to the dialog service system to trigger a response optimization operation, wherein the response optimization operation is used to update the retrieval weight of the heterogeneous database and the dynamic adjustment parameters of the context window.

2. The multi-round dialogue interaction method based on context reconstruction and multi-database retrieval according to claim 1 is characterized in that: The performing context reconstruction processing on the multi-round dialogue data set to generate context reconstruction features of each dialogue turn sequence includes: Extracting local semantic features of the user input sentence of the current dialogue round and feedback semantic features of the corresponding system response sentence; Traversing the historical dialogue turns of the dialogue turn sequence, performing long-range dependency analysis on the user input sentence of each historical dialogue turn, and generating a historical semantic chain feature, wherein the historical semantic chain feature includes a decay weight and a jump association strength of a historical semantic focus in time sequence; Performing consistency verification processing on the local semantic features, feedback semantic features, and historical semantic chain features to obtain a semantic focus offset of the current dialogue turn, wherein the semantic focus offset is used to quantify the difference in semantic coherence between the current dialogue turn and the historical dialogue turn; Dynamically intercepting the historical semantic chain features based on a preset context window threshold to generate the semantic dependency strength of the historical dialogue round, wherein the semantic dependency strength is used to characterize the semantic contribution of the intercepted historical semantic chain to the current dialogue round; The semantic focus offset and the semantic dependency strength are normalized and concatenated to generate context reconstruction features of the dialogue turn sequence.

3. The multi-round dialogue interaction method based on context reconstruction and multi-database retrieval according to claim 2 is characterized in that: The consistency verification process of the local semantic features, the feedback semantic features and the historical semantic chain features is performed to obtain the semantic focus offset of the current dialogue turn, including: Calculating the semantic alignment between the local semantic feature and the feedback semantic feature, wherein the semantic alignment is used to quantify the coverage of the system response sentence to the user input sentence; Traversing the historical semantic chain features of each historical dialogue turn, calculating the jump correlation between the local semantic features and the historical semantic chain features, wherein the jump correlation is used to characterize the implicit semantic connection between the current dialogue turn and the historical dialogue turn; Building a consistency scoring function based on the semantic alignment and the jump correlation, wherein the consistency scoring function is used to measure whether the current dialogue turn deviates from the core semantic path of the historical dialogue turn; The semantic focus offset is determined according to the output value of the consistency scoring function. If the output value is lower than the fifth threshold, it is determined that a semantic focus offset occurs in the current dialogue turn, and a corresponding semantic focus offset is generated based on the offset direction.

4. The multi-round dialogue interaction method based on context reconstruction and multi-database retrieval according to claim 3 is characterized in that: The constructing a consistency scoring function based on the semantic alignment and the jump association includes: Normalizing the semantic alignment to obtain a first sub-score; Performing time-attenuated weighted processing on the jump correlation to obtain a second sub-score; Linearly combine the first sub-score and the second sub-score according to the position weight of the current dialogue turn in the dialogue turn sequence to generate a consistency score; The weight coefficient of the time decay weighted processing is inversely proportional to the number of intervals between the historical dialogue rounds and the current dialogue round, and the position weight is proportional to the position number of the current dialogue round in the sequence.

5. The multi-round dialogue interaction method based on context reconstruction and multi-database retrieval according to claim 1 is characterized in that: The retrieving multi-database search results matching the context reconstruction feature from preset heterogeneous databases based on the dynamic context window includes: adjusting the interception range of the dynamic context window according to the semantic focus offset to determine the number of valid historical rounds of the current dialogue round; Based on the semantic dependency strength, weighted fusion is performed on the historical semantic chain features within the valid historical round number to generate a context retrieval vector; Invoke the distributed index service of the heterogeneous database to map the context search vector to the search space of multiple sub-databases, wherein the search space of each sub-database corresponds to a data partition of a response type; In the retrieval space of each sub-database, the cosine similarity between the context retrieval vector and the semantic embedding vector of the candidate response sentence is calculated, and a set of candidate response sentences having a similarity higher than a first threshold is screened out according to the cosine similarity; The confidence calibration process is performed on the screened candidate response sentence set to generate the response confidence of the candidate response sentence, and the confidence calibration process includes: performing weighted calculation based on the occurrence frequency, semantic conflict degree and user feedback score of the candidate response sentence in the historical dialogue rounds.

6. The multi-round dialogue interaction method based on context reconstruction and multi-database retrieval according to claim 5 is characterized in that: The calling of the distributed index service of the heterogeneous database to map the context search vector to the search space of multiple sub-databases includes: Determining, according to the semantic distribution feature of the context retrieval vector, a matching degree between the semantic distribution feature and the response type of each sub-database; Selecting a sub-database with a matching degree higher than a sixth threshold as a target search space, and allocating dynamic search resources to the target search space; In the target retrieval space, clustering the candidate response sentences based on the dimensional distribution of the context retrieval vector to generate a plurality of semantic clusters; Selecting a central candidate response sentence from each of the semantic clusters as a retrieval anchor point, and calculating a similarity difference between the context retrieval vector and the retrieval anchor point; If the similarity difference is less than the seventh threshold, all candidate response statements in the corresponding semantic cluster are added to the candidate response statement set.

7. The multi-round dialogue interaction method based on context reconstruction and multi-database retrieval according to claim 1 is characterized in that: The performing multi-strategy matching processing on the context reconstruction feature and the multi-database search result to generate a response generation strategy for the current dialogue round includes: Determining a semantic stability level of the current dialogue turn according to the semantic focus offset, wherein the semantic stability level is used to classify the semantic revision strength of the candidate response sentences; Prioritizing the candidate response statement set based on the response confidence to generate an initial response priority sequence; Dynamically adjusting the initial response priority sequence according to the semantic stability level to generate an adjusted response priority sequence, wherein if the semantic stability level is lower than a preset level threshold, the priority of the candidate response statement with the highest matching degree with the historical semantic chain feature is increased, and the priority of the candidate response statement with a semantic conflict degree higher than a second threshold is decreased; A response generation strategy is generated based on the adjusted response priority sequence and confidence threshold, wherein the response generation strategy includes: a semantic modification rule of the response statement, a priority update rule, and a confidence filtering rule.

8. The multi-round dialogue interaction method based on context reconstruction and multi-database retrieval according to claim 7 is characterized in that: Feeding back the response generation strategy to the dialogue service system to trigger a response optimization operation includes: Updating the semantic mapping relationship of the corresponding data partition in the heterogeneous database according to the semantic correction rule; Adjusting the retrieval weight allocation parameters of the distributed index service based on the priority update rule so that data partitions with higher priority response types obtain larger weight coefficients in subsequent retrievals; Recalibrating the initial confidence of each candidate response statement in the heterogeneous database according to the confidence filtering rule, including: downgrading the confidence of the candidate response statement whose user feedback score is lower than the third threshold, and removing the candidate response statement whose semantic conflict exceeds the fourth threshold from the retrieval results; The step of downgrading the confidence of the candidate response statements whose user feedback scores are lower than the third threshold includes: Obtaining explicit feedback data and implicit feedback data of the user for the candidate response statement in the historical dialogue rounds, wherein the explicit feedback data includes user ratings or correction instructions, and the implicit feedback data includes whether the user repeats the same semantic input in subsequent dialogue rounds; generating a comprehensive feedback score according to a weighted sum of the explicit feedback data and the implicit feedback data; If the comprehensive feedback score is lower than the third threshold, the confidence of the candidate response statement is downgraded to a preset minimum level and moved into the response pool to be verified; Semantic conflict analysis and context adaptability test are regularly performed on candidate response statements in the response pool to be verified, and if the test passes, the original confidence level is restored.

9. The multi-round dialogue interaction method based on context reconstruction and multi-database retrieval according to claim 6 is characterized in that: Determining the matching degree between the semantic distribution feature and the response type of each sub-database includes: Extracting a projection component of the context retrieval vector on a preset response type dimension, where the response type dimension includes a task type, a question-and-answer type, and a chat type; Calculating the Euclidean distance between the response type label vector of each sub-database and the projection component, wherein the response type label vector is composed of the historical response statement type distribution statistics of the sub-database; Normalizing the Euclidean distance to obtain an initial matching degree for each sub-database; Dynamically compensate the initial matching degree based on the semantic focus offset of the current dialogue round. If the semantic focus offset exceeds the eighth threshold, increase the matching degree compensation coefficient of the task-based sub-database and reduce the matching degree compensation coefficient of the chat-based sub-database. The matching degree between the semantic distribution feature and the response type of the sub-database is determined according to the compensated matching degree, and the matching degrees are arranged in descending order to select a target retrieval space.

10. A multi-round dialogue interaction system, characterized in that: It includes a processor and a computer-readable storage medium, wherein the computer-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by the processor, the multi-round dialogue interaction method based on context reconstruction and multi-library retrieval described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Automatic round description in multi-round conversations

    CN114762038A

  • Multi-round dialogue method and device, equipment and storage medium

    CN118093796A

  • Dynamic question and answer matching method and system based on context awareness

    CN119691133A

  • Context-based multi-turn dialogue method and storage medium

    US20210200961A1

Cited By

  • Intelligent number-asking rapid construction optimization system for power production scene

    CN120371980A

  • Hierarchical information deep mining and matching method applied to purchase service system

    CN120525451A

  • Hierarchical information deep mining and matching method applied to procurement service system

    CN120525451B

  • Text semantic understanding analysis method and system applied to smart campus platform

    CN120597894A

  • Text semantic understanding analysis method and system applied to smart campus platform

    CN120597894B