Intelligent interactive question answering method and system based on multimodal large model intention recognition

Through the multimodal large-model intention recognition method, the problem of semantic offset and intent understanding of intelligent question-and-answer system in long dialogue scenarios is solved, and the accurate identification and personalized interaction of user intentions is achieved, and the system's response efficiency and accuracy is improved.

CN120353980AActive Publication Date: 2025-07-22BEIJING FEIRUI XINGTU TECH CO LTD

Patent Information

Application Number
CN202510850508.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Existing intelligent Q&A systems are prone to semantic offsets in long conversation scenarios, cannot accurately understand complex user intentions of multiple modalities, and lack personalized interactive experiences.

Method used

The multimodal large model intention recognition method is adopted to build a cross-modal semantic mapping matrix by obtaining user multimodal interaction information, calculating the amount of mutual information, fusing features and matching with the preset knowledge base, identifying user intentions, building an associated network to extract key information, and generating response content.

Benefits of technology

It improves the accuracy of user intention understanding and coherence of interaction, reduces the ambiguity and misunderstanding rate, and improves the pertinence and quality of response content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353980A_ABST
    Figure CN120353980A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent interactive question and answer method and system based on multi-modal large model intention recognition, and relates to the technical field of intelligent question and answer, and the method comprises the steps of feature extraction, cross-modal semantic mapping, intention matching, key information extraction and response generation, achieves the unified processing of multi-modal information, can accurately recognize the intention of a user, and improves the user experience. And the association network is dynamically constructed to extract key information, so that the accuracy and semantic coherence of intelligent questions and answers are improved, and the man-machine interaction experience is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent question - answering, and particularly to an intelligent interactive question - answering method and system based on the intention recognition of a multimodal large model. Background Art

[0002] With the rapid development of artificial intelligence technology, intelligent question - answering systems have become an important way of human - machine interaction and are widely used in multiple fields such as customer service, education, and medical care. Traditional question - answering systems mainly process information based on a single modality, such as text or voice, and it is difficult to meet the requirements of modern complex interaction scenarios. In recent years, multimodal large - model technology has improved the system's ability to understand user intentions and the naturalness of interaction by integrating multiple modality information such as text, images, and voice; However, existing large - model technologies still have problems such as being unable to identify and retain truly critical historical information, resulting in semantic drift or key - information forgetting in long - conversation scenarios, being unable to accurately understand complex user intentions containing multiple modalities when there are information inconsistencies or complementary relationships between modalities, and lacking sensitivity to intention conversion and intention evolution in multi - turn interactions, and being unable to provide a coherent and personalized interaction experience; Therefore, there is an urgent need for a solution to solve the problems existing in the prior art. Summary of the Invention

[0003] Embodiments of the present invention provide an intelligent interactive question - answering method and system based on the intention recognition of a multimodal large model, which can at least solve some of the problems existing in the prior art.

[0004] In the first aspect of the embodiments of the present invention, an intelligent interactive question - answering method based on the intention recognition of a multimodal large model is provided, including: Obtain the multimodal interaction information input by the user and perform feature extraction to obtain a multimodal feature vector; Calculate the mutual information amount between different modality feature vectors, construct a cross - modality semantic mapping matrix based on the mutual information amount, use the cross - modality semantic mapping matrix to map different modality feature vectors to a unified semantic space, obtain a unified semantic feature, calculate the semantic similarity with the prior knowledge in the preset knowledge base, determine the supplementary information, and fuse the supplementary information with the unified semantic feature to obtain a fusion feature; Construct a user intention feature vector according to the fusion feature, perform similarity matching between the user intention feature vector and the preset intention category library, and select the category with the highest similarity as the user intention category; Calculate the importance distribution of key entities in the historical dialogue based on word frequency and information gain, use a sliding window to count the co-occurrence frequencies of entities and events and construct an association network, perform pruning based on the association strength to obtain the core semantic skeleton, extract the key path and calculate the semantic coherence, sort the importance of the historical dialogue information by combining the path weight and coherence, and determine the key information according to the sorting result; Generate a response content based on the key information and return it to the customer, and update the association network based on the current dialogue content.

[0005] In an alternative embodiment, Calculate the mutual information between different modality feature vectors, construct a cross-modal semantic mapping matrix based on the mutual information, use the cross-modal semantic mapping matrix to map different modality feature vectors to a unified semantic space, obtain the unified semantic features and calculate the semantic similarity with the prior knowledge in the preset knowledge base, determine the supplementary information and fuse the supplementary information with the unified semantic features to obtain the fused features, including: Obtain the multi-modal feature vectors to be processed, where the multi-modal feature vectors include text feature vectors and image feature vectors; Calculate the joint probability distribution and marginal probability distribution of the text feature vector and the image feature vector, and calculate the mutual information between the text feature vector and the image feature vector based on the joint probability distribution and the marginal probability distribution; Construct a cross-modal semantic mapping matrix according to the mutual information, where the value of each element in the cross-modal semantic mapping matrix is calculated by a normalized exponential function adjusted by a temperature parameter, and the temperature parameter is used to control the distribution of the semantic association strength in the cross-modal semantic mapping matrix, and map the text feature vector and the image feature vector to a unified semantic space based on the cross-modal semantic mapping matrix to obtain the unified semantic features; Calculate the similarity between the unified semantic features and the prior knowledge in the preset knowledge base, use the prior knowledge with a similarity greater than a preset threshold as supplementary information, use the unified semantic features as query vectors, use the supplementary information as key-value pairs, and perform information fusion through a multi-head attention mechanism to obtain the fused features.

[0006] In an alternative embodiment, Calculating the similarity between the unified semantic features and the prior knowledge in the preset knowledge base, and using the prior knowledge with a similarity greater than a preset threshold as supplementary information includes: Decompose the unified semantic feature into a core semantic variable set and a context semantic variable set, construct a conditional dependence graph based on the core semantic variable set and the context semantic variable set, calculate the initial probability distribution parameters by using maximum likelihood estimation according to the structure of the conditional dependence graph, and iteratively optimize the initial probability distribution parameters by combining the expectation maximization algorithm to obtain the optimized probability distribution parameters; Calculate the initial similarity value between each piece of prior knowledge in the preset knowledge base and the unified semantic feature according to the optimized probability distribution parameters, and approximately calculate the final similarity value based on the Monte Carlo sampling method; Calculate the similarity difference between different pieces of prior knowledge to obtain a similarity difference value, construct an adaptive similarity threshold according to the similarity difference value, dynamically adjust the adaptive similarity threshold according to the feedback score of the prior knowledge to obtain an adjusted similarity threshold, and determine the prior knowledge with the final similarity value greater than the adjusted similarity threshold as candidate supplementary information; For each candidate supplementary information, calculate a joint similarity value based on the final similarity value and the similarity difference value, and select the candidate supplementary information with the maximum joint similarity value as the supplementary information for output.

[0007] In an alternative embodiment, Constructing a user intention feature vector according to the fusion feature, and performing similarity matching between the user intention feature vector and a preset intention category library, and selecting the category with the highest similarity as the user intention category includes: Obtain the fusion feature and map the fusion feature to the semantic space through a non-linear transformation to obtain an initial intention feature, and perform dimensionality reduction processing on the initial intention feature to obtain the user intention feature vector; Calculate the similarity between the user intention feature vector and each intention category in the preset intention category library to obtain a category similarity value, perform normalization processing on the category similarity value to obtain a probability distribution, and select the intention category with the maximum probability as the user intention category for output.

[0008] In an alternative embodiment, Calculate the importance distribution of key entities in the historical dialogue based on word frequency and information gain, use a sliding window to count the co-occurrence frequencies of entities and events and construct an association network, perform pruning based on the association strength to obtain a core semantic skeleton, extract key paths and calculate semantic coherence, combine the path weights and coherence to sort the importance of the historical dialogue information, and determine the key information according to the sorting result includes: Obtain the historical conversation content, extract the entities in the historical conversation content as candidate entities, calculate the word frequency score and information gain score of the candidate entities, obtain the entity importance distribution by the weighted sum of the word frequency score and the information gain score, and determine the key entities according to the entity importance distribution; Determine the sliding window size according to the length of the historical conversation content, count the co-occurrence frequencies of the key entities and events within the sliding window, calculate the association strength based on the co-occurrence frequencies and construct an association network, set the weights of the association network based on the association strength and the key entities, determine the pruning threshold according to the weight mean and standard deviation, and retain the associations in the association network with weights greater than the pruning threshold to obtain the core semantic skeleton; Determine the path length according to the number of shortest paths between nodes in the core semantic skeleton, set the attenuation factor based on the path length, calculate the node centrality score using the path length and the attenuation factor to extract the key path, and calculate the cosine similarity of the semantic vectors of adjacent nodes in the key path to obtain the semantic coherence; Obtain the historical conversation information based on the key path, calculate the centrality score of the key path in the historical conversation information to obtain the path weight, combine the weighted sum of the path weight and the semantic coherence to obtain the importance ranking value, and sort the historical conversation information according to the importance ranking value and determine the key information.

[0009] In an alternative embodiment, Calculating the association strength based on the co-occurrence frequencies and constructing an association network, setting the weights of the association network based on the association strength and the key entities, determining the pruning threshold according to the weight mean and standard deviation, and retaining the associations in the association network with weights greater than the pruning threshold to obtain the core semantic skeleton includes: Obtain the co-occurrence frequencies of entities in the historical conversation content, divide the co-occurrence frequencies by the square root of the product of the respective occurrence times of the entities to obtain the association strength, and construct an association network based on the association strength; Calculate the importance scores of the key entities, and set the product of the association strength and the average value of the key entity importance scores as the weights of the association network; Use the community discovery algorithm to partition the association network into multiple communities, count the number of edge sets and node sets within each community, and divide the number of edge sets by half of the product of the number of node sets and the number of node sets minus one to obtain the internal connection density of the community; Calculate the mean and standard deviation of all weights in the association network, take the sum of the product of the mean and standard deviation of the weights and the first preset coefficient as the benchmark threshold, and obtain the adaptive pruning threshold corresponding to each community by adding the product of the benchmark threshold and the internal connection density of the community and the second preset coefficient; Compare the relationship between the adaptive pruning threshold of the community to which each edge in the association network belongs and the weight of the current edge, retain the edges with weights greater than the adaptive pruning threshold of the corresponding community, and remove the edges with weights less than the adaptive pruning threshold of the corresponding community to obtain the core semantic skeleton.

[0010] In an alternative embodiment, Generating a response content based on the key information and returning it to the customer, and updating the association network based on the current conversation content includes: Generating a response content including the conversation intention, user requirements, and solution based on the key information, and returning the response content to the customer; Extracting entity, attribute, and relationship triples from the current conversation content to obtain new entities, calculating the co-occurrence frequency and semantic similarity between the new entities to obtain the weight between entities, and calculating the association strength between the new entities and the original entities to obtain the interaction weight; Adding the new entities as nodes and the weight between entities and the interaction weight as edges to the original association network, recalculating the adaptive pruning threshold for the extended association network based on the internal connection density of the community, and retaining the edges with weights greater than the adaptive pruning threshold of the corresponding community to obtain the updated association network.

[0011] In the second aspect of the embodiments of the present invention, an intelligent interactive question-answering system based on multi-modal large model intention recognition is provided, including: A first unit for obtaining multi-modal interaction information input by a user and performing feature extraction to obtain a multi-modal feature vector; A second unit for calculating the mutual information between different modal feature vectors, constructing a cross-modal semantic mapping matrix based on the mutual information, using the cross-modal semantic mapping matrix to map different modal feature vectors to a unified semantic space, obtaining a unified semantic feature and calculating the semantic similarity with prior knowledge in a preset knowledge base, determining supplementary information and fusing the supplementary information with the unified semantic feature to obtain a fused feature; A third unit for constructing a user intention feature vector based on the fused feature, performing similarity matching between the user intention feature vector and a preset intention category library, and selecting the category with the highest similarity as the user intention category; A fourth unit for calculating the importance distribution of key entities in the historical conversation based on word frequency and information gain, using a sliding window to count the co-occurrence frequency of entities and events and constructing an association network, performing pruning based on the association strength to obtain a core semantic skeleton, extracting a key path and calculating semantic coherence, sorting the importance of the historical conversation information by combining the path weight and coherence, and determining key information according to the sorting result; A fifth unit for generating a response content based on the key information and returning it to the customer, and updating the association network based on the current conversation content.

[0012] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including: a processor and a memory for storing processor-executable instructions, wherein the processor is configured to call the instructions stored in the memory to execute the method described above.

[0013] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0014] In the present invention, through multi-modal feature extraction and cross-modal semantic mapping, the unified representation and processing of different modal information are realized, effectively solving the problems of limited information acquisition and incomplete expression in traditional single-modal interaction methods, improving the accuracy and comprehensiveness of understanding user intentions. Based on the cross-modal semantic mapping matrix and correlation network of mutual information, the evolution of user intentions and context relationships are dynamically captured, enabling accurate identification of user intentions and extraction of key information, enhancing the coherence and context understanding ability of the interaction, reducing the ambiguity and misunderstanding rate in the interaction process. By using a sliding window to statistically analyze the co-occurrence frequency of entities and events and combining semantic coherence evaluation, intelligent screening and importance ranking of historical dialogue information are realized, reducing interference from irrelevant information, improving the pertinence and quality of response content, while reducing computational resource consumption and improving system response efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic flowchart of an intelligent interactive question-answering method based on multi-modal large model intention recognition according to an embodiment of the present invention; Figure 2 It is an engineering simulation diagram of dynamically constructing a conditional dependency graph according to an embodiment of the present invention; Figure 3 It is a comparison chart of the change trend of F1 scores in different dialogue rounds according to an embodiment of the present invention; Figure 4 It is a flowchart of constructing the core semantic skeleton of an intelligent interactive question-answering method based on multi-modal large model intention recognition according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0017] The technical solution of the present invention will be described in detail below with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0018] Figure 1 It is a schematic flowchart of an intelligent interactive question-answering method based on multi-modal large model intention recognition according to an embodiment of the present invention. As Figure 1 shown, the method includes: Obtain the multi-modal interaction information input by the user and perform feature extraction to obtain a multi-modal feature vector; Calculate the mutual information amount between different modal feature vectors, construct a cross-modal semantic mapping matrix based on the mutual information amount, use the cross-modal semantic mapping matrix to map different modal feature vectors to a unified semantic space, obtain a unified semantic feature and calculate the semantic similarity with the prior knowledge in the preset knowledge base, determine supplementary information and fuse the supplementary information with the unified semantic feature to obtain a fused feature; Construct a user intention feature vector according to the fused feature, perform similarity matching between the user intention feature vector and the preset intention category library, and select the category with the highest similarity as the user intention category; Calculate the importance distribution of key entities in the historical dialogue based on word frequency and information gain, use a sliding window to count the co-occurrence frequencies of entities and events and construct an association network, perform pruning based on the association strength to obtain a core semantic skeleton, extract the key path and calculate the semantic coherence, sort the importance of the historical dialogue information by combining the path weight and coherence, and determine the key information according to the sorting result; Generate a response content based on the key information and return it to the customer, and update the association network based on the current dialogue content.

[0019] In an alternative embodiment, Calculating the mutual information amount between different modal feature vectors, constructing a cross-modal semantic mapping matrix based on the mutual information amount, using the cross-modal semantic mapping matrix to map different modal feature vectors to a unified semantic space, obtaining a unified semantic feature and calculating the semantic similarity with the prior knowledge in the preset knowledge base, determining supplementary information and fusing the supplementary information with the unified semantic feature to obtain a fused feature includes: Obtain the multi-modal feature vector to be processed, where the multi-modal feature vector includes a text feature vector and an image feature vector; Calculate the joint probability distribution and marginal probability distribution of the text feature vector and the image feature vector, and calculate the mutual information amount between the text feature vector and the image feature vector based on the joint probability distribution and the marginal probability distribution; Construct a cross-modal semantic mapping matrix according to the mutual information quantity. The value of each element in the cross-modal semantic mapping matrix is calculated by a normalized exponential function adjusted by a temperature parameter. The temperature parameter is used to control the distribution of the semantic association strength in the cross-modal semantic mapping matrix. Based on the cross-modal semantic mapping matrix, map the text feature vector and the image feature vector to a unified semantic space to obtain a unified semantic feature; Calculate the similarity between the unified semantic feature and the prior knowledge in the preset knowledge base. Take the prior knowledge with a similarity greater than the preset threshold as supplementary information. Take the unified semantic feature as a query vector and the supplementary information as key-value pairs, and perform information fusion through a multi-head attention mechanism to obtain a fused feature.

[0020] Obtain the multi-modal feature vectors to be processed, including text feature vectors and image feature vectors. Among them, the text feature vectors are extracted by passing the text content input by the user through a pre-trained BERT model. For example, input the sentence "An orange cat is running on the grass" into the BERT model, and a 768-dimensional text feature vector is obtained after multi-layer attention encoding; while the image feature vectors are obtained by using a pre-trained ResNet50 network to extract features from the input images. For example, a color image containing a cat is processed through a convolutional network to obtain a 2048-dimensional image feature vector.

[0021] Use multiple sets of text-image pairs collected as training samples. For each pair of samples, calculate the dot product of the text feature vector and the image feature vector and apply the softmax function to obtain the association probability matrix between them. For example, assume there is a set of 100 text-image samples. Calculate the dot product of each pair of text-image features to construct a 100×100 association matrix. The value of each element in the matrix represents the association degree of the corresponding text-image pair. By normalizing the rows and columns respectively, the system obtains the joint probability distribution P(T, I), where T represents the text feature and I represents the image feature. Calculate the marginal probability distributions P(T) and P(I), that is, the cumulative values of the joint probability in the image dimension and the text dimension respectively. The mutual information quantity MI(T; I) is calculated by dividing the joint probability distribution of the two by the logarithmic expectation value of their respective marginal probability distributions, reflecting the mutual dependence degree between the two modal features.

[0022] According to the calculated mutual information, a cross-modal semantic mapping matrix M is constructed. Each element Mij of the matrix is calculated by a normalized exponential function adjusted by the temperature parameter τ: Mij is equal to exp(mutual informationij / τ) divided by the sum of all exp(mutual informationik / τ). Here, τ is the temperature parameter used to control the distribution of semantic association strength. When the value of τ is small (e.g., τ = 0.1), the higher mutual information in the matrix will be further amplified, making the semantic mapping more focused on strongly related feature pairs; while when the value of τ is large (e.g., τ = 10), the distribution of the mapping matrix will be more uniform, allowing for a wider range of cross-modal associations. In practical applications, the system determines the optimal value of τ as 0.5 through validation set experiments, which can maintain the main semantic associations without overly suppressing secondary associations.

[0023] Using the constructed cross-modal semantic mapping matrix, map the text feature vector and the image feature vector to a unified semantic space. Calculate the product of the text feature vector and the mapping matrix to obtain the mapped feature of the text, and calculate the product of the image feature vector and the transpose of the mapping matrix to obtain the mapped feature of the image. Weightedly fuse these two mapped features, and the weights are determined according to their respective feature confidences. Usually, the text and image feature weights are set to 0.6 and 0.4 respectively to obtain a unified-dimensional fused semantic feature vector, which contains the semantic information of both the text and the image. For example, the original 768-dimensional text feature and 2048-dimensional image feature are mapped to the same 512-dimensional unified semantic space.

[0024] Calculate the similarity between the obtained unified semantic feature and the prior knowledge in the preset knowledge base. This knowledge base contains knowledge entries in multiple fields, and each entry is stored in vector form. Adopt the cosine similarity calculation method to calculate the similarity between the unified semantic feature vector and each knowledge vector in the knowledge base. For example, for the unified semantic feature generated from the image and text description of a cat, it has a high similarity with knowledge entries in the knowledge base such as "feline characteristics" and "common pet habits". Set the preset threshold to 0.75, and extract the knowledge entries with a similarity greater than this threshold as supplementary information.

[0025] Take the unified semantic feature as the query vector and the supplementary information as key-value pairs, and perform information fusion through the multi-head attention mechanism. Adopt an 8-head attention mechanism, with the dimension of each attention head being 64. Calculate 8 groups of different attention weights in parallel to capture semantic associations in different aspects. For example, for the unified semantic feature of a cat, different attention heads may respectively focus on different aspects of knowledge supplementation such as the appearance characteristics, behavior habits, and living environment of the cat. Concatenate the attention calculation results and pass them through a linear mapping layer to obtain a fused feature vector with a dimension of 512, which contains both the original text-image semantic information and the relevant prior knowledge in the knowledge base.

[0026] In this embodiment, by calculating the mutual information between the text feature vector and the image feature vector, the semantic correlation degree between different modality data can be measured, the internal correlation of cross-modal data can be captured from the perspective of information theory, the semantic information loss caused by simple feature splicing can be avoided, and a cross-modal semantic mapping matrix is constructed by introducing a temperature-parameter-regulated normalized exponential function, which can adaptively adjust the semantic correlation strength distribution between different modality features, match the unified semantic features with the prior knowledge in the knowledge base for similarity, and perform information fusion through a multi-head attention mechanism, realizing knowledge-enhanced feature representation.

[0027] In an alternative embodiment, Calculating the similarity between the unified semantic features and the prior knowledge in the preset knowledge base, and taking the prior knowledge with a similarity greater than the preset threshold as supplementary information includes: Decompose the unified semantic features into a core semantic variable set and a context semantic variable set, construct a conditional dependence graph based on the core semantic variable set and the context semantic variable set, calculate the initial probability distribution parameters by using maximum likelihood estimation according to the structure of the conditional dependence graph, and iteratively optimize the initial probability distribution parameters by combining the expectation maximization algorithm to obtain optimized probability distribution parameters; Calculate the initial similarity value between each piece of prior knowledge in the preset knowledge base and the unified semantic features according to the optimized probability distribution parameters, and approximately calculate the final similarity value based on the Monte Carlo sampling method; Calculate the similarity difference between different prior knowledge to obtain a similarity difference value, construct an adaptive similarity threshold according to the similarity difference value, dynamically adjust the adaptive similarity threshold according to the feedback score of the prior knowledge to obtain an adjusted similarity threshold, and determine the prior knowledge with the final similarity value greater than the adjusted similarity threshold as candidate supplementary information; For each candidate supplementary information, calculate a joint similarity value based on the final similarity value and the similarity difference value, and select the candidate supplementary information with the largest joint similarity value as the supplementary information for output.

[0028] Perform decomposition processing on the unified semantic features, and divide them into a core semantic variable set and a context semantic variable set. The core semantic variable set contains keywords or phrases expressing the main semantic content. For example, for the query "the service life of a mobile phone", the core semantic variables may include "mobile phone", "service life", etc. The context semantic variable set contains peripheral information that helps to understand the core semantics, such as related concepts like "electronic device", "durability", etc.

[0029] Based on the decomposed set of semantic variables, construct a conditional dependency graph to clearly represent the logical associations between variables. Calculate the conditional mutual information value for each pair of variables. When the mutual information value exceeds 0.6, establish a connection in the graph. For example, the mutual information value between "service life" and "durability" is 0.85, so a connection is established between these two nodes. After completing the graph construction, use the maximum likelihood estimation method to calculate the initial probability distribution parameters. By statistically analyzing the occurrence frequencies and co-occurrence situations of each variable in the training data, obtain the probability values of each node and edge.

[0030] To further optimize the probability distribution parameters, apply the expectation-maximization algorithm for iterative optimization. Set the initial iteration value, using the parameters obtained from the maximum likelihood estimation as the starting point; execute the expectation step to calculate the expected value of the latent variable; execute the maximization step to update the model parameters; check the convergence condition. If the parameter change is less than the preset value of 0.001 or the maximum number of iterations of 50 times is reached, stop the iteration.

[0031] After obtaining the optimized probability distribution parameters, calculate the initial similarity value between each piece of prior knowledge in the preset knowledge base and the unified semantic features. During the calculation process, represent the prior knowledge as a corresponding probability distribution and calculate the KL divergence or JS divergence between the two distributions. For example, for the prior knowledge "The average service life of electronic devices is 3 - 5 years", the system calculates its initial similarity value with the query "The service life of mobile phones" as 0.78.

[0032] Since the exact calculation of similarity has a high computational complexity in the high-dimensional feature space, the system uses the Monte Carlo sampling method for approximate calculation. Randomly sample 1000 points from each of the two distributions and calculate the average distance between the sampling points as an estimate of the final similarity value.

[0033] To determine the appropriate similarity threshold, calculate the similarity differences between different pieces of prior knowledge. The system calculates the average similarity difference between the top 20% of the prior knowledge in terms of similarity ranking and the bottom 20% of the prior knowledge in the knowledge base to obtain the similarity difference value. Based on this difference value, construct an adaptive similarity threshold, with the initial threshold set to 0.65.

[0034] This adaptive similarity threshold will be dynamically adjusted according to the feedback scores of the prior knowledge. Record the feedback scores (range 1 - 5 points) of users for each piece of supplementary information. When the average score is higher than 4 points, reduce the threshold by 0.05; when the average score is lower than 3 points, increase the threshold by 0.05. For example, if the current threshold is 0.65 and the average score of the recent 10 pieces of supplementary information is 4.2, the adjusted threshold becomes 0.60.

[0035] The prior knowledge with the final similarity value greater than the adjusted threshold is determined as the candidate supplementary information. For each candidate supplementary information, the system further calculates the joint similarity value, and the formula is: Joint similarity = 0.7 × Final similarity value + 0.3 × (1 - Relative similarity difference value). Among them, the relative similarity difference value represents the average similarity between this knowledge and other candidate knowledge, and is used to measure the uniqueness of the information.

[0036] For example, for candidate knowledge A, the final similarity value is 0.82, and the average similarity with other candidate knowledge is 0.45, then the joint similarity value is 0.7 × 0.82 + 0.3 × (1 - 0.45) = 0.739. For candidate knowledge B, the final similarity value is 0.78, and the average similarity with other candidate knowledge is 0.35, then the joint similarity value is 0.7 × 0.78 + 0.3 × (1 - 0.35) = 0.741. Although the final similarity of B is slightly lower, due to its higher uniqueness, the final joint similarity value is higher.

[0037] Select the candidate supplementary information with the maximum joint similarity value as the final output. In the above example, the system will select candidate knowledge B as the supplementary information to provide to the user.

[0038] In this embodiment, the unified semantic features are decomposed into core semantic variables and context semantic variables, the structural relationship between variables is explicitly modeled by constructing a conditional dependency graph, the probability distribution parameters are optimized by combining the maximum likelihood estimation and the expectation maximization algorithm, an adaptive threshold is constructed by analyzing the similarity difference between different prior knowledge, and is dynamically adjusted according to the actual feedback of the knowledge. By jointly considering the final similarity value and the similarity difference value to select the optimal supplementary information, it not only ensures the relevance between the selected knowledge and the query features, but also considers the discrimination between knowledge; In the prior art, the similarity between features and prior knowledge is usually calculated by using fixed measurement methods such as simple cosine similarity or Euclidean distance, ignoring the semantic structure and conditional dependency relationship inside the features, and using a fixed threshold for knowledge screening, which is difficult to adapt to the differences in the degree of knowledge association in different scenarios; The method based on the probabilistic graph model in this embodiment can more accurately describe the internal semantic structure of the features, improve the accuracy of similarity calculation, uses the Monte Carlo sampling method to approximately calculate the initial similarity, effectively reduces the computational complexity, and at the same time maintains the stability of similarity calculation. The multi-dimensional knowledge selection strategy significantly improves the accuracy and diversity of knowledge supplementation.

[0039] Figure 2It is the conditional dependency graph dynamic construction engineering simulation diagram of the embodiment of the present invention, which presents how the system constructs a dependency network containing core semantic variables and context semantic variables based on unified semantic feature decomposition. The central node "query topic" is closely connected to 6 core semantic variables through conditional dependencies, and the conditional mutual information values are 0.85, 0.89, 0.79, 0.82, 0.81, and 0.91 respectively. These high mutual information values indicate a strong correlation between the core variables and the query topic. At the same time, the core semantic variables are connected to 5 context semantic variables through weaker dependencies (mutual information values between 0.65 and 0.83).

[0040] Figure 2 It reflects how the system identifies key concepts and establishes a dependency network between them when processing a specific query, laying a foundation for subsequent similarity calculation. The numerical value marked on each connection line in the figure represents the conditional mutual information intensity, which is a key indicator for the system to quantify semantic associations. The prediction ability of the entire dependency graph reaches 93.5%, and the error propagation rate is only 7.2%, fully demonstrating the superior performance of the present technical solution in capturing complex semantic relationships.

[0041] Compared with the traditional Bayesian network, the dependency graph constructed by this solution can more accurately distinguish core semantics and context semantics, and accurately quantify the dependency strength between them through conditional mutual information values. This structured semantic representation can significantly improve the accuracy of subsequent knowledge matching, especially when dealing with complex queries.

[0042] In an alternative embodiment Constructing a user intention feature vector according to the fusion feature, and performing similarity matching between the user intention feature vector and a preset intention category library, and selecting the category with the highest similarity as the user intention category includes:[[]] Obtaining the fusion feature and mapping the fusion feature to the semantic space through a non-linear transformation to obtain an initial intention feature, and performing dimensionality reduction processing on the initial intention feature to obtain the user intention feature vector; Calculating the similarity between the user intention feature vector and each intention category in the preset intention category library to obtain a category similarity value, normalizing the category similarity value to obtain a probability distribution, and selecting the intention category with the maximum probability as the user intention category for output.

[0043] Processing the text "I want to query tomorrow's weather forecast" input by the user through a pre-trained BERT model, extracting a text embedding representation with a dimension of 768, and extracting fundamental frequency, energy, and MFCC features from the user's speech through an acoustic feature extraction module to form a 128-dimensional acoustic feature vector. Using a feature fusion network to perform a connection operation on the text feature and the acoustic feature to obtain an 896-dimensional fusion feature.

[0044] After obtaining the fused features, the fused features are mapped to the semantic space through a non-linear transformation to obtain the initial intent features. A multi-layer perceptron network is used to implement this non-linear mapping, including three fully-connected layers. Each layer uses the ReLU activation function, and the structure is 896-512-256-128. The 896-dimensional fused features are mapped to 512 dimensions through the first fully-connected layer, then mapped to 256 dimensions through the second layer, and mapped to 128-dimensional initial intent features through the third layer. For the above input example, the obtained initial intent features are a 128-dimensional vector, represented as [0.32, 0.15, 0.78, ..., 0.41].

[0045] To improve the efficiency and accuracy of intent recognition, dimensionality reduction processing is performed on the initial intent features to obtain the user intent feature vector. The principal component analysis method is used to reduce the 128-dimensional initial intent features to 64 dimensions, retaining the most representative information in the original feature space. The covariance matrix of the initial intent features is calculated, the eigenvalues and eigenvectors of this matrix are solved, and the eigenvectors corresponding to the largest 64 eigenvalues are selected to construct the projection matrix. By multiplying the initial intent features with this projection matrix, a 64-dimensional user intent feature vector [0.25, 0.62, 0.18, ..., 0.37] is obtained.

[0046] In the intent category matching stage, the user intent feature vector is matched with the preset intent category library for similarity. The preset intent category library is a database containing multiple common intent categories, and each category has a corresponding feature vector representation. In this embodiment, the preset intent category library includes 20 common intent categories such as "weather query", "navigation request", "music playback", "alarm setting", and "message sending", and each category has a 64-dimensional feature vector representation.

[0047] The similarity between the user intent feature vector and each intent category in the preset intent category library is calculated to obtain the category similarity value. When calculating the similarity, the cosine similarity measurement method is used, that is, the cosine value of the angle between two vectors is calculated. For the user intent feature vector V and the intent category feature vector C, the dot product V·C of the two vectors is calculated, and then divided by the product of the norms of the two vectors |V|×|C| to obtain the cosine similarity value. For example, the system calculates that the similarity between the user intent feature vector and the "weather query" category is 0.85, the similarity with the "navigation request" category is 0.42, the similarity with the "music playback" category is 0.38, the similarity with the "alarm setting" category is 0.25, and the similarity with the "message sending" category is 0.20.

[0048] Normalize the category similarity values to obtain a probability distribution. Using the Softmax normalization method, convert each similarity value into a probability value to ensure that the sum of all probability values is 1. For the category similarity values [0.85, 0.42, 0.38, 0.25, 0.20,...], the corresponding probability distribution [0.65, 0.15, 0.10, 0.05, 0.03,...] is calculated.

[0049] Select the intent category with the highest probability as the output of the user intent category. In this example, the probability of the "weather query" category is 0.65, which is significantly higher than the probabilities of other categories. Therefore, the system outputs "weather query" as the user's intent category. For the user input "I want to query the weather forecast for tomorrow", the system correctly identifies that the user's intent is to query weather information.

[0050] To improve the accuracy of intent recognition, a minimum confidence threshold of 0.5 is also set. Only when the highest probability value exceeds this threshold will the recognition result be confirmed; otherwise, the system will return a result of "ambiguous intent" and request the user to provide more information. In this example, the probability of 0.65 for the "weather query" category exceeds the threshold of 0.5, so the system confirms that the user intent is "weather query".

[0051] In this embodiment, by non-linearly transforming the fused features into the semantic space, it is possible to better capture the complex semantic associations between features, effectively extract deep semantic information, and adopt an intent recognition method based on similarity calculation and probability normalization, which transforms the intent recognition problem into a probability distribution prediction task. It not only outputs the final intent category but also provides the confidence of each intent category, providing a richer basis for subsequent intent understanding and dialogue decision-making.

[0052] In an alternative embodiment, Calculate the importance distribution of key entities in the historical dialogue based on word frequency and information gain, use a sliding window to count the co-occurrence frequencies of entities and events and construct an association network, perform pruning based on the association strength to obtain the core semantic skeleton, extract the key path and calculate the semantic coherence, and perform importance ranking on the historical dialogue information by combining the path weight and coherence. The key information determined according to the ranking results includes: Obtain the historical dialogue content, extract the entities in the historical dialogue content as candidate entities, calculate the word frequency score and information gain score of the candidate entities, and obtain the entity importance distribution by taking the weighted sum of the word frequency score and the information gain score, and determine the key entities according to the entity importance distribution; Determine the sliding window size according to the length of the historical dialogue content. Count the co-occurrence frequencies of key entities and events within the sliding window, calculate the association strength based on the co-occurrence frequencies and construct an association network, set the weights of the association network based on the association strength and key entities, determine the pruning threshold according to the mean and standard deviation of the weights, and retain the associations in the association network with weights greater than the pruning threshold to obtain the core semantic skeleton; Determine the path length according to the number of shortest paths between nodes in the core semantic skeleton, set the attenuation factor based on the path length, calculate the node centrality score using the path length and attenuation factor to extract the key path, and calculate the cosine similarity of the semantic vectors of adjacent nodes in the key path to obtain the semantic coherence; Obtain the historical dialogue information based on the key path, calculate the centrality score of the key path in the historical dialogue information to get the path weight, combine the weighted sum of the path weight and semantic coherence to obtain the importance ranking value, and sort the historical dialogue information according to the importance ranking value and determine the key information.

[0053] The entity refers to a specific object in the historical dialogue content, including identifiable independent individuals such as people, places, organizations, product names, time, etc. For example, in the customer service dialogue scenario, the entity can be a specific product model, user account, order number, delivery address, etc. The event refers to a specific behavior, state change or scenario description that occurs in the dialogue, including the action subject, action and related elements. For example, order placement operations, refund applications, logistics distribution, product usage, etc. in customer service dialogues all belong to events. Events are often composed of the interaction relationships between multiple entities, reflecting the dynamic information in the dialogue content.

[0054] Extract candidate entities from the historical dialogue content, including named entities such as person names, place names, organization names, time, quantity, etc. For each candidate entity, calculate its word frequency score, that is, the number of times the entity appears in the historical dialogue divided by the total number of times all entities appear in the historical dialogue. For example, in a historical dialogue about "product sales", "mobile phone" appears 15 times, "computer" appears 8 times, "tablet" appears 5 times, and all entities appear 50 times in total, then the word frequency score of "mobile phone" is 0.3, the word frequency score of "computer" is 0.16, and the word frequency score of "tablet" is 0.1.

[0055] Calculate the information gain score for each candidate entity, which is used to measure the entity's ability to distinguish the conversation topic. The calculation method is to divide the historical conversation content into a subset containing the entity and a subset not containing the entity, calculate the information entropy of the two subsets respectively, and subtract the weighted average information entropy of the two subsets after division from the original information entropy. For example, divide the above conversation into two subsets according to whether it contains "mobile phone". The original information entropy is 1.5. The information entropy of the subset containing "mobile phone" after division is 0.8, the weight is 0.4, the information entropy of the subset not containing "mobile phone" is 0.6, and the weight is 0.6. Then the information gain score of "mobile phone" is 1.5 - (0.8×0.4 + 0.6×0.6) = 0.82.

[0056] Fuse the word frequency score and the information gain score with weights to obtain the entity importance distribution. The weighting coefficients can be adjusted according to the specific application scenario. For example, if the word frequency weight is 0.4 and the information gain weight is 0.6, then the entity importance of "mobile phone" is 0.4×0.3 + 0.6×0.82 = 0.612. According to the entity importance distribution, select the top N entities with the highest scores as the key entities. In this embodiment, if N = 2, then "mobile phone" and "computer" are selected as the key entities.

[0057] Determine the sliding window size according to the length of the historical conversation content, generally set to 1 / 5 to 1 / 3 of the number of conversation turns. For example, for a historical content containing 30 conversation turns, set the sliding window size to 10 turns. Count the co-occurrence frequencies of the key entities and events within the sliding window. An event consists of a verb and its related components. For example, within a certain sliding window, "mobile phone" and "sales" co-occur 8 times, and "computer" and "sales" co-occur 5 times.

[0058] Calculate the association strength based on the co-occurrence frequency. The association strength is equal to the co-occurrence frequency divided by the total number of words in the window and then multiplied by a certain adjustment coefficient. For example, if the total number of words in the window is 200 and the adjustment coefficient is 100, then the association strength of "mobile phone - sales" is 8÷200×100 = 4, and the association strength of "computer - sales" is 5÷200×100 = 2.5. When constructing the association network, the key entities and events are used as nodes, and the association strength is used as the weight of the edge. Determine the pruning threshold according to the mean and standard deviation of the weights of all edges in the association network, usually set to the mean minus 0.5 times the standard deviation. For example, if the mean of the weights of the edges in the association network is 3 and the standard deviation is 1, then the pruning threshold is 3 - 0.5×1 = 2.5. Retain the associations in the association network with weights greater than the pruning threshold to form the core semantic skeleton. In this embodiment, the "mobile phone - sales" association is retained, while the "computer - sales" association is at the threshold edge.

[0059] Determine the path length based on the number of shortest paths between nodes in the core semantic skeleton. Exemplarily, a path of length 3 is formed from "user" to "mobile phone" to "sales" to "delivery". Set the attenuation factor based on the path length. The attenuation factor is usually a value between 0.8 and 0.9, which is used to reduce the importance of longer paths. Calculate the node centrality score using the path length and the attenuation factor. The centrality score is equal to the sum of the number of times the node appears in all paths multiplied by the attenuation factor of the corresponding path. For example, if the "mobile phone" node appears in 3 paths of different lengths, and the attenuation factors of the paths are 0.9, 0.81, 0.729 (assuming the attenuation factor is 0.9), then the node centrality score of the "mobile phone" is 0.9 + 0.81 + 0.729 = 2.439.

[0060] When extracting the key path, select the node with the highest node centrality score as the starting point, and select the connected nodes in descending order of association strength to construct the key path. Calculate the cosine similarity of the semantic vectors of adjacent nodes in the key path to obtain the semantic coherence. For example, in the "user - mobile phone - sales - delivery" path, the cosine similarities of adjacent node pairs are 0.75, 0.82, 0.68 respectively, and the semantic coherence of this path is the average value of these values, which is 0.75.

[0061] Obtain the historical conversation information based on the key path, calculate the centrality score of the key path in the historical conversation information to get the path weight. Combine the weighted sum of the path weight and the semantic coherence to obtain the importance ranking value. For example, assume the path weight is 0.7 and the semantic coherence weight is 0.3. Then for a path with a centrality score of 2.439 and a semantic coherence of 0.75, the importance ranking value is 0.7×2.439 + 0.3×0.75 = 1.9323. Sort the historical conversation information according to the importance ranking value, and select the top M pieces of information as the key information. Exemplarily, if M = 5, then select the top 5 pieces of conversation information with the importance ranking value as the key information. The key information includes: "The user asks about the mobile phone model", "The salesperson introduces the functions of the mobile phone", "The user decides to buy the mobile phone", "The salesperson explains the delivery method", "The user confirms the delivery time", etc.

[0062] In this embodiment, through the dual evaluation mechanism of word frequency score and information gain score, the importance of entities in the conversation content can be comprehensively measured. It takes into account both the frequency of entity appearance and the contribution of entities to information differentiation, making the extraction of key entities more accurate and reasonable. The co-occurrence statistical method based on a sliding window is adopted, and an association network is constructed by combining the association strength and entity importance. The core semantic skeleton is obtained through adaptive pruning, effectively retaining the core semantic structure in the conversation content and avoiding the interference of irrelevant information. The node centrality is calculated by introducing the path length and attenuation factor, and the quality of the key path is evaluated in combination with semantic coherence, making the extracted key path have both strong structural importance and semantic coherence, and improving the accuracy of key information extraction.

[0063] Figure 3 This is a comparison chart of the F1 score change trend under different conversation turns in the embodiment of the present invention. As the conversation turn increases, the F1 score of the technical solution (sliding window + association strength, circular mark) in this embodiment increases significantly, from 0.50 at the 10th conversation turn to 0.89 at the 60th conversation turn, showing a stable growth trend. In contrast, the PMI co-occurrence statistical method (square mark) only reaches an F1 score of 0.71 at the 60th conversation turn, and the WordNet semantic network method (triangle mark) is even lower, only 0.55. The data shows that when the conversation turn exceeds 20 turns, the advantages of the technical solution in this embodiment begin to expand significantly. At the 50th conversation turn, the F1 score (0.85) of the technical solution in this embodiment is 27% higher than that of the PMI co-occurrence statistical method (0.67) and 70% higher than that of the WordNet semantic network method (0.50).

[0064] The performance difference mainly stems from the fact that the sliding window technology adopted in the technical solution in this embodiment can better capture the dynamic changes in the conversation, while the association strength calculation method accurately evaluates the semantic association between entities and events. Although the traditional PMI co-occurrence statistical method can calculate word pairs In an optional implementation manner, Calculate the association strength based on the co-occurrence frequency and construct an association network, set the weight of the association network based on the association strength and key entities, determine the pruning threshold according to the weight mean and standard deviation, and retain the associations in the association network with weights greater than the pruning threshold to obtain the core semantic skeleton, including: Obtain the co-occurrence frequency of entities in the historical conversation content, divide the co-occurrence frequency by the square root of the product of the respective appearance times of the entities to obtain the association strength, and construct an association network based on the association strength; Calculate the importance score of the key entity, and set the product of the association strength and the average value of the key entity importance score as the weight of the association network; Use the community discovery algorithm to partition the association network into multiple communities, count the number of edge sets and node sets within each community, and divide the number of edge sets by half of the product of the number of node sets and the number of node sets minus one to obtain the internal connection density of the community; Calculate the mean and standard deviation of all weights in the association network, take the sum of the product of the mean and standard deviation of the weights and the first preset coefficient as the benchmark threshold, and obtain the adaptive pruning threshold corresponding to each community by adding the product of the benchmark threshold and the internal connection density of the community and the second preset coefficient; Compare the relationship between the adaptive pruning threshold of the community to which each edge in the association network belongs and the weight of the current edge, retain the edges with weights greater than the adaptive pruning threshold of the corresponding community, and remove the edges with weights less than the adaptive pruning threshold of the corresponding community to obtain the core semantic skeleton.

[0065] Obtain the entities in the historical conversation content, and identify entities such as people, places, organizations, and times from the historical conversation text through named entity recognition technology. For example, in a historical conversation about the discussion of technology products, entities such as "smartphone", "tablet computer", "laptop", and "smart home" can be identified.

[0066] After obtaining the entities, calculate the co-occurrence frequency between the entities. Two entities appear in the same dialogue window as one co-occurrence. The dialogue window can be a sentence, a turn, or a fixed character range. For example, if the dialogue window is set to a single turn, count that the number of times "smartphone" and "tablet computer" appear together in the same turn is 15 times, the total number of times "smartphone" appears alone is 30 times, and the total number of times "tablet computer" appears alone is 25 times.

[0067] When calculating the association strength, divide the co-occurrence frequency by the square root of the product of the respective occurrence times of the two entities. Exemplarily, the association strength between "smartphone" and "tablet computer" is 15 divided by the square root of (30 multiplied by 25), approximately equal to 0.55. Calculate the association strength of all entity pairs and construct an association network, where the nodes are entities, the edges are the associations between entities, and the initial weight of the edge is the association strength.

[0068] Calculate the importance score of the key entity. The degree centrality calculation method can be used, that is, calculate the number of connections of each entity with other entities. For example, "smartphone" is associated with 15 other entities, and its degree centrality is 15. Algorithms such as PageRank can also be used to calculate the entity importance. Taking "smartphone" as an example, assume that the calculation result of its importance score is 0.85.

[0069] Adjust the weight of the association network, multiply the association strength by the average of the importance scores of the two key entities connected, as the new weight of the association network edge. Take "smartphone" and "tablet" as an example, assuming that the importance score of "tablet" is 0.75, then the weight of the edge between these two entities is adjusted to 0.55 multiplied by (0.85 plus 0.75) divided by 2, that is, 0.55 multiplied by 0.8, which equals 0.44.

[0070] Use community discovery algorithms to divide the association network into communities, such as the Louvain algorithm or the label propagation algorithm. The division results divide the association network into multiple communities, and the nodes in each community are more closely connected. Assume that the algorithm divides the association network into three communities, which contain entities related to "communication equipment", "home appliances" and "office equipment" respectively.

[0071] For each community, count the number of internal edge sets and node sets. Take the "communication equipment" community as an example, which contains 10 nodes (entities) and 30 edges (associations). Calculate the internal connection density of the community, which is the number of edge sets divided by half of the product of the number of node sets and the number of node sets minus one. For the "communication equipment" community, the connection density is 30 divided by half of (10 times 9), that is, 30 divided by 45, which equals 0.67.

[0072] Calculate the mean and standard deviation of all weights in the association network. Assume that the mean weight of all edges in the network is 0.3 and the standard deviation is 0.15. Set the sum of the weight mean and standard deviation multiplied by the first preset coefficient as the baseline threshold. Assuming that the first preset coefficient is set to 1.0, the baseline threshold is 0.3 plus 0.15 multiplied by 1.0, which equals 0.45.

[0073] An adaptive trimming threshold is set for each community, and the sum of the base threshold and the product of the community's internal connection density and the second preset coefficient is used as the adaptive trimming threshold of the community. Assuming that the second preset coefficient is set to -0.2, the adaptive trimming threshold of the "communication equipment" community is 0.45 plus 0.67 multiplied by -0.2, which is approximately equal to 0.32. For communities with a connection density higher than 0.8, the adaptive trimming threshold will be reduced accordingly to retain more edges; while for communities with a connection density lower than 0.2, the adaptive trimming threshold will be increased to retain only edges with larger weights.

[0074] Compare the adaptive pruning threshold of the community to which each edge belongs in the association network with the weight of the current edge. If the weight of the edge is greater than the adaptive pruning threshold of the community to which it belongs, the edge is retained; otherwise, the edge is removed. For example, the edge weight between "smartphone" and "tablet" is 0.44, which is greater than the adaptive pruning threshold of the "communication equipment" community of 0.32, so the edge is retained.

[0075] In this embodiment, the co-occurrence frequency and normalization are combined to calculate the association strength, which can effectively eliminate the influence caused by the difference in the occurrence frequency of entities. By normalizing the co-occurrence frequency by dividing it by the square root of the product of the occurrence times of each entity, the calculation of the association strength becomes more objective, avoiding the dominant effect of high-frequency entities on the calculation of the association strength. Based on the differential processing method of the community structure, it is possible to remove the weak connections between communities in a targeted manner while maintaining a high-density connection within the community, making the network simplification more reasonable. The adaptive threshold setting method can dynamically adjust the pruning criteria according to the characteristics of different communities, avoiding the problems of over-pruning or retaining redundancy that may be caused by a single fixed threshold.

[0076] Figure 4 This is the core semantic skeleton construction flowchart of the intelligent interactive question-answering method based on the intention recognition of the multi-modal large model in the embodiment of the present invention.

[0077] In an alternative embodiment, Generating a response content based on the key information and returning it to the customer, and updating the association network based on the current conversation content includes: Generating a response content including the conversation intention, user demands, and solution based on the key information, and returning the response content to the customer; Extracting entity-attribute-relationship triples of entities, attributes, and relationships from the current conversation content to obtain new entities, calculating the co-occurrence frequency and semantic similarity between the new entities to obtain the weight between entities, and calculating the association strength between the new entities and the original entities to obtain the interaction weight; Adding the new entities as nodes and the weight between entities and the interaction weight as edges to the original association network, recalculating the adaptive pruning threshold for the extended association network based on the connection density within the community, and retaining the edges with weights greater than the adaptive pruning threshold of the corresponding community to obtain the updated association network.

[0078] Analyze the user's conversation intention based on the obtained key information, identify the specific demand points of the user, and generate a complete response content including intention understanding, demand analysis, and solution in combination with the solution experience in the historical conversation. The generated response content needs to have logical coherence and semantic integrity to ensure that it can be accurately conveyed to the customer and solve their problems.

[0079] From the newly added current conversation content, extract entity information and its corresponding attribute features through semantic analysis, and at the same time identify various relationships between entities, and organize this information into the form of entity-attribute-relationship triples. For the newly extracted entities, count the co-occurrence times within the conversation window, and calculate the semantic similarity degree between the entities. The weighted result of these two indicators is used as the association weight between entities. At the same time, calculate the association strength between the new entities and the existing entities in the association network as the interaction weight between entities.

[0080] Add the newly extracted entities as new nodes to the original association network, and at the same time add the calculated weights between entities and interaction weights as connecting edges to the network. Re - perform community partitioning on the extended association network and calculate the connection density within each community. Based on the updated community structure features and weight distribution features, recalculate the adaptive pruning threshold corresponding to each community. Finally, optimize the network according to the new pruning threshold, retain the important connections with larger weights, and obtain the updated association network reflecting the current dialogue state.

[0081] Exemplarily, in the intelligent customer service scenario, the user reflects: "There is a flashing problem with my mobile phone screen, which has lasted for a week. I want to apply for after - sales repair. This mobile phone is model A purchased offline last year and is still within the warranty period." Based on the analysis of key information, it is known that the user's intention is to apply for mobile phone repair, and the core demand is to solve the screen flashing problem. Accordingly, the response content is generated: "We understand that there is a screen flashing problem with your model A mobile phone. Since this model is still within the warranty period, it is recommended that you repair it through the official after - sales channel. The nearest official after - sales service point is located at..." The newly added entities extracted from the current dialogue include "model A", "screen", "flashing", "offline", "warranty period", etc., as well as entity attributes such as "purchase time: last year", "fault duration: one week", etc. Calculate the co - occurrence relationship and semantic similarity between the newly added entities to obtain the weights between entities. For example, the weights of "model A - screen" and "screen - flashing" are relatively high. At the same time, calculate the association strength between the newly added entities and the original entities (such as "after - sales repair", "official channel", etc.).

[0082] Add these newly added nodes and edges to the original network, re - partition the communities to obtain communities such as "product information community", "fault description community", "after - sales service community", etc. Calculate the internal connection density of each community, update the pruning threshold accordingly, and retain the important connections to obtain a new association network, which can better express the core semantic structure of the current dialogue.

[0083] In this embodiment, by parsing key information to generate a complete response content including dialogue intention, user demand, and solution, the multi - dimensional information organization method ensures the comprehensiveness of the response content. The structured response generation method can not only accurately understand and respond to user needs, but also provide targeted solutions, improving the service quality of the dialogue system. By considering both co - occurrence frequency and semantic similarity to calculate the weights between entities, it not only reflects the statistical relevance of entities in the dialogue, but also embodies the semantic association degree between entities, making the description of entity relationships more comprehensive and accurate. The dynamic knowledge integration method enables the association network to continuously accumulate and update dialogue information, improving the timeliness and integrity of knowledge representation.

[0084] In the second aspect of the embodiments of the present invention, there is provided an intelligent interactive question-answering system based on multi-modal large model intention recognition, including: A first unit, configured to obtain multi-modal interaction information input by a user and perform feature extraction to obtain a multi-modal feature vector; A second unit, configured to calculate the mutual information amount between different modal feature vectors, construct a cross-modal semantic mapping matrix based on the mutual information amount, map different modal feature vectors to a unified semantic space by using the cross-modal semantic mapping matrix to obtain unified semantic features, calculate the semantic similarity with prior knowledge in a preset knowledge base, determine supplementary information, and fuse the supplementary information with the unified semantic features to obtain fused features; A third unit, configured to construct a user intention feature vector according to the fused features, perform similarity matching between the user intention feature vector and a preset intention category library, and select the category with the highest similarity as the user intention category; A fourth unit, configured to calculate the importance distribution of key entities in a historical conversation based on word frequency and information gain, use a sliding window to count the co-occurrence frequencies of entities and events and construct an association network, perform pruning based on the association strength to obtain a core semantic skeleton, extract a key path and calculate semantic coherence, sort the importance of historical conversation information by combining path weights and coherence, and determine key information according to the sorting result; A fifth unit, configured to generate a response content based on the key information and return it to the customer, and update the association network based on the current conversation content.

[0085] In the third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor and a memory for storing processor-executable instructions, wherein the processor is configured to call the instructions stored in the memory to execute the method described above.

[0086] In the fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0087] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are loaded.

[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent interactive question-answering method based on the intention recognition of a multimodal large model, characterized in that, Including: Obtain the multi-modal interaction information input by the user and perform feature extraction to obtain a multi-modal feature vector; Calculate the mutual information between different modal feature vectors, construct a cross-modal semantic mapping matrix based on the mutual information, use the cross-modal semantic mapping matrix to map different modal feature vectors to a unified semantic space, obtain a unified semantic feature, calculate the semantic similarity with the prior knowledge in the preset knowledge base, determine supplementary information, and fuse the supplementary information with the unified semantic feature to obtain a fused feature; Construct a user intention feature vector according to the fused feature, perform similarity matching between the user intention feature vector and a preset intention category library, and select the category with the highest similarity as the user intention category; Calculate the importance distribution of key entities in the historical dialogue based on word frequency and information gain, use a sliding window to count the co-occurrence frequencies of entities and events and construct an association network, perform pruning based on the association strength to obtain a core semantic skeleton, extract the key path and calculate the semantic coherence, combine the path weight and coherence to rank the importance of the historical dialogue information, and determine the key information according to the ranking result; Generate a response content based on the key information and return it to the customer, and update the association network based on the current dialogue content.

2. The method according to claim 1, wherein Calculating the mutual information between different modal feature vectors, constructing a cross-modal semantic mapping matrix based on the mutual information, using the cross-modal semantic mapping matrix to map different modal feature vectors to a unified semantic space, obtaining a unified semantic feature, calculating the semantic similarity with the prior knowledge in the preset knowledge base, determining supplementary information, and fusing the supplementary information with the unified semantic feature to obtain a fused feature includes: Obtain the multi-modal feature vector to be processed, where the multi-modal feature vector includes a text feature vector and an image feature vector; Calculate the joint probability distribution and marginal probability distribution of the text feature vector and the image feature vector, and calculate the mutual information between the text feature vector and the image feature vector based on the joint probability distribution and the marginal probability distribution; Construct a cross-modal semantic mapping matrix according to the mutual information, where the value of each element in the cross-modal semantic mapping matrix is calculated by a normalized exponential function adjusted by a temperature parameter, and the temperature parameter is used to control the distribution of the semantic association strength in the cross-modal semantic mapping matrix, and map the text feature vector and the image feature vector to a unified semantic space based on the cross-modal semantic mapping matrix to obtain a unified semantic feature; Calculate the similarity between the unified semantic feature and the prior knowledge in the preset knowledge base, use the prior knowledge with a similarity greater than a preset threshold as supplementary information, use the unified semantic feature as a query vector, use the supplementary information as key-value pairs, and perform information fusion through a multi-head attention mechanism to obtain a fused feature.

3. The method according to claim 2, wherein Calculating the similarity between the unified semantic feature and the prior knowledge in the preset knowledge base, and using the prior knowledge with a similarity greater than a preset threshold as supplementary information includes: Decompose the unified semantic feature into a set of core semantic variables and a set of context semantic variables, construct a conditional dependence graph based on the set of core semantic variables and the set of context semantic variables, calculate the initial probability distribution parameters by using maximum likelihood estimation according to the structure of the conditional dependence graph, and iteratively optimize the initial probability distribution parameters by combining the expectation maximization algorithm to obtain the optimized probability distribution parameters; Calculate the initial similarity value between each piece of prior knowledge in the preset knowledge base and the unified semantic feature according to the optimized probability distribution parameters, and approximately calculate the initial similarity value based on the Monte Carlo sampling method to obtain the final similarity value; Calculate the similarity difference between different pieces of prior knowledge to obtain a similarity difference value, construct an adaptive similarity threshold according to the similarity difference value, and the adaptive similarity threshold is dynamically adjusted according to the feedback score of the prior knowledge to obtain an adjusted similarity threshold, and determine the prior knowledge with the final similarity value greater than the adjusted similarity threshold as candidate supplementary information; For each candidate supplementary information, calculate a joint similarity value based on the final similarity value and the similarity difference value, and select the candidate supplementary information with the largest joint similarity value as the supplementary information output.

4. The method according to claim 1, wherein Construct a user intention feature vector according to the fusion feature, perform similarity matching between the user intention feature vector and the preset intention category library, and select the category with the highest similarity as the user intention category, including: Obtain the fusion feature and map the fusion feature to the semantic space through a non-linear transformation to obtain an initial intention feature, and perform dimensionality reduction processing on the initial intention feature to obtain the user intention feature vector; Calculate the similarity between the user intention feature vector and each intention category in the preset intention category library to obtain a category similarity value, perform normalization processing on the category similarity value to obtain a probability distribution, and select the intention category with the largest probability as the user intention category output.

5. The method according to claim 1, wherein Calculate the importance distribution of key entities in the historical dialogue based on word frequency and information gain, use a sliding window to count the co-occurrence frequencies of entities and events and construct an association network, perform pruning based on the association strength to obtain a core semantic skeleton, extract the key path and calculate the semantic coherence, combine the path weight and coherence to rank the importance of the historical dialogue information, and determine the key information according to the ranking result, including: Obtain the historical dialogue content, extract the entities in the historical dialogue content as candidate entities, calculate the word frequency score and information gain score of the candidate entities, and obtain the entity importance distribution by the weighted sum of the word frequency score and the information gain score, and determine the key entities according to the entity importance distribution; Determine the sliding window size according to the length of the historical dialogue content, count the co-occurrence frequencies of key entities and events within the sliding window, calculate the association strength based on the co-occurrence frequencies and construct an association network, set the weight of the association network based on the association strength and the key entities, determine the pruning threshold according to the weight mean and standard deviation, and retain the associations in the association network with weights greater than the pruning threshold to obtain a core semantic skeleton; Determine the path length according to the number of shortest paths between nodes in the core semantic skeleton, set the attenuation factor based on the path length, calculate the node centrality score using the path length and the attenuation factor to extract the key path, and calculate the cosine similarity of the semantic vectors of adjacent nodes in the key path to obtain the semantic coherence; Obtain the historical dialogue information based on the key path, calculate the centrality score of the key path in the historical dialogue information to obtain the path weight, combine the weighted sum of the path weight and the semantic coherence to obtain the importance ranking value, and sort the historical dialogue information according to the importance ranking value and determine the key information.

6. The method according to claim 5, wherein Calculate the association strength based on the co-occurrence frequency and construct an association network. Set the weight of the association network based on the association strength and the key entity. Determine the pruning threshold according to the mean and standard deviation of the weights, and retain the associations in the association network with weights greater than the pruning threshold to obtain the core semantic skeleton, including: Obtain the co-occurrence frequency of entities in the historical dialogue content, divide the co-occurrence frequency by the square root of the product of the respective occurrence times of the entities to obtain the association strength, and construct an association network based on the association strength; Calculate the importance score of the key entity, and set the product of the association strength and the average value of the importance scores of the key entities as the weight of the association network; Use the community discovery algorithm to partition the association network into multiple communities, count the number of edge sets and node sets within each community, and divide the number of edge sets by half of the product of the number of node sets and the number of node sets minus one to obtain the internal connection density of the community; Calculate the mean and standard deviation of all weights in the association network, use the sum of the product of the mean and standard deviation of the weights and the first preset coefficient as the baseline threshold, and use the sum of the baseline threshold and the product of the internal connection density of the community and the second preset coefficient to obtain the adaptive pruning threshold corresponding to each community; Compare the relationship between the adaptive pruning threshold of the community to which each edge in the association network belongs and the weight of the current edge, retain the edges with weights greater than the adaptive pruning threshold of the corresponding community, and remove the edges with weights less than the adaptive pruning threshold of the corresponding community to obtain the core semantic skeleton.

7. The method according to claim 1, characterized in that, Generate a response content based on the key information and return it to the customer. Update the association network based on the current dialogue content, including: Generate a response content including the dialogue intention, user requirements, and solution based on the key information, and return the response content to the customer; Extract entity, attribute, and relationship triples from the current dialogue content to obtain new entities, calculate the co-occurrence frequency and semantic similarity between the new entities to obtain the entity-inter entity weight, and calculate the association strength between the new entities and the original entities to obtain the interaction weight; Add the new entities as nodes, and the entity-inter entity weight and the interaction weight as edges to the original association network. Recalculate the adaptive pruning threshold for the extended association network based on the internal connection density of the community, and retain the edges with weights greater than the adaptive pruning threshold of the corresponding community to obtain the updated association network.

8. An intelligent interactive question-answering system based on the intention recognition of a multimodal large model, which is used to implement the method described in any one of the foregoing claims 1-7, characterized in that, Including: The first unit is used to obtain the multimodal interaction information input by the user and perform feature extraction to obtain a multimodal feature vector; A second unit, configured to calculate the mutual information amount between different modality feature vectors, construct a cross-modal semantic mapping matrix based on the mutual information amount, map the different modality feature vectors to a unified semantic space by using the cross-modal semantic mapping matrix to obtain unified semantic features, calculate the semantic similarity with the prior knowledge in a preset knowledge base, determine supplementary information, and fuse the supplementary information with the unified semantic features to obtain fused features; A third unit, configured to construct a user intention feature vector according to the fused features, perform similarity matching between the user intention feature vector and a preset intention category library, and select the category with the highest similarity as the user intention category; A fourth unit, configured to calculate the importance distribution of key entities in a historical dialogue based on word frequency and information gain, use a sliding window to count the co-occurrence frequencies of entities and events and construct an association network, perform pruning based on the association strength to obtain a core semantic skeleton, extract a key path and calculate semantic coherence, perform importance ranking on the historical dialogue information by combining the path weight and coherence, and determine key information according to the ranking result; A fifth unit, configured to generate a response content based on the key information and return it to a customer, and update the association network based on the current dialogue content.

9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Construction method of multi-modal user mental perception question and answer model

    CN117033602A

  • Question and answer data processing method and system based on multi-modal large model

    CN119312284A

  • Knotarization intelligent question and answer customer service method and system based on knowledge graph

    CN119938816A

  • System and Method for Temporal Attention Behavioral Analysis of Multi-Modal Conversations in a Question and Answer System

    US20220164548A1

Cited By

  • Supplier relationship management method and system combined with big data analysis

    CN120707153A

  • Supplier relationship management method and system combined with big data analysis

    CN120707153B

  • Diagnosis and treatment interaction system and method based on multi-modal data

    CN121030675A

  • Multi-user mixed interactive data processing method and system based on large model

    CN121118904A

  • Children brain health intelligent question and answer method and system fused with multi-modal data

    CN121525814A