Dialogue scene data visualization analysis method and system based on knowledge graph
By using a knowledge graph-based method for visualizing and analyzing dialogue scenario data, a knowledge subgraph of the dialogue scenario is constructed, multi-dimensional analysis is performed, and visualization results are generated. This solves the problems of incomplete semantic understanding and insufficient dynamic tracking in existing technologies for dialogue scenario data analysis, and achieves efficient visualization and real-time optimization of dialogue scenarios.
Patent Information
- Application Number
- CN202511436769.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-10-09
AI Technical Summary
Existing dialogue data analysis technologies do not adequately handle the relationships between user intent, topic nodes, and context, resulting in incomplete semantic understanding. The analysis results lack semantic relationship visualization based on knowledge graphs, cannot intuitively display the dialogue flow path and hotspot distribution, and cannot achieve dynamic tracking and real-time feedback.
A visualization analysis method for dialogue scenario data based on knowledge graphs is used. This method acquires dialogue scenario data, performs word segmentation and semantic parsing, extracts keywords and contextual tags, constructs knowledge subgraphs, conducts multi-dimensional analysis, generates semantic relationship graphs, process path graphs and hotspot distribution graphs, and uses incremental updates to achieve dynamic adjustments.
It enables a visualized panoramic presentation of dialogue scenarios, enhances the ability to understand dialogue scenarios, helps discover core intentions, abnormal paths and user hotspots, and improves the efficiency and accuracy of dialogue optimization and decision support.
Smart Images

Figure CN121093979A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data visualization analysis, in particular to a dialogue scene data visualization analysis method and system based on a knowledge graph. BACKGROUND
[0002] In intelligent customer service, online education, financial consulting and other business scenarios, a large amount of dialogue data between users and systems is continuously generated and accumulated. How to efficiently analyze these dialogue scene data has become an important way to improve service quality, optimize interaction experience and assist business decision-making. Knowledge graph is widely used in dialogue systems due to its advantages in semantic modeling and association relationship expression, to support intent recognition, context understanding and multi-round interaction. However, with the increase of dialogue data scale and complexity, it is difficult to meet the deep understanding and multi-dimensional analysis needs of dialogue data by relying only on traditional text retrieval and statistical analysis methods.
[0003] The existing dialogue data analysis technology still has deficiencies in application. On the one hand, the association relationship between user intent, theme node and context in the dialogue scene is often not fully handled, resulting in incomplete semantic understanding, thereby affecting the analysis accuracy of dialogue data. On the other hand, the analysis results of the existing system are mostly presented in the form of static charts or reports, lacking semantic relationship visualization based on knowledge graph, and unable to intuitively display dialogue flow paths, semantic associations and hot spot distribution. At the same time, for the increasing new dialogue data, the existing method lags in updating and displaying analysis results, and cannot realize dynamic tracking and real-time feedback. Therefore, it is necessary to provide a dialogue scene data visualization analysis method and system based on a knowledge graph to solve the above problems. SUMMARY
[0004] To solve the above technical problems, a dialogue scene data visualization analysis method and system based on a knowledge graph are provided, which solves the deficiencies of the existing dialogue data analysis technology in application. On the one hand, the association relationship between user intent, theme node and context in the dialogue scene is often not fully handled, resulting in incomplete semantic understanding, thereby affecting the analysis accuracy of dialogue data. On the other hand, the analysis results of the existing system are mostly presented in the form of static charts or reports, lacking semantic relationship visualization based on knowledge graph, and unable to intuitively display dialogue flow paths, semantic associations and hot spot distribution. At the same time, for the increasing new dialogue data, the existing method lags in updating and displaying analysis results, and cannot realize dynamic tracking and real-time feedback.
[0005] To achieve the above purposes, the technical scheme adopted by the present application is as follows: A dialogue scene data visualization analysis method based on a knowledge graph, comprising: obtaining dialogue scene data, the dialogue scene data comprising user input information, system reply information and context attributes; performing word segmentation and semantic analysis on the dialogue scene data, extracting keywords, topic words and context labels, and extracting corresponding entity nodes and relationship edges from a pre-constructed knowledge graph to construct a knowledge subgraph corresponding to the dialogue scene; fusing the dialogue scene data and the knowledge subgraph to obtain dialogue semantic structure data, the dialogue semantic structure data comprising instantiated user intent nodes, instantiated topic nodes and context relationships thereof; performing multi-dimensional analysis based on the dialogue semantic structure data to obtain an analysis result set, the multi-dimensional analysis comprising node importance analysis, dialogue path analysis and intent clustering analysis; extracting node importance analysis results, dialogue path analysis results and intent clustering analysis results from the analysis result set, respectively constructing a dialogue semantic relationship graph, a dialogue flow path graph and a dialogue hotspot distribution graph, and performing unified coordinate mapping and weighted fusion on the dialogue semantic relationship graph, the dialogue flow path graph and the dialogue hotspot distribution graph to obtain visualized analysis results; when new dialogue scene data is input, incrementally updating the knowledge graph to extract new associated nodes and relationship edges, updating and recalculating local semantic structures, and dynamically adjusting the visualized analysis results.
[0006] In an optional embodiment, the word segmentation and semantic analysis on the dialogue scene data, the extraction of keywords, topic words and context labels, and the extraction of corresponding entity nodes and relationship edges from a pre-constructed knowledge graph to construct a knowledge subgraph corresponding to the dialogue scene specifically comprise: collecting artificial customer service dialogue records, online customer service system conversation records and multi-round historical dialogue data to obtain dialogue scene data, the dialogue scene data comprising user input information, system reply information and context attributes, the context attributes comprising dialogue round information, speaker marks, timestamps and dialogue scene labels; performing word segmentation on the dialogue scene data and normalizing the semantic words to obtain a word sequence; performing hierarchical semantic analysis on the word sequence to extract keywords, topic words and context labels; organizing the keywords, topic words and context labels into a keyword set, a topic word set and a context label set, respectively; pre-constructing a knowledge graph, the knowledge graph comprising user intent entities, topic entities and scene entities, as well as semantic relationships, context dependency relationships and topic association relationships; extract entity nodes and relation edges corresponding to the keyword set, the subject keyword set and the context label set from a pre-constructed knowledge graph; Based on the context label set, the extracted entity nodes and relation edges are pruned and semantically consistent, only the nodes and relation edges related to the current dialogue context and semantically consistent are retained, and a knowledge subgraph corresponding to the current dialogue scene is constructed, which is composed of user intent entities, theme entities and scene entities related to the current dialogue scene as nodes, and semantic relations, context dependency relations and theme association relations related to the current dialogue scene as relation edges.
[0007] In an optional embodiment, the dialogue scene data and the knowledge subgraph are fused to obtain dialogue semantic structure data, specifically including: According to the user input information in the dialogue scene data, the corresponding dialogue text is extracted; The semantic features of the dialogue text are extracted, and the semantic features are mapped into semantic vectors through a pre-trained semantic embedding model; The semantic vectors and the vector representation of the user intent entity in the knowledge subgraph are calculated for similarity, and based on the user intent entity with the highest similarity, an instantiated user intent node is generated; Based on the system reply information in the dialogue scene data, the reply text is obtained, and the subject keywords are extracted from the reply text; The subject keywords and the theme entities in the knowledge subgraph are matched in vocabulary, and the semantic consistency is verified, and based on the verification result, the corresponding instantiated theme node is generated; Based on the context attribute, the context dependency relation in the knowledge subgraph is combined, the time sequence dependency modeling method is adopted, the context relation between the instantiated user intent node and the instantiated theme node is established, and the potential semantic association between the instantiated user intent node and the instantiated theme node in the context relation is supplemented; The instantiated user intent node, the instantiated theme node and the context relation thereof are structured and coded to form dialogue semantic structure data.
[0008] In an optional embodiment, the dialogue semantic structure data is analyzed in multiple dimensions to obtain a set of analysis results, specifically including: Based on the instantiated user intent node and the instantiated theme node in the dialogue semantic structure data, the node connection information is obtained; Based on the node connection information, the semantic path co-occurrence between nodes is analyzed to obtain the path co-occurrence features of the nodes; The path co-occurrence features and the historical interaction features of instantiated user intent nodes and instantiated topic nodes are numerically processed to obtain node quantification indicators. The historical interaction features include the frequency of node occurrence in historical dialogues, the number of times nodes co-occur, and the frequency of node-triggered user operations. Numerical processing is performed on the contextual relationship attributes between instantiated nodes to obtain a quantitative index of relationship edges. The contextual relationship attributes include semantic relevance and contextual dependency. Calculate the connection strength of a node based on its node quantification index and its relation edge quantification index; Calculate the importance score of a node based on its connection strength and node quantification metrics. The node importance index is calculated based on the connection strength and importance scores; Based on the preset division rules of node importance indicators, nodes are divided into key nodes and auxiliary nodes, and the node importance analysis results are formed by comparing the distribution differences of node importance indicators. Based on the contextual relationships in the dialogue semantic structure data, extract the semantic path between the instantiated user intent node and the instantiated topic node; The complexity coefficient is obtained by evaluating factors related to the number of nodes, node types, and relationship types in the semantic path. The semantic path is calculated for path length, and the complexity of the path is evaluated by combining the complexity coefficient, thereby generating path feature indicators. The paths are classified according to their characteristic indicators to obtain the main paths and abnormal paths, thus forming the dialogue path analysis results. Semantic similarity is calculated for instantiated user intent nodes in the knowledge subgraph to obtain the intent similarity matrix; Clustering of instantiated user intent nodes is performed based on intent similarity matrix. The node importance index corresponding to each instantiated user intent node is used as the node weight. The intent clustering index is calculated based on the node weight of each instantiated user intent node in the clustering result. Intent distribution characteristics are extracted based on intent clustering indicators to form intent clustering analysis results; The results of node importance analysis, dialogue path analysis, and intent clustering analysis are combined to generate a set of analysis results for subsequent visualization. The formula for calculating the node importance index is as follows: ; In the formula, For the first The node importance index of each node. For the first The node and the first The connection strength of each node For the first The importance score of each node In order to be with the first The total number of nodes associated with each node; The formula for calculating the path index is as follows: ; In the formula, For the first Semantic path feature indicators Let k be the length of the k-th path. Let be the complexity coefficient of the k-th path; The formula for calculating the intent clustering index is as follows: ; In the formula, For the first The intention clustering metric for each cluster. For the first The node weight of each instantiated user intent node. For the first The total number of instantiated user intent nodes in each cluster.
[0009] In an optional embodiment, based on the analysis result set, the node importance analysis results, dialogue path analysis results, and intent clustering analysis results are extracted to construct a dialogue semantic relationship graph, a dialogue flow path graph, and a dialogue hotspot distribution graph, respectively. The dialogue semantic relationship graph, dialogue flow path graph, and dialogue hotspot distribution graph are then subjected to unified coordinate mapping and weighted fusion to obtain visualized analysis results, specifically including: Based on the set of analysis results, extract the node importance analysis results, dialogue path analysis results, and intent clustering analysis results; Based on the node importance analysis results, key nodes and auxiliary nodes are mapped to graph nodes, and the size, color and transparency of the graph nodes are adjusted with semantic weight as a reference to obtain the dialogue semantic relationship graph. Based on the dialogue path analysis results, the main path and abnormal path are mapped into a path graph structure, and the path lines are rendered differently based on the path length and complexity to obtain the dialogue flow path graph. Based on the intent clustering analysis results, the set of instantiated user intent nodes within each cluster is mapped to the corresponding hotspot distribution area; The hotspot areas are associated with the corresponding scene nodes in the knowledge subgraph. The scene nodes are displayed as background annotations and the instantiated user intent nodes they cover are shown by connecting them. For each hotspot distribution area, calculate the intent concentration and intent dispersion of instantiated user intent nodes within the area; The intention concentration and dispersion are taken as indexes, mapped into color depth or transparency level by heat zone mode, used to identify core intention area and edge intention area, thereby generating a dialogue hotspot distribution map; The dialogue semantic relationship graph, dialogue flow path graph and dialogue hotspot distribution map are projected into a unified visualization coordinate system, keeping the consistency of nodes, paths and hotspot areas; The overlapping areas of the three types of graphs are weighted and fused and difference-contrasted, generating a comprehensive visualization picture as the final visualization analysis result; The calculation formula of the intention concentration formula is: ; In the formula, is the intention concentration of the hotspot area , is the set of instantiated user intention nodes in the hotspot area , is the occurrence frequency of the instantiated user intention node , is the weight of the instantiated user intention node in the hotspot area, is the total number of nodes in the area; The calculation formula of the intention dispersion formula is: ; In the formula, is the intention dispersion of the hotspot area .
[0010] In an optional embodiment, when new dialogue scene data is input, incremental updating is adopted to extract new associated nodes and relationship edges in the knowledge graph, update and recalculate the local semantic structure, and dynamically adjust the visualization analysis result, specifically including: When new dialogue scene data is input, incremental updating is adopted to dynamically extract new associated nodes and relationship edges in the knowledge graph, and without reconstructing the overall semantic structure, selectively adjust the affected local semantic area to obtain updated dialogue semantic structure data; Based on the updated dialogue semantic structure data, the corresponding node importance indicators, path indicators and intention clustering indicators are recalculated, and incremental analysis results are generated through difference comparison; When generating the visualization analysis result, an incremental merging mechanism is introduced to only locally refresh the affected semantic nodes, paths and hotspot areas, while maintaining the visualization stability of the unchanged part; Through the incremental merging and local refreshing process, dynamic updating of the dialogue semantic relationship graph, dialogue flow path graph and dialogue hotspot distribution graph is realized, thereby guaranteeing real-time and continuity of the visualization result.
[0011] Further, a dialogue scene data visualization analysis system based on a knowledge graph is proposed, which is used to implement the analysis method of any one of the above, comprising: A master control module is configured to receive dialogue scene data and analysis results transmitted by the data transmission module of each functional module, process and analyze the received data, and control the operation of each functional module according to the processing result; A data acquisition module is configured to acquire dialogue scene data, including user input information, system reply information and context attributes, and transmit the data to the master control module; A knowledge graph processing module is configured to perform word segmentation and semantic analysis on the dialogue scene data, extract keywords, topic words and context labels, and extract corresponding entity nodes and relationship edges from a pre-constructed knowledge graph to construct a knowledge subgraph corresponding to the dialogue scene; A semantic structure generation module is configured to fuse the dialogue scene data and the knowledge subgraph to generate dialogue semantic structure data, including instantiated user intent nodes, instantiated topic nodes and context relationships thereof; An analysis module is configured to perform multi-dimensional analysis based on the dialogue semantic structure data, acquire node importance analysis results, dialogue path analysis results and intent clustering analysis results, and generate an analysis result set; A visualization module is configured to generate a dialogue semantic relationship graph, a dialogue flow path graph and a dialogue hotspot distribution graph based on the analysis result set, and perform unified coordinate mapping and weighted fusion on the three types of graphs to obtain a visualization analysis result; An incremental updating module is configured to dynamically update knowledge graph nodes and relationship edges, recalculate local semantic structures, and dynamically adjust the visualization analysis result when new dialogue scene data is input.
[0012] In an optional embodiment, the master control module comprises: A data receiving unit is configured to receive dialogue data and analysis results from the data acquisition module, the knowledge graph processing module, the semantic structure generation module, the analysis module and the visualization module; A data processing unit is configured to pre-process, structure code and extract features of the received dialogue data, and provide the processing result to the analysis module and the visualization module; A control unit is configured to control the running states of the data collection module, the knowledge graph processing module, the semantic structure generation module, the analysis module, the visualization module and the incremental updating module according to the analysis result and the visualization feedback.
[0013] In an optional embodiment, the visualization module comprises: A semantic relationship graph generation unit is configured to map the key nodes and the auxiliary nodes into graph nodes according to the node importance analysis result, adjust the size, color and transparency of the graph nodes with the semantic weight, and generate a dialogue semantic relationship graph. A dialogue flow path graph generation unit is configured to map the main path and the abnormal path into a path graph structure according to the dialogue path analysis result, differentially render the path lines based on the path length and complexity, and generate a dialogue flow path graph. A dialogue hotspot distribution graph generation unit is configured to map the instantiated user intent nodes in each cluster into a hotspot area according to the intent clustering analysis result, combine the node weight to calculate the intent concentration and dispersion, generate a dialogue hotspot distribution graph through the heat mapping, and generate a dialogue hotspot distribution graph. A visualization fusion unit is configured to project the dialogue semantic relationship graph, the dialogue flow path graph and the dialogue hotspot distribution graph into a unified coordinate system, perform weighted fusion and difference comparison, and generate a comprehensive visualization analysis result.
[0014] In an optional embodiment, the analysis module comprises: A node importance analysis unit is configured to calculate the node connection strength and the importance score based on the node connection information in the dialogue semantic structure data and the historical interaction features of the node itself, and generate a node importance analysis result. A dialogue path analysis unit is configured to extract the semantic path between the user intent nodes and the theme nodes, calculate the path length and complexity, identify the main path and the abnormal path, and generate a dialogue path analysis result. An intent clustering analysis unit is configured to calculate the semantic similarity of the user intent nodes, perform clustering analysis, generate an intent index, and output an intent clustering analysis result. An incremental analysis unit is configured to, when new dialogue scene data is input, only locally refresh the affected nodes, paths and clusters, generate an incremental analysis result, and update the visualization display.
[0015] Compared with the prior art, the present application has the following beneficial effects: The method and system for visualizing and analyzing dialogue scene data based on a knowledge graph proposed in the present application fuse user input information, system reply information and context attributes, construct a knowledge subgraph containing user intent, theme entity and context relationship, and realize the generation of instantiated nodes and relationships of dialogue semantic structure. In the analysis process, no manual experience or fixed rules are relied on, and the dialogue path, node importance and intent distribution characteristics are dynamically reflected. Through multi-dimensional analysis, a semantic relationship graph, a flow path graph and a hot spot distribution graph are generated, the dialogue scene is visually presented in a panoramic manner, the dialogue scene understanding capability is improved, the core intent, abnormal path and user focus are found, and the efficiency and accuracy of dialogue optimization and decision support are improved. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 A flowchart of the method for visualizing and analyzing dialogue scene data based on a knowledge graph proposed in the present application is shown. Figure 2 A flowchart of the semantic structure generation in the present application is shown. Figure 3 A flowchart of the visual analysis in the present application is shown. Figure 4 A system framework diagram of the system for visualizing and analyzing dialogue scene data based on a knowledge graph proposed in the present application is shown. DETAILED DESCRIPTION
[0017] The following description is used to disclose the present application so that those skilled in the art can implement the present application. The preferred embodiments in the following description are only used as examples, and other obvious modifications can be thought of by those skilled in the art.
[0018] REFERENCE Figure 1 - Figure 4 As shown in the figure, the method for visualizing and analyzing dialogue scene data based on a knowledge graph in the embodiment of the present application includes: Obtaining dialogue scene data, the dialogue scene data including user input information, system reply information and context attributes; Performing word segmentation and semantic analysis on the dialogue scene data, extracting keywords, theme words and context labels, and extracting entity nodes and relationship edges corresponding thereto from a pre-constructed knowledge graph, to construct a knowledge subgraph corresponding to the dialogue scene; Fusing the dialogue scene data and the knowledge subgraph, to obtain dialogue semantic structure data, the dialogue semantic structure data including instantiated user intent nodes, instantiated theme nodes and context relationships thereof; Based on the dialogue semantic structure data, performing multi-dimensional analysis to obtain a set of analysis results, the multi-dimensional analysis including node importance analysis, dialogue path analysis and intent clustering analysis; According to the analysis result set, the node importance analysis result, the dialogue path analysis result and the intent clustering analysis result are extracted, and a dialogue semantic relationship graph, a dialogue flow path graph and a dialogue hotspot distribution graph are constructed respectively, and the dialogue semantic relationship graph, the dialogue flow path graph and the dialogue hotspot distribution graph are uniformly mapped and weighted fused to obtain a visual analysis result; When new dialogue scene data is input, incremental updating is adopted to extract new associated nodes and relationship edges in the knowledge graph, update and recalculate the local semantic structure, and dynamically adjust the visual analysis result.
[0019] Further, the dialogue scene data is segmented and semantically analyzed to extract keywords, topic words and context labels, and entity nodes and relationship edges corresponding to the keywords, topic words and context labels are extracted from a pre-constructed knowledge graph to construct a knowledge subgraph corresponding to the dialogue scene, specifically including: The artificial customer service dialogue record, the online customer service system conversation record and the multi-round historical dialogue data are collected to obtain dialogue scene data, and the dialogue scene data includes user input information, system reply information and context attributes, and the context attributes include dialogue turn information, speaker mark, timestamp and dialogue scene label. Specifically, the artificial customer service dialogue record, the online customer service system conversation record and the multi-round historical dialogue data are collected to obtain dialogue scene data. In addition to containing user input information and system reply information, each piece of dialogue data also includes context attributes, which are used to describe the context information of the dialogue, and specifically include dialogue turn information, speaker mark, timestamp and dialogue scene label. Among them, the dialogue turn information represents the order of the current information in a dialogue, which can be directly obtained in chronological order from the original dialogue record or system log; the speaker mark is used to distinguish between users and customer service personnel, which can be extracted from the role field in the dialogue record; the timestamp indicates the specific time of information sending or receiving, which can be directly read from the time field of the dialogue record; the dialogue scene label is used to identify the business scene to which the dialogue content belongs, which can be obtained through the classification field in the dialogue record or combined with rules / artificial labeling. In this way, each piece of dialogue data has complete context attributes, providing a basis for subsequent knowledge graph construction, semantic structure generation and visual analysis.
[0020] The dialogue scene data is segmented and processed, and is combined with semantic word normalization to obtain a word sequence; The word sequence is hierarchically semantically analyzed to extract keywords, topic words and context labels; The keywords, topic words and context labels are respectively arranged into a keyword set, a topic word set and a context label set; A knowledge graph is pre-constructed, and the knowledge graph includes user intent entities, topic entities and scene entities, as well as semantic relationships, context dependency relationships and topic association relationships; extracting, from the pre-constructed knowledge graph, entity nodes and relation edges corresponding to the keyword set, the theme word set and the context label set; Based on the context label set, the extracted entity nodes and relation edges are pruned and semantically consistent, only the nodes and relation edges related to the current dialogue context and semantically consistent are retained, and a knowledge subgraph corresponding to the current dialogue scene is constructed, which is composed of user intent entities, theme entities and scene entities related to the current dialogue scene as nodes, and semantic relations, context dependency relations and theme association relations related to the current dialogue scene as relation edges. Specifically, after the preliminary arrangement of the entity nodes and relation edges extracted from the knowledge graph, the context label set is introduced to prune and constrain the semantic consistency. Specifically, first, according to the context label set corresponding to the dialogue data (such as "account information query", "account security management", "payment problem handling"), the candidate nodes and relation edges related to the current context semantics are filtered out; Then compare the semantic attributes of the candidate nodes with the context labels, if the semantics are inconsistent or lack of association, the node and its related relation edges will be pruned from the current subgraph; For nodes with ambiguous semantics or ambiguity, through the context dependency relation, only the nodes and relation edges with semantic consistency higher than the preset threshold are retained. Through this process, the constructed knowledge subgraph can focus on the semantic information related to the current dialogue scene while maintaining integrity, avoiding redundant node interference, thereby improving the accuracy of subsequent semantic structure generation and analysis.
[0021] Further, the dialogue scene data and the knowledge subgraph are fused to obtain dialogue semantic structure data, specifically including: According to the user input information in the dialogue scene data, the corresponding dialogue text is extracted; Perform semantic feature extraction on the dialogue text, and map the semantic features to semantic vectors through a pre-trained semantic embedding model; Calculate the similarity between the semantic vector and the vector representation of the user intent entity in the knowledge subgraph, and generate an instantiated user intent node based on the user intent entity with the highest similarity; Based on the system reply information in the dialogue scene data, the reply text is obtained, and the theme words are extracted from the reply text; Match the theme words with the theme entities in the knowledge subgraph, and perform semantic consistency test, and generate corresponding instantiated theme nodes based on the test results; Based on the context attribute, combined with the context dependency relation in the knowledge subgraph, a time sequence dependency modeling method is adopted to establish the context relationship between the instantiated user intent node and the instantiated theme node, and the potential semantic association between the instantiated user intent node and the instantiated theme node is supplemented in the context relationship. Understandably, the first step is to extract the corresponding dialogue text based on the user input information in the dialogue scenario data. For example, if a user inputs "I want to change the bound mobile phone number," the system will use this input as the dialogue text. Then, semantic features are extracted from this text to obtain lexical features, dependency relation features, etc., and these semantic features are mapped into semantic vector representations using pre-trained semantic embedding models (such as BERT, Word2Vec, etc.).
[0022] Subsequently, the semantic vector is compared with the vector representations of existing user intent entities in the knowledge subgraph to calculate their similarity, finding the user intent entity that is closest to the input semantics (e.g., "modify account information"). Based on the entity with the highest similarity, an instantiated user intent node corresponding to the current dialogue is generated, thereby clarifying the user's current intent.
[0023] Next, based on the system response information in the dialogue scenario data, the response text is obtained and the topic words are extracted. For example, when the system responds "Please provide a new mobile phone number," the extracted topic word is "mobile phone number." This topic word is then matched with the topic entities in the knowledge subgraph, and further semantic consistency is checked (e.g., through synonym expansion) to confirm that "mobile phone number" and "account binding information" are semantically consistent, thereby generating the corresponding instantiated topic node.
[0024] Building upon this foundation, based on the contextual attributes of the dialogue data (such as dialogue turn, timestamp, and speaker role), and combined with the contextual dependencies in the knowledge subgraph, a temporal dependency modeling approach is adopted to establish the contextual relationship between instantiated user intent nodes and instantiated topic nodes. This contextual relationship is used to represent the temporal dependency and logical connection between user intent and topic during the dialogue process; for example, the intent to "modify account information" leads to the topic of "phone number." Furthermore, the potential semantic associations between user intent nodes and topic nodes are supplemented into the contextual relationship. These potential semantic associations are used to characterize deeper connections not directly stated in the dialogue but which can be inferred semantically, such as "phone number belongs to account information" and "modifying phone number involves account security," thus enabling the constructed knowledge subgraph to more comprehensively reflect the semantic structure of the dialogue scenario.
[0025] The instantiated user intent nodes, instantiated topic nodes, and their contextual relationships are structured and encoded to form dialogue semantic structure data; Understandably, the identified instantiated user intent nodes, instantiated topic nodes, and the contextual relationships between them are structured and encoded. Specifically, the semantic feature vectors of the nodes are combined with contextual relationship triples into a unified graph structure representation, such as (instantiated user intent node: Modify account information, contextual relationship: introduction, instantiated topic node: phone number), and the relationship type is encoded independently. In this way, the original text information is transformed into computable structured data, which not only preserves semantic information but also facilitates subsequent node quantification, path analysis, cluster analysis, and visualization.
[0026] Furthermore, multi-dimensional analysis is performed based on the semantic structure data of the dialogue to obtain a set of analysis results, specifically including: Based on the instantiated user intent nodes and instantiated topic nodes in the dialogue semantic structure data, obtain node connection information; Based on the analysis of node connection information, the semantic path co-occurrence of nodes is analyzed to obtain the path co-occurrence features of nodes; The path co-occurrence features and the historical interaction features of instantiated user intent nodes and instantiated topic nodes are numerically processed to obtain node quantification indicators. The historical interaction features include the frequency of node occurrence in historical dialogues, the number of times nodes co-occur, and the frequency of node-triggered user operations. Numerical processing is performed on the contextual relationship attributes between instantiated nodes to obtain a quantitative index of relationship edges. The contextual relationship attributes include semantic relevance and contextual dependency. Specifically, firstly, node connection information is obtained based on the instantiated user intent nodes and instantiated topic nodes in the dialogue semantic structure data. This node connection information includes the adjacency relationships and multi-hop path information of the nodes. Then, semantic path analysis is performed on the connection information of each instantiated node, that is, traversing the node's adjacent nodes and multi-hop nodes, and counting the frequency and distribution of the node co-occurring with other nodes in the same semantic path. For example, when the "change phone number" intent node and the "phone number" topic node frequently co-occur in multiple semantic paths, it indicates a strong semantic co-occurrence relationship between them. Based on the above statistical results, the path co-occurrence characteristics of each instantiated node are obtained.
[0027] Then, interaction features of nodes are extracted from historical dialogue data, including the frequency of node appearance in the dialogue, the number of times it co-occurs with other nodes, and the frequency with which the node triggers user actions. Subsequently, the path co-occurrence features and historical interaction features are uniformly numerically processed: first, the original count values are smoothed and compressed to suppress the influence of extreme values (e.g., logarithmic compression is used for high-frequency values), and then each feature is normalized and mapped to ensure that the two are on a unified scale; for low-frequency or rare nodes, a smoothing factor is introduced for compensation to avoid instability caused by zero values. Finally, the path co-occurrence features and historical interaction features are weighted and merged according to a preset fusion strategy to obtain a single quantitative representation of each instantiated node, namely the node quantification index. This index retains the connectivity information of the node in the semantic path and integrates its interactive behavior features in historical conversations, providing input for subsequent analysis.
[0028] Finally, the contextual relationship attributes between instantiated nodes are numerically processed: First, the semantic tightness of the contextual relationship is obtained by statistically analyzing the common frequency and association of connected nodes in multi-turn dialogues, and then calibrated by combining manual annotation or historical statistical data to avoid distortion caused by model bias; at the same time, the temporal sequence and dependency frequency of nodes in multi-turn dialogues are analyzed to obtain the contextual dependency strength. For example, when an intent node frequently evokes specific topic nodes in the dialogue, the dependency strength of the contextual relationship is higher; then, the semantic tightness and contextual dependency strength are normalized and mapped, and attributes with large noise are smoothed and corrected. Finally, the two are synthesized according to the preset fusion method to obtain a quantitative index of contextual relationship that comprehensively reflects semantic relevance and dependency strength.
[0029] Calculate the connection strength of a node based on its node quantification index and its relation edge quantification index; Calculate the importance score of a node based on its connection strength and node quantification metrics. Understandably, for a given node, all its neighboring nodes are traversed, and the overall connection strength of the node is calculated by combining the quantitative indicators of the neighboring nodes with the corresponding quantitative indicators of the context relationship. This reflects the connectivity of the node in the semantic structure data. Subsequently, the importance score of the node is calculated. This score not only considers the node's own quantitative indicators but also combines the node's connection strength. The two are weighted and fused according to preset weights. This reflects both the importance of the node in historical dialogues and path co-occurrence and its connection value in the global semantic network. The final importance score serves as the core basis for calculating the overall importance index of the node and is used for the subsequent classification of key nodes and auxiliary nodes.
[0030] The node importance index is calculated based on the connection strength and importance scores; Based on a pre-defined rule for classifying nodes according to their importance indices, nodes are divided into key nodes and auxiliary nodes. The node importance analysis results are generated by comparing the distribution differences of these indices. A node is identified as a key node when its importance index exceeds a pre-defined threshold or when it occupies a high quantile in the importance index distribution; otherwise, it is classified as an auxiliary node. Key nodes typically play a core role in dialogue semantic understanding, directly influencing the dialogue flow and the determination of user intent. Auxiliary nodes, on the other hand, serve as supplementary information, used to refine contextual connections and semantic reasoning. This classification method highlights the core elements of the dialogue semantic structure while avoiding interference from redundant nodes, achieving efficient modeling of the dialogue semantic structure.
[0031] Based on the contextual relationships in the dialogue semantic structure data, extract the semantic path between the instantiated user intent node and the instantiated topic node; The complexity coefficient is obtained by evaluating factors related to the number of nodes, node types, and relationship types in the semantic path. The semantic path is calculated for path length, and the complexity of the path is evaluated by combining the complexity coefficient, thereby generating path feature indicators. The paths are classified according to their characteristic indicators to obtain the main paths and abnormal paths, thus forming the dialogue path analysis results. Specifically, starting from any user intent node, the system traverses adjacent topic nodes and multi-hop nodes along their context, recording the connection order between nodes to obtain a complete semantic path sequence. Then, the complexity of each semantic path is evaluated. During the evaluation, firstly, the number of nodes on the path is counted; more nodes indicate a richer information flow. Secondly, different weights are assigned to the path based on node type (e.g., intent node or topic node) to reflect the contribution of node type to path complexity. Finally, the type and direction of contextual relationships in the path are considered, such as the existence of unidirectional or circular dependencies, to make the path complexity more accurately reflect the actual semantic associations. Through the above evaluation, a complexity coefficient is obtained for each path. Next, the path length of each semantic path is calculated, typically the number of nodes or edges in the path, and combined with the complexity coefficient to generate a comprehensive path feature index. This index can simultaneously reflect the path's length and structural complexity, providing a quantitative basis for subsequent classification. Finally, the semantic paths are classified according to the path feature index. Specifically, paths with path feature indicators above a preset threshold or ranking in the top 30% of the overall path feature indicator distribution are identified as primary paths, representing the most frequently occurring node sequences in the dialogue that carry the core semantic flow. Paths with path feature indicators below the preset threshold or ranking in the bottom 30% of the overall path feature indicator distribution are identified as anomalous paths, potentially reflecting rare dialogue patterns or potential anomalies. By classifying and processing all semantic paths, a complete dialogue path analysis result is formed, providing quantitative basis for dialogue flow optimization, anomaly detection, and visualization.
[0032] Semantic similarity is calculated for instantiated user intent nodes in the knowledge subgraph to obtain the intent similarity matrix; Specifically, for each pair of instantiated user intent nodes, their semantic vector representations are extracted. The semantic similarity between the two nodes is calculated using cosine similarity or other similarity metrics, resulting in a value between zero and one, where one indicates complete semantic overlap and zero indicates complete semantic irrelevance. After calculating the pairwise semantic similarity for all nodes, an intent similarity matrix is constructed. Each row and column of the matrix corresponds to an instantiated user intent node, and any element in the matrix represents the semantic similarity between the corresponding two nodes. Due to the symmetry of semantic similarity, this matrix is typically symmetric. Using this intent similarity matrix, intent clustering analysis can be performed, grouping semantically similar user intent nodes into the same cluster to identify common user intent patterns. This clustering result can be used for subsequent intent distribution analysis and dialogue hotspot identification, providing a quantitative basis for visualization analysis.
[0033] Clustering of instantiated user intent nodes is performed based on intent similarity matrix. The node importance index corresponding to each instantiated user intent node is used as the node weight. The intent clustering index is calculated based on the node weight of each instantiated user intent node in the clustering result. Intent distribution characteristics are extracted based on intent clustering indicators to form intent clustering analysis results; The results of node importance analysis, dialogue path analysis, and intent clustering analysis are combined to generate a set of analysis results for subsequent visualization. The formula for calculating the node importance index is as follows: ; In the formula, For the first The node importance index of each node. For the first The node and the first The connection strength of each node For the first The importance score of each node In order to be with the first The total number of nodes associated with each node; The formula for calculating the path metric is: ; In the formula, For the first Semantic path feature indicators Let k be the length of the k-th path. Let be the complexity coefficient of the k-th path; The formula for calculating the intention clustering index is: ; In the formula, For the first The intention clustering metric for each cluster. For the first The node weight of each instantiated user intent node. For the first The total number of instantiated user intent nodes in each cluster.
[0034] Furthermore, based on the analysis results set, node importance analysis results, dialogue path analysis results, and intent clustering analysis results are extracted to construct a dialogue semantic relationship graph, a dialogue flow path graph, and a dialogue hotspot distribution graph, respectively. Then, a unified coordinate mapping and weighted fusion are performed on the dialogue semantic relationship graph, dialogue flow path graph, and dialogue hotspot distribution graph to obtain visualized analysis results, specifically including: Based on the set of analysis results, extract the node importance analysis results, dialogue path analysis results, and intent clustering analysis results; Based on the node importance analysis results, key nodes and auxiliary nodes are mapped to graph nodes, and the size, color and transparency of the graph nodes are adjusted with semantic weight as a reference to obtain the dialogue semantic relationship graph. Based on the dialogue path analysis results, the main path and abnormal path are mapped into a path graph structure, and the path lines are rendered differently based on the path length and complexity to obtain the dialogue flow path graph. Based on the intent clustering analysis results, the set of instantiated user intent nodes within each cluster is mapped to the corresponding hotspot distribution area; The hotspot areas are associated with the corresponding scene nodes in the knowledge subgraph. The scene nodes are displayed as background annotations and the instantiated user intent nodes they cover are shown by connecting them. For each hotspot distribution area, calculate the intent concentration and intent dispersion of instantiated user intent nodes within the area; Intent concentration and dispersion are used as indicators and mapped to color depth or transparency through heat maps to identify core intent areas and peripheral intent areas, thereby generating a dialogue hotspot distribution map. Specifically, based on the node importance analysis results, instantiated user intent nodes and topic nodes are distinguished into key nodes and auxiliary nodes, and these nodes are mapped to visual graphical nodes. During the mapping process, the size, color, and transparency of the nodes are adjusted according to their importance scores. For example, key nodes with high importance scores are displayed as larger, darker, and more transparent graphical nodes; auxiliary nodes are displayed as smaller, lighter, and less transparent nodes. Through this visual adjustment, users can intuitively distinguish between core and auxiliary nodes, thereby understanding the main elements in the semantic structure of the dialogue.
[0035] Using the results of dialogue path analysis, primary and abnormal paths are mapped to a path graph structure. When drawing the paths, the path lines are rendered differently based on their length, complexity, and frequency of occurrence. Specifically, longer or more complex paths can be displayed with bolder lines or different line styles; primary paths are displayed with highlighted colors or thick lines, while abnormal paths are represented with low-saturation colors or dashed lines. In this way, the path graph not only presents the structure of the dialogue flow but also intuitively reflects the importance and abnormality of each path.
[0036] Based on the intent clustering analysis results, the instantiated user intent nodes in each cluster are mapped to corresponding hotspot regions, and the intent concentration and dispersion are calculated by combining the node weights. Regions with high intent concentration are displayed as dark-colored, highly transparent hotspot regions, indicating that user intent is highly concentrated within these regions; regions with high intent dispersion are displayed as light-colored, low-transparency regions, indicating that user intent is more dispersed. Hotspot regions are associated with scene nodes in the knowledge subgraph, and the instantiated user intent nodes they cover are displayed by connecting them, thus forming an intuitive hotspot distribution map.
[0037] The dialogue semantic relationship graph, dialogue flow path graph, and dialogue hotspot distribution graph are projected onto a unified visualization coordinate system to maintain the consistency of nodes, paths, and hotspot areas. The overlapping areas of the three types of images are weighted, fused, and compared to generate a comprehensive visualization, which serves as the final visualization analysis result. The formula for calculating the concentration of intent is as follows: ; In the formula, Hotspots are divided into different regions The degree of intent concentration hotspot areas The collection of instantiated user intent nodes within. To instantiate the user intent node Frequency of occurrence The weights of instantiated user intent nodes within the hotspot area. This represents the total number of nodes within the region. The formula for calculating the intention dispersion is as follows: ; In the formula, hotspot areas The intentional dispersion.
[0038] Furthermore, when new dialogue scenario data is input, an incremental update method is used to extract newly added related nodes and relation edges in the knowledge graph, update and recalculate the local semantic structure, and dynamically adjust the visualization analysis results, specifically including: When new dialogue scenario data is input, the newly added related nodes and relation edges in the knowledge graph are dynamically extracted using an incremental update method. Without reconstructing the overall semantic structure, the affected local semantic regions are selectively adjusted to obtain the updated dialogue semantic structure data. Based on the updated dialogue semantic structure data, the corresponding node importance index, path index and intent clustering index are recalculated, and incremental analysis results are generated through difference comparison. When generating visualization analysis results, an incremental merging mechanism is introduced to only refresh the affected semantic nodes, paths and hotspot areas locally, while maintaining the visualization stability of the unchanged parts. Understandably, the visualization analysis results need to be dynamically updated when new dialogue scenario data is input. To avoid redrawing the entire graph with each update, the system introduces an incremental merging mechanism, only partially refreshing semantic nodes, paths, and hotspot areas affected by new or changed data. Specifically, by comparing the differences between the new data and the existing semantic structure, the system identifies the nodes and relationships that need updating, such as newly added user intent nodes, newly emerging topic nodes, or new semantic paths. Subsequently, only the quantitative indicators, path features, and intent clustering information of these nodes and paths are recalculated, and the corresponding graphic elements in the visualization are updated. For nodes, paths, and hotspot areas that have not changed, their original display status is maintained, including visual attributes such as position, size, color, and transparency, thus ensuring the stability and continuity of the entire visualization interface. Through this combination of partial refresh and overall stability, users will not experience interface jumps or redrawing delays due to partial updates when viewing the visualization results, while the changes brought about by new dialogue scenario data can be reflected in real time, thereby achieving dynamic, continuous, and intuitive visualization analysis.
[0039] Through this incremental merging and local refresh process, the dialogue semantic relationship graph, dialogue flow path graph, and dialogue hotspot distribution graph are dynamically updated, thereby ensuring the real-time and continuous nature of the visualization results.
[0040] Furthermore, a knowledge graph-based dialogue scenario data visualization and analysis system is proposed to implement the analysis methods described above, including: The main control module receives dialogue scenario data and analysis results transmitted by each functional module through the data transmission module, processes and analyzes the received data, and controls the operation of each functional module based on the processing results. The data acquisition module is used to acquire dialogue scenario data, including user input information, system response information and context attributes, and transmit the data to the main control module. The knowledge graph processing module is used to perform word segmentation and semantic parsing on dialogue scenario data, extract keywords, topic words and context labels, and extract corresponding entity nodes and relation edges from the pre-built knowledge graph to construct a knowledge subgraph corresponding to the dialogue scenario. The semantic structure generation module is used to fuse dialogue scenario data with knowledge subgraphs to generate dialogue semantic structure data, including instantiated user intent nodes, instantiated topic nodes and their contextual relationships. The analysis module is used to perform multi-dimensional analysis based on dialogue semantic structure data, obtain node importance analysis results, dialogue path analysis results, and intent clustering analysis results, and generate a set of analysis results. The visualization module is used to generate a dialogue semantic relationship diagram, a dialogue flow path diagram, and a dialogue hotspot distribution diagram based on the analysis result set, and to perform unified coordinate mapping and weighted fusion on the three types of diagrams to obtain the visualization analysis results. The incremental update module is used to dynamically update knowledge graph nodes and relation edges, recalculate local semantic structures, and dynamically adjust visualization analysis results when new dialogue scenario data is input.
[0041] Furthermore, the main control module includes: The data receiving unit is used to receive dialogue data and analysis results from the data acquisition module, knowledge graph processing module, semantic structure generation module, analysis module, and visualization module. The data processing unit is used to preprocess, structure, and extract features from the received dialogue data, and provides the processing results to the analysis and visualization modules. The control unit is used to control the operation status of the data acquisition module, knowledge graph processing module, semantic structure generation module, analysis module, visualization module, and incremental update module based on the analysis results and visualization feedback.
[0042] Furthermore, the visualization module includes: A semantic relationship graph generation unit is used to map key nodes and auxiliary nodes into graphical nodes based on the node importance analysis results, and adjust the size, color and transparency of the graphical nodes with semantic weights to generate a dialogue semantic relationship graph. The dialogue flow path diagram generation unit is used to map the main path and abnormal path into a path diagram structure based on the dialogue path analysis results, and to perform differentiated rendering of the path lines based on the path length and complexity, thereby generating a dialogue flow path diagram. A dialogue hotspot distribution map generation unit is used to map instantiated user intent nodes within each cluster to hotspot areas based on intent clustering analysis results, and calculate intent concentration and dispersion by combining node weights, and generate a dialogue hotspot distribution map through heat mapping. The visualization fusion unit is used to project the dialogue semantic relationship graph, dialogue flow path graph and dialogue hotspot distribution graph onto a unified coordinate system, and perform weighted fusion and difference comparison to generate a comprehensive visualization analysis result.
[0043] Furthermore, the analysis module includes: The node importance analysis unit is used to calculate the node connection strength and importance score based on the node connection information in the dialogue semantic structure data and the node's own historical interaction characteristics, and generate the node importance analysis results. The dialogue path analysis unit is used to extract the semantic path between the user intent node and the topic node, calculate the path length and complexity, identify the main path and abnormal path, and generate dialogue path analysis results. The intent clustering analysis unit is used to calculate the semantic similarity of user intent nodes, perform clustering analysis, generate intent metrics, and output intent clustering analysis results. The incremental analysis unit is used to partially refresh only the affected nodes, paths, and clusters when new dialogue scenario data is input, generating incremental analysis results to update the visualization.
[0044] In summary, the advantages of this invention are as follows: by integrating user input information, system response information, and contextual attributes, and combining them with a pre-constructed knowledge graph, it enables the dynamic instantiation of user intent nodes, topic nodes, and contextual relationships; based on multi-dimensional analysis of node importance, semantic paths, and intent clustering, it can generate semantic relationship graphs, dialogue flow path graphs, and hotspot distribution graphs, achieving panoramic visualization of the dialogue scenario; and it supports an incremental update mechanism, which can dynamically adjust the local semantic structure and visualization results when new dialogue data is input, thereby significantly improving the efficiency and accuracy of dialogue understanding, abnormal path identification, and core intent discovery.
[0045] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A knowledge graph-based method for visualizing and analyzing dialogue scenario data, characterized in that, include: Acquire dialogue scenario data, which includes user input information, system response information, and context attributes; The dialogue scenario data is segmented and semantically parsed to extract keywords, topic words and context labels. The corresponding entity nodes and relationship edges are extracted from the pre-built knowledge graph to construct a knowledge subgraph corresponding to the dialogue scenario. The dialogue scenario data is fused with the knowledge subgraph to obtain dialogue semantic structure data, which includes instantiated user intent nodes, instantiated topic nodes and their contextual relationships. Multi-dimensional analysis is performed based on dialogue semantic structure data to obtain a set of analysis results. The multi-dimensional analysis includes node importance analysis, dialogue path analysis, and intent clustering analysis. Based on the set of analysis results, the node importance analysis results, dialogue path analysis results, and intent clustering analysis results are extracted. Dialogue semantic relationship graph, dialogue flow path graph, and dialogue hotspot distribution graph are constructed respectively. The dialogue semantic relationship graph, dialogue flow path graph, and dialogue hotspot distribution graph are then subjected to unified coordinate mapping and weighted fusion to obtain visualized analysis results. When new dialogue scenario data is input, the newly added related nodes and relation edges in the knowledge graph are extracted using an incremental update method, the local semantic structure is updated and recalculated, and the visualization analysis results are dynamically adjusted.
2. The knowledge graph-based data visualization and analysis method for dialogue scenarios according to claim 1, characterized in that, The process of segmenting and semantically parsing the dialogue scenario data to extract keywords, topic terms, and context labels, and then extracting corresponding entity nodes and relation edges from a pre-built knowledge graph to construct a knowledge subgraph corresponding to the dialogue scenario, specifically includes: Collect human customer service dialogue records, online customer service system conversation records, and multi-round historical dialogue data to obtain dialogue scenario data. The dialogue scenario data includes user input information, system response information, and context attributes. The context attributes include dialogue round information, speaker markers, timestamps, and dialogue scenario tags. The dialogue scene data is segmented into words and normalized using synonyms to obtain a word sequence; Hierarchical semantic analysis is performed on word sequences to extract keywords, thematic terms, and contextual tags; Organize keywords, topic terms, and context tags into keyword sets, topic term sets, and context tag sets, respectively. A knowledge graph is pre-constructed, which includes user intent entities, topic entities, and scene entities, as well as semantic relationships, contextual dependencies, and topic associations. Based on the keyword set, topic term set, and context tag set, extract the corresponding entity nodes and relationship edges from the pre-built knowledge graph; Based on the context label set, the extracted entity nodes and relation edges are pruned and semantically consistent, retaining only nodes and relation edges that are relevant to the current dialogue context and semantically consistent. A knowledge subgraph corresponding to the current dialogue scenario is constructed. The knowledge subgraph consists of user intent entities, topic entities, and scene entities related to the current dialogue scenario as nodes, and semantic relations, contextual dependencies, and topic associations related to the current dialogue scenario as relation edges.
3. The knowledge graph-based method for visualizing and analyzing dialogue scenario data according to claim 1, characterized in that, The process of fusing dialogue scenario data with knowledge subgraphs to obtain dialogue semantic structure data specifically includes: Extract the corresponding dialogue text based on the user input information in the dialogue scenario data; Semantic features are extracted from the dialogue text, and the semantic features are mapped into semantic vectors through a pre-trained semantic embedding model; The semantic vector is compared with the vector representation of the user intent entity in the knowledge subgraph. Based on the user intent entity with the highest similarity, an instantiated user intent node is generated. Based on the system response information in the dialogue scenario data, obtain the response text and extract the topic words from the response text; The topic words are matched with the topic entities in the knowledge subgraph, and semantic consistency is checked. Based on the check results, the corresponding instantiated topic nodes are generated. Based on contextual attributes and combined with contextual dependencies in the knowledge subgraph, a temporal dependency modeling approach is adopted to establish contextual relationships between instantiated user intent nodes and instantiated topic nodes, and to supplement the contextual relationships with potential semantic associations between instantiated user intent nodes and instantiated topic nodes. The instantiated user intent nodes, instantiated topic nodes, and their contextual relationships are structured and encoded to form dialogue semantic structure data.
4. The knowledge graph-based data visualization and analysis method for dialogue scenarios according to claim 1, characterized in that, The multi-dimensional analysis based on dialogue semantic structure data yields a set of analysis results, specifically including: Based on the instantiated user intent nodes and instantiated topic nodes in the dialogue semantic structure data, obtain node connection information; Based on the analysis of node connection information, the semantic path co-occurrence of nodes is analyzed to obtain the path co-occurrence features of nodes; The path co-occurrence features and the historical interaction features of instantiated user intent nodes and instantiated topic nodes are numerically processed to obtain node quantitative indicators; Numerical processing is performed on the contextual relationship attributes between instantiated nodes to obtain a quantitative index of relationship edges. The contextual relationship attributes include semantic relevance and contextual dependency. Calculate the connection strength of a node based on its node quantification index and its relation edge quantification index; Calculate the importance score of a node based on its connection strength and node quantification metrics. The node importance index is calculated based on the connection strength and importance scores; Based on the preset division rules of node importance indicators, nodes are divided into key nodes and auxiliary nodes, and the node importance analysis results are formed by comparing the distribution differences of node importance indicators. Based on the contextual relationships in the dialogue semantic structure data, extract the semantic path between the instantiated user intent node and the instantiated topic node; The complexity coefficient is obtained by evaluating factors related to the number of nodes, node types, and relationship types in the semantic path. The semantic path is calculated for path length, and the complexity of the path is evaluated by combining the complexity coefficient, thereby generating path feature indicators. The paths are classified according to their characteristic indicators to obtain the main paths and abnormal paths, thus forming the dialogue path analysis results. Semantic similarity is calculated for instantiated user intent nodes in the knowledge subgraph to obtain the intent similarity matrix; Clustering of instantiated user intent nodes is performed based on intent similarity matrix. The node importance index corresponding to each instantiated user intent node is used as the node weight. The intent clustering index is calculated based on the node weight of each instantiated user intent node in the clustering result. Intent distribution characteristics are extracted based on intent clustering indicators to form intent clustering analysis results; The results of node importance analysis, dialogue path analysis, and intent clustering analysis are combined to generate a set of analysis results for subsequent visualization. The formula for calculating the node importance index is as follows: ; In the formula, For the first The node importance index of each node. For the first The node and the first The connection strength of each node For the first The importance score of each node In order to be with the first The total number of nodes associated with each node; The formula for calculating the path index is as follows: ; In the formula, For the first Semantic path feature indicators Let k be the length of the k-th path. Let be the complexity coefficient of the k-th path; The formula for calculating the intent clustering index is as follows: ; In the formula, For the first The intention clustering metric for each cluster. For the first The node weight of each instantiated user intent node. For the first The total number of instantiated user intent nodes in each cluster.
5. The knowledge graph-based method for visualizing and analyzing dialogue scenario data according to claim 1, characterized in that, Based on the analysis result set, the node importance analysis results, dialogue path analysis results, and intent clustering analysis results are extracted to construct a dialogue semantic relationship graph, a dialogue flow path graph, and a dialogue hotspot distribution graph, respectively. Then, a unified coordinate mapping and weighted fusion are performed on the dialogue semantic relationship graph, dialogue flow path graph, and dialogue hotspot distribution graph to obtain visualized analysis results, specifically including: Based on the set of analysis results, extract the node importance analysis results, dialogue path analysis results, and intent clustering analysis results; Based on the node importance analysis results, key nodes and auxiliary nodes are mapped to graph nodes, and the size, color and transparency of the graph nodes are adjusted with semantic weight as a reference to obtain the dialogue semantic relationship graph. Based on the dialogue path analysis results, the main path and abnormal path are mapped into a path graph structure, and the path lines are rendered differently based on the path length and complexity to obtain the dialogue flow path graph. Based on the intent clustering analysis results, the set of instantiated user intent nodes within each cluster is mapped to the corresponding hotspot distribution area; The hotspot areas are associated with the corresponding scene nodes in the knowledge subgraph. The scene nodes are displayed as background annotations and the instantiated user intent nodes they cover are shown by connecting them. For each hotspot distribution area, calculate the intent concentration and intent dispersion of instantiated user intent nodes within the area; Intent concentration and dispersion are used as indicators and mapped to color depth or transparency through heat maps to identify core intent areas and peripheral intent areas, thereby generating a dialogue hotspot distribution map. The dialogue semantic relationship graph, dialogue flow path graph, and dialogue hotspot distribution graph are projected onto a unified visualization coordinate system to maintain the consistency of nodes, paths, and hotspot areas. The overlapping areas of the three types of images are weighted, fused, and compared to generate a comprehensive visualization, which serves as the final visualization analysis result.
6. The knowledge graph-based method for visualizing and analyzing dialogue scenario data according to claim 1, characterized in that, When new dialogue scenario data is input, an incremental update method is used to extract newly added associated nodes and relation edges in the knowledge graph, update and recalculate the local semantic structure, and dynamically adjust the visualization analysis results, specifically including: When new dialogue scenario data is input, the newly added related nodes and relation edges in the knowledge graph are dynamically extracted using an incremental update method. Without reconstructing the overall semantic structure, the affected local semantic regions are selectively adjusted to obtain the updated dialogue semantic structure data. Based on the updated dialogue semantic structure data, the corresponding node importance index, path index and intent clustering index are recalculated, and incremental analysis results are generated through difference comparison. When generating visualization analysis results, an incremental merging mechanism is introduced to only refresh the affected semantic nodes, paths and hotspot areas locally, while maintaining the visualization stability of the unchanged parts. Through this incremental merging and local refresh process, the dialogue semantic relationship graph, dialogue flow path graph, and dialogue hotspot distribution graph are dynamically updated, thereby ensuring the real-time and continuous nature of the visualization results.
7. A knowledge graph-based dialogue scenario data visualization and analysis system, used to implement the analysis method as described in any one of claims 1-6, characterized in that, include: The main control module is used to receive dialogue scenario data and analysis results transmitted by each functional module through the data transmission module, process and analyze the received data, and control the operation of each functional module according to the processing results. The data acquisition module is used to acquire dialogue scenario data, including user input information, system response information and context attributes, and transmit the data to the main control module. The knowledge graph processing module is used to perform word segmentation and semantic parsing on dialogue scenario data, extract keywords, topic words and context labels, and extract corresponding entity nodes and relation edges from the pre-built knowledge graph to construct a knowledge subgraph corresponding to the dialogue scenario. A semantic structure generation module is used to fuse dialogue scenario data with knowledge subgraphs to generate dialogue semantic structure data, including instantiated user intent nodes, instantiated topic nodes and their contextual relationships. The analysis module is used to perform multi-dimensional analysis based on dialogue semantic structure data, obtain node importance analysis results, dialogue path analysis results and intent clustering analysis results, and generate a set of analysis results. The visualization module is used to generate a dialogue semantic relationship graph, a dialogue flow path graph, and a dialogue hotspot distribution graph based on the analysis result set, and to perform unified coordinate mapping and weighted fusion on the three types of graphs to obtain the visualization analysis results. The incremental update module is used to dynamically update knowledge graph nodes and relation edges, recalculate local semantic structures, and dynamically adjust visualization analysis results when new dialogue scenario data is input.
8. The knowledge graph-based dialogue scenario data visualization and analysis system according to claim 7, characterized in that, The main control module includes: The data receiving unit is used to receive dialogue data and analysis results from the data acquisition module, knowledge graph processing module, semantic structure generation module, analysis module, and visualization module. The data processing unit is used to preprocess, structure, and extract features from the received dialogue data, and provide the processing results to the analysis module and the visualization module. The control unit is used to control the operation status of the data acquisition module, knowledge graph processing module, semantic structure generation module, analysis module, visualization module, and incremental update module based on the analysis results and visualization feedback.
9. A knowledge graph-based dialogue scenario data visualization and analysis system according to claim 7, characterized in that, The visualization module includes: A semantic relationship graph generation unit is used to map key nodes and auxiliary nodes into graphical nodes based on the node importance analysis results, and adjust the size, color and transparency of the graphical nodes with semantic weights to generate a dialogue semantic relationship graph. The dialogue flow path diagram generation unit is used to map the main path and abnormal path into a path diagram structure based on the dialogue path analysis results, and to perform differentiated rendering of the path lines based on the path length and complexity, thereby generating a dialogue flow path diagram. A dialogue hotspot distribution map generation unit is used to map instantiated user intent nodes within each cluster to hotspot areas based on intent clustering analysis results, and calculate intent concentration and dispersion by combining node weights, and generate a dialogue hotspot distribution map through heat mapping. The visualization fusion unit is used to project the dialogue semantic relationship graph, dialogue flow path graph and dialogue hotspot distribution graph onto a unified coordinate system, and perform weighted fusion and difference comparison to generate a comprehensive visualization analysis result.
10. A knowledge graph-based dialogue scenario data visualization and analysis system according to claim 7, characterized in that, The analysis module includes: The node importance analysis unit is used to calculate the node connection strength and importance score based on the node connection information in the dialogue semantic structure data and the node's own historical interaction characteristics, and generate the node importance analysis results. The dialogue path analysis unit is used to extract the semantic path between the user intent node and the topic node, calculate the path length and complexity, identify the main path and abnormal path, and generate dialogue path analysis results. The intent clustering analysis unit is used to calculate the semantic similarity of user intent nodes, perform clustering analysis, generate intent indicators, and output intent clustering analysis results. The incremental analysis unit is used to partially refresh only the affected nodes, paths, and clusters when new dialogue scenario data is input, and generate incremental analysis results to update the visualization display.
Citation Information
Patent Citations
Multi-round dialogue method and system integrating knowledge graph and emotion supervision
CN111651609A
Topic recommendation method and device, electronic equipment and storage medium
CN115292460A
Customer service data quality inspection method and device based on dynamic reasoning, equipment and medium
CN120216707A
Customer data processing and insight system based on large language model
CN120705704A
Eco-friendly acrylic coating composition for packaging material structure penetration waterproofing and surface density enhancement
KR102851692B1
Cited By
Agricultural information service system based on big data
CN121502823A
Big data based agricultural information service system
CN121502823B