A method and system for extracting the theme of user question text
By analyzing and know-gram mapping of the problem text submitted by users, combining the graph neural network and adaptive weight allocation mechanism, structured description information is generated, and the problems of low efficiency and poor accuracy of user problem text theme extraction in the existing technology are solved, achieving a more efficient and accurate theme extraction effect.
Patent Information
- Application Number
- CN202510205806.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The prior art is inefficient and poorly accurate in extracting text topics of user problem, especially when dealing with polysynonyms and complex contexts.
By analyzing the received user-submitted problem text, the core intentions and association concepts are identified, and mapping is performed on the pre-constructed knowledge graph. Combining the graph neural network and the adaptive weight allocation mechanism, we determine the scope of topic expansion and generate structured description information representing the corresponding topics of the problem text.
It improves the efficiency and accuracy of user problem text topic extraction, enables a deeper understanding of user intentions, handles complex contexts, and generates more coherent and logically consistent topic descriptions.
Smart Images

Figure CN119692343B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of topic extraction, and in particular, to a method and system for extracting topics from user question texts. Background Art
[0002] In today's era of information explosion, users submit question texts through various online platforms to obtain the required information or solutions. For example, in fields such as customer service, intelligent question-and-answer systems, and search engine optimization, it is crucial to quickly and accurately understand the core intention and related concepts of user questions.
[0003] Currently, common methods for extracting topics from user question texts mainly include rule-based methods, traditional machine learning methods, and deep learning methods. The specific practices include: rule-based methods rely on predefined syntactic structures and vocabularies; traditional machine learning methods use feature engineering and classification algorithms for topic classification; while deep learning methods such as convolutional neural networks and recurrent neural networks and their variants perform topic modeling by automatically learning text representations. In addition, some studies have attempted to combine knowledge graphs to enhance the ability to understand texts.
[0004] Although the existing solutions can achieve topic extraction from user question texts to a certain extent, they still have some limitations: rule-based methods and traditional machine learning methods highly rely on manually designed rules and feature engineering, and it is difficult to adapt to constantly changing language patterns and complex semantic relationships; although many existing deep learning methods can capture the surface features of texts, they are not deep enough in mining deep semantic understanding and implicit relationships, especially performing poorly when dealing with polysemous words and complex contexts; the existing solutions usually adopt fixed weight allocation strategies and cannot dynamically adjust the importance of nodes according to specific problems, resulting in unreasonable resource allocation in some cases, affecting the effect and efficiency of topic extraction. Summary of the Invention
[0005] The embodiments of the present application provide a method and system for extracting topics from user question texts to solve the problems of low efficiency and poor accuracy in topic extraction of user question texts in the prior art.
[0006] In a first aspect, the embodiments of the present application provide a method for extracting topics from user question texts, including:
[0007] Based on the received question text submitted by the user, parsing and processing the question text to identify the core intention and related concepts;
[0008] Perform mapping processing on a pre-constructed knowledge graph according to the core intention and associated concepts, and determine the topic expansion range based on the semantic relevance between nodes in the knowledge graph and the implicit relationships between the nodes learned through a graph neural network;
[0009] Based on the topic expansion range, adopt an adaptive weight allocation mechanism combined with a link analysis algorithm to evaluate the importance of nodes within the topic expansion range and the centrality of the nodes in the graph neural network, so as to adjust the weights of each node;
[0010] Generate structured description information representing the topic corresponding to the problem text according to the adjusted node weights and the connection strength between the nodes calculated through point mutual information;
[0011] Use a multi-layer attention mechanism to further refine the structured description information to obtain an extraction result, where the extraction result refers to the summary output result of the topic corresponding to the problem text.
[0012] Optionally, generating structured description information representing the topic corresponding to the problem text according to the adjusted node weights and the connection strength between the nodes calculated through point mutual information includes:
[0013] Based on the adjusted node weights, perform a priority sorting process on the nodes in the knowledge graph to obtain core topic nodes;
[0014] Use point mutual information to quantify the semantic association strength between the core topic nodes and determine the tightness between the core topic nodes;
[0015] According to the tightness between the core topic nodes, use a clustering algorithm or a community discovery algorithm to group the core topic nodes to form topic clusters;
[0016] Perform a comprehensive analysis process on the core topic nodes within the topic cluster to extract topic labels that can summarize the core meaning of the topic cluster;
[0017] Combine the explicit concepts in the problem text with the topic labels, supplement potential implicit topics through a logical reasoning mechanism, and generate topic content corresponding to the implicit topics;
[0018] Based on the topic content, organize the explicit concepts, the topic labels, and the implicit topics to generate structured description information representing the topic corresponding to the problem text.
[0019] Optionally, combine the explicit concepts in the problem text with the topic tags, supplement potential implicit topics through a logical reasoning mechanism, and generate topic content corresponding to the implicit topics, including:
[0020] Based on the explicit concepts in the problem text and the topic tags, construct a preliminary topic framework;
[0021] Utilize the relevance between nodes and context information in the knowledge graph to perform expansion processing on the topic framework, and identify additional entities or concepts related to the explicit concepts and the topic tags;
[0022] Based on rule-based reasoning or a machine learning model, analyze the relationships between the additional entities or concepts to mine implicit topics;
[0023] Based on the implicit topics, the explicit concepts, and the topic tags, perform consistency processing through semantic fusion technology, and integrate the explicit concepts, the topic tags, and the implicit topics to generate topic content corresponding to the implicit topics.
[0024] Optionally, based on the topic expansion scope, adopt an adaptive weight allocation mechanism combined with a link analysis algorithm to evaluate the importance of nodes within the topic expansion scope and the centrality of the nodes in the graph neural network, so as to adjust the weights of each node, including:
[0025] Based on the topic expansion scope, use the adaptive weight allocation mechanism to initialize the basic weights of the nodes within the topic expansion scope, and calculate the local influence of the basic weights of each node within the topic expansion scope according to the semantic relevance between nodes in the knowledge graph and the implicit relationships learned through the graph neural network, to obtain local influence results;
[0026] Adopt the link analysis algorithm to perform iterative calculation processing on the scores of the nodes to obtain importance scores;
[0027] Combine the local influence results and the importance scores to comprehensively evaluate the importance of the nodes, so as to dynamically adjust the weights of each node, enabling the nodes to obtain higher weights.
[0028] Optionally, combine the influence results, the importance scores, and the centrality of the nodes in the graph neural network to comprehensively evaluate the importance of the nodes, so as to dynamically adjust the weights of each node, enabling the nodes to obtain higher weights, including:
[0029] Using the local influence results, calculate and process the basic influence scores of each node to obtain the basic influence scores;
[0030] Based on the basic influence scores and the importance scores, and in combination with a preset functional relationship, calculate and process the comprehensive evaluation scores of the nodes to obtain the comprehensive evaluation scores reflecting the importance of the nodes within the scope of the topic expansion;
[0031] Using the comprehensive evaluation scores, sort all the nodes, and set a threshold or a ranking interval to determine the nodes with high importance;
[0032] Apply a gain function to perform weight enhancement processing on the nodes with high importance, so that the nodes with high importance obtain higher weights.
[0033] Optionally, according to the core intention and associated concepts, perform mapping processing on a pre-constructed knowledge graph, and determine the scope of topic expansion based on the semantic relevance between nodes in the knowledge graph and the implicit relationships between the nodes learned through a graph neural network, including:
[0034] Using a bidirectional encoder representation model and a semantic matching algorithm, based on the core intention and the associated concepts, perform mapping processing on the concepts in the problem text in the pre-constructed knowledge graph to obtain a topic mapping result;
[0035] According to the topic mapping result, in combination with the semantic relevance between nodes in the knowledge graph, perform exploration processing on multi-step paths starting from the mapped nodes, and identify potential nodes related to the core intention to construct a graph neural network;
[0036] Using the graph neural network combined with a self-attention mechanism, analyze and process the implicit relationships between the potential nodes, and introduce a context-aware model and a Bayesian optimization technique to evaluate the relevance and importance of the potential nodes to the problem text, and dynamically adjust the relevance and importance scores of the potential nodes to obtain an adjustment result;
[0037] According to the adjustment result, determine the nodes that can be included in the scope of the topic expansion from the potential nodes, and in combination with a hierarchical clustering algorithm, group the nodes within the scope of the topic expansion to determine the scope of the topic expansion.
[0038] Optionally, the further refining process of the structured description information using a multi-layer attention mechanism to obtain an extraction result includes:
[0039] Using the constructed multi-layer attention network model to encode the obtained structured description information to obtain the encoded structured description information;
[0040] Based on the structured description information after the encoding process, use the first-layer attention mechanism to perform preliminary filtering and concentration processing on the structured description information after the encoding process, identify and strengthen the elements directly related to the core intention, and weaken or ignore the unrelated structured description information after the encoding process to obtain a preliminary extraction result;
[0041] According to the preliminary extraction result, in each subsequent layer of the attention mechanism, based on the output of the previous layer of the attention mechanism, perform layer-by-layer in-depth analysis processing on the structured description information to obtain the screening result of the multi-layer attention mechanism;
[0042] Based on the screening result of the multi-layer attention mechanism, process the output of the last layer of the attention mechanism to obtain an extraction result.
[0043] In a second aspect, an embodiment of the present application provides a system for extracting the theme of a user question text, including:
[0044] An analysis module, based on the received question text submitted by the user, performs analysis processing on the question text to identify the core intention and associated concepts;
[0045] A mapping module, according to the core intention and associated concepts, performs mapping processing on a pre-constructed knowledge graph, and determines the theme expansion range based on the semantic relevance between nodes in the knowledge graph and the implicit relationship between the nodes learned through a graph neural network;
[0046] An adjustment module, based on the theme expansion range, adopts an adaptive weight allocation mechanism combined with a link analysis algorithm to evaluate the importance of nodes within the theme expansion range and the centrality of the nodes in the graph neural network, so as to adjust the weights of the nodes;
[0047] A generation module, according to the adjusted node weights and the connection strength between the nodes calculated through point mutual information, generates structured description information representing the theme corresponding to the question text;
[0048] A processing module, uses a multi-layer attention mechanism to further refine the structured description information to obtain an extraction result, where the extraction result refers to the summary output result of the theme corresponding to the question text.
[0049] In the embodiments of the present application, based on the received problem text submitted by the user, the problem text is parsed and processed to identify the core intention and related concepts; according to the core intention and related concepts, mapping processing is performed on a pre-constructed knowledge graph, and the topic expansion range is determined based on the semantic relevance between nodes in the knowledge graph and the implicit relationship between the nodes learned through a graph neural network; based on the topic expansion range, an adaptive weight allocation mechanism is combined with a link analysis algorithm to evaluate the importance of nodes within the topic expansion range and the centrality of the nodes in the graph neural network, so as to adjust the weights of each node; according to the adjusted node weights and the connection strength between the nodes calculated through point mutual information, structured description information representing the topic corresponding to the problem text is generated; the structured description information is further refined using a multi-layer attention mechanism to obtain an extraction result, where the extraction result refers to the summary output result of the topic corresponding to the problem text.
[0050] The technical solution of the present application has the following beneficial effects:
[0051] By parsing and processing the received problem text submitted by the user, the present application can effectively identify the core intention and related concepts in the problem text. This ensures a more accurate and comprehensive understanding of the user's question, providing a solid foundation for subsequent steps. Using the pre-constructed knowledge graph for mapping processing can not only capture the explicit semantic relevance but also learn the implicit relationship between nodes through a graph neural network, thereby determining a wider and more accurate topic expansion range. By combining the adaptive weight allocation mechanism and the link analysis algorithm, the importance and centrality of nodes within the topic expansion range are evaluated and the weights of each node are adjusted. By quantifying the importance of nodes, the model can focus on the most critical information points, optimize resource allocation, and improve information processing efficiency. By generating structured description information representing the topic corresponding to the problem text according to the adjusted node weights and the connection strength between the nodes calculated through point mutual information, it ensures that the scattered information is integrated into a coherent whole, facilitating further analysis and refinement, while retaining the core intention and related concepts of the original text.
[0052] Furthermore, by performing priority sorting on the nodes in the knowledge graph, the core topic nodes can be accurately identified, ensuring that the model focuses on the most important and relevant nodes, thereby improving the accuracy of topic extraction. By quantifying the semantic association strength between the core topic nodes through point mutual information, the implicit relationships between the nodes can be explored more deeply, enhancing the ability to understand complex contexts. By determining the tightness between the core topic nodes, it is ensured that the generated topic description information has higher coherence and logical consistency. By grouping the core topic nodes through clustering algorithms or community discovery algorithms, topic clusters are formed. Ensuring that the scattered information is integrated into multiple meaningful topic sets for subsequent analysis. By organizing explicit concepts, topic tags, and implicit topics, structured description information is generated, providing users with a comprehensive and detailed perspective to ensure that users quickly grasp the core points of the problem.
[0053] Furthermore, through the explicit concepts and topic tags in the question text, a preliminary topic framework is constructed, providing a clear direction and basis for subsequent expansion processing, ensuring the pertinence and effectiveness of the subsequent steps. By identifying additional entities or concepts, the model can capture more information hidden in the question text, further enhancing the ability to understand the user's intention. By analyzing the relationships between the additional entities or concepts, it is ensured that the mined implicit topics are highly relevant to the original question text, avoiding the introduction of irrelevant information and improving the quality of topic extraction. Through semantic fusion technology, the implicit topics are processed for consistency with the explicit concepts and topic tags to ensure the semantic consistency and logical coherence of the generated topic content.
[0054] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 It is a flowchart of a method for extracting the topic of a user question text provided by an embodiment of the present application;
[0057] Figure 2 It is a schematic structural diagram of a system for extracting the topic of a user question text provided by an embodiment of the present application;
[0058] Figure 3 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. Specific Embodiments
[0059] In order to enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application.
[0060] In some processes described in the specification and claims of this application and the above-mentioned accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit that "first" and "second" are of different types.
[0061] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.
[0062] Figure 1 The following is a flowchart of a method for extracting the theme of a user's question text provided for the embodiments of this application, as Figure 1 shown, the method includes:
[0063] 101. Based on the received question text submitted by the user, perform parsing processing on the question text to identify the core intention and associated concepts;
[0064] In this step, perform parsing processing on the question text submitted by the user to identify the core intention and associated concepts therein. The parsing processing includes natural language processing techniques such as word segmentation, part-of-speech tagging, named entity recognition, syntactic analysis, etc., which are used to decompose the question text into structured data. This data not only contains lexical information, but also covers syntactic structures and semantic features, which are the basis for understanding the user's intention. The core intention refers to the main purpose or requirement of the user's question, and the associated concepts refer to other important information points closely related to the core intention.
[0065] In the application embodiment, in the customer service system, when the user asks "How to query my order status", the system first parses the question text into "How", "query", "my", "order", "status". Through part-of-speech tagging and named entity recognition, "order" is marked as a key entity, and "query" is regarded as the action word expressing the user's intention. Syntactic analysis further confirms that "query" is the main action and "order status" is the target object. Finally, the system identifies that the user's core intention is "query order status", and the associated concepts include "order number", "logistics information", etc. This parsing result provides a clear direction for the subsequent steps, ensuring that the system can accurately map to the relevant nodes in the knowledge graph.
[0066] 102. Based on the core intention and associated concepts, perform mapping processing on the pre-constructed knowledge graph, and determine the theme expansion range according to the semantic relevance between nodes in the knowledge graph and the implicit relationships between the nodes learned through the graph neural network;
[0067] In this step, perform mapping processing on the pre-constructed knowledge graph, and determine the theme expansion range according to the semantic relevance between nodes in the knowledge graph and the implicit relationships between the nodes learned through the graph neural network. The knowledge graph is a semantic network composed of entity nodes and the relationship edges between them, which can represent complex knowledge structures. Mapping processing refers to mapping the parsed text elements to the corresponding nodes in the knowledge graph, thereby establishing the connection between the text and the knowledge base. Theme expansion is based on these mapping results and uses the deep-level associations between nodes captured by the graph neural network to expand the scope of understanding of the original question.
[0068] In the application embodiment, the customer service system maps "order status" to the "order management" field in the knowledge graph and finds that the related nodes include "order details", "payment status", "delivery progress", etc. Through the learning of the graph neural network, the system identifies that "logistics information" and "return policy" are also important associated concepts. Therefore, the theme expansion range is not limited to the query of "order status", but also includes other factors that may affect order processing, such as logistics delay, return process, etc. This expansion ensures that the system can provide comprehensive and targeted answers to meet the diverse needs of users.
[0069] 103. Based on the theme expansion range, adopt an adaptive weight allocation mechanism combined with a link analysis algorithm to evaluate the importance of the nodes within the theme expansion range and the centrality of the nodes in the graph neural network, so as to adjust the weights of each node;
[0070] In this step, an adaptive weight assignment mechanism is adopted in combination with a link analysis algorithm to evaluate the importance and centrality of nodes within the scope of topic expansion, and to adjust the weights of each node. The adaptive weight assignment dynamically adjusts the weights of nodes according to their positions and connection situations in the knowledge graph, ensuring the effective allocation of resources. The link analysis algorithm is used to quantify the connection strength between nodes and evaluate the relative importance of nodes in the network. This process aims to highlight those nodes that are most crucial for understanding and solving the user's problem, while optimizing the overall processing efficiency.
[0071] In the application embodiment, in the customer service case, the system calculates that the "order status" node has a high centrality because it is closely connected to multiple other important nodes (such as "order details", "payment status"). Through link analysis, it is found that there is a strong connection between "logistics information" and "order status", indicating that these two concepts are closely related. Therefore, the system assigns higher weights to "order status" and "logistics information" to ensure that they play a more important role in subsequent topic refinement. At the same time, the system also appropriately reduces the weight of the "return policy" because although it is also a related concept, it is not the primary concern in the current query. This step ensures the effective allocation of resources and improves the processing efficiency.
[0072] 104. Generate structured description information representing the corresponding topic of the problem text based on the adjusted weights of the nodes and the connection strength between the nodes calculated through point mutual information;
[0073] In this step, based on the adjusted weights of the nodes and the connection strength between the nodes calculated through point mutual information, structured description information representing the corresponding topic of the problem text is generated. The structured description information is a highly generalized summary of the nodes and their mutual relationships within the scope of topic expansion, aiming to provide a clear and coherent topic framework. Point mutual information is used to measure the probability of co-occurrence of two nodes, thereby quantifying the semantic association strength between them, ensuring that the generated information is both concise and complete.
[0074] In the application embodiment, the customer service system selects "order status" and "logistics information" as core nodes and constructs structured description information based on their high connection strength. This information not only outlines the order status that the user is concerned about but also details the logistics progress, providing a complete solution. In addition, the system supplements the "return policy" as additional information because although its weight is slightly lower, it may also be the focus of the user's attention in some cases, enhancing the comprehensiveness and practicality of the answer. This structured description information provides a solid foundation for subsequent multi-layer attention mechanism refinement, ensuring the integrity and logic of the information.
[0075] 105. The structured description information is further refined by using a multi-layer attention mechanism to obtain an extraction result, where the extraction result refers to a summary output result of the topic corresponding to the question text.
[0076] In this step, the structured description information is further refined using a multi-layer attention mechanism to obtain the final summary output result. The multi-layer attention mechanism allows the model to understand information from different levels and angles, focusing on the most critical parts, thereby generating a highly summarized topic summary that closely revolves around the core intent. This mechanism improves the model's adaptability to different types of input data, ensuring that the output results are both accurate and easy to understand.
[0077] In the application embodiment, in the customer service scenario, the system uses a multi-layer attention mechanism to refine the previously generated structured description information. The first layer of attention emphasizes the two core concepts of "order status" and "logistics information", and the second layer delves into the specific connections between them, such as "reasons for logistics delays" and "estimated delivery time". In the end, the system generates a concise and clear summary: "Your order has been shipped and is currently in transit. It is expected to arrive within the next three days. If you have any questions, please refer to our return policy." This summary not only answers the user's questions, but also provides additional useful information, greatly improving the user experience. Through the continuous implementation of five steps, the system realizes the automation of the entire process from question parsing to final answer generation, ensuring efficient and accurate service response.
[0078] In summary, through the above steps 101-105, the present invention can achieve efficient and accurate topic extraction and content generation in the fields of customer service, and provide users with high-quality help and support. Each step not only works independently, but also closely connects to form a complete solution chain, ensuring seamless connection from problem understanding to answer provision.
[0079] Optionally, in step 104, structured description information representing the topic corresponding to the question text is generated according to the adjusted node weights and the connection strength between the nodes calculated by point mutual information, including:
[0080] Based on the adjusted node weights, the nodes in the knowledge graph are prioritized to obtain core topic nodes; the semantic association strength between the core topic nodes is quantified using point mutual information to determine the closeness between the core topic nodes; according to the closeness between the core topic nodes, a clustering algorithm or a community discovery algorithm is used to group the core topic nodes to form a topic cluster; a comprehensive analysis is performed on the core topic nodes in the topic cluster to extract topic labels that can summarize the core meaning of the topic cluster; the explicit concepts in the question text and the topic labels are combined to supplement potential implicit topics through a logical reasoning mechanism, and the topic content corresponding to the implicit topic is generated; based on the topic content, the explicit concepts, the topic labels and the implicit topics are organized to generate structured description information representing the topic corresponding to the question text.
[0081] Optionally, combining the explicit concepts in the question text with the topic labels, supplementing potential implicit topics through a logical reasoning mechanism, and generating topic content corresponding to the implicit topics, including:
[0082] Based on the explicit concepts and the topic labels in the question text, a preliminary topic framework is constructed; the topic framework is expanded by utilizing the correlation between nodes in the knowledge graph and contextual information, and additional entities or concepts related to the explicit concepts and the topic labels are identified; based on rule-based reasoning or machine learning models, the relationships between the additional entities or concepts are analyzed and processed to mine implicit topics; based on the implicit topics, the explicit concepts and the topic labels, consistency processing is performed through semantic fusion technology, and the explicit concepts, the topic labels and the implicit topics are integrated to generate topic content corresponding to the implicit topics.
[0083] In this step, in the knowledge graph, each node represents an entity or concept. Through the adjusted node weights, the system can identify which nodes are the most critical for the current problem. The priority sorting process sorts the nodes according to these weights, thereby screening out the most core topic nodes. Ensuring that subsequent analysis focuses on the most important information improves the processing efficiency and accuracy. Point mutual information is a statistical method for measuring the co-occurrence probability of two nodes, used to quantify the semantic association strength between them. Clustering algorithms or community discovery algorithms are used to group the core topic nodes with similar characteristics to form multiple topic clusters. Each cluster represents a set of highly related concepts, which helps to understand the topic of the problem text from different perspectives. This grouping method not only simplifies the complex information structure but also facilitates subsequent comprehensive analysis and refinement. The system integrates explicit concepts, topic labels, and implicit topics to generate a comprehensive and coherent structured description information, ensuring that the output result covers both the core content of the original problem text and the extended background information, providing a more complete and in-depth answer.
[0084] In the embodiments of the present application, based on the adjusted node weights, the system performs priority sorting on the nodes in the knowledge graph to screen out the most important batch of nodes as the core topic nodes. Using point mutual information, the system evaluates the semantic association strength between the core topic nodes to determine their tightness. By using clustering algorithms or community discovery algorithms, the system groups the core topic nodes to form multiple topic clusters, and each cluster represents a set of highly related concepts. Comprehensive analysis is performed on the core topic nodes within each topic cluster to extract topic labels that can summarize the core meaning of the cluster. Combining the explicit concepts and topic labels in the problem text, potential implicit topics are mined through a logical reasoning mechanism, and corresponding topic contents are generated. The explicit concepts, topic labels, and implicit topics are integrated together to generate a comprehensive and coherent structured description information and provided to the user.
[0085] The system first identifies "order status", "logistics information", "shipment", etc. as core topic nodes according to the adjusted node weights. Through point mutual information calculation, a very strong semantic association is shown between "order status" and "logistics information", indicating that these two concepts are closely related. The system uses a clustering algorithm to group nodes such as "order status", "logistics information", "shipment", etc. into the same topic cluster, named "order transportation". At the same time, another topic cluster "after-sales service" is identified, including nodes such as "return policy", "customer service contact information", etc. Through comprehensive analysis of the "order transportation" cluster, the topic label "logistics delay" is refined. This label accurately summarizes the main issues that users are concerned about, that is, the reasons for the failure to update logistics information in a timely manner. The system generates a structured description message: "Your order has been shipped, but the logistics information has not been updated yet. It may be due to improper selection of the logistics company or the impact of holidays, resulting in logistics delay. It is recommended that you contact the express company to query the specific reasons, or check the latest logistics progress on our website. If you have further questions, please refer to our after-sales service policy."
[0086] Optionally, in step 103, based on the topic expansion range, an adaptive weight allocation mechanism is adopted in combination with a link analysis algorithm to evaluate the importance of the nodes within the topic expansion range and the centrality of the nodes in the graph neural network, so as to adjust the weights of each node, including:
[0087] Based on the topic expansion range, using the adaptive weight allocation mechanism, the basic weights of the nodes within the topic expansion range are initialized, and according to the semantic relevance between the nodes in the knowledge graph and the implicit relationships between the nodes learned through the graph neural network, the local influence of the basic weight of each node within the topic expansion range is calculated to obtain the local influence result; the link analysis algorithm is used to perform iterative calculation processing on the scores of the nodes to obtain the importance scores; combining the local influence result and the importance scores, as well as the centrality of the nodes in the graph neural network, the importance of the nodes is comprehensively evaluated to dynamically adjust the weights of each node, so that the nodes obtain higher weights.
[0088] Optionally, combining the influence result and the importance score, comprehensively evaluate the importance of the nodes to dynamically adjust the weights of the nodes, so that the nodes obtain higher weights, including: using the local influence result, calculating the basic influence score of each node to obtain the basic influence score; based on the basic influence score and the importance score, combining a preset functional relationship, calculating the comprehensive evaluation score of the node to obtain a comprehensive evaluation score reflecting the importance of the node within the scope of the topic expansion; using the comprehensive evaluation score, sorting all the nodes, and setting a threshold or a ranking interval to determine the nodes with high importance; applying a gain function to enhance the weights of the nodes with high importance, so that the nodes with high importance obtain higher weights.
[0089] In this step, the adaptive weight allocation mechanism is a method for dynamically adjusting the node weights, aiming to optimize resource allocation according to the position and connection of nodes in the knowledge graph. The local influence refers to the degree of influence of a node within a specific topic expansion range. By combining the semantic relevance between nodes in the knowledge graph and the implicit relationships learned by the graph neural network, the system can quantify the local influence of each node within this range, making the evaluation more comprehensive and accurate. The comprehensive evaluation is a process of comprehensively evaluating the importance of nodes by combining the local influence result and the importance score. Through a preset functional relationship, the system calculates the comprehensive evaluation score of each node, and then determines its actual importance within the scope of the topic expansion. The dynamic weight adjustment is to increase the weights of the nodes with high importance according to the comprehensive evaluation results, ensuring that they play a more important role in subsequent processing. The gain function is used to further increase the weights of the nodes with high importance, so that they receive more attention in the overall structure. The design of the gain function can be adjusted according to specific application scenarios to achieve the best resource allocation effect.
[0090] In the embodiment of the present application, the system first assigns a basic weight to each node, which is based on the basic attributes of the node and serves as the basis for subsequent calculations. Through the semantic relevance between nodes in the knowledge graph and the implicit relationships learned by the graph neural network, the system calculates the local influence of each node within the scope of the topic expansion to ensure the comprehensiveness of the evaluation. Combining the local influence result and the importance score, the system comprehensively evaluates the importance of the nodes. Through a preset functional relationship, the comprehensive evaluation score of each node is calculated to reflect its actual importance within the scope of the topic expansion. The system applies a gain function to enhance the weights of the nodes with high importance. In this way, it is ensured that these nodes play a more important role in the subsequent topic refinement and content generation processes, improving the response efficiency and accuracy of the model.
[0091] Suppose the user asks: "The status of my order shows that it has been shipped, but the logistics information has not been updated for a long time. What should I do?" First, the system assigns a basic weight to each node in the knowledge graph (such as "order status", "logistics information", "shipment", "logistics company", "customer service contact information", etc.). Using the semantic relevance between nodes in the knowledge graph and the implicit relationships learned by the graph neural network, the system calculates the local influence of each node within the scope of the "order status not updated in a timely manner" problem. Combining the local influence results and importance scores, the system comprehensively evaluates the importance of each node. Through a preset functional relationship, the comprehensive evaluation score of each node is calculated. The results show that "logistics information" and "logistics company" perform outstandingly in both dimensions, so they obtain relatively high comprehensive evaluation scores. The system applies a gain function to enhance the weights of "logistics information" and "logistics company". In this way, in the subsequent topic extraction and content generation processes, these two nodes will receive more attention, ensuring that the system can provide more accurate service suggestions.
[0092] Optionally, in step 102, mapping processing is performed on the pre-constructed knowledge graph according to the core intention and associated concepts, and the topic expansion scope is determined based on the semantic relevance between nodes in the knowledge graph and the implicit relationships between the nodes learned through the graph neural network, including:
[0093] Using a bidirectional encoder representation model and a semantic matching algorithm, map the concepts in the question text in the pre-constructed knowledge graph to obtain a topic mapping result; according to the topic mapping result, explore multi-step paths starting from the mapped nodes to identify potential nodes related to the core intention to construct a graph neural network; use the graph neural network combined with a self-attention mechanism to analyze the implicit relationships between the potential nodes, and introduce a context-aware model and Bayesian optimization technology to evaluate the relevance and importance of the potential nodes to the question text, and dynamically adjust the relevance and importance scores of the potential nodes to obtain an adjustment result; according to the adjustment result, and combined with a hierarchical clustering algorithm, group the nodes within the topic expansion scope to determine the topic expansion scope.
[0094] In this step, the semantic matching algorithm is used to compare the concepts in the question text with the nodes in the knowledge graph, find the most similar or relevant mapping results, and accurately map the concepts in the question text to the corresponding nodes in the knowledge graph. The topic mapping result refers to the corresponding nodes in the knowledge graph of the concepts in the question text obtained through the semantic matching algorithm. Ensure that the core intention and related concepts of the question text can be accurately mapped to the relevant parts of the knowledge graph. The graph neural network is a neural network model specifically used to process graph-structured data and can capture the complex relationships between nodes. The context-aware model is used to consider the specific background information of the question text to improve the accuracy of mapping and analysis. The Bayesian optimization technique is used to evaluate the relevance and importance of potential nodes to the question text and dynamically adjust the scores to ensure that the finally selected nodes best meet the user's needs.
[0095] In the embodiments of the present application, through the semantic matching algorithm, the system maps the concepts in the question text to the corresponding nodes in the pre-constructed knowledge graph to obtain the topic mapping result, ensuring that the core intention and related concepts of the question text can be accurately mapped to the relevant parts of the knowledge graph. According to the topic mapping result, the system starts from the mapped nodes and conducts multi-step path exploration in the knowledge graph to identify potential nodes related to the core intention. By constructing a graph neural network and combining with the self-attention mechanism, the system deeply analyzes the implicit relationships between potential nodes. The self-attention mechanism allows the model to dynamically adjust its weights according to the importance of the nodes, thereby better understanding the relationships between nodes. By introducing the context-aware model, the system considers the specific background information of the question text to improve the accuracy of mapping and analysis. At the same time, using the Bayesian optimization technique, the system evaluates the relevance and importance of potential nodes to the question text. According to the adjusted relevance and importance scores, the system determines the nodes that can be included in the topic expansion range from the potential nodes. Then, combined with the hierarchical clustering algorithm, these nodes are grouped to form multiple topic clusters, and finally the topic expansion range is determined.
[0096] The customer service system first uses a semantic matching algorithm to map concepts such as "order status" and "logistics information" in the question text to corresponding nodes in a pre-constructed knowledge graph. For example, "order status" is mapped to relevant nodes in the "order management" domain, and "logistics information" is mapped to nodes in the "logistics service" domain. By constructing a graph neural network and combining with the self-attention mechanism, the system deeply analyzes the implicit relationships between the above potential nodes. The self-attention mechanism allows the model to dynamically adjust its weights according to the importance of the nodes, so as to better understand the relationships between the nodes. Then, combined with the hierarchical clustering algorithm, these nodes are grouped to form multiple topic clusters, and finally the topic expansion scope is determined. For example, "logistics company" and "transportation progress" are grouped into the same cluster, named "reasons for logistics delay", while "customer service contact information" is separately grouped into another cluster, named "help channels".
[0097] Optionally, the further refining process of the structured description information by using the multi-layer attention mechanism in step 105 to obtain the extraction result includes:
[0098] Using the constructed multi-layer attention network model to encode the obtained structured description information to obtain the encoded structured description information; based on the encoded structured description information, using the first-layer attention mechanism to perform preliminary filtering and concentration processing on the encoded structured description information, identifying and strengthening elements directly related to the core intention, weakening or ignoring the irrelevant encoded structured description information to obtain a preliminary extraction result; according to the preliminary extraction result, in each subsequent layer of the attention mechanism, based on the output of the previous layer of the attention mechanism, perform layer-by-layer in-depth analysis processing on the structured description information to obtain the screening result of the multi-layer attention mechanism; based on the screening result of the multi-layer attention mechanism, process the output of the last layer of the attention mechanism to obtain the extraction result.
[0099] In this step, the multi-layer attention network model is a deep learning architecture that can understand information from different levels and perspectives. Each layer of the attention mechanism focuses on a specific information level, ensuring that the final output result is both accurate and easy to understand. The preliminary filtering and concentration processing is to use the first-layer attention mechanism to perform the initial screening on the encoded structured description information. In this way, the system can quickly lock in the key information points and improve the processing efficiency. The layer-by-layer in-depth analysis processing means that in each subsequent layer of the attention mechanism, based on the output of the previous layer, further refine and deepen the understanding of the information, ensuring the comprehensiveness and accuracy of the information processing. The extraction result refers to the topic summary or key information finally obtained after being processed by the multi-layer attention mechanism. These results closely revolve around the user's core intention, provide concise and clear answers, and may contain additional relevant information to enhance the user experience.
[0100] In the embodiments of the present application, based on the structured description information, the system uses the first-layer attention mechanism for preliminary filtering and concentration processing. The preliminary extraction result centrally reflects the most critical part of the problem text, improving the efficiency of information processing. According to the preliminary extraction result, in each subsequent layer of the attention mechanism, the system performs a layer-by-layer in-depth analysis and processing of the structured description information based on the output of the previous layer. Each layer of the attention mechanism focuses on different information levels, gradually uncovering deeper relationships and implicit themes. In a progressive manner layer by layer, the system can capture more complex information structures, ensuring the comprehensiveness and accuracy of the final output. Based on the screening results of the multi-layer attention mechanism, the system processes the output of the final-layer attention mechanism to obtain the extraction result. These results closely revolve around the user's core intention, providing a concise and clear answer and possibly including additional relevant information to enhance the user experience.
[0101] The customer service system, based on the structured description information, uses the first-layer attention mechanism for preliminary filtering and concentration processing. For example, it identifies and strengthens the elements directly related to "order status" and "logistics information", and weakens or ignores irrelevant information such as "return policy". According to the preliminary extraction result, in each subsequent layer of the attention mechanism, the system performs a layer-by-layer in-depth analysis and processing of the structured description information based on the output of the previous layer. For example, the second-layer attention mechanism may find that "improper selection of logistics company" or "holidays impact" are important reasons for logistics delay. Based on the screening results of the multi-layer attention mechanism, the system processes the output of the final-layer attention mechanism to obtain the extraction result. Finally, a concise and clear summary is generated: "Your order has been shipped, but the logistics information has not been updated yet, which may be due to problems with the logistics company resulting in a delay. It is recommended that you contact the courier company to inquire about the specific reason, or check the latest logistics progress on our website. If you have further questions, please refer to our after-sales service policy or contact customer service for assistance."
[0102] In the process of extracting the theme of the user's problem text, it is crucial to determine the importance and weight of the nodes within the theme expansion range. Traditional weight assignment methods usually rely on static features such as frequency or predefined rules, but this method is difficult to adapt to complex and changing contexts and dynamic information requirements. Therefore, an adaptive weight assignment mechanism combined with a link analysis algorithm is introduced, aiming to dynamically adjust the node weights according to the centrality of the nodes in the graph neural network and other relevant factors.
[0103] Optionally, based on the theme expansion range, an adaptive weight assignment mechanism combined with a link analysis algorithm is adopted to evaluate the importance of the nodes within the theme expansion range and the centrality of the nodes in the graph neural network, so as to adjust the weights of each node, including:
[0104] Based on the subject expansion range and the obtained local influence results, the link analysis algorithm is adopted to iteratively calculate and process the scores of the nodes, and the importance scores of the nodes within the subject expansion range are obtained; among them, the importance Evaluation calculation formula:
[0105] ;
[0106] Among them, represents the centrality of the node within the subject expansion range, represents the link analysis score of the node within the subject expansion range, is a balance factor between 0 and 1, and are the exponential factors applied to the centrality and the link analysis score respectively, represents the variability score of the node , reflecting the stability of the node, is the coefficient used to adjust the variability score, represents the activity score of the node , is the coefficient used to adjust the activity score, is the importance;
[0107] represents that the centrality of the node in the graph neural network is calculated using the PageRank algorithm, and the centrality calculation formula:
[0108] ;
[0109] Among them, is the damping coefficient, is the total number of the nodes, is the set of all nodes pointing to the node , is the node , is the node 's out-link number, is an additional exponential factor, is a distance attenuation function used to judge the shortest path distance between the node and the node , , is to control the attenuation rate, represents the node Centrality;
[0110] Combining the local influence results, the importance scores, and the centrality of the nodes in the graph neural network, comprehensively evaluate the importance of the nodes to dynamically adjust the weights of each node and adjust the weights of the nodes of the weights ;
[0111] Weight calculation formula:
[0112] ;
[0113] Wherein, is the original weight of the node ; is a mixing ratio factor between 0 and 1, used to control the mixing ratio between the calculated importance and the original weight, is the exponential factor for the importance, is the importance, is the node of the weight.
[0114] The overall formula introduces node importance evaluation calculation and weight calculation, aiming to construct an efficient, flexible and intelligent node importance evaluation and weight adjustment framework to improve the system's understanding depth of user problems and service quality. The introduction of centrality calculation reflects the relative importance of nodes in the network. The overall formula design fully considers multiple aspects such as multi-dimensional evaluation, dynamic adaptability, balance adjustment, enhanced discrimination, smooth transition, context sensitivity, and application universality, enabling the system to better serve the accurate understanding and solution of user problems.
[0115] The following briefly explains the design reasons for each item in the formula:
[0116] Importance Design reason for the evaluation calculation formula: Importance The evaluation calculation formula can balance multiple factors, providing a comprehensive and dynamic evaluation method to ensure that the system can accurately identify and highlight the most important nodes and provide more accurate service recommendations for users.
[0117] Formula expression: Used to comprehensively evaluate the structural characteristics and stability of nodes in the network. Among them, is the centrality of the node , and its influence is adjusted by the exponential factor , reflects the variability score of the node, represents the node The variability score, which reflects the stability of the node, and its influence is controlled by an adjustment coefficient controls its influence; is used to comprehensively evaluate the connection strength and activity level of a node in the network, reflects the link analysis score of the node and adjusts its influence strength through an exponential factor ; reflects the activity score of the node and controls its influence through an adjustment coefficient ; Through a balance factor ensures an appropriate balance between the link analysis score and the activity score.
[0118] Centrality Design rationale for the calculation formula: Centrality reflects the relative importance of a node in the network. Nodes with high centrality often play a key role in the network and are important hubs for information dissemination.
[0119] Formula expression: Centrality In the calculation formula calculates the base probability of each node, that is, the initial weight assigned to each node at the beginning of each iteration. The damping coefficient is a parameter between 0 and 1, usually set to about 0.85. It represents the probability that a user continues to click on a link when browsing a web page or moving in the network. The total number of nodes represents the total number of nodes in the network; is mainly used to calculate the weight transfer obtained by the node from other nodes pointing to it, and accurately calculates the centrality score of the node ; is the set of all nodes pointing to the node ; is the out-degree of the node ; is an additional exponential factor, is a distance decay function used to judge the shortest path distance between the node and the node ; is to control the decay rate
[0120] Design rationale of the weight calculation formula: The weight formula plays a crucial role in adjusting the node weights. By combining the newly evaluated importance and the original weights, it achieves smooth and dynamic adjustment of the node weights. This not only improves the flexibility and adaptability of the system but also ensures the stability and continuity of the results, enabling the system to better serve complex and changing application scenarios.
[0121] Formula expression: Used to calculate the new weight contribution according to the importance of the node is a mixing ratio factor between 0 and 1, used to control the mixing ratio between the calculated importance and the original weight is the exponential factor for the importance is the importance; is used to retain the original weight of the node to ensure the stability and continuity of the system is the weight of the node is the node The following briefly explains how to obtain each parameter of the formula:
[0122] These parameters can be tuned through techniques such as experiments and Bayesian optimization to find the optimal values most suitable for specific application scenarios;
[0123] These scores can be automatically calculated from historical data, expert knowledge, or machine learning models
[0124] calculated using the PageRank algorithm calculated through a link analysis algorithm calculated according to the change frequency of node attributes calculated according to the activity level of the node over a period of time in the past. In the embodiments of this application, assume we have a customer service scenario where the user asks: "My order status shows shipped, but the logistics information has not been updated for a long time. What should I do?"
[0125] Initial settings:
[0126] The nodes involved in the knowledge graph include "order status", "logistics information", "logistics company", "customer service contact information"
[0127] The initial weight of each node
[0128] is set to 0.5
[0129] Parameter settings:
[0130]
[0130] Calculation process:
[0131] Calculate centrality :
[0132] Use the PageRank algorithm to calculate the centrality of each node. For example, the centrality of "Order Status" may be relatively high because it is a core concept;
[0133] Evaluate importance :
[0134] For the "Order Status" node (assumed to be node 1), substitute it into the formula to calculate the importance:
[0135] ;
[0136] Assume , then:
[0137] ;
[0138] Adjust weights :
[0139] For the "Order Status" node, substitute it into the formula to calculate the new weight:
[0140] ;
[0141] The system determines that the importance score of the "Order Status" node is approximately 0.916 and adjusts its weight to approximately 0.695, so that the "Order Status" node will receive more attention in the subsequent topic refinement and content generation processes, thus ensuring that the system can provide more accurate service recommendations.
[0142] In the process of extracting the theme of the user's question text, generating structured description information needs to comprehensively consider the node weights, the connection strength between nodes, and the dynamic characteristics of nodes. Traditional methods are usually based on static features, but this method is difficult to adapt to complex and changing contexts and dynamic information requirements. Therefore, point mutual information is introduced to calculate the connection strength between nodes, and combined with the adjusted node weights and other dynamic features to more accurately evaluate the importance of nodes, and finally generate structured description information representing the corresponding theme of the question text.
[0143] Optionally, generating structured description information representing the corresponding theme of the question text according to the adjusted node weights and the connection strength between nodes calculated by point mutual information further includes:
[0144] Generate the connection strength between nodes according to the adjusted node weights and point mutual information;
[0145] Connection strength Calculation formula:
[0146] ;
[0147] wherein, is the probability of simultaneous occurrence of the said node and the said node , and are respectively the probabilities of separate occurrence of the said node and the said node , is an enhancement coefficient for strengthening the influence of high co-occurrence frequency on the said connection strength, is an exponential factor for adjusting the influence degree of the said high co-occurrence frequency, is the connection strength, is the weight of the said node ;
[0148] According to the node weight, the connection strength between nodes, the variability and activity of the nodes, use a comprehensive score to evaluate the importance of the said node in the said structured description information,
[0149] Comprehensive score calculation formula:
[0150] ;
[0151] wherein, represents the set of nodes directly connected to the said node , is the exponential factor applied to the said point mutual information, is the exponential factor applied to the said connection strength, is a distance attenuation function, the shortest path distance between the said node and the said node , controls the attenuation rate, is a similarity enhancement function, based on the connection strength between the said node and the said node , wherein is the exponent controlling the influence degree of the said semantic similarity, is the coefficient for balancing the connection strength of the nodes, is the variability score of the said node , is the exponential factor of the said variability score, is the said node The activity score, is the exponential factor of the activity score, is the node 's popularity score, is the exponential factor of the popularity score, and are coefficients used to adjust the influence of the variability score, the activity score, and the popularity score respectively, is the comprehensive score;
[0152] Based on the comprehensive score , select the node with the highest score and its connection relationship to form the structured description information representing the corresponding theme of the problem text;
[0153] Calculation formula:
[0154] ;
[0155] Among them, is a preset connection strength threshold used to filter out weak connection strength, is the connection strength between the node and the node .
[0156] By combining the connection strength and the comprehensive score The overall design aims to provide a comprehensive and dynamic method to evaluate the importance of nodes. By combining the static structural features of nodes (such as weights, connection strengths) and dynamic features (such as variability, activity, popularity), the formula can balance multiple factors to ensure the comprehensiveness and accuracy of the evaluation results. In addition, by introducing parameters such as exponential factors, adjustment coefficients, and distance decay functions, the system is ensured to have high flexibility and adaptability, suitable for complex application scenarios in different fields.
[0157] The following briefly explains the design reasons for each item in the formula:
[0158] Connection strength Design reason: The connection strength measures the co-occurrence relationship between nodes through point mutual information, reflects the association strength between nodes, ensures that nodes with high co-occurrence frequencies contribute more to the connection strength, and highlights strong association relationships.
[0159] Formula expression: represents the ratio of the probability of the simultaneous occurrence of node and node to the product of their individual occurrence probabilities, measures the co-occurrence relationship between nodes through point mutual information, and reflects the association strength between nodes; To strengthen the influence of high co-occurrence frequency on connection strength by introducing an enhancement coefficient and an exponential factor , the formula can highlight those nodes with high co-occurrence frequency and strong association while maintaining a reasonable numerical range, making the evaluation results more comprehensive and accurate. is the probability of the simultaneous occurrence of the node and the node . and are the probabilities of the separate occurrences of the node and the node respectively. is an enhancement coefficient used to strengthen the influence of high co-occurrence frequency on the connection strength. is an exponential factor used to adjust the degree of influence of the high co-occurrence frequency. is the connection strength;
[0160] Design rationale of the comprehensive score calculation formula: The comprehensive score calculation formula combines node weights, connection strength, and dynamic features to comprehensively evaluate the importance of nodes, achieving an appropriate balance between the static and dynamic features of nodes, ensuring a smooth transition between old and new information, and avoiding drastic fluctuations in results due to frequent updates.
[0161] Formula expression: Used to evaluate the connection strength of node and its contribution to the comprehensive score. By combining node weights, connection strength, and dynamic features (such as distance attenuation, semantic similarity), it comprehensively evaluates the importance of node . By introducing a balance factor , an exponential factor and and other parameters, the formula can achieve a balance among multiple factors, ensuring the comprehensiveness and accuracy of the evaluation results; Used to accumulate the connection strengths of all nodes directly connected to node and consider the influence of distance attenuation and semantic similarity;
[0162] By combining the dynamic features of nodes, it comprehensively evaluates the importance of nodes. By introducing a balance factor and a regulation coefficient , , the formula can achieve a balance among multiple factors, ensuring the comprehensiveness and accuracy of the evaluation results. Among them, represents the set of nodes directly connected to the node . is the exponential factor applied to the point mutual information, is the exponential factor applied to the connection strength, is a distance attenuation function, where the shortest path distance between the node and the node , controls the attenuation rate, is a similarity enhancement function based on the connection strength between the node and the node , where is the exponent controlling the influence degree of the semantic similarity, is the coefficient balancing the connection strength of the nodes, is the variability score of the node , is the exponential factor of the variability score, is the activity score of the node , is the exponential factor of the activity score, is the popularity score of the node , is the exponential factor of the popularity score, and are the coefficients used to adjust the influence of the variability score, the activity score, and the popularity score respectively, is the comprehensive score;
[0163] Design reason for the structured description information: Through the calculation design of the structured description information, the system can effectively remove the weak connections that are irrelevant to the problem, thus generating a more concise and useful structured description information.
[0164] Formula expression: If , then the connection is retained, where is a preset connection strength threshold. By setting , the weak connections that have little impact on the overall structure can be effectively removed, thus focusing on the important node relationships, which helps to improve the efficiency of the system and the accuracy of the results. is the representative of the connection strength of the structured description information of the node and the node . The system can effectively remove the weak connections that are irrelevant to the problem, thus generating a more concise and useful structured description information.
[0165] The following is a brief explanation of the acquisition methods of each parameter in the formula:
[0166] Node weight : Calculated by combining link analysis algorithms through an adaptive weight allocation mechanism,
[0167] Connection strength : Calculated by point mutual information,
[0168] Variability score , Activity score and Popularity score : Derived from historical data analysis,
[0169] Exponential factor and Adjustment coefficient , , : Can be tuned through techniques such as experiments and Bayesian optimization,
[0170] Enhancement coefficient and Exponential factor : Used to adjust the impact of high co-occurrence frequencies and ensure the rationality of evaluation results,
[0171] Distance decay parameter and Similarity enhancement index : Used to adjust the impact of distance decay and semantic similarity and ensure the rationality of evaluation results.
[0172] In the embodiments of the present application, a logistics network composed of multiple logistics centers is considered. Each logistics center has its own weight , reflecting its importance or capacity; and the connection strength with other logistics centers , calculated based on historical data. Our goal is to evaluate the importance of each logistics center based on this information and select the most effective connections to form an optimized logistics network.
[0173] Parameter settings and assumptions:
[0174] Set the weights of logistics centers 1 to 5 to be , and 0.5.
[0175] Set the variability scores to be respectively , and 0.3.
[0176] Activity scores to be respectively , and 0.4.
[0177] Popularity scores to be respectively , and 0.6.
[0178] Connection strength The specific value will be calculated based on the actual co-occurrence probability and other factors. Here, we simplify and directly give some example values. For example, the connection strength between logistics center 1 and 2 is 0.8, and between 1 and 3 is 0.6, and so on. Set a connection strength threshold , that is, only retain the connections with a connection strength greater than 0.5; , and the other parameters are adjusted according to the actual situation. For the sake of simplicity in calculation, we will use default values or empirical estimates.
[0179] Calculate the comprehensive score :
[0180] For logistics center 1, we first calculate the sum of the weighted connection strengths between it and all other logistics centers, and then calculate according to the given comprehensive score formula . Suppose there are:
[0181] ;
[0182] ;
[0183] Distance attenuation function (assuming that the distance attenuation is not significant);
[0184] Similarity enhancement function (assuming linear);
[0185] ;
[0186] Here represents the node weights that may have been adjusted through certain means, and other exponential factors and coefficients also need to be determined according to the actual situation.
[0187] After calculating the comprehensive scores of each logistics center through the above formula, they can be sorted according to the scores, and the logistics center with the highest score can be selected as the key node, and retain those connections with a connection strength exceeding the threshold . The finally formed structured description information can help decision-makers understand which logistics centers and paths are the most important, and then guide practical operations such as resource allocation and path planning. The calculation results show that logistics center 4 has the highest comprehensive score, and the strong connections are mainly concentrated on logistics centers 1 and 2, indicating that establishing an efficient logistics channel among these three logistics centers may be the optimal choice.
[0188] Figure 2 This is a schematic structural diagram of a system for extracting the theme of user question texts provided by an embodiment of the present application. As Figure 2 shown, the device includes:
[0189] The parsing module 21 parses the problem text submitted by the user based on the received problem text, and identifies the core intention and related concepts;
[0190] The mapping module 22 performs a mapping process on the pre-constructed knowledge graph according to the core intention and related concepts, and determines the topic expansion range based on the semantic relevance between nodes in the knowledge graph and the implicit relationship between the nodes learned through the graph neural network;
[0191] The adjustment module 23 evaluates the importance of the nodes within the topic expansion range and the centrality of the nodes in the graph neural network by using an adaptive weight allocation mechanism in combination with a link analysis algorithm based on the topic expansion range, so as to adjust the weights of the nodes;
[0192] The generation module 24 generates structured description information representing the topic corresponding to the problem text according to the adjusted node weights and the connection strength between the nodes calculated through point mutual information;
[0193] The processing module 25 further refines the structured description information by using a multi-layer attention mechanism to obtain an extraction result, where the extraction result refers to the summary output result of the topic corresponding to the problem text.
[0194] Figure 2 The described system for extracting the topic of the user problem text can execute Figure 1 The method for extracting the topic of the user problem text shown in the embodiment, and its implementation principle and technical effects will not be elaborated. For the system for extracting the topic of the user problem text in the above embodiment, the specific ways for each module and unit to perform operations have been described in detail in the embodiment related to the method, and will not be elaborated here.
[0195] In a possible design, Figure 2 The system for extracting the topic of the user problem text shown in the embodiment can be implemented as a computing device, such as Figure 3 shown, and the computing device can include a storage component 31 and a processing component 32;
[0196] The storage component 31 stores one or more computer instructions, where the one or more computer instructions are called and executed by the processing component 32.
[0197] The processing component 32 is used for: parsing the problem text based on the received problem text submitted by the user, and identifying the core intention and related concepts;
[0198] Perform mapping processing on a pre-constructed knowledge graph according to the core intention and associated concepts, and determine the topic expansion range based on the semantic relevance between nodes in the knowledge graph and the implicit relationship between the nodes learned through a graph neural network;
[0199] Based on the topic expansion range, adopt an adaptive weight allocation mechanism combined with a link analysis algorithm to evaluate the importance of nodes within the topic expansion range and the centrality of the nodes in the graph neural network, so as to adjust the weights of each node;
[0200] Generate structured description information representing the topic corresponding to the problem text according to the adjusted node weights and the connection strength between the nodes calculated through point mutual information;
[0201] Use a multi-layer attention mechanism to further refine the structured description information to obtain an extraction result, where the extraction result refers to the summary output result of the topic corresponding to the problem text.
[0202] Among them, the processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.
[0203] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0204] Of course, the computing device may also necessarily include other components, such as input / output interfaces, display components, communication components, etc.
[0205] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module can be an output device, an input device, etc.
[0206] The communication component is configured to facilitate wired or wireless communication between the computing device and other devices, etc.
[0207] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server. The above-mentioned processing components, storage components, etc. can be basic server resources leased or purchased from a cloud computing platform.
[0208] The embodiments of the present application also provide a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above-mentioned Figure 1 method for extracting the theme of a user question text in the shown embodiment.
[0209] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0210] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0211] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a magnetic disk, an optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A topic extraction method for user question text, characterized in that: include: Based on the received question text submitted by the user, the question text is parsed to identify the core intent and related concepts; According to the core intent and related concepts, mapping processing is performed on a pre-constructed knowledge graph, and the scope of topic expansion is determined based on the semantic relevance between nodes in the knowledge graph and the implicit relationship between the nodes learned through the graph neural network; Based on the topic extension scope, an adaptive weight allocation mechanism combined with a link analysis algorithm is used to evaluate the importance of nodes within the topic extension scope and the centrality of the nodes in the graph neural network to adjust the weight of each node; Generate structured description information representing the topic corresponding to the question text according to the adjusted node weights and the connection strength between the nodes calculated by point mutual information; The structured description information is further refined by using a multi-layer attention mechanism to obtain an extraction result, wherein the extraction result refers to a summary output result of the topic corresponding to the question text; Based on the topic extension scope, an adaptive weight allocation mechanism combined with a link analysis algorithm is used to evaluate the importance of nodes within the topic extension scope and the centrality of the nodes in the graph neural network to adjust the weight of each node, including: Based on the topic extension scope, the basic weights of the nodes within the topic extension scope are initialized using an adaptive weight allocation mechanism, and the local influence of the basic weight of each node within the topic extension scope is calculated based on the semantic relevance between the nodes in the knowledge graph and the implicit relationship between the nodes learned by the graph neural network to obtain a local influence result; Using the link analysis algorithm, iteratively calculate the scores of the nodes to obtain importance scores; The importance of the nodes is comprehensively evaluated by combining the local influence results and the importance scores, as well as the centrality of the nodes in the graph neural network, so as to dynamically adjust the weights of the nodes so that the nodes obtain higher weights.
2. The method according to claim 1, characterized in that According to the adjusted node weights and the connection strengths between the nodes calculated by point mutual information, structured description information representing the topic corresponding to the question text is generated, including: Based on the adjusted node weights, the nodes in the knowledge graph are prioritized to obtain core topic nodes; Using point mutual information, the semantic association strength between the core topic nodes is quantified to determine the closeness between the core topic nodes; According to the closeness between the core topic nodes, a clustering algorithm or a community discovery algorithm is used to group the core topic nodes to form a topic cluster; Comprehensively analyzing and processing the core topic nodes in the topic cluster to extract topic tags that can summarize the core meaning of the topic cluster; Combining the explicit concepts in the question text with the topic labels, supplementing potential implicit topics through a logical reasoning mechanism, and generating topic content corresponding to the implicit topics; Based on the subject content, the explicit concepts, the subject tags and the implicit subjects are organized to generate structured description information representing the subject corresponding to the question text.
3. The method according to claim 2, characterized in that Combining the explicit concepts in the question text with the topic labels, supplementing potential implicit topics through a logical reasoning mechanism, and generating topic content corresponding to the implicit topics, including: Constructing a preliminary topic framework based on the explicit concepts in the question text and the topic labels; Using the associations and context information between nodes in the knowledge graph, the subject framework is expanded to identify additional entities or concepts related to the explicit concept and the subject label; Based on rule-based reasoning or machine learning models, the relationships between the additional entities or concepts are analyzed and processed to mine implicit topics; Based on the implicit topic, the explicit concept and the topic label, consistency processing is performed through semantic fusion technology, and the explicit concept, the topic label and the implicit topic are integrated to generate the topic content corresponding to the implicit topic.
4. The method according to claim 1, characterized in that Combining the influence result and the importance score, a comprehensive evaluation process is performed on the importance of the node to dynamically adjust the weight of each node so that the node obtains a higher weight, including: Using the local influence result, the basic influence score of each node is calculated to obtain a basic influence score; Based on the basic influence score and the importance score, combined with a preset functional relationship, the comprehensive evaluation score of the node is calculated to obtain a comprehensive evaluation score reflecting the importance of the node within the scope of the topic extension; Using the comprehensive evaluation score, all nodes are sorted, and a threshold or ranking interval is set to determine nodes with high importance; A gain function is applied to increase the weight of the high-importance node so that the high-importance node obtains a higher weight.
5. The method according to claim 1, characterized in that According to the core intent and related concepts, mapping processing is performed on the pre-constructed knowledge graph, and the scope of topic expansion is determined according to the semantic relevance between nodes in the knowledge graph and the implicit relationship between the nodes learned through the graph neural network, including: Using a bidirectional encoder representation model and a semantic matching algorithm, based on the core intent and the associated concepts, the concepts in the question text are mapped in a pre-constructed knowledge graph to obtain a topic mapping result; According to the topic mapping results, combined with the semantic relevance between nodes in the knowledge graph, multi-step paths starting from the mapped nodes are explored and processed to identify potential nodes related to the core intent, so as to construct a graph neural network; The graph neural network is combined with a self-attention mechanism to analyze and process the implicit relationship between the potential nodes, and a context-aware model and Bayesian optimization technology are introduced to evaluate the relevance and importance of the potential nodes to the question text, and the relevance and importance scores of the potential nodes are dynamically adjusted to obtain adjustment results; According to the adjustment result, nodes that can be included in the subject extension scope are determined from the potential nodes, and the nodes in the subject extension scope are grouped in combination with a hierarchical clustering algorithm to determine the subject extension scope.
6. The method according to claim 1, characterized in that The structured description information is further refined by using a multi-layer attention mechanism to obtain an extraction result, including: Using the constructed multi-layer attention network model, encoding the obtained structured description information to obtain the structured description information after encoding; Based on the structured description information after the encoding process, the structured description information after the encoding process is preliminarily filtered and condensed by using the first-layer attention mechanism, the elements directly related to the core intent are identified and strengthened, and the irrelevant structured description information after the encoding process is weakened or ignored, so as to obtain a preliminary extraction result; According to the preliminary extraction result, in each subsequent layer of the attention mechanism, based on the output of the previous layer of the attention mechanism, the structured description information is subjected to a layer-by-layer in-depth analysis and processing to obtain the multi-layer attention mechanism screening result; Based on the multi-layer attention mechanism screening results, the output of the attention mechanism of the final layer is processed to obtain the extraction result.
7. A topic extraction system for user question text, characterized in that: include: A parsing module, based on the received question text submitted by the user, parses the question text to identify the core intent and related concepts; A mapping module performs mapping processing on a pre-constructed knowledge graph according to the core intent and related concepts, and determines the scope of topic expansion according to the semantic relevance between nodes in the knowledge graph and the implicit relationship between the nodes learned through the graph neural network; An adjustment module, based on the topic extension scope, adopts an adaptive weight allocation mechanism combined with a link analysis algorithm to evaluate the importance of nodes within the topic extension scope and the centrality of the nodes in the graph neural network to adjust the weight of each node; A generation module, generating structured description information representing a topic corresponding to the question text according to the adjusted node weights and the connection strength between the nodes calculated by point mutual information; A processing module further refines the structured description information using a multi-layer attention mechanism to obtain an extraction result, wherein the extraction result refers to a summary output result of the topic corresponding to the question text; Based on the topic extension scope, an adaptive weight allocation mechanism combined with a link analysis algorithm is used to evaluate the importance of nodes within the topic extension scope and the centrality of the nodes in the graph neural network to adjust the weight of each node, including: Based on the topic extension scope, the basic weights of the nodes within the topic extension scope are initialized using an adaptive weight allocation mechanism, and the local influence of the basic weight of each node within the topic extension scope is calculated based on the semantic relevance between the nodes in the knowledge graph and the implicit relationship between the nodes learned by the graph neural network to obtain a local influence result; Using the link analysis algorithm, iteratively calculate the scores of the nodes to obtain importance scores; The importance of the nodes is comprehensively evaluated by combining the local influence results and the importance scores, as well as the centrality of the nodes in the graph neural network, so as to dynamically adjust the weights of the nodes so that the nodes obtain higher weights.
Citation Information
Patent Citations
Question and answer method and system based on text-knowledge extension graph collaborative reasoning network
CN116361438A