Knowledge graph retrieval and classification method based on improved semi-supervised classification model
By improving the semi-supervised classification model, combining GCN, GAT and JK mechanisms, using semantic matching algorithms and inverted index structures, the shortcomings of knowledge graphs in node classification accuracy and retrieval efficiency are solved, and efficient and accurate knowledge graph retrieval and classification are achieved.
Patent Information
- Application Number
- CN202510771831.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
The existing knowledge graph retrieval and classification methods have shortcomings in node classification accuracy and retrieval efficiency, which is difficult to meet the actual application needs, especially when large-scale graph processing, performance bottlenecks are prominent, and insufficient labeling data lead to supervision learning difficulties.
The improved semi-supervised classification model is adopted, combined with graph convolutional network (GCN), graph attention network (GAT) and jump connection (JK) mechanisms, and the graph nodes are preprocessed and classified. The similarity score is calculated using the semantic matching algorithm of Jaccard coefficient and Levenshtein distance similarity, and the graph node matching and extended query are carried out to establish an inverse index structure optimization search.
It improves the classification accuracy of graph nodes and query keyword matching accuracy, improves the coverage and efficiency of retrieval, reduces dependence on labeled data, and is suitable for real-time query of large-scale knowledge graphs.
Smart Images

Figure CN120296179A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and in particular, to a knowledge graph retrieval and classification method based on an improved semi-supervised classification model. Background Art
[0002] As an efficient knowledge representation method, the knowledge graph has been widely applied in the field of intelligent systems, covering multiple application scenarios such as search engines, recommendation systems, question answering systems, medical diagnosis, and financial risk control. The knowledge graph can systematically store and organize information through a graph structure composed of nodes (entities) and edges (relationships), and support in-depth understanding and reasoning of complex data.
[0003] However, common knowledge graph retrieval and classification methods still have certain limitations in practical applications. On the one hand, traditional classification models are usually based on manually defined rules or simple graph algorithms, making it difficult to fully exploit the complex relationships and structural information in graph data, resulting in insufficient accuracy of graph node classification. On the other hand, although the knowledge graph itself has efficient information representation capabilities, during the retrieval process, graph traversal methods, especially depth-first search and breadth-first search, although performing well in small-scale graphs, have a significantly increased time complexity when dealing with complex queries and are difficult to meet real-time requirements.
[0004] That is to say, common knowledge graph retrieval and classification methods still have obvious deficiencies in terms of node classification accuracy and retrieval efficiency, and more optimized solutions are urgently needed to improve their performance.
[0005] The above content is only used to assist in understanding the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0006] The main purpose of the present invention is to provide a knowledge graph retrieval and classification method based on an improved semi-supervised classification model, aiming to solve the problem that common knowledge graph retrieval and classification methods still have obvious deficiencies in terms of node classification accuracy and retrieval efficiency.
[0007] To achieve the above purpose, the present invention provides a knowledge graph retrieval and classification method based on an improved semi-supervised classification model, and the knowledge graph retrieval and classification method based on the improved semi-supervised classification model includes the following steps:
[0008] Based on the improved semi-supervised classification model, perform a preprocessing classification operation on the graph nodes in the knowledge graph to determine the node features of the graph nodes;
[0009] When a query keyword is received, a semantic matching algorithm combining the Jaccard coefficient and the Levenshtein distance similarity is used to calculate the similarity score between the query keyword and the node features;
[0010] According to the similarity score and a preset similarity threshold, it is determined whether a graph entity is matched;
[0011] If so, the graph nodes corresponding to the query keyword and their directly associated nodes are obtained. If not, the graph node with the highest matching degree is determined according to the similarity score;
[0012] According to the graph node with the highest matching degree, an extended query action is executed to obtain the graph nodes related to the query keyword and their attribute information.
[0013] Optionally, the step of performing a preprocessing classification operation on the graph nodes in the knowledge graph based on the improved semi-supervised classification model to determine the node features of the graph nodes includes:
[0014] Based on a graph convolutional network, a feature aggregation action is performed on the graph nodes to capture the global information of the graph nodes;
[0015] The global information is passed to a graph attention network to determine the attention weights corresponding to the node neighbors of each graph node; and,
[0016] Based on a skip mechanism, the node information at different levels in the global information is fused to determine the node features of each graph node.
[0017] Optionally, the graph nodes include labeled nodes and unlabeled nodes. The step of performing a preprocessing classification operation on the graph nodes in the knowledge graph based on the improved semi-supervised classification model to determine the node features of the graph nodes further includes:
[0018] Based on a label propagation mechanism, the information of the labeled nodes is diffused to the unlabeled nodes.
[0019] Optionally, before the step of performing a preprocessing classification operation on the graph nodes in the knowledge graph based on the improved semi-supervised classification model to determine the node features of the graph nodes, it further includes:
[0020] Training the improved semi-supervised classification model;
[0021] The step of training the improved semi-supervised classification model includes:
[0022] Performing supervised learning on the improved semi-supervised classification model based on labeled data and performing unsupervised learning on the improved semi-supervised classification model based on unlabeled data;
[0023] Among them, the quantity of the labeled data is less than that of the unlabeled data.
[0024] Optionally, the step of performing an extended query operation according to the graph node with the highest matching degree to obtain graph nodes related to the query keyword and their attribute information includes:
[0025] Perform a multi-hop path search according to the graph node with the highest matching degree to obtain an extended result, where the extended result includes graph nodes and relationships on all paths;
[0026] Based on the extended result, determine graph nodes related to the query keyword and their attribute information.
[0027] Optionally, the step of performing an extended query operation according to the graph node with the highest matching degree to obtain graph nodes related to the query keyword and their attribute information includes:
[0028] Perform a type classification operation on the query keyword to obtain a classification result corresponding to the query keyword;
[0029] Determine a corresponding query path according to the classification result;
[0030] Perform a query operation based on the query path to obtain a query result;
[0031] Gradually perform an extension action according to the query result to obtain graph nodes related to the query keyword and their attribute information.
[0032] Optionally, before the step of performing a preprocessing classification operation on graph nodes in the knowledge graph based on an improved semi-supervised classification model to determine the node features of the graph nodes, it further includes:
[0033] Obtain target data and store the target data in a graph database in the form of triples;
[0034] In the graph database, establish an inverted index structure based on node attributes and relationships to form the knowledge graph.
[0035] Optionally, the step of establishing an inverted index structure based on node attributes and relationships in the graph database to form the knowledge graph includes:
[0036] Obtain the node name, node type, node attributes of each graph node, and the relationship information between the graph nodes;
[0037] Based on the node name, node type, node attributes, and the relationship information between the graph nodes, establish the inverted index structure based on node attributes and relationships to form the knowledge graph.
[0038] In addition, to achieve the above object, the present invention also provides a retrieval and classification device, which includes a memory, a processor, and a knowledge graph retrieval and classification program based on an improved semi-supervised classification model stored on the memory and operable on the processor. When the knowledge graph retrieval and classification program based on the improved semi-supervised classification model is executed by the processor, it implements the steps of the knowledge graph retrieval and classification method based on the improved semi-supervised classification model as described above.
[0039] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a knowledge graph retrieval and classification program based on an improved semi-supervised classification model is stored. When the knowledge graph retrieval and classification program based on the improved semi-supervised classification model is executed by a processor, it implements the steps of the knowledge graph retrieval and classification method based on the improved semi-supervised classification model as described above.
[0040] The beneficial effects of the present invention are as follows: 1. An improved semi-supervised classification model is constructed through the feature aggregation mechanism of the graph convolutional network (GCN) and the graph attention network (GAT), and in combination with the jump connection (JK) mechanism, to classify the graph nodes in the knowledge graph. 2. Through a fuzzy matching algorithm (a semantic matching algorithm that combines the Jaccard coefficient and the Levenshtein distance similarity), the similarity score is calculated by combining the Jaccard coefficient and the Levenshtein distance similarity, and the most relevant graph nodes are selected for matching. 3. When the matching fails, an extended query is performed based on the similarity score, and the multi-hop path search is used to gradually expand the query range. Finally, according to the user query type, relevant nodes and their attribute information are returned through the dynamic path search strategy, thereby realizing efficient and accurate knowledge graph retrieval and classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention. To more clearly illustrate the technical solutions of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 It is a schematic diagram of the architecture of the hardware operating environment of the retrieval and classification device related to the embodiments of the present invention;
[0043] Figure 2Schematic flowchart of a knowledge graph retrieval and classification method based on an improved semi-supervised classification model according to an embodiment of the present invention;
[0044] Figure 3 Schematic flowchart of the refinement of step S100 in the knowledge graph retrieval and classification method based on an improved semi-supervised classification model according to an embodiment of the present invention;
[0045] Figure 4 Overall flowchart framework of the knowledge graph retrieval and classification method based on an improved semi-supervised classification model according to an embodiment of the present invention;
[0046] Figure 5 Schematic flowchart framework of determining whether a graph node is matched by a fuzzy matching algorithm in the knowledge graph retrieval and classification method based on an improved semi-supervised classification model according to an embodiment of the present invention;
[0047] Figure 6 Schematic flowchart framework of extended query and path search in the knowledge graph retrieval and classification method based on an improved semi-supervised classification model according to an embodiment of the present invention.
[0048] The realization, functional characteristics and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners
[0049] With the rapid development of big data and artificial intelligence technologies, the knowledge graph, as an efficient knowledge representation method, has been widely applied in multiple intelligent application fields, including but not limited to search engines, recommendation systems, intelligent question answering, medical auxiliary diagnosis, and financial risk control, etc. The knowledge graph realizes the systematic organization and efficient storage of information through the graph structure constructed by graph nodes (entities) and edges (relationships), and can not only support the in-depth understanding of complex data, but also has powerful logical reasoning capabilities. However, despite the expanding application scope of the knowledge graph, the existing retrieval and classification methods still face many technical challenges in practical applications.
[0050] Traditional knowledge graph retrieval and classification methods mainly adopt rule-based classification models or basic graph algorithms. These methods rely too much on manually defined rules and are difficult to fully exploit the complex relationships and structural information in graph data, resulting in generally low accuracy of entity classification. Especially when dealing with large-scale knowledge graphs, the performance bottlenecks of these methods are more prominent, and their classification effects often fail to meet the actual application requirements. In terms of retrieval efficiency, although the knowledge graph itself has efficient information representation capabilities, common graph traversal-based retrieval methods (such as depth-first search and breadth-first search) often exhibit high time complexity when dealing with complex queries and are difficult to meet the real-time requirements. In addition, insufficient labeled data is also a problem faced by current knowledge graph retrieval and classification. In professional field applications, it is difficult and costly to obtain labeled data, which makes it difficult for traditional supervised learning methods to effectively train models.
[0051] Regarding the complexity characteristics of the graph structure of the knowledge graph, the diversity of its graph node types and relationship types, and the multi-level relationship network among entities, although graph neural networks (GNNs) can effectively capture the local relationships between graph nodes through convolutional operations, with the increase in the network depth, the common "over-smoothing" phenomenon in graph convolutional networks (GCNs) will cause the node features to gradually converge, thereby resulting in information loss and a decline in model performance.
[0052] To solve the above problems, this application proposes a knowledge graph retrieval and classification method based on an improved semi-supervised classification model. By using the improved semi-supervised classification model, preprocessing classification operations are performed on the graph nodes in the knowledge graph to determine the node features of the graph nodes. When a query keyword is received, based on a semantic matching algorithm that combines the Jaccard coefficient and the Levenshtein distance similarity, the similarity score between the query keyword and the node features is calculated. According to the similarity score and a preset similarity threshold, it is determined whether a graph entity is matched. If not, then according to the similarity score, the graph node with the highest matching degree is determined. According to the graph node with the highest matching degree, an extended query action is performed to obtain the graph nodes related to the query keyword and their attribute information. The purpose is to improve the classification accuracy of graph nodes, improve the matching accuracy between query keywords and graph nodes, and enhance the coverage and accuracy of retrieval.
[0053] To better understand the above technical solution, the exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0054] As an implementation solution,Figure 1 This is a schematic diagram of the hardware operating environment architecture of the retrieval and classification device involved in the embodiment solution of the present invention.
[0055] As Figure 1 shown, the retrieval and classification device may include: a processor 101, such as a Central Processing Unit (CPU), a memory 102, and a communication bus 103. Among them, the memory 102 may be a high-speed Random Access Memory (RAM) memory, or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 102 may also be a storage device independent of the aforementioned processor 101. The communication bus 103 is used to implement the connection and communication between these components.
[0056] Those skilled in the art can understand that Figure 1 the structure shown in
[0057] As Figure 1 shown, in the memory 102 as a computer-readable storage medium, there may be included an operating system, a data storage module, a network communication module, a user interface module, and a knowledge graph retrieval and classification method program based on an improved semi-supervised classification model.
[0058] In Figure 1 the retrieval and classification device shown, the processor 101 and the memory 102 may be provided in the retrieval and classification device. The retrieval and classification device calls the knowledge graph retrieval and classification program stored in the memory 102 through the processor 101, and performs the following operations:
[0059] Based on the improved semi-supervised classification model, perform a preprocessing classification operation on the graph nodes in the knowledge graph to determine the node features of the graph nodes;
[0060] When a query keyword is received, based on a semantic matching algorithm that combines the Jaccard coefficient and the Levenshtein distance similarity, calculate the similarity score between the query keyword and the node features;
[0061] According to the similarity score and a preset similarity threshold, determine whether a graph entity is matched;
[0062] If so, obtain the graph node corresponding to the query keyword and its directly associated nodes. If not, then according to the similarity score, determine the graph node with the highest matching degree;
[0063] Execute an extended query action based on the graph node with the highest matching degree to obtain graph nodes related to the query keyword and their attribute information.
[0064] In an embodiment, the processor 101 may be configured to call the knowledge graph retrieval and classification program stored in the memory 102 based on the improved semi-supervised classification model, and perform the following operations:
[0065] Based on the graph convolutional network, perform a feature aggregation action on the graph nodes to capture the global information of the graph nodes;
[0066] Transmit the global information to the graph attention network to determine the attention weights corresponding to the node neighbors of each graph node; and,
[0067] Based on the skip mechanism, fuse the node information at different levels in the global information to determine the node features of each graph node.
[0068] In an embodiment, the processor 101 may be configured to call the knowledge graph retrieval and classification program stored in the memory 102 based on the improved semi-supervised classification model, and perform the following operations:
[0069] Based on the label propagation mechanism, spread the information of the labeled nodes to the unlabeled nodes.
[0070] In an embodiment, the processor 101 may be configured to call the knowledge graph retrieval and classification program stored in the memory 102 based on the improved semi-supervised classification model, and perform the following operations:
[0071] Train the improved semi-supervised classification model;
[0072] The steps of training the improved semi-supervised classification model include:
[0073] Perform supervised learning on the improved semi-supervised classification model based on the labeled data, and perform unsupervised learning on the improved semi-supervised classification model based on the unlabeled data;
[0074] Wherein, the number of the labeled data is less than the number of the unlabeled data.
[0075] In an embodiment, the processor 101 may be configured to call the knowledge graph retrieval and classification program stored in the memory 102 based on the improved semi-supervised classification model, and perform the following operations:
[0076] According to the graph node with the highest matching degree, perform a multi-hop path search to obtain an extended result, where the extended result includes graph nodes and relationships on all paths;
[0077] Based on the expansion result, determine the graph nodes related to the query keyword and their attribute information.
[0078] In an embodiment, the processor 101 may be configured to call the knowledge graph retrieval and classification program stored in the memory 102 based on an improved semi-supervised classification model, and perform the following operations:
[0079] Perform a type classification operation on the query keyword to obtain a classification result corresponding to the query keyword;
[0080] Determine a corresponding query path according to the classification result;
[0081] Execute a query action based on the query path to obtain a query result;
[0082] Gradually execute an expansion action according to the query result to obtain the graph nodes related to the query keyword and their attribute information.
[0083] In an embodiment, the processor 101 may be configured to call the knowledge graph retrieval and classification program stored in the memory 102 based on an improved semi-supervised classification model, and perform the following operations:
[0084] Obtain target data, and store the target data in a graph database in the form of triples;
[0085] In the graph database, establish an inverted index structure based on node attributes and relationships to form the knowledge graph.
[0086] In an embodiment, the processor 101 may be configured to call the knowledge graph retrieval and classification program stored in the memory 102 based on an improved semi-supervised classification model, and perform the following operations:
[0087] Obtain the node name, node type, node attributes of each of the graph nodes, and the relationship information between the graph nodes;
[0088] Based on the node name, node type, node attributes, and the relationship information between the graph nodes, establish the inverted index structure based on node attributes and relationships to form the knowledge graph.
[0089] Based on the hardware architecture of the above-mentioned retrieval and classification device, embodiments of the knowledge graph retrieval and classification method based on an improved semi-supervised classification model according to the present invention are proposed.
[0090] Refer to Figure 2 , in an embodiment, the knowledge graph retrieval and classification method based on an improved semi-supervised classification model includes the following steps:
[0091] Step S100: Based on the improved semi-supervised classification model, perform preprocessing classification operations on the graph nodes in the knowledge graph to determine the node features of the graph nodes.
[0092] In this embodiment, the improved semi-supervised classification model performs preprocessing classification on the graph nodes in the knowledge graph to improve the classification quality of the graph nodes and optimize the retrieval effect in subsequent queries. It should be noted that the improved semi-supervised model refers to the GAT-GCN-JK model, where GAT refers to the Graph Attention Network, GCN refers to the Graph Convolutional Network, and JK (Jumping Knowledge) is the jumping connection mechanism.
[0093] Optionally, before performing the preprocessing classification operation using the above improved semi-supervised model, it also includes training the model, that is, training the improved semi-supervised classification model. Among them, the steps of training the improved semi-supervised classification model include performing supervised learning on the improved semi-supervised classification model based on labeled data and performing unsupervised learning on the improved semi-supervised classification model based on unlabeled data to obtain the improved semi-supervised model. It should be noted that the number of the labeled data is less than the number of the unlabeled data. Understandably, use a small amount of labeled data for supervised learning and a large amount of unlabeled data for unsupervised learning, and enhance the generalization ability of the improved semi-supervised model through the label propagation mechanism; during the learning process of the improved semi-supervised model, the classification of nodes can be adjusted according to the structural relationship and attribute information between nodes.
[0094] Specifically, the optimization objective function of the improved semi-supervised classification model can be expressed as:
[0095]
[0096] In the formula, is the cross-entropy loss function, is the true label, is the predicted label, is the regularization term, is the regularization coefficient, N is the total number of nodes, and M is the number of model parameters.
[0097] During the process of the improved semi-supervised classification model performing preprocessing classification operations on the graph nodes, GCN is used to aggregate the features of the graph nodes, and GAT assigns corresponding attention weights to the node neighbors of each graph node on this basis. Finally, the jumping connection mechanism is used to fuse the node information from different levels. The improved semi-supervised classification model can effectively aggregate information from the node neighbors of the graph nodes, and the influence of different node neighbors can be adjusted according to the graph attention mechanism.
[0098] In this embodiment, in improving the semi-supervised classification model, by combining a graph convolutional network and a graph attention network and introducing a skip connection mechanism, the graph nodes in the graph spectrum are classified. The purpose of doing so is to improve the effect of graph node classification. In addition, the skip connection mechanism is adopted to fuse node information from different levels, which can mitigate the over-smoothing problem in deep networks, thereby improving the accuracy and robustness of node classification.
[0099] Optionally, in the knowledge graph, the graph nodes include labeled nodes and unlabeled nodes. The step of performing preprocessing classification on the graph nodes in the knowledge graph based on the improved semi-supervised classification model to determine the node features of the graph nodes further includes diffusing the information of the labeled nodes to the unlabeled nodes based on the label propagation mechanism.
[0100] It can be understood that most graph nodes in the knowledge graph are usually unlabeled (lacking class or attribute labels). Label propagation utilizes the known information of the labeled nodes and infers the labels of the unlabeled nodes through the connection relationships in the graph structure (such as attention weights, similarity of node neighbors), so as to expand the scale of training data and reduce the dependence on manual annotation. Moreover, label propagation iteratively updates the label distribution of nodes, making the labels of adjacent graph nodes tend to be consistent, thereby capturing the implicit associations between graph nodes and enhancing the rationality of classification, especially for knowledge graphs with dense relationships. In addition, through the diffusion process, the unlabeled nodes obtain temporary or probabilistic labels, and this information can be input into the subsequent classification model as supplementary features (such as label probability vectors), providing support for the model to more comprehensively understand the semantics and context relationships of graph nodes. And label propagation only requires a small amount of initial annotations to generalize to the entire knowledge graph, avoiding the high annotation cost of full-supervised learning. Its computational complexity is usually lower than end-to-end training and is suitable for quickly generating preliminary classification results in the preprocessing stage.
[0101] That is to say, through the transitivity of the graph structure, the limited labeled information is maximally utilized, enabling the unlabeled nodes to obtain pseudo-labels or feature enhancements that match their topological positions, thereby improving the accuracy and robustness of subsequent classification tasks. The improved semi-supervised classification model that combines GCN, GAT, and JK mechanisms improves the classification accuracy of graph nodes and can fully utilize unlabeled data for learning even in the case of scarce labeled data, thus providing more accurate node classification results in retrieval.
[0102] Further, before the step of performing preprocessing classification on the graph nodes using the improved semi-supervised classification model, it further includes obtaining target data and storing the target data in a graph database in the form of triples; then, in the graph database, establishing an inverted index structure based on node attributes and relationships to form the knowledge graph.
[0103] In this embodiment, a graph database refers to a database system specifically used for storing and processing graph-structured data. With nodes, edges, and attributes as the core elements, it intuitively expresses the association relationships between entities through a graph theory model. The triple form refers to the form of "entity - relationship - entity". It can be understood that this inverted index structure based on node attributes and relationships is used to index node attributes and relationships, thereby optimizing query performance.
[0104] As an alternative implementation, the knowledge graph is stored in the graph database Neo4j in triple form. Each triple represents the basic entity relationships in the knowledge graph. In this way, all entities and their relationships in the knowledge graph are stored in a structured manner, thus providing data support for subsequent retrieval and classification.
[0105] Optionally, the steps of establishing the above inverted index structure include obtaining the node names, node types, node attributes of each of the graph nodes, and the relationship information between the graph nodes; then, based on the node names, node types, node attributes, and the relationship information between the graph nodes, establishing the inverted index structure based on node attributes and relationships to form the knowledge graph.
[0106] Specifically, the inverted index structure can be represented by the following formula:
[0107]
[0108] where n represents a graph node, represents the relationship related to node n.
[0109] It can be understood that based on the characteristics prepared by the inverted index structure, by establishing the inverted index structure, it is possible to efficiently retrieve and query the graph nodes and their relationships related to the keywords, improve the speed and accuracy of graph queries, and thus optimize the subsequent retrieval and classification processes.
[0110] Step S200: When a query keyword is received, calculate the similarity score between the query keyword and the node features based on a semantic matching algorithm that combines the Jaccard coefficient and the Levenshtein distance similarity.
[0111] In this embodiment, the received query keywords are input by the user. The Jaccard coefficient (set overlap) is good at capturing the term overlap between keywords and node features. The Levenshtein distance (edit distance) is used to measure spelling or morphological similarity and solve spelling variants or abbreviations. The combination of the Jaccard coefficient and the Levenshtein distance covers exact and fuzzy matches at the lexical level, thereby achieving the purpose of reducing mis-matches caused by term differences.
[0112] Optionally, in the semantic matching algorithm, the calculation formula for the similarity score of semantic matching is:
[0113]
[0114] Where is the Jaccard coefficient, is the edit distance similarity, is the weight coefficient. It should be noted that . Preferably, takes 0.5 to obtain the highest information retrieval accuracy. By combining the similarities of the two with weights, different types of matching situations are fully considered.
[0115] It can be understood that the Jaccard coefficient is a measure to measure the similarity and difference between different finite sample sets. When solving the Jaccard coefficient between the query keyword w and the graph entity e, the two can be regarded as a character sequence set. First, calculate the number of identical characters Same(w, e) between the two, and then calculate the Jaccard coefficient through the Jaccard coefficient calculation formula. The Jaccard coefficient calculation formula is:
[0116]
[0117] Where the Sizeh function represents the number of different characters in the string, w is the query keyword, and e is the graph entity.
[0118] Optionally, the edit distance similarity The calculation formula is:
[0119]
[0120] In the formula, is the edit distance between the retrieval term and the node name, refers to the length of the word with more characters in the query keyword and the graph entity e, and is used to normalize the lengths of the two strings. is used to reflect the degree of difference between the two, and refers to the minimum number of operations to convert the string w into e. Specifically, it is calculated through the dynamic programming algorithm. The calculation formula is:
[0121]
[0122] Among them, i is the position of the character in the query term, j is the position of the character in the node.
[0123] By weighted combination of the Jaccard coefficient and the edit distance similarity, the purpose of improving the fault tolerance ability for synonyms, spelling mistakes, and fuzzy matching is achieved.
[0124] The semantic matching algorithm improves the matching accuracy between the query keyword and the graph nodes through the combination of the Jaccard coefficient and the edit distance, achieving the purpose of effectively dealing with fuzzy matching problems such as synonyms and spelling mistakes.
[0125] Step S300: Determine whether a graph entity is matched according to the similarity score and a preset similarity threshold.
[0126] In this embodiment, by quantitatively evaluating the semantic association strength between the query keyword and the graph nodes and combining with the preset similarity threshold, accurate matching and fault-tolerant retrieval are realized, ensuring that the query results are highly consistent with the user's intention, and problems caused by term differences and spelling mistakes are avoided through threshold boundary determination.
[0127] It should be noted that the preset similarity threshold can be adjusted according to the actual needs of different application scenarios. For example, in a medical graph, a relatively high preset similarity threshold needs to be adopted for disease names to control strict matching and avoid misdiagnosis.
[0128] Step S400: If so, obtain the graph node corresponding to the query keyword and its directly associated nodes; if not, determine the graph node with the highest matching degree according to the similarity score;
[0129] In this embodiment, different execution strategies are adopted based on different matching results. When the query keyword matches the graph node, return the graph node and its directly associated nodes; when the query keyword does not match the graph node, the most relevant node is selected as the "anchor point" based on the similarity score ranking for subsequent extended queries.
[0130] When the query keyword does not match the graph node, that is, when the similarity scores of all graph nodes are less than the preset similarity threshold, the graph node with the highest matching degree, that is, the graph node with the highest similarity score, is selected as the benchmark for subsequent extended queries to ensure that the extension starting point has semantic relevance.
[0131] Step S500: Perform an extended query action according to the graph node with the highest matching degree to obtain the graph nodes related to the query keyword and their attribute information.
[0132] In this embodiment, based on the graph node with the highest matching degree, an extended query action is performed to expand the query scope and obtain more relevant information.
[0133] Optionally, step S500 includes performing a multi-hop path search according to the graph node with the highest matching degree to obtain an extended result, where the extended result includes graph nodes and relationships on all paths; and determining graph nodes related to the query keyword and their attribute information based on the extended result. It should be noted that the maximum depth of the path search can be specified by the user.
[0134] It can be understood that the basic process of multi-hop path search includes anchor point positioning, path expansion, and result filtering. In anchor point positioning, the graph node with the highest similarity is used as the starting point; in path expansion, starting from the anchor point, traversing hop by hop along the edge relationships of the knowledge graph, where 1-hop neighbors are directly associated nodes and 2-hop neighbors are neighbors of neighbors; in result filtering, irrelevant nodes are filtered according to path weights, node types, or semantic relevance. It should be noted that when performing multi-hop path search, path depth control can be performed, that is, the maximum number of hops is limited, such as 3 hops, to avoid explosion of the search space.
[0135] Multi-hop path search expands the retrieval scope while ensuring efficiency through graph traversal with a limited number of steps to solve the problems of semantic gap and annotation sparsity in the knowledge graph.
[0136] Furthermore, the step of performing an extended query action based on the graph node with the highest matching degree to obtain graph nodes related to the query keyword and their attribute information further includes performing a type classification operation on the query keyword to obtain a classification result corresponding to the query keyword; then, determining a corresponding query path according to the classification result; and performing a query action based on the query path to obtain a query result; then, gradually performing an expansion action according to the query result to obtain graph nodes related to the query keyword and their attribute information. It should be noted that there are multiple groups of graph nodes related to the query keyword and their attribute information for the user to select.
[0137] Specifically, the formula used for extended query is:
[0138]
[0139] In the formula, is the set of neighbor nodes of node , is the attention coefficient of node to node , is the weight matrix of the th layer, Is a node In the Through recursive expansion, the relevance of the query node is gradually improved to ensure that highly relevant query results are returned.
[0140] Optionally, the optimization process of path search is expressed as:
[0141]
[0142] In the formula, Represents the path expansion result, are the nodes on the path, is the path depth. The path depth can be adjusted according to the query type entered by the user and the number of nodes that need to be returned.
[0143] Through the path search strategy, the query scope can be flexibly expanded, and relevant nodes and their attribute information can be dynamically returned to achieve the purpose of improving the coverage and accuracy of the retrieval.
[0144] As an optional implementation method, the improved semi-supervised classification model sorts the query results according to the node type and node attribute information during the query process to ensure that the returned graph nodes are more semantically relevant to the query keywords, further improving the accuracy of the retrieval.
[0145] As another optional implementation method, the query results finally returned are output in JSON format, including matching graph nodes, expansion nodes and their triple relationships, and attached classification label information of the graph nodes, to facilitate front-end display and user interaction.
[0146] In the technical solution provided in this embodiment, by combining the improved semi-supervised classification model of GCN, GAT and JK mechanism, the graph nodes are pre-processed and classified, and the node features of the graph nodes are determined. The node features formed in the classification stage not only contain the original attributes, but also integrate the contextual semantics inferred by label propagation, so that the node feature expression is more comprehensive, and the model's classification ability for sparse or long-tail nodes is enhanced, thereby achieving the purpose of improving the node classification accuracy.
[0147] Through the semantic matching algorithm combining the Jaccard coefficient and Levenshtein distance similarity, the similarity score between the query keywords and the node features is calculated to reduce the mismatch caused by terminology differences; and low-quality matches are filtered out by a preset similarity threshold, and then the node expansion query is performed by determining the graph node with the highest matching degree, avoiding the retrieval omissions caused by strict matching in traditional methods, thereby achieving the purpose of improving the query matching accuracy.
[0148] Improve the semi-supervised classification model, preprocess and classify the graph nodes, transform the classification task into iterative calculations on the graph, and reduce the computational burden of real-time retrieval; expand the query to directly associate the extraction of graph node and attribute information, make full use of the topological connectivity of the knowledge graph. Even if the query keyword exactly matches a certain entity, the results can still be expanded through associated nodes, improving the retrieval recall rate, and thus enhancing the efficiency and coverage of the retrieval.
[0149] That is to say, a knowledge graph retrieval and classification method based on an improved semi-supervised classification model provided in this embodiment uses graph structure information to complement insufficient annotations, improving the classification robustness; fuses multi-dimensional semantic similarities to solve the problem of term differences; breaks through the strict matching limit with an extended query mechanism, taking into account both efficiency and coverage.
[0150] In the embodiment, as Figure 3 shown, step S100 further includes:
[0151] Step S110: Based on the graph convolutional network, perform a feature aggregation operation on the graph nodes to capture the global information of the graph nodes;
[0152] Step S120: Transmit the global information to the graph attention network to determine the attention weights corresponding to the node neighbors of each graph node; and,
[0153] Step S130: Based on the skip mechanism, fuse the node information at different levels in the global information to determine the node features of each graph node.
[0154] In this embodiment, the graph convolutional network is used to aggregate the node information of the node neighbors of each graph node, thereby updating the features of the nodes. The basic operation of the GCN is the message passing mechanism, that is, the graph nodes aggregate information according to the features of their node neighbors. In each layer of graph convolution, the feature vector of the graph node is obtained by weighted summation of the feature vectors of its node neighbors. The feature update formula for each node is:
[0155]
[0156] In the formula represents the feature representation of the graph node at the layer. is the graph node 's set of node neighbors. and are respectively the degrees of the node and its neighbor node . is the weight matrix of the layer, is the node neighbor At the layer features. is the activation function ReLU. This equation indicates that the features of each graph node are updated by aggregating the node features of its node neighbors, and a regularization factor is used to adjust the impact of node degree on information propagation.
[0157] The attention mechanism is introduced through the graph attention network, enabling each graph node to assign different weights to different node neighbors based on the node features of its node neighbors. This mechanism can adaptively adjust the degree of influence of node neighbors on the information update of the current node. For graph node and its node neighbor , the representation of graph node is updated as:
[0158]
[0159] In the equation, is the updated feature representation of graph node . is the feature of node neighbor . is the weight matrix. is the weight calculated through the attention mechanism, representing the attention degree of graph node to node neighbor . The attention weight is calculated as follows:
[0160]
[0161] In the equation, is the parameter used to calculate the attention coefficient. LeakyReLU() is the activation function. When the input is positive, the output is the same as the input; when the input is negative, the output is the result of multiplying the negative value of the input by a very small constant (taking 0.01). represents the concatenation operation of feature vectors. The function ensures the normalization of the weight coefficients. Through this attention mechanism, GAT can adaptively adjust the weights according to the features of each node neighbor, enabling the model to capture important structural information in the graph data.
[0162] In this embodiment, the skip connection mechanism is used to fuse the node features of different layers, thereby enhancing the representation ability of graph nodes. Traditional GNNs usually can only use the node features of the current layer for classification, while the skip connection aggregates the node features of multiple layers together, improving the expressive power of the model. The final feature representation of graph nodes can be calculated as follows:
[0163]
[0164] Wherein, , , , are the feature representations of the graph nodes from the first layer to the L-th layer. Concat means concatenating these feature vectors together. The role of the skip connection is to merge the feature information of different layers, so that the nodes can obtain context information from different depths, which is particularly important for complex graph data.
[0165] During the node classification process, the optimization objective of the GCN-GAT-JK model is to improve the accuracy of node classification by minimizing the cross-entropy loss function. The optimization objective function of the model is as follows:
[0166]
[0167] Wherein, is the cross-entropy function of node i, is the true label of the node, is the predicted label of the node, is the regularization term, which is used to avoid overfitting. is the regularization coefficient, which is used to adjust the strength of regularization. N is the total number of nodes, and M is the number of model parameters. By minimizing this loss function, the model can learn the optimal feature representation of the nodes and improve the accuracy of node classification.
[0168] In the technical solution provided in this embodiment, the improved semi-supervised classification model can perform efficient node classification on graph data, and in a semi-supervised learning environment, it can use a small amount of labeled data and a large amount of unlabeled data for training, thereby improving the classification effect.
[0169] In the embodiment, as Figure 4 shown, the overall process of the knowledge graph retrieval and classification method based on the improved semi-supervised classification model is as follows: First, construct a knowledge graph; then, the improved semi-supervised classification model (GAT-GCN-JK) performs node classification; then, it is judged whether a graph node is matched through a fuzzy matching algorithm (semantic matching algorithm). If so, the matched graph node and its information are returned, and then the process ends; if not, an extended query is performed. In the extended query, the extended query action is executed according to the path depth; after the extended query is completed, the query result (related nodes and attributes) is returned, and then the process ends.
[0170] In this embodiment, as Figure 5As shown in the figure, the process of determining whether a graph node is matched by the fuzzy matching algorithm is as follows: after receiving the query keyword input by the user, calculate the similarity between the query keyword and the graph node; then, determine whether the similarity score reaches the preset similarity threshold. If so, return the matched graph node and information and output the query result; if not, perform an extended query, calculate the query result after node expansion, and then output the query result.
[0171] Further, as Figure 6 shown in the figure, the process of the above-mentioned extended query and path search is as follows: after obtaining the type of the query keyword input by the user, determine the query type and select the query path; then, extend the query path, dynamically adjust the path depth, and after returning more nodes according to the user's needs, return the result of the extended query (the graph nodes related to the query keyword and their attribute information).
[0172] The present invention also provides a computer-readable storage medium, which stores a knowledge graph retrieval and classification program based on an improved semi-supervised classification model. When the knowledge graph retrieval and classification program based on the improved semi-supervised classification model is executed by a processor, it implements each step of the knowledge graph retrieval and classification method based on the improved semi-supervised classification model as described in the above embodiments.
[0173] Among them, the computer-readable storage medium can be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.
[0174] It should be noted that since the storage medium provided in the embodiments of the present application is the storage medium used to implement the method of the embodiments of the present application, those skilled in the art can understand the specific structure and deformation of the storage medium based on the method introduced in the embodiments of the present application, so it will not be elaborated here. Any storage medium used in the method of the embodiments of the present application belongs to the scope to be protected by the present application.
[0175] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0176] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0177] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0179] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several means, several of these means can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0180] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0181] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.
Claims
1. A knowledge graph retrieval and classification method based on an improved semi-supervised classification model, characterized in that The method includes the following steps: Based on an improved semi-supervised classification model, perform a preprocessing classification operation on the graph nodes in the knowledge graph to determine the node features of the graph nodes; When a query keyword is received, calculate the similarity score between the query keyword and the node features based on a semantic matching algorithm that combines the Jaccard coefficient and the Levenshtein distance similarity; Determine whether a graph entity is matched according to the similarity score and a preset similarity threshold; If so, obtain the graph nodes corresponding to the query keyword and their directly associated nodes. If not, determine the graph node with the highest matching degree according to the similarity score; According to the graph node with the highest matching degree, perform an extended query action to obtain the graph nodes related to the query keyword and their attribute information.
2. The knowledge graph retrieval and classification method based on the improved semi-supervised classification model according to claim 1, characterized in that, The step of performing a preprocessing classification operation on the graph nodes in the knowledge graph based on the improved semi-supervised classification model to determine the node features of the graph nodes includes: Based on a graph convolutional network, perform a feature aggregation action on the graph nodes to capture the global information of the graph nodes; Transmit the global information to a graph attention network to determine the attention weights corresponding to the node neighbors of each graph node; and Based on a skip mechanism, fuse the node information at different levels in the global information to determine the node features of each graph node.
3. The knowledge graph retrieval and classification method based on the improved semi-supervised classification model according to claim 1, characterized in that The graph nodes include labeled nodes and unlabeled nodes. The step of performing a preprocessing classification operation on the graph nodes in the knowledge graph based on the improved semi-supervised classification model to determine the node features of the graph nodes further includes: Based on a label propagation mechanism, spread the information of the labeled nodes to the unlabeled nodes.
4. The knowledge graph retrieval and classification method based on the improved semi-supervised classification model according to claim 1, characterized in that, Before the step of performing a preprocessing classification operation on the graph nodes in the knowledge graph based on the improved semi-supervised classification model to determine the node features of the graph nodes, it further includes: Train the improved semi-supervised classification model; The step of training the improved semi-supervised classification model includes: Perform supervised learning on the improved semi-supervised classification model based on labeled data and perform unsupervised learning on the improved semi-supervised classification model based on unlabeled data; Among them, the number of the labeled data is less than the number of the unlabeled data.
5. The knowledge graph retrieval and classification method based on the improved semi-supervised classification model according to claim 1, wherein, The step of performing an extended query action according to the graph node with the highest matching degree to obtain the graph nodes related to the query keyword and their attribute information includes: According to the graph node with the highest matching degree, perform a multi-hop path search to obtain an extended result, where the extended result includes the graph nodes and relationships on all paths; Based on the extended result, determine the graph nodes related to the query keyword and their attribute information.
6. The method for knowledge graph retrieval and classification based on an improved semi-supervised classification model according to claim 1, wherein The step of performing an extended query action according to the graph node with the highest matching degree to obtain the graph nodes related to the query keyword and their attribute information includes: Perform a type classification operation on the query keyword to obtain the classification result corresponding to the query keyword; According to the classification result, determine the corresponding query path; Perform a query action based on the query path to obtain a query result; Gradually execute expansion actions according to the query results to obtain graph nodes related to the query keywords and their attribute information.
7. The knowledge graph retrieval and classification method based on the improved semi-supervised classification model according to claim 1, characterized in that, Before the step of performing a preprocessing classification operation on the graph nodes in the knowledge graph based on the improved semi-supervised classification model to determine the node features of the graph nodes, it further includes: Obtain target data and store the target data in a graph database in the form of triples; In the graph database, establish an inverted index structure based on node attributes and relationships to form the knowledge graph.
8. The knowledge graph retrieval and classification method based on the improved semi-supervised classification model according to claim 7, wherein, The step of establishing an inverted index structure based on node attributes and relationships in the graph database to form the knowledge graph includes: Obtain the node names, node types, node attributes of each graph node, and the relationship information between the graph nodes; Based on the node names, node types, node attributes, and the relationship information between the graph nodes, establish the inverted index structure based on node attributes and relationships to form the knowledge graph.
9. A retrieval and classification device, characterized in that, The retrieval and classification device includes: a memory, a processor, and a knowledge graph retrieval and classification program based on an improved semi-supervised classification model stored on the memory and executable on the processor. The knowledge graph retrieval and classification program based on the improved semi-supervised classification model is configured to implement the steps of the knowledge graph retrieval and classification method based on the improved semi-supervised classification model as described in any one of claims 1 to 8.
10. A readable storage medium, characterized in that, A knowledge graph retrieval and classification program based on an improved semi-supervised classification model is stored on the readable storage medium. When the knowledge graph retrieval and classification program based on the improved semi-supervised classification model is executed by a processor, it implements the steps of the knowledge graph retrieval and classification method based on the improved semi-supervised classification model as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Node relation graph processing method and device, equipment and storage medium
CN112989134A
Retrieval method and device based on knowledge graph, electronic equipment and storage medium
CN113761219A
Text knowledge multi-hop question and answer method based on implication tree form
CN118095427A
Semi-supervised entity alignment method based on multi-hop attention mechanism
CN118153679A
Project evaluation and review method and system fused with natural language processing
CN118780767A