Knowledge graph knowledge expansion method

By applying sub-graph knowledge relational attention completion network and contrast loss learning technology in the home appliance manufacturing knowledge graph, the problems of relationship omissions and data in the knowledge graph are solved, and efficient expansion of the knowledge graph and the accuracy of relationship prediction are achieved.

CN119917701AActive Publication Date: 2025-05-02OCEAN UNIV OF CHINA

Patent Information

Application Number
CN202510421139.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-02
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing home appliance manufacturing knowledge graph faces the problems of relationship omissions and incomplete data, resulting in poor application results in scenarios such as manufacturing Q&A, fault analysis, quality analysis and personalized recommendations.

Method used

A knowledge graph knowledge completion method based on sub-graph knowledge relational attention completion network combined with contrast loss learning technology is proposed. The prediction results are generated through the integration of K-hop neighborhood and sub-graph extraction of link length constraints, multi-layer relationship message delivery, link attention mechanism and gating mechanism, and the prediction results are generated, and the combination of comparison loss and cross-entropy loss is optimized.

Benefits of technology

It significantly improves the knowledge expansion efficiency of the home appliance manufacturing knowledge graph, enhances the accuracy of relationship prediction, reduces the cost of manual labeling, reduces the dependence of expert knowledge, and solves the problems of data sparseness and relationship omissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917701A_ABST
    Figure CN119917701A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph knowledge expansion method, which comprises the following steps of: extracting a sub-graph of a target entity through a knowledge graph sub-graph extraction method based on K-hop neighborhood and link length constraint to reduce resource consumption, and then extracting neighborhood relation representation information and neighborhood relation link representation information by utilizing an attention mechanism; fusing the two kinds of relationship representation information by utilizing a relationship fusion method based on a gating mechanism, and further performing relationship prediction completion; and finally, model optimization is performed by combining comparison loss and cross entropy loss, more negative samples are increased through a subgraph structure-based negative sampling method, feedback on sparse data is enriched, the defect of the cross entropy loss in generalization performance is made up, and effective expansion of the knowledge graph is realized. According to the method, the relationship prediction accuracy of the knowledge graph is improved, the data sparsity problem of the knowledge graph is effectively relieved, the manual annotation cost is remarkably reduced, and excessive dependence on expert knowledge is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of knowledge management, and specifically, relates to a knowledge graph knowledge expansion method. Background Art

[0002] Intelligent manufacturing is one of the key measures to enhance the core competitiveness of the manufacturing industry. Although online monitoring and prediction based on production big data have been applied in manufacturing scenarios, intelligent programs are still insufficient in scenarios that require complex decision-making. Most complex problems still rely on experts to solve them manually, which not only requires a lot of experience and knowledge, but is also costly.

[0003] High-quality knowledge graphs can comprehensively store the knowledge precipitation and accumulation formed by manufacturing enterprises in the long-term production and operation innovation activities, helping decision makers to correctly analyze enterprise data, so as to timely discover and solve problems in enterprise production management and predict industry development trends and opportunities. The home appliance industry is an important part of the manufacturing industry. Applying knowledge graphs to the field of home appliance manufacturing can not only process, organize and manage accumulated experience and case knowledge to form a knowledge base, but also conduct knowledge reasoning on similar or even repeated decision-making activities in the manufacturing process, discover new knowledge, and thus realize the accumulation, inheritance and reuse of manufacturing knowledge.

[0004] However, due to the limitations of annotation resources and technology, as well as the high cost of manually searching for all fact triples, almost all knowledge graphs currently face the challenges of missing relationships and incomplete data, and incomplete home appliance manufacturing knowledge graphs are difficult to meet the application requirements of scenarios such as manufacturing question answering, fault analysis, quality analysis, and personalized recommendations. Relationship completion technology aims to fill the missing connections between entities. It is an effective strategy to solve the incompleteness of knowledge graphs, expand knowledge graphs, and promote knowledge growth. However, the scenarios of the home appliance manufacturing industry are complex and changeable. The existing relationship completion models are difficult to accurately adapt to the situations of missing relationships and incomplete data in the field of home appliance manufacturing, and often ignore the rich reasoning patterns between entities, which affects the accuracy of relationship completion. Summary of the invention

[0005] In response to the problems of relationship omissions and incomplete data existing in the current knowledge graph applications in the manufacturing field, the present invention proposes a knowledge graph knowledge completion method based on a subgraph knowledge relationship attention completion network combined with contrastive loss learning technology, which improves the knowledge expansion efficiency of the home appliance manufacturing knowledge graph, expands its coverage, solves the problems of relationship omissions and incomplete data in the existing home appliance manufacturing knowledge graph, and can significantly reduce the cost of manual annotation and reduce excessive reliance on expert knowledge.

[0006] The present invention is implemented by the following technical solutions:

[0007] A knowledge graph knowledge expansion method is proposed, which is characterized by comprising:

[0008] S1: A knowledge graph subgraph extraction method based on K-hop neighborhood and link length constraints extracts subgraphs related to the target triples from the original knowledge graph;

[0009] S2: Utilize a multi-layer relational message passing strategy based on subgraph attention to aggregate the neighborhood relation representation information of target entity pairs;

[0010] S3: Using the neighborhood relationship representation information as a priori guidance, weights are assigned to different relationship links based on the link attention mechanism to obtain the neighborhood relationship link representation information of the target entity pair;

[0011] S4: A relationship information fusion method based on a gating mechanism is used to fuse the neighborhood relationship representation information and the neighborhood relationship link representation information to generate a prediction result;

[0012] S5: Introduce the contrast loss learning idea to optimize the model, and increase the number of negative samples based on the negative sampling method of the subgraph structure to enrich the sparse data feedback.

[0013] In some embodiments of the present invention, S1 specifically includes:

[0014] Generate candidate subgraphs using a breadth-first search algorithm; and,

[0015] A maximum link length threshold is set in the algorithm. When the length of any link in the search exceeds the maximum link length threshold, all connected edges are removed.

[0016] In some embodiments of the present invention, the subgraph attention-based multi-layer relationship message delivery strategy in S2 includes message delivery and message aggregation; wherein,

[0017] Messaging includes:

[0018] The edge weight calculation formula is used to calculate the importance weights of different edges:

[0019] ;

[0020] use Aggregate neighborhood messages and update node representations;

[0021] use Combine node messages and their own status to update the edge Hidden state;

[0022] Above, use represents the attention weight of edge e to the message of node n in the i-th iteration, represents the hidden state of edge e in the i-th iteration, represents the message received by node n in the i-th message transmission, obtained by aggregating all the connecting edges r of node n; a is the learnable attention vector, is to transform the hidden state of edge e With the characteristics of node n The concatenated vector, represents any edge in the neighborhood of the target node n; N(n) is the neighborhood set of node n; p and q are the two endpoints of edge e, For edge The set of adjacent points, [·] is the splicing function, w i and b i They represent the learnable transformation matrix and bias σ(·) is a nonlinear activation function. is the initial feature of edge e, which can be used as the one-hot unit vector of the relationship type to which e belongs;

[0023] Message aggregation includes:

[0024] based on and Calculate the final message of the entity pair (s,o);

[0025] Aggregate the final message to get the relation attribute embedding of the entity pair (s,o):

[0026] ;

[0027] Among them, k is the number of times the neighborhood relationship is transmitted in the subgraph.

[0028] In some embodiments of the present invention, S3 specifically includes:

[0029] For each relationship link P( ) to model and obtain the representation of the relationship link ;

[0030] Represent information C through relationship links and neighborhood relationships (s, o) The similarity of is used to calculate the attention weight of the link ;in, It is the aggregate representation of the relationship link from entity s to o.

[0031] In some embodiments of the present invention, the relationship information fusion method based on the gating mechanism in S4 specifically includes:

[0032] Gated fusion module receives and Two inputs; is the neighborhood relationship representation of the entity pair (s,o), is the neighborhood relationship link representation of the entity pair (s,o); Indicates that and Perform vector stitching;

[0033] based on Calculate the gate vector It is used to control the weighted ratio of domain relationship representation and neighborhood relationship link representation; where σ represents the activation function, W g and b g They are respectively the learnable weight matrix and bias term;

[0034] based on The neighborhood relationship representation information and the neighborhood relationship link representation information are fused to obtain the embedded representation; ⊙ represents the element-by-element product operation, (1−g) represents the weight adjustment.

[0035] In some embodiments of the present invention, in S4, a prediction result is generated based on the probability that a relationship r exists between the subject and the guest entities:

[0036] ,in, is the projection matrix used to map the joint representation of entity pairs into the relation space.

[0037] In some embodiments of the present invention, contrast loss and cross entropy loss are combined for model training in S5:

[0038] ; and is the weight used to control the relationship prediction loss; is the contrast loss, is the cross entropy loss;

[0039] The contrast loss is calculated as:

[0040] ; is the set of all triplets in the training set; D + and D - Represent positive and negative triplets respectively; γ is a boundary hyperparameter used to control the minimum gap between positive and negative sample pairs, and Represent the prediction scores of positive and negative example triplets respectively;

[0041] The cross entropy loss is calculated as:

[0042] ; T is the number of triple samples in a batch of training, R represents all the relations in the knowledge graph, Represents the true value label. If the triple (s, r, o) is valid, then =1, otherwise 0; is the model’s predicted probability for the triple (s, r, o).

[0043] In some embodiments of the present invention, the negative sampling method based on the subgraph structure in S5 includes:

[0044] For each positive sample, extract the entities and relations in the n-hop neighborhood associated with it from the knowledge graph;

[0045] In the n-hop neighborhood set of the positive sample, select entities with similar neighborhood structures to replace the main entity or guest entity in the positive sample, and select other relations with similar semantics to replace the original relations;

[0046] Verify the effectiveness of the generated negative samples;

[0047] Multiple negative samples are generated for each positive sample triple by combining different types of negative samples.

[0048] Compared with the prior art, the advantages and positive effects of the present invention are as follows: the knowledge graph knowledge expansion method proposed in the present invention first reduces resource consumption by extracting the subgraph of the target entity through the knowledge graph subgraph extraction method based on K-hop neighborhood and link length constraints, and then uses the attention mechanism to extract the neighborhood relationship representation information and the neighborhood relationship link representation information, and uses the relationship fusion method based on the gating mechanism to fuse these two types of relationship representation information, and then performs relationship prediction completion. Finally, by combining contrast loss and cross entropy loss for model optimization, more negative samples are added by introducing a negative sampling method based on the subgraph structure, enriching the feedback of sparse data, and making up for the shortcomings of cross entropy loss in generalization performance, the effective expansion of the knowledge graph is achieved. Compared with traditional methods, the present invention uses subgraph relationship representation information and contrast loss learning technology to improve the accuracy of relationship prediction of the knowledge graph, effectively alleviate the data sparsity problem of the knowledge graph, and can also effectively solve the problems of high cost of manual annotation and excessive reliance on expert knowledge.

[0049] Other features and advantages of the present invention will become more apparent after reading the detailed description of the embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0051] Figure 1 This is a schematic diagram of the execution steps of the knowledge graph knowledge expansion method proposed in the present invention;

[0052] Figure 2 This is an example of relational information in the home appliance manufacturing knowledge graph in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] For manufacturing enterprises, building a complete home appliance knowledge graph is crucial to achieving the mutual coordination of knowledge management and application in the entire chain cycle of electrical manufacturing. To address the problems of missing relationships and incomplete data in the graph, the industry usually uses knowledge graph relationship completion technology to expand the knowledge of the manufacturing knowledge graph. However, when the existing relationship completion technology is applied to the manufacturing field with sparse data, it often ignores the rich reasoning pattern relationship information between entities, thereby affecting the accuracy of relationship completion. To this end, the present invention proposes a knowledge graph knowledge expansion method that combines a subgraph knowledge relationship attention completion network with contrast loss learning. First, a subgraph knowledge relationship attention completion network is introduced to comprehensively mine the subgraph relationship information in the manufacturing knowledge graph, and the subgraph knowledge relationship information is integrated to perform relationship prediction and completion, thereby improving the efficiency of knowledge expansion; then, knowledge contrast learning is used to alleviate the problem of data sparsity.

[0055] Specifically, for a given entity pair, the knowledge graph subgraph extraction method based on K-hop neighborhood and link length constraints is first used to extract the subgraph related to the target triple from the original home appliance manufacturing knowledge graph; then the subgraph-based multi-layer relational message passing strategy is used to aggregate the neighborhood relation information in the subgraph to capture the characteristics and type information of the entity; then, the neighborhood relation information is used as a priori guidance, and weights are assigned to different relation links based on the link attention mechanism to obtain the neighborhood relation link representation information of the target entity pair; finally, the two types of relation representation information obtained in the first two steps are fused with the help of the relation fusion method based on the gating mechanism to complete the relationship; in terms of model optimization strategy, the cross entropy loss is combined with the contrast loss, and more negative examples are introduced through the negative sampling method based on the subgraph structure to enhance the feedback of sparse data in view of the data sparsity of the manufacturing scene. This method can effectively utilize the relation information around the entity to alleviate the knowledge incompleteness problem of the manufacturing knowledge graph.

[0056] Specifically, taking home appliance manufacturing as an example, the knowledge graph knowledge expansion method proposed in the present invention is as follows: Figure 1 As shown, the following steps are included:

[0057] S1: A knowledge graph subgraph extraction method based on K-hop neighborhood and link length constraints is used to extract subgraphs related to the target triples from the original knowledge graph.

[0058] The data and knowledge in the home appliance manufacturing knowledge graph are stored as triples (subject, relation, object), where s and o are nodes in the home appliance manufacturing knowledge graph, representing the entity knowledge of the subject and object respectively, and r is an edge, representing the relational knowledge from s to o.

[0059] Assuming that in the knowledge graph, the local subgraph around the target triple contains the logical basis required to infer the relationship between the two nodes, and that the information implying the target relationship exists in the link connecting the two target nodes. Based on this, the present invention extracts the subgraph of the target triple from the original home appliance manufacturing knowledge graph, and limits the link range of the subgraph with the link length as a constraint.

[0060] In the extraction of subgraphs, the breadth-first search (BFS) algorithm is first used to generate candidate subgraphs: starting from the entity pair (s, o), traverse layer by layer until the depth k, where k represents the search depth and is used to control the size of the subgraph and the information coverage. Each candidate subgraph contains nodes on all links connecting the target entity pair (s, o), which are the intersection of the k-hop neighborhood nodes of (s, o) (N k(s) ∩N k(o) ) is composed of. Among them, N k(s) and N k(o) Represents the set of nodes in the k-hop neighborhood of the target node.

[0061] Secondly, the present invention sets a maximum link length threshold L in the BFS algorithm, and adopts a subgraph screening strategy based on link length to reduce noise and control subgraph complexity. If the length of any link between (s, o) exceeds L, all edges connecting them are removed, thereby ensuring that the longest link length between any two nodes in the subgraph does not exceed L. In this way, not only can the prediction efficiency be improved, but also the model's ability to focus on local key information is enhanced, specifically including: (1) defining a link length; the link length usually refers to the shortest path length between nodes, that is, the minimum number of edges required between two nodes; (2) setting a link length threshold L; the link length threshold L is the upper limit of the link length set according to the requirements, for example, only retaining links with a length not exceeding 3; (3) screening subgraphs; screening subgraphs according to the link length threshold L, starting from a certain node, and expanding to all links with a length not exceeding the link length threshold L; (4) constructing subgraphs; combining the screened links and nodes into a new subgraph.

[0062] S2: Utilize a multi-layer relational message passing strategy based on subgraph attention to aggregate the neighborhood relation representation information of target entity pairs.

[0063] The surrounding relationships of entities can reveal the characteristics and category information of entities in the field of home appliance manufacturing, adding important value to relationship completion. Taking the triple (s, r, o) in the home appliance manufacturing knowledge graph as an example, if r represents "purchased from", then the surrounding of s may be associated with relationships such as "parts_use", "parts_included", and the surrounding of o may involve information such as "supplier_location" and "supplier_founder". In view of this, the present invention adopts a multi-layer relationship message passing strategy based on subgraph attention to integrate the neighborhood relationship characteristics of the target entity. This strategy introduces an attention mechanism to dynamically adjust the importance of neighborhood edges to node information aggregation, thereby capturing the most relevant relationship features and partially improving the modeling capabilities of key relationship characteristics.

[0064] The present invention is based on a subgraph attention multi-layer relationship message passing strategy, which includes two steps: message passing and message aggregation. It is a message feature mechanism based on the joint features of neighborhood edges and nodes, and is used to calculate the neighborhood relationship information representation of the target node. In this process, an attention mechanism is introduced, so that the method can dynamically assign weights to the neighborhood edges of each node in an alternating message passing framework, so as to dynamically adjust the importance of neighborhood information, thereby enhancing the modeling capability of key relationship types.

[0065] (1) Message passing.

[0066] The edge weight calculation formula (1) is used to calculate the importance weights of different edges, so that the model can dynamically focus on the most relevant edge information:

[0067] (1)

[0068] Among them, represents the attention weight of edge e to the message of node n in the i-th iteration, represents the hidden state of edge e in the i-th iteration, a is the learnable attention vector, is to transform the hidden state of edge e With the characteristics of node n The vector after connection (concatenation), [·] is the concatenation function; Represents any edge in the neighborhood of the target node n; N(n) is the neighborhood set of node n.

[0069] Formula (2) is used to aggregate neighborhood messages and update node representation:

[0070] (2)

[0071] It represents the message received by node n in the i-th message transmission, which is obtained by aggregating all the connected edges e of node n.

[0072] Formula (3) is used to combine the node message and its own state to update the edge Hidden state:

[0073] (3)

[0074] p and q are the two endpoints of edge e, For edge The set of adjacent points, w i and b i They represent the learnable transformation matrix and bias σ(·) which is a nonlinear activation function; It is the initial feature of edge e and can be used as the one-hot unit vector of the relationship type to which e belongs.

[0075] (2) Message aggregation.

[0076] Assuming that the neighborhood relationship in the above equation is passed k times in the subgraph, based on the above message passing method, it can be calculated that the final messages of the entity pair (s, o) are , aggregate them to get the relation attribute embedding of entity pair (s,o), and use express.

[0077] (4)

[0078] (5)

[0079] (6)

[0080] S3: Using the neighborhood relationship representation information as a priori guidance, weights are assigned to different relationship links based on the link attention mechanism to obtain the neighborhood relationship link representation information of the target entity pair.

[0081] As a bridge between two entities in the home appliance manufacturing knowledge graph, the relationship link is crucial for the relationship completion task. For example, when identifying similar home appliances such as refrigerators, Figure 2 As shown, although the entities disinfection cabinet and microwave oven have the same relational background {contains, belongs to}, the link leading to "refrigerator" shows differences: {(belongs to, belongs to), (contains, contains)} vs {(contains, contains)}. This subtle difference in the link enables the model to infer that "refrigerator" and "disinfection cabinet" belong to the same category, thereby highlighting the key role of relational links in predicting the relationship type between entities. Based on this, the present invention assumes that the relational link between the entity pair (s, o) contains the clues required for prediction, but the entity pair (s, o) often has multiple relational links, and the importance of each link is different, and some of the links are logically not related to the predicted relationship r. In view of this, the present invention introduces the previously determined entity neighborhood relationship information C(s,o) As a priori reference for the importance of relationship links, the contribution of each relationship link is refined according to the neighborhood relationship information.

[0082] In order to effectively capture link information, the present invention introduces a link attention mechanism, which helps the model select the most relevant link among multiple link relationships by dynamically adjusting the importance weight of each link relationship. ) to model and obtain the representation of the link , and represents information C through the relationship with the neighborhood (s, o) The similarity of is used to calculate the attention weight of the link .

[0083] In view of the sequential nature of links, when modeling relational links, the present invention adopts LSTM (Long Short-Term Memory) to learn relational link information.

[0084] (7)

[0085] Relationship links of different relationships correspond to different attention scores. The present invention calculates the corresponding weight for each link according to the subject-object entity relationship information:

[0086] (8) (9)

[0087] in, It is the aggregated representation of the relationship link from entity s to o. (s,o) As the prior information of the link between two entities, it helps to identify the importance of the relationship link.

[0088] S4: The relationship information fusion method based on the gating mechanism is used to fuse the neighborhood relationship representation information and the neighborhood relationship link representation information to generate the prediction result.

[0089] Through the aforementioned steps S12 and S13, the present invention obtains the neighborhood relationship representation of the entity pair (s, o) respectively and neighborhood relationship link representation This step fuses the relationship representations obtained by the previous modules to generate a prediction result, which is the missing relationship, thereby achieving relationship completion.

[0090] In order to more flexibly capture the importance of neighborhood information and relationship links in different situations, the present invention designs a relationship information fusion method based on a gating mechanism to dynamically adjust the contribution of these two types of information. Specifically, by calculating the gating vector g to control the neighborhood relationship representation and link representation The weighted ratio of is used to generate the fused embedding representation.

[0091] Specifically, the gated fusion module receives and Two inputs are weighted (i.e., predicted relationship distribution) through activation functions, and the weighted sum of the two parts is output. The gating mechanism can dynamically adjust the fusion ratio of neighborhood relationship representation and link representation according to different relationship types. It is particularly suitable for sparse knowledge graphs, where some paths may contribute more than other paths. The gating mechanism enhances the expressiveness of the model through this dynamic adjustment. The application model is as follows:

[0092] (10)

[0093] Among them, σ represents the activation function (such as Sigmoid function), Indicates that and Perform vector concatenation, W g and b g are the learnable weight matrix and bias term respectively. Through this formula, we can capture and The dynamic correlation between them and provide a basis for subsequent fusion.

[0094] (11)

[0095] in, is the embedded representation obtained by fusing neighborhood information and link information. ⊙ represents the element-by-element product operation, and (1−g) represents the The formula implements the weight adjustment based on the gate vector and The weighted fusion can dynamically balance the contribution of neighborhood information and link information.

[0096] Generate prediction results based on the probability of relationship r between subject and object entities:

[0097] (12)

[0098] is the projection matrix used to map the joint representation of entity pairs into the relation space.

[0099] Neighborhood relationship information reveals the attribute or category information of the entity itself by capturing the relationship type of the neighboring edges of a given entity pair; while the relationship link describes the relative position of the two entities in the knowledge graph, which can reflect the relationship between the entities in the home appliance manufacturing knowledge graph. In general, neighborhood relationship information and relationship links are the two key factors in knowledge graph relationship completion.

[0100] S5: Introduce the contrast loss learning idea to optimize the model, and increase the number of negative samples based on the negative sampling method of the subgraph structure to enrich the sparse data feedback.

[0101] The present invention introduces the idea of ​​contrastive loss learning for model optimization and combines contrastive loss with cross entropy loss for model training. This takes into account both the prediction accuracy of positive samples and the discrimination ability of negative samples, thereby enhancing the performance of the model in the task of completing the relationship of the home appliance manufacturing knowledge graph.

[0102] Since the knowledge graph of home appliance manufacturing usually contains only a limited number of positive samples, in order to ensure that the model has good generalization ability, a specific technology needs to be used to generate negative samples as a data enhancement strategy during the entire training process to improve the model training effect. That is, the positive samples taken from the training set are recorded as D+, and then a specific negative sampling strategy is used to generate negative samples D-. At present, the most widely used negative sampling method is to randomly select negative samples by replacing the head entity or the tail entity in the triple. However, this random sampling method has many limitations when generating negative samples, such as: poor quality of negative samples, weak adaptability to sparse knowledge graphs, easy generation of overly simple negative samples, failure to fully utilize graph structure information, and the risk of introducing erroneous negative samples. These drawbacks will significantly reduce the training effect of the model in the knowledge graph relationship completion task. In order to solve the above problems, the present invention proposes a negative sampling method based on a subgraph structure to further improve the learning effect and generalization ability of the model.

[0103] For a given positive sample triple, the goal of negative sampling is to generate a negative sample triple that is similar to the positive sample in graph structure but does not represent the true relationship. The specific steps of negative sampling based on subgraph structure are as follows:

[0104] First, for each positive sample (subject, relation, object), the entities and relations in the 3-hop neighborhood associated with it are extracted from the knowledge graph.

[0105] Next, within the 3-hop neighborhood set of the positive sample, entities with similar neighborhood structures are selected to replace the main entity or object entity in the positive sample, and other relations with similar semantics are selected to replace the original relations.

[0106] Then, the generated negative samples are verified for validity, that is, checking whether the negative samples already exist in the positive sample set of the knowledge graph to ensure that they do not belong to the positive sample set.

[0107] Finally, by combining different types of negative sample strategies, multiple negative samples are generated for each positive sample triple, enabling the model to more accurately identify subtle differences in relationships and improve the generalization ability of the model.

[0108] From the training set D T Take out the positive sample and record it as D + , and then use a specific negative sampling strategy to generate negative samples D - .exist On the new data set composed of ) is calculated as:

[0109] (13)

[0110] in is the set of all triplets in the training set; D+ and D - denote positive and negative triplets respectively; is a boundary hyperparameter, which is usually a value greater than 0 and is used to control the minimum gap between positive and negative sample pairs. and Represent the prediction scores of positive and negative example triplets respectively.

[0111] Cross entropy loss is used to measure the difference between the probability distribution predicted by the model and the probability distribution of the true label. The calculation formula of cross entropy loss is as follows:

[0112] (14)

[0113] Among them, T is the number of triple samples in a batch of training, and R represents all relations in KG. Represents the true value label. If the triple (s, r, o) is valid, then =1 if the value is set to 0, otherwise it is 0. is the predicted probability of the model for the triple (s, r, o), which can be calculated by softmax.

[0114] Finally, the model is optimized by combining contrast loss and cross entropy loss, namely:

[0115] (15)

[0116] in, and Used to control the weights between relationship prediction losses. Combining contrast loss and cross entropy loss can improve the feature representation ability and classification boundary distinction while enhancing the generalization ability and robustness of the model, which is especially suitable for tasks that require fine distinction and metric learning.

[0117] Compared with the general knowledge graph relationship completion method, the knowledge graph knowledge expansion method proposed in the above-mentioned present invention completes the relationship by fusing the knowledge relationship information of the subgraph, and extracts the local subgraph node set containing the target triplet for model training based on the k-hop neighborhood and path length constraints, thereby effectively avoiding the problem of resource loss; it uses the local subgraph structure information around the target entity to complete the relationship without introducing additional new information, thereby avoiding the high cost of manual labeling and over-reliance on expert knowledge, and significantly improving the accuracy of relationship completion.

[0118] There are many implementation methods of the present invention, and all technical solutions formed by equivalent transformation or equivalent transformation fall within the protection scope of the present invention.

[0119] It should be pointed out that the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by ordinary technicians in this technical field within the essential scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A knowledge graph knowledge expansion method, characterized in that: include: S1: A knowledge graph subgraph extraction method based on K-hop neighborhood and link length constraints extracts subgraphs related to the target triples from the original knowledge graph; S2: Utilize a multi-layer relational message passing strategy based on subgraph attention to aggregate the neighborhood relation representation information of target entity pairs; S3: Using the neighborhood relationship representation information as a priori guidance, weights are assigned to different relationship links based on the link attention mechanism to obtain the neighborhood relationship link representation information of the target entity pair; S4: A relationship information fusion method based on a gating mechanism is used to fuse the neighborhood relationship representation information and the neighborhood relationship link representation information to generate a prediction result; S5: Introduce the contrast loss learning idea to optimize the model, and increase the number of negative samples based on the negative sampling method of the subgraph structure to enrich the sparse data feedback.

2. The knowledge graph knowledge expansion method according to claim 1, characterized in that: S1 specifically includes: Generate candidate subgraphs using a breadth-first search algorithm; and, A maximum link length threshold is set in the algorithm. When the length of any link in the search exceeds the maximum link length threshold, all connected edges are removed.

3. The knowledge graph knowledge expansion method according to claim 1, characterized in that: The multi-layer relational message passing strategy based on subgraph attention in S2 includes message passing and message aggregation; among them, Messaging includes: The edge weight calculation formula is used to calculate the importance weights of different edges: ; use Aggregate neighborhood messages and update node representations; use Combine node messages and their own status to update the edge Hidden state; Above, use represents the attention weight of edge e to the message of node n in the i-th iteration, represents the hidden state of edge e in the i-th iteration, represents the message received by node n in the i-th message transmission, obtained by aggregating all the connected edges e of node n; a is a learnable attention vector, is to transform the hidden state of edge e With the characteristics of node n The concatenated vector, represents any edge in the neighborhood of the target node n; N(n) is the neighborhood set of node n; p and q are the two endpoints of edge e, is the set of adjacent points of edge e, [·] is the splicing function, w i and b i They represent the learnable transformation matrix and bias σ(·) which is a nonlinear activation function; is the initial feature of edge e, which is the one-hot unit vector of the relation type to which e belongs; Message aggregation includes: based on and Calculate the final message of the entity pair (s,o); Aggregate the final message to get the relation attribute embedding of the entity pair (s,o): ; Among them, k is the number of times the neighborhood relationship is transmitted in the subgraph.

4. The knowledge graph knowledge expansion method according to claim 1, characterized in that: S3 specifically includes: For each relationship link P( ) to model and obtain the representation of the relationship link ; Represent information C through relation links and neighborhood relations (s, o) The similarity of is used to calculate the attention weight of the link ;in, It is the aggregate representation of the relationship link from entity s to o.

5. The knowledge graph knowledge expansion method according to claim 1, characterized in that: The relationship information fusion method based on the gating mechanism in S4 specifically includes: Gated fusion module receives and Two inputs; is the neighborhood relationship representation of the entity pair (s,o), is the neighborhood relationship link representation of the entity pair (s,o); Indicates that and Perform vector stitching; based on Calculate the gate vector It is used to control the weighted ratio of domain relationship representation and neighborhood relationship link representation; where σ represents the activation function, W g and b g They are respectively the learnable weight matrix and bias term; based on The neighborhood relationship representation information and the neighborhood relationship link representation information are fused to obtain the embedded representation; ⊙ represents the element-by-element product operation, (1−g) represents the weight adjustment.

6. The knowledge graph knowledge expansion method according to claim 5 is characterized in that: S4 generates prediction results based on the probability of the relationship r between the subject and the object entities: ;in, is the projection matrix used to map the joint representation of entity pairs into the relation space.

7. The knowledge graph knowledge expansion method according to claim 1, characterized in that: In S5, contrast loss and cross entropy loss are combined for model training: ; and is the weight used to control the relationship prediction loss; is the contrast loss, is the cross entropy loss; The contrast loss is calculated as: ;D + and D - denote positive and negative triplets respectively; is a boundary hyperparameter used to control the minimum gap between positive and negative sample pairs, and Represent the prediction scores of positive and negative example triplets respectively; The cross entropy loss is calculated as: ; T is the number of triple samples in a batch of training, R represents all the relations in the knowledge graph, Represents the true value label. If the triple (s, r, o) is valid, then =1, otherwise 0; is the model’s predicted probability for the triple (s, r, o).

8. The knowledge graph knowledge expansion method according to claim 1, characterized in that: The negative sampling methods based on subgraph structure in S5 include: For each positive sample, extract the entities and relations in the n-hop neighborhood associated with it from the knowledge graph; In the n-hop neighborhood set of the positive sample, select entities with similar neighborhood structures to replace the main entity or guest entity in the positive sample, and select other relations with similar semantics to replace the original relations; Verify the effectiveness of the generated negative samples; Multiple negative samples are generated for each positive sample triple by combining different types of negative samples.

Citation Information

Patent Citations

  • Knowledge graph completion method based on multi-granularity hierarchy and dynamic embedding

    CN116842199A

  • Small sample knowledge graph completion method based on embedding fusion and data enhancement

    CN119537600A

  • Knowledge graph completion method and device for relieving sparsity

    CN119623598A

  • Induction link prediction method based on relation message passing and interaction information maximization

    CN119721229A

  • Relation-enhancement knowledge graph embedding method and system

    US20230297553A1

Cited By

  • Knowledge graph entity completion method

    CN120851176A

  • A knowledge graph entity completion method

    CN120851176B

  • Microbial secondary metabolite retrieval method based on knowledge graph

    CN121478844A

  • Context self-adaptive knowledge graph question and answer method, equipment and medium

    CN122412559A