A method for expanding knowledge in a knowledge graph

By applying sub-graph knowledge relational attention completion network and contrast loss learning technology in the home appliance manufacturing knowledge graph, the problems of relationship omissions and data in the knowledge graph are solved, and efficient expansion of the knowledge graph and the accuracy of relationship prediction are achieved.

CN119917701BActive Publication Date: 2025-06-13OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510421139.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-13
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing home appliance manufacturing knowledge graph faces the problems of relationship omissions and incomplete data, resulting in poor application results in scenarios such as manufacturing Q&A, fault analysis, quality analysis and personalized recommendations.

Method used

A knowledge graph knowledge completion method based on sub-graph knowledge relational attention completion network combined with contrast loss learning technology is proposed. The prediction results are generated through the integration of K-hop neighborhood and sub-graph extraction of link length constraints, multi-layer relationship message delivery, link attention mechanism and gating mechanism, and the prediction results are generated, and the combination of comparison loss and cross-entropy loss is optimized.

Benefits of technology

It significantly improves the knowledge expansion efficiency of home appliance manufacturing knowledge graph, enhances the accuracy of relationship prediction, reduces manual labeling costs, reduces dependence on expert knowledge, and effectively solves the problems of data sparseness and relationship omissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917701B_ABST
    Figure CN119917701B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for knowledge expansion of a knowledge graph. By using a method for extracting a subgraph of the knowledge graph based on K-hop neighborhood and link length constraints to extract the subgraph of the target entity, resource consumption is reduced. Subsequently, the attention mechanism is used to extract neighborhood relationship representation information and neighborhood relationship link representation information, and a relationship fusion method based on a gating mechanism is used to fuse these two relationship representation information, and then relationship prediction and completion are carried out. Finally, the model is optimized by combining contrastive loss and cross-entropy loss, and more negative samples are added through a negative sampling method based on the subgraph structure, enriching the feedback on sparse data and making up for the deficiency of the cross-entropy loss in generalization performance, realizing the effective expansion of the knowledge graph. The present invention improves the relationship prediction accuracy of the knowledge graph, effectively alleviates the data sparsity problem of the knowledge graph, significantly reduces the manual annotation cost, and reduces the over-reliance on expert knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of knowledge management, and specifically, relates to a method for knowledge expansion of a knowledge graph. Background Art

[0002] Intelligent manufacturing is one of the key measures to enhance the core competitiveness of the manufacturing industry. At present, although online monitoring based on production big data and prediction has been applied in manufacturing scenarios, in scenarios that require complex decision-making, the degree of intelligence is still insufficient. Most complex problems still rely on experts to solve manually, which not only requires a large amount of experience and knowledge, but also has a high cost.

[0003] A high-quality knowledge graph can comprehensively store the knowledge precipitation and accumulation formed during the long-term production, operation, and innovation activities of manufacturing enterprises, helping decision-makers correctly analyze enterprise data, so as to timely discover and solve problems in enterprise production management and predict industry development trends and opportunities. The home appliance industry is an important part of the manufacturing industry. Applying the knowledge graph to the home appliance manufacturing field can not only process, organize, and manage the accumulated experience and case knowledge to form a knowledge base, but also perform knowledge reasoning on similar or even repetitive decision-making activities during the manufacturing process to discover new knowledge, so as to realize the accumulation, inheritance, and reuse of manufacturing knowledge.

[0004] However, due to the limitations of annotation resources and technologies, as well as the high cost of manually searching for all fact triples, almost all current knowledge graphs are facing the challenges of missing relationships and incomplete data. An incomplete home appliance manufacturing knowledge graph is difficult to meet the application requirements of scenarios such as manufacturing Q&A, fault analysis, quality analysis, and personalized recommendation. The relationship completion technology aims to fill in the missing connections between entities, and is an effective strategy to solve the incompleteness of the knowledge graph, expand the knowledge of the knowledge graph, and promote knowledge growth. However, the home appliance manufacturing scenario is complex and changeable. The existing relationship completion models are difficult to accurately adapt to the situations of missing relationships and incomplete data in the home appliance manufacturing field, and often ignore the rich reasoning patterns between entities, affecting the accuracy of relationship completion. Summary of the Invention

[0005] In view of the problems of missing relationships and incomplete data existing in the application of the knowledge graph in the current manufacturing field, the present invention proposes a knowledge graph knowledge completion method based on a subgraph knowledge relationship attention completion network combined with contrastive loss learning technology, which improves the knowledge expansion efficiency of the home appliance manufacturing knowledge graph, expands its coverage, solves the problems of missing relationships and incomplete data in the existing home appliance manufacturing knowledge graph, and can significantly reduce the manual annotation cost and reduce the over-reliance on expert knowledge.

[0006] The present invention is implemented by adopting the following technical solutions:

[0007] A knowledge graph knowledge expansion method is proposed, which is characterized by including:

[0008] S1: A knowledge graph subgraph extraction method based on K-hop neighborhood and link length constraints is used to extract a subgraph related to the target triple from the original knowledge graph;

[0009] S2: A multi-layer relational message passing strategy based on subgraph attention is used to aggregate the neighborhood relationship representation information of the target entity pair;

[0010] S3: Using the neighborhood relationship representation information as a prior guidance, weights are assigned to different relationship links based on the link attention mechanism to obtain the neighborhood relationship link representation information of the target entity pair;

[0011] S4: A relational information fusion method based on a gating mechanism is used to fuse the neighborhood relationship representation information and the neighborhood relationship link representation information to generate a prediction result;

[0012] S5: The contrastive loss learning idea is introduced to optimize the model, and the negative sample quantity is increased based on the subgraph structure negative sampling method to enrich the sparse data feedback.

[0013] In some embodiments of the present invention, S1 specifically includes:

[0014] The breadth-first search algorithm is used to generate candidate subgraphs; and,

[0015] A maximum link length threshold is set in the algorithm. When any link length in the search exceeds the maximum link length threshold, all connected edges are removed.

[0016] In some embodiments of the present invention, the multi-layer relational message passing strategy based on subgraph attention in S2 includes message passing and message aggregation; among them,

[0017] Message passing includes:

[0018] The edge weight calculation formula is used to calculate the importance weights of different edges:

[0019] ;

[0020] Adopt Aggregate neighborhood messages and update the node representation;

[0021] Adopt Combine the node message and its own state to update the edge Hidden state;

[0022] Above, use To represent the attention weight of edge e to the node n message in the i-th iteration, To represent the hidden state of edge e in the i-th iteration, Denote the message received by node n in the i-th message passing, obtained by aggregating all connected edges r of node n; a is a learnable attention vector. is the hidden state of edge e connected (concatenated) with the feature of node n to form a vector. Denote any edge in the neighborhood of the target node n; N(n) is the neighborhood set of node n; p and q are the two endpoints of edge e. is the edge and the set of adjacent nodes of [·] is the concatenation function, w and b i represent the learnable transformation matrix and bias respectively. σ(·) is a non-linear activation function. is the initial feature of edge e, which can be used as a one-hot unit vector of the relationship type to which e belongs.

[0023] Message aggregation includes:

[0024] Based on and calculate the final message of the entity pair (s, o).

[0025] Aggregate the final message to obtain the relationship attribute embedding of the entity pair (s, o):

[0026] ;

[0027] where k is the number of times the neighborhood relationship is passed in the subgraph.

[0028] In some embodiments of the present invention, S3 specifically includes:

[0029] Model each relationship link P( ) to obtain the representation of the relationship link ;

[0030] Calculate the attention weight of the link through the similarity between the relationship link and the neighborhood relationship representation information C (s, o) ; where is the aggregated representation of the relationship link from entity s to o.

[0031] In some embodiments of the present invention, the relationship information fusion method based on the gating mechanism in S4 specifically includes:

[0032] The gating fusion module receives and as two inputs; is the neighborhood relationship representation of the entity pair (s, o), is the neighborhood relationship link representation of the entity pair (s, o); It means to and perform vector concatenation;

[0033] Based on calculate the gating vector to control the weighted ratio of the domain relationship representation and the neighborhood relationship link representation; where, σ represents the activation function, W g and b g are the learnable weight matrix and bias term respectively;

[0034] Based on fuse the neighborhood relationship representation information and the neighborhood relationship link representation information to obtain the embedding representation; ⊙ represents the element-wise product operation, and (1−g) represents the weight adjustment of .

[0035] In some embodiments of the present invention, in S4, a prediction result is generated based on the probability that there is a relationship r between the subject and object entities:

[0036] , where is the projection matrix used to map the joint representation of the entity pair to the relationship space.

[0037] In some embodiments of the present invention, in S5, the contrastive loss and the cross-entropy loss are combined for model training:

[0038] ; and are the weights used to control the relationship prediction losses; is the contrastive loss, is the cross-entropy loss;

[0039] The contrastive loss is calculated as:

[0040] ; is the set of all triples in the training set; D + and D - represent positive and negative triples respectively; γ is the margin hyperparameter used to control the minimum gap between positive and negative sample pairs, and represent the prediction scores of positive and negative example triples respectively;

[0041] The cross-entropy loss is calculated as:

[0042] ; T is the number of triple samples in a batch of training, R represents all relationships in the knowledge graph, represents the true label. If the triple (s,r,o) is valid, then = 1, otherwise 0; is the predicted probability of the model for the triple (s, r, o).

[0043] In some embodiments of the present invention, the negative sampling method based on the subgraph structure in S5 includes:

[0044] For each positive sample, extract the entities and relationships within the n-hop neighborhood associated with it from the knowledge graph;

[0045] Within the n-hop neighborhood set of the positive sample, select entities with similar neighborhood structures to replace the main entity or the object entity in the positive sample, and select other semantically similar relationships to replace the original relationship;

[0046] Verify the validity of the generated negative samples;

[0047] Generate multiple negative samples for each positive sample triple by combining different types of negative samples.

[0048] Compared with the prior art, the advantages and positive effects of the present invention are as follows: The knowledge expansion method of the knowledge graph proposed by the present invention first extracts the subgraph of the target entity through the knowledge graph subgraph extraction method based on K-hop neighborhood and link length constraints to reduce resource consumption. Subsequently, the attention mechanism is used to extract the neighborhood relationship representation information and the neighborhood relationship link representation information, and the relationship fusion method based on the gating mechanism is used to fuse these two relationship representation information, and then relationship prediction and completion are carried out. Finally, the model is optimized by combining the contrast loss and the cross-entropy loss, and more negative samples are added by introducing the negative sampling method based on the subgraph structure, enriching the feedback on sparse data and making up for the deficiency of the cross-entropy loss in generalization performance, realizing the effective expansion of the knowledge graph. Compared with the traditional method, the present invention uses the subgraph relationship representation information and the contrast loss learning technology, improves the relationship prediction accuracy of the knowledge graph, effectively alleviates the data sparsity problem of the knowledge graph, and can also effectively solve the problems of high cost of manual annotation and excessive dependence on expert knowledge.

[0049] After reading the detailed description of the embodiments of the present invention in conjunction with the accompanying drawings, other features and advantages of the present invention will become clearer. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0051] Figure 1 Schematic diagram of the execution steps of the knowledge expansion method of the knowledge graph proposed by the present invention;

[0052] Figure 2 This is an example of the relationship information in the home appliance manufacturing knowledge graph in the embodiments of the present invention. Detailed implementation manners

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0054] For manufacturing enterprises, constructing a complete home appliance knowledge graph is crucial for realizing the mutual coordination of knowledge management and application in the entire life cycle of home appliance manufacturing. In response to the problems of missing relationships and incomplete data in the graph, the industry usually applies knowledge graph relationship completion technology to expand the knowledge of the manufacturing knowledge graph. However, when the existing relationship completion technology is applied to the data-sparse manufacturing field, it often ignores the rich inference pattern relationship information between entities, thus affecting the accuracy of relationship completion. For this reason, the present invention proposes a knowledge graph knowledge expansion method that combines a subgraph knowledge relationship attention completion network and contrastive loss learning. First, a subgraph knowledge relationship attention completion network is introduced to comprehensively mine the subgraph relationship information in the manufacturing knowledge graph, and the subgraph knowledge relationship information is fused for relationship prediction and completion, thereby improving the efficiency of knowledge expansion; then, knowledge contrast learning is used to alleviate the data sparsity problem.

[0055] Specifically, for a given entity pair, first, a knowledge graph subgraph extraction method based on K-hop neighborhood and link length constraint is used to extract the subgraph related to the target triple from the original home appliance manufacturing knowledge graph; then, a multi-layer relationship message passing strategy based on subgraph attention is used to aggregate the neighborhood relationship information in the subgraph to capture the characteristics and type information of the entity; again, using the neighborhood relationship information as a priori guidance, based on the link attention mechanism, weights are assigned to different relationship links to obtain the neighborhood relationship link representation information of the target entity pair; finally, a relationship fusion method based on a gating mechanism is used to fuse the two types of relationship representation information obtained in the previous two steps to complete the relationship; in the model optimization strategy, the cross-entropy loss and the contrastive loss are combined. In response to the data sparsity of the manufacturing scenario, more negative examples are introduced through a negative sampling method based on the subgraph structure to enhance the feedback on sparse data. This method can effectively utilize the relationship information around the entity to alleviate the knowledge incompleteness problem of the manufacturing knowledge graph.

[0056] Specifically, taking home appliance manufacturing as an example, the knowledge graph knowledge expansion method proposed by the present invention, as Figure 1 shown, includes the following steps:

[0057] S1: A method for extracting a subgraph of a knowledge graph based on K-hop neighborhood and link length constraints, which extracts a subgraph related to a target triple from the original knowledge graph.

[0058] In the home appliance manufacturing knowledge graph, data and knowledge are stored as triples (subject, relation, object), where s and o are nodes in the home appliance manufacturing knowledge graph, representing the entity knowledge of the subject and object respectively, and r is the edge, representing the relational knowledge of the entity pointing from s to o.

[0059] It is assumed that in the knowledge graph, the local subgraph around the target triple contains the logical basis for inferring the relationship between the two nodes, and moreover, the information suggesting the target relationship exists in the link connecting the two target nodes. Accordingly, the present invention extracts the subgraph of the target triple from the original home appliance manufacturing knowledge graph and limits the link range of the subgraph with the link length as a constraint.

[0060] In the extraction of the subgraph, first, the breadth-first search (BFS) algorithm is used to generate candidate subgraphs: starting from the entity pair (s, o), traversing layer by layer until the depth k, where k represents the search depth and is used to control the size and information coverage range of the subgraph. Each candidate subgraph contains the nodes on all the links connecting the target entity pair (s, o), and is composed of the intersection (N k(s) ∩N k(o) ) of the k-hop neighborhood nodes of (s, o). Among them, N k(s) and N k(o) represent the node sets of the target nodes within the k-hop neighborhood.

[0061] Secondly, the present invention sets a maximum link length threshold L in the BFS algorithm and adopts a subgraph screening strategy based on link length to reduce noise and control the subgraph complexity. If the length of any link between (s, o) exceeds L, then all the edges connecting them are removed, so as to ensure that the longest link length between any two nodes in the subgraph does not exceed L. In this way, not only can the prediction efficiency be improved, but also the focusing ability of the model on local key information can be enhanced. Specifically, it includes: (1) defining a link length; this link length generally refers to the shortest path length between nodes, that is, the minimum number of edges required between two nodes; (2) setting a link length threshold L; this link length threshold L is the upper limit of the link length set according to requirements, for example, only retaining links with a length not exceeding 3; (3) screening the subgraph; screening the subgraph according to the link length threshold L, starting from a certain node and expanding to all links with a length not exceeding the link length threshold L; (4) constructing the subgraph; combining the screened links and nodes into a new subgraph.

[0062] S2: Using a multi-layer relational message passing strategy based on subgraph attention to aggregate the neighborhood relationship representation information of the target entity pair.

[0063] The peripheral relationships of entities can reveal the characteristics and category information of entities in the field of home appliance manufacturing, adding important value to relationship completion. Taking the triple (s, r, o) in the home appliance manufacturing knowledge graph as an example, if r represents "purchased from", then the periphery of s may be associated with relationships such as "component_usage" and "component_included", while the periphery of o may involve information such as "supplier_location" and "supplier_founders". In view of this, the present invention adopts a multi-layer relationship message passing strategy based on subgraph attention to integrate the neighborhood relationship features of target entities. This strategy introduces an attention mechanism to dynamically adjust the importance of neighborhood edges for node information aggregation, thereby capturing the most relevant relationship features and further enhancing the modeling ability for key relationship intensities.

[0064] The multi-layer relationship message passing strategy based on subgraph attention of the present invention includes two steps: message passing and message aggregation. It is a message feature mechanism based on the joint features of neighborhood edges and nodes, used to calculate the neighborhood relationship information representation of target nodes. In this process, an attention mechanism is introduced, enabling the method to dynamically assign weights to the neighborhood edges of each node in the alternating message passing framework, thereby dynamically adjusting the importance of neighborhood information and enhancing the modeling ability for key relationship types.

[0065] (1) Message passing.

[0066] The edge weight calculation formula (1) is used to calculate the importance weights of different edges, enabling the model to dynamically focus on the most relevant edge information:

[0067] (1)

[0068] Among them, represents the attention weight of edge e on the message of node n in the i-th iteration, represents the hidden state of edge e in the i-th iteration, a is a learnable attention vector, is the vector obtained by connecting (concatenating) the hidden state of edge e with the feature of node n, [·] is the concatenation function; represents any edge in the neighborhood of target node n; N(n) is the neighborhood set of node n.

[0069] The formula (2) is used to aggregate neighborhood messages and update the node representation:

[0070] (2)

[0071] represents the message received by node n in the i-th message passing, obtained by aggregating all the connected edges e of node n.

[0072] Update the edge by combining the node message and its own state using formula (3) Hidden state:

[0073] (3)

[0074] p and q are the two endpoints of edge e, is the set of adjacent nodes of the edge w i and b i represent the learnable transformation matrix and bias respectively. σ(·) is a non-linear activation function; is the initial feature of edge e and can be used as a one-hot unit vector of the relationship type to which e belongs.

[0075] (2) Message aggregation.

[0076] Assume that the neighborhood relationship in the above equation is passed k times in the subgraph. Based on the above message passing method, the final messages of the entity pair (s, o) can be calculated as . Aggregate them to obtain the relationship attribute embedding of the entity pair (s, o), denoted by .

[0077] (4)

[0078] (5)

[0079] (6)

[0080] S3: Using the neighborhood relationship representation information as prior guidance, assign weights to different relationship links based on the link attention mechanism to obtain the neighborhood relationship link representation information of the target entity pair.

[0081] Relationship links, as the bridge between two entities in the home appliance manufacturing knowledge graph, are crucial for the relationship completion task. For example, when identifying similar home appliances to a refrigerator, as Figure 2 shown, although the entities disinfection cabinet and microwave oven have the same relationship background {contain, belong to}, the links leading to "refrigerator" show differences: {(belong to, belong to), (contain, contain)} vs {(contain, contain)}. This subtle difference in the links enables the model to infer that "refrigerator" and "disinfection cabinet" belong to the same category, thus highlighting the key role of relationship links in predicting the relationship type between entities. Based on this, the present invention assumes that the relationship links between entity pairs (s, o) contain clues required for prediction. However, there are often multiple relationship links between entity pairs (s, o), and the importance of each link varies, and some of the links are not logically related to the predicted relationship r. In view of this, the present invention introduces the previously determined entity neighborhood relationship information C(s,o) As a prior reference for the importance of relationship links, that is, to refine the contribution of each relationship link according to neighborhood relationship information.

[0082] To effectively capture link information, the present invention introduces a link attention mechanism, which helps the model select the most relevant one among multiple relationship links by dynamically adjusting the importance weights of each relationship link. In the model, for each relationship link P( ) is modeled to obtain the representation of the link , and the attention weight of the link is calculated through the similarity with the neighborhood relationship representation information C (s, o) . .

[0083] In view of the sequential nature of the link, when modeling the relationship link, the present invention uses LSTM (Long Short-Term Memory) to learn the relationship link information.

[0084] (7)

[0085] Relationship links of different relationships correspond to different attention scores. The present invention calculates the corresponding weights for each link according to the subject-object entity relationship information:

[0086] (8)

[0087] (9)

[0088] Among them, is the aggregated representation of the relationship link from entity s to o. The neighborhood relationship information C (s,o) is used as the prior information of the link between two entities to help identify the importance degree of the relationship link.

[0089] S4: Adopt a relationship information fusion method based on a gating mechanism to fuse the neighborhood relationship representation information and the neighborhood relationship link representation information to generate a prediction result.

[0090] Through the foregoing steps S12 and S13, the present invention respectively obtains the neighborhood relationship representation and the neighborhood relationship link representation of the entity pair (s, o). This step fuses the relationship representations obtained by the previous modules to generate a prediction result, which is the missing relationship, thereby realizing relationship completion.

[0091] To more flexibly capture the importance of neighborhood information and relationship links in different situations, the present invention designs a relationship information fusion method based on a gating mechanism for dynamically adjusting the contributions of these two types of information. Specifically, by calculating the gating vector g to control the neighborhood relationship representation And link representation The weighted ratio of, thus generating the fused embedding representation.

[0092] Specifically, the gating fusion module receives and Two inputs, generate weights (i.e., predict the relationship distribution) through the activation function, and perform weighted sum on the two parts to obtain the output. The gating mechanism can dynamically adjust the fusion ratio of the neighborhood relationship representation and the link representation according to different relationship types, and is especially suitable for sparse knowledge graphs, where the contributions of some paths may be greater than those of other paths. The gating mechanism enhances the expressive power of the model through this dynamic adjustment. The application model is as follows:

[0093] (10)

[0094] Among them, σ represents the activation function (such as the Sigmoid function), Represents concatenating and Vector concatenation, W g and b g Are the learnable weight matrix and bias term respectively. Through this formula, the dynamic correlation between and Can be captured, and a basis for subsequent fusion is provided.

[0095] (11)

[0096] Among them, Is the embedding representation obtained by fusing neighborhood information and link information. ⊙ represents element-wise product operation, and (1−g) represents the weight adjustment of . This formula realizes the weighted fusion of and Based on the gating vector, and can dynamically balance the contributions of neighborhood information and link information.

[0097] Generate the prediction result based on the probability that there is a relationship r between the subject and object entities:

[0098] (12)

[0099] Is the projection matrix used to map the joint representation of the entity pair to the relationship space.

[0100] The neighborhood relationship information reveals the attribute or category information of the entity itself by capturing the relationship types of the adjacent edges of the given entity pair; while the relationship link depicts the relative positions of the two entities in the knowledge graph, and can reflect the relationships between the entities in the household appliance manufacturing knowledge graph. Generally speaking, the neighborhood relationship information and the relationship link are two key factors in knowledge graph relationship completion.

[0101] S5: Introduce the idea of contrastive loss learning to optimize the model, and increase the number of negative samples based on the subgraph structure-based negative sampling method to enrich the sparse data feedback.

[0102] The present invention introduces the idea of contrastive loss learning for model optimization, combines contrastive loss with cross-entropy loss for model training, which not only considers the prediction accuracy of positive samples but also the discrimination ability of negative samples, thereby enhancing the performance of the model in the task of relationship completion of the home appliance manufacturing knowledge graph.

[0103] Since the home appliance manufacturing knowledge graph usually only contains a limited number of positive samples, to ensure that the model has good generalization ability, during the entire training process, specific techniques need to be adopted to generate negative samples as a data augmentation strategy to improve the model training effect. That is: the positive samples taken from the training set are denoted as D+, and subsequently, a specific negative sampling strategy is adopted to generate negative samples D-. Currently, the most widely used negative sampling method is to randomly select negative samples by replacing the head entity or tail entity in the triple. However, this random sampling method has many limitations when generating negative samples, such as: poor quality of negative samples, weak adaptability to sparse knowledge graphs, easy to generate overly simple negative samples, failure to fully utilize graph structure information, and the risk of introducing incorrect negative samples, etc. These drawbacks will significantly reduce the training effectiveness of the model in the task of knowledge graph relationship completion. To solve the above problems, the present invention proposes a subgraph structure-based negative sampling method to further improve the learning effect and generalization ability of the model.

[0104] For a given positive sample triple, the goal of negative sampling is to generate negative sample triples that are similar to the positive sample in the graph structure but do not represent real relationships. The specific steps of subgraph structure-based negative sampling are as follows:

[0105] First, for each positive sample (subject, relation, object), extract the entities and relationships within its 3-hop neighborhood associated in the knowledge graph.

[0106] Next, within the 3-hop neighborhood set of the positive sample, select entities with similar neighborhood structures to replace the main entity or the guest entity in the positive sample, and select other relationships with similar semantics to replace the original relationship.

[0107] Then, verify the validity of the generated negative samples, that is, check whether the negative samples already exist in the positive sample set of the knowledge graph to ensure that they do not belong to the positive sample set.

[0108] Finally, by combining different types of negative sample strategies, generate multiple negative samples for each positive sample triple, enabling the model to more accurately identify the subtle differences in relationships and improve the generalization ability of the model.

[0109] Extract the positive samples from the training set D T and denote them as D + . Subsequently, use a specific negative sampling strategy to generate negative samples D - . On the new dataset formed by , the contrastive loss function ( ) is calculated as:

[0110] (13)

[0111] where is the set of all triples in the training set; D+ and D - represent positive and negative triples respectively; is the margin hyperparameter, which is usually a value greater than 0 and is used to control the minimum gap between positive and negative sample pairs. and represent the predicted scores of positive and negative example triples respectively.

[0112] The cross-entropy loss is used to measure the difference between the probability distribution predicted by the model and the probability distribution of the true labels. The calculation formula of the cross-entropy loss is as follows:

[0113] (14)

[0114] where T is the number of triple samples in a batch of training, and R represents all relationships in the KG. represents the true label. If the triple (s, r, o) is valid, then = 1, otherwise 0. is the predicted probability of the model for the triple (s, r, o), which can be obtained by softmax calculation.

[0115] Finally, optimize the model by combining the contrastive loss and the cross-entropy loss, that is:

[0116] (15)

[0117] where and are used to control the weights between the relationship prediction losses. Combining the contrastive loss and the cross-entropy loss can improve the feature representation ability and the classification boundary discrimination while enhancing the generalization ability and robustness of the model, especially suitable for tasks that require fine discrimination and metric learning.

[0118] The knowledge graph knowledge expansion method proposed by the present invention, compared with the general knowledge graph relation completion method, performs relation completion by integrating the knowledge relation information of the subgraph. By extracting the local subgraph node set containing the target triple for model training based on k-hop neighborhood and path length constraints, the problem of resource consumption is effectively avoided; relation completion is performed using the local subgraph structure information around the target entity without introducing new information additionally, avoiding the problems of high manual annotation cost and over-reliance on expert knowledge, and significantly improving the accuracy of relation completion.

[0119] There are still various implementation manners of the present invention. All technical solutions formed by equivalent transformation or equivalent change fall within the protection scope of the present invention.

[0120] It should be noted that the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by those of ordinary skill in the art within the substantial scope of the present invention should also belong to the protection scope of the present invention.

Claims

1. A method for expanding knowledge of household appliance manufacturing knowledge graph, characterized in that: include: S1: A knowledge graph subgraph extraction method based on K-hop neighborhood and link length constraints is used to extract subgraphs related to the target triples from the original knowledge graph of home appliance manufacturing; S2: Utilize the subgraph attention-based multi-layer relational message passing strategy to aggregate the neighborhood relation representation information of home appliance manufacturing target entity pairs; S3: Using the neighborhood relationship representation information as a priori guidance, weights are assigned to different relationship links based on the link attention mechanism to obtain the neighborhood relationship link representation information of the home appliance manufacturing target entity pair; S4: Using a gating mechanism-based relationship information fusion method to fuse neighborhood relationship representation information and neighborhood relationship link representation information to generate a prediction result; the prediction result is a relationship missing from the home appliance manufacturing knowledge graph; S5: Introduce the contrast loss learning idea to optimize the model, and increase the number of negative samples based on the negative sampling method of the subgraph structure to enrich the sparse data feedback.

2. The method for expanding knowledge of household appliance manufacturing knowledge graph according to claim 1, characterized in that: S1 specifically includes: Generate candidate subgraphs using a breadth-first search algorithm; and, A maximum link length threshold is set in the algorithm. When the length of any link in the search exceeds the maximum link length threshold, all connected edges are removed.

3. The method for expanding knowledge of household appliance manufacturing knowledge graph according to claim 1, characterized in that: The multi-layer relational message passing strategy based on subgraph attention in S2 includes message passing and message aggregation; among them, Messaging includes: The edge weight calculation formula is used to calculate the importance weights of different edges: ; use Aggregate neighborhood messages and update node representations; use Combine node messages and their own status to update the edge Hidden state; Above, use represents the attention weight of edge e to the message of node n in the i-th iteration, represents the hidden state of edge e in the i-th iteration, represents the message received by node n in the i-th message transmission, obtained by aggregating all the connected edges e of node n; a is a learnable attention vector, is to transform the hidden state of edge e With the characteristics of node n The concatenated vector, represents any edge in the neighborhood of the target node n; N(n) is the neighborhood set of node n; p and q are the two endpoints of edge e, is the set of adjacent points of edge e, [·] is the splicing function, w i and b i They represent the learnable transformation matrix and bias σ(·) which is a nonlinear activation function; is the initial feature of edge e, which is the one-hot unit vector of the relation type to which e belongs; Message aggregation includes: based on and Calculate the final message of the entity pair (s,o); Aggregate the final message to get the relation attribute embedding of the entity pair (s,o): ; Among them, k is the number of times the neighborhood relationship is transmitted in the subgraph.

4. The method for expanding knowledge of household appliance manufacturing knowledge graph according to claim 1, characterized in that: S3 specifically includes: For each relationship link P( ) to model and obtain the representation of the relationship link ; Represent information C through relation links and neighborhood relations (s, o) The similarity of is used to calculate the attention weight of the link ;in, It is the aggregate representation of the relationship link from entity s to o.

5. The method for expanding knowledge of household appliance manufacturing knowledge graph according to claim 1, characterized in that: The relationship information fusion method based on the gating mechanism in S4 specifically includes: Gated fusion module receives and Two inputs; is the neighborhood relationship representation of the entity pair (s,o), is the neighborhood relationship link representation of the entity pair (s,o); Indicates that and Perform vector stitching; based on Calculate the gate vector It is used to control the weighted ratio of domain relationship representation and neighborhood relationship link representation; where σ represents the activation function, W g and b g They are respectively the learnable weight matrix and bias term; based on The neighborhood relationship representation information and the neighborhood relationship link representation information are fused to obtain the embedded representation; ⊙ represents the element-by-element product operation, (1−g) represents the weight adjustment.

6. The method for expanding knowledge of household appliance manufacturing knowledge graph according to claim 5, characterized in that: S4 generates prediction results based on the probability of the relationship r between the subject and the object entities: ;in, is the projection matrix used to map the joint representation of entity pairs into the relation space.

7. The method for expanding knowledge of household appliance manufacturing knowledge graph according to claim 1, characterized in that: In S5, contrast loss and cross entropy loss are combined for model training: ; and is the weight used to control the relationship prediction loss; is the contrast loss, is the cross entropy loss; The contrast loss is calculated as: ;D + and denote positive and negative triplets respectively; is a boundary hyperparameter used to control the minimum gap between positive and negative sample pairs, and Represent the prediction scores of positive and negative example triplets respectively; The cross entropy loss is calculated as: ; T is the number of triple samples in a batch of training, R represents all the relations in the knowledge graph, Represents the true value label. If the triple (s, r, o) is valid, then =1, otherwise 0; is the model’s predicted probability for the triple (s, r, o).

8. The method for expanding knowledge of household appliance manufacturing knowledge graph according to claim 1, characterized in that: The negative sampling methods based on subgraph structure in S5 include: For each positive sample, extract the entities and relations in the n-hop neighborhood associated with it from the knowledge graph; In the n-hop neighborhood set of the positive sample, select entities with similar neighborhood structures to replace the main entity or guest entity in the positive sample, and select other relations with similar semantics to replace the original relations; Verify the effectiveness of the generated negative samples; Multiple negative samples are generated for each positive sample triple by combining different types of negative samples.

Citation Information

Patent Citations

  • Small sample knowledge graph completion method based on embedding fusion and data enhancement

    CN119537600A

  • Induction link prediction method based on relation message passing and interaction information maximization

    CN119721229A