Matrix representation and Grassmann manifold optimization-based knowledge graph embedding and link prediction optimization method
Through the method based on matrix representation and Grassmann manifold optimization, knowledge graph embedding and link prediction are optimized, and the information integrity, semantic accuracy and rule consistency problems of traditional methods when dealing with complex relationships are solved, and efficient and accurate knowledge graph embedding and link prediction are achieved.
Patent Information
- Application Number
- CN202510303011.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional knowledge graph embedding methods are difficult to effectively deal with complex and diverse relationships, and there are problems of insufficient information integrity, semantic accuracy deviation and rule consistency, resulting in low link prediction accuracy and high operation and maintenance costs.
Using a method based on matrix representation and Grassmann manifold optimization, a similarity algorithm based on input templates, Grassmann manifold optimization algorithm coding, adaptive iterative detection, semantic accuracy and deviation control, and rule template matching is constructed to optimize knowledge graph embedding and link prediction.
It improves the accuracy and information integrity of the embedding matrix, ensures the consistency between the model output and the gold label, reduces operation and maintenance costs, and improves the stability of the model and the accuracy of link prediction.
Smart Images

Figure CN120258099A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of Knowledge Graph Embedding (KGE) and Link Prediction, and particularly relates to a method based on matrix representation and Grassmann manifold optimization. Specifically, an optimized method for Knowledge Graph Embedding and Link Prediction based on matrix representation and Grassmann manifold optimization is proposed. Background Art
[0002] With the rapid development of information technology, knowledge graphs have become important tools for understanding and representing complex relationships, showing extensive application value in many fields. Traditional knowledge graph embedding methods face many challenges when dealing with large-scale and complex relationship networks, especially in link prediction tasks. Knowledge graph embedding and link prediction methods based on matrix representation and Grassmann manifold optimization have shown their unique advantages. By transforming entities and relationships into matrix forms and using Grassmann manifold optimization techniques, the performance of the model in link prediction tasks has been significantly improved. This method not only enhances the ability to understand complex relationships in knowledge graphs but also opens up a new path for the intelligent processing of knowledge graphs. Currently, knowledge graph datasets are relatively scarce due to problems such as the difficulty of construction and the continuous emergence of relationship types. For example, the representation of symmetric, asymmetric, inverse, etc. relationships in multi-relational knowledge graphs, and the effects of traditional knowledge graph embedding methods in dealing with these complex relationships are not ideal.
[0003] Traditional knowledge graph embedding methods are mainly designed for simple relationship patterns and cannot effectively handle the large number of complex and diverse relationships emerging in knowledge graphs. These traditional methods often rely on designing rules or features for different types of relationships, being limited to the fixed pattern matching of static features, lacking generality and applicability, and involving huge workloads. Although the embedding vectors of traditional methods can represent known relationships, they cannot identify and process "customized" relationships and non-fixed fields. These non-standard relationships are difficult to quantitatively analyze in a standard way, easily misleading traditional knowledge graph embedding methods and reducing the accuracy of link prediction. In addition, traditional methods are difficult to capture complex relationships and implicit information, unable to fully understand the underlying intentions and motivations of entities and relationships. Constructing and maintaining a large feature library and corresponding rules requires a large amount of human resources input and continuous manual adjustment of rules or features, thus reducing the embedding quality of the continuously emerging knowledge graph information.
[0004] The knowledge graph embedding method based on matrix representation and Grassmann manifold optimization shows good adaptability and generality in dealing with different types of relationships and "customized" and "non-standard" vocabulary in the knowledge graph, and can handle various complex relationship behaviors and scenarios. Compared with traditional algorithms, this method has excellent context understanding and the ability to handle complex relationships, and can propose suggestions for filling the knowledge gaps exposed by the current knowledge graph in a user-friendly manner. It is not restricted by the static interpretation of the fixed pattern matching of keywords in traditional algorithms, and its diverse and expressive matrix representation makes the embedding of the knowledge graph more vivid and intuitive. This generative embedding is very friendly to non-professional data analysts in understanding a large amount of complex knowledge graph information, and greatly reduces the R & D and operation and maintenance costs. Due to its strong generalization ability and iterative correction and learning, this method can better adapt to new knowledge graph information without frequent model or algorithm updates.
[0005] However, the knowledge graph embedding method based on matrix representation and Grassmann manifold optimization has some deficiencies in some aspects: First, in terms of information integrity and standardization, although the generated embeddings are diverse and expressive, sometimes they may not comprehensively cover all relevant important key information points. Especially when the relationship pattern is novel or complex, it is easy to miss key details. Second, in terms of semantic accuracy and deviation control, when dealing with knowledge graphs containing special relationships or more complex ones, the large model may misinterpret the context, resulting in embeddings that do not match the semantics of the original knowledge graph. For example, a normal relationship may be misjudged as an abnormal relationship. At the same time, due to the limitations or unevenness of the training data, or over-generalization, the large model may ignore special situations in specific contexts and have a deviation in the understanding of certain relationships, resulting in inaccurate key details in the answers it generates; Third, in terms of rule consistency, the diversity of the large model may lead to its neglect of the importance of organizing key information according to a preset template, which may lead to inconsistencies and chaos in the output format, especially in scenarios involving standardized reports or highly structured information. Therefore, this increases the difficulty of the data analyst's quick response understanding and operation of information to a certain extent. Summary of the Invention
[0006] To solve the problems of the prior art, the present invention proposes an optimization method for knowledge graph embedding and link prediction based on matrix representation and Grassmann manifold optimization, which can effectively solve the problems of embedding representation and link prediction of complex relationship networks and efficiently screen and evaluate the optimal link prediction results.
[0007] The present invention is realized through the following technical solutions: An optimization method for knowledge graph embedding and link prediction based on matrix representation and Grassmann manifold optimization, the steps of which are as follows:
[0008] Step 1: Develop an input template: Divide the knowledge graph dataset D into several groups, each group containing a set number of entities, and set a matrix representation template for each group of datasets as the input for different models.
[0009] The specific method is as follows:
[0010] Develop a matrix representation template for entities and relationships as the input;
[0011] Divide the knowledge graph dataset D into n groups of datasets D based on every m entities k , where Y k =(E k , R k , T k ) is used as the gold label, and set the matrix representation template <Y k , Q> as the input for different models;
[0012] Among them, Y k ∈D k ,, k = 1, 2, 3...; E k is the entity set in the k-th group of datasets; R k is the relationship set in the k-th group of datasets; T k is the tail entity set in the k-th group of datasets; Q is a new entity or relationship.
[0013] Step 2: Obtain the embedding matrices output by the model: Use different models Mn, input the matrix representation template into the model, and obtain the embedding matrices (E', R', T') of entities, relationships, and tail entities.
[0014] The specific method is as follows:
[0015] Obtain the corresponding embedding matrices (E', R', T') of the model output through the matrix representation of entities and relationships;
[0016] Select the model Mn (n = 1, 2, 3, 4...), input the matrix representation template into the model Mn in sequence, and obtain the embedding matrices (E', R', T') output by the model Mn after matrix operations. Among them, E' represents the embedding matrix of entities, R' represents the embedding matrix of relationships, and T' represents the embedding matrix of tail entities. The specific method is as follows:
[0017] prompt is input to the model M1, and (E1', R1', T1') is output; input to the model M2, and (E2', R2', T2') is output; input to the model M3, and (E3', R3', T3') is output; input to the model M4, and (E4', R4', T4') is output; and so on. prompt is input to the model Mn, and (E', R', T') is output.
[0018] Step 3 Grassmann manifold optimization algorithm encoding: Use the Grassmann manifold optimization algorithm to encode the (E', R', T') matrix generated by each model Mn to generate a spatial encoding vector V' that can be understood by a computer.
[0019] The specific method is as follows:
[0020] For the (E', R', T') embedding matrix generated by each model Mn, after encoding through the Grassmann manifold optimization algorithm, a multi-dimensional spatial vector encoding V' (i = 1, 2, 3...) is generated;
[0021] The (E', R', T') generated by each model Mn all belong to the matrix category. To input these matrices as variables into the composite similarity evaluation criterion function, use the Grassmann manifold optimization algorithm to encode the matrices to generate a spatial encoding vector V' (i = 1, 2, 3...) that can be understood by a computer. The encoding calculation formula is as follows:
[0022]
[0023] Among them,
[0024] V' is the multi-dimensional spatial vector encoding output after the (E', R', T') generated by model Mn passes through the Grassmann manifold optimization algorithm Encode;
[0025] λ j , u j , v j are respectively the singular values and corresponding singular vectors of the matrix;
[0026] d is the dimension of the matrix.
[0027] Step 4 Iterative retrieval and completion of information integrity: Use the adaptive iterative detection algorithm to perform iterative retrieval and completion of the information integrity of (E', R', T') and calculate the matching score.
[0028] The specific method is as follows:
[0029] Perform iterative retrieval and completion of the information integrity of (E', R', T') through the information-corrected adaptive iterative detection algorithm function to obtain the matching score S km ;
[0030] Through the information-corrected adaptive iterative detection algorithm, calculate the proportion of the number of times a specified feature co-occurs in the total number of mentions to measure the matching degree of the feature in the two matrices. Iteratively perform iterative retrieval and completion of the information integrity of (E', R', T') in turn until the maximum number of iterations is reached or all features have been retrieved, and then the iteration ends;
[0031] Calculate the matching score S km , and the specific calculation formula is as follows:
[0032]
[0033] Among them,
[0034] F is the list of feature sets of the adaptive iterative detection algorithm for information correction, including entity features and relationship features;
[0035] and respectively represent the co-occurrence times and total mention times of feature f in V i and V * .
[0036] Step 5 Calculate the semantic similarity: Calculate the semantic similarity score between the (E', R', T') output by the model and the gold label through the semantic precision and deviation control algorithm.
[0037] The specific method is:
[0038] Calculate the semantic similarity between the (E', R', T') output by the model Mn and the gold label Y~(E*, R*, T*) given by the dataset D through the semantic precision and deviation control algorithm to obtain the matching score S sm ;
[0039] Calculate the semantic similarity between the (E', R', T') output by the model Mn and the gold label Y~(E*, R*, T*) given by the dataset D through the semantic precision and deviation control algorithm to obtain the semantic similarity score;
[0040] Determine the representative vocabulary list or phrase for the feature after preprocessing. For a certain feature, use the pre-trained word embedding model to calculate the average semantic similarity score between the matrix given by the model and the gold label on this feature, and then take the weighted average of all features in the feature set as the overall semantic similarity score: The specific calculation formula is as follows:
[0041]
[0042] Among them,
[0043] T is the set of fuzzy semantic features, including entity features and relationship features;
[0044] R t is the relationship set corresponding to feature t;
[0045] W r is the vocabulary list of relationship r;
[0046] w * is a representative word of feature t;
[0047] sim(w, w * ) is the average semantic similarity score between word w and all relevant words of feature t in V;
[0048] Z is a normalization factor.
[0049] Step 6 calculates the template matching degree: Calculate the template matching degree between the (E', R', T') output by the model and the gold label through a similarity algorithm based on rule-based template matching, and set a threshold g3. If the matching degree is lower than g3, then regenerate (E', R', T').
[0050] The specific method is as follows:
[0051] Calculate the template matching degree between the (E', R', T') output by model Mn and the gold label Y~(E, R*, T*) provided by dataset D through a similarity algorithm based on rule-based template matching, and obtain a matching score S tm , set a threshold g3, where g3 is the threshold of the similarity algorithm based on rule-based template matching. If S tm < g3, then let model Mn regenerate (E', R', T') with reference to the gold label Y template;
[0052] Calculate the number or proportion of matching templates between the (E', R', T') output by model Mn and the gold label Y~(E*, R*, T*) provided by dataset D through a similarity algorithm based on rule-based template matching, and take the weighted average of all feature template matching scores as the overall template matching score to obtain a matching score S tm , set a threshold g3. If S tm < g3, then let model Mn regenerate (E', R', T') with reference to the template of the gold label Y;
[0053] The specific calculation formula is as follows:
[0054]
[0055] Among them,
[0056] C is the set of expression templates of features;
[0057] c * is the expression template of the feature in the gold label;
[0058] match(c, c * ) is the matching degree between feature c and template c * ;
[0059] α c and βc is the template matching weight of features, which is used to adjust the relative importance of different features in template matching.
[0060] Step 7 constructs a composite similarity evaluation criterion: Combine the information-corrected adaptive iterative detection, semantic accuracy and deviation control, and the results of rule-based template matching to construct a composite similarity evaluation criterion, and calculate the composite similarity evaluation score.
[0061] The specific method is as follows:
[0062] Construct a composite similarity evaluation criterion through the information-corrected adaptive iterative detection algorithm, the semantic accuracy and deviation control algorithm, and the similarity algorithm based on rule-based template matching, and obtain the total evaluation score S of the output of model Mn and the gold label Y~(E, R*, T*) i (i = 1, 2, 3...);
[0063] Construct a composite similarity evaluation criterion through the information-corrected adaptive iterative detection algorithm, the semantic accuracy and deviation control algorithm, and the similarity algorithm based on rule-based template matching, and obtain the total evaluation score S of the output of model Mn and the gold label Y~(E*, R*, T*) i (i = 1, 2, 3...).
[0064] The specific calculation formula is as follows:
[0065] S i = γ km ·S km + γ sm ·S sm + γ tm ·S tm
[0066] Among them,
[0067] S i is the composite similarity evaluation score;
[0068] γ km 、γ sm 、γ tm are the weights based on information-corrected adaptive iterative detection, semantic accuracy and deviation control, and rule-based template matching respectively, reflecting the relative importance of each matching method in the overall similarity evaluation, and satisfying: γ km + γ sm + γ tm = 1.
[0069] Step 8 determines whether it exceeds the threshold G: Set the threshold G, and determine whether the composite similarity evaluation score exceeds G. If it does not exceed, a new round of iteration is performed; if the number of iterations exceeds 5 times or the evaluation score exceeds G, the iteration ends.
[0070] The specific method is as follows:
[0071] Set a threshold G, which is used to judge whether the overall composite similarity evaluation score S i meets the ideal requirement. Judge whether S i exceeds the threshold G. If S i < G, repeat steps 5-7 for a new round of iteration. If the number of iterations > 5, or there is an S i > G, the iteration ends;
[0072] Horizontal level comparison:
[0073] Compare the interpretation scores of different methods Mn for the same knowledge graph information, and calculate their difference di;
[0074] Set a threshold g. If di exceeds the threshold range, it indicates that the output differences between methods are large, and a standard historical data set needs to be introduced for correction; if di is within the threshold range, it is considered that the outputs of each method tend to be stable and the accuracy is the highest. Finally, select (E', R', T') corresponding to the maximum score Smax as the reliable result and output it to the user.
[0075] The beneficial effects of the present invention are as follows: It can significantly improve the accuracy and information integrity of the embedding matrix, ensure a high degree of consistency between the model output and the gold label. Through the composite similarity evaluation criterion, the performance of the model is further optimized. At the same time, the adaptive iteration mechanism effectively reduces errors and improves the stability of the model. In addition, the introduction of horizontal comparison and threshold setting enhances the reliability of the results, making this method perform excellently in semantic understanding and link prediction of knowledge graphs, and providing more accurate and reliable technical support for the application of knowledge graphs. Brief Description of the Drawings
[0076] Figure 1 is the overall process of the present invention;
[0077] Figure 2 is the framework diagram of the composite similarity evaluation criterion. Detailed Embodiment
[0078] An optimization method for knowledge graph embedding and link prediction based on matrix representation and Grassmann manifold optimization includes the following steps:
[0079] Specifically describe the present invention in combination with the accompanying drawings.
[0080] The process of the knowledge graph embedding optimization algorithm is as Figure 1 shown, which includes the specific algorithm detail process:
[0081] 1. Set the input template.
[0082] (1) Divide the knowledge graph dataset D into n groups of datasets D according to every m entities as a standard. k , Y k = (E k , R k , T k ) is used as the gold label.
[0083] (2) Set the matrix representation template: <Y k , Q>
[0084] Among them, Y k ∈ D k , k = 1, 2, 3...; E k is the entity set in the k-th group of datasets; R k is the relationship set in the k-th group of datasets; T k is the tail entity set in the k-th group of datasets; Q is a new entity or relationship.
[0085] 2. Select the method Mn (n = 1, 2, 3, 4...), and input the matrix representation template to the method Mn in turn, and obtain (E', R', T') output by the method Mn after matrix operation. The matrix representation template is input to the method M1, and (E1', R1', T1') is output; the matrix representation template is input to the method M2, and (E2', R2', T2') is output; the matrix representation template is input to the method M3, and (E3', R3', T3') is output; the matrix representation template is input to the method M4, and (E4', R4', T4') is output;...... The matrix representation template is input to the method Mn, and (E', R', T') is output.
[0086] Among them, E' represents the embedding matrix of entities; R' represents the embedding matrix of relationships; T' represents the embedding matrix of tail entities.
[0087] 3. Matrix encoding Each (E', R', T') generated by the method Mn belongs to the matrix category. To be able to input it as a variable into the composite similarity evaluation standard function, it needs to be encoded through the Grassmann manifold optimization algorithm to generate a spatial encoding vector V' (i = 1, 2, 3...) that can be understood by the computer. Its encoding calculation formula is as follows:
[0088]
[0089] Among them: V' is the multi-dimensional space vector encoding output after encoding (E', R', T') generated by the method Mn through the Grassmann manifold optimization algorithm Encode; are the singular values and corresponding singular vectors of the matrix respectively; d is the dimension of the matrix.
[0090] 4. Adaptive Iterative Detection Algorithm for Information Correction.
[0091] (1) First, through the adaptive iterative detection algorithm for information correction, calculate the proportion of the co-occurrence times of specified features in the total mention times to measure the matching degree of features in the two matrices.
[0092] (2) In an iterative manner, retrieve and complete the information integrity of (E', R', T') in sequence until the maximum number of iterations is reached or all features have been retrieved, and then the iteration ends.
[0093] (3) Calculate the matching score S km , and the specific calculation formula is as follows:
[0094]
[0095] Among them, F is the feature set list (such as entity features, relationship features, etc.); and respectively represent the co-occurrence times and total mention times of feature f in V i and V * .
[0096] 5. Control Algorithm for Semantic Precision and Deviation
[0097] (1) Calculate the semantic similarity between (E', R', T') output by the calculation method Mn and the golden label Y~(E*, R*, T*) given by the dataset D through the control algorithm for semantic precision and deviation to obtain the semantic similarity score.
[0098] (2) Determine the representative vocabulary list or phrase for features obtained after preprocessing. For a certain feature, use the pre-trained word embedding model to calculate the average semantic similarity score between the matrix given by the calculation method and the golden label on this feature, and then take the weighted average for all features in the feature set as the overall semantic similarity score. The specific calculation formula is as follows:
[0099]
[0100] Among them, T is the fuzzy semantic feature set ({entity, relation, tail entity}); Rt is the relation set corresponding to feature t; W r is the vocabulary list of relation r; w * is the representative vocabulary of feature t; sim(w, w * ) is the average semantic similarity score between vocabulary w and all relevant vocabularies of feature t in V*; Z is the normalization factor.
[0101] 6. Similarity Algorithm Based on Rule Template Matching.
[0102] (1) Calculate the number or proportion of the matching templates of (E', R', T') output by the similarity algorithm calculation method Mn based on rule template matching with the golden label Y~(E*, R*, T*) provided by the dataset D, and take the weighted average of all feature template matching scores as the overall template matching score to obtain the matching score S tm , set the threshold g3, if S tm < g3, then let the method Mn regenerate (E', R', T') with reference to the golden label Y template;
[0103] (2) The specific calculation formula is as follows:
[0104]
[0105] Among them, C is the set of expression templates of features; c * is the expression template of the feature in the golden label; match(c, c * ) is the matching degree between the feature c and the template c * ; α c and β c are the template matching weights of the features, used to adjust the relative importance of different features in template matching.
[0106] 7. Construct a composite similarity evaluation criterion.
[0107] (1) Construct a composite similarity evaluation criterion through the information-corrected adaptive iterative detection algorithm, the semantic accuracy and deviation control algorithm, and the similarity algorithm based on rule template matching, and obtain the total evaluation score S i (i = 1, 2, 3...).
[0108] (2) The specific calculation formula is as follows:
[0109] S i = γ km ·S km + γ sm ·S sm + γ tm ·S tm
[0110] Among them, S i is the composite similarity evaluation score; γ km , γ sm , γ tm are the weights of the information-corrected adaptive iterative detection, semantic accuracy and deviation control, and rule template matching-based respectively, reflecting the relative importance of each matching method in the overall similarity evaluation, and satisfying: γ km + γ sm + γtm = 1
[0111] 8. Iteration end condition.
[0112] First, set a threshold G and judge the composite similarity evaluation score S i whether it exceeds the threshold G. If S i < G (convergence), then repeat the iterative processes of information integrity, semantic similarity, and rule template matching respectively. If the number of iterations > 5, or there exists S i > G, then the iteration ends.
[0113] 9. Horizontal level comparison.
[0114] (1) Conduct a horizontal level comparison of the composite similarity evaluation scores between methods Mn, that is, calculate the difference di between the interpretation scores S of method Mn for the same knowledge graph information i among them;
[0115] (2) Set a threshold g. If di > g ∪ di
Claims
1. An optimization method for knowledge graph embedding and link prediction based on matrix representation and Grassmann manifold optimization, characterized in that The steps are as follows: Step 1: Develop an input template: Divide the knowledge graph dataset D into several groups, each group containing a set number of entities, and set a matrix representation template for each group of datasets as the input for different models; Step 2: Obtain the embedding matrices output by the models: Use different models Mn to input the matrix representation template into the models to obtain the embedding matrices (E', R', T') of entities, relationships, and tail entities; Step 3: Grassmann manifold optimization algorithm encoding: Use the Grassmann manifold optimization algorithm to encode the (E', R', T') matrices generated by each model Mn to generate a spatial encoding vector V' that can be understood by the computer; Step 4: Iterative retrieval and completion of information integrity: Use an adaptive iterative detection algorithm to retrieve and complete the information integrity of (E', R', T') and calculate the matching score; Step 5: Calculate semantic similarity: Calculate the semantic similarity score between the (E', R', T') output by the model and the gold label through a semantic accuracy and deviation control algorithm; Step 6: Calculate the template matching degree: Calculate the template matching degree between the (E', R', T') output by the model and the gold label through a similarity algorithm based on rule template matching, and set a threshold g3. If the matching degree is lower than g3, then regenerate (E', R', T'); Step 7: Construct a composite similarity evaluation criterion: Combine the results of adaptive iterative detection with information correction, semantic accuracy and deviation control, and rule template matching-based results to construct a composite similarity evaluation criterion and calculate the composite similarity evaluation score; Step 8: Determine whether the threshold G is exceeded: Set a threshold G and determine whether the composite similarity evaluation score exceeds G. If it does not exceed, then perform a new round of iteration; If the number of iterations exceeds 5 times or the evaluation score exceeds G, then the iteration ends.
2. The optimization method for knowledge graph embedding and link prediction with matrix representation and Grassmann manifold optimization according to claim 1, wherein In the said Step 1, the specific method is: Develop a matrix representation template of entities and relationships as the input; Divide the knowledge graph dataset D into n groups of datasets D with every m entities as a standard k , where Y k =(E k , R k , T k ) is used as the gold label, and set the matrix representation template <Y k , Q> as the input of different models; Among them, Y k ∈ D k , where k = 1, 2, 3...; E k is the entity set in the k-th dataset; R k is the relationship set in the k-th dataset; T k is the tail entity set in the k-th dataset; Q is a new entity or relationship.
3. The optimization method for knowledge graph embedding and link prediction with matrix representation and Grassmann manifold optimization according to claim 1, wherein In the said Step 2, the specific method is: Obtain the corresponding embedding matrices (E', R', T') output by the model through the matrix representation of entities and relationships; Select models Mn (n = 1, 2, 3, 4...), input the matrix representation template into model Mn in sequence, and obtain the embedding matrices (E', R', T') output by model Mn after matrix operations. Among them, E' represents the embedding matrix of entities, R' represents the embedding matrix of relationships, and T' represents the embedding matrix of tail entities. The specific method is: prompt is input to model M1, and (E1', R1', T1') is output; input to model M2, and (E2', R2', T2') is output; input to model M3, and (E3', R3', T3') is output; input to model M4, and (E4', R4', T4') is output; and so on. prompt is input to model Mn, and (E', R', T') is output.
4. An optimization method for knowledge graph embedding and link prediction with matrix representation and Grassmann manifold optimization according to claim 1, characterized in that, In the said Step 3, the specific method is: For the (E', R', T') embedding matrices generated by each model Mn, generate a multi-dimensional space vector encoding V' (i = 1, 2, 3...) after encoding through the Grassmann manifold optimization algorithm; The (E', R', T') generated by each model Mn belongs to the matrix category. To use these matrices as variables that can be input into the composite similarity evaluation criterion function, the Grassmann manifold optimization algorithm is used to encode the matrices, generating a spatial encoding vector V' (i = 1, 2, 3...) that can be understood by the computer. The encoding calculation formula is as follows: Where, V' is the multi-dimensional spatial vector encoding output after encoding (E', R', T') generated by model Mn through the Grassmann manifold optimization algorithm Encode; λ j , u j , v j are the singular values of the matrix and the corresponding singular vectors, respectively; d is the dimension of the matrix.
5. An optimization method for knowledge graph embedding and link prediction with matrix representation and Grassmann manifold optimization according to claim 1, characterized in that, In step 4, the specific method is as follows: The information integrity of (E', R', T') is iteratively retrieved and complemented through the adaptive iterative detection algorithm function with information correction to obtain the matching score S km ; Through the information-corrected adaptive iterative detection algorithm, calculate the proportion of the co-occurrence times of the specified features in the total mention times to measure the matching degree of the features in the two matrices. In an iterative manner, retrieve and complete the information integrity of (E', R', T') in turn until the maximum number of iterations is reached or all features have been retrieved, and then the iteration ends; Calculate the matching score S km , and the specific calculation formula is as follows: Where, F is the feature set list of the information-corrected adaptive iterative detection algorithm, including entity features and relationship features; and respectively represent the co-occurrence times and total mention times of feature f in V i and V * respectively.
6. The optimization method for knowledge graph embedding and link prediction with matrix representation and Grassmann manifold optimization according to claim 1, characterized in that, In step 5, the specific method is as follows: Calculate the semantic similarity between the output (E', R', T') of the model Mn and the gold label Y~(E*, R*, T*) given by the dataset D through the control algorithm of semantic precision and deviation, and obtain the matching score S sm ; Calculate the semantic similarity between (E', R', T') output by model Mn and the golden label Y~(E*, R*, T*) given by dataset D through the semantic accuracy and deviation control algorithm to obtain the semantic similarity score; Determine the representative vocabulary list or phrase for the feature after preprocessing. For a certain feature, use the pre-trained word embedding model to calculate the average semantic similarity score between the matrix given by the model and the golden label on this feature, and then take the weighted average of all features in the feature set as the overall semantic similarity score: The specific calculation formula is as follows: Where, T is the fuzzy semantic feature set, including entity features and relationship features; R t is the set of relationships corresponding to feature t; w r is the vocabulary list of relationship r; w * is a representative term of feature t; sim(w, w * ) is the average semantic similarity score of the vocabulary w and the feature t among all relevant vocabularies in V; Z is the normalization factor.
7. An optimization method for knowledge graph embedding and link prediction with matrix representation and Grassmann manifold optimization according to claim 1, characterized in that In step 6, the specific method is as follows: Calculate the template matching degree between (E', R', T') output by model Mn and the gold label Y~(E, R*, T*) provided by dataset D through a similarity algorithm based on rule template matching, and obtain the matching score S tm , set a threshold g3, where g3 is the threshold of the similarity algorithm based on rule template matching. If S tm < g3, then let model Mn regenerate (E', R', T') with reference to the gold label Y template; Calculate the number or proportion of matching templates between the (E', R', T') output by the model Mn and the golden label Y~(E*, R*, T*) provided by the dataset D through a similarity algorithm based on rule template matching, and take the weighted average of all feature template matching scores as the overall template matching score to obtain the matching score S tm , set a threshold g3, if S tm < g3, then let the model Mn regenerate (E', R', T') with reference to the template of the golden label Y; The specific calculation formula is as follows: Where, C is the set of expression templates for features; c * is an expression template for the features in the gold label; match(c,c * ) is the matching degree between feature c and template c * ; α c and β c are the template matching weights of the features, used to adjust the relative importance of different features in template matching.
8. An optimization method for knowledge graph embedding and link prediction with matrix representation and Grassmann manifold optimization according to claim 1, characterized in that, In step 7, the specific method is as follows: Construct a composite similarity evaluation criterion by jointly using an adaptive iterative detection algorithm with information correction, a control algorithm for semantic accuracy and deviation, and a similarity algorithm based on rule template matching, and obtain the total evaluation score S of the output of model Mn and the gold label Y~(E, R*, T*) i (i = 1, 2, 3...); Construct a composite similarity evaluation criterion by jointly using an adaptive iterative detection algorithm with information correction, a control algorithm for semantic accuracy and deviation, and a similarity algorithm based on rule template matching, and obtain the total evaluation score S of the output of model Mn and the gold label Y~(E*, R*, T*) i (i = 1, 2, 3...); The specific calculation formula is as follows: S i = γ km ·S km + γ sm ·S sm + γ tm ·S tm Where, S i is the composite similarity evaluation score; γ km and γ sm and γ tm are the adaptive iterative detection based on information correction, the control of semantic accuracy and deviation, and the weight based on rule template matching respectively, which reflect the relative importance of each matching method in the overall similarity evaluation and satisfy: γ km + γ sm + γ tm = 1 9. An optimization method for knowledge graph embedding and link prediction with matrix representation and Grassmann manifold optimization according to claim 1, characterized in that, In step 8, the specific method is as follows: Set a threshold G, where G is a threshold used to determine whether the overall composite similarity evaluation score S i meets the ideal requirements, and determine whether S i exceeds the threshold G. If S i < G, repeat steps 5 to 7 for a new round of iteration. If the number of iterations > 5, or there exists S i > G, the iteration ends; Horizontal-level comparison: Compare the interpretation scores of different methods Mn for the same knowledge graph information and calculate their difference di; Set a threshold g. If di exceeds the threshold range, it indicates that the output differences between methods are large and a standard historical dataset needs to be introduced for correction; if di is within the threshold range, it is considered that the outputs of each method tend to be stable and the accuracy is the highest. Finally, select the (E', R', T') corresponding to the maximum score Smax as the reliable result and output it to the user.
Citation Information
Cited By
Intelligent proposition method based on knowledge graph and pre-training generation model
CN121279406A