Open knowledge graph standardization framework based on comparative learning
By introducing contrast learning and multi-view fusion clustering technology into open knowledge graph normalization, the existing methods have solved the problem of high computational complexity and inability to process large-scale data, and efficient and accurate knowledge graph normalization is achieved.
Patent Information
- Application Number
- CN202510077695.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-30
AI Technical Summary
The existing open knowledge graph normalization method has problems with high computational complexity and long run time when processing large-scale data, and cannot effectively consider the ambiguity of nouns and relational phrases in different contexts.
A framework for open knowledge graph normalization based on contrast learning is proposed, including a comparison learning unit, a context view embedding unit, a fact view embedding unit, a data enhancement unit and a multi-view fusion clustering unit. Through contrast learning, the embedding of nouns and relational phrases is optimized, and clustered using multi-view information.
It significantly reduces the running time and improves the accuracy of standardization, allowing the framework to be used on a large-scale open knowledge graph, with practical application value.
Smart Images

Figure CN120069023A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of knowledge graphs, and more specifically, to an open knowledge graph normalization framework based on contrastive learning. Background Art
[0002] Open knowledge graphs are widely used in the fields of natural language processing, information retrieval, and knowledge graphs. However, knowledge extraction usually relies on open information extraction technology. When constructing a knowledge graph using these technologies, a large number of knowledge triples (noun phrases, relation phrases, noun phrases) are extracted from the original text. However, nouns and relation phrases are usually not normalized, and synonyms with different expressions increase redundancy.
[0003] To solve the problem of synonymous phrase clustering in open knowledge graphs, achieve the normalization of knowledge graphs, and improve their practicality, a series of open knowledge graph normalization methods have emerged; traditional text similarity-based and rule-based methods can only make judgments based on the surface semantics of nouns or relation words, with limited application scope, and cannot distinguish between nouns or relations that are textually similar but have different actual meanings, so the effect cannot meet the application requirements; with the development of deep learning technology, some open knowledge graph normalization methods have made progress based on deep learning technology. However, in different contexts, nouns and relation phrases may have different meanings, and these context information needs to be considered during the normalization process. However, these methods only use single-view information and ignore the complementary information between the fact view and the context view; recently, multi-view open knowledge graph normalization methods using multi-view information have emerged, but these methods are highly complex, require a large amount of computing power, and have a long running time, and can only meet the research needs in laboratory scenarios and cannot meet the growing needs of massive data on the Internet.
[0004] Based on this, on the basis of multi-view open knowledge graph normalization methods, the present invention invented a contrastive learning technique applicable to open knowledge graph normalization tasks, as well as the CHFK-means clustering algorithm, and proposed the CMVC+ framework, which improves the normalization accuracy while significantly reducing the running time, enabling it to be used on large-scale open knowledge graphs and having practical application value. Summary of the Invention
[0005] This application aims to solve at least one of the technical problems existing in the prior art. To this end, this application proposes an open knowledge graph normalization framework based on contrastive learning, including a contrastive learning unit, a context view embedding unit, a fact view embedding unit, a data augmentation unit, and a multi-view fusion clustering unit. The contrastive learning unit determines positive and negative sample pairs based on seed pairs, calculates the contrastive loss to optimize the embeddings of noun phrases and relation phrases. The positive samples are generated through the transitivity of the seed pairs, and the negative samples include hard negative samples and in-batch negative samples. The hard negative samples are selected as the M ones with the highest similarity to the current noun phrase from different synonym clusters, and the in-batch negative samples are calculated based on other noun phrases in the mini-batch except itself. The context view embedding unit uses a pre-trained language model to encode the source context into the context embeddings of noun phrases and fine-tunes the pre-trained language model through a contrastive learning algorithm. Its context view loss function is equal to the contrastive learning loss. The fact view embedding unit projects fact triples into fact embeddings using a knowledge graph embedding model, minimizes the margin-based loss function based on the training data, and optimizes the fact embeddings in combination with the contrastive learning loss. The fact view loss function is the sum of the original KEM loss and the contrastive learning loss. The data augmentation unit collects synonymous named entity pairs from external resources as seed pairs, exchanges the corresponding parts of the fact triples in the original training data using the seed pairs to generate augmented training data to increase the density of the training data.
[0006] The multi-view fusion clustering unit, based on the co-EM algorithm, realizes the mutual reinforcement between the fact view and the context view by alternately executing the M-step and the E-step, and fuses the multi-view clustering results according to the CHF index. In the M-step, the cluster center embeddings are calculated according to the clustering results of the other view, and in the E-step, the noun phrases are assigned to the most similar clusters to minimize the loss function. After convergence, the final clustering result is determined by combining the weighted cosine similarity with the CHF index as the weight.
[0007] In addition, an open knowledge graph normalization framework based on contrastive learning according to an embodiment of this application further has the following additional technical features: Preferably, in the contrastive learning unit, when calculating the contrastive loss, it is necessary to calculate the embedding similarities of positive and negative sample pairs. The embedding similarity calculation of positive sample pairs is based on the embeddings of noun phrases in a specific view, and the embedding similarity of negative sample pairs is the sum of the embedding similarities of all negative samples.
[0008] Preferably, in the context view embedding unit, the pre-trained language model includes but is not limited to BERT, XLNet, and ELMo, which can encode the source context into context embeddings. The contrastive learning algorithm fine-tunes the pre-trained language model by optimizing the contrastive loss function to improve the quality of context embeddings.
[0009] Preferably, in the fact view embedding unit, the knowledge graph embedding model includes, but is not limited to, TransE, HolE, and RotatE. High-quality fact embeddings are learned by minimizing the margin-based loss function, and the embedding effect is further optimized by combining the contrastive learning loss.
[0010] Preferably, in the data augmentation unit, the seed pairs are automatically collected from external resources, and the corresponding entities in the fact triples containing the seed pair entities are replaced in the original training data to generate new triples and add them to the augmented training data set, thereby increasing the number of training data instances.
[0011] Preferably, in the multi-view fusion clustering unit, when calculating the cluster center embedding in the M step, based on the embedding of the noun phrase in the current view and the L1 norm, and in the E step, the noun phrase is assigned to the most similar cluster in the current view based on the cosine similarity. The loss function is minimized by iteratively executing the M step and the E step until convergence.
[0012] Preferably, when processing open knowledge graph data, first, the data augmentation unit expands the training data, then the context view embedding unit and the fact view embedding unit extract the context and fact view embeddings respectively, and then the embedding is optimized by the contrastive learning unit. Finally, the multi-view fusion clustering unit performs clustering to obtain a normalized result.
[0013] Preferably, each unit works collaboratively. Among them, the contrastive learning unit provides a means to optimize the embedding for the context view embedding unit and the fact view embedding unit. The data augmentation unit provides richer training data for the fact view embedding unit. The multi-view fusion clustering unit integrates the information of each view to obtain the final clustering result.
[0014] Preferably, when processing open knowledge graph data with different structures and sources, the framework can adaptively adjust the parameters and operations of each unit, including knowledge graph data in different domains and different languages, to achieve effective normalization processing.
[0015] Preferably, parallel computing and distributed computing technical means can be adopted in the implementation of each unit in the framework to improve the processing efficiency, and at the same time ensure the accuracy and stability of the normalized result.
[0016] A contrastive learning-based open knowledge graph normalization framework according to an embodiment of the present application has the beneficial effects that: 1. The contrastive learning unit optimizes the embeddings of noun phrases and relation phrases through unique positive and negative sample generation and contrastive loss calculation, enabling the framework to excel in entity and relation normalization. Experiments on multiple datasets show that the average F1 value of entity normalization has a significant improvement compared to advanced methods, and the average F1 value of relation normalization also exceeds many baseline methods. It effectively solves the problem of synonymous phrase clustering and improves the accuracy and usability of the knowledge graph; 2. The data augmentation unit automatically collects seed pairs to generate augmented training data, increasing the data density and providing richer information for the fact view embedding unit, which helps to learn high-quality fact embeddings. The multi-view fusion clustering unit fuses multi-view information based on the co-EM algorithm. The coordinated operations of its M-step and E-step and the use of the CHF index improve the clustering quality while reducing unnecessary calculations, significantly enhancing the overall performance of the framework and enabling it to adapt to the processing of large-scale knowledge graph data; 3. The various units in the framework work together and can adaptively adjust parameters and operations. They can process knowledge graph data with different structures, sources, domains, and languages, demonstrating good generality. Parallel computing, distributed computing, and other technologies can be adopted during the implementation of each unit to further improve the processing efficiency. While ensuring the accuracy and stability of the normalization results, the running time is significantly shortened, overcoming many problems of existing methods in processing large-scale data. Brief Description of the Drawings
[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a schematic diagram of the overall framework of an open knowledge graph normalization framework based on contrastive learning according to an embodiment of the present application. Detailed Embodiments
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.
[0020] The following details a contrastive learning-based open knowledge graph normalization framework of the present invention through specific embodiments, asFigure 1 As shown in Figure 1 , an open knowledge graph normalization framework based on contrastive learning includes a contrastive learning unit, a context view embedding unit, a fact view embedding unit, a data augmentation unit, and a multi-view fusion clustering unit.
[0021] 1. Example Preparation Data Collection: Obtain fact triple data and related source context information from open knowledge graph data sources to ensure data integrity and accuracy. At the same time, automatically collect synonymous named entity pairs as seed pairs from external knowledge bases, text corpora, and other resources for subsequent data augmentation and contrastive learning.
[0022] Model Selection and Parameter Setting: Determine pre-trained language models (such as BERT, XLNet, or ELMo) and knowledge graph embedding models (such as TransE, HolE, or RotatE), and set the hyperparameters of the models according to specific requirements and hardware conditions, such as learning rate, number of iterations, hidden layer dimension, etc.
[0023] 2. Implementation of the Contrastive Learning Unit Positive and Negative Sample Generation: Based on the collected seed pairs, determine positive sample pairs according to the transitivity rule, that is, consider noun phrases belonging to the same seed pair as positive samples. For negative samples, select the M most similar noun phrases from different synonymous clusters as hard negative samples, and at the same time use other noun phrases in the mini-batch except itself as in-batch negative samples.
[0024] Contrastive Loss Calculation: Calculate the embedding similarity of positive and negative sample pairs, and calculate the contrastive loss function according to the set temperature hyperparameter. By optimizing the contrastive loss function, adjust the embeddings of noun phrases and relation phrases to make the embeddings of positive sample pairs closer and the embeddings of negative sample pairs farther apart.
[0025] Specifically, taking the noun phrase subi as an example, the process is also applicable to the relation phrase reli and the noun phrase obji. By distinguishing positive and negative sample pairs among noun phrases, we define the contrastive loss: where S = {sub1, ……, subi, ……} is the set of noun phrases, and subi+ represents the set of positive samples of subi.
[0026] In our module, positive samples are generated through the transitivity of seed pairs. For example, if noun phrases sub1 and sub2 belong to the same seed pair, they are synonymous and point to the same entity. If sub2 and sub3 belong to another seed pair, then by transitivity, sub1 and sub3 also form a positive sample pair. These noun phrases that point to the same entity form synonymous clusters, such as cluster1 and cluster2. Each pair of noun phrases in the same cluster can be used as a positive sample pair.
[0027] For each positive sample pair (subi, subi+), we calculate its contrastive loss: where, represents the embedding similarity between subi and its positive sample subi+ in view v, τ is the temperature hyperparameter, is the sum of the embedding similarities between subi and all its negative samples.
[0028] Negative samples include hard negative samples and in-batch negative samples, which are defined as follows: Different from previous open knowledge graph normalization methods, we use seed pairs to generate negative sample pairs. We assume that noun phrases in different synonymous clusters point to different entities. Therefore, noun phrases sub1, sub2, sub3 in cluster1 are candidate hard negative samples for noun phrase sub4 in cluster2. We select the top M with the highest similarity from these candidates as the final hard negative samples.
[0029] Through this process, we obtain the hard negative sample subi- of subi and calculate its embedding similarity with the hard negative sample: For the embedding similarity of in-batch negative samples, we define it as follows: where B represents the size of the mini-batch, subi is a noun phrase different from subk and is in the same mini-batch. We assume that each noun phrase can act as an in-batch negative sample, that is, learn to separate the other B - 1 noun phrases.
[0030] 3. Implementation of the context view embedding unit Context embedding extraction: Input the source context into the selected pre-trained language model to obtain the initial context embedding. The pre-trained language model encodes the source context into a vector representation as the embedding of the noun phrase in the context view.
[0031] Model Fine-tuning: Using the contrastive learning algorithm, the pre-trained language model is fine-tuned with the contrastive loss as the objective function. During the fine-tuning process, the model parameters are updated according to the contrastive loss, thereby improving the quality of the context embedding and making it better reflect the semantic information of the noun phrase in the context environment.
[0032] Specifically, to utilize the semantic distribution features of the source context where the factual triples are located, we adopt pre-trained language models (PLMs), such as BERT, XLNet, and ELMo, to extract context embeddings. Specifically, for a certain noun phrase subi, we input its source context csubi into the PLM and regard the output embedding as its context embedding. . As long as the source context can be encoded into context embeddings, most PLMs can be applied to this framework.
[0033] When there is enough task-specific labeled data for fine-tuning, PLMs can usually achieve excellent performance on this task. However, these labeled data often require a large amount of labor costs, and in the context view, the source context is initially unlabeled, which poses a challenge to the task. A feasible solution is to generate pseudo-labels based on the collected seed pairs. However, since the seed pairs are automatically collected and not manually verified, they may contain errors, thus making the generated pseudo-labels noisy.
[0034] To solve the above problems and generate high-quality context embeddings, we use the contrastive learning algorithm to fine-tune the PLM. The loss function of its context view is defined as: Compared with the traditional iterative clustering method, the contrastive learning module we proposed has significant advantages in execution efficiency. This module directly optimizes the distance between positive and negative sample pairs, avoiding the large amount of calculations required for multiple rounds of clustering, thereby improving the normality of the effect and significantly reducing the execution time.
[0035] 4. Implementation of the Factual View Embedding Unit Factual Embedding Generation: Use the selected knowledge graph embedding model to project the factual triples into factual embeddings. The knowledge graph embedding model maps them to a low-dimensional vector space according to the semantic information of the entities and relationships in the triples.
[0036] Loss Function Optimization: Calculate the margin-based loss function based on the training data, and at the same time combine the contrastive learning loss to jointly optimize the parameters of the knowledge graph embedding model. By minimizing the loss function, more accurate factual embeddings are learned, enabling them to better represent the structural information in the factual triples.
[0037] Specifically, common models include TransE, HolE, RotatE, etc. Most KEMs are applicable to this framework as long as they can encode fact triples into corresponding fact embeddings.
[0038] Specifically, given a fact triple , KEM can project subi, reli, and obji into their corresponding fact embeddings, denoted as .
[0039] Suppose the scoring function used by KEM is , which is used to evaluate the credibility of fact triples. To learn high-quality fact embeddings, KEM usually minimizes a margin-based loss function based on the training data (i.e., a set of fact triples). The specific form is as follows: where represents ; γ > 0 is a margin hyperparameter; represents all available positive training data (positive fact triples) in OKB, while represents the negative training data generated by negative sampling, defined as follows: where N is the set of all named entities in T.
[0040] The formula shows that when generating negative samples, only a single replacement is made for the subject or object of the positive fact triple, rather than replacing both at the same time. In the fact view, we apply this contrastive learning module to fact embeddings and define the loss function of the fact view as: This loss function combines the original KEM loss with the contrastive learning loss, enabling the embeddings of noun phrases to not only reflect the co-occurrence relationship of triples but also be further optimized through the comparison of positive and negative sample pairs, thereby improving the expressive power of the embeddings.
[0041] 5. Implementation of the data augmentation unit Utilization of seed pairs: Utilize the collected seed pairs to perform entity replacement operations in the fact triples of the original training data. For each seed pair, find the fact triples containing the entities in the seed pair and replace one of the entities with the other to generate new triples.
[0042] Generation of augmented data: Add the generated new triples to the augmented training data set, thereby increasing the quantity and diversity of the training data, improving the density of the training data, and providing richer learning information for the knowledge graph embedding model.
[0043] Specifically, in an open knowledge graph, fact triples are often sparse, which poses challenges for KEM (Knowledge Embedding Model) to learn high-quality fact embeddings. To increase the density of training data and expand the number of data instances, we propose a data augmentation operation that utilizes the prior knowledge embedded in seed pairs to generate additional augmented training data by swapping the corresponding parts of the seed pairs in fact triples. This method can be regarded as a strategy for optimizing training data in terms of the quantity of data.
[0044] Specifically, assume that the original training data of OKB (i.e., all available fact triples in OKB) is , and the augmented training data is denoted as , and the latter is generated from through a data augmentation operator.
[0045] To implement data augmentation, we can automatically collect a set of synonymous named entity pairs (i.e., seed pairs) from external resources as high-quality prior knowledge. Let represent a seed pair, where and are synonymous named entities referring to the same entity. The set of collected seed pairs is denoted as . Given the set of seed pairs and the original training data , the augmented training data is generated as follows: For each seed pair , find the fact triples in the original training data that contain or ; Replace in the triple with or replace with to generate new triples; Add the generated triples to the augmented data set .
[0046] The formula is as follows: Through this method, the number of fact triples in the augmented training data increases significantly, thereby improving the density of training data and providing more effective information for KEM to learn high-quality fact embeddings.
[0047] 6. Implementation of the multi-view fusion clustering unit Initialization: Set initial clustering parameters, such as the initial values of cluster centers, the number of target clusters, etc.
[0048] Iteratively execute the M-step and the E-step: M-step: Based on the clustering results from another view, calculate the cluster center embedding for each cluster in the current view. When calculating, based on the embedding of noun phrases in the current view and the L1 norm, sum and normalize the embeddings of noun phrases belonging to the same cluster.
[0049] E-step: Based on cosine similarity, assign noun phrases to the most similar cluster in the current view. Calculate the cosine similarity between each noun phrase and the cluster centers, and assign it to the cluster with the highest similarity.
[0050] Convergence judgment: Calculate the loss function and determine whether the convergence condition is satisfied (such as the change in the loss function value is less than the set threshold or the maximum number of iterations is reached). If not converged, continue to execute the M-step and the E-step; if converged, proceed to the next step.
[0051] Determine the final clustering result: Calculate the consensus mean and the CHF index. The CHF index is calculated based on factors such as the number of noun phrases within the cluster, the distance between the cluster center and the mean, etc. Using the CHF index as the weight, combined with the weighted cosine similarity, assign noun phrases to the final clusters to obtain a normalized clustering result.
[0052] Specifically, in the specific embodiments of this application, we propose a new multi-view CHFK-Means algorithm, which is specifically used for the normalization of open knowledge graphs to evaluate and optimize the clustering quality in a more refined manner. Specifically, we fuse knowledge in two views - the fact view and the context view - to improve the clustering effect.
[0053] Inspired by the multi-view spherical K-Means algorithm, we designed the multi-view CHFK-Means method based on the co-EM algorithm. To quantify the clustering quality, we adopted the well-known statistical metric - the Caliński-Harabasz (CH) index, and innovatively proposed a new fine-grained extended form called the CHF index to distinguish the quality of each cluster in different views, rather than treating all views and clusters equally.
[0054] In the algorithm design, we alternately execute the M-step (maximization step) and the E-step (expectation step) of the two views to achieve mutual reinforcement between the views, transfer the clustering results between the two views, and use the view fusion algorithm based on the CHF index to fuse the clustering results of multiple views.
[0055] The following details the specific process of this algorithm: M-step: Input the clustering results from another view , where is the j-th cluster and K is the number of target clusters.
[0056] The cluster center embedding of the jth cluster in view v is calculated according to the following formula: in, is the view-specific embedding of the noun phrase subi in view v, represents the L1 norm.
[0057] Step E: Given the cluster center embeddings computed in the M step , assign the noun phrase subi to the most similar cluster of view v: in, is the cosine similarity, S is the full set of noun phrases. All embeddings are assumed to be of unit length.
[0058] After completing one M-step and one E-step of view v, the clustering result is passed to another view to continue iteration, alternating between the M-step and the E-step to minimize the following loss function: When the above process converges, the final clustering result of each view can be obtained. .
[0059] Based on CHF view fusion algorithm: Since the clustering results of different views may conflict, we design a strategy based on the consistency of cluster centers to integrate the final results.
[0060] To eliminate conflicts, we compute the consensus mean of the jth cluster in view v based on the distance between the noun phrase and its nearest cluster center embedding: : In addition, to quantify the clustering quality of clusters in a view, we define the jth cluster in view v as Calculate the CHF index: in, is the number of noun phrases in the cluster, is the total number of noun phrases.
[0061] The higher the CHF index, the closer the cluster is to other clusters. We use the CHF index as the weight of each cluster and combine it with the weighted cosine similarity to assign noun phrases to the final cluster to obtain the final clustering result. .
[0062] Each of these clusters : 。
[0063] 7. Overall framework operation process First, the original training data is augmented by the data augmentation unit to improve the data quality. Then, the context view embedding unit and the fact view embedding unit are used to extract embedding information from different views respectively, and the contrast learning unit is used to optimize the embedding effect. Finally, the multi-view fusion clustering unit integrates the information of the two views and obtains the final normalized result through iterative clustering, realizing the effective clustering and normalization of noun phrases and relation phrases in the open knowledge graph.
[0064] In the specific embodiments of this application, the experimental results of all methods on the entity normalization task are shown in Table 1, Table 1 Experimental results table of entity normalization Generally speaking, it can be seen from Table 1 that the CMVC+ framework we proposed is always superior to all competing baseline methods in terms of the average F1 value on the three datasets, which verifies the effectiveness of CMVC+ in the entity normalization task.
[0065] Compared with the state-of-the-art baseline CMVC, the average F1 values of CMVC+ on the ReVerb45K, NYTimes2018, and OPIEC59K datasets are increased by 1.1%, 1.4%, and 2.3% respectively. This shows that the newly proposed multi-view CHFK-Means clustering algorithm and contrast learning module in our work play an effective role and make great contributions in improving the NP normalization performance.
[0066] Table 2 Experimental results table of relation normalization It can be seen from Table 2 that all baseline results are directly cited from CMVC. It can be seen that the CMVC+ proposed in this paper exceeds all baseline methods in terms of the average F1 value on the three datasets.
[0067] Compared with CMVC, our new framework CMVC+ further improves the average F1 value by more than 1% on the three datasets by using the newly proposed multi-view CHFK-Means clustering algorithm and contrast learning module, indicating its superiority in the RP normalization task.
[0068] Table 3 Comparison table of running time To evaluate the efficiency of CMVC+ in terms of execution time, we compared it with the previously proposed framework CMVC. The main difference between the two is that CMVC+ uses the multi-view CHFK-Means clustering algorithm and the contrastive learning module, while CMVC uses the multi-view K-Means clustering algorithm and the iterative clustering process.
[0069] On each dataset, the execution times for noun phrase and relation phrase normalization are shown in Table 3.
[0070] As can be seen from Table 3, the execution time of CMVC+ is significantly lower than that of CMVC. On the ReVerb45K dataset, the execution time of CMVC is 78 minutes, while that of CMVC+ is only 57 minutes. On the NYTimes2018 dataset, the execution time of CMVC is 96 minutes, while that of CMVC+ is reduced to 73 minutes. On the OPIEC59K dataset, the execution time of CMVC is 157 minutes, while that of CMVC+ is reduced to 113 minutes.
[0071] The high efficiency of CMVC+ is mainly attributed to two aspects. First, the contrastive learning module used by CMVC+ avoids multiple calls to the K-Means algorithm during the iterative clustering process by directly optimizing the quality of the embeddings. Second, the multi-view CHFK-Means clustering algorithm used by CMVC+ is more efficient than the previous multi-view K-Means algorithm because it reduces unnecessary computations by finely measuring the clustering quality in different views.
[0072] Therefore, the experimental results show that while maintaining high performance, CMVC+ significantly improves the execution efficiency of the open knowledge graph normalization task.
[0073] The above embodiments are only used to illustrate the specific embodiments of the present invention and are not limited thereto. For those skilled in the art, according to the idea of the present invention, various similar deformations and transformations can be made, and these deformations and transformations should be regarded as the protection scope of the present invention.
Claims
1. A framework for normalizing open knowledge graphs based on contrastive learning, characterized in that: include: Contrastive learning unit, determines positive and negative sample pairs based on seed pairs, and calculates contrastive loss to optimize the embedding of noun phrases and relation phrases, where positive samples are transitively generated through seed pairs, and negative samples include difficult negative samples and in-batch negative samples. Difficult negative samples select the M most similar to the current noun phrase from different synonymous clusters, and in-batch negative samples are calculated based on other noun phrases in the mini-batch except themselves; The context-view embedding unit encodes the source context into context embeddings of noun phrases using a pre-trained language model and fine-tunes the pre-trained language model via a contrastive learning algorithm, with a context-view loss function equal to the contrastive learning loss; The fact view embedding unit uses the knowledge graph embedding model to project the fact triples into fact embeddings, minimizes the boundary-based loss function based on the training data, and optimizes the fact embeddings in combination with the contrastive learning loss. The fact view loss function is the sum of the original KEM loss and the contrastive learning loss. The data augmentation unit collects synonymous named entity pairs from external resources as seed pairs, uses the seed pairs to exchange the corresponding parts of the fact triples in the original training data, and generates augmented training data to improve the density of training data; The multi-view fusion clustering unit is based on the co-EM algorithm. It realizes the mutual reinforcement between the fact view and the context view by alternately executing the M step and the E step. The multi-view clustering results are fused according to the CHF index. The M step calculates the cluster center embedding according to the clustering result of another view, and the E step assigns the noun phrases to the most similar cluster to minimize the loss function. After convergence, the CHF index is used as the weight combined with the weighted cosine similarity to determine the final clustering result.
2. The open knowledge graph normalization framework based on contrastive learning as claimed in claim 1, characterized in that: In the contrastive learning unit, when calculating the contrastive loss, it is necessary to calculate the embedding similarity of the positive sample pair and the negative sample pair, where the embedding similarity of the positive sample pair is calculated based on the embedding of the noun phrase in a specific view, and the embedding similarity of the negative sample pair is the sum of the embedding similarities of all negative samples.
3. The open knowledge graph normalization framework based on contrastive learning as claimed in claim 1, characterized in that: In the context view embedding unit, the pre-trained language model includes but is not limited to BERT, XLNet and ELMo, which can encode the source context into context embedding, and the contrastive learning algorithm fine-tunes the pre-trained language model by optimizing the contrastive loss function to improve the quality of context embedding.
4. The open knowledge graph normalization framework based on contrastive learning as claimed in claim 1, characterized in that: In the fact view embedding unit, the knowledge graph embedding models include but are not limited to TransE, HolE and RotatE, which learn high-quality fact embedding by minimizing the boundary-based loss function, and further optimize the embedding effect by combining contrastive learning loss.
5. The open knowledge graph normalization framework based on contrastive learning as claimed in claim 1, characterized in that: In the data enhancement unit, the seed pairs are automatically collected from external resources, and the corresponding entities in the fact triples containing the seed pair entities are replaced in the original training data to generate new triples to be added to the enhanced training data set, thereby increasing the number of training data instances.
6. The open knowledge graph normalization framework based on contrastive learning as claimed in claim 1, characterized in that: In the multi-view fusion clustering unit, the M step calculates the cluster center embedding based on the embedding and L1 norm of the noun phrase in the current view, and the E step assigns the noun phrase to the most similar cluster in the current view based on the cosine similarity. The loss function is minimized by iteratively executing the M step and the E step until convergence.
7. The open knowledge graph normalization framework based on contrastive learning as claimed in claim 1, characterized in that: When processing open knowledge graph data, the training data is first expanded by the data augmentation unit, and then the context view embedding unit and the fact view embedding unit are used to extract the context and fact view embeddings respectively, and then the embeddings are optimized by the comparative learning unit, and finally clustered by the multi-view fusion clustering unit to obtain the normalized results.
8. The open knowledge graph normalization framework based on contrastive learning as claimed in claim 1, characterized in that: The various units work together, in which the contrastive learning unit provides a means of optimizing embedding for the context view embedding unit and the fact view embedding unit, the data enhancement unit provides richer training data for the fact view embedding unit, and the multi-view fusion clustering unit integrates the information of each view to obtain the final clustering result.
9. The open knowledge graph normalization framework based on contrastive learning as claimed in claim 1, characterized in that: When processing open knowledge graph data of different structures and sources, the framework can adaptively adjust the parameters and operations of each unit, including knowledge graph data from different fields and languages, to achieve effective normalized processing.
10. The open knowledge graph normalization framework based on contrastive learning as claimed in claim 1, characterized in that: During the implementation process, each unit in the framework can use parallel computing and distributed computing technology to improve processing efficiency while ensuring the accuracy and stability of the normalized results.