Data processing method and device, electronic equipment, medium and computer program product

By jointly learning triplets in natural language processing, combining relational words, structures and attribute embedding, the problems of information transmission errors and entity recognition difficulties are solved, and the processing accuracy is significantly improved.

CN119962651APending Publication Date: 2025-05-09CHINA MOBILE SHANGHAI ICT CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510127830.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In natural language processing, information transfer errors lead to low accuracy, especially when identifying the same entity in different triples.

Method used

By combining the embeddings, structural embeddings and attribute embeddings of relational words in triplets, the triplets are jointly learned to identify different expressions of the same entity.

Benefits of technology

It effectively improves the accuracy of triple embedding, reduces the error in context information transmission, and improves the accuracy of natural language processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962651A_ABST
    Figure CN119962651A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, electronic equipment, a medium and a computer program product, and the data processing method comprises the steps: determining a first embedding learning mode of a triple based on a first relation word of each triple in a triple set; each triple comprises a relation triple and an attribute triple; based on each attribute triple in the triple set, determining a second embedding learning mode of the triple; based on each relation triple in the triple set, determining a third embedding learning mode of the triple; and through the first embedded learning mode, the second embedded learning mode and the third embedded learning mode, carrying out joint embedded learning on a triple.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and in particular relates to a data processing method, device, electronic device, medium and computer program product. Background Art

[0002] Natural Language Processing (NLP) is an important branch of artificial intelligence, which aims to enable computers to understand, generate and process human natural language. At present, the natural language processing process is usually divided into multiple stages, such as text preprocessing, syntactic analysis, semantic analysis, information extraction, context understanding, etc. Since the information extracted in each stage is different, and the information extracted in each stage will enter the next stage for further application and processing, when there is an error in the information obtained in one stage, it will affect the accuracy of the processing in the subsequent stages, resulting in transmission errors, resulting in low accuracy in the natural language processing process. Summary of the invention

[0003] The embodiments of the present application provide a data processing method, device, electronic device, medium and computer program product, which improve the accuracy of natural language processing. By combining the embedding of relational words in triples, structural embedding and attribute embedding, the embedding of triples is jointly learned, which can effectively identify different manifestations of the same entity in triples, improve the accuracy of triple embedding, and further improve the accuracy of natural language processing.

[0004] The present application provides a data processing method, the method comprising:

[0005] Determine a first embedding learning mode of the triple based on the first relation word of each triple in the triple set; each triple includes a relation triple and an attribute triple;

[0006] Based on each attribute triple in the triple set, determine a second embedding learning method for the triple; based on each relationship triple in the triple set, determine a third embedding learning method for the triple;

[0007] The triples are jointly embedded and learned by using the first embedding learning method, the second embedding learning method, and the third embedding learning method.

[0008] In some embodiments, the first embedding learning method of determining the triple based on the first relation word of each triple in the triple set includes: based on each triple in the triple set, obtaining the first target type of the first head entity before the first relation word, and the second target type of the tail word after the first relation word; based on the first relation word, the first target type, and the second target type, determining the first embedding learning method of the triple; the tail vocabulary includes the first tail entity or the first attribute value.

[0009] It can be seen that the first target type helps determine the correlation between any two first head entities, and the second target type helps determine the correlation between any two tail words. By determining the first embedding learning method of the first relational word based on the first target type and the second target type, it is helpful to obtain the correlation between any two first relational words and improve the accuracy of the embedding of the first relational words.

[0010] In some embodiments, the obtaining of the first target type of the first head entity before the first relationship word and the second target type of the tail word after the first relationship word includes: obtaining a first type set of the first head entity before the first relationship word and a second type set of the tail word after the first relationship word; wherein the first type set includes a plurality of preset types corresponding to the first head entity; the second type set includes a plurality of preset types corresponding to the tail word; the first target type is determined by weighting the vector of each type in the first type set and the weight of each type in the first type set; the second target type is determined by weighting the vector of each type in the second type set and the weight of each type in the second type set.

[0011] It can be seen that obtaining the target type by combining multiple preset types is conducive to accurately judging the similarity between any two first head entities, and accurately judging the similarity between any two tail words. Further based on the similarity results between the first head entities and the similarity results between the tail words, it is conducive to obtaining the similarity between any two first relational words, thereby improving the accuracy of embedding relational words in triples.

[0012] In some embodiments, before determining the first target type and determining the second target type, the method also includes: determining the weight of the first type based on the number of the first type in the triple set type; the first type is any one type in the first type set and the second type set; the triple set type includes multiple types corresponding to each first head entity and multiple types corresponding to each tail word.

[0013] It can be seen that determining the weights of different types based on the number of type occurrences and determining the target type based on multiple types and the weights corresponding to each type is conducive to improving the accuracy of the target type, and further helps to improve the accuracy of the first relation word embedding.

[0014] In some embodiments, before determining the first target type and determining the second target type, the method further includes: based on a first correlation between a pre-labeled type and an entity, deleting types in the first type set and the second type set whose first correlation with the entity is less than a first correlation threshold.

[0015] It can be seen that by deleting types whose first correlation with the entity is less than the first correlation threshold, it is helpful to improve the efficiency of determining the target type through multiple types, which is further helpful to improve data processing efficiency.

[0016] In some embodiments, the first embedding learning method of determining the triple based on the first relation word, the first target type, and the second target type includes: constructing a first damaged triple corresponding to each triple based on the first relation word, the first target type, and the second target type; each first damaged triple includes a first relation word, a third target type, and a fourth target type; in each triple, a first vector is obtained by summing the vector of the first target type and the vector of the first relation word; a first difference between the first vector and the vector of the second target type is calculated; in each first damaged triple, a second vector is obtained by summing the vector of the third target type and the vector of the first relation word; a second difference between the second vector and the vector of the fourth target type is calculated; the first difference corresponds to the second difference one-to-one; a first objective function is constructed based on the difference between the first difference and the corresponding second difference, and the first embedding learning method of the triple is determined with the goal of minimizing the first objective function.

[0017] It can be seen that the method provided in this embodiment is conducive to improving the accuracy of embedding the first relational word in the triple. By implementing the first embedding learning method for embedding relational words, combining the second embedding learning method and the third embedding learning method, the triple is jointly embedded, which is conducive to improving the accuracy of triple embedding.

[0018] In some embodiments, each attribute triple includes a second head entity, a second relational word, and a second attribute value; the second embedding learning method of determining the triple based on each attribute triple in the triple set includes: splitting the vector of the second attribute value in each attribute triple to obtain a first N-tuple sequence of each attribute triple; N is a positive integer greater than or equal to 2; constructing a damaged attribute triple corresponding to each attribute triple; wherein each damaged attribute triple includes a third head entity, the second relational word, and a second N-tuple sequence; in each attribute triple, A third vector is obtained by summing the vector of the second head entity and the vector of the second relation word; a third difference between the third vector and the first N-tuple sequence is calculated; in each damaged attribute triple, a fourth vector is obtained by summing the vector of the third head entity and the vector of the second relation word; a fourth difference between the fourth vector and the second N-tuple sequence is calculated; the third difference corresponds one-to-one to the fourth difference; a second objective function is constructed based on the difference between the third difference and the corresponding fourth difference, and a second embedding learning method for the triple is determined with the goal of minimizing the second objective function.

[0019] It can be seen that by splitting the vector of the second attribute value in the triple, the split vector is conducive to accurately judging the similarity between any two second attribute values. Further determining the second embedding learning method corresponding to the second attribute value based on the N-tuple sequence, combining the first embedding learning method and the third embedding learning method to jointly embed the triple, is conducive to improving the accuracy of triple embedding.

[0020] In some embodiments, each of the relationship triples includes a fourth head entity, a third relationship word, and a second tail entity; the third embedding learning method of the triples is determined based on each relationship triple in the triple set, including: constructing a damaged relationship triple corresponding to each relationship triple; the damaged relationship triple represents a triple obtained by replacing the fourth head entity or the second tail entity in each relationship triple with an erroneous entity; based on each damaged relationship triple and each relationship triple, a third objective function is constructed to minimize the third objective function and determine the third embedding learning method of the triples; the third embedding learning method is determined by the first embedding learning method, the second embedding learning method, and the Before the third embedding learning method is used to perform joint embedding learning on the triples, the method also includes: determining a first similarity between the second objective function and the third objective function; constructing a fourth objective function based on the first similarity; the value of the fourth objective function is negatively correlated with the first similarity; determining a fourth embedding learning method with the goal of minimizing the fourth objective function; performing joint embedding learning on the triples through the first embedding learning method, the second embedding learning method, and the third embedding learning method, includes: performing joint embedding learning on the triples through the first embedding learning method, the second embedding learning method, the third embedding learning method, and the fourth embedding learning method.

[0021] It can be seen that the fourth objective function is constructed by the third embedding learning method representing entity embedding and the second embedding learning method representing attribute embedding, and the fourth embedding learning method is determined based on the fourth objective function, which is conducive to capturing the similarities of entities between different triples through entity embedding and attribute embedding, and realizing the unification of entities between different triples.

[0022] The present application also provides a data processing device, the device comprising:

[0023] The first processing module is used to determine a first embedding learning mode of a triple based on a first relation word of each triple in the triple set; each triple includes a relation triple and an attribute triple; based on each attribute triple in the triple set, determine a second embedding learning mode of the triple; based on each relation triple in the triple set, determine a third embedding learning mode of the triple;

[0024] The second processing module is used to perform joint embedding learning on triples through the first embedding learning method, the second embedding learning method, and the third embedding learning method.

[0025] An embodiment of the present application provides an electronic device, the electronic device comprising a processor and a memory for storing a computer program that can be run on the processor; wherein:

[0026] The processor is used to run the computer program to execute any one of the above data processing methods.

[0027] An embodiment of the present application provides a computer storage medium on which a computer program is stored. When the computer program is executed by a processor, any of the above-mentioned data processing methods is implemented.

[0028] An embodiment of the present application provides a computer program product, including a computer program, which implements any of the above-mentioned data processing methods when executed by a processor.

[0029] The embodiments of the present application provide a data processing method, apparatus, electronic device, medium and computer program product. Based on the data processing method given in the embodiments of the present application, triples are jointly embedded through a first embedding learning method for implementing first relational word embedding in each triple, a second embedding learning method for implementing attribute embedding, and a third embedding learning method for implementing structure embedding. The joint embedding learning effectively reduces the error in context information transmission, and the accuracy of determining the same entity in different triples is improved through the three embedding learning methods, which is further conducive to achieving the unification of entities in different triples, improving the accuracy of triple embedding, and improving the accuracy of natural language processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A flow chart of a data processing method provided in an embodiment of the present application;

[0031] Figure 2 A flow chart of an embedding method provided in an embodiment of the present application;

[0032] Figure 3 A flowchart of relational word unification and entity unification provided for an embodiment of the present application;

[0033] Figure 4 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;

[0034] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0035] Natural language processing is an important branch of artificial intelligence, which aims to enable computers to understand, generate and process human natural language. The main technical methods of natural language processing include large language model (LLM), lexical analysis, grammatical analysis, semantic analysis, entity recognition, sentiment analysis, etc. Large language model is one of the most basic concepts in natural language processing, which is used to predict the probability of a word or phrase appearing in a given context. High-quality data annotation and training play a key role in evaluating large language models. At present, common data annotation and training methods are: 1. Fine-grained sentiment analysis method based on supervised learning, which uses syntactic spanning tree to generate feature sets of associated evaluation objects, and then uses machine learning methods for training. 2. Fine-grained sentiment analysis method based on deep learning, which first obtains the word embedding of each word, and then uses structures such as recurrent networks or convolutional networks to perform information recognition data training in multiple stages. 3. Natural language text triple extraction method based on attention mechanism, which first performs dual-path encoding on a single word, and then uses a two-layer bidirectional long short-term memory network to identify and extract the associated information in the adjacency matrix of the graph convolutional network. 4. The triplet extraction method based on the joint approach is a pipeline-based distributed extraction method that uses two modules, encoder and decoder, to label and train data.

[0036] The problems faced in the above implementation method mainly include the following two points: one is that information recognition and data extraction are not comprehensive enough. If different information is extracted in multiple stages, there will be information transmission errors between tasks because the mutual assistance and improvement effect between tasks in multiple stages is indirect; the other is that the accuracy is not high in dealing with slight changes in context and low-resource languages.

[0037] In order to solve the above problems, the embodiment of the present application proposes a data processing method, which determines the first embedding learning method of the first relational word through the first relational word in each triple. Through the first embedding learning method of characterizing the embedding of relational words, combined with the second embedding learning method of characterizing the embedding of triple attributes, and the third embedding learning method of characterizing the embedding of triple structures, by calculating the similarity between entities and the similarity between relational words, different entities or entities in different knowledge graphs (KG) can be embedded into the same vector space, and finally, through the embedded entities and relational words, the missing entities (head entities or tail entities) or attribute values ​​can be predicted.

[0038] The following is a further detailed description of the embodiments of the present application in conjunction with the accompanying drawings and examples. It should be understood that the embodiments provided herein are only used to explain the embodiments of the present application and are not intended to limit the embodiments of the present application. In addition, the embodiments provided below are partial embodiments for implementing the present application, rather than providing all embodiments for implementing the present application. In the absence of conflict, the technical solutions recorded in the embodiments of the present application can be implemented in any combination.

[0039] It should be noted that, in the embodiments of the present application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a method or device including a series of elements includes not only the elements explicitly recorded, but also includes other elements not explicitly listed, or also includes elements inherent to the implementation of the method or device. In the absence of further restrictions, an element defined by the sentence "includes..." does not exclude the presence of other related elements (such as steps in the method or units in the device, such as a unit in the device may be a part of a circuit, a part of a processor, a part of a program or software, etc.) in the method or device including the element.

[0040] The data processing method provided in the embodiment of the present application includes a series of steps, but the data processing method provided in the embodiment of the present application is not limited to the recorded steps. Similarly, the data processing device provided in the embodiment of the present application includes a series of modules, but the device provided in the embodiment of the present application is not limited to including the modules explicitly recorded, and may also include modules required to obtain relevant information or perform processing based on information.

[0041] The present application embodiment provides a data processing method, such as Figure 1 As shown, Figure 1 A flow chart of a data processing method is shown. Figure 1 The data processing method shown includes:

[0042] Step 101: Based on the first relation word of each triple in the triple set, determine a first embedding learning mode of the triple; each triple includes a relation triple and an attribute triple.

[0043] Triples include attribute triples and relation triples. Attribute triples are usually composed of entities, relation words and attribute values. For example, in the triple <apple, color, red>, apple is the entity, color is the relation word, and red is the attribute value. Relation triples are usually composed of head entities, relation words and tail entities. For example, in the triple <cat, chase, mouse>, cat is the head entity, chase is the relation word, and mouse is the tail entity. In general, triples include entities, relation words and attribute values. Entities are divided into head entities and tail entities according to their relationship with relation words.

[0044] In natural language processing, triples are the basic units that constitute the knowledge graph KG. By extracting triples from text, a knowledge graph containing rich entities and relationships can be constructed to provide support for subsequent semantic search, intelligent answering and other tasks. In the process of obtaining triples, the original text can be preprocessed by data cleaning, word segmentation, part-of-speech tagging, etc., to remove irrelevant characters in the original text, and the named entity recognition technology can be used to identify entities with specific meanings in the preprocessed text, such as names, place names, organization names, etc. On the basis of entity recognition, the relationship between entities is identified in the preprocessed text through relationship extraction technology, and finally the corresponding triples are generated through the identified entities and relationships. In order to improve the accuracy of data processing, multiple different types of texts can be used to extract triples in the embodiment of the present application, and then embedded learning is performed based on the extracted triples. The language of different types of texts should be representative and diverse.

[0045] In the process of obtaining triples, the obtained entities can be used as nodes in the knowledge graph, and the obtained relationship words can be used as edges of the knowledge graph to construct knowledge graphs corresponding to different texts.

[0046] The triple set in this embodiment includes triples extracted from different contexts or different texts. Specifically, the triple set in this embodiment can come from two or more knowledge graphs, and the triples in the two knowledge graphs can be combined in their original form. Each entity, relational word, and attribute value in the triple set are in their original form, and entity unification has not been performed. Among them, for different entities or attribute values ​​in each knowledge graph, the type, characteristics, part of speech, and other information of the entity or attribute value can be pre-annotated through the word segmentation model.

[0047] After obtaining the triple set, multiple subsets can be obtained based on the triples in the triple set. The subsets in the triple set can include the verb triple Tp set, the attribute triple Ta set, and the relation triple Tr set. The triples in the verb triple Tp set are composed of the first head entity before the first relation word, the first relation word, and the first tail entity or the first attribute value after the first relation word. The verb triple Tp set can also be composed of each triple in the triple set; the triples in the attribute triple Ta set are composed of the second head entity, the second relation word, and the second attribute value; the triples in the relation triple Tr set are composed of the fourth head entity, the third relation word, and the second tail entity. In practical applications, the triples in the verb triple Tp set may have the same triples as the triples in the attribute triple Ta set, and the triples in the verb triple Tp set may have the same triples as the triples in the relation triple Tr set.

[0048] In the embodiment of the present application, the first relational word represents the relational word in each triplet used to determine the first embedding learning method, the second relational word represents the relational word in each attribute triplet used to determine the second embedding learning method, and the third relational word represents the relational word in each relation triplet used to determine the third embedding learning method. In the triplet set, the first relational word and the second relational word may be the same relational word; the first relational word and the third relational word may also be the same relational word; the second relational word and the third relational word come from different triples.

[0049] In the embodiment of the present application, the first relational terms in each triple may be the same or different; the second relational terms in each attribute triple Ta may be the same or different; the third relational terms in each relation triple Tr may be the same or different. The first head entity in each verb triple Tp, the second head entity in each attribute triple Ta, and the fourth head entity in each relation triple Tr may be the same or different; the tail word in each verb triple Tp, the second attribute value in each attribute triple Ta, and the second tail entity in each relation triple Tr may be the same or different.

[0050] In practical applications, semantic analysis can be performed on the first head entity before each first relation word, and semantic analysis can be performed on the tail word after each first relation word. The similarity between any two first head entities and the similarity between any two tail words can be determined through semantic analysis. Based on the similarity between any two first head entities and the similarity between any two tail words, the similarity between any two first relation words can be determined; based on the similarity between any two first relation words, the first embedding learning method can be determined.

[0051] Alternatively, based on the knowledge graph embedding technology, the first head entity, the first relation word, and the tail word in the knowledge graph can be embedded into a continuous vector space, so as to retain the structural information in the knowledge graph while facilitating calculation, and further determine the first embedding learning method. For example, each verb triple Tp composed of <first head entity, first relation word, tail word> can be embedded based on any one of the translation models (such as TransE, TransH, TransR, etc.), decomposition-based models (such as DistMult, ComplEx, etc.), or neural network-based models (such as R-MeN, ConvE, CapsE, etc.).

[0052] Step 102: Based on each attribute triple in the triple set, determine a second embedding learning method for the triple; based on each relationship triple in the triple set, determine a third embedding learning method for the triple.

[0053] In this step, the fourth head entity, the third relation word, and the second tail entity in each relation triple Tr can be obtained based on the set of relation triples Tr. After obtaining the relation triple Tr determined by the fourth head entity, the third relation word, and the second tail entity, each relation triple Tr consisting of <fourth head entity, third relation word, second tail entity> can be embedded based on any one of the models in the translation model (such as TransE, TransH, TransR, etc.), the decomposition-based model (such as DistMult, ComplEx, etc.), or the neural network-based model (such as R-MeN, ConvE, CapsE, etc.). Here, the third embedding learning method should select the same model for embedding learning as the first embedding learning method and the second embedding learning method.

[0054] It can be seen that by embedding the relation triple Tr composed of <fourth head entity, third relation word, second tail entity>, the similarity of the structural information composed of entities and relations in the triple is determined, so that entities with similar structures have the same entity representation, and the learning of structural embedding is realized. Taking the TransE embedding learning method as an example, each relation triple Tr includes a head entity (head, h), a relation word composed of a predicate (predicate, p), and a tail entity (tail, t), that is, each relation triple Tr can be expressed as<h,p,t> In the form of, the vector of the head entity h of each relation triple Tr should be similar to the vector representation of the tail entity t after being translated by the vector of the relation word r, that is, h+p≈t. In order to learn structural embeddings, TransE minimizes the margin-based objective function JSE:

[0055] JSE=∑max{0,[γ+f(tr)-f(tr′)] (1)

[0056] f(tr)=||h+pt||2 (2)

[0057] Tr = {<h,p,t> |<h,p,t> ∈G} (3)

[0058] Tr′={<h′,p,t> |h′∈∈}∪{<h,p,t′> |t′∈∈} (4)

[0059] Among them, ||x||2 represents the L2 norm of vector x, γ is a marginal hyperparameter, and γ is a constant representing the distance between positive samples (relation triples) and negative samples (damaged relation triples), which is used to ensure that the score of the relation triple is lower than the score of the damaged relation triple. Tr represents the set of relation triples, Tr′ is the set of damaged relation triples, ∈ is the entity set in G, and G represents the set of triples. The damaged relation triple Tr′ is used as a negative sample, which represents the triple obtained by replacing the fourth head entity h or the second tail entity t in the relation triple Tr.

[0060] In this step, the second head entity, the second relation word, and the second attribute value can also be obtained based on the attribute triple Ta set. After determining the attribute triple Ta consisting of the second head entity, the second relation word, and the second attribute value, each attribute triple Ta constructed by <second head entity, second relation word, second attribute value> can be embedded based on any one of the models in the translation model (such as TransE, TransH, TransR, etc.), the decomposition-based model (such as DistMult, ComplEx, etc.), or the neural network-based model (such as R-MeN, ConvE, CapsE, etc.). Here, the second embedding learning method should select the same model for embedding learning as the first embedding learning method and the third embedding learning method.

[0061] It can be seen that this step can also use the attribute value information in the attribute triple Ta to enrich the entity representation. Each attribute triple Ta is composed of a head entity h, a relation word p, and an attribute value (value, v). Based on the above formula, by<h,p,v> Attribute embedding helps entities with similar attributes have similar representations in the vector space.

[0062] Step 103: Perform joint embedding learning on the triples through the first embedding learning method, the second embedding learning method, and the third embedding learning method.

[0063] Through the first embedding learning method of determining the relevance of the first relation word in the triple, the second embedding learning method of determining the attribute embedding in the triple, and the third embedding learning method of determining the structural relationship of the entity in the triple, the three learning methods are combined for joint learning, and the representation of the entity in the vector space is made more accurate through joint learning.

[0064] In a specific implementation, in the process of joint embedding learning based on the first embedding learning method, the second embedding learning method, and the third embedding learning method, a unified objective function can be determined, and the loss can be considered simultaneously based on the unified objective function. By minimizing the objective function, it is ensured that the vector representations of the structural information, attribute information, and relational word information can be simultaneously brought close, thereby achieving accurate embedding of the triples.

[0065] An embodiment of the present application provides a data processing method, which improves the embedding accuracy of the triples by respectively determining a first embedding learning method, a second embedding learning method, and a third embedding learning method, and determining the embedding of the triples through joint embedding learning.

[0066] In practical applications, steps 101 to 103 can be implemented based on a processor, and the processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor.

[0067] In some embodiments, the above-mentioned first embedding learning method of determining the triple based on the first relation word of each triple in the triple set includes: in each triple in the triple set, obtaining the first target type of the first head entity before the first relation word, and the second target type of the tail word after the first relation word; based on the first relation word, the first target type, and the second target type, determining the first embedding learning method of the triple; the tail word includes the first tail entity or the first attribute value.

[0068] In this embodiment, the first head entity, the first relational word, and the tail word in each verb triple Tp are obtained in the verb triple Tp set. The first embedding learning method determined in this embodiment is mainly used for embedding learning of relational words in the triple set, that is, it is mainly used for learning to determine the similarity between each relational word in the triple set.

[0069] The first relational word is used to describe the relationship between the first head entity and the tail word in the triple. Therefore, in order to determine the similarity of each first relational word in the triple and realize the correct embedding of the first relational word, it is necessary to consider the relationship between the first head entity and the tail word connected to it in the embedding of the first relational word. In different KGs, there may be different first head entities in different KGs, or different tail words in different KGs, but in different KGs, some first head entities refer to the same real-life affairs, or, in different KGs, some tail words refer to the same real-life affairs. When the entity or attribute value of the triples in the triple set is not unified, for example, the triples in the triple set come from different knowledge graphs KG1 and KG2, the entity corresponding to the cat in KG1 is "cat", and the entity corresponding to the cat in KG2 is "MAO" or "cat". In this case, the entity "cat" in the triple set is different from the entity "MAO" or "cat", but they both refer to the animal cat in reality. In this case, the first relation words corresponding to "cat" in KG1 and KG2 may be similar. It can be seen that the description forms of the same entity in the real world may be different in KG, so it is necessary to identify the essential relationship between entities in different KGs, and then determine the similarity between the first relation words in each KG.

[0070] In order to achieve accurate embedding of the relational words, the first target type of the first head entity and the second target type of the tail word are extracted in this embodiment, and the relational word proximity graph is constructed by replacing the first head entity with the first target type and the tail word with the second target type. The relational word proximity graph has the first target type, the first relational word, and the second target type as nodes, and the relationship between the first head entity, the first relational word, and the tail word as edges.

[0071] In this embodiment, the first target type and the second target type can be specifically determined in the knowledge graph. The first target type of the first head entity and the second target type of the tail word can be determined in the knowledge graph by querying the value of the rdfs:type attribute corresponding to the entity or attribute value.

[0072] After constructing the neighbor graph of the relational words, <first target type, first relational word, second target type> can be used as a new triple to determine the first embedding learning method. For example, the triple <first target type, first relational word, second target type> can be embedded in any model based on translation models (such as TransE, TransH, TransR, etc.), decomposition-based models (such as DistMult, ComplEx, etc.), or neural network-based models (such as R-MeN, ConvE, CapsE, etc.).

[0073] It can be seen that the method provided in this embodiment can realize the embedding learning of the first relational word based on the type of the first head entity and the type of the tail word, which is further conducive to obtaining the similarity of the relational words in the triple.

[0074] In some embodiments, the above-mentioned acquisition of the first target type of the first head entity before the first relationship word, and the second target type of the tail word after the first relationship word, includes: acquiring a first type set of the first head entity before the first relationship word, and a second type set of the tail words after the first relationship word; wherein the first type set includes multiple types corresponding to the preset first head entity; the second type set includes multiple types corresponding to the preset tail words; the first target type is determined by weighting the vector of each type in the first type set and the weight of each type in the first type set; the second target type is determined by weighting the vector of each type in the second type set and the weight of each type in the second type set.

[0075] In this embodiment, each entity or attribute value in the triple set is preset with multiple types, and the type of the entity or attribute value indicates the meaning of the entity or attribute value in different triples. For example, in the triple <fishing village, located in, southeast>, the type set of the attribute value "southeast" can be {place, location, direction}. For different entities or attribute values, according to their meanings in different contexts, multiple corresponding types are preset, that is, the preset type set corresponding to the attribute value "southeast" can be {place, location, direction}.

[0076] After determining the first head entity and the tail word in the verb triple Tp, obtain the first type set corresponding to the first head entity and the second type set corresponding to the tail word. Since the type sets contain multiple preset types, based on the vectors of each type in the multiple types and the preset weights of each type, the vectors of the multiple types corresponding to the first head entity are weighted summed, and the vector obtained after the weighted summation is used as a pseudo type, and the pseudo type is determined as the first target type corresponding to the first head entity; the vectors of the multiple types corresponding to the tail word are weighted summed, and the vector obtained after the weighted summation is used as a pseudo type, and the pseudo type is determined as the second target type corresponding to the tail word, so as to obtain the most unique type corresponding to each entity or attribute value.

[0077] When the triples in the triple set come from different knowledge graphs KG, due to the different contexts corresponding to different KGs, the type sets corresponding to the same entity or attribute value in different KGs may also be different. Therefore, in different KGs, the first head entity and tail word before and after each first relation word may be replaced by different pseudo-types. Based on the method given in this embodiment, the first target type and the second target type are obtained by aggregating the types of the first head entity and tail word before and after the relation word, and the first head entity before the first relation word and the tail word after the first relation word are replaced by the target type obtained by aggregation, which can intuitively reflect the key features of the first head entity and tail word in different knowledge graphs and realize accurate embedding of the first relation word.

[0078] In some embodiments, before determining the first target type and determining the second target type, the above method also includes: determining the weight of the first type based on the number of the first type in the triple set type; the first type is any one type in the first type set and the second type set; the triple set type includes multiple types corresponding to each first head entity, and multiple types corresponding to each tail word.

[0079] Specifically, multiple types corresponding to each entity and attribute value in the triple set can be summarized to obtain a total set of types, the total set of types contains multiple different types, and the total set of types contains multiple repeated types.

[0080] For example, Table 1 shows a subset G1 of triples of KG1:

[0081] Table 1

[0082] <G1:240111203, Population Quantity, 1595> <G1:240111203, Label, "Fishing Village"> <G1:240111203, Geographical latitude, 90.9988888889> <G1:240111203, Person in Charge, "XXX"> <G1:240111203, at, G1:110000>

[0083] Table 2 shows a subset G2 of triples of KG2:

[0084] Table 2

[0085] <G2: Fishing village, label, "Fishing village"> <G2: Fishing village, geographical latitude, 90.9989> <G2: Fishing village, total population, 1595> <G2: Fishing village, located in, China> <G2: Fishing village, the town it belongs to, Shipu Town>

[0086] Among them, "a village called Fishing Village is located in China" is a real-world entity that exists in two different KGs (KG1 and KG2). As shown in Table 1 and Table 2. The entity is represented in KG1 as "G1:240111203", but in KG2 as "G2: Fishing Village". The number of these entities in each KG and the corresponding types are complementary, so the two KGs can be merged to obtain a set of triples.

[0087] Taking the triples in G1 and G2 shown in Table 1 and Table 2 as an example, the type set of the first tail entity "China" is preset as U1 = {affairs, country, location, place}, the type set of the first tail entity "Shipu Town" is preset as U2 = {location, place}, and the type set of the first tail entity "1595" is preset as U3 = {quantity, value, coordinates}, etc. The types of each head entity and tail word in G1 and G2 are summarized accordingly, and the total set of types is obtained as U = U1∪U2∪U3∪…∪UN.

[0088] When the weight of the type "position" needs to be determined, the weight of "position" is determined based on the number of occurrences of "position" in the total set U of types.

[0089] In the specific implementation process, weights can also be assigned to types based on their particularity. For example, the type "country" has higher particularity than other types, so a larger weight can be assigned to the type "country". Therefore, when determining the target type of the first tail entity "China", the target type corresponding to "China" can better reflect the type characteristics of "country" rather than the type characteristics of "affairs".

[0090] Based on the method for determining the weight of a type given in this embodiment, another method for determining the weight of a type is given below. Assume that the type set of a given first tail entity has a total of m types preset, including a set of m types The weight w corresponding to each type can be determined according to formula (5): i :

[0091]

[0092] in, is a vector representation of type, l i Indicates the specificity level of the preset type, that is, according to the preset classification standard, the more special the type, the higher the specificity level. i The larger the value of , for example, the specificity level of "country" is higher than that of "transaction"; a represents the number of types in Ux, that is, the size of m; r i Indicates the number of times this type appears in the KG, that is The number of types in the total set U corresponding to KG; k i Indicates that it contains type z i For example, the type of entity or attribute value in KG1 includes "position", and the type of entity or attribute value in KG2 also includes "position", that is, the k corresponding to type "position" i Equal to 2. The weights are converted into probability distributions for each type in Ux through the softmax function.

[0093] After obtaining the weight of each type through formula (5), the target type u corresponding to the set Ux of m types is determined based on formula (6):

[0094]

[0095] Based on formula (5) and formula (6), the target type of each entity or attribute value can be obtained.

[0096] In some embodiments, before determining the first target type and determining the second target type, the above method also includes: based on the first correlation between the pre-labeled type and the entity, in the first type set and the second type set, deleting the type whose first correlation with the entity is less than a first correlation threshold.

[0097] Specifically, the attention algorithm can be used to ignore potential type noise, that is, the attention algorithm is used to remove types whose first correlation with the entity is less than the first correlation threshold. For example, the type set of the first tail entity "China" is preset as U1 = {business, country, location, place}, where the type "business" ∈ U1, but in the triple <fishing village, located, China>, the type "business" is not related to the first tail entity "China", so a smaller first correlation can be annotated for the type "business" in U1. In this case, the type "business" as a type of noise can be ignored for the embedding of the first relation word "located", so the type "business" can be ignored when determining the target type before and after "located".

[0098] Based on the methods of calculating target types given by the above formulas (5) and (6), another method of calculating target types is given below. Formula (7) gives a method of calculating the i-th type z i The attention weight Method:

[0099]

[0100] Among them, w z Representation Type The trainable weight matrix, the final target matrix is ​​obtained by vectorizing all types representing different meanings Weighted addition yields:

[0101]

[0102] It can be seen that by determining the target type through the weighted method given in the above embodiment and embedding the first relational word based on the target type, it is possible to embed the relational words in different KGs into the same vector space, thereby improving the accuracy of embedding the first relational word.

[0103] In some embodiments, the above-mentioned first embedding learning method for determining triples based on the first relation word, the first target type, and the second target type includes: constructing a first damaged triple corresponding to each triple based on the first relation word, the first target type, and the second target type; each first damaged triple includes the first relation word, the third target type, and the fourth target type; in each triple, obtaining a first vector by summing the vector of the first target type and the vector of the first relation word; calculating a first difference between the first vector and the vector of the second target type; in each first damaged triple, obtaining a second vector by summing the vector of the third target type and the vector of the first relation word; calculating a second difference between the second vector and the vector of the fourth target type; the first difference corresponds to the second difference one-to-one; constructing a first objective function based on the difference between the first difference and the corresponding second difference, and determining the first embedding learning method for the triple with the goal of minimizing the first objective function.

[0104] This embodiment further provides a method for specifically determining the first embedding learning mode. The first damaged triple in this embodiment represents an erroneous triple obtained by replacing the first target type or the second target type corresponding to any triple after obtaining the first target type and the second target type.

[0105] Taking any triple as an example, first based on the first relational word, the first target type, and the second target type, the third target type or the fourth target type can be obtained by replacing the first target type or the second target type with an incorrect target type, and a first damaged triple is constructed based on the third target type, the first relational word, and the fourth target type.

[0106] Here, when the first target type is replaced, the fourth target type is the same as the second target type, and when the second target type is replaced, the third target type is the same as the first target type. Specifically, the first target type or the second target type can be replaced by the type of other entities in the triple set or the type of other attribute values.

[0107] Since there is a logical connection between the first target type, the first relation word, and the second target type, ideally, it is expected that the first target type + the first relation word ≈ the second target type. That is, the embedding of the second target type of the tail word should satisfy the embedding of the first target type plus the embedding of the first relation word. Therefore, the translation property of h+p≈t should be satisfied during modeling, where h corresponds to the first target type and t corresponds to the second target type. Assuming that the first target type h is replaced by the third target type h′, the corresponding modeling of the first damaged triple should satisfy the property of h′+p≈t.

[0108] Based on the method given in this embodiment, the first objective function J is constructedpe As shown in formula (9):

[0109]

[0110] Among them, f(tp) represents the verb triple T based on any verb p After extracting the first target type and the second target type, the first difference f(tp′) is obtained, which represents the difference between the verb triple T and the verb triple T. p The corresponding first damaged triplet T p The second difference obtained by ′ can be specifically calculated by using formula (10) and formula (11), and the first difference and the second difference can be determined by using the L2 norm:

[0111] f(tp)=||u hp +pu tp ||2 (10)

[0112] f(tp′)=||u hp +pu tp′ ||2 (11)

[0113] Wherein, when the second target type is replaced, u in formula (10) hp Indicates the first target type, u tp represents the second target type, u in formula (11) hp Indicates the first target type, u tp′ Corresponding to the method given in this embodiment, when the second target type is replaced, the first target type in formula (11) corresponds to the first damaged triplet T p The third target type in '. hp and u tp Specifically, it can be obtained by formula (6) or formula (8).

[0114] To minimize J pe The value of is the target and determines the first embedding learning method.

[0115] In practical applications, the embedding learning method can be determined based on TransE. TransE is a method that embeds entities and relationships into a low-dimensional vector space and performs embedding on the triples in the vector space.<h,p,t> TransE represents the relationship between a pair of entities as a transformation between entity embeddings:<h,p,t> The formed triple should have the property of h+p≈t.

[0116] In some embodiments, each attribute triple includes a second head entity, a second relational word, and a second attribute value; the above-mentioned second embedding learning method of determining the triple based on each attribute triple in the triple set includes: splitting the vector of the second attribute value in each attribute triple to obtain a first N-tuple sequence of each attribute triple; N is a positive integer greater than or equal to 2; constructing a damaged attribute triple corresponding to each attribute triple; wherein each damaged attribute triple includes a third head entity, a second relational word, and a second N-tuple sequence; in each attribute triple, a third vector is obtained by summing the vector of the second head entity and the vector of the second relational word; a third difference between the third vector and the first N-tuple sequence is calculated; in each damaged attribute triple, a fourth vector is obtained by summing the vector of the third head entity and the vector of the second relational word; a fourth difference between the fourth vector and the second N-tuple sequence is calculated; the third difference corresponds one-to-one to the fourth difference; a second objective function is constructed based on the difference between the third difference and the corresponding fourth difference, and the second embedding learning method of the triple is determined with the goal of minimizing the second objective function.

[0117] This embodiment further provides a method for attribute embedding. The damaged attribute triple in this embodiment represents an erroneous triple corresponding to the attribute triple, that is, a negative sample corresponding to the attribute triple. When constructing the damaged attribute triple, it can be ensured that the second head entity or the second attribute value in the attribute triple of each positive sample is replaced only once.

[0118] In the triple, the second relational term represents the conversion relationship from the second head entity to the second attribute value. In a set of triples, especially in a set of triples composed of two different KGs, an attribute value may have multiple representations. For example, the dimensions "50.8889" and "50.88898546" are both attributes of an entity, and "Mr. Z" and "Mr. W" are both titles of people. In this embodiment, the relationship of the attribute triple is established as: h+p=fa(v). Among them, fa(v) represents a combinatorial function, v represents the vector of attribute values, v={c1,c2,c3,…,ct}, and v consists of characters. Specifically, v represents the vector of the second attribute value, and c1,c2,c3,…,ct is the first N-tuple sequence obtained by splitting the vector v. In practical applications, fa(v) can be calculated based on the combinatorial function of N-gram:

[0119]

[0120] Where N is the maximum value of n used in the N-gram combination, c j represents the characters in v, t represents the number of characters in v, and l is the length of the vector of the second attribute value. Based on the method given in this embodiment, the second objective function is constructed:

[0121]

[0122] Where f(ta) represents the third difference corresponding to any attribute triple, f(ta′) represents the fourth difference corresponding to the damaged attribute triple corresponding to the attribute triple, Ta represents the attribute triple set consisting of the second head entity, the second relation word, and the first N-tuple sequence, T a ' represents the set of damaged attribute triples obtained by replacing the second head entity or the first N-tuple sequence. When only the second head entity is replaced, the second N-tuple sequence is the same as the first N-tuple sequence; when only the first N-tuple sequence is replaced, the second head entity is the same as the third head entity. f(ta) in formula (13) can be obtained based on formula (14):

[0123] f(ta)=||h+p-fa(v)||2;T a ={<h,p,v> ∈G U} (14)

[0124] T a ′={<h′,p,v> |h′∈∈ u}∪{<h,p,v′ > |v′∈A u} (15)

[0125] Among them, T a Represents each correct attribute triple, T a It is composed of the second head entity, the second relation word, and the first N-tuple sequence. a ′ represents each damaged attribute triple, and f(ta) is the confidence score calculated based on the embedding of the second head entity h, the embedding of the second relation word p, and the vector representation of the attribute value calculated using the function fa(v). U Represents a set of attribute triples, A u Contains G u The triple consisting of entities and attribute values ​​excluding the second attribute value in the currently embedded attribute triple; ∈ u Contains G u The triple consisting of entities and attribute values ​​excluding the second head entity in the currently embedded attribute triple.

[0126] In practical applications, we can also add weight factors Adjust the objective function J in formula (13) ce , as shown in formula (16):

[0127]

[0128] Among them, the weight factor It can be determined based on formula (17):

[0129]

[0130] Among them, count(r) is the number of occurrences of the relation word r, and |T| is the total number of triples in the merged KG, that is, the total number of triples in the triple set.

[0131] To minimize J ce The value of is the target and determines the second embedding learning method.

[0132] In some embodiments, each relation triple includes a fourth head entity, a third relation word, and a second tail entity; the above-mentioned method of determining the third embedding learning method of the triple based on each relation triple in the triple set includes: constructing a damaged relation triple corresponding to each relation triple; the damaged relation triple represents a triple obtained by replacing the fourth head entity or the second tail entity in each relation triple with an erroneous entity; based on each damaged relation triple and each relation triple, constructing a third objective function, and determining the third embedding learning method of the triple with the goal of minimizing the third objective function; the above-mentioned method of determining the third embedding learning method of the triple through the first embedding learning method and the second embedding learning method Learning method, and a third embedding learning method, before jointly embedding learning the triples, the method also includes: determining a first similarity between the second objective function and the third objective function; constructing a fourth objective function based on the first similarity; the value of the fourth objective function is negatively correlated with the first similarity; determining a fourth embedding learning method with the goal of minimizing the fourth objective function; the above-mentioned joint embedding learning of the triples through the first embedding learning method, the second embedding learning method, and the third embedding learning method, includes: jointly embedding learning of the triples through the first embedding learning method, the second embedding learning method, the third embedding learning method, and the fourth embedding learning method.

[0133] Corresponding to the method given in the above embodiment, based on the relation triple Tr composed of the fourth head entity, the third relation word, and the second tail entity, a damaged relation triple corresponding to each relation triple Tr is constructed. The damaged relation triple in this embodiment represents an erroneous triple corresponding to the relation triple Tr, that is, a negative sample corresponding to the relation triple Tr. When constructing the damaged relation triple, it can be ensured that the fourth head entity or the second tail entity in the relation triple Tr of each positive sample is replaced only once.

[0134] When only the fourth head entity is replaced, the tail entity in the damaged relation triple is the same as the second tail entity; when only the second tail entity is replaced, the head entity in the damaged relation triple is the same as the fourth head entity.

[0135] Based on the weight factor given by formula (17) The third objective function can be expressed as:

[0136]

[0137] Among them, T r Represents each relation triple, T r ′ represents each damage relation triple. f(tr) can be obtained based on formula (2).

[0138] It can be seen that the third objective function J se It mainly focuses on the structural embedding of the form of <entity, relation word, entity>. The second objective function J ce It mainly focuses on the attribute embedding of the form of <entity, relation word, attribute value>. In order to further improve the accuracy of triple embedding and determine which entities in different triples or different KGs refer to the same entities in reality, this embodiment further provides a method for determining entity similarity. ce And the third objective function J se Translate to the same vector space and calculate the second objective function J ce And the third objective function J se The fourth objective function is further determined by the first similarity.

[0139] In the specific implementation process, the second objective function J can be determined based on the cosine similarity. ce And the third objective function J se The first similarity of can be further used as the fourth objective function based on formula (19):

[0140]

[0141] Among them, cos(J se ,J ce ) is the second objective function J ce And the third objective function J se The cosine similarity of entities can be captured through formula (19).

[0142] Based on the above formula (9), formula (16), formula (18), and formula (19), the objective function of joint learning is obtained:

[0143] J=J pe +J se +J ce +J sim (20)

[0144] With the goal of minimizing the objective function J in formula (20), the triples are jointly embedded and learned.

[0145] The embodiment of the present application provides a data processing method. When embedding the triples in the two KGs given in Table 1 and Table 2 above, based on the method given in the above embodiment, the same entities in Table 1 and Table 2 can be identified, so that the same entities in Table 1 and Table 2 can be effectively fused, and the same entities are merged into one KG. Figure 2 The figure shows a flowchart of embedding the triple data in Table 1 and Table 2 based on the method given in the embodiment of the present application, including:

[0146] Step 201: Triple merging.

[0147] Merge the triples in Table 1 and Table 2 to obtain the merged triple set G U = G1 ∪ G2. At this time, the entities in G U have not been unified yet. Based on the triples in G U three types of triple sets are obtained: the verb triple set Tp, the attribute triple set Ta, and the relationship triple set Tr. For example, <village, located in, country> represents a relationship triple, where "village" is the head entity and "country" is the tail entity, and a head entity or a tail entity is preset with multiple entity types.

[0148] Step 202: Triple embedding.

[0149] Specifically, it includes Step 2021: Constructing the verb triple set.

[0150] The triples in the verb triple set Tp are all composed of the first head entity, the first relational word, and the tail vocabulary. Combining the triples shown in Table 1 and Table 2, the verb triple set Tp may specifically include: <village, G1: in, country>, <village, G2: located in, country>, <village, G1: population, integer>, etc.

[0151] Step 2022: Relational word embedding.

[0152] Based on the method given in the above embodiment, obtain the first target type of the first head entity in each verb triple and the second target type of the tail vocabulary, and realize the embedding of the first relational word through the first objective function given by formula (9).

[0153] Step 2023: Constructing the attribute triple set.

[0154] The triples in the attribute triple set Ta are all composed of the second head entity, the second relational word, and the second attribute value. Combining the triples shown in Table 1 and Table 2, the attribute triple set Ta may specifically include: <G1: 240111203, label, fishing village>, <G2: fishing village, label, fishing village>, etc.

[0155] Step 2024: Attribute Embedding.

[0156] Based on the method given in the above embodiments, obtain the first N - tuple sequence corresponding to the third attribute value in each attribute triple, and implement attribute embedding through the second objective function given by formula (16).

[0157] Step 2025: Construct a set of relation triples.

[0158] Each triple in the set of relation triples Tr is composed of a fourth head entity, a third relation word, and a second tail entity. Combining the triples shown in Table 1 and Table 2, the set of relation triples Tr may specifically include: <G1:240111203, G1: is in, G1: 110000>, <G2: fishing village, G2: is located in, China>, etc.

[0159] Step 2026: Structure Embedding.

[0160] Based on the method given in the above embodiments, implement the structure embedding of the triples through the third objective function given by formula (18).

[0161] Step 203: Joint Learning.

[0162] Based on the objective function given by formula (20) above, perform joint embedding learning on each triple with the aim of minimizing the objective function.

[0163] Based on Figure 2 the shown embedding method flow chart, embed the triples in Table 1 and Table 2 to obtain the merged triple G3 shown in Table 3:

[0164] Table 3

[0165] <G1:240111203, Population, 1595> <G1:240111203, Label, "Fishing Village"> <G1:240111203, Geographical Latitude, 90.9988888889> <G1:240111203, Person in Charge, "XXX"> <G1:240111203, at, 110000> <G1:240111203, Town where located, Shipu Town>

[0166] It can be seen that although the head entities in the two subsets G1 and G2 of Table (1) and Table (2) have different forms, they both point to the same entity "fishing village", that is, "G1:240111203" and "G2: fishing village" actually refer to the same thing. Through the method given in the embodiments of the present application, such entities can be identified and given a unified identity identifier so that the two knowledge graphs can be merged together through the unified identity identifier. In Table 3, "G1:240111203" is used as the unified identity identifier of the entity "fishing village", and this entity has a set of attributes, which is the union of the attributes of the two knowledge graphs.

[0167] For Figure 2 the shown embedding method, Figure 3 further shows the flow chart of relation word unification and entity unification, including:

[0168] Step 301: Obtain triples in different KGs.

[0169] Obtain triples in Table 1 and Table 2. For example, obtain the triple <G1: 240111203, G1: at, G1: 110000> in G1 and the triple <G2: Fishing Village, G2: located in, G2: China> in G2.

[0170] Step 302: Calculate similar relation words.

[0171] Calculate the similarity between "G1: at" and "G2: located in", that is, based on the method given in the above embodiment, based on the first target type of "G1: 240111203" before "G1: at" and the second target type of "G1: 110000" after "G1: at", perform the embedding of "G1: at". Based on the first target type of "G2: Fishing Village" before "G2: located in" and the second target type of "G2: China" after "G2: located in", perform the embedding of "G2: located in". Determine the similarity between "G1: at" and "G2: located in" based on the embedded vectors.

[0172] Step 303: Unify relation words.

[0173] The relation word embeddings of the same relation in two KGs should have similar embeddings in the same vector space. For example, "G1: at" and "G2: located in" should have close embeddings. At this time, "G1: at" and "G2: located in" can be unified.

[0174] Step 304: Calculate entity similarity.

[0175] Based on the triples in Table 1 and Table 2 obtained in Step 301, for example, the triple <G1: 240111203, G1: at, G1: 110000> in G1 and the triple <G2: Fishing Village, G2: located in, G2: China> in G2, calculate the similarity between "G1: 240111203" and "G2: Fishing Village", and calculate the similarity between "G1: 110000" and "G2: China". Specifically, the structural embedding can be determined based on the third objective function given in the above embodiment to calculate the entity similarity. If entity 1 in G1 and entity 2 in G2 correspond to the same real-world entity, entity 1 should have a similar embedding to entity 2 in the same vector space. For example, "G1: 240111203" in Table 1 and "G2: Fishing Village" in Table 2 should have similar embeddings.

[0176] Step 305: Unify entities.

[0177] Merge and unify "G1: 240111203" in Table 1 and "G2: Fishing Village" in Table 2 that have similar embeddings.

[0178] The embodiment of the present application provides a data processing method that can achieve entity unification based on natural language processing of a low corpus, that is, identify whether the entities corresponding to the same real world in two knowledge graphs (KGs) are the same. In order to accurately judge the similarity of entities, the embodiment of the present application combines the embedding, attribute embedding and structure embedding of relational words in triples, first constructs a relational word proximity graph, captures the relationship of relational words as entity types, and then uses relational word embedding to jointly learn structural embedding and attribute embedding, uses a combination function to encode attribute values, calculates the minimized objective function to judge the similarity between two entities, and moves the entity embeddings in the two KGs to the same vector space. The method given in the embodiment of the present application can perceive subtle changes in semantics in different knowledge graphs. After the knowledge graphs are merged and embedded based on the embedding method given in the embodiment of the present application, the relational structure missing from the context can be effectively identified through the embedded knowledge graph.

[0179] In the embedding method of relational words given in the embodiment of the present application, the relational words are captured as the process of entity type relations, multiple entity types in the relational word proximity graph are aggregated into vector representations of relational words using pseudo-type embedding, entity type weights are evaluated and potential type noise is ignored. This method not only considers the weights of unique types in entity types in verb triples, but also considers the negligible type noise that may exist in entity types, solves the problem of transitive errors between multiple tasks, and improves embedding accuracy and embedding efficiency.

[0180] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.

[0181] Based on the data processing method proposed in the above embodiment, the present application embodiment also provides a data processing device, such as Figure 4 As shown, the data processing device includes:

[0182] The first processing module 401 is used to determine the first embedding learning method of the triple based on the first relation word of each triple in the triple set; each triple includes a relation triple and an attribute triple; based on each attribute triple in the triple set, determine the second embedding learning method of the triple; based on each relation triple in the triple set, determine the third embedding learning method of the triple.

[0183] The second processing module 402 is used to perform joint embedding learning on the triples through the first embedding learning method, the second embedding learning method, and the third embedding learning method.

[0184] The data processing device provided in the embodiment of the present application can effectively reduce the error of context information transmission through joint embedding learning. By combining the three embedding learning methods, the accuracy of determining the same entity in different triples is improved, which is further conducive to achieving the unification of entities in different triples, improving the accuracy of triple embedding, and improving the accuracy of natural language processing.

[0185] In practical applications, the first processing module 401 and the second processing module 402 can be implemented based on a processor and a communication device.

[0186] In some embodiments, the first processing module 401 is specifically used to obtain, in each triple in the triple set, a first target type of a first head entity before a first relation word, and a second target type of a tail word after the first relation word; based on the first relation word, the first target type, and the second target type, determine a first embedding learning method for the triple; the tail word includes a first tail entity or a first attribute value.

[0187] In some embodiments, the first processing module 401 is specifically used to obtain a first type set of first head entities before the first relationship word, and a second type set of tail words after the first relationship word; wherein the first type set includes multiple types corresponding to the preset first head entity; the second type set includes multiple types corresponding to the preset tail words; the first target type is determined by weighting the vector of each type in the first type set and the weight of each type in the first type set; the second target type is determined by weighting the vector of each type in the second type set and the weight of each type in the second type set.

[0188] In some embodiments, before determining the first target type and determining the second target type, the first processing module 401 is also used to determine the weight of the first type based on the number of the first type in the triple set type; the first type is any one type in the first type set and the second type set; the triple set type includes multiple types corresponding to each first head entity and multiple types corresponding to each tail word.

[0189] In some embodiments, before determining the first target type and determining the second target type, the first processing module 401 is also used to delete types whose first correlation with the entity is less than a first correlation threshold in the first type set and the second type set based on the first correlation between the pre-labeled types and the entities.

[0190] In some embodiments, the first processing module 401 is specifically used to construct a first damaged triple corresponding to each triple based on the first relation word, the first target type, and the second target type; each first damaged triple includes the first relation word, the third target type, and the fourth target type; in each triple, the first vector is obtained by the sum of the vector of the first target type and the vector of the first relation word; the first difference between the first vector and the vector of the second target type is calculated; in each first damaged triple, the second vector is obtained by the sum of the vector of the third target type and the vector of the first relation word; the second difference between the second vector and the vector of the fourth target type is calculated; the first difference corresponds to the second difference one by one; a first objective function is constructed based on the difference between the first difference and the corresponding second difference, and the first embedding learning method of the triple is determined with the goal of minimizing the first objective function.

[0191] In some embodiments, each attribute triple includes a second head entity, a second relational word, and a second attribute value; the first processing module 401 is specifically used to split the vector of the second attribute value in each attribute triple to obtain a first N-tuple sequence of each attribute triple; N is a positive integer greater than or equal to 2; construct a damaged attribute triple corresponding to each attribute triple; wherein each damaged attribute triple includes a third head entity, a second relational word, and a second N-tuple sequence; in each attribute triple, a third vector is obtained by summing the vector of the second head entity and the vector of the second relational word; a third difference between the third vector and the first N-tuple sequence is calculated; in each damaged attribute triple, a fourth vector is obtained by summing the vector of the third head entity and the vector of the second relational word; a fourth difference between the fourth vector and the second N-tuple sequence is calculated; the third difference corresponds one-to-one to the fourth difference; a second objective function is constructed based on the difference between the third difference and the corresponding fourth difference, and a second embedding learning method for the triple is determined with the goal of minimizing the second objective function.

[0192] In some embodiments, each relation triple includes a fourth head entity, a third relation word, and a second tail entity; the first processing module 401 is specifically used to construct a damaged relation triple corresponding to each relation triple; the damaged relation triple represents a triple obtained by replacing the fourth head entity or the second tail entity in each relation triple with an erroneous entity; based on each damaged relation triple and each relation triple, a third objective function is constructed, and a third embedding learning method for the triple is determined with the goal of minimizing the third objective function; before the triple is jointly embedded by the first embedding learning method, the second embedding learning method, and the third embedding learning method, the first processing module 401 is also used to determine a first similarity between the second objective function and the third objective function; a fourth objective function is constructed based on the first similarity; the value of the fourth objective function is negatively correlated with the first similarity; the fourth embedding learning method is determined with the goal of minimizing the fourth objective function; the second processing module 402 is specifically used to jointly embed the triple by the first embedding learning method, the second embedding learning method, the third embedding learning method, and the fourth embedding learning method.

[0193] It should be noted that the description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the same method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0194] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to enable a computer device (which can be a terminal, a server, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0195] An embodiment of the present application also provides an electronic device. Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the electronic device 50 may include:

[0196] The memory 501 is used to store executable instructions.

[0197] The processor 502 is used to implement any of the above-mentioned data processing methods when executing the executable instructions stored in the memory 501.

[0198] The processor 502 may be at least one of an ASIC, a DSP, a DSPD, a PLD, a FPGA, a CPU, a controller, a microcontroller, and a microprocessor.

[0199] The above-mentioned computer-readable storage medium or memory 501 can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM) and other memories; it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0200] An embodiment of the present application further provides a computer storage medium, on which computer executable instructions are stored, and the computer executable instructions are used to implement any one of the data processing methods provided in the above embodiments.

[0201] Correspondingly, an embodiment of the present application further provides a computer program product, which includes computer executable instructions, and the computer executable instructions are used to implement any one of the data processing methods provided in the above embodiments.

[0202] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0203] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0204] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0205] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0206] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0207] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0208] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A data processing method, characterized in that: The method comprises: Determine a first embedding learning mode of the triple based on the first relation word of each triple in the triple set; each triple includes a relation triple and an attribute triple; Based on each attribute triple in the triple set, determine a second embedding learning method for the triple; based on each relationship triple in the triple set, determine a third embedding learning method for the triple; The triples are jointly embedded and learned by using the first embedding learning method, the second embedding learning method, and the third embedding learning method.

2. The method according to claim 1, characterized in that The determining of a first embedding learning mode of a triple based on a first relational word of each triple in the triple set includes: In each triple in the triple set, a first target type of a first head entity before the first relation word and a second target type of a tail word after the first relation word are obtained; based on the first relation word, the first target type, and the second target type, a first embedding learning method of the triple is determined; the tail word includes a first tail entity or a first attribute value.

3. The method according to claim 2, characterized in that The obtaining of the first target type of the first head entity before the first relational word and the second target type of the tail word after the first relational word includes: Acquire a first type set of a first head entity before a first relational word, and a second type set of tail words after the first relational word; wherein the first type set includes a plurality of preset types corresponding to the first head entity; and the second type set includes a plurality of preset types corresponding to the tail words; Determine the first target type by weighting the vector of each type in the first type set and the weight of each type in the first type set; The second target type is determined by weighting the vector of each type in the second type set and the weight of each type in the second type set.

4. The method according to claim 3, characterized in that Before determining the first target type and determining the second target type, the method further includes: Based on the number of first types in a triple set type, a weight of the first type is determined; the first type is any one type in the first type set and the second type set; the triple set type includes multiple types corresponding to each first head entity and multiple types corresponding to each tail word.

5. The method according to claim 3, characterized in that: Before determining the first target type and determining the second target type, the method further includes: According to the first correlation between the pre-labeled types and the entity, in the first type set and the second type set, types whose first correlation with the entity is less than a first correlation threshold are deleted.

6. The method according to claim 2, characterized in that The determining a first embedding learning method of a triple based on the first relation word, the first target type, and the second target type includes: Based on the first relation word, the first target type, and the second target type, construct a first damaged triple corresponding to each triple; each first damaged triple includes the first relation word, the third target type, and the fourth target type; In each triple, a first vector is obtained by summing the vector of the first target type and the vector of the first relational word; and a first difference between the first vector and the vector of the second target type is calculated; In each first damaged triplet, a second vector is obtained by summing the vector of the third target type and the vector of the first relational word; a second difference between the second vector and the vector of the fourth target type is calculated; the first difference corresponds to the second difference one by one; A first objective function is constructed based on the difference between the first difference and the corresponding second difference, and a first embedding learning method for the triple is determined with the goal of minimizing the first objective function.

7. The method according to claim 6, characterized in that Each attribute triple includes a second head entity, a second relation word, and a second attribute value; and determining a second embedding learning method of the triple based on each attribute triple in the triple set includes: Splitting the vector of the second attribute value in each attribute triplet to obtain a first N-tuple sequence of each attribute triplet; N is a positive integer greater than or equal to 2; Constructing a damaged attribute triple corresponding to each attribute triple; wherein each damaged attribute triple includes a third head entity, the second relational word, and a second N-tuple sequence; In each attribute triple, a third vector is obtained by summing the vector of the second head entity and the vector of the second relation word; and a third difference between the third vector and the first N-tuple sequence is calculated; In each damaged attribute triple, a fourth vector is obtained by summing the vector of the third head entity and the vector of the second relation word; a fourth difference between the fourth vector and the second N-tuple sequence is calculated; the third difference corresponds to the fourth difference one by one; A second objective function is constructed based on the difference between the third difference and the corresponding fourth difference, and a second embedding learning method for the triple is determined with the goal of minimizing the second objective function.

8. The method according to claim 7, characterized in that Each of the relation triples includes a fourth head entity, a third relation word, and a second tail entity; and determining a third embedding learning method of the triple based on each relation triple in the triple set includes: Constructing a damaged relation triple corresponding to each relation triple; the damaged relation triple represents a triple obtained by replacing the fourth head entity or the second tail entity in each relation triple with an erroneous entity; Based on each damaged relation triple and each relation triple, a third objective function is constructed, and a third embedding learning method of the triple is determined with the goal of minimizing the third objective function; Before performing joint embedding learning on triples by using the first embedding learning method, the second embedding learning method, and the third embedding learning method, the method further includes: Determining a first similarity between the second objective function and the third objective function; Constructing a fourth objective function based on the first similarity; a value of the fourth objective function is negatively correlated with the first similarity; Determining a fourth embedding learning method with the goal of minimizing the fourth objective function; The step of jointly embedding the triples by using the first embedding learning method, the second embedding learning method, and the third embedding learning method includes: The triples are jointly embedded and learned by the first embedding learning method, the second embedding learning method, the third embedding learning method, and the fourth embedding learning method.

9. A data processing device, characterized in that: The device comprises: The first processing module is used to determine a first embedding learning mode of a triple based on a first relation word of each triple in the triple set; each triple includes a relation triple and an attribute triple; based on each attribute triple in the triple set, determine a second embedding learning mode of the triple; based on each relation triple in the triple set, determine a third embedding learning mode of the triple; The second processing module is used to perform joint embedding learning on triples through the first embedding learning method, the second embedding learning method, and the third embedding learning method.

10. An electronic device, characterized in that: The electronic device comprises a processor and a memory for storing a computer program that can be run on the processor; wherein, The processor is configured to run the computer program to perform the method according to any one of claims 1 to 8.

11. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

12. A computer program product, comprising a computer program, characterized in that The computer program implements the method according to any one of claims 1 to 8 when executed by a processor.