Knowledge graph representation method and system based on FastText-TransH algorithm
The FastText-TransH algorithm generates word vectors and combines entity category information and relational hyperplanar projection mechanism, which solves the shortcomings of existing algorithms in complex entity relationships and long text entity semantic extraction, and improves the representation accuracy and interpretability of the knowledge graph.
Patent Information
- Application Number
- CN202510855053.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The existing knowledge graph embedding algorithms have shortcomings in handling complex entity relationships, semantic extraction of long text entity entities, and the use of entity category information, resulting in the distribution of entities in vector space that cannot fully reflect the correlation and differences between categories, reducing the accuracy and interpretability of knowledge graph representation.
The FastText-TransH algorithm is used to train the text data in the talent knowledge graph, generate word vectors, and combine entity category information and relational hyperplanar projection mechanism to optimize vector representation, and use dynamic weighting mechanisms and loss functions to improve the convergence and accuracy of the model.
It significantly improves the representation ability and reasoning efficiency of large-scale knowledge graphs, can distinguish different categories of entities and their relationships more clearly, and improves the accuracy and interpretability of knowledge graph representation.
Smart Images

Figure CN120354925A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graph representation, and in particular to a knowledge graph representation method and system based on the FastText-TransH algorithm. Background Art
[0002] In the field of talent knowledge graphs, knowledge graph embedding technology is crucial for accurately representing entities and relationships. Some existing classical embedding algorithms have many limitations when dealing with talent knowledge graphs.
[0003] The traditional TransE algorithm, although having the advantages of simplicity and efficiency, is based on the assumption that entities and relationships are in the same semantic space. This makes it difficult to accurately express the semantic differences of entities under different relationships when facing complex many-to-many relationships and long-text entities with rich semantics. For example, in a talent knowledge graph, in scenarios of many-to-many relationships such as various professional skills and project collaborations, and for long-text patent name entities like "Research and Development of a New Communication Technology Based on the Principle of Quantum Entanglement", the TransE algorithm performs poorly.
[0004] The TransR algorithm introduces the concept of a relationship space, but there are still deficiencies in mining deep semantic features. Moreover, these traditional methods generally ignore lexical-level information and entity category information. The long-text entities in talent knowledge graphs contain a large amount of professional knowledge and semantic connotations, and traditional methods cannot effectively convert the key information and semantic associations in them into vector representations accurately, thus affecting the performance of knowledge graphs in knowledge completion, reasoning, intelligent retrieval, etc.
[0005] In addition, the entities in talent knowledge graphs have different category attributes, such as paper and monograph categories, patent categories, award and honor categories, etc. Entities of the same category have certain similarities in semantics and features. Traditional embedding methods fail to make full use of this category information, resulting in the distribution of entities in the vector space not being able to fully reflect the relevance and differences between categories, reducing the accuracy and interpretability of knowledge graph representation. Summary of the Invention
[0006] The purpose of the present invention is to provide a knowledge graph representation method and system based on the FastText-TransH algorithm, aiming to solve the deficiencies of existing classical algorithms in dealing with complex entity relationships, long-text entity semantic extraction, and using entity category information, and to generate vector representations of corresponding entities and relationships efficiently and accurately.
[0007] In a first aspect, the present invention provides a knowledge graph representation method based on the FastText-TransH algorithm, and the method includes: Use the FastText algorithm to train the text data in the talent knowledge graph, and obtain word vectors corresponding to each word in the text data according to the training results; Input the word vectors corresponding to each word in the text data into the TransH algorithm for training to obtain an improved TransH algorithm, specifically as follows: Obtain the triple data of the talent knowledge graph. The triple data includes the word sets of the head entity, relation, and tail entity in the triples belonging to the talent knowledge graph, and obtain the initial vector of the triple by weighting the word vectors of all words in the word set; Obtain the categories to which the head entity and the tail entity belong, encode the categories, obtain the head entity vector and the tail entity vector according to the encoding results and the initial vector, and calculate the Euclidean distance according to the head entity vector and the tail entity vector; Construct a loss function according to the Euclidean distance, calculate the loss value according to the loss function, and judge whether the TransH algorithm converges according to the loss value. If it converges, the training ends; Input the triple data to be analyzed into the improved TransH algorithm to obtain the target vectors corresponding to the head entity and the tail entity in the triple data to be analyzed respectively.
[0008] In a second aspect, the present invention provides a knowledge graph representation system based on the FastText-TransH algorithm. The system includes: A word vector acquisition module for using the FastText algorithm to train the text data in the talent knowledge graph, and obtaining word vectors corresponding to each word in the text data according to the training results; A model training module for inputting the word vectors corresponding to each word in the text data into the TransH algorithm for training to obtain an improved TransH algorithm, specifically as follows: Obtain the triple data of the talent knowledge graph. The triple data includes the word sets of the head entity, relation, and tail entity in the triples belonging to the talent knowledge graph, and obtain the initial vector of the triple by weighting the word vectors of all words in the word set; Obtain the categories to which the head entity and the tail entity belong, encode the categories, obtain the head entity vector and the tail entity vector according to the encoding results and the initial vector, and calculate the Euclidean distance according to the head entity vector and the tail entity vector; Construct a loss function according to the Euclidean distance, calculate the loss value according to the loss function, and judge whether the TransH algorithm converges according to the loss value. If it converges, the training ends; The triple data analysis module is used to input the triple data to be analyzed into the improved TransH algorithm to obtain target vectors corresponding to the head entity and the tail entity in the triple data to be analyzed respectively.
[0009] Thirdly, the present invention provides a storage medium storing one or more programs, which when executed by a processor implement the above-mentioned knowledge graph representation method based on the FastText-TransH algorithm.
[0010] Fourthly, the present invention provides an electronic device, which includes a memory and a processor, wherein: The memory is used for storing a computer program; When the processor executes the computer program stored on the memory, it implements the above-mentioned knowledge graph representation method based on the FastText-TransH algorithm.
[0011] In summary, according to the above-mentioned knowledge graph representation method based on the FastText-TransH algorithm, by constructing a knowledge graph embedding algorithm that fuses structural and semantic information, combining the relationship projection model with the entity category enhancement mechanism, the representation ability and reasoning efficiency of large-scale knowledge graphs are significantly improved. By fusing the relationship hyperplane projection mechanism of TransH with the sub-word semantic information of entity names and introducing entity category constraints, when dealing with complex graphs of the association between talents and achievements, this framework improves the semantic generalization ability for low-frequency entities (such as "high-reliability systems") through the sub-word sharing mechanism of FastText, uses the relationship hyperplane of TransH to solve the modeling problem of many-to-many relationships (such as "talent - publication - paper"), and optimizes the semantic distribution of the vector space through entity category information (such as distinguishing "talent category" and "achievement category"). It enables the model to more clearly distinguish different category entities and their relationships, comprehensively improves the overall performance in large-scale knowledge graphs, and effectively improves the accuracy and interpretability of knowledge graph representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a flowchart of the knowledge graph representation method based on the FastText-TransH algorithm proposed in an embodiment of the present invention; Figure 2 It is a refinement diagram of step S102 proposed in an embodiment of the present invention; Figure 3 It is a structural schematic diagram of the knowledge graph representation system based on the FastText-TransH algorithm proposed in an embodiment of the present invention.
[0013] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. Detailed implementation manners
[0014] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art in the field to which the present invention pertains. The words such as "including" used herein are intended to mean that the elements or items appearing before the word cover the elements or items listed after the word and their equivalents, without excluding other elements or items.
[0015] As Figure 1 shown, an embodiment of the present invention provides a knowledge graph representation method based on the FastText-TransH algorithm. The method includes steps S101 to S103, where: Step S101: Use the FastText algorithm to train the text data in the talent knowledge graph, and obtain word vectors corresponding to each word in the text data according to the training results; FastText is a powerful word vector learning algorithm, which has significant advantages in processing text data. Traditional word vector algorithms (such as Word2Vec) usually regard each word as an independent unit for learning, and cannot provide effective vector representations for words that do not appear in the training data (out-of-vocabulary words). FastText introduces sub-word information, which enables it to show strong capabilities when processing languages such as Chinese. Taking the Chinese word "artificial intelligence" as an example, FastText will split it into multiple character n-gram sub-words. When n = 2, the possible sub-words are "artificial" and "intelligence", and when n = 3, there is "artificial intelligence", etc. During the learning process, FastText not only learns the vector of the whole word "artificial intelligence", but also learns the vectors of these sub-words. This way enables FastText to generate reasonable vector representations for in-vocabulary words. For example, when encountering an out-of-vocabulary word "artificial intelligence technology", since it contains known sub-words such as "artificial" and "intelligence", FastText can use the vector combinations of these sub-words to approximately represent the vector of "artificial intelligence technology".
[0016] It should be noted that after obtaining the text data, the text data in the talent knowledge graph is first preprocessed, such as word segmentation and stop word removal. Then, the FastText algorithm is used to train the processed text data. After the FastText algorithm training is completed, for any word in the text, it is input into the trained FastText algorithm to obtain the corresponding word vector.
[0017] Exemplarily, for the talent knowledge graph, a large amount of semantic information is contained in patents and paper abstracts. These abstract texts are collected and processed using the jieba tool for word segmentation, and then meaningless words and stop words in the sequence are removed. For example, an abstract text is: "In this study, we combined artificial intelligence technology to achieve fuzzy recognition of road license plates." Using the Jieba tool for word segmentation, we get: "In / this / study / , / we / combined / artificial / intelligence / to achieve / road / license plate / fuzzy / recognition". After removing punctuation marks, meaningless words and stop words (such as: "In", "in", "to achieve", etc.), it becomes: "this / study / combined / artificial / intelligence / technology / to achieve / road / license plate / fuzzy / recognition". Then, the processed text data is used to train word vectors using the FastText algorithm, and finally, the word vector representation of the used words is obtained through the trained model. For example, the word vector representation of "intelligence" is "[0.00235796, 0.09234943, 0.1239492,..., 0.043636239]". It should be noted that the word vectors obtained through the above steps are the word vectors of the words contained in patents, paper treatises, etc., rather than just the word vectors representing the entities in the knowledge graph.
[0018] In addition, after the FastText training is completed, for any word in the text, the corresponding word vector is obtained. For example, input: "intelligent recognition", using the above-trained FastText algorithm, generate the corresponding word vector, output: "[0.123929396, 0.042232494, 0.00422322,..., 0.10300029]".
[0019] Step S102: Input the word vectors corresponding to each word in the text data into the TransH algorithm for training to obtain an improved TransH algorithm; It should be pointed out that as Figure 2 shown, this step includes steps S1021 to S1023, where: Step S1021: Obtain the triple data of the talent knowledge graph. The triple data includes the word sets of the head entity, relation, and tail entity in the triples belonging to the talent knowledge graph, and obtain the initial vector of the triple by weighting the word vectors of all words in the word sets. In some embodiments, first divide the triple data into positive example triples and negative example triples, and then set the head entity in the positive example triples , relation , and tail entity are each composed of words, and , , . In addition, the initial vectors of the entity and the relation are obtained by weighted combination of the FastText word vectors of the words they contain: ; where are the word sets corresponding to the head entity, relation, and tail entity respectively, , , are the 1st, 2nd, and mth words in the word set corresponding to the head entity respectively, are the 1st, 2nd, and nth words in the word set corresponding to the relation respectively, are the 1st, 2nd, and pth words in the word set corresponding to the tail entity respectively, are the initial vectors corresponding to the head entity, relation, and tail entity respectively, are the weight coefficient and word vector of the ith word in the word set corresponding to the head entity respectively, are the weight coefficient and word vector of the jth word in the word set corresponding to the relation respectively, are the weight coefficient and word vector of the kth word in the word set corresponding to the tail entity respectively. Regarding the weight coefficients of the corresponding words, they will be learned in the subsequent model training to adjust the contributions of different words to the entity and relation representations. In addition, these weight coefficients have no specific fixed range before training and theoretically belong to the real number set R. In actual training, to make the model stable and reasonable, after operations such as normalization, they will be constrained in the interval [0,1].
[0020] In summary, regarding traditional knowledge graph embedding methods, especially when only using TransH, it should be noted that when traditional methods handle the vector representations of entities and relationships, they mainly rely on the structural information of the knowledge graph and insufficiently mine the text semantic information contained in entities and relationships. This single reliance on structural information in the processing method does not fully consider the rich semantic details in the text and the complementarity between different modal information, resulting in the generated vector representations being unable to accurately reflect the true semantics of entities and relationships in complex knowledge representation scenarios and performing poorly in tasks such as knowledge reasoning and link prediction. In this step, the weighted fusion formula based on FastText realizes the deep fusion of the structural information of the knowledge graph and the text semantic information by introducing a dynamic weight mechanism. Varying with entity characteristics: For different types of entities, the richness of their text descriptions and the importance of structural features vary. For entities with detailed text descriptions, such as patent names, paper names, award names, etc., the text semantic weight in the formula increases, enabling the model to more fully capture the unique semantic information contained in their texts; for entities dominated by structural features, such as some abstract concepts or numerical entities, the structural information weight is dominant to ensure that key structural features are not overlooked. Varying with relationship types: Different relationships play different roles in the knowledge graph, and their dependencies on text semantics and structural information also vary. When dealing with relationships with close semantic associations (such as "belong to", etc.), the text semantic weight is increased to better match the semantic associations; for relationships emphasizing structural associations (such as "contain", etc.), the structural information weight is higher to highlight the structural-level connections.
[0021] Step S1022: Obtain the categories to which the head entity and the tail entity belong, encode the categories, obtain the head entity vector and the tail entity vector according to the encoding results and the initial vectors, and calculate the Euclidean distance based on the head entity vector and the tail entity vector; It should be noted that traditional entity vector representation fusion methods often adopt simple concatenation or fixed-weight addition. This method does not fully consider the contribution differences of different information sources to entity representation and the feature changes of entities in different contexts. The simple fusion method cannot effectively capture the dynamic semantic information of entities and is difficult to accurately represent entities in complex knowledge graph scenarios, resulting in performance limitations in tasks such as knowledge reasoning and link prediction. Based on this, in this step, the entities in the knowledge graph are first classified. Let the category set be , the category to which the head entity belongs is , the one-hot encoded vector representation of the category to which the head entity belongs is , the category to which the tail entity belongs is , and the one-hot encoded vector representation of the category to which the tail entity belongs is ; Exemplarily, for instance, entity classifications in the knowledge graph include personnel information, awards and honors, patents, papers and treatises, research projects, schools, majors, etc., which belong to the category set C. Each category has a one-hot encoded vector. For example, the one-hot encoded vector for personnel information is [1, 0, 0, 0, 0, 0, 0], and the one-hot encoded vector for awards and honors is [0, 1, 0, 0, 0, 0, 0].
[0022] Fuse the initial vector with the one-hot encoded vector representation: ; where, and are the updated head entity vector and tail entity vector respectively, and and and and and are all hyperparameters, is element-wise multiplication.
[0023] According to the above fusion formula, as the importance of the information source changes, adjust the weights and and and and and through training, and it is possible to flexibly allocate the contribution of each part to the updated vector according to the importance of different information sources (such as the text semantic information from FastText and the structural information contained in the entity initial vector) under different tasks or entity characteristics. Combining with non-linear interaction, the element-wise multiplication operation between vectors introduces non-linear interaction. This operation is not a simple linear superposition, but explores the deep semantic association between the initial vector and the FastText semantic vector, can discover the synergistic effect between different information source features, generate a more discriminative and expressive entity vector representation, and thus better capture the complex relationships between entities in knowledge graph related tasks and improve the overall performance.
[0024] In addition, has a value range of [0.1, 1], which refers to the weight proportion of the initial vector itself in the updated vector, and the larger the value, the greater the influence of the initial vector. usually has a value range of [0, 1], which is used to measure the contribution degree of the category one-hot encoded vector to the updated vector, and the larger the value, the more important the category information is in the vector representation. generally has a value in [0, 1], which determines the weight of the element-wise multiplication result of the initial vector and the category one-hot encoded vector, and the larger the value, the greater the influence of this interaction on the updated vector.
[0025] In addition, in some embodiments, weight matrices of different types of relationships are also used to capture the semantic features between entities under different relationship types. A set of relationship types is defined, such as "personnel information - type relationships", "award and honor - type relationships", etc. For each relationship, there is a corresponding relationship type. A weight matrix is introduced for different types of relationships to adjust the entity projection. Each relationship corresponds to a hyperplane normal vector and a relationship vector.
[0026] Specifically, the updated head - entity vector and the tail - entity vector are projected onto the hyperplane corresponding to the relationship: ; where and are the projection vectors corresponding to the head entity and the tail entity respectively, is the hyperplane normal vector corresponding to the relationship, is the weight matrix corresponding to the type of the relationship, and the parameters in the weight matrix are updated by minimizing the loss function.
[0027] In addition, in some embodiments, the expression of the weight matrix is as follows: .
[0028] is the element in the d - th row and d - th column of the weight matrix, and the value range of any element is .
[0029] The traditional TransH algorithm defines a hyperplane for each relationship, and the hyperplane is determined by the normal vector It is determined that the entity vector is projected onto the hyperplane through the dot product operation with the normal vector. Only the projection position of the entity on the hyperplane is considered, and the influence of different relationship types on the entity projection is not distinguished. It is assumed that different relationships belong to the same hyperplane and different entities are in the same semantic space. However, in this embodiment, the differences in relationship types are considered, and weight matrices are introduced for different types of relationships to adjust the entity projection. When projecting the entity, not only the normal vector of the hyperplane is used for projection calculation, but also the weight matrix corresponding to the relationship type is combined to adjust the updated entity vector, so that the projected vector representation can capture the semantic features between entities under different relationship types, allowing for more flexible entity projection methods for different relationship types. In addition, by introducing the relationship type weight matrix, differential adjustment can be performed on the entity projection for different relationship types (such as "personnel information type relationships", "award and honor type relationships", etc.), mining the subtle differences in entity semantics under different relationship types, and making the vector representations of entities in different relationship types better reflect their true semantics. For example, in "personnel information type relationships", the semantics related to the basic attributes of personnel may be more concerned; in "award and honor type relationships", the semantics related to honors are more emphasized, so that the model can capture semantics more accurately.
[0030] In addition, the Euclidean distance needs to be calculated according to the following formula: ; where, is the Euclidean distance in the positive example triple, d is the dimension of the vector, is the relationship vector in the positive example triple, , are the i-th elements of the projection vectors corresponding to the head entity and the tail entity respectively.
[0031] In addition, it should be noted that the way to obtain the Euclidean distance in the negative example triple is exactly the same as that in the positive example triple, except that the objects processed are different. Therefore, in this embodiment, the Euclidean distance in the negative example triple will not be described again.
[0032] Step S1023: Construct a loss function according to the Euclidean distance, calculate the loss value according to the loss function, and determine whether the TransH algorithm converges according to the loss value. If it converges, the training ends; In this step, the loss function is specifically constructed according to the following formula: ; where, L is the loss value, N is the number of triples, is the margin hyperparameter, and its value range is [0.5, 2], which is used to control the score gap between positive and negative examples, is the Euclidean distance in the counterexample triple, is the intra-class average distance, is the inter-class average distance, G is the set of positive example triples in the knowledge graph, is the set of counterexample triples in the knowledge graph, 、 、 are respectively the head entity vector, relation vector, and tail entity vector in the negative example triple, 、 are respectively the head entity vector and tail entity vector in the positive example triple, and are both hyperparameters, and their value ranges are both in [0.01, 1], is the positive-taking function.
[0033] The intra-class average distance is calculated according to the following formula: ; The inter-class average distance is calculated according to the following formula: ; Among them, is the intra-class average distance, n is the total number of entity categories in the triple data, is the category set of the i-th entity, is the category set is the number of entities in the set, is the i-th entity category set is the mean vector of all entity vectors in the set, and are respectively the k-th dimensional components of the mean vectors of the i-th class and the j-th class, is the entity is the value of the k-th dimension in the vector representation of, e is a specific entity in the set , is the inter-class average distance, d is the dimension of the entity vector.
[0034] In summary, since traditional loss functions are mainly based on the maximum margin idea, which optimizes the vector representations of entities and relationships by constraining the distances between positive and negative examples on the hyperplane, focusing on the structural rationality of entities and relationships in triples, but not considering the intra-class aggregation and inter-class separation characteristics of entity semantics. When dealing with entities with similar semantics, their vector representations may lack sufficient distinguishability, which is not conducive to accurately capturing semantic details. Based on this, in this embodiment, a part of the score difference between positive and negative examples based on the margin is incorporated into the loss function of the traditional TransH, combined with the intra-class and inter-class distance terms, to form a more comprehensive loss calculation method. While increasing the score gap between positive and negative examples, with the constraints of intra-class and inter-class distances, the model can distinguish positive and negative examples not only based on structure but also considering aggregation and separation at the semantic level. For example, in complex knowledge reasoning tasks, it can more accurately judge the rationality of triples and improve the performance of the model in tasks such as link prediction and knowledge reasoning.
[0035] In addition, the loss value is calculated through the loss function designed in this step, the parameters are updated using the Adam optimizer, and then it is judged whether the loss converges. If it converges, the training ends, and thus the improved TransH algorithm is obtained.
[0036] Step S103: Input the triple data to be analyzed into the improved TransH algorithm to obtain target vectors corresponding to the head entity and the tail entity in the triple data to be analyzed respectively.
[0037] It should be noted that this target vector refers to the projection vector, and the triple data to be analyzed is similar to the triple data in the talent knowledge graph during the training process.
[0038] In summary, according to the above knowledge graph representation method based on the FastText-TransH algorithm, by constructing a knowledge graph embedding algorithm that integrates structural and semantic information, combining the relationship projection model with the entity category enhancement mechanism, the representation ability and reasoning efficiency of large-scale knowledge graphs are significantly improved. By fusing the relationship hyperplane projection mechanism of TransH with the sub-word semantic information of entity names and introducing entity category constraints, when dealing with complex graphs related to talents and achievements, this framework improves the semantic generalization ability for low-frequency entities (such as "high-reliability systems") through the sub-word sharing mechanism of FastText, uses the relationship hyperplane of TransH to solve the modeling problem of many-to-many relationships (such as "talent - publication - paper"), and optimizes the semantic distribution of the vector space through entity category information (such as distinguishing "talent category" and "achievement category"). This enables the model to more clearly distinguish different category entities and their relationships, comprehensively improve the overall performance in large-scale knowledge graphs, and effectively improve the accuracy and interpretability of knowledge graph representation.
[0039] On the other hand, the present invention also provides a storage medium, on which one or more programs are stored, and when the program is executed by a processor, the above-mentioned knowledge graph representation method based on the FastText-TransH algorithm is implemented.
[0040] On the other hand, the present invention also provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored on the memory to implement the above-mentioned knowledge graph representation method based on the FastText-TransH algorithm.
[0041] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus or device), or in combination with these instruction execution systems, apparatus or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus or device.
[0042] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0043] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0044] Although the embodiments of the present invention have been described in detail above, it will be apparent to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are all within the scope and spirit of the present invention as described in the claims. Moreover, the present invention described herein can have other embodiments and can be implemented or realized in various ways.
Claims
1. A knowledge graph representation method based on the FastText-TransH algorithm, characterized in that The method includes: Training the text data in the talent knowledge graph using the FastText algorithm, and obtaining word vectors corresponding to each word in the text data according to the training results; Inputting the word vectors corresponding to each word in the text data into the TransH algorithm for training to obtain an improved TransH algorithm, specifically as follows: Obtaining triple data about the talent knowledge graph, where the triple data includes word sets of the head entity, relation, and tail entity in the triples belonging to the talent knowledge graph, and obtaining an initial vector for the triple by weighting the word vectors of all words in the word set; Obtaining the categories to which the head entity and the tail entity belong, encoding the categories, obtaining the head entity vector and the tail entity vector according to the encoding results and the initial vector, and calculating the Euclidean distance based on the head entity vector and the tail entity vector; Constructing a loss function according to the Euclidean distance, calculating a loss value according to the loss function, and determining whether the TransH algorithm converges according to the loss value. If it converges, the training ends; Inputting the triple data to be analyzed into the improved TransH algorithm to obtain target vectors corresponding to the head entity and the tail entity in the triple data to be analyzed respectively.
2. The knowledge graph representation method based on the FastText-TransH algorithm according to claim 1, wherein The step of obtaining triple data about the talent knowledge graph, where the triple data includes word sets of the head entity, relation, and tail entity in the triples belonging to the talent knowledge graph, and obtaining an initial vector for the triple by weighting the word vectors of all words in the word set includes: Dividing the triple data into positive example triples and negative example triples; Let the head entity in the positive example triple , the relation and the tail entity be composed of words respectively, and , , ; Obtaining the initial vector of the positive example triple according to the following formula: ; Among them, are respectively the word sets corresponding to the head entity, relation, and tail entity, , , are respectively the 1st, 2nd, and m-th words in the word set corresponding to the head entity, are respectively the 1st, 2nd, and n-th words in the word set corresponding to the relation, are respectively the 1st, 2nd, and p-th words in the word set corresponding to the tail entity, are respectively the initial vectors corresponding to the head entity, relation, and tail entity, are respectively the weight coefficient and word vector of the i-th word in the word set corresponding to the head entity, are respectively the weight coefficient and word vector of the j-th word in the word set corresponding to the relation, are respectively the weight coefficient and word vector of the k-th word in the word set corresponding to the tail entity.
3. The knowledge graph representation method based on the FastText-TransH algorithm according to claim 2, characterized in that The step of obtaining the categories to which the head entity and the tail entity belong, encoding the categories, and obtaining the head entity vector and the tail entity vector according to the encoding results and the initial vector includes: Let the category set be , the category to which the head entity belongs is , and the one-hot encoded vector representation of the category to which the head entity belongs is , the category to which the tail entity belongs is , and the one-hot encoded vector representation of the category to which the tail entity belongs is ; Fusing the initial vector with the one-hot encoding vector representation; ; Among them, , are the updated head entity vector and tail entity vector respectively, , , , , , are all hyperparameters, is element-wise multiplication.
4. The knowledge graph representation method based on the FastText-TransH algorithm according to claim 3, characterized in that, The step of calculating the Euclidean distance based on the head entity vector and the tail entity vector includes: The updated head entity vector and the tail entity vector are projected onto the hyperplane corresponding to the relation: ; Among them, , are projection vectors corresponding to the head entity and the tail entity respectively, is the hyperplane normal vector corresponding to the relation, is the weight matrix corresponding to the type to which the relation belongs; Calculating the Euclidean distance according to the following formula: ; Among them, is the Euclidean distance in the positive example triple, d is the dimension of the vector, is the relationship vector in the positive example triple, , are respectively the i-th elements of the projection vectors corresponding to the head entity and the tail entity.
5. The knowledge graph representation method based on the FastText-TransH algorithm according to claim 4, characterized in that Constructing a loss function according to the following formula: ; Where L is the loss value and N is the number of triples, is the margin hyperparameter used to control the score gap between positive and negative examples, is the Euclidean distance in the negative example triples, is the intra-class average distance, is the inter-class average distance, G is the set of positive example triples in the knowledge graph, is the set of negative example triples in the knowledge graph, and and are the head entity vector, relation vector, and tail entity vector in the negative example triples respectively, and are the head entity vector and tail entity vector in the positive example triples respectively, and are both hyperparameters, is the positive-taking function.
6. The knowledge graph representation method based on the FastText-TransH algorithm according to claim 5, wherein Calculating the average intra-class distance according to the following formula: ; Calculating the average inter-class distance according to the following formula: ; Among them, is the average intra-class distance, n is the total number of entity categories in the triple data, is the category set of the i-th entity, is the category set the number of entities in, is the i-th entity category set the mean vector of all entity vectors in, and are the k-th dimensional components of the mean vectors of the i-th and j-th classes respectively, is the entity the value of the k-th dimension in the vector representation of, e is a specific entity in the set in, is the average inter-class distance, d is the dimension of the entity vector.
7. A knowledge graph representation system based on the FastText-TransH algorithm, characterized in that, The system includes: A word vector acquisition module for training the text data in the talent knowledge graph using the FastText algorithm and obtaining word vectors corresponding to each word in the text data according to the training results; A model training module for inputting the word vectors corresponding to each word in the text data into the TransH algorithm for training to obtain an improved TransH algorithm, specifically as follows: Obtaining triple data about the talent knowledge graph, where the triple data includes word sets of the head entity, relation, and tail entity in the triples belonging to the talent knowledge graph, and obtaining an initial vector for the triple by weighting the word vectors of all words in the word set; Obtain the categories to which the head entity and the tail entity belong, encode the categories, obtain the head entity vector and the tail entity vector based on the encoding result and the initial vector, and calculate the Euclidean distance according to the head entity vector and the tail entity vector; Construct a loss function according to the Euclidean distance, calculate the loss value according to the loss function, and determine whether the TransH algorithm converges according to the loss value. If it converges, the training ends; A triple data analysis module, configured to input the triple data to be analyzed into the improved TransH algorithm to obtain target vectors corresponding to the head entity and the tail entity in the triple data to be analyzed respectively.
8. A storage medium, characterized in that, The storage medium stores one or more programs, which when executed by a processor implement the knowledge graph representation method based on the FastText-TransH algorithm according to any one of claims 1-6.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, wherein: The memory is used for storing a computer program; The processor is configured to implement the knowledge graph representation method based on the FastText-TransH algorithm according to any one of claims 1-6 when executing the computer program stored on the memory.
Citation Information
Patent Citations
Knowledge graph representation learning method and system in combination with entity description
CN106886543A
Knowledge-graph representation learning method based on multi-objective optimization
CN107885759A
Knowledge-graph representation learning method based on multiple semantemes
CN107885760A
A knowledge map representation learning method based on structure information and text description
CN109299284A
Knowledge graph representation learning method fusing entity description and types
CN111753101A