Knowledge Graph Entity Linking Method, Device, Computer Equipment and Storage Medium
By generating an entity link model based on training data of problem samples and knowledge graph entities, the problem of poor performance of entity consistency model in question-and-answer scenarios is solved, and higher entity link accuracy is achieved.
Patent Information
- Application Number
- CN202310522687.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-10
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-05-10
AI Technical Summary
In the Q&A scenario, the existing entity consistency model is not effective, the entity link accuracy is low, and it is difficult to effectively use the information in the knowledge graph for accurate linking.
By obtaining training data positive samples based on problem samples, entity mention samples, knowledge graph entity positive samples and adjacency sub-graph samples, combining knowledge graph entity negative samples, training using BERT model and multi-layer perceptron models, generating entity link models, and using user problems and candidate knowledge graph entities for entity linking.
It improves the accuracy and effectiveness of the entity link model, improves the entity consistency model effect in the Q&A scenario, and enhances the accuracy of entity links.
Smart Images

Figure CN116561339B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing, and particularly to a method, apparatus, computer device, and readable storage medium for knowledge graph entity linking. Background Art
[0002] Entity linking refers to linking the entity mentions that appear in the text with the corresponding entities in the knowledge graph, and it is an important part of many information extraction tasks and natural language understanding tasks, such as knowledge graph update, question answering based on the knowledge graph, search engines, etc. Due to the diversity of natural language descriptions, eliminating the ambiguity of entity mentions is the main work in the entity linking task.
[0003] Entity linking methods aim to establish a text comparison model between entity mentions and knowledge graph entities. Most of the existing methods focus on the semantic information of the text descriptions of knowledge graph entities, or model the consistency of entities in the document according to the association relationships of entity nodes in the knowledge graph. The latter can utilize the large amount of graph structures and stored knowledge in the knowledge graph to obtain information outside the user's question to help eliminate ambiguity and effectively improve the accuracy of entity linking. However, in the question answering scenario, most user questions involve a small number of entities, and it is difficult to effectively establish an entity consistency model. Therefore, the knowledge graph is difficult to be utilized, and the accuracy of entity linking is relatively low.
[0004] Regarding the problem that the entity consistency model has poor effects and the entity linking accuracy is relatively low in the question answering scenario in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] In the present embodiment, a method, apparatus, computer device, and readable storage medium for knowledge graph entity linking are provided to solve the problem that the entity consistency model has poor effects and the entity linking accuracy is relatively low in the question answering scenario in the related art.
[0006] In a first aspect, in the present embodiment, a method for knowledge graph entity linking is provided, and the method includes:
[0007] Based on question samples, entity mention samples, knowledge graph entity positive samples, and knowledge graph entity adjacent sub-graph samples, training data positive samples are obtained. The entity mention samples are obtained based on the question samples, the knowledge graph entity positive samples are obtained based on the annotated entities of the entity mention samples in the knowledge graph, and the knowledge graph entity adjacent sub-graph samples are obtained based on the entity relationships of the knowledge graph entity positive samples in the knowledge graph;
[0008] Based on the problem samples, the entity mention samples, the knowledge graph entity negative samples, and the corresponding knowledge graph entity adjacent sub-graph samples, training data negative samples are obtained. The knowledge graph entity negative samples are randomly obtained based on entities in the knowledge graph that have no annotation relationship with the entity mention samples;
[0009] Based on the training data positive samples and the training data negative samples, the entity linking initial model is trained to obtain an entity linking model;
[0010] The user question, the entity mention, the candidate knowledge graph entities, and the corresponding knowledge graph entity adjacent sub-graphs are input into the trained entity linking model to determine the target knowledge graph entity linked to the entity mention; the entity mention is obtained based on the user question, the candidate knowledge graph entities are obtained based on the entity mention, and the knowledge graph entity adjacent sub-graphs are obtained based on the entity relationships of the candidate knowledge graph entities in the knowledge graph.
[0011] In some embodiments, the training of the entity linking initial model based on the training data positive samples and the training data negative samples includes:
[0012] The training data positive samples and the training data negative samples are input into the entity linking initial model, and sample entity mention vectors and sample entity graph convolutional vectors are output;
[0013] Based on the sample entity mention vectors, the sample entity graph convolutional vectors, and the pre-obtained sample marking parameters, the loss function of the entity linking initial model is determined;
[0014] The entity linking initial model is trained based on the loss function.
[0015] In some embodiments, the entity linking initial model includes a text embedding module and an entity graph network embedding module. The inputting of the training data positive samples and the training data negative samples into the entity linking initial model to output sample entity mention vectors and sample entity graph convolutional vectors includes:
[0016] The problem samples and the knowledge graph entity samples are input into the text embedding module, and the sample entity mention vectors and the sample question vectors are output. The knowledge graph entity samples include the knowledge graph entity positive samples and the knowledge graph entity negative samples;
[0017] Based on the sample question vectors, the attention weights of the entity sample vectors in the entity graph network embedding module are obtained;
[0018] The knowledge graph entity adjacent sub-graph samples are input into the entity graph network embedding module, and the sample entity graph convolutional vectors are output based on the attention weights.
[0019] In some of these embodiments, the text embedding module includes a BERT model, and inputting the problem sample and the knowledge graph entity sample into the text embedding module and outputting the sample entity mention vector and the sample problem vector includes:
[0020] Concatenating the problem sample, the knowledge graph entity sample, the CLS flag bit, and the SEP flag bit based on a predetermined format, and inputting the result into the BERT model;
[0021] Determining the sample problem vector based on the first output vector of the BERT model corresponding to the CLS flag bit.
[0022] In some of these embodiments, the text embedding module further includes a multi-layer perceptron model, and inputting the problem sample and the knowledge graph entity sample into the text embedding module and outputting the sample entity mention vector and the sample problem vector further includes:
[0023] Obtaining a second output vector of the BERT model corresponding to the start position of the entity mention sample, and a third output vector of the BERT model corresponding to the end position of the entity mention sample;
[0024] Concatenating the first output vector, the second output vector, and the third output vector, and inputting the result into the multi-layer perceptron model to obtain the sample entity mention vector.
[0025] In some of these embodiments, inputting the knowledge graph entity adjacent subgraph sample into the entity graph network embedding module and outputting the sample entity graph convolutional vector based on the attention weights includes:
[0026] Initializing the knowledge graph entity adjacent subgraph sample based on the first output vector to obtain a corresponding entity sample vector;
[0027] Inputting the entity sample vector into the entity graph network embedding module and outputting the sample entity graph convolutional vector based on the attention weights.
[0028] In some of these embodiments, before concatenating the problem sample, the knowledge graph entity sample, the CLS flag bit, and the SEP flag bit based on a predetermined format and inputting the result into the BERT model, the method further includes:
[0029] Pre-training the initial BERT model to obtain pre-trained model parameters;
[0030] Establishing the BERT model based on the pre-trained model parameters.
[0031] Second aspect, in this embodiment, a knowledge graph entity linking device is provided, and the device includes:
[0032] A first acquisition module, configured to acquire positive training data based on a question sample, an entity mention sample, a knowledge graph entity positive sample, and a knowledge graph entity adjacent sub-graph sample, where the entity mention sample is acquired based on the question sample, the knowledge graph entity positive sample is acquired based on an annotated entity of the entity mention sample in the knowledge graph, and the knowledge graph entity adjacent sub-graph sample is acquired based on an entity relationship of the knowledge graph entity positive sample in the knowledge graph;
[0033] A second acquisition module, configured to acquire negative training data based on the question sample, the entity mention sample, a knowledge graph entity negative sample, and a corresponding knowledge graph entity adjacent sub-graph sample, where the knowledge graph entity negative sample is randomly acquired based on an entity in the knowledge graph that has no annotation relationship with the entity mention sample;
[0034] A training module, configured to train an initial entity linking model based on the positive training data and the negative training data to obtain an entity linking model;
[0035] A determination module, configured to input a user question, an entity mention, a candidate knowledge graph entity, and a corresponding knowledge graph entity adjacent sub-graph into the trained entity linking model to determine a target knowledge graph entity linked to the entity mention; the entity mention is acquired based on the user question, the candidate knowledge graph entity is acquired based on the entity mention, and the knowledge graph entity adjacent sub-graph is acquired based on an entity relationship of the candidate knowledge graph entity in the knowledge graph.
[0036] Third aspect, in this embodiment, a computer device is provided, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps of the knowledge graph entity linking method described in the first aspect above.
[0037] Fourth aspect, in this embodiment, a readable storage medium is provided, on which a program is stored, and when the program is executed by a processor, the steps of the knowledge graph entity linking method described in the first aspect above are implemented.
[0038] Compared with the related technologies, the knowledge graph entity linking method provided in this embodiment obtains positive training data samples based on question samples, entity mention samples, knowledge graph entity positive samples, and knowledge graph entity adjacent sub-graph samples, and uses the knowledge graph entities linked to the entity mention samples through manual annotation as the sample data for positive training, ensuring the accuracy of the positive training data samples and the effectiveness of training; obtains negative training data samples based on question samples, entity mention samples, knowledge graph entity negative samples, and corresponding knowledge graph entity adjacent sub-graph samples, that is, randomly selects entities as knowledge graph entity negative samples after removing the entity positive samples in the knowledge graph, improving the randomness of the negative training data samples and the effectiveness of training; trains the entity linking initial model based on the positive training data samples and the negative training data samples to obtain an entity linking model, improving the linking accuracy of the entity linking model; inputs the user question, entity mention, candidate knowledge graph entity, and corresponding knowledge graph entity adjacent sub-graph into the trained entity linking model to determine the target knowledge graph entity linked to the entity mention, and obtains the knowledge graph entity most similar to the entity mention as the linking target, improving the linking effect of the entity linking model, and solving the problems of poor entity consistency model effect and low entity linking accuracy in the question-and-answer scenario in the related technologies.
[0039] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0041] Figure 1 is a structural block diagram of a terminal of the knowledge graph entity linking method according to some embodiments of the present application;
[0042] Figure 2 is a flowchart of the knowledge graph entity linking method according to some embodiments of the present application;
[0043] Figure 3 is a flowchart of training the entity linking initial model based on training data samples according to some embodiments of the present application;
[0044] Figure 4 is a flowchart of outputting sample entity mention vectors and sample entity graph convolution vectors based on training data samples according to some embodiments of the present application;
[0045] Figure 5It is a flowchart for obtaining sample problem vectors based on problem samples and knowledge graph entity samples in some embodiments of the present application;
[0046] Figure 6 It is a flowchart for obtaining sample entity mention vectors based on problem samples and knowledge graph entity samples in some embodiments of the present application;
[0047] Figure 7 It is a flowchart for obtaining sample entity graph convolution vectors based on knowledge graph entity adjacency subgraph samples in some embodiments of the present application;
[0048] Figure 8 It is a schematic structural diagram of an entity linking model in some preferred embodiments of the present application;
[0049] Figure 9 It is a flowchart of a knowledge graph entity linking method in some preferred embodiments of the present application;
[0050] Figure 10 It is a structural block diagram of a knowledge graph entity linking device in some embodiments of the present application. Detailed implementation manners
[0051] To understand the purpose, technical solution and advantages of the present application more clearly, the present application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0052] Unless otherwise defined, technical terms or scientific terms involved in this application shall have the general meanings understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "an", "one kind", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variants thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in this application do not limit to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The term "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third" and the like involved in this application only distinguish similar objects and do not represent a specific sorting of the objects.
[0053] The knowledge graph entity linking method provided by the embodiments of this application can be executed in a terminal, a computer, or a similar computing device. When this method is applied to a terminal, a computer, or a similar computing device, Figure 1 is a hardware structure block diagram of the terminal of the knowledge graph entity linking method of some embodiments of this application. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 and a memory 104 for storing data. Among them, the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a field programmable gate array FPGA. The above terminal may also include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may also include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown.
[0054] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the knowledge graph entity linking method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some embodiments, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0055] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by a communication provider. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0056] In this embodiment, a knowledge graph entity linking method is provided. Figure 2 It is a flowchart of the knowledge graph entity linking method in some embodiments of the present application, as Figure 2 shown, and the process includes the following steps:
[0057] Step S201, based on the problem sample, entity mention sample, knowledge graph entity positive sample, and knowledge graph entity adjacent subgraph sample, obtain the training data positive sample. The entity mention sample is obtained based on the problem sample, the knowledge graph entity positive sample is obtained based on the labeled entity of the entity mention sample in the knowledge graph, and the knowledge graph entity adjacent subgraph sample is obtained based on the entity relationship of the knowledge graph entity positive sample in the knowledge graph.
[0058] In this embodiment, the problem sample refers to the existing user problems obtained by means such as collection. The entity mention sample refers to the entity mentions extracted from the problem sample by means of existing NER models, etc. A problem sample may include one or more entity mention samples. The knowledge graph entity positive sample refers to the knowledge graph entity obtained in the knowledge graph by manual annotation according to the entity mention sample, and the knowledge graph entity corresponds to the entity mention sample. The knowledge graph entity adjacent subgraph sample refers to the adjacent subgraph formed by the one-hop adjacent nodes of the knowledge graph entity positive sample in the knowledge graph.
[0059] Specifically, the knowledge graph G can be represented as G(V, E), where V represents nodes (entities) and E represents edges (relationships) between nodes. Define the adjacent subgraph of entity e i in the knowledge graph G as the subgraph composed of the one-hop adjacent nodes of e i , denoted as Ge i .
[0060] For the user question q, the NER model obtains the set of entity mentions M = {m i} i=1,...,n , the set of knowledge graph entities E = {e i} i=1,...,n that are labeled and linked, and forms a quadruple <q, m i , e i , Ge i > i=1,...,n as the positive sample of the training data.
[0061] Step S202: Based on the problem sample, the entity mention sample, the negative sample of the knowledge graph entity, and the corresponding adjacent subgraph sample of the knowledge graph entity, obtain the negative sample of the training data. The negative sample of the knowledge graph entity is randomly obtained based on the entities in the knowledge graph that have no labeled relationship with the entity mention sample.
[0062] Randomly extract entities from the knowledge graph and filter out the set of entities E that appear in the positive sample, denoted as E* = {e j} j=1,...,n , where Form a quadruple of the user question, the entity mention, the randomly extracted knowledge graph entity, and its adjacent subgraph as the negative sample of the training data.
[0063] Step S203: Train the initial entity linking model based on the positive sample of the training data and the negative sample of the training data to obtain the entity linking model.
[0064] Input the quadruple data of the positive sample of the training data and the negative sample of the training data into the initial entity linking model for training. Calculate the loss function of the model according to the output of the initial entity linking model, and adjust the model parameters according to the loss function to finally obtain the entity linking model. The initial entity linking model can be different types of neural network models or a combination of multiple neural network models, including CNN, RNN, Transformer, BERT, GPT, multi-layer perceptron, etc.
[0065] Step S204: Input the user question, entity mention, candidate knowledge graph entity, and the corresponding knowledge graph entity adjacent subgraph into the trained entity linking model to determine the target knowledge graph entity linked to the entity mention. The entity mention is obtained based on the user question, the candidate knowledge graph entity is obtained based on the entity mention, and the knowledge graph entity adjacent subgraph is obtained based on the entity relationships of the candidate knowledge graph entity in the knowledge graph.
[0066] When a new user question is input, entity mentions in the user question can be extracted through an entity recognition model such as an NER model. According to the entity mentions and a text similarity evaluation algorithm, such as the BM25 algorithm, multiple knowledge graph entities similar to the entity mentions are obtained as a candidate entity set, and then the corresponding knowledge graph adjacent subgraphs of the candidate knowledge graph entities are obtained respectively.
[0067] Construct a quadruple from the user question, entity mention, candidate knowledge graph entity, and the corresponding knowledge graph entity adjacent subgraph, input it into the trained entity linking model, obtain the vector representation of the entity mention and the set of graph convolution vector representations of the candidate knowledge graph entities, and select the candidate entity with the highest cosine similarity to the vector representation of the entity mention from the set of graph convolution vector representations as the linked target entity.
[0068] Through the above steps S201 - S204, by using the question samples, entity mention samples, knowledge graph entity positive samples, and knowledge graph entity adjacent subgraph samples, positive training data samples are obtained. The knowledge graph entities linked to the entity mention samples through manual annotation are used as the sample data for positive training, ensuring the accuracy of the positive training data samples and the effectiveness of training. By using the question samples, entity mention samples, knowledge graph entity negative samples, and the corresponding knowledge graph entity adjacent subgraph samples, negative training data samples are obtained, that is, after removing the entity positive samples in the knowledge graph, entities are randomly selected as the knowledge graph entity negative samples, improving the randomness of the negative training data samples and the effectiveness of training. By training the initial entity linking model based on the positive training data samples and negative training data samples, an entity linking model is obtained, improving the linking accuracy of the entity linking model. By inputting the user question, entity mention, candidate knowledge graph entity, and the corresponding knowledge graph entity adjacent subgraph into the trained entity linking model to determine the target knowledge graph entity linked to the entity mention, and obtaining the knowledge graph entity most similar to the entity mention as the linking target, the linking effect of the entity linking model is improved, solving the problems of poor entity consistency model effect and low entity linking accuracy in the relevant art in the question - answering scenario.
[0069] In some of the embodiments, Figure 3 is a flowchart of training the initial entity linking model based on training data samples in some embodiments of the present application, asFigure 3 As shown in the figure, the process includes the following steps:
[0070] Step S301: Input the positive training data samples and negative training data samples into the initial entity linking model, and output the sample entity mention vector and the sample entity graph convolution vector.
[0071] After inputting the quadruple data of the positive training data samples and negative training data samples into the initial entity linking model, through the processing of the initial entity linking model, the semantic information contained in the problem sample, entity mention sample, and knowledge graph entity sample in the quadruple is extracted, and the sample entity mention vector is output. This sample entity mention vector is the representation vector mapped by the entity mention sample, contains the context information in the problem sample, and expresses the semantic information corresponding to the entity mention sample in the form of a vector.
[0072] The initial entity linking model processes the knowledge graph entity adjacent subgraph samples in the quadruple data and outputs the sample entity graph convolution vector. This sample entity graph convolution vector is a graph convolution vector extracted according to the one-hop adjacent entity relationship of the knowledge graph entity sample in the knowledge graph, and reflects the semantic information corresponding to the knowledge graph entity sample.
[0073] Step S302: Based on the sample entity mention vector, the sample entity graph convolution vector, and the pre-acquired sample marking parameter, determine the loss function of the initial entity linking model.
[0074] The sample entity mention vector and the sample entity graph convolution vector respectively express the semantic information contained in the entity mention sample and the corresponding knowledge graph entity sample. Therefore, the similarity between the two vectors can reflect the semantic similarity between the entity mention sample and the corresponding knowledge graph entity sample, and can also reflect the training effect of the initial entity linking model. The sample marking parameter is used to reflect the difference in the training effect of the positive training data samples and negative training data samples on the entity linking model. When the knowledge graph entity sample in the quadruple data is a positive sample, the sample marking parameter is equal to 1, otherwise it is 0.
[0075] Specifically, the loss function of the initial entity linking model can be expressed by the following formula:
[0076] L = ||cosine(V e , V mention ) - y||,
[0077] where cosine(V1, V2) is used to calculate the cosine similarity between V1 and V2, V mention is the sample entity mention vector, V e is the sample entity graph convolution vector, and y is the sample marking parameter.
[0078] Step S303: Train the initial entity linking model based on this loss function.
[0079] Input the training data (including positive and negative sample pairs) into the initial entity linking model for training, and adjust the initial model parameters in the direction of minimizing the loss function. When the value of the loss function is less than a pre-set threshold, the entity linking model training is completed, and the model parameters are stored.
[0080] Through the above steps S301 - S303, by inputting the positive training data samples and negative training data samples into the initial entity linking model, the sample entity mention vector and the sample entity graph convolution vector are output, and the semantic information contained in the entity mention samples and the corresponding knowledge graph entity samples is respectively converted into corresponding vector representations; by determining the loss function of the initial entity linking model based on the sample entity mention vector, the sample entity graph convolution vector, and the pre-obtained sample marking parameters, the similarity between the sample entity mention vector and the sample entity graph convolution vector is measured through the loss function; by training the initial entity linking model based on this loss function, the model parameters corresponding to the minimum value of the loss function are obtained, and the trained entity linking model is obtained, improving the link accuracy of the knowledge graph entities based on the user's question in the question and answer scenario, and improving the answer accuracy based on the user's question.
[0081] In some of these embodiments, the initial entity linking model includes a text embedding module and an entity graph network embedding module. Figure 4 It is a flowchart of outputting the sample entity mention vector and the sample entity graph convolution vector based on the positive and negative samples of the training data in some embodiments of the present application. As Figure 4 shown, this process includes the following steps:
[0082] Step S401: Input the question sample and the knowledge graph entity sample into the text embedding module, and output the sample entity mention vector and the sample question vector. The knowledge graph entity sample includes the knowledge graph entity positive sample and the knowledge graph entity negative sample.
[0083] The text embedding module is a model that maps the entity mention samples in the question sample to the corresponding sample entity mention vectors through text embedding. This module can be composed of different models of natural language processing and combinations of different models, such as CoVe, Transformer, BERT, roberta, etc. The input of the text embedding module includes the question sample, the entity mention sample, and the knowledge graph entity sample, where the entity mention sample is input into the text embedding module included in the question sample. Corresponding to the question sample and the entity mention sample, the output of the text embedding module includes the sample entity mention vector and the sample question vector, where the sample question vector refers to the representation vector mapped by the question sample.
[0084] Step S402: Based on the sample problem vector, obtain the attention weights of the entity sample vectors in the entity graph network embedding module.
[0085] The entity graph network embedding module is a model that maps the entity graph network into corresponding representation vectors through entity graph embedding. The models used in this module can include GCN, deepWalk, SDNE, node2vec, etc. In this embodiment, the entity graph network embedding module is composed of multiple layers of CA-GCN modules and an attention mechanism is added. This attention mechanism can assign different weights to different entity relationships in the knowledge graph entity adjacency subgraph samples through weight adjustment. The attention weights can be obtained through the sample problem vector output by the text embedding module to reflect the influence of the context information in the problem sample on the importance of different entity relationships in the knowledge graph.
[0086] Specifically, for each CA-GCN module, its forward propagation formula can be expressed by the following formula:
[0087]
[0088] Where, represents the entity sample e in the knowledge graph entity adjacency subgraph sample i 's vector representation at the l-th layer, σ() is the activation function, R represents the set of entity relationships in the knowledge graph, and N i represents the set of adjacent nodes of the entity sample e i , represents the set of adjacent nodes of the entity sample e i linked by the entity relationship r, is the number of adjacent nodes of the entity sample e i linked by the entity relationship r, represents the weight matrix corresponding to the entity relationship r, which is a trainable parameter of the model, represents the weight matrix of node self-linking, which is a trainable parameter of the model, is the cross-attention matrix that fuses the problem sample and the knowledge graph information, and · represents element-wise multiplication.
[0089] Among them, the cross-attention matrix can be calculated first by the following formula:
[0090]
[0091] Where, j ∈ N i , is the vector representation corresponding to the problem sample, represents the entity sample e in the knowledge graph entity adjacency subgraph sample j 's vector representation at the l-th layer.
[0092] Then, perform normalization on and finally
[0093] After passing through the L-layer CA-GCN module, the output x of entity e at the L-th layer L is used as its graph convolutional vector V e .
[0094] Step S403: Input the knowledge graph entity adjacent subgraph sample into the entity graph network embedding module, and output the sample entity graph convolutional vector based on the attention weights.
[0095] The input of the entity graph network embedding module is the knowledge graph entity adjacent subgraph sample, which is the adjacent subgraph composed of the one-hop adjacent nodes corresponding to the positive or negative samples of the knowledge graph entities. The output of the entity graph network embedding module is the graph convolutional vector corresponding to the knowledge graph entity adjacent subgraph sample.
[0096] Through the above steps S401 - S403, by inputting the question sample and the knowledge graph entity sample into the text embedding module, the sample entity mention vector and the sample question vector are output. The knowledge graph entity sample includes the positive sample and the negative sample of the knowledge graph entity. The text embedding module maps the semantic information in the question sample and the knowledge graph entity sample into corresponding vectors; based on the sample question vector, the attention weights of the entity sample vectors in the entity graph network embedding module are obtained, that is, the importance of the entity relationships in the knowledge graph entity adjacent subgraph sample is adjusted through the context information in the question sample; by inputting the knowledge graph entity adjacent subgraph sample into the entity graph network embedding module and outputting the sample entity graph convolutional vector based on the attention weights, the sample entity graph convolutional vector can more accurately represent the entity relationships in the knowledge graph and improve the link accuracy of the entity link model.
[0097] In some embodiments, the text embedding module includes a BERT model. Figure 5 is a flowchart for obtaining the sample question vector based on the question sample and the knowledge graph entity sample in some embodiments of the present application, as Figure 5 shown, and this process includes the following steps:
[0098] Step S501: Concatenate the question sample, the knowledge graph entity sample, the CLS flag bit, and the SEP flag bit based on a predetermined format, and input them into the BERT model.
[0099] The BERT model is a context-based text embedding model, and its main structure is the Transformer. The CLS flag bit and the SEP flag bit are semantic identifiers for the input content of the BERT model. Among them, the CLS flag bit is used to identify whether there is a context relationship between two input sentences, while the SEP flag bit is an identifier used to separate two sentences.
[0100] Specifically, for the training data quadruple <q, m, e, Ge>, q and e in it are concatenated as [CLS]q[SEP]e, and then it can be converted into a numerical value by the bag-of-words model as the input of the BERT model. That is, the BERT model converts each word in the text into a one-dimensional vector, that is, a word vector. In addition, the model input also includes a text vector and a position vector. The BERT model takes the sum of the word vector, the text vector, and the position vector as the model input. The model output is the vector representation of each input word after fusing the full-text semantic information. Further, in order to utilize existing more general semantic information, the BERT model can also load pre-trained Chinese model parameters.
[0101] Step S502, based on the first output vector of the BERT model corresponding to the CLS flag bit, determine the sample problem vector.
[0102] The output vector of the BERT model corresponds to the input vector word by word. Among them, the CLS flag bit is used as the model input, and the first output vector of the corresponding BERT model is the semantic representation vector of the text. This vector can be used as the vector representation V corresponding to the problem sample. query This vector can guide the entity graph network embedding module to learn the importance degree of adjacent nodes through the attention mechanism.
[0103] Through the above steps S501 - S502, by concatenating the problem sample, the knowledge graph entity sample, the CLS flag bit, and the SEP flag bit based on a predetermined format, inputting them into the BERT model, obtaining the semantic information of the problem sample and the knowledge graph entity sample through the BERT model, and mapping them into corresponding vectors; by determining the sample problem vector based on the first output vector of the BERT model corresponding to the CLS flag bit, and using this sample problem vector to adjust the attention weights of the entity relationships in the knowledge graph entity adjacent subgraph sample, the training effect and link accuracy of the entity link model are improved.
[0104] In some of these embodiments, the text embedding module further includes a multi-layer perceptron model. Figure 6 It is a flowchart of obtaining the sample entity mention vector based on the problem sample and the knowledge graph entity sample in some embodiments of the present application, as Figure 6 shown, and this process includes the following steps:
[0105] Step S601: Obtain the second output vector of the BERT model corresponding to the start position of the entity mention sample and the third output vector of the BERT model corresponding to the end position of the entity mention sample.
[0106] The start position of the entity mention sample is the start position of the entity mention sample in the question sample input to the BERT model, and the end position of the entity mention sample is the end position of the entity mention sample in the question sample input to the BERT model. Since the input and output of the BERT model are word-by-word corresponding, the output vectors corresponding to the start position and the end position are also at the same sequence position.
[0107] Step S602: Concatenate the first output vector, the second output vector, and the third output vector, and input them into a multi-layer perceptron model to obtain the sample entity mention vector.
[0108] Input the first, second, and third output vectors output by the BERT model into the multi-layer perceptron model, and further process the vectors corresponding to the question sample and the entity mention sample, and finally obtain the sample entity mention vector.
[0109] Through the above steps S601 - S602, by obtaining the second output vector of the BERT model corresponding to the start position of the entity mention sample and the third output vector of the BERT model corresponding to the end position of the entity mention sample, select the vectors corresponding to the start position and the end position of the entity mention sample for further processing; by concatenating the first output vector, the second output vector, and the third output vector and inputting them into a multi-layer perceptron model to obtain the sample entity mention vector, the accuracy of the semantic information representation of the sample entity mention vector is improved.
[0110] In some of these embodiments, Figure 7 is a flowchart of obtaining the sample entity graph convolution vector based on the knowledge graph entity adjacent subgraph sample in some embodiments of the present application, as Figure 7 shown, this process includes the following steps:
[0111] Step S701: Initialize the knowledge graph entity adjacent subgraph sample based on the first output vector to obtain the corresponding entity sample vector.
[0112] During the model training process, it is necessary to first convert the entities in the knowledge graph entity adjacent subgraph sample into corresponding input vectors through vector initialization before inputting them into the entity graph network embedding module for training. The conversion method can be random initialization, or it can be initialized using pre-trained word vectors or sentence vectors. In this embodiment, the first output vector can be used to initialize the entities to increase the semantic information of the entity mention sample during the vector initialization process.
[0113] Step S702: Input the entity sample vector into the entity graph network embedding module, and output the sample entity graph convolution vector based on the attention weights.
[0114] Through the above steps S701 - S702, by initializing the knowledge graph entity adjacency subgraph sample based on the first output vector, the corresponding entity sample vector is obtained, and the semantic information of the entity mention sample is added during the vector initialization process; by inputting the entity sample vector into the entity graph network embedding module and outputting the sample entity graph convolution vector based on the attention weights, the graph convolution vector in the same vector space as the vector representation of the entity mention sample is obtained, improving the accuracy of the entity linking model.
[0115] In some of these embodiments, it also involves the establishment process of the BERT model. Before splicing the question sample, the knowledge graph entity sample, the CLS flag bit, and the SEP flag bit in a predetermined format and inputting them into the BERT model, the establishment process of the BERT model includes the following steps:
[0116] Step S11: Pre - train the BERT initial model to obtain the pre - trained model parameters.
[0117] The pre - training of the BERT model means training the BERT initial model with a large amount of Chinese training data. When the pre - training effect meets the expected requirements, the pre - trained model parameters are obtained. It is also possible to directly obtain the pre - trained model parameters from the Internet or a pre - trained model library.
[0118] Step S12: Establish the BERT model based on the pre - trained model parameters.
[0119] Load the pre - trained model parameters into the BERT model of the above - mentioned embodiment. This BERT model can be directly obtained based on the structure of the BERT initial model, or obtained by modifying the BERT initial model according to actual needs.
[0120] Through the above steps S11 - S12, by pre - training the BERT initial model to obtain the pre - trained model parameters, the time and computing resources required for subsequent training of the entity linking model are saved, and the training efficiency is improved; by establishing the BERT model based on the pre - trained model parameters, the training effect of the BERT model is improved.
[0121] The following describes and illustrates this embodiment through preferred embodiments. Figure 8 It is a schematic diagram of the entity linking model structure of this preferred embodiment. As Figure 8As shown in the figure, the entity link model of the present preferred embodiment includes a text embedding module 81 and an entity graph network embedding module 82. The text embedding module 81 includes a BERT model 811 and a multi-layer perceptron (MLP) 812. The entity graph network embedding module 82 includes multiple CA-GCN modules. The input of the BERT model 811 is the question sample q and the knowledge graph entity sample e, which are concatenated with the CLS flag bit and the SEP flag bit in a predetermined format. Among them, the question sample q includes N characters such as Tok1 to TokN, and the knowledge graph entity sample includes M characters such as Tok1 to TokM. The output of the BERT model 811 is a vector corresponding to the input character by character, including a C vector and an S vector corresponding to the CLS flag bit and the SEP flag bit respectively, and vectors such as T1 to TN, T1' to TM' corresponding to the characters. The input of the entity graph network embedding module 82 is the knowledge graph entity adjacency subgraph sample Ge, and the output is the sample entity graph convolution vector Ve.
[0122] Figure 9 is the flowchart of the knowledge graph entity link method of the present preferred embodiment. As Figure 9 shown, this process includes the following steps:
[0123] Step S901, collect the user's question, obtain the entity mentions in the user's question according to the existing NER model, manually label the knowledge graph entities corresponding to the entity mentions, and form a quadruple of the user's question, entity mentions, knowledge graph entities, and knowledge graph entity adjacency subgraphs as the positive sample of the training data;
[0124] Step S902, randomly extract knowledge graph entities from the knowledge graph entities, filter out the already linked knowledge graph entities, and form a quadruple with the user's question, entity mentions, and knowledge graph entity adjacency subgraphs as the negative sample of the training data;
[0125] Step S903, concatenate the user's question and the knowledge graph entity in the training data quadruple as [CLS] user's question [SEP] knowledge graph entity, and convert it into a numerical value using the bag-of-words model as the input of the BERT model. The BERT model loads the pre-trained Chinese model parameters;
[0126] Step S904, concatenate the vector output by Bert at the [CLS] position and the output vectors corresponding to the start and end positions of the entity mentions, and input them into the multi-layer perceptron to obtain the vector representation of the entity mentions;
[0127] Step S905, use the vector output by Bert at the [CLS] position as the initial vector of the entity in the knowledge graph entity adjacency subgraph, initialize the entity, and input it into the entity graph network embedding module;
[0128] Step S906: Use the vector output by Bert at the [CLS] position as the vector representation of the user question;
[0129] Step S907: Calculate the cross-attention matrix that fuses the user question and knowledge graph information based on the vector representation of the user question and the vector representations of the entities in the knowledge graph entity adjacency subgraph;
[0130] where j ∈ N i
[0131] Step S908: Normalize the cross-attention matrix;
[0132]
[0133] Step S909: For each CA-GCN module in the entity graph network embedding module, obtain the vector representation of the entity node according to the forward propagation formula;
[0134]
[0135] where, represents the vector representation of entity node e i at the l-th layer, σ() is the activation function, R represents the set of relationships in the knowledge graph, N i represents the set of adjacent nodes of entity e i , represents the set of adjacent nodes of entity e i linked by relationship r, is the number of adjacent nodes of entity e i linked by relationship r, represents the weight matrix corresponding to relationship r, which is a trainable parameter of the model, represents the weight matrix of node self-linking, which is a trainable parameter of the model. · represents element-wise multiplication.
[0136] Step S910: After passing through L layers of CA-GCN modules, use the output of entity e at the L-th layer as its graph convolution vector;
[0137] Step S911: Obtain the loss function based on the vector representation of the entity mention, the graph convolution vector, and the sample marking parameter, and train the initial entity linking model according to the loss function;
[0138] L = ||cosine(V e , V mention ) - y||,
[0139] Among them, cosine(V1, V2) is to calculate the cosine similarity between V1 and V2, and y is the positive and negative sample label of the data. If the entity e and the entity mention in the quadruple data are a pair of positive samples, it is 1, otherwise it is 0.
[0140] Step S912, when the value of the loss function satisfies the preset threshold range, stop training, obtain the entity linking model, and store the model parameters.
[0141] Step S913, receive user input, use the existing NER model to obtain the entity mentions in the user's question, use the existing BM25 algorithm to obtain the candidate knowledge graph entity set of the entity mentions, respectively obtain the adjacent subgraphs corresponding to the candidate entities, and form quadruples as the input of the entity linking model.
[0142] Step S914, according to the output of the entity linking model, obtain the vector representation of the entity mention and the set of candidate entity graph convolution vector representations, and select the candidate entity with the highest cosine similarity to the vector representation of the entity mention as the entity to be linked.
[0143] Through the above steps S901 to S914, using the knowledge graph entities linked to the entity mention samples through manual annotation as the positive samples of the training data, and randomly extracting entities as the negative samples of the training data after removing the positive samples of the entities, the effectiveness and accuracy of the model training are improved; taking advantage of the BERT model and the multi-layer perceptron in text semantic information extraction, as well as the advantage of the CA-GCN model in semantic information extraction of entity adjacent graphs, respectively obtain the sample entity mention vector and the sample entity graph convolution vector, and measure the similarity between the two through the loss function, improving the link accuracy of the knowledge graph entities based on the user's question in the question-and-answer scenario; adjusting the importance degree of the entity relationship in the knowledge graph entity adjacent subgraph sample through the sample question vector to generate attention weight data, and outputting the sample entity graph convolution vector based on the attention weight, making the sample entity graph convolution vector more accurately represent the entity relationship in the knowledge graph; initializing the adjacent subgraph sample based on the vector output by Bert at the [CLS] position to obtain the graph convolution vector in the same vector space as the vector representation of the entity mention sample, improving the accuracy of the entity linking model; obtaining the knowledge graph entity most similar to the entity mention as the link target according to the comparison of the vector similarity between each candidate knowledge graph entity and the entity mention by the entity linking model, improving the effect of entity linking in the question-and-answer scenario.
[0144] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0145] In some embodiments, the present application further provides a knowledge graph entity linking device, which is used to implement the above-mentioned embodiments and preferred implementation manners. Those that have been described will not be repeated here. The following terms "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. In some embodiments, Figure 10 is a structural block diagram of the knowledge graph entity linking device in this embodiment, as Figure 10 shown. The device includes:
[0146] A first acquisition module 1101, configured to acquire a positive training data sample based on a problem sample, an entity mention sample, a knowledge graph entity positive sample, and a knowledge graph entity adjacent sub-graph sample. The entity mention sample is acquired based on the problem sample, the knowledge graph entity positive sample is acquired based on the annotated entity of the entity mention sample in the knowledge graph, and the knowledge graph entity adjacent sub-graph sample is acquired based on the entity relationship of the knowledge graph entity positive sample in the knowledge graph;
[0147] A second acquisition module 1102, configured to acquire a negative training data sample based on a problem sample, an entity mention sample, a knowledge graph entity negative sample, and a corresponding knowledge graph entity adjacent sub-graph sample. The knowledge graph entity negative sample is randomly acquired based on an entity in the knowledge graph that has no annotation relationship with the entity mention sample;
[0148] A training module 1103, configured to train an initial entity linking model based on the positive training data sample and the negative training data sample to obtain an entity linking model;
[0149] A determination module 1104, configured to input a user problem, an entity mention, a candidate knowledge graph entity, and a corresponding knowledge graph entity adjacent sub-graph into the trained entity linking model to determine a target knowledge graph entity linked to the entity mention; the entity mention is acquired based on the user problem, the candidate knowledge graph entity is acquired based on the entity mention, and the knowledge graph entity adjacent sub-graph is acquired based on the entity relationship of the candidate knowledge graph entity in the knowledge graph.
[0150] The knowledge graph entity linking device of this embodiment obtains positive training data through the first acquisition module 1101 based on problem samples, entity mention samples, knowledge graph entity positive samples, and knowledge graph entity adjacent subgraph samples, and uses the knowledge graph entities linked to the entity mention samples through manual annotation as the sample data for positive training, ensuring the accuracy of the positive training data and the effectiveness of training; through the second acquisition module 1102, based on problem samples, entity mention samples, knowledge graph entity negative samples, and corresponding knowledge graph entity adjacent subgraph samples, negative training data is obtained, that is, entities are randomly selected as knowledge graph entity negative samples after excluding the positive samples of entities in the knowledge graph, improving the randomness of the negative training data and the effectiveness of training; through the training module 1103, the entity linking initial model is trained based on the positive training data and the negative training data to obtain an entity linking model, improving the accuracy of the training data of the entity linking model; through the determination module 1104, the user question, entity mention, candidate knowledge graph entities, and corresponding knowledge graph entity adjacent subgraphs are input into the trained entity linking model to determine the target knowledge graph entity linked to the entity mention, and the knowledge graph entity most similar to the entity mention is obtained as the linking target, improving the linking effect of the entity linking model and solving the problems of poor entity consistency model effect and low entity linking accuracy in the question and answer scenario in the related technology.
[0151] In some embodiments, the training module includes an input sub-module, a determination sub-module, and a training sub-module. The input sub-module is used to input the positive training data and the negative training data into the entity linking initial model and output the sample entity mention vector and the sample entity graph convolution vector; the determination sub-module is used to determine the loss function of the entity linking initial model based on the sample entity mention vector, the sample entity graph convolution vector, and the pre-obtained sample marking parameters; the training sub-module is used to train the entity linking initial model based on the loss function.
[0152] The knowledge graph entity linking device of this embodiment inputs positive training data samples and negative training data samples into an initial entity linking model through an input sub-module, and outputs sample entity mention vectors and sample entity graph convolution vectors, converting the semantic information contained in the entity mention samples and the corresponding knowledge graph entity samples into corresponding vector representations respectively; through a determination sub-module, based on the sample entity mention vectors, sample entity graph convolution vectors, and pre-acquired sample marking parameters, it determines the loss function of the initial entity linking model, and measures the similarity between the sample entity mention vectors and the sample entity graph convolution vectors through the loss function; through a training sub-module, it trains the initial entity linking model based on the loss function, obtains the model parameters corresponding to the minimum value of the loss function, and obtains a trained entity linking model, improving the linking accuracy of knowledge graph entities based on user questions in a question-and-answer scenario and improving the answer accuracy based on user questions.
[0153] In some embodiments, the initial entity linking model includes a text embedding module and an entity graph network embedding module. The input sub-module includes a first input unit, an acquisition unit, and a second input unit. The first input unit is used to input question samples and knowledge graph entity samples into the text embedding module, and outputs sample entity mention vectors and sample question vectors. The knowledge graph entity samples include positive knowledge graph entity samples and negative knowledge graph entity samples; the acquisition unit is used to obtain the attention weights of the entity sample vectors in the entity graph network embedding module based on the sample question vectors; the second input unit is used to input knowledge graph entity adjacency sub-graph samples into the entity graph network embedding module and output sample entity graph convolution vectors based on the attention weights.
[0154] The knowledge graph entity linking device of this embodiment inputs question samples and knowledge graph entity samples into the text embedding module through the first input unit, and outputs sample entity mention vectors and sample question vectors. The knowledge graph entity samples include positive knowledge graph entity samples and negative knowledge graph entity samples. The text embedding module maps the semantic information in the question samples and knowledge graph entity samples into corresponding vectors; the acquisition unit obtains the attention weights of the entity sample vectors in the entity graph network embedding module based on the sample question vectors, that is, adjusts the importance degree of the entity relationships in the knowledge graph entity adjacency sub-graph samples through the context information in the question samples; the second input unit inputs the knowledge graph entity adjacency sub-graph samples into the entity graph network embedding module and outputs sample entity graph convolution vectors based on the attention weights, making the sample entity graph convolution vectors more accurately represent the entity relationships in the knowledge graph and improving the linking accuracy of the entity linking model.
[0155] In some embodiments, the text embedding module includes a BERT model. The first input unit includes a first input subunit and a determination subunit. The first input subunit is configured to splice a question sample, a knowledge graph entity sample, a CLS flag bit, and a SEP flag bit based on a predetermined format and input them into the BERT model. The determination subunit is configured to determine a sample question vector based on a first output vector of the BERT model corresponding to the CLS flag bit.
[0156] In the knowledge graph entity linking device of this embodiment, the first input subunit splices a question sample, a knowledge graph entity sample, a CLS flag bit, and a SEP flag bit based on a predetermined format and inputs them into the BERT model. The BERT model is used to obtain semantic information of the question sample and the knowledge graph entity sample and map them into corresponding vectors. The determination subunit determines a sample question vector based on a first output vector of the BERT model corresponding to the CLS flag bit, and uses the sample question vector to adjust the attention weights of entity relationships in the knowledge graph entity adjacent subgraph sample, improving the training effect and linking accuracy of the entity linking model.
[0157] In some embodiments, the text embedding module further includes a multi-layer perceptron model. The first input unit further includes an acquisition subunit and a second input subunit. The acquisition subunit is configured to acquire a second output vector of the BERT model corresponding to the start position of the entity mention sample and a third output vector of the BERT model corresponding to the end position of the entity mention sample. The second input subunit is configured to splice the first output vector, the second output vector, and the third output vector and input them into the multi-layer perceptron model to obtain a sample entity mention vector.
[0158] In the knowledge graph entity linking device of this embodiment, the acquisition subunit acquires a second output vector of the BERT model corresponding to the start position of the entity mention sample and a third output vector of the BERT model corresponding to the end position of the entity mention sample, and selects the vectors corresponding to the start position and the end position of the entity mention sample for further processing. The second input subunit splices the first output vector, the second output vector, and the third output vector and inputs them into the multi-layer perceptron model to obtain a sample entity mention vector, improving the accuracy of semantic information representation of the sample entity mention vector.
[0159] In some embodiments, the second input unit includes an initialization subunit and a third input subunit. The initialization subunit is configured to initialize the knowledge graph entity adjacent subgraph sample based on the first output vector to obtain a corresponding entity sample vector. The third input subunit is configured to input the entity sample vector into the entity graph network embedding module and output a sample entity graph convolution vector based on the attention weights.
[0160] The knowledge graph entity linking device of this embodiment initializes the knowledge graph entity adjacent subgraph samples based on the first output vector through the initialization subunit to obtain corresponding entity sample vectors, and adds the semantic information of the entity mention samples during the vector initialization process; the third input subunit inputs the entity sample vectors into the entity graph network embedding module, and outputs the sample entity graph convolution vectors based on the attention weights to obtain the graph convolution vectors in the same vector space as the vector representation of the entity mention samples, improving the accuracy of the entity linking model.
[0161] In some embodiments, the knowledge graph entity linking device further includes a third acquisition module and a building module. The third acquisition module is used to pre-train the BERT initial model to obtain pre-trained model parameters; the building module is used to build a BERT model based on the pre-trained model parameters.
[0162] The knowledge graph entity linking device of this embodiment pre-trains the BERT initial model through the third acquisition module to obtain pre-trained model parameters, saving the time and computing resources required for subsequent training of the entity linking model and improving the training efficiency; the building module builds a BERT model based on the pre-trained model parameters, improving the training effect of the BERT model.
[0163] In some embodiments, the present application also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the knowledge graph entity linking method of the above embodiment.
[0164] Optionally, the computer device may further include a transmission device and input / output devices. Among them, the transmission device is connected to the processor, and the input / output devices are connected to the processor.
[0165] In addition, in combination with the knowledge graph entity linking method provided in the above embodiment, a storage medium may also be provided in this embodiment to implement it. A program is stored on the storage medium; when the program is executed by the processor, it implements any one of the knowledge graph entity linking methods in the above embodiment.
[0166] It should be noted that the specific examples in this embodiment may refer to the examples described in the above embodiment and optional implementation manners, and will not be repeated in this embodiment.
[0167] It should be understood that the specific embodiments described here are only used to explain this application, rather than to limit it. According to the embodiments provided by the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0168] Obviously, the accompanying drawings are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar situations based on these drawings without creative efforts. Additionally, it can be understood that although the work done during the development here may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in the present application are only routine technical means and should not be regarded as insufficient disclosure of the present application.
[0169] The term "embodiment" in this application means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in this application can be combined with other embodiments without conflict.
[0170] The above-described embodiments only express several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A knowledge graph entity linking method, characterized in that, The method includes: Based on problem samples, entity mention samples, knowledge graph entity positive samples, and knowledge graph entity adjacent sub-graph samples, obtaining positive training data samples. The entity mention samples are obtained based on the problem samples. The knowledge graph entity positive samples are obtained based on the annotated entities of the entity mention samples in the knowledge graph. The knowledge graph entity adjacent sub-graph samples are obtained based on the entity relationships of the knowledge graph entity positive samples in the knowledge graph. Based on the problem samples, the entity mention samples, knowledge graph entity negative samples, and corresponding knowledge graph entity adjacent sub-graph samples, obtaining negative training data samples. The knowledge graph entity negative samples are randomly obtained based on entities in the knowledge graph that have no annotation relationship with the entity mention samples. Training an initial entity linking model based on the positive training data samples and negative training data samples to obtain an entity linking model. Inputting a user question, an entity mention, candidate knowledge graph entities, and corresponding knowledge graph entity adjacent sub-graphs into the trained entity linking model to determine a target knowledge graph entity linked to the entity mention. The entity mention is obtained based on the user question. The candidate knowledge graph entities are obtained based on the entity mention. The knowledge graph entity adjacent sub-graphs are obtained based on the entity relationships of the candidate knowledge graph entities in the knowledge graph.
2. The method according to claim 1, characterized in that, The training of the initial entity linking model based on the positive training data samples and negative training data samples includes: Inputting the positive training data samples and negative training data samples into the initial entity linking model to output sample entity mention vectors and sample entity graph convolutional vectors. Based on the sample entity mention vectors, sample entity graph convolutional vectors, and pre-obtained sample marking parameters, determining a loss function of the initial entity linking model. Training the initial entity linking model based on the loss function.
3. The method according to claim 2, wherein The initial entity linking model includes a text embedding module and an entity graph network embedding module. The inputting of the positive training data samples and negative training data samples into the initial entity linking model to output sample entity mention vectors and sample entity graph convolutional vectors includes: Inputting the problem samples and knowledge graph entity samples into the text embedding module to output the sample entity mention vectors and sample problem vectors. The knowledge graph entity samples include the knowledge graph entity positive samples and knowledge graph entity negative samples. Based on the sample problem vectors, obtaining attention weights of entity sample vectors in the entity graph network embedding module. Inputting the knowledge graph entity adjacent sub-graph samples into the entity graph network embedding module and outputting the sample entity graph convolutional vectors based on the attention weights.
4. The method according to claim 3, characterized in that, The text embedding module includes a BERT model. The inputting of the problem samples and knowledge graph entity samples into the text embedding module to output the sample entity mention vectors and sample problem vectors includes: Concatenating the problem samples, knowledge graph entity samples, CLS flag bits, and SEP flag bits in a predetermined format and inputting them into the BERT model. Determine the sample problem vector based on the first output vector of the BERT model corresponding to the CLS flag bit.
5. The method according to claim 4, characterized in that, The text embedding module further includes a multi-layer perceptron model. The step of inputting the problem sample and the knowledge graph entity sample into the text embedding module and outputting the sample entity mention vector and the sample problem vector further includes: Obtain a second output vector of the BERT model corresponding to the start position of the entity mention sample, and a third output vector of the BERT model corresponding to the end position of the entity mention sample. Concatenate the first output vector, the second output vector, and the third output vector, and input them into the multi-layer perceptron model to obtain the sample entity mention vector.
6. The method according to claim 4, characterized in that, The step of inputting the knowledge graph entity adjacent subgraph sample into the entity graph network embedding module and outputting the sample entity graph convolution vector based on the attention weight includes: Initialize the knowledge graph entity adjacent subgraph sample based on the first output vector to obtain a corresponding entity sample vector. Input the entity sample vector into the entity graph network embedding module, and output the sample entity graph convolution vector based on the attention weight.
7. The method according to claim 4, characterized in that, Before concatenating the problem sample, the knowledge graph entity sample, the CLS flag bit, and the SEP flag bit in a predetermined format and inputting them into the BERT model, the method further includes: Perform pre-training on the initial BERT model to obtain pre-trained model parameters. Establish the BERT model based on the pre-trained model parameters.
8. A knowledge graph entity linking device, characterized in that, The apparatus includes: A first acquisition module, configured to acquire a positive training data sample based on a problem sample, an entity mention sample, a positive knowledge graph entity sample, and a knowledge graph entity adjacent subgraph sample, where the entity mention sample is acquired based on the problem sample, the positive knowledge graph entity sample is acquired based on an annotated entity of the entity mention sample in the knowledge graph, and the knowledge graph entity adjacent subgraph sample is acquired based on an entity relationship of the positive knowledge graph entity in the knowledge graph. A second acquisition module, configured to acquire a negative training data sample based on the problem sample, the entity mention sample, a negative knowledge graph entity sample, and a corresponding knowledge graph entity adjacent subgraph sample, where the negative knowledge graph entity sample is randomly acquired based on an entity in the knowledge graph that has no annotation relationship with the entity mention sample. A training module, configured to train an initial entity linking model based on the positive training data sample and the negative training data sample to obtain an entity linking model. A determination module, configured to input a user question, an entity mention, a candidate knowledge graph entity, and a corresponding knowledge graph entity adjacent subgraph into the trained entity linking model to determine a target knowledge graph entity linked to the entity mention; the entity mention is acquired based on the user question, the candidate knowledge graph entity is acquired based on the entity mention, and the knowledge graph entity adjacent subgraph is acquired based on an entity relationship of the candidate knowledge graph entity in the knowledge graph.
9. A computer device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the knowledge graph entity linking method according to any one of claims 1 to 7.
10. A readable storage medium, on which a program is stored, characterized in that, When the program is executed by the processor, the steps of the knowledge graph entity linking method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Automatic question and answer processing method and device, equipment and storage medium
CN113360616A
Multi-knowledge graph question and answer model training method and system and storage medium
CN116011548A