Chinese short text joint entity disambiguation method and system based on graph convolutional neural network

Through the PinSageGCN model based on graph convolution neural network, the accuracy and computing resource consumption of entity disambiguation in Chinese short texts are solved, efficient entity disambiguation task is realized, and the accuracy and robustness of entity disambiguation in Chinese short texts are improved.

CN120493913AActive Publication Date: 2025-08-15BEIJING INSTITUTE OF PETROCHEMICAL TECHNOLOGY

Patent Information

Application Number
CN202510582533.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-15
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The entity disambiguation task in Chinese short text faces the problems of limited length and sparse semantic features. Traditional methods are difficult to accurately deal with entity ambiguity in short text, and the deep learning model has high computational complexity and high resource consumption.

Method used

The PinSageGCN model based on graph convolution neural network is used to calculate the similarity between entity reference and candidate entities through BERT+BiLSTM, combined with LTP word segmentation and dependency syntax analysis, a candidate entity sorting graph is constructed, the node features are encoded using word2vec, and multi-layer iterative training is performed through the heterogeneous PinSageGCN model, and the similarity model is fused for entity disambiguation.

Benefits of technology

It improves the accuracy of entity disambiguation, shows strong robustness, can handle noisy data and handles larger data sets without increasing computing resource consumption, overcoming memory and video memory limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493913A_ABST
    Figure CN120493913A_ABST
Patent Text Reader

Abstract

The invention discloses a Chinese short text joint entity disambiguation method and system based on a graph convolutional neural network, and the method comprises the steps: obtaining an original short text, obtaining a candidate entity set referred by each entity, and generating a training data set; training the training data set to obtain a similarity score of the entity reference and each candidate entity; the method comprises the following steps: performing word segmentation and dependency syntactic analysis on an original short text, converting the original short text into graph data according to an analysis result, adding candidate entities as nodes into the graph data to obtain an adjacent matrix, and encoding each node by using word2vec to obtain a feature matrix; constructing a GCN entity disambiguation model, fusing the similarity score as a weight into the adjacency matrix, and inputting the feature matrix and the adjacency matrix with the weight into the GCN entity disambiguation model for training; and outputting a feature vector of each node, obtaining a global feature score of each candidate entity node by using a full connection layer, and sorting according to the global feature scores to complete entity disambiguation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and in particular relates to a Chinese short text joint entity disambiguation method and system based on a graph convolutional neural network. Background Art

[0002] With the development of big data technologies, the amount of information in today's society is exploding, leading to information overload and posing challenges for users in obtaining the information they need and accurately. Due to the diverse nature of natural language, including polysemy and synonymy, accurately acquiring target information requires processing a large amount of false, redundant, and noisy information. However, computers often cannot identify the objects or concepts referred to by entity names in text, making it difficult for humans to deeply understand the complexity of natural language expressions. For example, when the word "Nezha" appears in a text, a human can accurately understand that it refers to the movie "Nezha 2: The Devil Child Comes into the World," not the Nezha from the game or the Third Prince Nezha from Journey to the West. However, computers struggle to do the same. This often impacts the effectiveness of a range of natural language processing tasks, including question understanding, knowledge point interpretation, and knowledge base construction.

[0003] Entity disambiguation has always been a task of common concern in fields such as natural language processing, information retrieval, and text intelligent mining. It is the task of associating the referential terms (entity references) representing entities in the text with entity entries in a specific knowledge base, that is, mapping the entity references mentioned in the text to the correct entries in the specified knowledge base. When faced with chaotic and disordered data information, people are eager to accurately describe various entities and concepts in the real world and achieve mapping from unstructured information to structured semantic descriptions. The entity disambiguation task provides an effective technical means for analyzing large amounts of unstructured text information and plays an important role in many natural language processing tasks, as follows:

[0004] (1) Information retrieval

[0005] With the increasing abundance of internet data resources, information retrieval applications such as Baidu and Google have made it easier for people to access information. However, traditional search engines, which rely solely on keywords to mechanically return search results, are no longer able to meet user needs. Furthermore, the inherent diversity and ambiguity of natural language pose significant challenges to retrieving information with high accuracy. Entity linking technology effectively eliminates ambiguity in search results, allowing precise matching of entities and related content within knowledge bases or knowledge graphs. This not only enables search engines to accurately understand users' true intent but also provides more precise search results, driving information retrieval technology towards intelligence.

[0006] (2) Knowledge base and knowledge graph completion

[0007] The world is changing rapidly, and network data is constantly increasing. This means that new entities and the relationships between entities are also increasing. This situation has led to incompleteness in many knowledge bases and knowledge graphs. Entity linking technology can combine unstructured short text and structured knowledge bases or knowledge graphs. By adding new knowledge to knowledge bases and knowledge graphs, it effectively completes them and promotes the application and development of knowledge bases and knowledge graphs in various fields.

[0008] (3) Intelligent Question and Answer

[0009] Many intelligent question-answering systems are constantly evolving, leveraging advancements in artificial intelligence. These systems allow users to enter text to ask questions. Since users typically ask concise questions, the input text is often short. The system then provides an answer. During this process, algorithms analyze the short text provided by the user, using entity linking technology to disambiguate entities within it, thereby helping to more accurately find answers. This demonstrates the importance of short text entity linking in intelligent question-answering systems.

[0010] While researchers have achieved significant results with long texts like news, the vast majority of information on social media exists in the form of short texts, such as search queries, news headlines, and social media comments. These short texts contain irregular expressions and arbitrary wording, and Chinese sentences lack the whitespace delimiters found in English. Therefore, correctly understanding short texts and extracting useful knowledge from them remains a challenge. The challenge of entity disambiguation in Chinese short texts stems primarily from their limited length and sparse semantic features. Therefore, traditional entity disambiguation methods are difficult to directly and effectively apply to short texts.

[0011] Early entity disambiguation models used a simple language model, the bag-of-words model. Many researchers used it to represent the context of entity references and the description text of candidate entities as vectors, and then used the cosine similarity algorithm to calculate similarity. However, the bag-of-words model typically uses one-hot encoding, which is very sparse and does not consider grammar and word order, resulting in poor semantic representation. Subsequently, the Word2Vec method was developed, which uses dense vectors to represent words in text. Its advantage is that semantically similar words have closer vectors, and the relationships between some words can be represented by vector operations. However, this method lacks positional modeling and cannot consider the correlation between all words in the entire sentence. With the development of deep learning, this technology has been frequently applied to various natural language processing tasks, such as neural networks like RNN and BERT and pre-trained language models. However, the stacking of too many models increases computational complexity, sacrificing significant computing resources while slightly improving efficiency. Furthermore, many deep learning methods for entity disambiguation fail to consider the correlation between entity references and candidate entities and contextual terms, resulting in room for improvement in disambiguation accuracy. Summary of the Invention

[0012] To address the problems of the prior art, the present invention provides a method and system for joint entity disambiguation of Chinese short texts based on graph convolutional neural networks. This method uses a heterogeneous PinSage graph convolutional neural network model (PinSageGCN), trained based on four different types of edge relationships, and integrates the prediction results of similarity models to complete the entity disambiguation task. The model exhibits strong robustness when processing noisy data. Even if there is a certain degree of uncertainty in the input data, the GCN in the present invention can still provide relatively accurate results. Furthermore, the model overcomes the limitations of memory and video memory during training and can process larger data sets without consuming a lot of computing resources.

[0013] To achieve the above object, the present invention provides the following solutions:

[0014] A Chinese short text joint entity disambiguation method based on graph convolutional neural network, the method comprising:

[0015] Obtain the original short text containing entity references, and obtain the candidate entity set of each entity reference based on the existing knowledge base to generate a training dataset;

[0016] The training dataset is trained using a BERT+BiLSTM model to obtain a similarity score between the entity reference and each candidate entity in the candidate entity set;

[0017] Use LTP to perform word segmentation and dependency syntactic analysis on the original short text, convert the analysis results into graph data, and add the candidate entities as nodes to the graph data to obtain the adjacency matrix of the complete graph data, and then use word2vec to encode each node to obtain the feature matrix of the complete graph data;

[0018] Calculate the cosine similarity between each connected node and incorporate it into the adjacency matrix as the weight of the edge relationship to build a PinSageGCN entity disambiguation model. Input the feature matrix and the adjacency matrix with weights into the PinSageGCN entity disambiguation model to start training and update the node features.

[0019] The PinSageGCN entity disambiguation model is used to obtain the features of each node after multiple layers of iteration. Based on the four edge relationships, four PinSageGCN entity disambiguation models, namely heterogeneous PinSageGCN models, are used to complete the entity disambiguation task.

[0020] Preferably, the method of using the BERT+BiLSTM model to train the training dataset to obtain a similarity score between the entity reference and each candidate entity in the candidate entity set includes:

[0021] Given an entity reference m and context Z, for each candidate entity e∈E m , concatenate Z+M, candidate entity description C, and two special tags [CLS] and [SEP] from the BERT vocabulary as the input sequence; [CLS] indicates the beginning of the sequence, and [SEP] separates the input of different segments; after being concatenated into a string, it is sent to BERT for word segmentation and encoded into a vector T n ; Superimpose the BiLSTM model for feature extraction, learn the semantic order of entity reference context and candidate entity description text, and obtain the BiLSTM output vector p n :

[0022] p n =BiLSTM(T n )=BiLSTM(BERT([CLS]text[SEP]C mi [SEP]))

[0023] Among them, text represents a short text Z+M, C mi A candidate entity description text representing an entity reference m;

[0024] Calculate text and C mi The similarity between them is used to generate feature vectors and pass them through the fully connected layer, Dropout, and sigmoid for binary classification to generate the final context similarity S txt (m, e):

[0025] S ctxt (m,e)=sigmod(Dropout(Dense(p n ))).

[0026] Preferably, the graph data is: candidate entity ranking graph G=(V, E), where V represents a node set and E represents an edge set.

[0027] Preferably, the method of adding the candidate entity as a node to the graph data to obtain the adjacency matrix of the complete graph data includes:

[0028]

[0029] Where i and j represent the two nodes in the i-th row and j-th column of the adjacency matrix A.

[0030] Preferably, the method of encoding each node using word2vec to obtain a feature matrix of the complete graph data includes:

[0031] The entity reference, context words, and candidate entities of each entity reference obtained from the knowledge base are taken as nodes, and the node vectors are initialized for the nodes. The pre-trained Word2vec model is used to convert the context words z i ∈Z,m i ∈M includes context words and entity references, and the conversion word vector is recorded as w m ,w z :

[0032]

[0033] Assume that the candidate entity description text consists of j words and is converted into a vector representation, denoted as w1,w2,...,w j , that is, each word in the candidate entity description text has a corresponding word vector, and then the vector representations of all words are summed and averaged to obtain a candidate entity vector representation result w c :

[0034]

[0035] The entity reference, context words and node vectors of candidate entities are uniformly represented as node features of the node set V of the candidate entity ranking graph G in Represents the initial eigenvector of a node.

[0036] Preferably, the method for calculating the cosine similarity between each connected node includes:

[0037]

[0038] Where h i and h j Represents the feature vectors of node i and node j.

[0039] Preferably, the feature matrix and the adjacency matrix with weights are input into PinSageGCN to start training, and the method for updating node features includes:

[0040] Obtain the neighbor nodes of the node according to the obtained adjacency matrix A'. For each node i, j represents a certain neighbor node, and the neighbor node set is N(i):

[0041] N(i)={j|A' ij ≠0};

[0042] At the same time, the information about the relationship between nodes in the graph structure and the semantic information in the node feature vector are used to calculate the node neighbor message distribution and aggregation vector n i ;For each node i, the set of neighboring nodes is N(i);;

[0043]

[0044] Among them, Dense is a fully connected layer network, γ() represents aggregation;

[0045] The feature vector of node i is also updated in the fully connected layer:

[0046]

[0047] Among them, s i The feature vector representing the node self-update;

[0048] Aggregate neighbor information n i and the node's own encoding i Add and regularize to get the new node vector representation, which is the iteration of the node during the training process:

[0049]

[0050] in, Represents the intermediate result of node update features.

[0051] Preferably, the features of each node after multi-layer iteration are obtained through the PinSageGCN entity disambiguation model, and the method for completing the entity disambiguation task by using four PinSageGCN entity disambiguation models, i.e., heterogeneous PinSageGCN models, based on four edge relationships includes:

[0052] Based on the four types of edge relationships, the heterogeneous PinSageGCN method is used for training, and the node features of each node under the four edge relationships are output. The various features of the nodes are fused and spliced, and the fully connected layer is used to obtain the feature scores of the candidate entity nodes. Then, combined with the similarity scores, joint learning is performed to obtain the joint disambiguation score of each candidate entity node. The nodes are sorted according to the joint disambiguation score, and the PinSageGCN entity disambiguation model training is supervised by annotations to finally complete entity disambiguation.

[0053] The present invention also provides a Chinese short text joint entity disambiguation system based on a graph convolutional neural network, the system is used to implement any one of the methods, the system includes: an acquisition module, a training module, an encoding module, an update module and a sorting module;

[0054] The acquisition module is used to acquire the original short text containing entity references, and obtain a candidate entity set for each entity reference based on the existing knowledge base to generate a training data set;

[0055] The training module is used to train the training data set using the BERT+BiLSTM model to obtain a similarity score between the entity reference and each candidate entity in the candidate entity set;

[0056] The encoding module is used to perform word segmentation and dependency syntax analysis on the original short text using LTP, convert the analysis results into graph data, and add the candidate entities as nodes to the graph data to obtain an adjacency matrix of the complete graph data, and then use word2vec to encode each node to obtain a feature matrix of the complete graph data;

[0057] The update module is used to calculate the cosine similarity between each connected node, incorporate it into the adjacency matrix as the weight of the edge relationship, build a PinSageGCN entity disambiguation model, input the feature matrix and the adjacency matrix with weights into the PinSageGCN entity disambiguation model to start training, and update the node features;

[0058] The sorting module is used to obtain the features of each node after multi-layer iteration through the PinSageGCN entity disambiguation model, and complete the entity disambiguation task based on four edge relationships using four PinSageGCN entity disambiguation models, namely heterogeneous PinSageGCN models.

[0059] Compared with the prior art, the present invention has the following beneficial effects:

[0060] The present invention first uses the pre-trained BERT model and BiLSTM model to calculate the similarity between entity reference and candidate entity, fully captures more semantic information, reduces the sparsity of vectors, and improves the ability of semantic representation. In addition, the dependency syntactic analysis method is used to obtain the correlation between words in short texts, thereby constructing a graph data set, and connecting candidate entities. The relationship between context words, entity references and candidate entities is reflected in the graph data set, and the relationship between the three is fully considered, so that the neural network deeply extracts their inherent features to improve the accuracy of entity disambiguation. Finally, the heterogeneous PinSage graph convolutional neural network model (PinSageGCN) is used to train based on four different types of edge relationships and fuse the predicted results of the similarity model to complete the entity disambiguation task. The model shows strong robustness when processing noisy data. Even if there is a certain degree of uncertainty in the input data in the present invention, GCN can still give relatively accurate results, and the model overcomes the limitations of memory and video memory during training and can process larger data sets without consuming a lot of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0062] Figure 1 1 is a schematic diagram of the overall steps of a Chinese short text joint entity disambiguation method based on a graph convolutional neural network according to an embodiment of the present invention;

[0063] Figure 2 Detailed schematic diagram of a Chinese short text joint entity disambiguation process based on a graph convolutional neural network according to an embodiment of the present invention;

[0064] Figure 3 This is a schematic diagram of a candidate entity recall and labeling process according to an embodiment of the present invention;

[0065] Figure 4 This is a schematic diagram of a similarity calculation model architecture according to an embodiment of the present invention;

[0066] Figure 5 This is a sample diagram of entity disambiguation candidate entity ranking according to an embodiment of the present invention;

[0067] Figure 6 This is a schematic diagram of a process for constructing a candidate entity ranking graph for entity disambiguation according to an embodiment of the present invention;

[0068] Figure 72 is a schematic diagram of a PinSageGCN entity disambiguation model architecture according to an embodiment of the present invention;

[0069] Figure 8 This is a schematic diagram of candidate entity ranking according to an embodiment of the present invention.

[0070] Figure 9 This is a schematic diagram of the deployment of an entity disambiguation operating environment in an embodiment of the present invention. DETAILED DESCRIPTION

[0071] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0072] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0073] Example 1

[0074] like Figure 1 As shown, this embodiment provides a Chinese short text joint entity disambiguation method based on a graph convolutional neural network, the method comprising:

[0075] Obtain the original short text containing entity references, and obtain the candidate entity set of each entity reference based on the existing knowledge base to generate a training dataset;

[0076] The training dataset is trained using a BERT+BiLSTM model to obtain a similarity score between the entity reference and each candidate entity in the candidate entity set;

[0077] Use LTP to perform word segmentation and dependency syntactic analysis on the original short text, convert it into graph data based on the analysis results, add the candidate entities as nodes to the graph data, obtain the adjacency matrix of the complete graph data, and then use word2vec to encode each node to obtain the feature matrix of the complete graph data;

[0078] Calculate the cosine similarity between each connected node and incorporate it into the adjacency matrix as the weight of the edge relationship to build a PinSageGCN entity disambiguation model. Input the feature matrix and the adjacency matrix with weights into the PinSageGCN entity disambiguation model to start training and update the node features.

[0079] The PinSageGCN entity disambiguation model is used to obtain the features of each node after multiple layers of iteration. Based on the four edge relationships, four PinSageGCN entity disambiguation models, namely heterogeneous PinSageGCN models, are used to complete the entity disambiguation task.

[0080] like Figure 2 The figure shows a schematic diagram of the Chinese short text joint entity disambiguation process based on graph convolutional neural network of the present invention. The specific implementation process of each step of the present invention is as follows.

[0081] (1) Obtain the original short text containing entity references from the entity disambiguation dataset, and obtain the candidate entity set for each entity reference based on the existing knowledge base, including the annotation of whether each candidate entity is the correct mapping entity of the entity reference, thereby generating a training dataset.

[0082] (2) Using BERT (Bidirectional Encoder Representations from Transformer) and BiLSTM (Bi-directional Long Short-Term Memory) models to train the training dataset, a similarity score between the entity reference and each candidate entity in the candidate entity set is obtained.

[0083] (3) Use the natural language processing toolkit LTP (Language Technology Platform) to perform word segmentation and dependency syntax analysis on the original short text, convert it into graph data based on the analysis results, and add the candidate entities as nodes to the graph data to obtain the adjacency matrix of the complete graph data. Then use word2vec (Word To Vector) to encode each node, that is, encode the word segmentation results of the original short text and the candidate entity description text to obtain the feature matrix of the complete graph data, ready for subsequent input into the PinSageGCN model.

[0084] (4) Calculate the cosine similarity between each connected node as the weight of the edge relationship, build a PinSageGCN entity disambiguation model, input the feature matrix and the adjacency matrix with weights into PinSageGCN to start training, and update the node features.

[0085] (5) Based on the four types of edge relationships, the heterogeneous PinSageGCN method is used for training, and the node features of each node under the four edge relationships are output. The various features of the nodes are fused and spliced, and the fully connected layer is used to obtain the feature scores of the candidate entity nodes. Then, the similarity scores are combined for joint learning to obtain the joint disambiguation score of each candidate entity node. The nodes are sorted according to the joint disambiguation score, and the PinSageGCN entity disambiguation model is trained with annotation supervision to finally complete the entity disambiguation.

[0086] In this embodiment, an example of a short text to be disambiguated is given, as shown below:

[0087] {"text_id":"139","text":"Time will never wait for people, but people will always chase time.","mention_data":[{"kb_id":"277073","mention":"Time","offset":"0"}]}

[0088] The above data example is a short text to be disambiguated, and mention_data contains all entity references m i , m i ∈M, followed by kb_id, which is the correct corresponding entity in the knowledge base.

[0089] Here is an example of a knowledge base entity data entry, as shown below:

[0090] {"alias":["Victory"],"subject_id":"10001","subject":"Victory","type":["Thing"],"data":[{"predicate":"Abstract","object":"League of Legends Victory series skin is one of the commemorative limited series skins produced by Riot Games. The commemorative limited series skins produced by Riot Games also include the League of Legends Champion Series skins, the MSI Mid-Season Championship Conqueror Series and the League of Legends World Championship Champion Series skins. At the end of each season, Riot Games will produce a Victory series skin as a season reward to recognize those who have fought hard in the qualifying matches. Gold-ranked players."}, {"predicate":"Producer","object":"RiotGames"}, {"predicate":"Foreign name","object":"Victorious"}, {"predicate":"Source","object":"League of Legends"}, {"predicate":"Chinese name","object":"Victory"}, {"predicate":"Attributes","object":"Virtual"}, {"predicate":"Description","object":"Limited skins from the Victory series of the game League of Legends"}]}.

[0091] Where subject is the entity name, subject_id indicates the location of the entity in the knowledge base, alias is the alias, and the content after data is the description of the entity.

[0092] like Figure 3 The figure shows a schematic diagram of the candidate entity recall and labeling process of the present invention. First, the candidate entity description text is checked. If the content is incomplete, it is eliminated. The description content of the remaining entities in the knowledge base is extracted and spliced using the dictionary key-value pair method as the entity description text, also called the candidate entity description text C; secondly, the repetitiveness of the entity name and alias in the knowledge base is integrated to ensure that different expressions of the same entity in short texts can also correctly recall the candidate entity.

[0093] The mention_data in the dataset is filtered by empty judgment to remove training data with incorrect format and no candidate entities. The entity number kb_id is processed again to remove entities that have no mapping relationship with the entity set to ensure the validity of the correct entity. A dictionary is constructed by combining the short text, a string of a correct entity and its number. Then, the entity reference string is matched with the knowledge base entity name through character matching to obtain several candidate entities, kb_id and description text C. iFinally, several recalled candidate entities are labeled according to kb_id. Candidate entities with consistent kb_id are marked as 1, indicating correctly disambiguated candidate entities, and the rest are marked as 0, indicating incorrect candidate entities.

[0094] In this embodiment, if Figure 4 As shown, given an entity reference m (m∈M) and its context Z, for each candidate entity e∈E m , concatenate Z+M (i.e. a short text, Z represents all context words, M represents all entity references in the short text), candidate entity description C, and two special tags [CLS] and [SEP] from the BERT vocabulary as the input sequence. [CLS] indicates the beginning of the sequence, and [SEP] separates different segments of the input. After being concatenated into a string, it is sent to BERT for word segmentation and encoded into a vector T n On this basis, the BiLSTM model is superimposed for feature extraction, focusing on learning the semantic order of entity reference context and candidate entity description text, and obtaining the BiLSTM output vector p n .

[0095] p n =BiLSTM(T n )=BiLSTM(BERT([CLS]text[SEP]C mi [SEP]))

[0096] Where text represents a short text (Z+M), C mi Represents a candidate entity description text of an entity reference m, and calculates the sum of text and C mi The similarity (matching degree) between them is calculated, and the feature vector is generated and overfitting is reduced through the fully connected layer and Dropout. Sigmoid is used for binary classification to generate the final context similarity S. txt (m, e), as shown in the following formula.

[0097] S ctxt (m,e)=sigmod(Dropout(Dense(p n )))

[0098] During training, supervised learning is performed using the annotations described above. After training, the similarity calculation model is optimized. Finally, the trained model is used to output a score for all candidate entities referred to by each entity. This score is used to sort the candidate entities.

[0099] In this embodiment, this part converts the short text into a graph in combination with candidate entities, which includes three types of nodes and four types of relationship edges. Taking the short text to be disambiguated as "Nezha's box office continues to rise because Taiyi Zhenren and Shen Gongbao are the funny ones," as an example, it is converted into a candidate entity ranking graph. The sample figure is shown in Figure 5.

[0100] Among them, "Nezha," "Taiyi Zhenren," and "Shen Gongbao" are entity references, and the remaining words are context. Each entity reference is connected to its corresponding candidate entity. The purpose of constructing this candidate entity ranking graph is to fully establish the connection between the three nodes of entity reference, candidate entity, and context words, so that the neural network can fully extract their inherent characteristics, thereby ranking the candidate entities and completing the entity disambiguation task.

[0101] Table 1 Description table of candidate entity ranking graph construction

[0102]

[0103]

[0104] For a short text, it contains a context consisting of multiple words. i ∈Z, multiple entities refer to m i ∈M, each entity reference has a set of candidate entities C m Now we construct an undirected graph with three types of nodes and four types of relationship edges. The meaning of each node and edge is shown in Table 1. The specific contents are as follows:

[0105] The node set is Z∪M∪C1∪C2∪...∪C k , where k represents the number of all candidate entities referred to by all entities in a text.

[0106] Relationship Edge:

[0107] 1. The nodes in the text that have a relationship with the results obtained through dependency syntactic analysis, that is, between an entity reference and a context, or between two entity references that have a relationship, or between two contexts. It can be expressed as: <m i ,m j > or <m i ,z j > or <z i ,z j >

[0108] 2. The relationship between the entity referent node and the corresponding candidate entity node. It is expressed as <m i ,e i >,e i ∈C mi ;

[0109] 3. The relationship between candidate entities referred to by different entities (considering co-occurrence information). It is expressed as <e i ,e j >,e i ∈C mi , e j ∈C mj , i≠j;

[0110] 4. Context words of candidate entities and short texts. Represented as <e i ,z j >,e i ∈C mi .

[0111] According to the above description, the candidate entity ranking graph G is composed of three types of nodes and four types of relationship edges. The specific graph construction process is as follows: Figure 6 shown.

[0112] First, we perform dependency syntactic analysis on the short text to be disambiguated. While segmenting the short text, we correct the segmentation results based on the entity references obtained from the Chinese short text data to avoid separating the entity references. The segmentation results should include both the entity references and the context words. We then perform dependency syntactic analysis on the segmentation results using the LTP natural language processing toolkit. The results of the dependency syntactic analysis include the indices of the words that each word depends on in the sentence. This information is subsequently used to determine the connectivity of each node in the graph data, which is a candidate entity ranking graph G = (V, E), where V represents the node set and E represents the edge set.

[0113] After that, the entity reference, context words, and candidate entities of each entity reference obtained from the knowledge base in the previous candidate entity recall module are used as nodes. The node vectors of these three types of nodes are initialized using the pre-trained Word2vec model to convert the context words z i ∈Z,m i ∈M includes context words and entity references, and the conversion word vector is recorded as w m ,w z .

[0114] w m =Word2vec(m)

[0115] w z =Word2vec(z)

[0116] When converting the candidate entity description text into a vector, the averaging method is used, that is, the vectors of all words in the candidate entity description text are summed, and then the average of all word vectors is calculated. Assume that the candidate entity description text consists of j words and is converted into a vector representation using the above formula, denoted as w1, w2, ..., w j, that is, each word in the candidate entity description text has a corresponding word vector, and the vector representations of all words are summed and averaged to obtain a candidate entity vector representation result w c , can be calculated by the following formula:

[0117]

[0118] After obtaining the entity reference, context words and node vectors of candidate entities, they are uniformly represented as node features of the node set V of the candidate entity ranking graph G That is the following formula:

[0119]

[0120] According to the dependency syntactic analysis results and the correspondence between entity references and candidate entities, the four edge relationships of the candidate entity ranking graph G mentioned above are established.

[0121] In this embodiment, after obtaining the candidate entity ranking graph G, the node features are also obtained. in and Represents the initial feature vectors of nodes i and j.

[0122] The adjacency matrix A is obtained based on G. The adjacency matrix A represents the connection relationship between the nodes in the candidate entity ranking G in matrix form. The adjacency matrix A is composed of 0 and 1, 1 indicates that the two nodes are connected, and 0 indicates that the two nodes are not connected. The connection status in the node set V is determined by traversing the edge set E. There are n nodes in G, that is, A is an n×n matrix represented as follows:

[0123]

[0124] Where i and j represent the two nodes in the i-th row and j-th column of the adjacency matrix A. After obtaining the adjacency matrix A, the cosine similarity of each pair of connected nodes is calculated and normalized. This is used as the weight representation of the edge relationship and is used in the adjacency matrix to represent the strength of the relationship between nodes. The formula for calculating cosine similarity is as follows:

[0125]

[0126] Where h i and h j Represents the feature vectors of nodes i and j (using the initial feature vector and Calculate similarity), w ij Represents the cosine similarity calculation result of node i and node j, and replaces the calculation result with A of the adjacency matrix ij At, the weight of the edge is represented by A' ijThis operation is performed on all connected nodes to obtain the weighted adjacency A'.

[0127] Sort node features of graph G according to candidate entities And the weighted adjacency matrix A is input into the PinSageGCN entity disambiguation model, the model architecture is as follows Figure 7 shown.

[0128] After obtaining the node features of the candidate entity ranking graph data And the corresponding weighted adjacency matrix A', the above-constructed graph neural network model is used for encoding.

[0129] During the encoding training process of the candidate entity ranking graph, PinSageGCN fully considers the importance of neighbor nodes in the graph and obtains the neighbor nodes of the node according to the obtained adjacency matrix A'. For each node i, j, it represents a neighbor node. The set of neighbor nodes is N(i), and the formula is as follows:

[0130] N(i)={j|A' ij ≠0}

[0131] Thus, the information of the relationship between nodes in the graph structure and the semantic information in the node feature vector are simultaneously used to calculate the node neighbor message distribution and aggregation vector n i ; For each node i, the set of neighbor nodes is N(i).

[0132]

[0133] Among them, Dense is the fully connected layer network, and γ() represents aggregation. In addition, the feature vector of node i is also updated in the fully connected layer, using s i The feature vector representing the node self-update is shown below.

[0134]

[0135] Aggregate neighbor information n i and the node's own encoding i Add and regularize to get the new node vector representation, which is the iteration of the node during the training process.

[0136]

[0137] In summary, PinSageGCN updates node features by adding neighbor aggregation information to the node's own code and regularizing it. After one layer of PinSageGCN, in Represents the intermediate result of node update features.

[0138] In the process of constructing the candidate entity ranking graph, the model fully completes the information interaction of each connected node by connecting four relationships: between entity reference and context, between entity reference and its corresponding candidate entity, between candidate entities with different entity references, and between entity reference and context.

[0139] The PinSageGCN model is used to obtain the features of each node after multiple layers of iteration. Next, four PinSageGCN models (heterogeneous PinSageGCN models) will be used based on four edge relationships to complete the entity disambiguation task.

[0140] In this embodiment, in the present invention, heterogeneous PinSageGCN is used to complete the entity disambiguation task, such as Figure 8 As shown, according to four different types of edges, four PinSageGCN models are used to train the node feature vectors of the four edge relationships respectively, and each outputs the node feature vector H p , where p represents the feature vector of the node trained under the pth edge relationship, p∈[1,2,34]. The model outputs the feature vectors of all n nodes in is the vector representation of node i under the relationship edge of p, L represents the number of layers of the graph neural network in the PinSageGCN model, and then the vectors of each node under various edge relationships are spliced together, that is, for each Further use the fully connected layer and activation function to convert its nonlinear transformation into the global feature score of the node's graph neural network

[0141]

[0142] Then, the similarity scores of the entity reference and candidate entities calculated by the similarity model are fused, and joint learning is performed through the fully connected layer to obtain the final disambiguation score. All candidate entities are sorted according to the score to obtain the final disambiguation result. The formula is as follows.

[0143]

[0144] Model based on The result is used to calculate the loss and annotate y i For supervised learning, the loss function is as follows.

[0145]

[0146] where y i ∈{1,-1} indicates whether the node is the correct target entity corresponding to the entity reference, k is the number of candidate entities, when node i is a candidate entity node, Ki =1, if it is the other two types of nodes (entity reference node and context node), then K i =0 means not participating in the loss calculation.

[0147] In this embodiment, if Figure 9 FIG. 1 shows a schematic diagram of an entity disambiguation operating environment deployment of the present invention, wherein a terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in each embodiment of the aforementioned graph convolutional neural network-based Chinese short text joint entity disambiguation method are implemented.

[0148] Deployed in this environment, the computer program is mainly divided into three important stages, namely the data preparation stage, the model training stage, and the model use stage. According to the stage, the computer program can be further divided into a data preprocessing module, an entity disambiguation similarity calculation module, a candidate entity ranking graph construction module, a PinSageGCN model construction module, and a candidate entity ranking module. The specific functions of each module are detailed in the above description. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 9 This is merely an example of a terminal device and does not limit the terminal device. The terminal device may include more or fewer components than shown in the figure.

[0149] Example 2

[0150] The present invention also provides a Chinese short text joint entity disambiguation system based on a graph convolutional neural network, the system is used to implement the method described in any one of the first embodiments, the system comprising: an acquisition module, a training module, an encoding module, an update module, and a sorting module;

[0151] The acquisition module is used to obtain the original short text containing entity references and obtain the candidate entity set of each entity reference based on the existing knowledge base to generate a training dataset;

[0152] A training module, configured to train the training dataset using a BERT+BiLSTM model to obtain a similarity score between the entity reference and each candidate entity in the candidate entity set;

[0153] An encoding module is used to perform word segmentation and dependency syntactic analysis on the original short text using LTP, convert the analysis results into graph data, add candidate entities as nodes to the graph data, obtain an adjacency matrix of the complete graph data, and then use word2vec to encode each node to obtain a feature matrix of the complete graph data;

[0154] An update module is used to calculate the cosine similarity between each connected node, incorporate it into the adjacency matrix as the weight of the edge relationship, build a PinSageGCN entity disambiguation model, input the feature matrix and the adjacency matrix with weights into the PinSageGCN entity disambiguation model to start training, and update the node features;

[0155] The sorting module is used to obtain the features of each node after multi-layer iteration through the PinSageGCN entity disambiguation model, and complete the entity disambiguation task based on four edge relationships using four PinSageGCN entity disambiguation models, namely heterogeneous PinSageGCN models.

[0156] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A Chinese short text joint entity disambiguation method based on graph convolutional neural network, characterized by: The method comprises: Obtain the original short text containing entity references, and obtain the candidate entity set of each entity reference based on the existing knowledge base to generate a training dataset; The training dataset is trained using a BERT+BiLSTM model to obtain a similarity score between the entity reference and each candidate entity in the candidate entity set; Use LTP to perform word segmentation and dependency syntactic analysis on the original short text, convert the analysis results into graph data, and add the candidate entities as nodes to the graph data to obtain the adjacency matrix of the complete graph data, and then use word2vec to encode each node to obtain the feature matrix of the complete graph data; Calculate the cosine similarity between each connected node and incorporate it into the adjacency matrix as the weight of the edge relationship to build a PinSageGCN entity disambiguation model. Input the feature matrix and the adjacency matrix with weights into the PinSageGCN entity disambiguation model to start training and update the node features. The PinSageGCN entity disambiguation model is used to obtain the features of each node after multiple layers of iteration. Based on the four edge relationships, four PinSageGCN entity disambiguation models, namely heterogeneous PinSageGCN models, are used to complete the entity disambiguation task.

2. The method according to claim 1, characterized in that The method of training the training dataset using the BERT+BiLSTM model to obtain a similarity score between the entity reference and each candidate entity in the candidate entity set includes: Given an entity reference m and context Z, for each candidate entity e∈E m , concatenate Z+M, candidate entity description C, and two special tags [CLS] and [SEP] from the BERT vocabulary as the input sequence; [CLS] indicates the beginning of the sequence, and [SEP] separates the input of different segments; after being concatenated into a string, it is sent to BERT for word segmentation and encoded into a vector T n ; Superimpose the BiLSTM model for feature extraction, learn the semantic order of entity reference context and candidate entity description text, and obtain the BiLSTM output vector p n : p n =BiLSTM(T n )=BiLSTM(BERT([CLS]text[SEP]C mi [SEP])) Among them, text represents a short text Z+M, C mi A candidate entity description text representing an entity reference m; Calculate text and C mi The similarity between them is used to generate feature vectors and pass them through the fully connected layer, Dropout, and sigmoid for binary classification to generate the final context similarity S txt (m, e): S ctxt (m,e)=sigmod(Dropout(Dense(p n )))。 3. The method according to claim 1, characterized in that The graph data is: candidate entity ranking graph G=(V, E), where V represents a node set and E represents an edge set.

4. The method according to claim 3, characterized in that Methods for adding candidate entities as nodes to the graph data and obtaining the adjacency matrix of the complete graph data include: Where i and j represent the two nodes in the i-th row and j-th column of the adjacency matrix A.

5. The method according to claim 4, characterized in that Methods for using word2vec to encode each node and obtain the feature matrix of the complete graph data include: The entity reference, context words, and candidate entities of each entity reference obtained from the knowledge base are taken as nodes, and the node vectors are initialized for the nodes. The pre-trained Word2vec model is used to convert the context words z i ∈Z,m i ∈M includes context words and entity references, and the conversion word vector is recorded as w m ,w z : Assume that the candidate entity description text consists of j words and is converted into a vector representation, denoted as w1,w2,...,w j , that is, each word in the candidate entity description text has a corresponding word vector, and then the vector representations of all words are summed and averaged to obtain a candidate entity vector representation result w c : The entity reference, context words and node vectors of candidate entities are uniformly represented as node features of the node set V of the candidate entity ranking graph G in Represents the initial eigenvector of a node.

6. The method according to claim 1, wherein The method for calculating the cosine similarity between each connected node includes: Where h i and h j Represents the feature vectors of node i and node j.

7. The method according to claim 6, characterized in that The feature matrix and the adjacency matrix with weights are input into PinSageGCN to start training. The method for updating node features includes: Obtain the neighbor nodes of the node according to the obtained adjacency matrix A'. For each node i, j represents a certain neighbor node, and the neighbor node set is N(i): N(i)={j|A′ ij ≠0}; At the same time, the information about the relationship between nodes in the graph structure and the semantic information in the node feature vector are used to calculate the node neighbor message distribution and aggregation vector n i ;For each node i, the set of neighboring nodes is N(i); Among them, Dense is a fully connected layer network, γ() represents aggregation; The feature vector of node i is also updated in the fully connected layer: Among them, s i The feature vector representing the node self-update; Aggregate neighbor information n i and the node's own encoding i Add and regularize to get the new node vector representation, which is the iteration of the node during the training process: in, Represents the intermediate result of node update features.

8. The method according to claim 1, characterized in that The PinSageGCN entity disambiguation model is used to obtain the features of each node after multiple layers of iteration. Based on four edge relationships, four PinSageGCN entity disambiguation models, namely heterogeneous PinSageGCN models, are used to complete the entity disambiguation task. The following methods are used: Based on the four types of edge relationships, the heterogeneous PinSageGCN method is used for training, and the node features of each node under the four edge relationships are output. The various features of the nodes are fused and spliced, and the fully connected layer is used to obtain the feature scores of the candidate entity nodes. Then, combined with the similarity scores, joint learning is performed to obtain the joint disambiguation score of each candidate entity node. The nodes are sorted according to the joint disambiguation score, and the PinSageGCN entity disambiguation model training is supervised by annotations to finally complete entity disambiguation.

9. A Chinese short text joint entity disambiguation system based on graph convolutional neural network, the system is used to implement the method of any one of claims 1-8, characterized in that: The system includes: an acquisition module, a training module, an encoding module, an updating module and a sorting module; The acquisition module is used to acquire the original short text containing entity references, and obtain a candidate entity set for each entity reference based on the existing knowledge base to generate a training data set; The training module is used to train the training data set using the BERT+BiLSTM model to obtain a similarity score between the entity reference and each candidate entity in the candidate entity set; The encoding module is used to perform word segmentation and dependency syntax analysis on the original short text using LTP, convert the analysis results into graph data, and add the candidate entities as nodes to the graph data to obtain an adjacency matrix of the complete graph data, and then use word2vec to encode each node to obtain a feature matrix of the complete graph data; The update module is used to calculate the cosine similarity between each connected node, incorporate it into the adjacency matrix as the weight of the edge relationship, build a PinSageGCN entity disambiguation model, input the feature matrix and the adjacency matrix with weights into the PinSageGCN entity disambiguation model to start training, and update the node features; The sorting module is used to obtain the features of each node after multi-layer iteration through the PinSageGCN entity disambiguation model, and complete the entity disambiguation task based on four edge relationships using four PinSageGCN entity disambiguation models, namely heterogeneous PinSageGCN models.

Citation Information

Patent Citations

  • Cooperative disambiguation method based on deep semantic neighbor and multivariate entity association

    CN112883199A

  • Short text entity disambiguation method based on multi-task learning

    CN115081445A

  • Entity linking method and related equipment

    CN115238080A

  • Short text named entity disambiguation method based on multi-information fusion

    CN117709349A

  • NLP-based entity recognition and disambiguation

    WO2009052277A1

Cited By

  • Author name disambiguation method and related device

    CN120725014A