Entity linking using graph neural networks
The three-party graph is generated through the graph neural network and entity linking is solved, and the ambiguity and inconsistency in entity links is achieved, efficient and accurate entity matching is achieved, and it is suitable for processing entity names in a large number of information sources.
Patent Information
- Application Number
- CN202280099988.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art has ambiguity and inconsistency problems in entity links. String similarity matching technology cannot accurately handle the potential relationship of entity names, resulting in wrong entity matching and lack of reliable ground truth data sets, making it difficult to train machine learning models.
The graph neural network is used for entity linking, and unknown and known names are obtained from the information source and database through the extraction module. After tokenization, three-party graphs are generated. The graph neural network model is used for entity matching. The recommended module is accurately allocated based on similarity scores.
It realizes more accurate and efficient entity links, can automatically process a large number of information sources, reduce manual intervention, and improves the accuracy and efficiency of entity matching.
Smart Images

Figure CN120266128A_ABST
Abstract
Description
Technical Field
[0001] At least some aspects of the present disclosure relate to natural language processing, such as cross-party entity linking using graph neural networks. Background Art
[0002] In the field of natural language processing (NLP), entity linking (sometimes referred to as named entity linking or named entity matching) generally involves determining which words or strings of words recited in text refer to specific entities. Thus, entity linking can involve assigning a unique identifier to a word or string of words. In many cases, the unique identifier can be an individual or an organization, such as a company, foundation, charity, or government organization.
[0003] Entity linking can be a valuable tool for associating information with a specific individual or organization. In some aspects, this is because entity linking can enable a company to leverage the vast amount of information accessible via the Internet to evaluate an entity. For example, various private and public information sources (such as news articles, wikis, social media, and other databases and publications) can contain information related to an entity. This information may be relevant to evaluating the entity's reputation, assessing the risk of doing business with the entity, evaluating the entity's financial performance, and so on. However, since an evaluating company may be concerned with potentially millions of information sources and potentially thousands or even millions of entities, it may be technically difficult to manually identify the information that may be relevant. Thus, NLP and entity linking can be used as an automated way to identify specific entities mentioned in these information sources.
[0004] However, there are several technical challenges associated with entity linking. As an example, names used to refer to specific entities often are ambiguous and inconsistent. A news article about an entity with the legal name "United Airlines, Inc." may instead recite the name "United" in the text or even the title of the article. Thus, for example, a computer-implemented process may mistake the name "United" for other entities, such as "United Health Care" and "United Technology Corp.". As another example, it is difficult to create a ground truth or labeled dataset for reliably training a machine learning model to perform entity linking at scale.
[0005] Some methods of entity linking employ string similarity matching techniques. These string similarity matching techniques generally involve generating a string similarity score (e.g., Hamming distance, Jaro-Winkler distance) by comparing strings in the name of an entity (e.g., an unknown name) extracted from an information source with strings in the name of an entity (e.g., a known name) stored in a database. Then, the unknown name is assigned to the known name with the highest string similarity score.
[0006] However, when used for entity linking, string similarity matching techniques can be problematic. In some respects, this is because string similarity matching techniques are sensitive to data quality and generally cannot take into account potential relationships between entity names. For example, using a metric to calculate string similarity known to those skilled in the art, the unknown entity name "China Eastern Airlines Yunnan Company" may have a string similarity score of 0.70 with "Air China", a string similarity score of 0.75 with "China Eastern Air", and a string similarity score of 0.76 with "China Yunnan Hotel Corp.". Thus, the string similarity matching technique may incorrectly determine that "China Eastern Airlines Yunnan Company" refers to "China Yunnan Hotel Corp.".
[0007] Accordingly, there is a need for systems and methods that can perform entity linking automatically, accurately, and efficiently. The present disclosure provides a solution for performing entity linking by employing a graph neural network. SUMMARY OF THE INVENTION
[0008] In one aspect, the present disclosure provides a computer-implemented method for entity linking. According to the method, an extraction module retrieves known names and unknown names. The known names are entity names stored in a database, and the unknown names are extracted from an information source. A tokenization module tokenizes the known names and the unknown names. A graph generation module identifies candidates from the known names. The graph generation module generates a tripartite graph that includes a first layer of nodes corresponding to the unknown names, a second layer of nodes corresponding to words of the unknown names and the candidates, and a third layer of nodes corresponding to the candidates. A recommendation module applies the tripartite graph to a graph neural network model. The recommendation module assigns the unknown name to one of the known names based on applying the tripartite graph to the graph neural network model.
[0009] On the one hand, the present disclosure provides an entity linking system. The entity linking system may include an extraction module, a tokenization module, a graph generation module, a graph neural network, and a recommendation module. The extraction module may be configured to extract unknown names from an information source and extract known names from a database. The tokenization module may be configured to tokenize the unknown names and tokenize each of the known names. The graph generation module may be configured to identify candidates from the known names and generate a tripartite graph based on the unknown names and the candidates. The tripartite graph may include a first layer of nodes corresponding to the unknown names, a second layer of nodes corresponding to the words of the unknown names and the candidates, and a third layer of nodes corresponding to the candidates. The graph neural network may be configured to generate an unknown name embedding and a candidate embedding based on the tripartite graph. The unknown name embedding corresponds to the first layer of nodes, and the candidate embedding corresponds to the third layer of nodes. The recommendation module may be configured to determine a similarity score between the unknown name embedding and the candidate name embedding. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The various features and advantages of the embodiments described herein can be understood as follows in conjunction with the accompanying drawings according to the following description:
[0011] Figure 1 is a diagram showing a system for entity linking using a graph neural network according to at least one aspect of the present disclosure.
[0012] Figures 2A - 2C is a tripartite graph according to several aspects of the present disclosure that can be applied to a graph neural network for entity linking.
[0013] Figure 3 is a logical flow chart of a method for entity linking according to at least one aspect of the present disclosure.
[0014] Figure 4 is a logical flow chart of a method for assigning an unknown name to one of a set of known names by applying a tripartite graph to a graph neural network model according to at least one aspect of the present disclosure.
[0015] Figure 5 is a logical flow chart of a method for training a graph neural network model by specifying positive samples, hard negative samples, and random negative samples according to at least one aspect of the present disclosure.
[0016] Figure 6 shows a block diagram of a computer device according to at least one aspect of the present disclosure.
[0017] Figure 7 shows a block diagram of a system including a host according to at least one aspect of the present disclosure.
[0018] Throughout several views, corresponding reference numerals indicate corresponding parts. The examples set forth herein illustrate various aspects of the present disclosure in one form, and such examples should not be construed as limiting the scope of the present disclosure in any way. Detailed Description
[0019] The applicant of the present application owns the following U.S. provisional patent applications currently filed therewith, the disclosures of which are incorporated herein by reference in their entireties:
[0020] · U.S. Provisional Patent Application Docket No. 220265P (6063US01 / 220265P) entitled ENTITY LINKING USING SUBGRAPH MATCHING.
[0021] Before explaining various forms of entity linking using graph neural networks, it should be noted that the illustrative forms disclosed herein are not limited in application or use to the details of the construction and arrangement of the components shown in the figures and the description. The illustrative forms can be implemented or incorporated in other forms, variations, and modifications and can be practiced or carried out in various ways. Additionally, unless otherwise indicated, the terms and expressions utilized herein are chosen for the purpose of convenience to the reader in describing the illustrative forms and not for the purpose of limiting them.
[0022] As used herein, the term "computing device" or "computer device" can refer to one or more electronic devices configured to communicate directly or indirectly with one or more networks or over one or more networks. The computing device can be a mobile device, a desktop computer, etc. Additionally, the term "computer" can refer to any computing device that includes the necessary components for sending, receiving, processing, and / or outputting data and typically includes a display device, a processor, a memory, an input device, a network interface, etc.
[0023] As used herein, the term "server" can include one or more computing devices, which can be individual stand-alone machines located at the same or different locations, can be owned or operated by the same or different entities, and can further be one or more clusters of distributed computers or "virtual" machines housed within a data center. Those skilled in the art should understand and appreciate that the functions performed by a single "server" may be spread across multiple different computing devices for various reasons. As used herein, "server" is intended to refer to all such scenarios and should not be construed as or limited to a particular configuration. The term "server" may also refer to or include one or more processors or computers, storage devices, or similar computer arrangements that operate or facilitate communication and processing among multiple parties in a network environment such as the Internet, but it will be understood that communication may be facilitated through one or more public or private network environments and various other arrangements are possible. Additionally, multiple computers (e.g., servers or other computerized devices) that communicate directly or indirectly in a network environment can constitute a "system". As used herein, a reference to a "server" or "processor" may refer to a previously described server and / or processor, a different server and / or processor, and / or a combination of servers and / or processors that are described as performing steps or functions.
[0024] As used herein, the term "system" can refer to one or more computing devices or a combination of computing devices (e.g., processors, servers, client devices, software applications, modules, such components, etc.). For example, a system can include multiple computing devices that include software applications, where the multiple computing devices are connected via a network.
[0025] As used herein, the term "module" can refer to a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution.
[0026] As used herein, the term "entity" can refer to or include an individual, a company, a business-related organization, a non-profit organization, a government organization, a charity, an educational institution, or any other type of individual, group of individuals, or organization.
[0027] As used herein, the term "word" can refer to a string. For example, a word can refer to a string that is not separated by spaces. The string can include one or more characters. The one or more characters can include letters, numbers, and / or symbols.
[0028] As used herein, the term "name" can refer to a word or a string of words. For example, a name can refer to a word or a string of words that identify an entity.
[0029] As used herein, the term "known", when used to refer to a name and / or an entity (e.g., a known name, a known entity name, a known entity), may mean that the name and / or the entity has been assigned or otherwise associated with a specific unique entity. For example, a database storing a set of known names may be used to assign each known name in the set of known names to refer to a specific entity.
[0030] As used herein, the term "unknown", when used to refer to a name and / or an entity (e.g., an unknown name, an unknown entity name, an unknown entity), may mean any name and / or entity mentioned or otherwise described in an information source. An unknown name, an unknown entity name, and / or an unknown entity may or may not have been assigned to a known name or a known entity. As an example, a name that is the target of an entity linking process for assigning a name to a specific known entity and / or a specific known name may be referred to as an unknown name. As another example, a name may have been extracted from an information source and assigned to a known name through an entity linking process. Even if the name has been assigned to a known name, the name may still be referred to as an unknown name. Similarly, an entity described by a name may still be referred to as an unknown entity.
[0031] As used herein, the term "tokenization" may refer to the process of classifying and / or separating a string into one or more segments. For example, a name may be tokenized based on words (e.g., word tokenization), sub-words (e.g., sub-word or n-gram tokenization), characters (e.g., character tokenization), etc.
[0032] Entity linking is the basis for any organization to effectively utilize data outside the organization. Given the heterogeneous nature of the data and the environmental quality of the external data, it is quite technically challenging for an organization to utilize external data. According to the present disclosure, machine learning is used to effectively utilize external data. On the one hand, as described in more detail hereinafter, machine learning techniques are used to extract structural information from external entity names based on various factors such as words included in the external entity names. On the one hand, the machine learning techniques include a machine learning model to learn when a matching link should or should not be in an embedding space. The machine learning techniques according to the present disclosure provide a model for learning from positive link samples and negative link samples. The following description provides a technical solution for an organization to utilize external data.
[0033] In various aspects, the present disclosure provides solutions that can perform entity linking by using artificial neural networks (e.g., graph neural networks) for processing data that can be represented as a graph outside an organization. For example, a convolutional neural network can be applied in this context to a graph structured as node layers. Performing entity linking using a graph neural network can provide various technical benefits. For example, the systems and methods disclosed herein can allow a computer to perform entity linking more accurately and efficiently by: (i) generating a tripartite that includes nodes corresponding to an unknown name, a known name (e.g., a candidate), and words of the unknown name and the known name; and (ii) assigning the unknown name to one of the known names by applying the tripartite graph to a graph neural network model.
[0034] As another example, the systems and methods disclosed herein can allow automatic identification of a particular entity cited in one or more of potentially millions of various private and / or public information sources accessible via the Internet, thereby performing entity linking at a scale that cannot be practically performed in human thinking.
[0035] As yet another example, the systems and methods described herein can perform entity linking in an unconventional manner by: (i) tokenizing the unknown name and tokenizing the known names; (ii) identifying one or more candidates from the known names; (iii) generating a tripartite graph that includes a first layer of nodes corresponding to the unknown name, a second layer of nodes corresponding to words of the unknown name and the one or more candidates, and one or more third layer of nodes corresponding to the one or more candidates; and (iv) assigning the unknown name to one of the known names by applying the tripartite graph to a graph neural network (e.g., a graph convolutional network). Additionally, the systems and methods described herein can perform entity linking in an unconventional manner by supervising a graph neural network model (e.g., a graph convolutional network) based on positive samples, hard negative samples, and / or random negative samples.
[0036] Figure 1FIG. 100 shows an entity linking system 130 in accordance with at least one aspect of the present disclosure. The entity linking system 130 may include various modules such as an extraction module 132, a tokenization module 134, a graph generation module 136, a graph neural network 138 (GNN), a first natural language processing module 140 (NLP1), a second natural language processing module 142 (NLP2), a recommendation module 144, and / or a training module 146. Although the modules of the entity linking system 130 are described below as performing various functions separately, any module may be configured to perform any combination of the functions described herein. Similarly, multiple modules may be combined into a single module to perform any combination of the functions described herein, and / or a single module may be split into multiple sub-modules, where each of the sub-modules performs any of the functions described herein.
[0037] The entity linking system 130 is configured to access one or more information sources 1101, 1102, 1103,..110 n (collectively referred to as information sources 110) via the network 120 or otherwise communicate with the one or more information sources. The network 120 may include any kind of wired network, remote wireless network, and / or short-range wireless network. For example, the network 120 may include an internal network, a local area network (LAN), Wi-Fi, a cellular network, a private network, the Internet, a cloud computing network, and / or a combination of these or other types of networks. The information sources 110 may include any type of information source and combination of information sources, including text-based data accessible via the network 120. For example, the information sources 110 may include various private and public information sources such as news articles, Wikipedia, social media, and / or other databases and publications accessible via the Internet.
[0038] In Figure 1 a non-limiting aspect, the entity linking system 130 is also configured to access a known entity database 150 via the network 120 or otherwise communicate with the known entity database. In other aspects, the known entity database 150 may be included as part of the entity linking system 130 (e.g., stored on the same server or combination of servers as the entity linking system 130). The known entity database 150 may include data related to a plurality of known entities. For example, the known entity database 150 may include a list of names of known entities. As another example, the known entity database 150 may include profiles of known entities, which include the names of the known entities and other information related to the known entities, such as resume information, financial information, industry classification, parent company information, subsidiary company information, geographical information, etc. The names of the known entities stored in the known entity database 150 are sometimes referred to herein as "known names".
[0039] The extraction module 132 of the entity linking system 130 can be configured to extract information from the information source 110 and / or the known entity database 150. For example, on the one hand, the extraction module 132 can be configured to detect the text included in the information source 110 and extract entity names from the text. The extraction module 132 can employ various techniques to detect and extract entity names, such as rule-based named entity recognition (NER) techniques (e.g., techniques employed by the General Architecture for Text Engineering (GATE) and rule-based NER, known as DrNER, etc.) and / or machine learning-based NER techniques (e.g., techniques employed by the OpenNLP named entity recognizer and name finder, the free and open-source library for natural language processing in Python spaCy, and the named entity recognizer SemiNER, etc.). In some aspects, the extraction module 132 can be configured to detect the occurrence of entity names within the information source 110 without identifying the entity names or assigning the entity names to specific known entities. Thus, the entity names detected and extracted by the extraction module 132 from the information source 110 are sometimes referred to as "unknown names" herein.
[0040] In some aspects, the extraction module 132 can be configured to extract a set of known names from the known entity database 150. For example, on the one hand, the extraction module 132 can be configured to retrieve a list of known names from the known entity database 150. On the other hand, in the case where the known entity database 150 includes profiles of known entities, the extraction module 132 can be configured to extract and / or retrieve a set of known names from the profiles.
[0041] The tokenization module 134 of the entity linking system 130 can be configured to tokenize the known names and / or unknown names retrieved by the extraction module 132. In some aspects, the known names and / or unknown names can be tokenized according to the words included in each of the known names and / or unknown names. In other aspects, the known names and / or unknown names can be tokenized based on some other criteria, such as sub-word (e.g., n-gram) tokenization or character tokenization. The tokenization module 134 can employ various NLP tokenization techniques to tokenize the known names and / or unknown names (e.g., white space tokenization, Keras tokenization, natural language toolkit (NLTK) word tokenization, and spaCy tokenizer, etc.). The words generated by tokenizing the unknown names are sometimes referred to as "unknown words" herein. Similarly, the words generated by tokenizing the known names are sometimes referred to as "known words" herein.
[0042] The graph generation module 136 of the entity linking system 130 can be configured to generate a tripartite graph based on the unknown names retrieved by the extraction module 132, the set of known names retrieved by the extraction module 132, and the words (e.g., known words and unknown words) generated by the tokenization module 134. The structure of the tripartite graph is generally configured to be applied to the GNN 138 such that the unknown name can be assigned to one of the known names in the set of known names. In some aspects, the graph generation module 136 can be configured to generate a tripartite graph having a structure similar to the tripartite graphs 200A-C respectively shown in Figures 2A - 2C .
[0043] For example, now mainly referring to Figure 2A and also referring to Figure 1 , the graph generation module 136 can be configured to identify one or more candidates 2281, 2282,... 228 n (collectively referred to as candidates 228) from the set of known names retrieved by the extraction module 132. The candidates 228 are generally known names that the unknown name 222 may refer to. The graph generation module 136 can employ various techniques to identify the candidates 228. On the one hand, the graph generation module 136 can be configured to identify the candidates 228 by comparing the words of the unknown name 222 (e.g., unknown words 2241, 2242,... 224 n ) generated by the tokenization module 134 with the words of the known names (known words) generated by the tokenization module 134. Any known name having at least one known word that is the same or similar to one of the unknown words in the unknown words 224 can be designated as a candidate 228. The graph generation module 136 can determine whether a known word is the same or similar to an unknown word based on a string similarity score (e.g., Hamming distance, jaro - winkler distance, etc.). For example, if the string similarity score of a word pair is not less than 0.7 (such as not less than 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or not less than 0.99), then the graph generation module 136 can determine that the known word is the same or similar to the unknown word. On the other hand, the graph generation module 136 can be configured to identify each known name in the known names extracted from the known entity database 150 as a candidate 228. In this regard, the one or more candidates 228 include all the known names extracted from the known entity database 150. The known words included in the candidates 228 are sometimes referred to herein as "candidate words" (e.g., candidate words 2261, 2262, 2263,... 226 m ).
[0044] Still mainly referring to Figure 2A and also referring to Figure 1, the tripartite graph 200A can be structured to have a first - layer node 210, a second - layer node 212, and a third - layer node 214. The first - layer node 210 corresponds to an unknown name 222 retrieved by the extraction module 132 from the information source 110. The second - layer node 212 corresponds to an unknown word 224 and candidate words 226 generated by the tokenization module 134. The third - layer node 214 corresponds to candidates 228 identified by the graph generation module 136 from the set of known names. In addition, each first - layer node 210 is connected by an edge 223 to each second - layer node in the second - layer nodes 212 corresponding to the unknown word 224. Each third - layer node in the third - layer nodes 214 is connected by an edge 227 to each second - layer node in the second - layer nodes 212 corresponding to the candidate words 226 (which can also be unknown words 224) included in the corresponding candidates 228 to which the third - layer node 214 corresponds. As described above, on the one hand, the candidates 228 identified by the graph generation module 136 can include all known names extracted from the known entity database 150. Thus, in this regard, the candidate words 226 are all known words. Therefore, the tripartite graph 200A can include second - layer nodes 212 corresponding to each of the known words and third - layer nodes 214 corresponding to each of the known names.
[0045] Specifically, some candidate words can be the same as the unknown word, while other candidate words can be different from the unknown word. In the case where one or more candidate words are the same as the unknown word, the tripartite graph 200A includes one second - layer node 212 corresponding to the overlapping candidate / unknown word. For example, one of the second - layer nodes in the second - layer nodes 212 is shown as corresponding to the unknown word n 224 n . However, this second - layer node 212 also corresponds to the candidate words included in candidate 1 2281 and candidate n 228 n . Thus, the second - layer node 212 corresponding to the unknown word n 224 n is not only connected by an edge 223 to the first - layer node 210, but is also connected by an edge 227 to the third - layer nodes 214 corresponding to candidate 1 2281 and candidate n 228 n . The unknown word n 224 n is a word included in each of the unknown name 222, candidate 1 2281, and candidate n 228 n .
[0046] Now mainly referring to Figure 2B and also referring to Figure 1, the tripartite graph 200B is filled with an example unknown name 222, an example unknown word 224, an example candidate word 226, and an example candidate 228. In Figure 2B a non-limiting aspect of, the unknown name 222 is "China Eastern Airlines Yunnan Company". This unknown name may have been extracted by the extraction module 132 from the information source 110 (such as an online news article). The unknown words 224 generated by tokenizing "China Eastern Airlines Yunnan Company" are "China" 2241, "Eastern" 2243, "Airlines" 2244, "Yunnan" 2242, and "Company" 2245. Thus, the tripartite graph 200B includes second-layer nodes 212 corresponding to each of these unknown words 224. In addition, the second-layer nodes 212 corresponding to the unknown words 224 are connected to the first-layer nodes 210 by edges 223.
[0047] Still mainly referring to Figure 2B And also referring to Figure 1 , the tripartite graph 200B includes third-layer nodes 214 corresponding to the candidates "China Eastern Airlines" 2281 and "Yunnan Hotel" 2282. These candidates 228 may have been identified by the graph generation module 136 from a set of known names because tokenizing each of these names generates at least one word that is the same as or similar to the unknown word 224. For example, tokenizing "China Eastern Airlines" can generate the words "China", "Eastern", and "Airlines", so the known name "China Eastern Airlines" can be selected as a candidate because the words "China" and "Eastern" are also generated by tokenizing the unknown name "China Eastern Airlines Yunnan Company".
[0048] Still mainly referring to Figure 2B And also referring to Figure 1 , each candidate word among the candidate words that has not been included as a second-layer node 212 of the unknown word 224 is included as a second-layer node 212 of the candidate word 226. For example, "Hotel" 2262 is a candidate word that is different from one of the unknown words 224. Thus, a second-layer node 212 corresponding to "Hotel" 2262 is added. Similarly, "Airlines" 2261 is a candidate word that is different from one of the unknown words 224. Thus, a second-layer node 212 corresponding to "Airlines" 2261 is added. The third-layer nodes 214 are connected by edges 227 to each of the second-layer nodes 212 corresponding to the words included in the name of the corresponding candidate. For example, the third-layer node 214 corresponding to the candidate "China Eastern Airlines" 2281 is connected by an edge 227 to each of the second-layer nodes 212 corresponding to the words "China", "Eastern", and "Airlines".
[0049] In some aspects, in addition to the unknown words 224 and / or candidate words 226 second-layer nodes 212 selected to be included in the tripartite graph 200B as described above, the skip neighbors of these words (e.g., all one-hop neighbors; all one-hop neighbors and two-hop neighbors; all one-hop neighbors, two-hop neighbors, and three-hop neighbors, etc.) can be included in the tripartite graph 200B as additional candidate words 226 second-layer nodes 212. For example, the skip neighbors of the unknown words 226 and / or candidate words 226 can be identified based on the unknown word embeddings and / or candidate word embeddings.
[0050] In some aspects, in addition to the candidates 228 selected to be included in the tripartite graph 200B as described above, the skip neighbors of the candidates (e.g., all one-hop neighbors; all one-hop neighbors and two-hop neighbors; all one-hop neighbors, two-hop neighbors, and three-hop neighbors, etc.) can be included in the tripartite graph 200B as additional candidates 228 third-layer nodes 214. For example, the skip neighbors of the candidates 226 can be identified based on the candidate embeddings. As used herein, the "skip neighbor" of a node (e.g., the first node) can refer to another node (e.g., the second node) that is directly connected to the node (e.g., the first node) through an edge or indirectly connected through more than one edge. For example, the first node can be directly connected to the second node through an edge. The second node is a one-hop neighbor of the first node. The third node can be directly connected to the second node through another edge but not directly connected to the first node. The third node is a two-hop neighbor of the first node.
[0051] Referring again primarily to Figure 1 And also referring to Figure 2A , the recommendation module 144 of the GNN 138 and / or the entity linking system 130 can be configured to perform entity linking based on the tripartite graph generated by the graph generation module 136. To perform entity linking, the GNN 138 can be configured to generate embeddings corresponding to the nodes of the tripartite graph 200A (e.g., unknown name embeddings corresponding to the first-layer nodes 210, word embeddings corresponding to the second-layer nodes 212, and candidate embeddings corresponding to the third-layer nodes 214). In addition, as will be explained in detail below, the GNN 138 can be trained such that unknown names and candidate names that refer to the same unique entity will have similar representations in the embedding space. For example, the GNN 138 can be any type of GNN, such as a graph convolutional network (GCN).
[0052] The recommendation module 144 can be configured to assign an unknown entity to one of the known entities (e.g., one of the candidates) by comparing the unknown name embedding generated by the GNN 138 with each of the candidate embeddings generated by the GNN 138. For example, the recommendation module 144 can be configured to determine a similarity score between the unknown name embedding and each of the candidate name embeddings. In some aspects, the recommendation module 144 can assign the unknown name to one of the known names based on the unknown embedding / candidate embedding pair with the highest similarity score. In addition to or instead of the above, if the corresponding unknown embedding / candidate embedding pair has a similarity score that meets a predetermined threshold (such as a similarity score not less than 0.7, 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or not less than 0.99), then the recommendation module 144 can assign the unknown name to one of the known names. If the similarity score does not meet the predetermined threshold, the recommendation module 144 can refrain from assigning the unknown name to any of the known names.
[0053] In Figure 1 non-limiting aspects, the recommendation module 144 is shown as being separate from the GNN 138. However, in other aspects, the recommendation module 144 can be included as a layer of the GNN 138. The recommendation module 144 can employ various techniques to determine the similarity score between the unknown name embedding and the candidate embeddings. For example, the recommendation module 144 can employ a trained regression model and / or a trained classification model to determine the similarity score between the unknown name embedding and each of the candidate embeddings.
[0054] Referring again primarily to Figure 2A and also referring to Figure 1 , the tripartite graph 200A can be used to show how the GNN 138 and the recommendation module 144 can perform entity linking based on the tripartite graph. For example, the GNN 138 can be a GCN. In addition, the following functions can represent the embeddings generated by applying the tripartite graph 200A to the GNN 138:
[0055] GCN(“unknown name”);
[0056] GCN(“candidate 1”);
[0057] GCN(“candidate 2”);
[0058] GCN(“candidate n”).
[0059] The following equation represents an example similarity score that can be determined by the recommendation module 144 by comparing the embeddings, where Sim represents the recommendation module 144 for comparing the embeddings:
[0060] Sim{GCN("Unknown Name"), GCN("Candidate 1") = 0.95
[0061] Sim{GCN("Unknown Name"), GCN("Candidate 2") = 0.25
[0062] Sim{GCN("Unknown Name"), GCN("Candidate n") = 0.30
[0063] As indicated by the equation above, the embedding pair corresponding to Candidate 1 has the highest similarity score (e.g., 0.95). Thus, in some aspects, the recommendation module 144 can link the unknown name and / or assign the unknown name to Candidate 1 based on this pair having the highest similarity score. In other aspects, in the case where the recommendation module 144 requires a minimum similarity score for assignment (such as a similarity score not less than 0.97), the similarity score of 0.95 may not meet the threshold, and thus the recommendation module 144 may not make an assignment.
[0064] Referring again to Figure 1 , various techniques can be used to train the GNN 138 for entity linking. In some aspects, training the GNN 138 can include using various NLP models to initialize the embeddings. Thus, the entity linking system 130 can include a first NLP1 140 and / or a second NLP2 142. The first NLP1 140 can be configured to initialize the embeddings of the nodes corresponding to names (e.g., unknown name, candidate) in the tripartite graph. For example, the first NLP1 140 can employ an NLP model such as BERT (Bidirectional Encoder Representations from Transformers). The second NLP2 142 can be configured to initialize the embeddings of the nodes corresponding to words (e.g., unknown word, candidate word). For example, the NLP2 142 can employ an NLP model such as Word2Vec, WordPiece, etc.
[0065] Still referring to Figure 1 , various techniques (such as ground truth or labeled datasets for correct entity linking) that can be used to train the GNN 138 may not be available for training the GNN 138. In this case, the training module 146 can be used to train the GNN 138 using the biased random walk technique. For example, the training module 146 can be used to supervise the GNN 138 during training by specifying positive samples, hard negative samples, and / or random negative samples within the tripartite graph. Any of the positive samples, hard negative samples, and / or random negative samples can be identified by comparing the unknown word with one of the candidates. The candidate compared with the unknown word during supervised training is sometimes referred to herein as the "target candidate".
[0066] The training module 146 can be used to specify a target candidate having one or more words in common with an unknown name (e.g., at least one of the candidate words of the target candidate is the same as one of the unknown words of the unknown name). Additionally, the training module 146 can be used to specify one of the one or more words common to both the unknown name and the target candidate as the "target word". The training module 146 can be configured to compare all unknown words that are not the target word with all candidate words of the target candidate that are not the target word to determine whether to specify the unknown word and target candidate pair as a positive sample, a hard negative sample, or a random negative sample.
[0067] In some aspects, the training module 146 can be configured to specify a positive sample by determining that all unknown words that are not the target word and all candidate words that are not the target word included in the target candidate have a string similarity score (e.g., Hamming distance, jaro - winkler distance) not less than a predetermined threshold (such as not less than 0.7, 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or not less than 0.99).
[0068] In some aspects, the training module 146 can be configured to specify a hard negative sample by determining that a first subset of all unknown words that are not the target word and a first subset of all candidate words that are not the target word included in the target candidate have a string similarity score not less than a first predetermined threshold (such as not less than 0.7, 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or not less than 0.99), while a second subset of all unknown words that are not the target word and a second subset of all candidate words that are not the target word included in the target candidate have a string similarity score not greater than a second predetermined threshold (such as not greater than 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, or not greater than 0.1).
[0069] In some aspects, the training module 146 can be configured to specify a random negative sample by determining that all unknown words that are not the target word and all candidate words that are not the target word included in the target candidate have a string similarity score not greater than a predetermined threshold (such as not greater than 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, or not greater than 0.1).
[0070] Now mainly referring to Figure 2C And also referring to Figure 1, The tripartite graph 200C shows examples of positive samples, hard negative samples, and random negative samples that can be specified by the training module 146 to train the GNN 138. In the tripartite graph 200C, the unknown name 222 is "The Delta Airlines". In some aspects, common words included in "The Delta Airlines", such as "The", can be excluded from consideration during the specification of training samples. Any of the candidates 228 can be a target candidate. For example, both "Delta Dental" 2281 and the unknown name 222 of "The Delta Airlines" include the target word "Delta". Identifying all the unknown words in the unknown word "The Delta Airlines" that are not "Delta" results in the underlined word "Airline" (excluding "The"). Identifying all the words in the target candidate "Delta Dental" that are not "Delta" results in the underlined word "Dental". The string similarity between "Airline" and "Dental" may be relatively low (e.g., not greater than the predetermined threshold of the random negative samples mentioned above). Therefore, the training module 146 can specify "The Delta Airlines" and "Delta Dental" as random negative samples 234. The pair of "The Delta Airlines" and "South Airlines" can also be a random negative sample 236, where "Airline" is the target word, and the string similarity between "Delta" and "South" is relatively low.
[0071] As another example, both "Delta Airline" 2282 and the unknown name 222 of "The Delta Airlines" include the target word "Delta". Identifying all the words in the target candidate "Delta Airline" that are not "Delta" results in the underlined word "Airline". The string similarity between "Airlines" and "Airline" may be relatively high (e.g., not less than the predetermined threshold of the positive samples mentioned above). Therefore, the training module 146 can specify "The Delta Airlines" and "Delta Airline" as positive samples 230.
[0072] As another example, both "South Delta Airlines" 2283 and "Delta Airlines" unknown name 222 include the target word "Delta". Identifying all words that are not "Delta" in the target candidate "South Delta Airlines" produces the underlined words "South" and "Airlines". The string similarity between "Airlines" and "Airlines" may be relatively high (e.g., not less than the first predetermined threshold of hard negative samples mentioned above). However, the string similarity between "Airlines" and "South" may be relatively low (e.g., not greater than the second predetermined threshold of hard negative samples mentioned above). Therefore, the training module 146 can designate "Delta Airlines" and "South Delta Airlines" as hard negative samples 232.
[0073] Figure 3 is a logic flow diagram of a method 300 for entity linking according to at least one aspect of the present disclosure. The method 300 may be Figure 1 The method 300 may be practiced by any combination of the entity linking system 130 and / or components of the entity linking system 130. According to the method 300, the extraction module 132 retrieves 302 a set of known names from the known entity database 150. The known names may be entity names stored in the known entity database 150. The tokenization module 134 tokenizes 304 each of the known names. In addition, according to the method 300, the extraction module 132 retrieves 306 unknown names from the information source 110. The tokenization module 134 tokenizes 308 the unknown names extracted from the information source 110.
[0074] Still refer to Figure 3 According to the method 300, the graph generation module 136 identifies 310 one or more candidates from the set of known names. In addition, the graph generation module 136 generates 312 a tripartite graph. The tripartite graph includes a first layer node corresponding to the unknown name, a second layer node corresponding to the unknown name and the one or more candidates, and one or more third layer nodes corresponding to the one or more candidates. In addition, the recommendation module 144 assigns 314 the unknown name to one of the known names by applying the tripartite graph to the GNN 138.
[0075] On the one hand, according to method 300, identifying 310 one or more candidates from the set of known names may include identifying each known name in the set of known names as a candidate. In this regard, the one or more candidates include all known names in the set of known names. Further, in this regard, the tripartite graph includes a first layer of nodes corresponding to the unknown name, a second layer of nodes corresponding to the words of the unknown name and the known names, and a third layer of nodes corresponding to the known names.
[0076] Figure 4 is a logical flow diagram of method 400 for assigning an unknown name to one of the known names in a set of known names by applying a tripartite graph to a graph neural network model according to at least one aspect of the present disclosure. In some aspects, method 400 may be included as part of the functionality of assigning an unknown name to 314 one of the known names described above with respect to Figure 3 the known names. Thus, the tripartite graph may include a first node, a second node, and one or more third nodes. Method 400 may be practiced by the entity linking system 130 and / or any combination of components of the entity linking system 130 described above with respect to Figure 1 the entity linking system 130. According to method 400, the GNN 138 generates 402 an unknown name embedding. The unknown name embedding corresponds to the first node of the tripartite graph. The GNN 138 may also generate 404 word embeddings. Each word embedding in the word embeddings corresponds to one of the second nodes in the second layer of nodes of the tripartite graph. Further, the GNN 138 generates 406 one or more candidate embeddings. Each candidate embedding in the one or more candidate embeddings corresponds to one of the one or more third nodes in the tripartite graph. In some aspects, prior to generating 402 the unknown name embedding (e.g., during training), the word embeddings and the one or more candidate embeddings are generated 404, 406 respectively by the GNN 138. Thus, in aspects where the entity linking system 130 is used to assign multiple different instances of an unknown name to known names, only the unknown name embedding 402 is newly generated (e.g., not the word embeddings or the one or more candidate embeddings) for each unknown name being identified. Referring again to the aspects of method 400, where a single unknown name is being assigned to one of the known names in the set of known names, the recommendation module 144 determines 408 a similarity score between the unknown name embedding and each of the one or more candidate embeddings.
[0077] Figure 5 is a logical flow diagram of method 500 for training a graph neural network model by specifying positive samples, hard negative samples, and random negative samples. Method 500 may be used to train the tripartite graph to be applied by the recommendation module 144, thereby assigning an unknown name to 314 as described above with respect toFigure 3 A GNN 138 with one of the known names in the known names described above. Thus, method 500 can be applicable to a tripartite graph structure that includes first-layer nodes corresponding to an unknown name, second-layer nodes corresponding to words of the known name and one or more candidates, and one or more third-layer nodes corresponding to the one or more candidates. The unknown name may include one or more unknown words, and each of the candidates may include one or more candidate words. Method 500 can be practiced by any combination of the entity linking system 130 described above and / or components of the entity linking system 130. Figure 1 The entity linking system 130 and / or any combination of components of the entity linking system 130 described above.
[0078] Still referring to Figure 5 , according to method 500, the training module 146 identifies 502 target words. The target words are included in both the one or more unknown words of the unknown name and the one or more candidate words of one of the one or more candidates. In addition, the training module 146 identifies 504 the target candidate from the one or more candidates. The target candidate includes the target words. Still according to method 500, the training module 146 identifies 506 all one or more unknown words that are not target words, and identifies 508 all one or more candidate words that are not target words and are included in the target candidate. In addition, the training module 146 compares 510 all one or more unknown words that are not target words with all one or more candidate words that are not target words and are included in the target candidate.
[0079] Still referring to Figure 5 , based on the comparison 510 function, if the training module 146 determines 512 that all one or more unknown words that are not target words and all one or more candidate words that are not target words and are included in the target candidate have a string similarity score not less than a predetermined threshold (e.g., not less than 0.7, 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98 or not less than 0.99), then the training module 146 designates 518 a positive sample.
[0080] Still referring to Figure 5, based on the comparison 510 function, if the training module 146 determines 514 that: (i) a first subset of the one or more unknown words that are not the target word and a second subset of the one or more candidate words that are not the target word and are included in the target candidate have a first string similarity score that is not less than a first predetermined threshold (e.g., not less than 0.7, 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or not less than 0.99), and (ii) a third subset of the one or more unknown words that are not the target word and a fourth subset of the one or more candidate words that are not the target word and are included in the target candidate have a second string similarity score that is not greater than a second predetermined threshold (e.g., not greater than 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, or not greater than 0.1), then the training module 146 designates 520 hard negative samples.
[0081] Still referring to Figure 5 , based on the comparison 510 function, if the training module 146 determines 516 that the string similarity score of all one or more unknown words that are not the target word and all one or more candidate words that are not the target word and are included in the target candidate is not greater than a predetermined threshold (e.g., not greater than 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, or not greater than 0.1), then the training module 146 designates 522 random negative samples.
[0082] Referred to herein Figure 1 The entity linking system 130 and modules described herein can operate on one or more computer devices to facilitate the functions described herein. Additionally, the one or more computer devices can use any suitable number of subsystems to facilitate the functions described herein. For example, Figure 6 is a block diagram of a computer device 3000 having a data processing subsystem or component according to at least one aspect of the present disclosure. Figure 6The subsystems shown are interconnected via a system bus 3010. Additional subsystems such as a printer 3018, a keyboard 3026, a fixed disk 3028 (or other memory containing computer-readable media), a monitor 3022 coupled to a display adapter 3020, etc. are shown. Peripheral devices and input / output (I / O) devices coupled to an I / O controller 3012 (which can be a processor or other suitable controller) can be connected to the computer system by any number of means known in the art, such as a serial port 3024. For example, the serial port 3024 or an external interface 3030 can be used to connect the computer device to a wide area network (such as the Internet), a mouse input device, or a scanner. The interconnection via the system bus allows the central processor 3016 to communicate with each subsystem and allows control of the execution of instructions from the system memory 3014 or the fixed disk 3028, as well as information exchange between subsystems. The system memory 3014 and / or the fixed disk 3028 can be embodied as computer-readable media.
[0083] Figure 7 FIG. 4 is a schematic representation of an example system 4000 including a host 4002 according to at least one aspect of the present disclosure, within which a set of instructions for performing any one or more of the methods discussed herein can be executed. In various aspects, the host 4002 operates as a stand-alone device or can be connected (e.g., networked) to other machines. In a networked deployment, the host 4002 can operate in the function of a server or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The host 3002 can be a computer or computing device, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a portable music player (e.g., a portable hard disk audio device, such as a Moving Picture Experts Group Audio Layer 3 (MP3) player), a network appliance, a network router, a switch, or a bridge, or any machine capable of executing a set of instructions (sequentially or otherwise) specifying actions to be taken by that machine. Additionally, although only a single machine is illustrated, the term "machine" should also be understood to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein.
[0084] Example system 4000 includes a host 4002 that runs a main operating system (OS) 4004 on one or more processors / processor cores 4006 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both) and various memory nodes 4008. The main OS 4004 may include a hypervisor 4010 that is capable of controlling functions and / or communicating with virtual machines (“VMs”) 4012 running on machine-readable media. The VM 4012 may also include a virtual CPU or vCPU 4014. The memory nodes 4008 may be linked or attached to virtual memory nodes or vNodes 4016. When the memory nodes 4008 are linked or attached to their corresponding vNodes 4016, data can then be directly mapped from the memory nodes 4008 to their corresponding vNodes 4016.
[0085] All of the various components shown in the host 4002 may be connected to and communicate with each other, either directly or through a bus (not shown) or through other coupling or communication channels or mechanisms. The host 4002 may further include a video display, an audio device, or other peripheral devices 4018 (e.g., a liquid crystal display (LCD), an alphanumeric input device (including, e.g., a keyboard), a cursor control device (e.g., a mouse), a voice recognition or biometric authentication unit, an external drive, a signal generation device (e.g., a speaker)), a permanent storage device 4020 (also referred to as a disk drive unit), and a network interface device 4022. The host 4002 may further include a data encryption module (not shown) for encrypting data. The components provided in the host 4002 are components that are typically present in a computer system that may be adapted to be used with aspects of the present disclosure and are intended to represent a broad class of such computer components known in the art. Thus, the system 4000 may be a server, a minicomputer, a mainframe computer, or any other computer system. The computer may also include different bus configurations, networking platforms, multiprocessor platforms, etc. A variety of operating systems may be used, including UNIX, LINUX, WINDOWS, QNX ANDROID, IOS, CHROME, TIZEN, and other suitable operating systems.
[0086] The disk drive unit 4024 can also be a solid state drive (SSD), a hard disk drive (HDD), or other drives including a computer or machine-readable medium having stored thereon one or more sets of instructions and data structures (e.g., data / instruction 4026) embodying or utilizing any one or more of the methods or functions described herein. The data / instruction 4026 can also reside, in whole or at least in part, within the main memory portion of the memory node 4008 and / or within the processor 4006 during execution thereof by the host 4002. The data / instruction 4026 can be further transmitted or received via the network interface device 4022 over the network 4028 by utilizing any one of a number of well-known transport protocols (e.g., Hyper Text Transfer Protocol (HTTP)).
[0087] The processor 4006 and the memory node 4008 can also include machine-readable media. The term "computer-readable medium" or "machine-readable medium" should be regarded as including a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) storing one or more sets of instructions. The term "computer-readable medium" should also be regarded as including any medium capable of storing, encoding, or carrying a set of instructions for execution by the host 4002 and causing the host 4002 to perform any one or more of the methods of this application, or any medium capable of storing, encoding, or carrying the data structures utilized by or associated with this set of instructions. Thus, the term "computer-readable medium" should be regarded as including, but not limited to, solid state memories, optical and magnetic media, and carrier signals. Such media can also include, but not limited to, hard disks, floppy disks, flash memory cards, digital video discs, random access memory (RAM), read-only memory (ROM), etc. The exemplary aspects described herein can be implemented in an operating environment comprising software installed on a computer, software installed in hardware, or a combination of software and hardware.
[0088] Those skilled in the art will recognize that an Internet service can be configured to provide Internet access to one or more computing devices coupled to the Internet service, and the computing devices can include one or more processors, buses, memory devices, display devices, input / output devices, etc. In addition, those skilled in the art can appreciate that the Internet service can be coupled to one or more databases, repositories, servers, etc., which can be used to implement any aspect of the present disclosure as described herein.
[0089] Computer program instructions can also be loaded onto a computer, a server, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer or other programmable apparatus provide a process for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0090] For example, suitable networks can include any one or more of the following or interface with any one or more of the following: a local intranet; a PAN (Personal Area Network); a LAN (Local Area Network); a WAN (Wide Area Network); a MAN (Metropolitan Area Network); a Virtual Private Network (VPN); a Storage Area Network (SAN); a Frame Relay connection; an Advanced Intelligent Network (AIN) connection; a Synchronous Optical Network (SONET) connection; a digital T1, T3, E1, or E3 line; a Digital Data Service (DDS) connection; a DSL (Digital Subscriber Line) connection; an Ethernet connection; an ISDN (Integrated Services Digital Network) line; a dial-up port (such as a V.90, V.34, or V.34bis analog modem connection); a cable modem; an ATM (Asynchronous Transfer Mode) connection; or an FDDI (Fiber Distributed Data Interface) or CDDI (Copper Distributed Data Interface) connection. Additionally, the communication can also include a link to any wireless network among various wireless networks, including a WAP (Wireless Application Protocol), GPRS (General Packet Radio Service), GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access), or TDMA (Time Division Multiple Access), a cellular telephone network, a GPS (Global Positioning System), CDPD (Cellular Digital Packet Data), a RIM (Research in Motion, Limited) duplex paging network, a Bluetooth radio, or a radio frequency network based on IEEE 802.11. The network can further include any one or more of the following or interface with any one or more of the following: an RS-232 serial connection, an IEEE-1394 (FireWire) connection, a Fibre Channel connection, an IrDA (Infrared) port, an SCSI (Small Computer System Interface) connection, a USB (Universal Serial Bus) connection, or other wired or wireless, digital or analog interfaces or connections, a mesh or networking.
[0091] Generally, a cloud-based computing environment is a resource that typically combines the computing power of large groupings of processors (such as within a web server) and / or combines the storage capacity of large groupings of computer memory or storage devices. A system that provides cloud-based resources can be utilized only by its owner, or such a system can be accessed by external users who deploy applications within the computing infrastructure to obtain the benefits of large computing or storage resources.
[0092] For example, a cloud is formed by a network of web servers that includes multiple computing devices, such as host 4002, where each server 4030 (or at least a plurality of them) provides processor and / or storage resources. These servers manage workloads provided by multiple users (e.g., cloud resource customers or other users). Typically, each user's workload demand for the cloud changes in real time, sometimes significantly. The nature and extent of these changes usually depend on the type of business associated with the user.
[0093] It should be noted that any hardware platform suitable for performing the processes described herein is suitable for use with the technology. As used herein, the terms "computer-readable storage medium" and "computer-readable storage media" refer to any one or more media that participate in providing instructions to a CPU for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical or magnetic disks, such as fixed disks. Volatile media includes dynamic memory, such as system RAM. Transmission media includes coaxial cables, copper wire, and fiber optics, among others, which include wires that form an aspect of a bus. Transmission media can also take the form of acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media include, for example, floppy disks, hard disks, magnetic tape, any other magnetic media, CD-ROM discs, digital video discs (DVDs), any other optical media, any other physical media with a pattern of marks or holes, RAM, PROM, EPROM, EEPROM, FLASH EPROM, any other memory chip or data exchange adapter, a carrier wave, or any other media from which a computer can read.
[0094] Various forms of computer-readable media can participate in carrying one or more sequences of one or more instructions to a CPU for execution. A bus carries data to system RAM, and the CPU retrieves instructions from the system RAM and executes the instructions. Instructions received by the system RAM can optionally be stored on a fixed disk before or after being executed by the CPU.
[0095] The computer program code for performing operations on aspects of the technology of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages (such as Java, Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" programming language, Go, Python, or other programming languages including assembly language). The program code may be executed entirely on the user's computer, partially on the user's computer; as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).
[0096] Examples of the methods and systems according to various aspects of the present disclosure are provided below in the following numbered clauses. One aspect of any method and / or system in the method and / or system may include any one or more than one of the numbered clauses described below and any combination of the numbered clauses.
[0097] Clause 1. A computer-implemented method, comprising: retrieving, by an extraction module, known names from a database, wherein the known names are entity names stored in the database; tokenizing, by a tokenization module, the known names; retrieving, by the extraction module, unknown names extracted from an information source; tokenizing, by the tokenization module, the unknown names; identifying, by the graph generation module, candidates from the known names; generating, by the graph generation module, a tripartite graph that includes a first layer of nodes corresponding to the unknown names, a second layer of nodes corresponding to words contained in the unknown names and the candidates, and a third layer of nodes corresponding to the candidates; applying, by a recommendation module, the tripartite graph to a graph neural network model; and assigning, by the recommendation module, the unknown names to one of the known names based on the application of the tripartite graph to the graph neural network model by the recommendation module.
[0098] Clause 2. The computer-implemented method according to clause 1, wherein the unknown names include unknown words, and the candidates include candidate words.
[0099] Clause 3. The computer-implemented method according to any one of clauses 1 to 2, wherein the candidate words are the same as the unknown words.
[0100] Clause 4. The computer-implemented method according to any one of Clauses 1 to 2, wherein identifying the candidate from the known names includes identifying a known name among the known names that includes the unknown word.
[0101] Clause 5. The computer-implemented method according to any one of Clauses 1 to 4, wherein at least one of a positive sample, a negative sample, or a random negative sample is specified by a training module; and the graph neural network model is supervised by the training module based on the at least one of the specified positive sample, negative sample, or random negative sample.
[0102] Clause 6. The computer-implemented method according to any one of Clauses 1 to 5, wherein the unknown name includes an unknown word, the candidate includes a candidate word, and specifying the at least one of the positive sample, negative sample, or random negative sample includes identifying a target word in the unknown name and the candidate by the training module.
[0103] Clause 7. The computer-implemented method according to any one of Clauses 1 to 6, wherein specifying the positive sample further includes the training module determining that the unknown word that is not the target word and the candidate word that is not the target word have a string similarity score not less than a predetermined threshold.
[0104] Clause 8. The computer-implemented method according to any one of Clauses 1 to 7, wherein specifying the hard negative sample further includes: the training module determining that a first subset of the unknown words that are not the target word and a second subset of the candidate words that are not the target word have a first string similarity score not less than a first predetermined threshold; and the training module determining that a third subset of the unknown words that are not the target word and a fourth subset of the candidate words that are not the target word have a second string similarity score not greater than a second predetermined threshold.
[0105] Clause 9. The computer-implemented method according to any one of Clauses 1 to 8, wherein specifying the random negative sample further includes: the training module determining that the unknown word that is not the target word and the candidate word that is not the target word have a string similarity score not greater than a predetermined threshold.
[0106] Clause 10. The computer-implemented method according to any one of Clauses 1 to 9, wherein the graph neural network model is a graph convolutional network model.
[0107] Clause 11. The computer-implemented method according to any one of Clauses 1 to 10, wherein the assigning the unknown name to one of the known names based on applying the tripartite graph to the graph neural network model includes: generating, by the graph neural network model: an unknown name embedding corresponding to the first-layer nodes; a word embedding corresponding to the second-layer nodes; and a candidate embedding corresponding to the third-layer nodes; and determining, by the recommendation module, a similarity score between the unknown name embedding and the candidate embedding.
[0108] Clause 12. The computer-implemented method according to any one of Clauses 1 to 11, wherein the determining the similarity score between the unknown name embedding and the candidate embedding includes applying, by the recommendation module, the unknown name embedding and the candidate embedding to a trained regression model.
[0109] Clause 13. The computer-implemented method according to any one of Clauses 1 to 11, wherein the determining the similarity score between the unknown name embedding and the candidate embedding includes applying, by the recommendation module, the unknown name embedding and the candidate embedding to a trained classification model.
[0110] Clause 14. The computer-implemented method according to any one of Clauses 1 to 13, wherein the words included in the unknown name and the candidate are strings not separated by spaces.
[0111] Clause 15. The computer-implemented method according to any one of Clauses 1 to 14, wherein the tokenizing of the known name includes performing word tokenization of the known name by the tokenization module, and wherein the tokenizing of the unknown name includes performing the word tokenization of the unknown name by the tokenization module.
[0112] Clause 16. An entity linking system, comprising: an extraction module configured to extract unknown names from an information source and extract known names from a database, where the known names are names of entities stored in the database; a tokenization module configured to tokenize the unknown names and tokenize the known names; a graph generation module configured to identify candidates from the known names and generate a tripartite graph based on the unknown names and the candidates, where the tripartite graph includes a first layer of nodes corresponding to the unknown names, a second layer of nodes corresponding to words contained in the unknown names and the candidates, and a third layer of nodes corresponding to the candidates; a graph neural network configured to generate an unknown name embedding and a candidate embedding based on the tripartite graph, where the unknown name embedding corresponds to the first layer of nodes, and where the candidate embedding corresponds to the third layer of nodes; and a recommendation module configured to determine a similarity score between the unknown name embedding and the candidate embedding.
[0113] Clause 17. The entity linking system according to Clause 16, where the graph neural network is a graph convolutional network.
[0114] Clause 18. The entity linking system according to any one of Clauses 16 to 17, where the words contained in the unknown names and the candidates include strings not separated by spaces.
[0115] Clause 19. The entity linking system according to any one of Clauses 16 to 18, where the words contained in the unknown names and the candidates include unknown words included in the unknown names and candidate words included in the candidates.
[0116] Clause 20. The entity linking system according to any one of Clauses 16 to 19, where the graph convolutional network is trained by identifying at least one of positive samples, negative samples, or random negative samples.
[0117] In addition, it should be understood that any one or more of the forms, forms of expressions, and examples described below can be combined with any one or more of the other forms, forms of expressions, and examples described below.
[0118] Any software component or function described in this application can be implemented as software code executed by a processor using any suitable computer language (e.g., Python, Java, C++, or Perl) and using, for example, conventional or object-oriented techniques. The software code can be stored on a computer-readable medium as a series of instructions or commands, such as random access memory (RAM), read-only memory (ROM), magnetic media (such as a hard disk drive or a floppy disk), or optical media (such as a CD-ROM). Any such computer-readable medium can reside on or within a single computing device and can be on or within different computing devices in a system or network.
[0119] Although several forms have been shown and described, the applicant does not intend to limit or restrict the scope of the appended claims to such details. Many modifications, variations, changes, substitutions, combinations, and equivalents of those forms can be made without departing from the scope of the present disclosure, and those many modifications, variations, changes, substitutions, combinations, and equivalents will occur to those skilled in the art. Additionally, the structure of each element associated with those forms can alternatively be described as a means for providing the function performed by the element. Moreover, where the material of certain components is disclosed, other materials can be used. Accordingly, it should be understood that the foregoing description and the appended claims are intended to cover all such modifications, combinations, and variations that fall within the scope of the disclosed forms. The appended claims are intended to cover all such modifications, variations, changes, substitutions, modifications, and equivalents.
[0120] The foregoing detailed description has set forth various forms of apparatuses and / or processes by use of block diagrams, flowcharts, and / or examples. In instances in which such block diagrams, flowcharts, and / or examples contain one or more functions and / or operations, those skilled in the art will understand that each function and / or operation within such block diagrams, flowcharts, and / or examples can be implemented, individually and / or jointly, by a variety of hardware, software, firmware, or virtually any combination thereof. Those skilled in the art will recognize that some aspects of the forms disclosed herein can be equivalently implemented, in whole or in part, in an integrated circuit as one or more computer programs running on one or more computers (e.g., one or more programs running on one or more computer systems), one or more programs running on one or more processors (e.g., one or more programs running on one or more microprocessors), firmware, or virtually any combination thereof, and, in accordance with the present disclosure, designing the circuitry and / or writing the code for the software and / or firmware would be well within the skill of those in the art. Additionally, those skilled in the art will appreciate that the mechanisms of the subject matter described herein can be distributed in a variety of forms as one or more program products, and that the illustrative forms described herein apply regardless of the particular type of signal-bearing medium used to actually effectuate such distribution.
[0121] Instructions for programming logic to perform the various disclosed aspects can be stored in a memory within the system, such as dynamic random access memory (DRAM), cache, flash memory, or other storage devices. Additionally, the instructions can be distributed via a network or by means of other computer-readable media. Thus, machine-readable media can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), but are not limited to floppy disks, optical disks, compact disc read-only memory (CD-ROM) and magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable storage devices for transmitting information over the Internet in the form of electrical, optical, acoustic, or other propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). Thus, non-transitory computer-readable media include any type of tangible machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0122] As used in any aspect herein, the term "control circuit" can refer to, for example, hardwired circuitry, programmable circuitry (e.g., a computer processor, processing unit, processor, microcontroller, microcontroller unit, controller, digital signal processor (DSP), programmable logic device (PLD), programmable logic array (PLA), or field programmable gate array (FPGA) including one or more individual instruction processing cores), state machine circuitry, firmware storing instructions executed by the programmable circuitry, and any combination thereof. The control circuit can be embodied jointly or separately as circuitry forming part of a larger system, e.g., an integrated circuit (IC), an application specific integrated circuit (ASIC), a system on a chip (SoC), a desktop computer, a laptop computer, a tablet computer, a server, a smart phone, etc. Thus, as used herein, "control circuit" includes, but is not limited to: circuitry having at least one discrete circuit; circuitry having at least one integrated circuit; circuitry having at least one application specific integrated circuit; circuitry forming a general purpose computing device configured by a computer program (e.g., a general purpose computer configured by a computer program that at least partially performs the processes and / or apparatus described herein, or a microprocessor configured by a computer program that at least partially performs the processes and / or apparatus described herein); circuitry forming a memory device (e.g., various forms of random access memory); and / or circuitry forming a communication device (e.g., a modem, a communication switch, or an optoelectronic device). Those skilled in the art will recognize that the subject matter described herein can be implemented in analog or digital form or some combination thereof.
[0123] As used in any aspect of this disclosure, the term "logic" can refer to an app, software, firmware, and / or circuitry configured to perform any of the foregoing operations. Software can be embodied as a software package, code, instructions, instruction set, and / or data recorded on a non-transitory computer-readable storage medium. Firmware can be embodied as code, instructions, or instruction set and / or data hard-coded (e.g., non-volatile) in a memory device.
[0124] As used in any aspect of this disclosure, the terms "component", "system", "module", etc. can refer to a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution.
[0125] As used in any aspect of this disclosure, "algorithm" refers to a self-consistent sequence of steps that produces a desired result, where "step" refers to the manipulation of physical quantities and / or logical states, which may (but need not) take the form of electrical or magnetic signals capable of being stored, transmitted, combined, compared, and otherwise manipulated. Common usage refers to these signals as bits, values, elements, symbols, characters, terms, numbers, etc. These terms and similar terms can be associated with appropriate physical quantities and are merely convenient labels applied to these quantities and / or states.
[0126] The network may include a packet-switched network. The communication devices may be capable of communicating with each other using a selected packet-switched network communication protocol. An example communication protocol may include an Ethernet communication protocol, which may be capable of permitting communication using the Transmission Control Protocol / Internet Protocol (TCP / IP). The Ethernet protocol may conform to or be compatible with the Ethernet standard named "IEEE 802.3 Standard" published by the Institute of Electrical and Electronics Engineers (IEEE) in December 2008 and / or subsequent versions of this standard. Alternatively or additionally, the communication devices may be capable of communicating with each other using the X.25 communication protocol. The X.25 communication protocol may conform to or be compatible with the standards promulgated by the International Telecommunication Union - Telecommunication Standardization Sector (ITU-T). Alternatively or additionally, the communication devices may be capable of communicating with each other using the Frame Relay communication protocol. The Frame Relay communication protocol may conform to or be compatible with the standards promulgated by the Consultative Committee for International Telegraph and Telephone (CCITT) and / or the American National Standards Institute (ANSI). Alternatively or additionally, the transceivers may be capable of communicating with each other using the Asynchronous Transfer Mode (ATM) communication protocol. The ATM communication protocol may conform to or be compatible with the ATM standard named "ATM-MPLS Network Interworking 2.0" published by the ATM Forum in August 2001 and / or subsequent versions of this standard. Of course, different and / or later-developed connection-oriented network communication protocols are also contemplated herein.
[0127] Unless otherwise specifically stated, as will be apparent from the foregoing disclosure, it should be understood that throughout the foregoing disclosure, discussions using terms such as "processing," "computing," "calculating," "determining," "displaying," etc., refer to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical (electronic) quantities within the registers and memories of the computer system and transforms it into other data similarly represented as physical quantities within the memories or registers or other such information storage, transmission, or display devices of the computer system.
[0128] One or more components may be referred to herein as “configured to,” “configurable to,” “operable / operable to,” “adapted / suitable for,” “capable of,” “in accordance with,” etc. Unless the context otherwise requires, one of ordinary skill in the art will recognize that “configured to” can generally cover active state components and / or inactive state components and / or standby state components.
[0129] One of ordinary skill in the art will recognize that, in general, the terms used herein, and especially the terms in the appended claims (e.g., the subject matter of the appended claims) are generally intended to be “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “including but not limited to,” etc.). One of ordinary skill in the art will further understand that if a specific number of introduced claim recitations is intended, such intent will be expressly recited in the claim, and in the absence of such recitation, no such intent exists. For example, for purposes of illustration, the following appended claims may contain the use of introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed as implying that the introduction of a claim recitation by the indefinite article “a” or “an” limits any particular claim containing such introduced claim recitation to a claim only containing one such recitation, even when the same claim includes an introductory phrase “one or more” or “at least one” as well as an indefinite article such as “a” or “an” (e.g., “a” and / or “an” should generally be interpreted to mean “at least one” or “one or more”); this also holds for the use of definite articles used to introduce claim recitations.
[0130] In addition, even if a specific number recited in an introduced claim is explicitly recited, those skilled in the art will recognize that such a recitation should generally be construed to mean at least the recited number (e.g., a simple recitation of "two recitations" without other modifiers generally means at least two recitations, or two or more recitations). Further, in cases where a convention such as "at least one of A, B, and C, etc." is used, generally, such a construction is intended to be such that those skilled in the art will understand the meaning of the convention (e.g., a "system having at least one of A, B, and C" will include, but not be limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In cases where a convention such as "at least one of A, B, or C, etc." is used, generally, such a construction is intended to be such that those skilled in the art will understand the meaning of the convention (e.g., a "system having at least one of A, B, or C" will include, but not be limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Those skilled in the art will further understand that, generally, unless the context otherwise indicates, whether in the specification, claims, or drawings, separate words and / or phrases presenting two or more alternative terms should be understood to contemplate the possibility of including one of the terms, any one of the terms, or both of the terms. For example, the phrase "A or B" will generally be understood to include the possibility of "A" or "B" or "A and B".
[0131] Regarding the appended claims, those skilled in the art will understand that the operations recited therein can generally be performed in any order. Moreover, although various operation flowcharts are presented in sequence, it should be understood that the various operations can be performed in other orders different from the presented order, or the various operations can be performed simultaneously. Unless the context otherwise provides, examples of such alternative orderings can include overlapping, interleaving, interrupting, reordering, incrementing, preparatory, supplementary, simultaneous, reversing, or other variant orderings. Further, unless the context otherwise provides, terms such as "in response to", "associated with", or other past tense adjectives generally are not intended to exclude such variants.
[0132] It is noted that any reference to "an aspect", "one aspect", "an example", "one example", etc. means that the specific features, structures, or characteristics described in connection with the aspect are included in at least one aspect. Thus, the phrases "in one aspect", "in one aspect", "in an example", and "in one example" that appear throughout the specification do not necessarily all refer to the same aspect. Further, in one or more aspects, the specific features, structures, or characteristics can be combined in any suitable manner.
[0133] Any patent application, patent, non-patent publication, or other publicly available material cited in this specification and / or listed in any application data sheet is hereby incorporated by reference into this document, provided that the incorporated material is not inconsistent with this document. Thus, and to the extent necessary, the disclosure set forth herein supersedes any conflicting material incorporated by reference into this document. Any material or portion thereof that is alleged to be incorporated by reference into this document but which conflicts with the existing definitions, statements, or other publicly available material set forth herein will be incorporated only to the extent that no conflict arises between the incorporated material and the existing publicly available material.
[0134] In summary, many benefits resulting from the adoption of the concepts described herein have been described. The foregoing description has been presented in one or more forms for purposes of illustration and description. It is not intended to be exhaustive or limited to the exact forms disclosed. Modifications or variations are possible in light of the above teachings. The one or more forms have been selected and described in order to illustrate the principles and practical applications, so as to enable one of ordinary skill in the art to utilize the various forms and make various modifications suitable for the particular purposes contemplated. The scope is intended to be defined by the claims submitted herewith.
[0135] The above description is illustrative and not intended to be limiting. Many variations of the claimed subject matter will become apparent to those skilled in the art upon review of this disclosure. Accordingly, the scope of this disclosure should not be determined with reference to the above description, but rather should be determined with reference to the pending claims and their full scope or equivalents.
[0136] All patents, patent applications, publications, and descriptions mentioned above are hereby incorporated by reference in their entirety for all purposes. No admission is made that any of the content is prior art.
Claims
1. A computer-implemented method, comprising: Retrieving, by an extraction module, known names from a database, where the known names are entity names stored in the database; Tokenizing, by a tokenization module, the known names; Retrieving, by the extraction module, unknown names extracted from an information source; Tokenizing, by the tokenization module, the unknown names; Identifying, by a graph generation module, candidates from the known names; Generating, by the graph generation module, a tripartite graph that includes a first layer of nodes corresponding to the unknown names, a second layer of nodes corresponding to words contained in the unknown names and the candidates, and a third layer of nodes corresponding to the candidates; Applying, by a recommendation module, the tripartite graph to a graph neural network model; And Assigning, by the recommendation module, the unknown name to one of the known names based on the application of the tripartite graph to the graph neural network model by the recommendation module.
2. The computer-implemented method according to claim 1, wherein the unknown name includes unknown words, and the candidates include candidate words.
3. The computer-implemented method according to claim 2, wherein the candidate words are the same as the unknown words.
4. The computer-implemented method according to claim 2, wherein identifying the candidates from the known names includes identifying a known name among the known names that includes the unknown words.
5. The computer-implemented method according to claim 1, further comprising: Designating, by a training module, at least one of positive samples, hard negative samples, or random negative samples; and Supervising, by the training module, the graph neural network model based on the designation of the at least one of the positive samples, the hard negative samples, or the random negative samples.
6. The computer-implemented method according to claim 5, wherein the unknown name includes unknown words, the candidates include candidate words, and designating the at least one of the positive samples, the hard negative samples, or the random negative samples includes identifying, by the training module, target words in the unknown names and the candidates.
7. The computer-implemented method according to claim 6, wherein designating the positive samples further includes determining, by the training module, that the unknown words that are not the target words and the candidate words that are not the target words have a string similarity score not less than a predetermined threshold.
8. The computer-implemented method according to claim 6, wherein designating the hard negative samples further includes: Determining, by the training module, that a first subset of the unknown words that are not the target words and a second subset of the candidate words that are not the target words have a first string similarity score not less than a first predetermined threshold; and Determining, by the training module, that a third subset of the unknown words that are not the target words and a fourth subset of the candidate words that are not the target words have a second string similarity score not greater than a second predetermined threshold.
9. The computer-implemented method according to claim 6, wherein specifying the random negative samples further comprises: determining, by the training module, that an unknown word that is not the target word and a candidate word that is not the target word have a string similarity score not greater than a predetermined threshold.
10. The computer-implemented method according to claim 1, wherein the graph neural network model is a graph convolutional network model.
11. The computer-implemented method according to claim 1, wherein assigning the unknown name to one of the known names based on applying the tripartite graph to the graph neural network model comprises: generating, by the graph neural network model: an unknown name embedding corresponding to the first layer nodes; a word embedding corresponding to the second layer nodes; and a candidate embedding corresponding to the third layer nodes; and determining, by the recommendation module, a similarity score between the unknown name embedding and the candidate embedding.
12. The computer-implemented method according to claim 11, wherein determining the similarity score between the unknown name embedding and the candidate embedding comprises applying, by the recommendation module, the unknown name embedding and the candidate embedding to a trained regression model.
13. The computer-implemented method according to claim 11, wherein determining the similarity score between the unknown name embedding and the candidate embedding comprises applying, by the recommendation module, the unknown name embedding and the candidate embedding to a trained classification model.
14. The computer-implemented method according to claim 1, wherein the words in the unknown name and the candidate contain strings not separated by spaces.
15. The computer-implemented method according to claim 1, wherein tokenizing the known name comprises performing word tokenization of the known name by the tokenization module, and wherein tokenizing the unknown name comprises performing word tokenization of the unknown name by the tokenization module.
16. An entity linking system, comprising: an extraction module configured to extract unknown names from an information source and extract known names from a database, wherein the known names are names of entities stored in the database; a tokenization module configured to tokenize the unknown names and tokenize the known names; a graph generation module configured to identify candidates from the known names and generate a tripartite graph based on the unknown names and the candidates, wherein the tripartite graph comprises first layer nodes corresponding to the unknown names, second layer nodes corresponding to words contained in the unknown names and the candidates, and third layer nodes corresponding to the candidates; a graph neural network configured to generate an unknown name embedding and a candidate embedding based on the tripartite graph, wherein the unknown name embedding corresponds to the first layer nodes and the candidate embedding corresponds to the third layer nodes; and A recommendation module, the recommendation module being configured to determine a similarity score between the unknown name embedding and the candidate embedding.
17. The entity linking system according to claim 16, wherein the graph neural network is a graph convolutional network.
18. The entity linking system according to claim 17, wherein the words included in the unknown name and the candidate contain strings not separated by spaces.
19. The entity linking system according to claim 18, wherein the words included in the unknown name and the candidate contain unknown words included in the unknown name and candidate words included in the candidate.
20. The entity linking system according to claim 19, wherein the convolutional graph network is trained by identifying at least one of positive samples, hard negative samples, or random negative samples.
Citation Information
Patent Citations
Improvement in book-cases
US220265A