Entity Linking Using Subgraph Matching

CN120035824A8Pending Publication Date: 2025-07-18VISA INTERNATIONAL SERVICE ASSOCIATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380063044.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2023-09-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has ambiguity and inconsistency in the entity linking process, making it difficult to accurately determine the entities mentioned in the information source and their related information, and the string similarity matching technology is sensitive to data quality and cannot effectively deal with information sources that lack titles or do not contain entity names in the titles.

Method used

The sub-graph matching technology is used to combine the graph neural network, and the information source and known entity attribute sets are extracted, unknown entity diagrams and known entity diagrams are generated, and graph embedding is generated using the graph neural network model, and the information source is allocated to the corresponding known entity based on the embedding space similarity score.

Benefits of technology

It realizes automatic, accurate and efficient entity linking, can process complex and changeable information source data, and improves the accuracy of entity identification and information association.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035824A8_ABST
    Figure CN120035824A8_ABST
Patent Text Reader

Abstract

Systems and methods for entity linking using graph neural networks are disclosed. In one aspect, a method for entity linking may include extracting a first set of attributes of an unknown entity from an information source, and retrieving a second set of attributes of known entities from a database, wherein each of the second set of attributes corresponds to one of the known entities. The method may further include: generating an unknown entity graph based on the first set of attributes; generating a known entity graph based on the second set of attributes; generating an unknown entity graph embedding by applying the unknown entity graph to a graph neural network; and generating a known entity graph embedding by applying the known entity graph to the graph neural network. The method may further include assigning the information source to one of the known entities based on the unknown entity graph embedding and the known entity graph embedding.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of and priority under 35 U.S.C. §119(e) to U.S. Provisional Application Serial No. 63 / 377,662, filed on September 29, 2023, entitled “ENTITY LINKING USING AGRAPH NEURAL NETWORK,” the contents of which are hereby incorporated by reference in their entirety. Technical Field

[0003] At least some aspects of the present disclosure relate to entity linking, such as entity linking using subgraph matching and natural language processing. Background Art

[0004] In the field of natural language processing (NLP), entity linking, sometimes referred to as named entity linking or named entity matching, generally involves determining that a word or string of words described in a text refers to a specific entity. Thus, entity linking may involve assigning a unique identity to a word or string of words. In many cases, a unique identity may be a person or an organization, such as a company, foundation, charity, or government organization.

[0005] Entity linking can be a valuable tool for associating information with a specific person or organization. In some aspects, this is because entity linking can enable companies to evaluate other entities using a large amount of information accessible via the Internet. For example, various private and public information sources, such as news articles, wikis, social media, and other databases and publications, can contain information related to an entity. This information may be related to assessing the reputation of an entity, assessing the risk of doing business with an entity, assessing the financial performance of an entity, etc. However, since there may be millions of information sources, and there may be thousands or even millions of entities of interest to the assessment company, it may be difficult to manually identify information that may be relevant. Therefore, NLP and entity linking can be used as an automatic way to identify specific entities mentioned in these information sources and / or determine that a specific information source such as a news article is about a specific entity.

[0006] However, there are several challenges associated with entity linking. As an example, there is often ambiguity and inconsistency in the names used to refer to specific entities. A news article about an entity whose legal name is "United Airlines, Inc." may instead state the name "United" in the body of the article or even in the title of the article. As a result, a computer-implemented process may mistake the name "United" for other entities, such as "United Health Care" and "United Technology Corp.", etc.

[0007] Furthermore, even if the names mentioned in the information source are correctly linked to a specific entity, the computer-implemented process may have difficulty determining whether the information source primarily provides information about the specific entity. For example, an online news article may report on various operational problems that a company is experiencing. In addition to reciting the company name, the article may also recite multiple other entity names related to the company, such as subsidiary names, parent company names, customer names, supplier names, names of company executives, etc. Although the computer-implemented process is able to link these names to a unique entity, the computer-implemented process may not be able to reliably determine which unique entity is experiencing the operational problems reported by the article.

[0008] Some approaches to entity linking employ string similarity matching techniques. These string similarity matching techniques typically involve generating a string similarity score (e.g., Hamming distance, Jaro-Winkler distance) by comparing strings in entity names extracted from an information source (e.g., unknown names) with strings in entity names stored in a database (e.g., known names). The unknown name is then assigned to the known name with the highest string similarity score. For example, an unknown name can be extracted from the title of a news article. Using string similarity matching techniques, the unknown name can be linked to the known names stored in the database. Therefore, based on the name in the title of the article, it can be determined that the article is about a specific entity.

[0009] However, when used to link an information source to a specific entity, string similarity matching techniques may be problematic. In some aspects, this may be problematic because some information sources do not have a title from which an entity name can be extracted. In other aspects, this may be problematic because some information sources may not include an entity name in its title. In other aspects, this may be problematic because string similarity matching techniques are sensitive to data quality, and the potential relationship of entity names cannot usually be considered. For example, using a metric to calculate string similarity known to those skilled in the art, the unknown entity name "China Eastern Airlines Yunnan Co., Ltd." narrated in the title of a news article can have a string similarity score of 0.70 with "China Airlines", a string similarity score of 0.75 with "China Eastern Airlines" and a string similarity score of 0.76 with "China Yunnan Hotel Company". Therefore, string similarity matching techniques may mistakenly determine that the article provides information about "China Yunnan Hotel Company".

[0010] Therefore, there is a need for a system and method that can automatically, accurately and efficiently perform entity linking. The present disclosure provides a solution for performing entity linking using subgraph matching technology. Summary of the invention

[0011] In one aspect, the present disclosure provides a computer-implemented method for entity linking. The method may include extracting a first attribute set from an information source by an extraction module, and retrieving a second attribute set from a database including known entities by the extraction module. The first attribute set corresponds to an unknown entity and includes a first attribute. The second attribute set corresponds to one of the known entities and includes a second attribute. The method may further include generating an unknown entity graph and a known entity graph by a graph generation module. The unknown entity graph includes a first node corresponding to the first attribute, and the known entity graph includes a second node corresponding to the second attribute. The method may further include: generating an unknown entity graph embedding by a graph neural network model by applying the unknown entity graph to the graph neural network model; generating a known entity graph embedding by the graph neural network model by applying the known entity graph to the graph neural network model; and generating an embedding space similarity score based on the unknown entity graph embedding and the known entity graph embedding by an embedding similarity module. The information source is assigned to one of the known entities based on the embedding space similarity score by a recommendation module.

[0012] In one aspect, the present disclosure provides a computer-implemented method for entity linking. The method may include extracting a first attribute set from an information source by an extraction module, and retrieving a second attribute set from a database including known entities by the extraction module. The first attribute set corresponds to an unknown entity and includes a first attribute. The second attribute set corresponds to one of the known entities and includes a second attribute. The method may further include generating an unknown entity graph and a known entity graph by a graph generation module. The unknown entity graph includes a first node corresponding to the first attribute, and the known entity graph includes a second node corresponding to the second attribute. The method may further include: generating an unknown entity graph embedding by a graph neural network model by applying the unknown entity graph to the graph neural network model; generating a known entity graph embedding by the graph neural network model by applying the known entity graph to the graph neural network model; and generating an embedding space similarity score based on the unknown entity graph embedding and the known entity graph embedding by an embedding similarity module.

[0013] The above method may further include identifying attribute pairs by a string similarity module, and generating string similarity scores corresponding to the attribute pairs by the string similarity module. Each of the attribute pairs includes a first attribute of the first attribute and a corresponding second attribute of the second attribute. Each of the string similarity scores is based on an attribute pair of the attribute pairs. The method may further include determining an overall similarity score based on the embedding space similarity score and the string similarity score by a recommendation module. The information source is assigned to one of the known entities by the recommendation module based on the overall similarity score.

[0014] In one aspect, the present disclosure provides an entity linking system. The entity linking system may include a graph generation module, a graph neural network, an embedding space similarity module, and a recommendation module. The graph generation module may be configured to generate an unknown entity graph and a known entity graph. The unknown entity graph is based on a first attribute set including a first attribute, wherein the first attribute is extracted from an information source, wherein the first attribute set corresponds to an unknown entity, and wherein the unknown entity graph includes a first node corresponding to the first attribute. The known entity graph is based on a second attribute set including a second attribute, wherein the second attribute is retrieved from a database including known entities, wherein the second attribute set corresponds to a known entity in the known entities, and wherein the known entity graph includes a second node corresponding to the second attribute. The graph neural network may be configured to generate an unknown entity graph embedding based on the unknown entity graph, and to generate a known entity graph embedding based on the known entity graph. The embedding space similarity module may be configured to generate an embedding space similarity score based on the unknown entity graph embedding and the known entity graph embedding. The recommendation module may be configured to assign an information source to a known entity in the known entities based on the embedding space similarity score. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] According to the following description in conjunction with the accompanying drawings, various features and advantages of the embodiments described herein can be understood as follows:

[0016] Figure 1 is a diagram illustrating an entity linking system according to at least one aspect of the present disclosure.

[0017] Figures 2A-2C is an example entity graph that can be applied to a graph neural network for entity linking using subgraph matching in accordance with aspects of the present disclosure.

[0018] Figure 3 According to at least one aspect of the present disclosure, Figure 1 Flowchart of a method for entity linking performed by an entity linking system.

[0019] Figure 4A-4B is a logic flow diagram of a method for entity linking using subgraph matching according to aspects of the present disclosure.

[0020] Figure 5 A block diagram of a computer device according to at least one aspect of the present disclosure is shown.

[0021] Figure 6 A block diagram of a system including a host according to at least one aspect of the present disclosure is shown.

[0022] Corresponding reference numerals indicate corresponding parts throughout the several views.The exemplifications set forth herein illustrate various aspects of the disclosure in one form, and such exemplifications should not be construed as limiting the scope of the disclosure in any way. DETAILED DESCRIPTION

[0023] The applicant of the present application owns the following international patent applications, the disclosures of which are incorporated herein by reference in their entirety:

[0024] International patent application No. PCT / US2022 / 077290, filed on September 29, 2022, and titled “Entity Linking Using Graph Neural Networks.”

[0025] Before explaining various forms of entity linking using subgraph matching, it should be noted that the illustrative forms disclosed herein are not limited in application or use to the details of construction and arrangement of the components shown in the drawings and description. The illustrative forms may be implemented or incorporated in other forms, variations and modifications, and may be practiced or implemented in various ways. In addition, unless otherwise indicated, the terms and expressions employed herein are selected for the purpose of describing the illustrative forms for the convenience of the reader, and not for the purpose of limiting them.

[0026] As used herein, the term "computing device" or "computer device" may refer to one or more electronic devices configured to communicate directly or indirectly with or on one or more networks. The computing device may be a mobile device, a desktop computer, etc. In addition, the term "computer" may refer to any computing device that includes necessary components for sending, receiving, processing and / or outputting data and typically includes a display device, a processor, a memory, an input device, a network interface, etc.

[0027] As used herein, the term "server" may include one or more computing devices, which may be individual stand-alone machines located in the same or different locations, may be owned or operated by the same or different entities, and may also be one or more clusters of distributed computers or "virtual" machines housed in a data center. It should be understood and appreciated by those skilled in the art that the functions performed by one "server" may be spread across multiple different computing devices for various reasons. As used herein, "server" is intended to refer to all such scenarios and should not be interpreted as or limited to a specific configuration. The term "server" may also refer to or include one or more processors or computers, storage devices, or similar computer arrangements that operate or facilitate communication and processing by multiple parties in a network environment such as the Internet, but it should be understood that communication may be facilitated by one or more public or private network environments, and various other arrangements are possible. In addition, multiple computers, such as servers or other computerized devices, that communicate directly or indirectly in a network environment may constitute a "system". As used herein, reference to a "server" or "processor" may refer to a previously described server and / or processor, different servers and / or processors, and / or a combination of servers and / or processors that are described as performing steps or functions.

[0028] As used herein, the term "system" may refer to one or more computing devices or a combination of computing devices (e.g., processors, servers, client devices, software applications, modules, components of these computing devices, etc.). For example, a system may include multiple computing devices that include a software application, where the multiple computing devices are connected via a network.

[0029] As used herein, the term "module" may refer to a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution.

[0030] As used herein, the term "entity" may refer to or include an individual, a corporation, a business-related organization, a nonprofit organization, a governmental organization, a charitable organization, an educational institution, or any other type of individual, group of individuals, or organization.

[0031] As used herein, the term "word" may refer to a string of characters. For example, a word may refer to a string of characters that are not separated by spaces. The string of characters may include one or more characters. The one or more characters may include letters, numbers, and / or symbols.

[0032] As used herein, the term "name" may refer to a word or a word string. For example, a name may be a word or a word string that refers to an entity.

[0033] As used herein, the term "known" when used to refer to a name and / or an entity (e.g., a known name, a known entity name, a known entity) may mean that the name and / or entity has been assigned or otherwise associated with a particular unique entity. For example, a database storing a set of known names may be used to designate each known name as referring to a particular entity.

[0034] As used herein, the term "unknown" when used to refer to a name and / or entity (e.g., an unknown name, an unknown entity name, an unknown entity) may mean any name and / or entity mentioned or otherwise described in an information source. An unknown name, an unknown entity name, and / or an unknown entity may or may not have been assigned to a known name or known entity. As an example, a name that is the target of an entity linking process for assigning a name to a specific known entity and / or a specific known name may be referred to as an unknown name. As another example, a name may have been extracted from an information source and assigned to a known name via an entity linking process. Even if a name has been assigned to a known name, it may be referred to as an unknown name. Similarly, an entity described by a name may still be referred to as an unknown entity.

[0035] As used herein, "attributes" or "entity attributes" of an entity may refer to any type of information that can be used to characterize and / or describe an entity. For example, attributes of an entity may include a name (e.g., legal name, stock symbol, dba, alias, nickname), biographical information, organizational information (e.g., registration status, legal entity type, such as corporation, LLC, LLP, non-profit, etc.), financial information, industry classification, industry sub-classification, parent company information, subsidiary information, competitor information, customer information, supplier information, personnel information (e.g., number of employees, names of executives, names of board members), and / or geographic information (e.g., headquarters location, office locations, market locations, other locations where the entity is known to conduct business).

[0036] Entity linking is fundamental to any organization's effective use of data outside the organization. Given the heterogeneity of data and the environmental quality of external data, it is technically quite challenging for organizations to utilize external data. According to the present disclosure, machine learning is used to effectively utilize external data. In one aspect, as described in more detail below, machine learning techniques are used to extract structural information from external name entities based on various factors such as geography, industry, etc. In one aspect, the machine learning techniques include machine learning models to learn when matching links should or should not be in the embedding space. The following description provides a technical solution for organizations to utilize external data.

[0037] The present disclosure provides a solution that can perform entity linking using subgraph matching, such as assigning information sources (eg, online news articles) to specific known entities using subgraph matching. Performing entity linking using subgraph matching can provide various technical benefits. For example, the systems and methods disclosed herein may allow a computer to perform entity linking more accurately and efficiently in an unconventional manner by: (i) extracting a first attribute set from an information source by an extraction module, wherein the first attribute set corresponds to an unknown entity and includes a first attribute; (ii) retrieving a second attribute set from a database including information related to known entities by the extraction module, wherein the second attribute set corresponds to one of the known entities and includes a second attribute; (iii) generating an unknown entity graph including a first node corresponding to the first attribute by a graph generation module; (iv) generating a known entity graph including a second node corresponding to the second attribute by the graph generation module; (v) generating an unknown entity graph embedding by applying the unknown entity graph to a graph neural network model; (vi) generating a known entity graph embedding by applying the known entity graph to a graph neural network model; (vii) generating an embedding space similarity score by an embedding similarity module based on the unknown entity graph embedding and the known entity graph embedding; and (viii) assigning the information source to one of the known entities based on the embedding space similarity score by a recommendation module.

[0038] As another example, in addition to the above elements (i)-(vii), the systems and methods disclosed herein may also allow a computer to perform entity linking more accurately and efficiently in an unconventional manner by: (ix) identifying attribute pairs by a string similarity module, wherein each of the attribute pairs includes one of the first attributes and a corresponding one of the second attributes; (x) generating string similarity scores by the string similarity module, wherein each of the string similarity scores corresponds to one of the attribute pairs in the attribute pairs; (xi) determining, by a recommendation module, an overall similarity score based on the embedding space similarity score and the string similarity score; and (xii) assigning, by the recommendation module, an information source to one of the known entities based on the overall similarity score.

[0039] In addition, the systems and methods disclosed herein integrate the above elements (i)-(vii) and (i)-(xi) into practical applications by: (viii) assigning an information source to one of the known entities based on an embedding space similarity score; and / or (xii) assigning an information source to one of the known entities based on an overall similarity score.

[0040] Additionally, the systems and methods disclosed herein may allow one or more of potentially millions of various private and / or public information sources (e.g., news articles, wikis, social media, online databases) that provide information about a particular entity to be automatically assigned to that particular entity, thereby performing entity linking at a scale that is not practical to perform in the human mind.

[0041] Figure 1 1 is a diagram 100 illustrating an entity linking system 160 according to at least one non-limiting aspect of the present disclosure. The entity linking system 160 may include various modules, such as an extraction module 162, a graph generation module 164, a graph neural network 166 (GNN), an embedding similarity module 168, a string similarity module 170, a recommendation module 172, a natural language processing (NLP) module 174, and / or a training module 176. Although the modules of the entity linking system 160 are described below as performing various functions individually, any module may be configured to perform any combination of the functions described herein. Similarly, multiple modules may be combined into a single module to perform any combination of the functions described herein, and / or a single module may be split into multiple sub-modules, each of which performs any function described herein.

[0042] The entity linking system 160 is configured to access the information source 110 via the network 120 1 , 110 2 , 110 3 ,……,110 n (collectively referred to as information sources 110) or otherwise communicate with the information sources. Network 120 may include any type of wired, long-range wireless and / or short-range wireless network. For example, network 120 may include an internal network, a local area network (LAN), Wi-Fi, a cellular network, a private network, the Internet, a cloud computing network, and / or a combination of these or other types of networks. Information sources 110 may include any type of information source and combination of information sources, including text-based data, image data, video data, and / or other multimedia-based data accessible via network 120. For example, information sources 110 may include various private and public information sources, such as news articles, wikis, social media, and / or other databases and publications accessible via network 120 (e.g., the Internet).

[0043] The entity linking system 160 is further configured to access or otherwise communicate with the known entity database 150 via the network 120. Figure 1 In non-limiting aspects of the invention, the known entity database 150 is separate from the entity linking system 160. In other aspects, the known entity database 150 may be included as part of the entity linking system 160 (e.g., stored on the same server or combination of servers as the entity linking system 160). The known entity database 150 may include data related to a plurality of known entities. For example, the known entity database 150 may include a list of names of known entities. As another example, the known entity database 150 may include various attributes of each known entity, such as entity name, biographical information, organizational information, financial information, industry classification, industry segment classification, parent company information, subsidiary information, competitor information, customer information, supplier information, personnel information, and / or geographic information.

[0044] The extraction module 162 of the entity linking system 160 can be configured to extract information from the information source 110 and / or the known entity database 150. For example, in one aspect, the extraction module 162 can be configured to detect, classify and / or extract attributes related to the unknown entity from text, images, videos and / or any other type of multimedia data included in the information source 110. The entity attributes extracted from the information source 110 can be similar to the entity attributes of that type stored in the known entity database. For example, the extraction module 162 can be configured to extract attributes such as entity name, biographical information, organizational information, financial information, industry classification, industry sub-classification, parent company information, subsidiary information, competitor information, customer information, supplier information, personnel information and / or geographic information from the information source 110. The collection of attributes of the unknown entity extracted from one information source 110 (e.g., a single news article) is sometimes referred to as an "attribute set".

[0045] The extraction module 162 can adopt various techniques to detect, classify and / or extract entity attributes from text-based data, such as rule-based named entity recognition (NER) techniques (e.g., techniques adopted by the General Architecture for Text Engineering (GATE) and a rule-based NER called DrNER, etc.) and / or machine learning-based NER techniques (e.g., techniques adopted by the OpenNLP named entity recognizer and name finder, the free open source library for natural language processing in Python spaCy, and the named entity recognizer SemiNER, etc.).

[0046] The extraction module 162 can adopt various technologies to detect, classify and / or extract entity attributes from multimedia-based data (e.g., videos, images), such as object recognition technology and image recognition technology using computer vision, machine learning (e.g., technology using support vector machine (SVM), feature bag model, Viola-Jones algorithm, etc.), and deep learning (e.g., technology using only look once (YOLO), single shot detector (SSD) and convolutional neural network (CNN) model).

[0047] In some aspects, the extraction module 162 may be configured to detect, classify, extract, and / or retrieve entity attributes stored in the known entity database 150. For example, the extraction module 162 may be configured to retrieve a set of attributes for each known entity in the known entity database 150. Each set of attributes for a particular known entity from the known entity database 150 is sometimes referred to as an “attribute set.”

[0048] The graph generation module 164 of the entity linking system 160 can be configured to generate an entity graph based on the attribute sets extracted and / or retrieved by the extraction module 162. Each entity graph generated by the graph generation module 164 can generally include nodes corresponding to attributes in one of the attribute sets. For example, in one aspect, the graph generation module 164 can be configured to generate an unknown entity graph having nodes corresponding to attributes in the attribute set of unknown entities extracted from the information source 110. In another aspect, the graph generation module 164 can be configured to generate a known entity graph, wherein each of the known entity graphs respectively has nodes corresponding to attributes in one of the attribute sets of known entities retrieved from the known entity database 150.

[0049] Figures 2A-2C According to some aspects of the present disclosure, Figure 1 An example entity graph generated by the graph generation module 164 is shown in FIG. Figure 2A , entity graph 200A is shown as having nodes 204. Each of nodes 204 corresponds to an attribute in attribute set 202 (eg, attribute set includes attribute 1 202 1 ,property 2 202 2 ,property 3 202 3 ,property 4 202 4 ,property 5 202 5 ,property 6 202 6 ,property 7 202 7 ,property8 202 8 , ... and attributes n 202 n ). The entity graph 200A may be an unknown entity graph or a known entity graph. Therefore, the attributes 202 may correspond to the attributes of an unknown entity extracted from the information source 110, or to the attributes of a known entity retrieved from the known entity database 150. In addition, any number (e.g., any positive integer) of attributes may be included in the attribute set. Therefore, the entity graph 200A may include any number (e.g., any positive integer) of nodes.

[0050] Still refer to Figure 2A , the structure of the entity graph 200A is determined by the placement of edges 206. Typically, each node 204 is connected to at least one other node 204 via an edge 206. Some nodes 204 may be connected to multiple other nodes 204 via multiple edges. For example, corresponding to an attribute 2 202 2 The node 204 is connected only via the edge 206 to the node corresponding to the attribute 1 202 1 Node 204. In contrast, corresponding to the attribute 1 202 1 The node 204 is connected via edge 206 to the node corresponding to the attribute 2 202 2 ,property 3 202 3 ,property 4 202 4 ,property 7 202 7 ,property 8 202 8 and properties n 202 n Node 204. Although Figure 2A One particular entity graph structure is depicted in the non-limiting aspects of FIG. 2 , but the entity graph 200A may have any structure in which each node 204 is connected to one or more of any other nodes 204 via edges 206 .

[0051] Main references Figure 2A And also refer to Figure 1 In some aspects, the graph generation module 164 can be configured to determine the structure of the entity graph 200A based on the types of attributes included in the attribute set. Specifically, including different types of attributes in the attribute set can enable the graph generation module 164 to implement a specific organization of the nodes 204 and the edges 206. For example, the attributes 1 202 1The various nodes 204 corresponding to the various other attributes 202 may be connected to the nodes corresponding to the attributes via edges 206 in a spoked configuration. 1 202 1 As another example, the attribute 4 202- 4 Can be the industry of an unknown entity (e.g., airline industry). In addition, the attribute 5 202- 5 and properties 6 202- 6 It can be a subdivision within an industry (e.g., low cost, regional). 5 202- 5 and properties 6 202- 6 The node 204 can be connected via an edge 206 to the node corresponding to the attribute 4 202 4 Node.

[0052] Now the main reference Figure 2B And also refer to Figure 1 , the illustrative unknown entity graph 200B is populated with attributes 202 extracted from the information source 110 1-n . In this example, information source 110 is an online news article about the airline Vueling. The headline of the online news article describes Vueling as a "Spanish low-cost airline" and states that Vueling is canceling "February flights to Ukraine." The text of the article further describes Vueling as "IAG's Spanish low-cost airline" and reports that Vueling has canceled eight flights from Paris to Kiev due to tensions between Russia and Ukraine. The text of the article also explains that Vueling is the only Spanish airline with "direct flights to the Ukrainian capital" from Paris, while "IAG's other airlines, including British Airways and Iberia, do not fly directly to the Ukrainian capital." In addition, the article includes an image of travelers standing in line at an airport ticket counter.

[0053] Still the main reference Figure 2B And also refer to Figure 1 Based on the text and multimedia data included in the descriptive online article about Vueling, the extraction module 162 can be configured to detect, classify and extract the attribute set including: "Vueling" (entity name) 202 1 , “IAG” (parent entity) 202 2 , "Airline" (Industry) 202 3, "Low Cost" (Industry Segment) 202 4 , "British Airways" (co-reality) 202 5 , "Russia" (geographic location) 202 6 , "Paris" (flight origin) 202 7 , "Airport" (classified multimedia scene) 202 8 , "Ukraine" (flight destination) 202 9 , ..., and "Spain" (geographic location) 202 n The graph generation module 164 may be configured to generate a graph based on the extracted attributes 202 1-n An unknown entity graph 200B is generated, wherein the unknown entity graph 200B includes nodes 204 connected by edges 206, and wherein each node 204 corresponds to an extracted attribute 202 1-n An attribute in .

[0054] Now the main reference Figure 2C And also refer to Figure 1 , the illustrative known entity graph 200C is populated with attributes 202 retrieved from the known entity database 150 1-n In this example, the known entity database 150 includes attributes of the entity "Vueling Airlines". Therefore, the extraction module 162 can be configured to detect, classify, extract and / or retrieve the attribute set of the entity "Vueling Airlines" including the following: "Vueling Airlines" (sub-entity name) 202 1 , "Travel" (Industry) 202- 2 , "IAG" (parent entity name) 202 3 , "Aer Lingus" (fruiting body name) 202 4 , "British Airways" (sub-entity name) 202 5 , "Airlines" (industry segment) 202 6 , "France" (geographic location) 202 7 , "Spain" (geographic location) 202 8 , ..., and "Italy" (geographic location) 202 n The graph generation module 164 may be configured to generate a known entity graph 200C based on the retrieved attributes 202 , wherein the known entity graph 200C includes nodes 204 connected by edges 206 , and wherein each node corresponds to one of the extracted attributes 202 .

[0055] Reference again Figure 1, various combinations of the GNN 166, embedding similarity module 168, string similarity module 170, and recommendation module 172 of the entity linking system 160 can be configured to perform entity linking based on the unknown entity graph and the known entity graph generated by the graph generation module 164. To illustrate this feature, Figure 3 According to at least one aspect of the present disclosure, a Figure 1 Flowchart 300 of a method for entity linking performed by the entity linking system 160 of the embodiment of the present invention. Although flowchart 300 depicts a method for entity linking based on an attribute set of an unknown entity and an attribute set of a single known entity, a person of ordinary skill in the art should understand that the method for entity linking is performed by comparing the attribute set of the unknown entity with the attribute sets of multiple known entities (e.g., by repeating the method depicted in flowchart 300 for multiple known entities).

[0056] Main references Figure 3 And also refer to Figure 1 , flowchart 300 depicts an unknown entity graph 200D generated by graph generation module 164 and a known entity graph 200E generated by graph generation module 164. Unknown entity graph 200D may be similar to entity graph 200A and / or unknown entity graph 200B. Thus, unknown entity graph 200D may include attributes (e.g., U) corresponding to an unknown entity (e.g., U) extracted by extraction module 162 from an information source 110. 1 , U 2 , U 3 , U 4 , U 5 ,……,U n ) nodes. Similarly, known entity graph 200E may be similar to entity graph 200A and / or known entity graph 200C. Thus, known entity graph 200E may include attributes (e.g., K) corresponding to known entities (e.g., K) retrieved by extraction module 162 from known entity database 150. 1 , K 2 , K 3 , K 4 , K 5 ,……,K n ) node.

[0057] Still reference Figure 1 and Figure 3, GNN 166 can be configured to generate unknown entity graph embedding 302 (e.g., Vec(U)) based on unknown entity graph 200D, and generate known entity graph embedding 304 (e.g., Vec(K)) based on known entity graph 200E. In addition, as explained in detail below, GNN 166 can be trained so that unknown entity graph embedding 302 and known entity graph embedding 304 generated from entity graphs including attributes of the same entity will have similar representations in the embedding space. GNN 166 can be any type of GNN, such as a graph convolutional network (GCN).

[0058] Still reference Figure 1 and Figure 3 , the embedding similarity module 168 can be configured to generate an embedding space similarity score 306 (e.g., g(Vec(U), Vec(K))) based on the unknown graph embedding 302 and the known entity graph embedding 304. The embedding similarity module 168 can use various techniques to generate the embedding space similarity score 306. For example, the embedding similarity module 168 can be a deep neural network (DNN) model, such as a multi-layer DNN. As explained in detail below, the DNN can be trained to generate an embedding space similarity score 306 corresponding to the degree of similarity between the attribute set of the unknown entity and the attribute set of the known entity (based on the similarity of the graph embedding 302 and the known entity graph embedding 304). In addition, as described above, the method depicted in the flowchart 300 can be repeated for multiple known entities. Therefore, multiple known entity graphs 200E, multiple known entity graph embeddings 304, and multiple embedding space similarity scores 306 can be generated, wherein each of the embedding space similarity scores 306 is based on one of the known entity graph embeddings 302 and the known entity graph embeddings 304.

[0059] In some aspects, the embedding similarity module 168 can be configured to assign the information source 110 (e.g., an online news article) from which the attribute of the unknown entity is extracted to one of the known entities based on the embedding space similarity score 306. For example, in one aspect, the embedding similarity module 168 can assign the information source 110 to one of the known entities based on the unknown entity graph embedding 302 / known entity graph embedding 304 pair with the highest embedding space similarity score 306. In addition to or in lieu of the foregoing, the embedding similarity module 168 can assign the information source 110 to one of the known entities if the corresponding unknown entity graph embedding 302 / known entity graph embedding 304 pair has an embedding space similarity score 306 that satisfies a predetermined threshold, such as an embedding space similarity score 306 of not less than 0.7, 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or not less than 0.99. If none of the embedding space similarity scores 306 satisfy a predetermined threshold, the embedding similarity module 168 may not assign the information source 110 to any known entity.

[0060] In other aspects, the recommendation module 172 of the entity linking system 160 can be configured to determine an overall similarity score 310. Furthermore, the recommendation module 172 can be configured to assign the information source 110 (e.g., an online news article) from which the attribute of the unknown entity was extracted to one of the known entities based on the overall similarity score 310. Each of the overall similarity scores 310 can be based on one of the embedding space similarity scores 306 and other metrics, such as a set of string similarity scores 308, as further explained below.

[0061] Still refer to Figure 1 and Figure 3 , and return to Figure 3 In a non-limiting aspect of the present invention, wherein the attribute set of the unknown entity is compared to the attribute set of a single known entity, the string similarity module 170 can be configured to generate a string similarity score set 308 (e.g., f(U)) based on the attribute set of the unknown entity and the attribute set of the known entity. 1 ,K 1 )、f(U 2 ,K 2 )、……、f(U n ,K n )). To generate the string similarity score set 308, the string similarity module 170 may be configured to identify attribute pairs (e.g., (U 1 ,K 1 )、(U 2 ,K 2 ),……,(U n ,Kn )), where each of the attribute pairs includes one of the attributes of the unknown entity and one of the attributes of the known entity.

[0062] In some aspects, the string similarity module 170 may identify attribute pairs based on the attribute types included in the attribute set of the unknown entity and the attribute types included in the attribute set of the known entity. For example, the attribute set of the unknown entity may include attributes of the entity name (e.g., 1 ), and the attribute set of a known entity may include an attribute of the entity name (e.g., K 1 ). Therefore, the attribute U 1 and K 1 As another example, the attribute set of the unknown entity may include attributes of industry classification (e.g., U 2 ), and the attribute set of known entities may include attributes of industry classification (e.g., K 2 ). Therefore, the attribute U 2 and K 2 Can be identified as attribute pairs.

[0063] Each of the attributes of the known entity and each of the attributes of the unknown entity may include one or more terms. Therefore, based on the terms included in each attribute pair, the string similarity module 170 can be configured to generate a string similarity score (e.g., Hamming distance, Jerod-Winkler distance). The set of attribute pairs corresponding to the attribute set of the unknown entity and the attribute set of a particular known entity is sometimes referred to herein as an "attribute pair set." Similarly, the set of string similarity scores corresponding to a particular attribute pair set is sometimes referred to herein as a "string similarity score set" (e.g., string similarity score set 308).

[0064] Still reference Figure 1 and Figure 3 , the recommendation module 172 can be configured to determine an overall similarity score 310 (e.g., Sim(U, K)) based on the embedding space similarity score 306 and the string similarity score set 308. The recommendation module 172 can use various techniques to determine the overall similarity score 310. For example, the recommendation module 172 can use a regression model and / or a classification model that is configured to be based on training parameters corresponding to the embedding space similarity score 306 (e.g., parameter α 0 ) and a training parameter (e.g., parameter α) corresponding to each of the string similarity scores in the string similarity score set 308. 1 , α 2 ,……,α n ) and calculate the overall similarity score 310. Therefore, the overall similarity score 310 can be calculated based on the following equation:

[0065] Sim(U,K)=α 0 g(Vec(U),Vec(K))+α 1 f(U 1 ,K 1 )+α 2 f(U 2 ,K 2 )+...+α n f(U n ,K n )

[0066] As explained in detail below, the parameters of the regression model and / or classification model employed by the recommendation module 172 (e.g., α 0 , α 1 , α 2 ,……,α n ), such that the overall similarity score 310 corresponds to how similar the set of attributes of the unknown entity and the set of attributes of the known entity are to each other (based on the embedding space similarity score 306 and based on the set of string similarity scores 308).

[0067] As described above, the method depicted in flowchart 300 can be repeated for multiple known entities. Thus, multiple embedding space similarity scores 306, multiple string similarity score sets 308, and multiple overall similarity scores 310 can be generated, wherein each of the overall similarity scores 310 is based on an attribute of the unknown entity and an attribute of one of the known entities. Thus, also as described above, recommendation module 172 can be configured to assign the information source 110 (e.g., an online news article) from which the attribute of the unknown entity is extracted to one of the known entities based on the overall similarity score 310. For example, in one aspect, recommendation module 172 can assign information source 110 to the known entity corresponding to the highest overall similarity score 306. In addition to or in lieu of the foregoing, if one of the known entities has a corresponding overall similarity score 306 that satisfies a predetermined threshold, such as an overall similarity score 310 of not less than 0.7, 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or not less than 0.99, then the recommendation module 172 may assign the information source 110 to the known entity. If none of the overall similarity scores 310 satisfies the predetermined threshold, then the recommendation module 172 may not assign the information source 110 to any of the known entities.

[0068] Thus, it can be appreciated from various aspects of the present disclosure that the entity linking system 160 can be configured to perform entity linking based on the embedding space similarity score 306 without generating the string similarity score set 308. In other aspects, the entity linking system 160 can be configured to perform entity linking based on both the embedding space similarity score 306 and the string similarity score set 308 by determining the overall similarity score 310. Both approaches can provide various technical benefits. For example, both approaches can accurately and efficiently link information sources (e.g., online news articles) to specific entities in an unconventional manner by: generating an unknown entity graph including nodes corresponding to attributes of an unknown entity; generating a known entity graph, wherein each of the known entity graphs includes a node corresponding to an attribute of one of the known entities; generating an unknown entity graph embedding by applying the unknown entity graph to a graph neural network; generating a known entity graph embedding by applying each of the known entity graphs to a graph neural network; and generating an embedding space similarity score, wherein each of the embedding space similarity scores is based on an unknown entity graph embedding and one of the known entity graph embeddings. The latter method can provide an unconventional and robust method for entity linking by additionally performing the following operations: identifying a set of attribute pairs; generating a set of string similarity scores corresponding to the set of attribute pairs; and determining an overall similarity score based on the embedding space similarity score and the set of string similarity scores.

[0069] Reference again Figure 1 , various techniques may be used to train the GNN 166, the embedding similarity module 168 (e.g., a DNN), and / or the recommendation module 172 (e.g., a regression model, a classification model, etc.). In some aspects, training the GNN 166 may include initializing nodes of the unknown entity graph embedding and nodes of the known entity graph embedding using various NLP models. Thus, the entity linking system 160 may include an NLP module 174 configured to perform this embedding initialization. For example, the NLP module 140 may employ an NLP model such as an encoder representation of a bidirectional transformer (BERT).

[0070] Still refer to Figure 1As well as various techniques that can be used to train GNN 166, embedding similarity module 168 and / or recommendation module 172, in some aspects, entity linking system 160 may include training module 176 to train various parameters of these modules and / or neural networks using labeled data. For example, GNN 166 may be a GCN. Training module 176 may be configured to train parameters (e.g., kernel) of GCN based on labeled data (e.g., an entity graph labeled as a specific known entity). As another example, embedding similarity module 168 may adopt DNN. Training module 176 may be configured to train parameters (e.g., weights of each layer) of DNN based on labeled data. As another example, recommendation module 172 may adopt a regression model and / or a classification model. Training module 176 may be configured to train parameters of a regression model and / or a classification model based on labeled data. In some aspects, training module 176 may be used to train parameters of GNN 166, embedding similarity module 168 and recommendation module 172 end-to-end (e.g., based on a set of attributes labeled as a specific known entity at the same time). In other aspects, the training module 176 may be used to train any of the GNN 166, the embedding similarity module 168, and / or the recommendation module 172 in separate stages.

[0071] Figure 4A-4B 4 is a logical flow diagram of a method 400 for entity linking using subgraph matching according to several aspects of the present disclosure. The method 400 may be as described above with respect to Figure 1 The described entity linking system 160 and / or any combination of components of the entity linking system 160 may be practiced. Figure 1 , 3 4A, according to method 400, extraction module 162 extracts 402 a first attribute set from information source 110. The first attribute set corresponds to an unknown entity and includes a first attribute. Additionally, graph generation module 164 generates 404 an unknown entity graph, such as unknown entity graph 200D. Unknown entity graph 200D includes a first node corresponding to the first attribute. GNN 166 generates 406 unknown entity graph embedding 302 by applying unknown entity graph 200D to a graph neural network.

[0072] Still reference Figure 1 , 34A, according to method 400, the extraction module 162 retrieves 408 a second attribute set from one of the known entity databases 150 including information related to the known entity. Each of the second attribute sets corresponds to one of the known entities and includes a second attribute. In addition, the graph generation module 164 generates 410 a known entity graph 200E. Each of the known entity graphs 200E includes a second node corresponding to a second attribute in one of the second attribute sets. The GNN 166 generates 412 a known entity graph embedding 304 by applying the known entity graph 200E to a graph neural network.

[0073] Still refer to Figure 1 , 3 According to method 400, embedding similarity module 168 generates 414 embedding space similarity scores 306. Each of embedding space similarity scores 306 is based on unknown entity graph embedding 320 and one of known entity graph embeddings 304. In one aspect of method 400, embedding space similarity scores 306 include a highest embedding space similarity score. Recommendation module 172 assigns 416A information source 110 to one of known entities based on the highest embedding space similarity score.

[0074] Reference now Figure 1 , 3 4B, according to another aspect of method 400, identifying 416B a set of attribute pairs. Each of the attribute pair sets corresponds to a second attribute set in the first attribute set and the second attribute set. In addition, each of the attribute pair sets includes an attribute pair, and each of the attribute pairs includes a first attribute in the first attribute and a corresponding second attribute in the second attribute. The string similarity module 170 generates 418 a set of string similarity scores 308 corresponding to the attribute pair sets. Each of the string similarity score sets 308 includes a string similarity score. Each of the string similarity scores is based on an attribute pair in the attribute pair. The recommendation module 172 determines 420 an overall similarity score 310. Each of the overall similarity scores 310 corresponds to a known entity in the known entities and is based on an embedding space similarity score in the embedding space similarity score 306 and a string similarity score set in the string similarity score set 308. In addition, the overall similarity score 310 includes a highest overall similarity score. In this aspect of the method 400 , the recommendation module 172 assigns 422 the information source 110 to one of the known entities based on the highest overall string similarity score.

[0075] References in this article Figure 1The described systems and modules can operate on one or more computer devices to facilitate the functions described herein. In addition, one or more computer devices can use any suitable number of subsystems to facilitate the functions described herein. For example, Figure 5 is a block diagram of a computer device 3000 having data processing subsystems or components according to at least one aspect of the present disclosure. Figure 5 The subsystems shown in the figure are interconnected via a system bus 3010. Additional subsystems such as a printer 3018, a keyboard 3026, a fixed disk 3028 (or other memory including computer readable media), a monitor 3022 coupled to a display adapter 3020, and the like are shown. Peripheral devices and input / output (I / O) devices coupled to an I / O controller 3012 (which may be a processor or any suitable controller) may be connected to the computer system by any number of means known in the art (e.g., a serial port 3024). For example, a serial port 3024 or an external interface 3030 may be used to connect a computer device to a wide area network (e.g., the Internet), a mouse input device, or a scanner. Interconnection via a system bus allows a central processor 3016 to communicate with each subsystem and allows control of the execution of instructions from the system memory 3014 or the fixed disk 3028 and the exchange of information between subsystems. The system memory 3014 and / or the fixed disk 3028 may embody computer readable media.

[0076] Figure 6 4000 is a schematic diagram of an exemplary system 4000 including a host 4002 according to at least one aspect of the present disclosure, in which an instruction set for performing any one or more methods discussed herein can be executed. In various aspects, the host 4002 operates as a standalone device or can be connected (e.g., networked) to other machines. In a network deployment, the host 4002 can operate as a server or client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The host 3002 can be a computer or computing device, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a portable music player (e.g., a portable hard disk audio device, such as a Moving Picture Experts Group Audio Layer 3 (MP3) player), a network appliance, a network router, a switch or a bridge, or any machine capable of executing an instruction set (sequentially or otherwise) specifying the actions to be taken by the machine. Further, while a single machine is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.

[0077] The example system 4000 includes a host 4002, running a host operating system 4004 (OS) on processor(s) / processor core(s) 4006 (e.g., central processing unit (CPU), graphics processing unit (GPU), or both) and various memory nodes in a host memory node 4008. The host OS 4004 may include a hypervisor 4010 capable of controlling functions and / or communicating with a virtual machine (“VM”) 4012 running on a machine-readable medium. The VM 4012 may also include a virtual CPU or vCPU 4014. The memory node 4008 may be linked or pinned to a virtual memory node or vNode 4016. When a memory node 4008 is linked or pinned to a corresponding vNode 4016, data may then be mapped directly from the memory node 4008 to its corresponding vNode 4016.

[0078] All the various components shown in the host 4002 can be connected to and connected to each other, or communicate with each other via a bus (not shown) or via other coupling or communication channels or mechanisms. The host 4002 may also include a video display, an audio device or other peripheral device 4018 (e.g., a liquid crystal display (LCD), (one or more) alphanumeric input devices (including, for example, a keyboard), a cursor control device (e.g., a mouse), a voice recognition or biometric verification unit, an external drive, a signal generating device (e.g., a speaker)), a permanent storage device 4020 (also known as a disk drive unit) and a network interface device 4022. The host 4002 may also include a data encryption module (not shown) for encrypting data. The components provided in the host 4002 are components that are typically present in a computer system that may be suitable for use with aspects of the present disclosure, and are intended to represent a broad category of such computer components known in the art. Therefore, the system 4000 may be a server, a minicomputer, a mainframe computer, or any other computer system. The computer may also include different bus configurations, network platforms, multi-processor platforms, and the like. Various operating systems may be used, including UNIX, LINUX, WINDOWS, QNX ANDROID, IOS, CHROME, TIZEN, and other suitable operating systems.

[0079] The disk drive unit 4024 may also be a solid state drive (SSD), a hard disk drive (HDD), or other drive including a computer or machine readable medium on which is stored one or more sets of instructions and data structures (e.g., data / instructions 4026) embodying or utilizing any one or more of the methods or functions described herein. The data / instructions 4026 may also reside completely or at least partially within the main memory node 4008 and / or within the processor(s) 4006 during execution by the host 4002. The data / instructions 4026 may be further transmitted or received over the network 4028 via a network interface device 4022 utilizing any of a number of well-known transmission protocols (e.g., Hyper Text Transfer Protocol (HTTP)).

[0080] (One or more) processors 4006 and memory nodes 4008 may also include machine-readable media. The term "computer-readable medium" or "machine-readable medium" should be understood to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated cache memory and server) that stores one or more sets of instructions. The term "computer-readable medium" should also be understood to include any medium capable of storing, encoding, or carrying an instruction set for execution by the host 4002 and causing the host 4002 to perform any one or more of the methods of the present application, or any medium capable of storing, encoding, or carrying a data structure utilized by this instruction set or a data structure associated with this instruction set. Therefore, the term "computer-readable medium" should be understood to include (but not limited to) solid-state memory, optical and magnetic media, and carrier signals. This medium may also include (but not limited to) hard disks, floppy disks, flash memory cards, digital video disks, random access memory (RAM), read-only memory (ROM), etc. The exemplary aspects described herein may be implemented in an operating environment that includes software installed on a computer, installed in hardware, or a combination of software and hardware.

[0081] Those skilled in the art will recognize that an Internet service can be configured to provide Internet access to one or more computing devices coupled to the Internet service, and that the computing devices may include one or more processors, buses, memory devices, display devices, input / output devices, etc. In addition, those skilled in the art will appreciate that the Internet service can be coupled to one or more databases, repositories, servers, etc., which can be used to implement any of the various aspects of the present disclosure described herein.

[0082] The computer program instructions may also be loaded onto a computer, server, other programmable data processing device or other apparatus so that a series of operating steps are performed on the computer, other programmable device or other apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0083] For example, a suitable network may include or interface with any one or more of the following: a local intranet, a PAN (personal area network), a LAN (local area network), a WAN (wide area network), a MAN (metropolitan area network), a virtual private network (VPN), a storage area network (SAN), a frame relay connection, an advanced intelligent network (AIN) connection, a synchronous optical network (SONET) connection, a digital T1, T3, E1 or E3 line, a digital data service (DDS) connection, a DSL (digital subscriber line) connection, an Ethernet connection, an ISDN (integrated services digital network) line, a dial-up port (e.g., a V.90, V.34 or V.34bis analog modem connection), a cable modem, an ATM (asynchronous transfer mode) connection, or an FDDI (fiber distributed data interface) or CDDI (copper distributed data interface) connection. In addition, communications may include links to any of a variety of wireless networks, including WAP (Wireless Application Protocol), GPRS (General Packet Radio Service), GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access) or TDMA (Time Division Multiple Access), cellular telephone networks, GPS (Global Positioning System), CDPD (Cellular Digital Packet Data), RIM (Research In Motion Ltd.) duplex paging network, Bluetooth radio, or IEEE 802.11 based radio frequency networks. The network may also include any one or more of the following or interface with any one or more of the following: RS-232 serial connection, IEEE-1394 (FireWire) connection, Fiber Channel connection, IrDA (infrared) port, SCSI (Small Computer System Interface) connection, USB (Universal Serial Bus) connection or other wired or wireless, digital or analog interface or connection, mesh or Network connection.

[0084] Generally speaking, a cloud-based computing environment is a resource that typically combines the computing power of large groups of processors (e.g., within a web server) and / or the storage capacity of large groups of computer memory or storage devices. Systems that provide cloud-based resources may be employed solely by their owners, or such systems may be accessed by external users who deploy applications within the computing infrastructure to gain the benefits of large computing or storage resources.

[0085] For example, a cloud is formed by a network of web servers including multiple computing devices (e.g., host 4002), where each server 4030 (or at least a plurality of them) provides processor and / or storage resources. These servers manage workloads provided by multiple users (e.g., cloud resource clients or other users). Typically, each user's workload requirements for the cloud change in real time, sometimes even greatly. The nature and extent of these changes typically depend on the type of business associated with the user.

[0086] It is noteworthy that any hardware platform suitable for performing the processing described herein is suitable for use with the technology. As used herein, the terms "computer-readable storage medium" and "computer-readable storage medium" refer to any one or more media that participate in providing instructions to the CPU for execution. This medium can take a variety of forms, including (but not limited to) non-volatile media, volatile media, and transmission media. Non-volatile media include (for example) optical or magnetic disks, such as fixed disks. Volatile media include dynamic memory, such as system RAM. Transmission media include coaxial cables, copper wires, and optical fibers, among others, including wires that include one aspect of a bus. Transmission media can also take the form of sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer readable media include, for example, floppy disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, digital video disks (DVDs), any other optical media, any other physical media with markings or patterns of holes, RAM, PROM, EPROM, EEPROM, FLASH EPROM, any other memory chip or data exchange adapter, carrier wave, or any other medium from which a computer can read.

[0087] Various forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to the CPU for execution. The bus carries the data to the system RAM, from which the CPU retrieves the instructions and executes them. The instructions received by the system RAM may optionally be stored on a fixed disk before or after execution by the CPU.

[0088] The computer program code for performing operations on aspects of the present technology can be written in any combination of one or more programming languages, including object-oriented programming languages ​​(e.g., Java, Smalltalk, C++, etc.) and conventional procedural programming languages ​​(e.g., "C" programming language, Go, Python, or other programming languages ​​including assembly language). The program code can be executed entirely on the user's computer, partially on the user's computer; as a stand-alone software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or an external computer can be connected (e.g., using an Internet service provider to connect via the Internet).

[0089] Examples of systems and methods according to various aspects of the present disclosure are provided below in the following numbered clauses. An aspect of any method and / or system may include any one or more than one and any combination of the numbered clauses described below.

[0090] Clause 1. A computer-implemented method, comprising: extracting a first attribute set from an information source by an extraction module, wherein the first attribute set corresponds to an unknown entity and wherein the first attribute set includes a first attribute; generating an unknown entity graph including a first node corresponding to the first attribute by a graph generation module; retrieving a second attribute set from a database including known entities by the extraction module, wherein the second attribute set corresponds to one of the known entities and wherein the second attribute set includes a second attribute; generating a known entity graph including a second node corresponding to the second attribute by the graph generation module; generating an unknown entity graph embedding by a graph neural network model by applying the unknown entity graph to the graph neural network model; generating a known entity graph embedding by the graph neural network model by applying the known entity graph to the graph neural network model; generating an embedding space similarity score by an embedding similarity module based on the unknown entity graph embedding and the known entity graph embedding; and assigning the information source to one of the known entities based on the embedding space similarity score by a recommendation module.

[0091] Clause 2. The method according to clause 1, wherein the graph neural network model is a graph convolutional network model.

[0092] Clause 3. The method of any one of clauses 1 to 2, wherein extracting the first set of attributes from the information source comprises extracting, by the extraction module, the first attributes from at least one of a news article or a web page.

[0093] Clause 4. A method according to any one of clauses 1 to 3, wherein extracting the first set of attributes from the information source includes extracting, by the extraction module, at least one of: the name of the unknown entity; the industry associated with the unknown entity; geographic information related to the unknown entity; the name of an entity that is not the unknown entity; a classification of images associated with the unknown entity; or a classification of videos associated with the unknown entity.

[0094] Clause 5. A method according to any one of clauses 1 to 4, wherein retrieving the second set of attributes from the database includes retrieving, by the extraction module, at least one of: the name of a known entity among the known entities; an industry associated with the known entity among the known entities; geographic information related to the known entity among the known entities; the name of a parent entity of the known entity among the known entities; the name of a child entity of the known entity among the known entities; or the name of an entity known to transact business with the known entity among the known entities.

[0095] Clause 6. A method according to any one of clauses 1 to 5, wherein generating the embedding space similarity score comprises: applying the unknown entity graph embedding and the known entity graph embedding to a deep neural network model by the embedding similarity module.

[0096] Clause 7. A computer-implemented method, comprising: extracting a first attribute set from an information source by an extraction module, wherein the first attribute set corresponds to an unknown entity and wherein the first attribute set includes a first attribute; generating an unknown entity graph including a first node corresponding to the first attribute by a graph generation module; retrieving a second attribute set from a database including known entities by the extraction module, wherein the second attribute set corresponds to one of the known entities and wherein the second attribute set includes a second attribute; generating a known entity graph including a second node corresponding to the second attribute by the graph generation module; generating an unknown entity graph embedding by a graph neural network model by applying the unknown entity graph to the graph neural network model; generating a known entity graph embedding by the graph neural network model by applying the known entity graph to the graph neural network model; generating an embedding space similarity score based on the unknown entity graph embedding and the known entity graph embedding by an embedding similarity module; determining an overall similarity score based on the embedding space similarity score by a recommendation module; and assigning the information source to one of the known entities among the known entities based on the overall similarity score by the recommendation module.

[0097] Clause 8. A method according to Clause 7, wherein each of the first attributes includes at least one word, and wherein each of the second attributes includes at least one word, the method further comprising: identifying attribute pairs by a string similarity module, wherein each of the attribute pairs includes a first attribute of the first attributes and a corresponding second attribute of the second attributes; and generating string similarity scores by the string similarity module, wherein each of the string similarity scores is based on one of the attribute pairs, and wherein the overall similarity score is further based on the string similarity scores.

[0098] Clause 9. A method according to any one of clauses 7 to 8, wherein determining the overall similarity score comprises: applying the embedding space similarity score and the string similarity score to at least one of a regression model or a classification model by the recommendation module.

[0099] Clause 10. A method according to any one of clauses 7 to 9, wherein generating the embedding space similarity score comprises: applying the unknown entity graph embedding and the known entity graph embedding to a deep neural network model by the embedding similarity module.

[0100] Clause 11. The method according to any one of clauses 7 to 10 further comprises: a training module end-to-end training the graph neural network model, the deep neural network model, and at least one of the regression model or the classification model based on labeled data.

[0101] Clause 12. A method according to any one of clauses 7 to 11, wherein the graph neural network model is a graph convolutional network model.

[0102] Clause 13. The method of any one of clauses 7 to 12, wherein extracting the first set of attributes from the information source comprises extracting the first attributes from at least one of a news article or a web page.

[0103] Clause 14. A method according to any one of clauses 7 to 13, wherein extracting the first set of attributes from the information source includes extracting at least one of: the name of the unknown entity; the industry associated with the unknown entity; geographic information related to the unknown entity; the name of an entity that is not the unknown entity; a classification of images associated with the unknown entity; or a classification of videos associated with the unknown entity.

[0104] Clause 15. A method according to any one of clauses 7 to 14, wherein retrieving the second set of attributes from the database includes retrieving at least one of the following: the name of a known entity among the known entities; an industry associated with the known entity among the known entities; geographic information related to the known entity among the known entities; the name of a parent entity of the known entity among the known entities; the name of a child entity of the known entity among the known entities; or the name of an entity known to transact business with the known entity among the known entities.

[0105] Item 16. An entity linking system, comprising: a graph generation module, configured to: generate an unknown entity graph based on a first attribute set including a first attribute, wherein the first attribute is extracted from an information source, wherein the first attribute set corresponds to an unknown entity, and wherein the unknown entity graph includes a first node corresponding to the first attribute; and generate a known entity graph based on a second attribute set including a second attribute, wherein the second attribute is retrieved from a database including known entities, wherein the second attribute set corresponds to one of the known entities, and wherein the known entity graph includes a second node corresponding to the second attribute; a graph neural network, configured to: generate an unknown entity graph embedding based on the unknown entity graph; and generate a known entity graph embedding based on the known entity graph; an embedding space similarity module, configured to generate an embedding space similarity score based on the unknown entity graph embedding and the known entity graph embedding; and a recommendation module, configured to assign the information source to one of the known entities based on the embedding space similarity score.

[0106] Clause 17. The system according to clause 16 further comprises: a string similarity module configured to generate string similarity scores based on attribute pairs, wherein each of the attribute pairs includes a first attribute of the first attributes and a corresponding second attribute of the second attributes, and wherein each of the string similarity scores is based on an attribute pair of the attribute pairs; and wherein the recommendation module is further configured to assign the information source to one of the known entities based on the string similarity score and the embedding space similarity score.

[0107] Clause 18. A system according to any one of clauses 16 to 17, wherein the graph neural network is a graph convolutional network, wherein the embedding space similarity module comprises a deep neural network, and wherein the recommendation module comprises at least one of a regression model or a classification model.

[0108] Clause 19. A system according to any one of clauses 16 to 18, wherein the information source is at least one of a news article or a web page, and wherein the first set of attributes includes at least one of: the name of the unknown entity; the industry associated with the unknown entity; geographic information associated with the unknown entity; the name of an entity that is not the unknown entity; a classification of images associated with the unknown entity; or a classification of videos associated with the unknown entity.

[0109] Clause 20. A system according to any one of clauses 16 to 19, wherein the second set of attributes includes at least one of the following: the name of a known entity among the known entities; an industry associated with the known entity among the known entities; geographic information related to the known entity among the known entities; the name of a parent entity of the known entity among the known entities; the name of a child entity of the known entity among the known entities; or the name of an entity known to transact business with the known entity among the known entities.

[0110] Furthermore, it should be understood that any one or more of the forms, formal expressions, and examples described below may be combined with any one or more of the other forms, formal expressions, and examples described below.

[0111] Any software component or function described in this application can be implemented as a software code executed by a processor using any suitable computer language (e.g., Python, Java, C++, or Perl), using, for example, conventional or object-oriented techniques. The software code can be stored as a series of instructions or commands on a computer-readable medium such as a random access memory (RAM), a read-only memory (ROM), a magnetic medium (such as a hard drive or floppy disk), or an optical medium (such as a CD-ROM). Any such computer-readable medium can reside on or within a single computing device, and can exist on or within different computing devices within a system or network.

[0112] Although several forms have been shown and described, the applicant has no intention to limit or restrict the scope of the appended claims to such details. Without departing from the scope of the present disclosure, many modifications, variations, changes, substitutions, combinations and equivalents to these forms can be implemented and will be thought of by those skilled in the art. In addition, the structure of each element associated with the described form can be alternatively described as a member for providing the function performed by the element. In addition, in the case of disclosing the material of certain parts, other materials can be used. Therefore, it should be understood that the aforementioned description and the appended claims are intended to cover all such modifications, combinations and variations that fall within the scope of the disclosed form. The appended claims are intended to cover all such modifications, variations, changes, substitutions, modifications and equivalents.

[0113] The foregoing detailed description has been described in various forms of devices and / or processes by using block diagrams, flow charts and / or examples. To the extent that such block diagrams, flow charts and / or examples contain one or more functions and / or operations, those skilled in the art will understand that each function and / or operation within such block diagrams, flow charts and / or examples can be implemented individually and / or collectively by a wide range of hardware, software, firmware or any combination thereof. Those skilled in the art will recognize that some aspects of the forms disclosed herein can be implemented in whole or in part in an integrated circuit equivalently as one or more computer programs running on one or more computers (e.g., as one or more programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or as any combination thereof, and recognize that according to the present disclosure, designing circuit systems and / or writing codes for software and or firmware will be well within the technical scope of those skilled in the art. In addition, those skilled in the art will understand that the mechanisms of the subject matter described herein can be distributed as one or more program products in various forms, and will understand that the illustrative forms of the subject matter described herein are applicable regardless of the specific type of signal-bearing medium used to actually perform the distribution.

[0114] Instructions for programming the logic to perform various disclosed aspects may be stored in a memory (e.g., dynamic random access memory (DRAM), cache memory, flash memory, or other storage device) in the system. In addition, the instructions may be distributed via a network or through other computer-readable media. Thus, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), but is not limited to a floppy disk, an optical disk, a compact disk, a read-only memory (CD-ROM) and a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic or optical card, a flash memory, or a tangible machine-readable storage device for transmitting information via an electrical, optical, acoustic, or other form of propagation signal (e.g., a carrier wave, an infrared signal, a digital signal, etc.) via the Internet. Thus, a non-transitory computer-readable medium includes any type of tangible machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0115] As used in any aspect of this document, the term "control circuitry" may refer to, for example, hard-wired circuitry, programmable circuitry (e.g., a computer processor including one or more separate instruction processing cores, a processing unit, a processor, a microcontroller, a microcontroller unit, a controller, a digital signal processor (DSP), a programmable logic device (PLD), a programmable logic array (PLA), or a field programmable gate array (FPGA)), state machine circuitry, firmware that stores instructions executed by the programmable circuitry, and any combination thereof. The control circuitry may be collectively or individually embodied as circuitry that forms part of a larger system, such as an integrated circuit (IC), an application specific integrated circuit (ASIC), a system on a chip (SoC), a desktop computer, a laptop computer, a tablet computer, a server, a smartphone, etc. Thus, as used herein, "control circuitry" includes, but is not limited to, circuitry having at least one discrete circuit, circuitry having at least one integrated circuit, circuitry having at least one application specific integrated circuit, circuitry forming a general purpose computing device configured by a computer program (e.g., a general purpose computer configured by a computer program that at least partially performs the processes and / or devices described herein, or a microprocessor configured by a computer program that at least partially performs the processes and / or devices described herein), circuitry forming a memory device (e.g., in the form of random access memory), and / or circuitry forming a communication device (e.g., a modem, a communication switch, or an optical electrical device). Those skilled in the art will recognize that the subject matter described herein may be implemented in an analog or digital manner, or some combination thereof.

[0116] As used in any aspect of this document, the term "logic" may refer to an application, software, firmware, and / or circuitry configured to perform any of the foregoing operations. Software may be embodied as a software package, code, instructions, instruction sets, and / or data recorded on a non-transitory computer-readable storage medium. Firmware may be embodied as code, instructions, or instruction sets and / or data hard-coded (e.g., non-volatile) in a memory device.

[0117] As used in any aspect herein, the terms "component," "system," "module," and the like may refer to a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution.

[0118] As used in any aspect herein, an "algorithm" refers to a self-consistent sequence of steps leading to a desired result, where "steps" refer to manipulations of physical quantities and / or logical states, although not necessarily in the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. These signals are often referred to as bits, values, elements, symbols, characters, terms, numbers, etc. These and similar terms may be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities and / or states.

[0119] The network may include a packet-switched network. The communication devices may be able to communicate with each other using a selected packet-switched network communication protocol. An exemplary communication protocol may include an Ethernet communication protocol, which may be able to permit communication using a transmission control protocol / internet protocol (TCP / IP). The Ethernet protocol may conform to or be compatible with the Ethernet standard entitled "IEEE 802.3 Standard (IEEE 802.3Standard)" issued by the Institute of Electrical and Electronics Engineers (IEEE) in December 2008 and / or subsequent versions of this standard. Alternatively or in addition, the communication devices may be able to communicate with each other using an X.25 communication protocol. The X.25 communication protocol may conform to or be compatible with a standard promulgated by the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T). Alternatively or in addition, the communication devices may be able to communicate with each other using a frame relay communication protocol. The frame relay communication protocol may conform to or be compatible with standards promulgated by the Consultative Committee for International Telegraph and Telephone (CCITT) and / or the American National Standards Institut (ANSI). Alternatively or in addition, the transceivers may be capable of communicating with each other using an asynchronous transfer mode (ATM) communication protocol. The ATM communication protocol may conform to or be compatible with the ATM standard entitled "ATM-MPLS Network Interworking 2.0" published by the ATM Forum in August 2001 and / or subsequent versions of this standard. Of course, different and / or developed connection-oriented network communication protocols are also contemplated herein.

[0120] Unless otherwise specifically stated, as is apparent from the foregoing disclosure, it should be understood that throughout the foregoing disclosure, discussions using terms such as "processing," "computing," "calculating," "determining," "displaying," etc. refer to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical (electronic) quantities within the computer system registers and memories and transforms it into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission, or display devices.

[0121] One or more components may be referred to herein as "configured to," "configurable to," "operable / operable," "suitable / adaptable to," "capable of," "compliant / compliant with," etc. Those skilled in the art will recognize that "configured to" may generally encompass active state components and / or inactive state components and / or standby state components, unless the context requires otherwise.

[0122] Those skilled in the art will recognize that, in general, the terms used herein and particularly in the appended claims (e.g., the appended claim bodies) are generally intended as "open" terms (e.g., the term "including" should be interpreted as "including but not limited to", the term "having" should be interpreted as "having at least", the term "includes" should be interpreted as "includes but is not limited to", etc.). Those skilled in the art will further understand that if a specific number of an introduced claim statement is intended, such intent will be expressly stated in the claim, and in the absence of such a statement, no such intent is present. For example, to aid understanding, the following appended claims may contain the use of the introductory phrases "at least one" and "one or more" to introduce claim statements. However, the use of such phrases should not be interpreted as meaning that introducing a claim recitation by the indefinite article "a" or "an" limits any particular claim containing such cited claim recitation to claims containing only one such recitation, even when the same claim includes the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should generally be interpreted as meaning "at least one" or "one or more"); the same is true for the use of definite articles used to introduce claim recitations.

[0123] Additionally, even if a specific number of an introduced claim recitation is explicitly stated, one skilled in the art will recognize that such recitation should generally be interpreted to mean at least the stated number (e.g., simply stating "two recitations" without other modifiers generally means at least two recitations, or two or more recitations). Furthermore, in those cases where a convention similar to "at least one of A, B, and C, etc. is used, generally such construction is intended in the sense that one skilled in the art would understand the convention (e.g., "a system having at least one of A, B, and C" would include but is not limited to systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In those cases where a convention similar to "at least one of A, B, or C, etc. is used, generally such construction is intended in the sense that one skilled in the art would understand the convention (e.g., "a system having at least one of A, B, or C" would include but is not limited to systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those skilled in the art that, whether in the description, claims or drawings, distinguishing words and / or phrases that generally present two or more alternative terms should be understood to include the possibility of one, any or both of the terms, unless the context dictates otherwise. For example, the phrase "A or B" will generally be understood to include the possibility of "A" or "B" or "A and B".

[0124] With regard to the appended claims, it will be appreciated by those skilled in the art that the operations described therein can be performed in any order in general. In addition, although various operational flow charts are presented in one or more sequences, it should be understood that various operations can be performed in other orders other than those described, or can be performed simultaneously. Examples of such alternative arrangements may include overlapping, interlaced, interrupted, reordered, incremental, preparatory, supplementary, simultaneous, reversed or other variations of the arrangement, unless the context otherwise provides. In addition, terms such as "in response to", "related to..." or other past tense adjectives are generally not intended to exclude such variants, unless the context otherwise provides.

[0125] It is worth noting that any reference to "one aspect," "an aspect," "an example," "an example," etc. means that a particular feature, structure, or characteristic described in conjunction with the aspect is included in at least one aspect. Thus, the phrases "in one aspect," "in an aspect," "in an example," and "in an example" appearing in various places throughout this specification are not necessarily all referring to the same aspect. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more aspects.

[0126] Any patent application, patent, non-patent publication, or other public material cited in this specification and / or listed in any application data sheet is incorporated herein by reference to the extent that the incorporated material is not inconsistent therewith. Thus, and to the extent necessary, the disclosure as expressly set forth herein supersedes any conflicting material incorporated by reference. Any material or portion thereof that is stated to be incorporated herein by reference but conflicts with existing definitions, statements, or other public materials set forth herein will be incorporated only to the extent that there is no conflict between the incorporated material and the existing public materials.

[0127] In summary, many benefits have been described that result from adopting the concepts described herein. The foregoing description of one or more forms has been presented for purposes of illustration and description. It is not intended to be exhaustive or limited to the precise forms disclosed. In view of the above teachings, modifications or variations are possible. One or more forms are selected and described to illustrate the principles and practical applications, thereby enabling a person of ordinary skill in the art to utilize various forms and make various modifications according to the specific use contemplated. The claims submitted herein are intended to limit the overall scope.

[0128] The above description is illustrative rather than restrictive. After reading this disclosure, many variations of the required subject matter will become apparent to those skilled in the art. Therefore, the scope of the present disclosure should not be determined with reference to the above description, but should be determined with reference to the pending claims and their full scope or equivalent.

[0129] All patents, patent applications, publications, and descriptions mentioned above are incorporated by reference in their entirety for all purposes. No admission is made that they are prior art.

Claims

1. A computer-implemented method, include: extracting, by an extraction module, a first attribute set from an information source, wherein the first attribute set corresponds to an unknown entity, and wherein the first attribute set includes a first attribute; generating, by a graph generation module, an unknown entity graph including a first node corresponding to the first attribute; retrieving, by the extraction module, a second set of attributes from a database comprising known entities, wherein the second set of attributes corresponds to one of the known entities, and wherein the second set of attributes comprises a second attribute; generating, by the graph generation module, a known entity graph including a second node corresponding to the second attribute; generating, by a graph neural network model, an unknown entity graph embedding by applying the unknown entity graph to the graph neural network model; generating, by the graph neural network model, a known entity graph embedding by applying the known entity graph to the graph neural network model; generating, by an embedding similarity module, an embedding space similarity score based on the unknown entity graph embedding and the known entity graph embedding; as well as The information source is assigned to one of the known entities by a recommendation module based on the embedding space similarity score.

2. The method according to claim 1, wherein the graph neural network model is a graph convolutional network model. 3 . The method of claim 1 , wherein extracting the first set of attributes from the information source comprises extracting, by the extraction module, the first attributes from at least one of a news article or a web page.

4. The method of claim 1, wherein extracting the first attribute set from the information source comprises extracting, by the extraction module, at least one of: the name of the unknown entity; The industry associated with the unknown entity; Geographic information associated with the unknown entity; the name of an entity that is not the unknown entity; a classification of images associated with the unknown entity; or A classification of the video associated with the unknown entity.

5. The method of claim 1 , wherein retrieving the second set of attributes from the database comprises retrieving, by the extraction module, at least one of: a name of one of the known entities; an industry associated with one of the known entities; geographic information associated with one of the known entities; the name of a parent entity of one of the known entities; a name of a sub-entity of a known entity of the known entities; or The name of an entity known to transact business with one of the known entities.

6. The method of claim 1, wherein generating the embedding space similarity score include: The unknown entity graph embedding and the known entity graph embedding are applied to a deep neural network model by the embedding similarity module.

7. A computer-implemented method, include: extracting, by an extraction module, a first attribute set from an information source, wherein the first attribute set corresponds to an unknown entity, and wherein the first attribute set includes a first attribute; generating, by a graph generation module, an unknown entity graph including a first node corresponding to the first attribute; retrieving, by the extraction module, a second set of attributes from a database comprising known entities, wherein the second set of attributes corresponds to one of the known entities, and wherein the second set of attributes comprises a second attribute; generating, by the graph generation module, a known entity graph including a second node corresponding to the second attribute; generating, by a graph neural network model, an unknown entity graph embedding by applying the unknown entity graph to the graph neural network model; generating, by the graph neural network model, a known entity graph embedding by applying the known entity graph to the graph neural network model; generating, by an embedding similarity module, an embedding space similarity score based on the unknown entity graph embedding and the known entity graph embedding; determining, by a recommendation module, an overall similarity score based on the embedding space similarity score; as well as The information source is assigned, by the recommendation module, to one of the known entities based on the overall similarity score.

8. The method of claim 7, wherein each of the first attributes comprises at least one word, and wherein each of the second attributes comprises at least one word, the method further comprising: include: identifying, by a string similarity module, attribute pairs, wherein each of the attribute pairs includes a first one of the first attributes and a corresponding one of the second attributes; as well as String similarity scores are generated by the string similarity module, wherein each of the string similarity scores is based on one of the attribute pairs, and wherein the overall similarity score is further based on the string similarity scores.

9. The method of claim 8, wherein determining the overall similarity score include: The embedding space similarity score and the string similarity score are applied by the recommendation module to at least one of a regression model or a classification model.

10. The method of claim 9, wherein generating the embedding space similarity score include: The unknown entity graph embedding and the known entity graph embedding are applied to a deep neural network model by the embedding similarity module.

11. The method according to claim 10, further comprising: include: The graph neural network model, the deep neural network model, and at least one of the regression model or the classification model are trained end-to-end by a training module based on labeled data.

12. The method according to claim 7, wherein the graph neural network model is a graph convolutional network model.

13. The method of claim 7, wherein extracting the first set of attributes from the information source comprises extracting, by the extraction module, the first attributes from at least one of a news article or a web page.

14. The method of claim 7, wherein extracting the first set of attributes from the information source comprises extracting, by the extraction module, at least one of: the name of the unknown entity; The industry associated with the unknown entity; Geographic information associated with the unknown entity; the name of an entity that is not the unknown entity; a classification of images associated with the unknown entity; or A classification of the video associated with the unknown entity.

15. The method of claim 7, wherein retrieving the second set of attributes from the database comprises retrieving, by the extraction module, at least one of: a name of one of the known entities; an industry associated with one of the known entities; geographic information associated with one of the known entities; the name of a parent entity of one of the known entities; a name of a sub-entity of a known entity of the known entities; or The name of an entity known to transact business with one of the known entities.

16. An entity linking system, include: A graph generation module that can be used to: generating an unknown entity graph based on a first attribute set including a first attribute, wherein the first attribute is extracted from an information source, wherein the first attribute set corresponds to an unknown entity, and wherein the unknown entity graph includes a first node corresponding to the first attribute; and generating a known entity graph based on a second set of attributes including a second attribute, wherein the second attribute is retrieved from a database including known entities, wherein the second set of attributes corresponds to one of the known entities, and wherein the known entity graph includes a second node corresponding to the second attribute; The graph neural network is configured to: generating an unknown entity graph embedding based on the unknown entity graph; and generating a known entity graph embedding based on the known entity graph; an embedding space similarity module configured to generate an embedding space similarity score based on the unknown entity graph embedding and the known entity graph embedding; as well as A recommendation module is configured to assign the information source to one of the known entities based on the embedding space similarity score.

17. The entity linking system according to claim 16, further comprising: include: a string similarity module configured to generate string similarity scores based on attribute pairs, wherein each of the attribute pairs includes a first one of the first attributes and a corresponding one of the second attributes, and wherein each of the string similarity scores is based on one of the attribute pairs; and The recommendation module is further configured to assign the information source to one of the known entities based on the string similarity score and the embedding space similarity score.

18. The entity linking system of claim 17, wherein the graph neural network is a graph convolutional network, wherein the embedding space similarity module comprises a deep neural network, and wherein the recommendation module comprises at least one of a regression model or a classification model.

19. The entity linking system of claim 16, wherein the information source is at least one of a news article or a web page, and wherein the first attribute set includes at least one of: the name of the unknown entity; The industry associated with the unknown entity; Geographic information associated with the unknown entity; the name of an entity that is not the unknown entity; a classification of images associated with the unknown entity; or A classification of the video associated with the unknown entity.

20. The entity linking system according to claim 16, wherein the second attribute set includes at least one of the following: a name of one of the known entities; an industry associated with one of the known entities; geographic information associated with one of the known entities; the name of a parent entity of one of the known entities; a name of a sub-entity of a known entity of the known entities; or The name of an entity known to transact business with one of the known entities.