Entity linking method and device, equipment and storage medium
By using a layer-by-layer classification process and character convolution network in the entity linking process, the problem of difficult to distinguish between character position exchange and character block omission in the prior art is solved, and a more efficient and accurate entity linking effect is achieved.
Patent Information
- Application Number
- CN202510043739.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to distinguish between character position exchange and character block omission during the entity linking process, resulting in inefficiency and limited generalization capabilities of entity linking.
The layer-by-layer link entity classification process is adopted. First, the initial classifier is used to identify the easy-to-classify situations, and then the character class variation situation is further divided into character position exchange or character block omission through a character convolution network based on auxiliary features such as Damerau editing distance and overlapping word proportion.
It improves the accuracy and efficiency of entity links, can more effectively distinguish complex character mutations, and improves the performance of entity links.
Smart Images

Figure CN119988642A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to an entity linking method, device, equipment and storage medium. Background Art
[0002] Entity linking is a key task in natural language processing, which aims to link entities extracted from text (which may be non-standard) to entity entries in a standard knowledge base. This linking helps users to have a deeper language understanding and information retrieval.
[0003] In the related technologies, early entity linking mainly relied on predefined rules and dictionary matching, such as using named entity recognition tools to extract entities from texts, and then linking entities to standard knowledge bases through string matching and other computational methods. This method is inefficient, relies on a large amount of labeled data, and has limited generalization capabilities. In recent years, entity linking methods based on language models or machine learning have gradually become mainstream, which can capture more comprehensive features to improve the performance of entity linking. However, in practical applications, they still cannot distinguish between some different situations. For example, character position exchange and character block omission are more complicated than abbreviations. In addition to being more ambiguous and complex, if some characters are missing from an entity surface and character positions are exchanged at the same time, it is difficult to distinguish between the two situations.
[0004] Based on the above analysis of the development status of this technology field, the existing technology lacks a solution for using machine learning algorithms and auxiliary information to perform detailed classification of character variations in the entity linking process to form an accurate knowledge base. Summary of the invention
[0005] The purpose of the present invention is to provide an entity linking method, device, equipment and storage medium, aiming to solve the above-mentioned problems in the prior art.
[0006] According to a first aspect of an embodiment of the present invention, there is provided an entity linking method, comprising:
[0007] Obtain standard knowledge graphs and multi-source documents, and identify candidate entities in multi-source documents, including:
[0008] Use the natural language entity recognition model BERT to identify candidate entities in multi-source documents.
[0009] Match the candidate entities to the corresponding standard entities in the standard knowledge graph, filter the candidate entities according to the matching results, and obtain the linked entities, including:
[0010] Determine whether the candidate entity is exactly the same as the standard entity through a string matching algorithm;
[0011] Eliminate the candidate entities that are exactly the same in the standard knowledge graph, and use the remaining candidate entities as link entities.
[0012] Inputting the link entities into the encoder, extracting entity features corresponding to each link entity through the encoder, inputting the entity features into the initial classifier to output the classification results of each link entity, and obtaining the link entities whose classification results are character class variations as the first set, specifically including:
[0013] The encoder extracts entity features of linked entities including n-gram features, part-of-speech tags, character types, and character type ratios;
[0014] The entity features are input into a pre-established initial classifier, wherein the initial classifier is a support vector machine, and the classification results of each linked entity are output through the support vector machine, wherein the classification results include aliases, name abbreviations, English abbreviations, pinyin abbreviations and character class variations.
[0015] Acquiring auxiliary features based on the first set and the corresponding standard entities, inputting the first set, the corresponding standard entities and the auxiliary features into the character convolution network, and updating the classification results of the linked entities in the first set through the output of the character convolution network, specifically including:
[0016] Calculating the Damerau edit distance between the linked entities in the first set and the corresponding standard entities;
[0017] Count the number of identical characters between the link entity and the corresponding standard entity in the first set, and calculate the ratio of the number of identical characters to the character length of the link entity as the overlapping character ratio;
[0018] Damerau edit distance and overlapping word ratio are used as auxiliary features.
[0019] Obtain a CharCNN convolutional model as a character convolutional network, wherein the character convolutional network includes a character embedding layer, a convolutional layer, a pooling layer, and a fully connected layer;
[0020] Use the separator concatenation method to concatenate the linked entities in the first set with the corresponding standard entities into a long sequence; input the long sequence into the character embedding layer, convert the long sequence into a dense vector representation through the character embedding layer, input the dense vector representation into the convolution layer to capture the character relationship, and input the character relationship into the pooling layer to obtain local features;
[0021] The auxiliary features and local features are input into the fully connected layer, and the classification results of character class variations are updated to character position exchange or character block omission.
[0022] Establish an index based on the classification results, and supplement the standard knowledge graph with links based on the index, including:
[0023] The classification result is used as the key of the index under the corresponding standard entity, and the link entity is used as the value of the index;
[0024] Add the index to the corresponding standard entity in the standard knowledge graph, use the index key as the connecting edge, and use the index value as the node to supplement the link.
[0025] According to a second aspect of an embodiment of the present invention, there is provided a physical linking device, comprising:
[0026] The initial module is used to obtain standard knowledge graphs and multi-source documents and identify candidate entities in multi-source documents. Specifically, it is used to:
[0027] Use the natural language entity recognition model BERT to identify candidate entities in multi-source documents.
[0028] The screening module is used to match the standard entities corresponding to the candidate entities in the standard knowledge graph, and screen the candidate entities according to the matching results to obtain the linked entities. It is specifically used for:
[0029] Determine whether the candidate entity is exactly the same as the standard entity through a string matching algorithm;
[0030] Eliminate the candidate entities that are exactly the same in the standard knowledge graph, and use the remaining candidate entities as link entities.
[0031] The initial classification module is used to input the link entity into the encoder, extract the entity features corresponding to each link entity through the encoder, input the entity features into the initial classifier to output the classification results of each link entity, and obtain the link entity whose classification result is a character class variation as the first set, which is specifically used for:
[0032] The encoder extracts entity features of linked entities including n-gram features, part-of-speech tags, character types, and character type ratios;
[0033] The entity features are input into a pre-established initial classifier, wherein the initial classifier is a support vector machine, and the classification results of each linked entity are output through the support vector machine, wherein the classification results include aliases, name abbreviations, English abbreviations, pinyin abbreviations and character class variations.
[0034] The character convolution module is used to obtain auxiliary features based on the first set and the corresponding standard entities, input the first set, the corresponding standard entities and the auxiliary features into the character convolution network, and update the classification results of the linked entities in the first set through the output of the character convolution network, specifically for:
[0035] Calculating the Damerau edit distance between the linked entities in the first set and the corresponding standard entities;
[0036] Count the number of identical characters between the link entity and the corresponding standard entity in the first set, and calculate the ratio of the number of identical characters to the character length of the link entity as the overlapping character ratio;
[0037] Damerau edit distance and overlapping word ratio are used as auxiliary features.
[0038] Obtain a CharCNN convolutional model as a character convolutional network, wherein the character convolutional network includes a character embedding layer, a convolutional layer, a pooling layer, and a fully connected layer;
[0039] Use the separator concatenation method to concatenate the linked entities and the corresponding standard entities in the first set into a long sequence;
[0040] The long sequence is input into the character embedding layer, which converts the long sequence into a dense vector representation. The dense vector representation is input into the convolution layer to capture the character relationship, and the character relationship is input into the pooling layer to obtain local features.
[0041] The auxiliary features and local features are input into the fully connected layer, and the classification results of character class variations are updated to character position exchange or character block omission.
[0042] The index link module is used to build an index based on the classification results and to supplement the standard knowledge graph with links according to the index. It is specifically used for:
[0043] The classification result is used as the key of the index under the corresponding standard entity, and the link entity is used as the value of the index;
[0044] Add the index to the corresponding standard entity in the standard knowledge graph, use the index key as the connecting edge, and use the index value as the node to supplement the link.
[0045] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the entity linking method provided in the first aspect of the present disclosure are implemented.
[0046] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which a program for implementing information transmission is stored, and when the program is executed by a processor, the steps of the entity linking method provided in the first aspect of the present disclosure are implemented.
[0047] The technical solution provided by the embodiment of the present invention includes the following beneficial effects: using a layer-by-layer linked entity classification process, first using an initial classifier to identify situations that are easy to classify, and then based on auxiliary features that can more deeply explore the connections, using a character convolutional network to further divide the character class variations into character position exchanges and character block omissions, thereby improving the entity linking effect.
[0048] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0050] Figure 1 is a flow chart of an entity linking method according to an embodiment of the present invention;
[0051] Figure 2 Schematic diagram of CharCNN convolution model processing in an embodiment of the present invention;
[0052] Figure 3 is a schematic diagram of an index link according to an embodiment of the present invention;
[0053] Figure 4 is a schematic diagram of a physical linking device according to an embodiment of the present invention;
[0054] Figure 5 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0055] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will be combined with the drawings in one or more embodiments of this specification to clearly and completely describe the technical solutions in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this document.
[0056] For the convenience of the following description, the technical terms involved in the embodiments of the present invention are explained as follows:
[0057] (1) Alias: refers to an alias for a standard entity. There is no abbreviation relationship between the two. For example, "RQ-1" is an alias for the U.S. military's unmanned reconnaissance aircraft "Predator";
[0058] The abbreviation is a combination of the abbreviations of the various parts of the standard entity, including the following three cases:
[0059] (2) Name abbreviation: refers to the abbreviation of common words;
[0060] (3) English abbreviations: For example, “STS” is the abbreviation of “Space Transportation System”;
[0061] (4) Pinyin abbreviation: For example, "TianGong-1" is the pinyin abbreviation of "Tiangong-1";
[0062] (5) Character variation: It is divided into two cases: character position exchange and character block omission. The former indicates that there is a position exchange phenomenon between characters. For example, the standard entity "precision tracking and aiming system" is often called "tracking and aiming precision system". The latter indicates that a part of the entity is used to replace the complete entity. For example, the standard entity "LR-115 liquid hydrogen engine" is often called "LR-115".
[0063] Method Embodiment
[0064] According to an embodiment of the present invention, a method for entity linking is provided. Figure 1 is a flow chart of the entity linking method according to an embodiment of the present invention. Figure 1 As shown, the entity linking method according to an embodiment of the present invention specifically includes:
[0065] In step S110, a standard knowledge graph and multi-source documents are obtained, and candidate entities in the multi-source documents are identified, specifically including:
[0066] The standard knowledge graph is a pre-established structured, semantically rich data representation that aims to describe entities and relationships between entities in the real world in a unified format and standard. The nodes in the standard knowledge graph are standard entities, that is, standardized terminology expressions.
[0067] However, in actual applications, expressions are often not particularly standardized, and entity linking is needed to help users understand what standard entities correspond to non-standard expressions, or to further understand the relationship between non-standard and standard expressions.
[0068] The multi-source documents are documents related to the fields described by the standard knowledge graph, so as to extract entities of the same type or the same as those in the standard knowledge graph; the trained natural language entity recognition model BERT (Bidirectional Encoder Representations from Transformers) is used to identify candidate entities in the multi-source documents. The BERT model is widely used in natural language processing and can effectively identify entities in multi-source documents;
[0069] Preferably, the extraction parameters in BERT are set to filter out entities that meet the requirements. Taking the military field as an example, it is hoped that equipment names or action names with clear meanings are filtered out, rather than some meaningless names or other types of entities.
[0070] The extracted entities are called "alternative entities" because not all identified entities need to be linked to the standard knowledge base. Some of the entities themselves are expressions of standard entities, which can be directly eliminated in subsequent steps.
[0071] In step S120, matching the candidate entities with the corresponding standard entities in the standard knowledge graph, filtering the candidate entities according to the matching results, and obtaining the linked entities, specifically includes:
[0072] Match the corresponding standard entity through the string matching algorithm, and judge whether the candidate entity is exactly the same as the standard entity through the matching result. Identical means that all characters and sequences are exactly the same. If they are exactly the same, it proves that the candidate entity is a standard entity.
[0073] Eliminate the candidate entities that are exactly the same in the standard knowledge graph, and use the remaining candidate entities as link entities;
[0074] Preferably, since BERT has extracted candidate entities with relatively high quality, in the application scenario of the embodiment of the present invention, there are no completely redundant candidate entities, that is, all can be matched to corresponding standard entities.
[0075] In step S130, the link entities are input into an encoder, entity features corresponding to each link entity are extracted by the encoder, the entity features are input into an initial classifier to output classification results of each link entity, and link entities whose classification results are character class variations are obtained as a first set, specifically including:
[0076] Extracting entity features of linked entities including but not limited to n-gram features, part-of-speech tags, character types, and character type ratios through an encoder, wherein n-gram features refer to dividing an entity into n strings, and the embodiment of the present invention does not limit the form of the encoder;
[0077] Character types, such as English characters and Chinese characters, character type ratios, such as the ratio of English characters to the entire link entity, the ratio of uppercase English characters to English characters, etc., can be set according to actual conditions;
[0078] The entity features used in the initial classification process are relatively superficial and can quickly obtain classification results;
[0079] Input entity features into a pre-established initial classifier, where the initial classifier is a support vector machine. The support vector machine is a multi-classification model that can map entity features to a high-dimensional space and then explore the relationship with the type. The support vector machine outputs the classification results of each linked entity, where the classification results include aliases, name abbreviations, English abbreviations, pinyin abbreviations, and character class variations;
[0080] In an embodiment of the present invention, the two situations of character position exchange and character block omission are superimposed into character class variation. This is because if the linked entity has both superficial character exchange and character omission compared to the standard entity, such as "DF31 ballistic missile" and "missile DF31", it is difficult to obtain a relatively accurate classification result for the two based on relatively superficial features and using a lightweight initial classifier.
[0081] In step S140, auxiliary features are obtained based on the first set and the corresponding standard entities, the first set, the corresponding standard entities and the auxiliary features are input into the character convolution network, and the classification results of the linked entities in the first set are updated through the output of the character convolution network, specifically including:
[0082] Calculating the Damerau edit distance between the linked entities in the first set and the corresponding standard entities;
[0083] Damerau edit distance is a variant of the traditional edit distance, which is used to measure the similarity between two entity strings. The traditional edit distance considers three operation costs: cost (insertion) = cost (deletion) = 1, cost (replacement) = 2. Damerau edit distance adds the "swap" operation, which has a cost of 1.
[0084] Count the number of identical characters between the link entity and the corresponding standard entity in the first set, and calculate the ratio of the number of identical characters to the character length of the link entity as the overlapping character ratio;
[0085] Damerau edit distance and overlapping word ratio are used as auxiliary features;
[0086] The Damerau edit distance and the percentage of overlapping characters can provide complementary information. The Damerau edit distance focuses on the changes in character order, while the percentage of overlapping characters focuses more on the existence or non-existence of characters. By combining the two into auxiliary features, the model can fully understand the difference between the two character variants.
[0087] Obtain a CharCNN convolutional model (Character-level Convolutional Neural Network) as a character convolutional network, wherein the character convolutional network includes a character embedding layer, a convolution layer, a pooling layer, and a fully connected layer;
[0088] Use the separator concatenation method to concatenate the linked entities in the first set with the corresponding standard entities into a long sequence. The separator can be <sep>It means that, in the embodiment of the present invention, in each long sequence, the link entity is before the separator, and the corresponding standard entity is after the separator;
[0089] The original CharCNN convolutional model often only inputs a single sequence, which the model processes character by character and uses the delimiter concatenation method to concatenate long sequences, which can effectively capture the deeper relationship between the linked entity and the corresponding standard entity;
[0090] The long sequence is input into the character embedding layer, which converts the long sequence into a dense vector representation. The dense vector representation is input into the convolution layer to capture the character relationship, and the character relationship is input into the pooling layer to obtain local features.
[0091] In the convolutional layer, the convolution kernel is set to a preset size, such as a 5-gram or 7-gram size, so that the convolution kernel can cross the separator and capture the character patterns in the linked entity and the corresponding standard entity at the same time;
[0092] The convolution layer focuses on local features, while the pooling layer can capture the global relationship between two entities to a certain extent, especially when the maximum pooling is used, the most representative feature value will be selected;
[0093] The auxiliary features and local features are concatenated and input into the fully connected layer. That is to say, the auxiliary features are not input into the character embedding layer at the beginning to avoid the interference of feature extraction caused by the early introduction of auxiliary features. The fully connected layer outputs the prediction results of character position exchange or character block omission;
[0094] The classification results of character class variation are updated to character position exchange or character block omission, which effectively distinguishes the character class variation situations. Figure 2 Schematic diagram of CharCNN convolution model processing in an embodiment of the present invention. Figure 2 As shown, the comparison processing method optimized by CharCNN is demonstrated.
[0095] In step S150, an index is established based on the classification result, and a standard knowledge graph is linked based on the index, specifically including:
[0096] The classification result is used as the key of the index under the corresponding standard entity, and the link entity is used as the index value. Taking the standard entity "National Defense Meteorological Satellite Program" as an example, the index example is shown in Table 1:
[0097] Table 1. Index example
[0098]
[0099] The index is added to the corresponding standard entity in the standard knowledge graph, with the index key as the connecting edge and the index value as the node for supplementary linking. That is, the entity "Defense Meteorological Satellite Program" and the entity "DMSP" are linked to the standard knowledge graph. Other types are not extracted from the multi-source documents, so the keys of other indexes do not exist and are empty.
[0100] Figure 3 Schematic diagram of index link of an embodiment of the present invention. Figure 3 As shown in the figure, the entity linking effect of the standard entity "National Defense Meteorological Satellite Program" in the standard knowledge graph is shown. Figure 3 Only local effects in the knowledge graph are shown.
[0101] In summary, in response to the existing problems, the entity linking method of this invention uses a layer-by-layer linked entity classification process. First, an initial classifier is used to identify situations that are easy to classify. The initial classifier is lightweight and can obtain preliminary classification results based on surface features; then based on auxiliary features that can dig deeper into the connection, the Damerlau edit distance and the proportion of overlapping characters in the auxiliary features can provide complementary information, which can enable the model to fully understand the difference between the two character variants. The character convolution network CharCNN is used to further divide the character class variation into character position exchange and character block omission, avoiding the difficulty in distinguishing the above two situations on the entity surface, thereby improving the entity linking effect; the delimiter splicing method is used to splice the linked entity and the corresponding standard entity into a long sequence, and then input into the character convolution network to solve the problem that the character convolution network model can only process a single sequence; the convolution layer and pooling layer of the character convolution network can capture the relationship between the linked entity and the standard entity to varying degrees; the auxiliary features are introduced in the fully connected layer to avoid the feature extraction interference caused by the early introduction of auxiliary features; overall, by marking the linked entities on the standard knowledge graph, the usability of the knowledge graph is enhanced.
[0102] Device Embodiment
[0103] According to an embodiment of the present invention, a physical linking device is provided. Figure 4 is a schematic diagram of a physical linking device according to an embodiment of the present invention. Figure 4 As shown, the entity linking device according to an embodiment of the present invention specifically includes:
[0104] The initial module 40 is used to obtain the standard knowledge graph and multi-source documents, identify candidate entities in the multi-source documents, and specifically to:
[0105] Use the natural language entity recognition model BERT to identify candidate entities in multi-source documents.
[0106] The screening module 42 is used to match the standard entity corresponding to the candidate entity in the standard knowledge graph, and screen the candidate entity according to the matching result to obtain the link entity, which is specifically used for:
[0107] Determine whether the candidate entity is exactly the same as the standard entity through a string matching algorithm;
[0108] Eliminate the candidate entities that are exactly the same in the standard knowledge graph, and use the remaining candidate entities as link entities.
[0109] The initial classification module 44 is used to input the link entities into an encoder, extract entity features corresponding to each link entity through the encoder, input the entity features into an initial classifier to output the classification results of each link entity, and obtain the link entities whose classification results are character class variations as the first set, which is specifically used to:
[0110] The encoder extracts entity features of linked entities including n-gram features, part-of-speech tags, character types, and character type ratios;
[0111] The entity features are input into a pre-established initial classifier, wherein the initial classifier is a support vector machine, and the classification results of each linked entity are output through the support vector machine, wherein the classification results include aliases, name abbreviations, English abbreviations, pinyin abbreviations and character class variations.
[0112] The character convolution module 46 is used to obtain auxiliary features based on the first set and the corresponding standard entities, input the first set, the corresponding standard entities and the auxiliary features into the character convolution network, and update the classification results of the linked entities in the first set through the output of the character convolution network, specifically for:
[0113] Calculating the Damerau edit distance between the linked entities in the first set and the corresponding standard entities;
[0114] Count the number of identical characters between the link entity and the corresponding standard entity in the first set, and calculate the ratio of the number of identical characters to the character length of the link entity as the overlapping character ratio;
[0115] Damerau edit distance and overlapping word ratio are used as auxiliary features.
[0116] Obtain a CharCNN convolutional model as a character convolutional network, wherein the character convolutional network includes a character embedding layer, a convolutional layer, a pooling layer, and a fully connected layer;
[0117] Use the separator concatenation method to concatenate the linked entities and the corresponding standard entities in the first set into a long sequence;
[0118] The long sequence is input into the character embedding layer, which converts the long sequence into a dense vector representation. The dense vector representation is input into the convolution layer to capture the character relationship, and the character relationship is input into the pooling layer to obtain local features.
[0119] The auxiliary features and local features are input into the fully connected layer, and the classification results of character class variations are updated to character position exchange or character block omission.
[0120] The index link module 48 is used to establish an index based on the classification results and to supplement the link standard knowledge graph according to the index, specifically for:
[0121] The classification result is used as the key of the index under the corresponding standard entity, and the link entity is used as the value of the index;
[0122] Add the index to the corresponding standard entity in the standard knowledge graph, use the index key as the connecting edge, and use the index value as the node to supplement the link.
[0123] In summary, in response to the existing problems, the entity linking device of this invention uses a layer-by-layer linked entity classification process. First, an initial classifier is used to identify situations that are easy to classify. The initial classifier is lightweight and can obtain preliminary classification results based on surface features; then based on auxiliary features that can dig deeper into the connection, the Damerlau edit distance and the proportion of overlapping characters in the auxiliary features can provide complementary information, which can enable the model to fully understand the difference between the two character variants. The character convolutional network CharCNN is used to further divide the character class variation into character position exchange and character block omission, avoiding the difficulty in distinguishing the above two situations on the entity surface, thereby improving the entity linking effect; the delimiter splicing method is used to splice the linked entity and the corresponding standard entity into a long sequence, and then input into the character convolutional network to solve the problem that the character convolutional network model can only process a single sequence; the convolutional layer and pooling layer of the character convolutional network can capture the relationship between the linked entity and the standard entity to varying degrees; the auxiliary features are introduced in the fully connected layer to avoid the feature extraction interference caused by the early introduction of auxiliary features; overall, by marking the linked entities on the standard knowledge graph, the usability of the knowledge graph is enhanced.
[0124] Electronic device embodiment
[0125] Figure 5 Schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device 500 may include at least one processor 510 and a memory 520. The processor 510 may execute instructions stored in the memory 520. The processor 510 is connected to the memory 520 through a data bus. In addition to the memory 520, the processor 510 may also be connected to an input device 530, an output device 540, and a communication device 550 through a data bus.
[0126] The processor 510 may be any conventional processor, such as a commercially available CPU. The processor may also include a graphics processor (Graphic Process Unit, GPU), a field programmable gate array (Field Programmable Gate Array, FPGA), a system on chip (System on Chip, SOC), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC) or a combination thereof.
[0127] The memory 520 may be implemented by any type of volatile or nonvolatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0128] In the embodiment of the present disclosure, executable instructions are stored in the memory 520, and the processor 510 can read the executable instructions from the memory 520 and execute the instructions to implement all or part of the steps of any entity linking method in the above exemplary embodiments.
[0129] Computer Readable Storage Medium Embodiments
[0130] In addition to the above-mentioned methods and apparatus, an exemplary embodiment of the present disclosure may also be a computer program product or a computer-readable storage medium storing the computer program product, wherein the computer product includes computer program instructions that can be executed by a processor to implement all or part of the steps described in any of the entity linking methods in the above-mentioned exemplary embodiments.
[0131] The computer program product may be written in any combination of one or more programming languages to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, etc., and also conventional procedural programming languages such as "C" language or similar programming languages and scripting languages (e.g., Python). The program code may be executed entirely on the user computing device, partially on the user computing device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0132] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of readable storage media include: a static random access memory (SRAM) with one or more wires electrically connected, an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk, or any suitable combination of the above.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.< / sep>
Claims
1. An entity linking method, characterized in that: include: Obtaining a standard knowledge graph and multi-source documents, and identifying candidate entities in the multi-source documents; Matching the candidate entity to a standard entity corresponding to the standard knowledge graph, and filtering the candidate entity according to the matching result to obtain a linked entity; Input the link entities into an encoder, extract entity features corresponding to each link entity through the encoder, input the entity features into an initial classifier to output classification results of each link entity, and obtain link entities whose classification results are character class variations as a first set; Acquire auxiliary features based on the first set and the corresponding standard entity, input the first set, the corresponding standard entity and the auxiliary features into a character convolutional network, and update the classification results of the linked entities in the first set through the output of the character convolutional network; An index is established based on the classification result, and the standard knowledge graph is supplemented and linked according to the index.
2. The method according to claim 1, characterized in that The identifying candidate entities in the multi-source documents specifically includes: using a natural language entity recognition model BERT to identify candidate entities in the multi-source documents.
3. The method according to claim 1, characterized in that The matching of the candidate entity to the standard entity corresponding to the candidate entity in the standard knowledge graph, and filtering the candidate entity according to the matching result to obtain the link entity specifically includes: Determining whether the candidate entity is completely identical to the standard entity by a string matching algorithm; Eliminate the candidate entities that are exactly the same in the standard knowledge graph, and use the remaining candidate entities as link entities.
4. The method according to claim 1, characterized in that: The extracting entity features corresponding to each linked entity by the encoder and inputting the entity features into the initial classifier to output the classification results of each linked entity specifically includes: The encoder extracts entity features of linked entities including n-gram features, part-of-speech tags, character types, and character type ratios; The entity features are input into a pre-established initial classifier, wherein the initial classifier is a support vector machine, and the classification results of each linked entity are output through the support vector machine, wherein the classification results include aliases, name abbreviations, English abbreviations, pinyin abbreviations and character class variations.
5. The method according to claim 1, characterized in that The acquiring of auxiliary features based on the first set and the corresponding standard entity specifically includes: Calculating the Damerau edit distance between the linked entities in the first set and the corresponding standard entities; Counting the number of identical characters between the link entity and the corresponding standard entity in the first set, and calculating the ratio of the number of identical characters to the character length of the link entity as the overlapping character ratio; The Damerlau edit distance and the overlapping character ratio are used as auxiliary features.
6. The method according to claim 1, characterized in that The step of inputting the first set, the corresponding standard entity and the auxiliary feature into a character convolution network, and updating the classification result of the linked entity in the first set through the output of the character convolution network specifically includes: Obtaining a CharCNN convolutional model as a character convolutional network, wherein the character convolutional network includes a character embedding layer, a convolutional layer, a pooling layer, and a fully connected layer; Use the separator concatenation method to concatenate the linked entities and the corresponding standard entities in the first set into a long sequence; Input the long sequence into the character embedding layer, convert the long sequence into a dense vector representation through the character embedding layer, input the dense vector representation into the convolution layer to capture the character relationship, and input the character relationship into the pooling layer to obtain local features; The auxiliary features and the local features are input into the fully connected layer, and the classification result of the character class variation is updated to character position exchange or character block omission.
7. The method according to claim 1, characterized in that The step of establishing an index based on the classification result and supplementing the link to the standard knowledge graph according to the index specifically includes: The classification result is used as the key of the index under the corresponding standard entity, and the link entity is used as the value of the index; The index is added under the corresponding standard entity in the standard knowledge graph, with the key of the index as the connecting edge and the value of the index as the node for supplementary linking.
8. A physical linking device, characterized in that: include: An initial module, configured to obtain a standard knowledge graph and multi-source documents, and identify candidate entities in the multi-source documents; A screening module, used for matching the standard entity corresponding to the candidate entity in the standard knowledge graph, screening the candidate entity according to the matching result, and obtaining a linked entity; An initial classification module, used for inputting the link entities into an encoder, extracting entity features corresponding to each link entity through the encoder, inputting the entity features into an initial classifier to output classification results of each link entity, and obtaining link entities whose classification results are character class variations as a first set; A character convolution module, used to obtain auxiliary features based on the first set and the corresponding standard entity, input the first set, the corresponding standard entity and the auxiliary features into a character convolution network, and update the classification results of the linked entities in the first set through the output of the character convolution network; An index linking module is used to establish an index based on the classification result and to supplement the link to the standard knowledge graph according to the index.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the entity linking method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an implementation program for information transmission, and when the program is executed by a processor, the steps of the entity linking method according to any one of claims 1 to 7 are implemented.