Method, apparatus, electronic device, and storage medium for entity linking

By comprehensively calculating the word similarity, word similarity and semantic similarity of short text, the accuracy of similarity calculation in short text entity links is solved, and more efficient entity links are achieved.

CN114969358BActive Publication Date: 2025-07-22IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210499774.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-09
Publication Date
2025-07-22
Estimated Expiration
2042-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the entity linking problem of short texts, especially in the case of short length, scarce features, and poor context, it is impossible to accurately calculate the entity similarity.

Method used

The comprehensive calculation method of word similarity, word similarity and semantic similarity is used to determine the entity similarity between the entity to be linked and the candidate entity. By calculating the word similarity, word similarity and semantic similarity between the entity to be linked and each candidate entity in the entity library, the highest candidate entity is selected as the linked entity.

Benefits of technology

It improves the similarity calculation accuracy of short text entity links, meets the needs of short text entity links, and ensures the accuracy and efficiency of entity links.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114969358B_ABST
    Figure CN114969358B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, electronic device, and storage medium for entity linking. The method includes calculating the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library respectively; determining the entity similarity between the entity to be linked and each candidate entity according to the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity; and determining the candidate entity with the highest entity similarity to the entity to be linked as the linked entity corresponding to the entity to be linked. The present application can determine the entity similarity between the entity to be linked and the candidate entity from three dimensions of character similarity, word similarity, and semantic similarity, effectively improving the accuracy of similarity calculation for short texts and meeting the entity linking requirements of short texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and particularly to a method, apparatus, electronic device, and storage medium for entity linking. Background Art

[0002] A knowledge graph is a way to represent the physical world and the cognitive world through abstract symbols linked by graphs, and serves as a bridge for different individuals to understand the world and exchange information. In the process of constructing a knowledge graph, entity linking is a relatively crucial step, and its main function is to correctly link entity referring terms to unambiguous candidate entities in the knowledge base.

[0003] Based on entity similarity algorithms is a common method in the entity linking process. However, due to the characteristics of short texts such as short length, scarce features, and lack of rich context, it is difficult to effectively calculate similarity in the entity linking process of short texts through entity similarity algorithms, and it cannot meet the entity linking requirements. Summary of the Invention

[0004] Based on the above requirements, this application proposes a method, apparatus, electronic device, and storage medium for entity linking, which can be used to overcome the problem of inability to meet the entity linking requirements of short texts in the prior art.

[0005] The technical solutions proposed in this application are as follows:

[0006] On the one hand, this application provides a method for entity linking, including:

[0007] Calculating the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library respectively;

[0008] Determining the entity similarity between the entity to be linked and each candidate entity according to the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity;

[0009] Determining the candidate entity with the highest entity similarity to the entity to be linked as the linked entity corresponding to the entity to be linked.

[0010] On the other hand, this application provides an apparatus for entity linking, including:

[0011] A calculation module for calculating the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library respectively;

[0012] A first determination module for determining the entity similarity between the entity to be linked and each candidate entity according to the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity;

[0013] A second determination module, configured to determine, as the linked entity corresponding to the entity to be linked, the candidate entity with the highest entity similarity to the entity to be linked.

[0014] On the other hand, the present application provides an electronic device, including:

[0015] A memory and a processor;

[0016] Wherein, the memory is used to store a program;

[0017] The processor is configured to implement the entity linking method according to any one of the above by running the program in the memory.

[0018] On the other hand, the present application provides a storage medium, including: a computer program is stored on the storage medium, and when the computer program is executed by a processor, each step of the entity linking method according to any one of the above is implemented.

[0019] In the entity linking method of the present application, the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library are calculated respectively; according to the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity, the entity similarity between the entity to be linked and each candidate entity is determined; the candidate entity with the highest entity similarity to the entity to be linked is determined as the linked entity corresponding to the entity to be linked. In the solution of the present application, the similarity between the entity to be linked and the candidate entity is determined from three dimensions of character similarity, word similarity, and semantic similarity, effectively improving the accuracy of similarity calculation.

[0020] For short texts, even if short texts have characteristics such as short length, scarce features, and lack of rich context, the solution of the present application can also determine the entity similarity between the entity to be linked and the candidate entity from three dimensions of character similarity, word similarity, and semantic similarity, effectively improving the accuracy of similarity calculation for short texts and meeting the entity linking requirements of short texts. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0022] Figure 1 It is a flowchart of an entity linking method provided by an embodiment of the present application;

[0023] Figure 2It is a schematic flowchart of the process of removing word segments in the entity to be linked that have little impact on the semantic features of the entity to be linked provided by an embodiment of the present application;

[0024] Figure 3 It is a schematic flowchart of the process of determining the elimination degree of each target word segment provided by an embodiment of the present application;

[0025] Figure 4 It is a schematic flowchart of the process of determining the order weight of the order in which each target word segment is located in the target entity provided by an embodiment of the present application;

[0026] Figure 5 It is a schematic flowchart of the calculation process of semantic similarity provided by an embodiment of the present application;

[0027] Figure 6 It is a schematic diagram of extracting key entity word segments provided by an embodiment of the present application;

[0028] Figure 7 It is a schematic flowchart of the process of differential training of similar entities provided by an embodiment of the present application;

[0029] Figure 8 It is a schematic diagram of constructing a differential training sample of similar entities provided by an embodiment of the present application;

[0030] Figure 9 It is a schematic flowchart of the process of similarity calculation training provided by an embodiment of the present application;

[0031] Figure 10 It is a schematic diagram of setting an adapter component provided by an embodiment of the present application;

[0032] Figure 11 It is a schematic diagram of the structure of an adapter component provided by an embodiment of the present application;

[0033] Figure 12 It is a schematic diagram of the structure of a device for entity linking provided by an embodiment of the present application;

[0034] Figure 13 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0035] The technical solution of the embodiment of the present application is applicable to the application scenario of short text entity linking. By adopting the technical solution of the embodiment of the present application, it is possible to determine the entity similarity between the entity to be linked and the candidate entities based on three dimensions of character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library, effectively improving the accuracy of similarity calculation for short texts and meeting the requirements of short text entity linking.

[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0037] This embodiment proposes a method for entity linking. Refer to Figure 1 As shown, the method includes:

[0038] S101. Calculate the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library respectively.

[0039] The above-mentioned entity library refers to a structured semantic knowledge base, which is used to quickly describe concepts in the physical world and their mutual relationships. For example, a knowledge graph is a kind of semantic knowledge base. A large amount of knowledge is stored in the entity library, and this knowledge is generally represented in the form of triples, <eh, r, et>, where eh represents the head entity, r represents the relationship between entities, and et represents the tail entity.

[0040] The construction process of the entity library mainly includes the following steps:

[0041] Information extraction: Extract entities, attributes, and the mutual relationships between entities from various types of data sources, and on this basis, form an ontological knowledge expression; Entity alignment: After obtaining new knowledge, it is necessary to integrate it to eliminate contradictions and ambiguities. For example, some entities may have multiple expressions, and a certain specific appellation may correspond to multiple different entities, etc.; Knowledge processing: For the fused new knowledge, after quality evaluation, the qualified part can be added to the entity library to ensure the quality of the entity library.

[0042] Entity linking is an important step in entity alignment. Entity linking refers to the operation of linking the entity object extracted from the text, that is, the entity to be linked in this embodiment, to the corresponding correct entity object in the entity library. Its basic idea is to first select a group of candidate entities from the entity library according to the given entity to be linked, and then link the entity to be linked to the correct entity object through similarity calculation.

[0043] The above-mentioned candidate entities are entities in the entity library that have several identical characters with the entity to be linked. The specific number of identical characters can be set according to the actual situation. For example, in a Chinese language environment, entities in the entity library that have at least 2 characters identical to the entity to be linked can be determined as candidate entities. This embodiment does not make any limitations.

[0044] The character similarity between the to-be-linked entity and each candidate entity in the entity library refers to the character overlap degree between the to-be-linked entity and each candidate entity in the entity library. The character similarity between the to-be-linked entity and each candidate entity in the entity library can be calculated in the following way:

[0045] Calculate the ratio of the first character count to the second character count, where the first character count is the number of overlapping characters between the to-be-linked entity and the candidate entity, and the second character count is the total number of characters of the to-be-linked entity and the candidate entity. Take the ratio of the first character count to the second character count as the character similarity between the to-be-linked entity and the candidate entity.

[0046] Exemplarily, if calculating the character similarity between "XX Company" and "XX Co., Ltd." (X represents a character), 4 characters in "XX Company" overlap with 4 characters in "XX Co., Ltd.", then the number of overlapping characters between "XX Company" and "XX Co., Ltd." is 8, the total number of characters is 10, and the character similarity between "XX Company" and "XX Co., Ltd." is 80%.

[0047] The word similarity between the to-be-linked entity and each candidate entity in the entity library refers to the word segmentation overlap degree between the to-be-linked entity and each candidate entity in the entity library. The word similarity between the to-be-linked entity and each candidate entity in the entity library can be calculated in the following way:

[0048] Perform word segmentation processing on the to-be-linked entity and each candidate entity in the entity library respectively. The Language Technology Platform (LTP) can be used to perform word segmentation processing on the to-be-linked entity and each candidate entity in the entity library. Calculate the ratio of the first word segmentation count to the second word segmentation count, where the first word segmentation count is the number of overlapping word segmentations between the to-be-linked entity and each candidate entity, and the second word segmentation count is the total number of word segmentations of the to-be-linked entity and the candidate entity. Take the ratio of the first word segmentation count to the second word segmentation count as the word segmentation similarity between the to-be-linked entity and the candidate entity.

[0049] Exemplarily, if calculating the word similarity between "XX Company" and "XX Co., Ltd." in the above embodiment, "XX Company" is segmented into 2 word segmentations "XX" and "Company" through LTP, and "XX Co., Ltd." is segmented into 3 word segmentations "XX", "Limited", and "Company" through LTP. 2 word segmentations in "XX Company" overlap with 2 word segmentations in "XX Co., Ltd.", then the number of overlapping word segmentations between "XX Company" and "XX Co., Ltd." is 4, and the total number of word segmentations between "XX Company" and "XX Co., Ltd." is 5. Then the word similarity between "XX Company" and "XX Co., Ltd." is 80%.

[0050] The semantic similarity between the to-be-linked entity and each candidate entity in the entity library refers to extracting the semantic features of the to-be-linked entity and each candidate entity in the entity library, and calculating the similarity between the semantic features of the to-be-linked entity and the semantic features of each candidate entity in the entity library. The semantic similarity between the to-be-linked entity and each candidate entity in the entity library can be calculated in the following ways:

[0051] The semantic vectors of the to-be-linked entity and each candidate entity can be determined respectively, and the similarity between the semantic vector of the to-be-linked entity and the semantic vector of each candidate entity can be calculated as the semantic similarity between the to-be-linked entity and the candidate entity;

[0052] The LTP word segmentation technology can also be used to perform word segmentation processing on the to-be-linked entity and each candidate entity in the entity library respectively, extract the key entity word segments of the to-be-linked entity and each candidate entity, and then convert the key entity word segments of the to-be-linked entity and the key entity word segments of the candidate entity into vectors. Based on the distance between the vectors corresponding to the key entity word segments of the to-be-linked entity and the key entity word segments of the candidate entity, the semantic similarity between the to-be-linked entity and each candidate entity in the entity library can be determined.

[0053] S102. Determine the entity similarity between the to-be-linked entity and each candidate entity according to the character similarity, word similarity, and semantic similarity between the to-be-linked entity and each candidate entity.

[0054] Based on the character similarity, word similarity, and semantic similarity between the to-be-linked entity and each candidate entity obtained in the above steps, further determine the entity similarity between the to-be-linked entity and each candidate entity.

[0055] In an optional embodiment, the maximum similarity among the character similarity, word similarity, and semantic similarity between the to-be-linked entity and each candidate entity can be taken as the entity similarity between the to-be-linked entity and the candidate entity.

[0056] In another optional embodiment, the average value of the character similarity, word similarity, and semantic similarity between the to-be-linked entity and each candidate entity can also be calculated as the entity similarity between the to-be-linked entity and the candidate entity.

[0057] In another optional embodiment, weights can also be assigned to the character similarity, word similarity, and semantic similarity between the to-be-linked entity and each candidate entity, and then the weighted sum of the character similarity, word similarity, and semantic similarity between the to-be-linked entity and each candidate entity can be calculated as the entity similarity between the to-be-linked entity and the candidate entity. Among them, the weights of the character similarity, word similarity, and semantic similarity can be adjusted and set according to the actual situation, and this embodiment does not make any limitations.

[0058] S103. Determine the candidate entity with the highest entity similarity to the entity to be linked as the linked entity corresponding to the entity to be linked.

[0059] In the embodiments of the present application, the entity similarity between the entity to be linked and each candidate entity is obtained, and the candidate entity with the highest entity similarity to the entity to be linked can be used as the linked entity corresponding to the entity to be linked.

[0060] Regarding whether to link the entity to be linked with the linked entity, the following judgment needs to be made:

[0061] If the entity similarity of the linked entity is greater than the preset first similarity, the entity to be linked can be linked with the linked entity; if the entity similarity of the linked entity is less than the preset first similarity and greater than the preset second similarity, manual discrimination is required, and it is determined whether to link the entity to be linked with the linked entity according to the result of manual discrimination. If the result of manual discrimination is to agree to link, the entity to be linked is linked with the linked entity. If the result of manual discrimination is not to agree to link, no linking operation is performed. If the result of manual discrimination is not to agree to link and indicates addition, the entity to be linked is added to the entity library; if the entity similarity of the linked entity is less than the preset second similarity, no linking operation is performed and manual discrimination is carried out. If the result of manual discrimination is to indicate addition, the entity to be linked is added to the entity library.

[0062] In the solution of the present application, the similarity between the entity to be linked and the candidate entity is determined from three dimensions: character similarity, word similarity, and semantic similarity, effectively improving the accuracy of similarity calculation. For short texts, even if the short texts have characteristics such as short length, scarce features, and lack of rich context, the solution of the present application can also determine the entity similarity between the entity to be linked and the candidate entity from the three dimensions of character similarity, word similarity, and semantic similarity, effectively improving the accuracy of similarity calculation for short texts and meeting the entity linking requirements of short texts.

[0063] In another embodiment of the present application, the entity library in the above embodiment is an entity library that has been pre-cleaned according to a preset entity cleaning method. Before calculating the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library in the steps of the above embodiment, the following steps are further included:

[0064] Clean the entity to be linked according to the preset entity cleaning method.

[0065] The entity library in this embodiment is an entity library that has been pre-cleaned according to a preset entity cleaning method. Before calculating the character similarity, word similarity, and semantic similarity between the entity to be linked and any candidate entity in the entity library, the entity to be linked is also cleaned according to the same entity cleaning method.

[0066] The above cleaning of entities refers to normalizing the entities. The cleaning steps mainly include removing special characters, normalizing numbers and letters, and removing duplicate suffixes, etc., so as to obtain regular entities and entity libraries.

[0067] It should be noted that the entities in the entity library can be cleaned before calculating the character similarity, word similarity, and semantic similarity between the entity to be linked and any candidate entity in the entity library each time. However, since there are a large number of entities stored in the entity library, in order to reduce the calculation amount and improve the calculation speed, in this embodiment, it is preferred to clean the entities in the entity library once in advance. Each time entity linking is performed, only the entity to be linked needs to be cleaned before calculating the character similarity, word similarity, and semantic similarity between the entity to be linked and any candidate entity in the entity library.

[0068] It should also be noted that the cleaning of the entity library in this embodiment is completed on the basis of the copied entity library, and the original entity library is retained. When entity linking is performed, the entity to be linked is linked to the original entity library. Correspondingly, the cleaned entity to be linked is linked to the copied entity library. If it is necessary to add the entity to be linked to the entity library, the entity to be linked needs to be added to the original entity library. Correspondingly, the cleaned entity to be linked is added to the copied entity library.

[0069] In this embodiment, cleaning the entity to be linked and the entities in the entity library in the same way can ensure the consistency of the entity to be linked and the entities in the entity library, and improve the accuracy of calculating the entity similarity between the entity to be linked and the candidate entity.

[0070] Further, the steps of the above embodiment clean the entity to be linked according to a preset entity cleaning method, including:

[0071] Remove special characters in the entity to be linked. In this embodiment, special character removal algorithms are used to remove special characters in the entity to be linked, such as "-", "\", "·", and punctuation marks such as space characters. Exemplarily, after removing the special character "·" from "XXX·XXX" (X represents a character), it becomes "XXXXXX".

[0072] Normalize the numbers and letters in the entity to be linked. In this embodiment, the numbers and letters in the entity to be linked are adjusted to a fixed format. For example, numbers are uniformly in Arabic numerals, and letters are uniformly in lowercase letters.

[0073] Eliminate the word segments in the entity to be linked that have little impact on the semantic features of the entity to be linked. Specifically, in the entity to be linked, there may be word segments that have little impact on the semantic features of the entity to be linked. The existence of these word segments will not only increase the computational amount but also may affect the accuracy of entity similarity. For example, the suffix word segments that repeatedly appear in certain types of entities to be linked have little impact on the semantic features of the entity to be linked. For example, in the organization name, suffixes such as "Company" and "Limited Company". Therefore, in this embodiment, the word segments in the entity to be linked that have little impact on the semantic features of the entity to be linked are eliminated.

[0074] Exemplarily, the entity to be linked can be cleaned in the order of eliminating special characters in the entity to be linked → normalizing numbers and letters in the entity to be linked → eliminating the word segments in the entity to be linked that have little impact on the semantic features of the entity to be linked.

[0075] Furthermore, as Figure 2 shown, the word segments in the entity to be linked that have little impact on the semantic features of the entity to be linked can be eliminated through the following steps:

[0076] S201. Perform word segmentation on the entity to be linked to obtain multiple target word segments of the entity to be linked.

[0077] In the embodiment of the present application, perform word segmentation on the entity to be linked to obtain multiple target word segments of the entity to be linked. Exemplarily, the LTP word segmentation technology can be used to perform word segmentation on the entity to be linked to obtain multiple target word segments of the entity to be linked.

[0078] S202. Determine the elimination degree of each target word segment.

[0079] The above-mentioned elimination degree refers to the degree to which each target word segment can be eliminated. The higher the elimination degree of the target word segment, the easier it is for the target word segment to be eliminated.

[0080] In an alternative embodiment, the importance degree of the target word segment in the entity to be linked can be determined through the term frequency–inverse document frequency (TF-IDF) technology. The lower the importance degree of the target word segment in the entity to be linked, the higher the elimination degree of the target word segment. Furthermore, the target word segment with a higher elimination degree can be eliminated.

[0081] The TF-IDF technology is a commonly used technology in information retrieval and data mining. TF represents the term frequency, and the calculation formula for the TF value of each target word segment is:

[0082] TF = S1 / S2

[0083] Among them, S1 represents the number of times the target word segment appears in the entity library, and S2 represents the total number of word segments in the entity library.

[0084] IDF represents the inverse document frequency. The calculation formula for the IDF value of each target word segment is:

[0085] IDF = log(L1 / L2 + 1)

[0086] Among them, L1 represents the total number of entities in the entity library, and L2 represents the number of entities in the entity library that contain the target word segment.

[0087] Calculate the product of the TF value and the IDF value to obtain the TF-IDF value of the target word segment, that is, the importance of the target word segment in the entity to be linked. The lower the importance of the target word segment in the entity to be linked, the higher the elimination degree of the target word segment. The reciprocal of the TF-IDF value can be taken as the elimination degree of the corresponding target word segment, and then the target word segments with a higher elimination degree can be eliminated.

[0088] In addition, the positions of the target word segments in the entity to be linked that have little influence on the semantic features of the entity to be linked are generally relatively fixed. For example, the target word segments at the suffix positions of organization names, weapon names, etc. have little influence on the semantic features of the entity to be linked. For example, in an organization name, the target word segments at the suffix positions such as "Limited" and "Company" can be eliminated. The above method of determining the elimination degree of the target word segment based on the frequency of the target word segment appearing in the entity library does not consider the influence of the position of the target word segment on the elimination degree. Therefore, in another alternative embodiment, the elimination degree of the target word segment can be determined by combining the position of the target word segment in the entity to be linked on the basis of the TF-IDF technology.

[0089] S203. Eliminate the target word segments in the entity to be linked whose elimination degree is greater than the first preset elimination degree threshold to obtain the cleaned entity to be linked.

[0090] In the embodiments of the present application, the target word segments in the entity to be linked whose elimination degree is greater than the first preset elimination degree threshold are eliminated to obtain the cleaned entity to be linked. It should be noted that the actual value of the first preset elimination degree threshold can be adjusted and set according to the actual situation, and this embodiment does not make a limitation.

[0091] This embodiment can eliminate the target word segments in the entity to be linked that have little influence on the semantic features of the entity to be linked, and improve the calculation accuracy of the entity similarity between the entity to be linked and the candidate entity.

[0092] It should be noted that the cleaning process of each entity in the entity library is the same as the cleaning process of the entity to be linked in the above embodiment. Those skilled in the art can refer to the cleaning process of the entity to be linked, and this embodiment will not be elaborated here.

[0093] In another embodiment of the present application, as Figure 3 shown, the steps of the above embodiment determine the elimination degree of each target word segment, and can be specifically implemented through the following steps:

[0094] S301. Based on the target entities in the entity library, count the entity order frequencies of each target word segment.

[0095] The above target entities are the entities in the entity library that contain the target word segments; the above entity order frequencies are the frequencies in the entity library of the order in which the target word segments are located in the target entities.

[0096] Exemplarily, the entity to be linked is "XX Co., Ltd.", and the entity to be linked is segmented to obtain 4 target word segments: "XX", "Stock", "Limited", and "Company". For the target word segment "Stock", the target entities are all the entities in the entity library that contain the two characters "Stock".

[0097] It should be noted that for entities to be linked such as organization names and weapon names, the target word segments that have little impact on the semantic features of the entities to be linked generally appear at the end of the entities to be linked such as organization names and weapon names. For example, in an organization name, target word segments such as "Limited" and "Company" that have little impact on the semantic features of the organization name generally appear at the end of the organization name. "Company" is generally the last word segment, and "Limited" is generally the second-to-last or third-to-last word segment.

[0098] If the order is sorted starting from the last target word segment of the target entity, then the positions of the target word segments that have little impact on the semantic features of the entity to be linked will be concentrated in the first, second, or third positions of all the word segments of the target entity, which is more convenient for calculating the entity order frequencies of the target word segments; if the order is sorted starting from the first target word segment of the target entity, due to the different lengths of the names of the entities to be linked, the positions of the target word segments that have little impact on the semantic features of the entity to be linked may be scattered in the fourth, fifth, sixth, or even later positions of all the word segments of the target entity, and the order of the target word segments in the target entity is not concentrated, resulting in inconvenient calculation of the entity order frequencies of the target word segments.

[0099] Therefore, in order to facilitate the statistics of the entity order frequencies of each target word segment, in this embodiment, the order is sorted starting from the last word segment of the target entity. For example, if the word segments of the target entity are "XX", "Stock", "Limited", and "Company", then "Company" is in the first order, "Limited" is in the second order, "Stock" is in the third order, and "XX" is in the fourth order.

[0100] Exemplarily, if the entity to be linked is "XX Co., Ltd.", the entity to be linked is segmented, and 4 target segments "XX", "Stock", "Limited", and "Company" are obtained. For the target segment "Stock", if in the target entities corresponding to "Stock", there are 10 target entities where "Stock" is in the second order and 20 target entities where "Stock" is in the third order, then the entity order frequency of the target segment "Stock" in the second order in the target entities is 20, and the entity order frequency of the target segment "Stock" in the third order in the target entities is 30.

[0101] It should be noted that in the above embodiments, the description that there are 10 target entities where "Stock" is in the second order and 20 target entities where "Stock" is in the third order in the entity library is only used to exemplarily assist in explaining the meaning of the entity order frequency and does not form any limitation.

[0102] S302. Determine the order weight of each target segment in the target entity.

[0103] In the embodiments of the present application, it is also necessary to further determine the order weight of each target segment in the target entity.

[0104] The order weight of each target segment in the target entity is used to represent the possibility of the target segment located in this order being removed. The greater the order weight of a certain order of a target segment in the target entity, the greater the possibility of the target segment located in this order being removed. In this embodiment, the order sorting starts from the last segment of the target entity. According to the sorting method of this embodiment, the earlier the order of each target segment in the target entity, the smaller the influence of the target segment on the semantic features of the entity to be linked, the greater the possibility of its being removed, and correspondingly, the greater the order weight of the order where the target segment is located.

[0105] If there is such a target segment that is not only in the front order in the target entity but also has a high entity order frequency, it means that the target segment not only has a small influence on the semantic features of the entity to be linked and has a high possibility of being removed, but also appears repeatedly many times in the entity library and belongs to a repeatedly occurring suffix. Then such a target segment should correspond to a higher removal degree to facilitate its removal.

[0106] Exemplarily, it is possible to first determine the order weight corresponding to each order of the target segments in the entity to be linked, and then determine the order weight of each target segment in the target entity according to the order weight corresponding to each order of the target segments in the entity to be linked.

[0107] S303. Use the order weight of the order in which each target word segment is located in the target entity as the weight of the corresponding entity order frequency, and calculate the weighted sum of the entity order frequencies corresponding to each target word segment.

[0108] The entity order frequency of each target word segment represents the frequency corresponding to the order in which each target word segment is located in the target entity. The magnitude of the order weight of the order in which each target word segment is located in the target entity represents the likelihood of the target word segment at that order being removed. In the embodiments of the present application, given the entity order frequency of each target word segment and the order weight of the order in which each target word segment is located in the target entity, the removal degree of each target word segment in the entity to be linked can be calculated. For example, calculate the weighted sum of the entity order frequencies corresponding to each target word segment to obtain the removal degree of that target word segment.

[0109] Specifically, the weighted sum of the entity order frequencies corresponding to each target word segment can be calculated. Use the order weight of the order in which each target word segment is located in the target entity as the weight of the corresponding entity order frequency. The calculation formula is as follows:

[0110] F = S1×D1 + S2×D2 + …… + S n ×D n

[0111] Among them, F represents the weighted sum of the entity order frequencies corresponding to each target word segment, D n represents the entity order frequencies corresponding to that target word segment, and S n represents the weight of the entity order frequencies corresponding to that target word segment.

[0112] Exemplarily, if the entity to be linked is "XX Co., Ltd.", perform word segmentation on the entity to be linked to obtain 4 target word segments: "XX", "stock", "limited", and "company"; assign a weight of 4 to the first order where "company" is located, assign a weight of 3 to the second order where "limited" is located, assign a weight of 2 to the third order where "stock" is located, and assign a weight of 1 to the fourth order where "XX" is located;

[0113] If in the target entities corresponding to the target participle "company", "company" is in the first order in all cases, and the number of target entities is 200, then the frequency of the entity order of "company" in the first order among the target entities is 200; in the target entities corresponding to the target participle "limited", "limited" is in the second and third orders, the number of target entities where "limited" is in the second order is 100, and the number of target entities where "limited" is in the third order is 100, then the frequency of the entity order of "limited" in the second order among the target entities is 100, and the frequency of the entity order of "limited" in the third order is 100; in the target entities corresponding to the target participle "stock", "stock" is in the third and fourth orders, the number of target entities where "stock" is in the third order is 90, and the number of target entities where "stock" is in the fourth order is 110, then the frequency of the entity order of "stock" in the third order among the target entities is 90, and the frequency of the entity order of "stock" in the fourth order is 110; in the target entities corresponding to the target participle "XX", "XX" is in the fourth and fifth orders, the number of target entities where "XX" is in the fourth order is 9, and the number of target entities where "XX" is in the fifth order is 5, then the frequency of the entity order of "XX" in the fourth order among the target entities is 9, and the frequency of the entity order of "XX" in the fifth order is 5.

[0114] Then there is:

[0115] The order weight of the target participle "company" in the first order among the target entities is the same as the order weight of the first order in the entity to be linked "XX Co., Ltd.", which is 4; the order weight of the target participle "limited" in the second order among the target entities is the same as the order weight of the second order in the entity to be linked "XX Co., Ltd.", which is 3; the order weight of the target participle "limited" in the third order among the target entities is the same as the order weight of the third order in the entity to be linked "XX Co., Ltd.", which is 2; the order weight of the target participle "stock" in the third order among the target entities is the same as the order weight of the third order in the entity to be linked "XX Co., Ltd.", which is 2; the order weight of the target participle "stock" in the fourth order among the target entities is the same as the order weight of the fourth order in the entity to be linked "XX Co., Ltd.", which is 1; the order weight of the target participle "XX" in the fourth order among the target entities is the same as the order weight of the fourth order in the entity to be linked "XX Co., Ltd.", which is 1; the target participle "XX" is in the fifth order among the target entities, and there is no same order situation in the entity to be linked, so this order does not allocate an order weight and does not participate in the subsequent calculation process.

[0116] Further, the weighted sum of the order frequencies of each entity corresponding to the target participle "company" is:

[0117] 200 × 4 = 800

[0118] The weighted sum of the order frequencies of each entity corresponding to the target participle "limited" is:

[0119] 100 × 3 + 100 × 2 = 500

[0120] The weighted sum of the order frequencies of each entity corresponding to the target participle "stock" is:

[0121] 90 × 2 + 110 × 1 = 290

[0122] The weighted sum of the order frequencies of each entity corresponding to the target participle "XX" is:

[0123] 9 × 1 = 9

[0124] S304. Use the word segmentation order weight of each target participle in the entity to be linked to correct the weighted sum corresponding to the target participle, and obtain the elimination degree corresponding to the target participle.

[0125] The word segmentation order weight of each target participle in the entity to be linked can be used to correct the weighted sum corresponding to the target participle, and the elimination degree corresponding to the target participle can be obtained. The purpose of using the word segmentation order weight of each target participle in the entity to be linked to correct the weighted sum corresponding to the target participle is to increase the influence of the word segmentation order weight of each target participle in the entity to be linked on the elimination degree, and then try to eliminate the target participles with a higher position in the entity to be linked. The correction formula is as follows:

[0126] Score T = F × h

[0127] Among them, Score T represents the elimination degree of each target participle, F represents the weighted sum of the order frequencies of each entity corresponding to the target participle, and h represents the word segmentation order weight of the target participle in the entity to be linked.

[0128] In the previous example, the elimination degree corresponding to the target participle "company" is:

[0129] 800 × 4 = 2400

[0130] The elimination degree corresponding to the target participle "limited" is:

[0131] 500 × 3 = 1500

[0132] The elimination degree corresponding to the target participle "stock" is:

[0133] 290 × 2 = 580

[0134] The rejection degree corresponding to the target participle "XX" is as follows:

[0135] 9×1=9

[0136] It can be clearly determined from the above examples that after correcting the weighted sum corresponding to the target participle using the participle order weight of each target participle in the entity to be linked, the rejection degree of the target participles at the front position in the entity to be linked increases significantly, effectively widening the gap between the rejection degree of the target participles at the front position and the rejection degree of the target participles at the rear position in the entity to be linked, and then trying to remove the target participles at the front position in the entity to be linked as much as possible.

[0137] In this embodiment, the target participles with a rejection degree greater than the first preset rejection degree threshold are removed from the entity to be linked. If the first preset rejection degree threshold is 500, then in the above embodiment, "company", "limited", and "corporation" are all removed, and the cleaned entity to be linked obtained is "XX". That is to say, after cleaning the entity to be linked "XX Co., Ltd.", "XX" is obtained.

[0138] In another alternative embodiment, the TF-IDF technology can also be improved in combination with the position of the target participles in the entity to be linked, and then the rejection degree of the target participles is calculated.

[0139] Specifically, based on the weighted sum of the order frequencies of each entity corresponding to the target participle, the position word frequency based on the position of the target participle is calculated:

[0140] TF W =F / S2

[0141] Among them, TF W represents the position word frequency of the target participle, F represents the weighted sum of the order frequencies of each entity corresponding to the target participle, and S2 represents the total number of participles in the entity library.

[0142] Calculate the product of the TF W value and the IDF value, and the TF W -IDF value of the target participle can be obtained. It should be noted that the calculation method of the IDF value is the same as the calculation method of the IDF value recorded when determining the rejection degree of the target participle through the TF-IDF value in the above embodiment. Those skilled in the art can refer to the records in the above embodiment and will not be elaborated here. Further, the TF W -IDF value corresponding to the target participle can be corrected using the participle order weight of each target participle in the entity to be linked to obtain the rejection degree of the corresponding target participle.

[0143] In addition, in this embodiment, the order sorting is performed starting from the last word segment of the entity to be linked. According to the arrangement order of this embodiment, generally only the first few target entities of the entity to be linked need to be excluded. In order to further save the calculation amount and improve the calculation speed, the exclusion degrees of only the first few target word segments of the entity to be linked can be calculated, and the target word segments with the exclusion degree greater than the first preset exclusion degree threshold are excluded from the entity to be linked, and the last few target word segments of the entity to be linked are not excluded. Specifically, the number of target word segments for which the exclusion degree needs to be calculated can be determined according to the number of word segments of the entity to be linked. Generally, the more the number of word segments of the entity to be linked, the more the number of target word segments for which the exclusion degree needs to be calculated. For example, if there are at least 5 target word segments in the entity to be linked, the exclusion degrees of the first three target word segments can be calculated; if the number of target word segments is between 3 and 5, the exclusion degrees of the first two target word segments can be calculated; if the number of target word segments is less than 3, only the exclusion degree of the first target word segment can be calculated.

[0144] In this embodiment, the gap between the exclusion degree of the target word segments at the front position in the entity to be linked and the exclusion degree of the target word segments at the rear position in the entity to be linked can be effectively widened, so as to exclude the target word segments at the front position in the entity to be linked as much as possible.

[0145] It should also be noted that the process of calculating the exclusion degree for each entity in the entity library is the same as the process of calculating the exclusion degree for the entity to be linked in the above embodiment. Those skilled in the art can refer to the process of calculating the exclusion degree for the entity to be linked, and this embodiment will not be elaborated here.

[0146] In another embodiment of the present application, as Figure 4 shown, the order weights of the positions of each target word segment in the target entity determined in the steps of the above embodiment can be specifically implemented through the following steps:

[0147] S401. Determine the order weight of the position of each target word segment in the entity to be linked according to the rule that the order weight of the position of the target word segment in the entity to be linked is negatively correlated with the order of the target word segment in the entity to be linked.

[0148] In the embodiments of the present application, in order to facilitate the statistics of the entity order frequencies of each target word segment, the order sorting is performed starting from the last word segment of the target entity. Correspondingly, in the entity to be linked, the order sorting is performed starting from the last target word segment of the entity to be linked.

[0149] According to the sorting rule of the target participles in the entity to be linked in this embodiment, the earlier the target participle is in the entity to be linked, the smaller the influence of the target participle on the semantic features of the entity to be linked. To increase the elimination degree of the target participles in the front position of the entity to be linked, the participle order weight assigned to the target participles in the front position of the entity to be linked should be greater than the participle order weight assigned to the target participles in the rear position of the entity to be linked. It should be noted that the order weight of each target participle in the entity to be linked can be configured as an integer or a fraction, which is not limited in this embodiment.

[0150] Exemplarily, the entity to be linked is "XX Co., Ltd.", and after performing participle processing on the entity to be linked, 4 target participles "XX", "stock", "limited", and "company" are obtained. According to the sorting rule of the target participles in the entity to be linked in this embodiment, the target participle "company" is in the first order, and the order weight corresponding to the first order is the highest. The target participle "limited" is in the second order, and the order weight corresponding to the second order is the second highest. The target participle "stock" is in the third order, and the order weight corresponding to the third order is the third highest. The target participle "XX" is in the fourth order, and the order weight corresponding to the fourth order is the smallest. For example, the weight 4 is assigned to the first order, the weight 3 is assigned to the second order, the weight 2 is assigned to the third order, and the weight 1 is assigned to the fourth order.

[0151] S402. For each target participle, respectively determine the order weight in the entity to be linked as the order weight of the same order where the target participle is located in the target entity.

[0152] In this embodiment, the calculation method of the order weight of each target participle in the order where it is located in the target entity is the same, and it is all:

[0153] Determine the order where the target participle is located in the target entity. If there is also the same order in the entity to be linked, further determine the order weight corresponding to the same order in the entity to be linked, and assign the above-mentioned order weight to the order where the target participle is located in the target entity.

[0154] Exemplarily, if the entity to be linked is "XX Co., Ltd.", and after performing participle processing on the entity to be linked, 4 target participles "XX", "stock", "limited", and "company" are obtained, the weight 4 is assigned to the first order, the weight 3 is assigned to the second order, the weight 2 is assigned to the third order, and the weight 1 is assigned to the fourth order. For the target participle "stock", if in the target entity corresponding to "stock", "stock" is only located in the second order and the third order, then the second order weight 3 in the entity to be linked can be assigned to the second order, and the third order weight 2 in the entity to be linked can be assigned to the third order.

[0155] In addition, if the length of the target entity is relatively long and the number of word segments of the target entity is greater than the number of target word segments in the entity to be linked, it is possible that after determining the order of the target word segment in the target entity, there is no same order in the entity to be linked. In such a case, the order of the target word segment in the target entity no longer assigns an order weight, that is, it does not participate in the subsequent calculation process.

[0156] Exemplarily, if the entity to be linked is "XX Co., Ltd.", after performing word segmentation on the entity to be linked, 4 target word segments "XX", "stock", "limited", and "company" are obtained. Taking "XX" as an example, there is such a target entity "XXYYYYY Co., Ltd.", which is segmented into 6 word segments "XX", "YYY", "YY", "limited", "liability", and "company". Among them, the word segment "XX" at the sixth order in the target entity is the same as the target word segment "XX" at the fourth order in the entity to be linked. The order of "XX" in the target entity is the sixth order, while there are only 4 target word segments in the entity to be linked and no sixth order. This is the situation described in this embodiment where after determining the order of the target word segment in the target entity, there is no same order in the entity to be linked. In such a case, the sixth order of "XX" in the target entity no longer assigns a weight, that is, it does not participate in the subsequent calculation process.

[0157] Specifically, in the embodiments of the present application, mainly the target word segments at the front order in the entity to be linked are removed. For example, in an organization name, target word segments such as "limited" and "company" that have little impact on the semantic features of the organization name. The target word segments at the front order in the entity to be linked are almost unlikely to have the situation where after determining the order of the target word segment in the target entity, there is no same order in the entity to be linked. The target word segments that are prone to this situation are at a relatively rear order in the entity to be linked, and such target word segments have a greater impact on the semantic features of the linked word segments and should not be removed but should be retained as much as possible. Further, in the embodiments of the present application, the removal degree of each target word segment is determined according to the entity order frequency of each target word segment and the order weight of the order in which each target word segment is located in the target entity. If the above situation occurs, the order of the target word segment in the target entity no longer assigns an order weight, that is, it does not participate in the subsequent calculation process, which will reduce the removal degree of the target word segment, so that the target word segment can be retained as much as possible and not be removed.

[0158] In this embodiment, for each target word segment, the order weight in the entity to be linked is respectively determined as the order weight of the same order in which the target word segment is located in the target entity. On the basis of combining the position of the target word segment, the removal degree of the target entity can be determined more accurately.

[0159] In another embodiment of the present application, asFigure 5 As shown in the above, the steps of the above embodiments calculate the semantic similarity between the entity to be linked and the candidate entity, which can be specifically implemented through the following steps:

[0160] S501. Respectively determine the key entity segmentations of the entity to be linked and the candidate entity.

[0161] The above-mentioned key entity segmentation refers to the segmentation in the entity to be linked that has a greater impact on the semantic features of the entity to be linked. In the embodiments of the present application, it is necessary to respectively determine the key entity segmentations of the entity to be linked and the candidate entity.

[0162] Exemplarily, if the entity to be linked and the candidate entity have been cleaned, then in this step, according to the elimination degrees of the target segmentations in the entity to be linked and the candidate entity, respectively determine the key entity segmentations of the entity to be linked and the candidate entity on the basis of the cleaned entity to be linked and the candidate entity. Exemplarily, the target segmentations with elimination degrees greater than the second preset elimination degree threshold in the entity to be linked and the candidate entity can be randomly covered to obtain the key entity segmentations of the entity to be linked and the candidate entity. The above-mentioned second preset elimination degree threshold can be set according to the actual situation, and this embodiment does not make a limitation. However, it should be noted that the second preset elimination degree threshold should be less than the first preset elimination degree threshold in the above embodiments.

[0163] Among them, the calculation method of the elimination degrees of the target segmentations in the entity to be linked and the candidate entity is the same as the calculation method of the elimination degree in the above embodiments, and those skilled in the art can refer to the records of the above embodiments and will not be elaborated here.

[0164] If the entity to be linked and the candidate entity have not been cleaned, then in this step, according to the elimination degrees of the target segmentations in the entity to be linked and the candidate entity, eliminate the target segmentations with elimination degrees greater than the first preset elimination degree threshold, and then further extract the key entity segmentations of the entity to be linked and the candidate entity according to the elimination degrees of the target segmentations in the entity to be linked and the candidate entity.

[0165] S502. Determine the first similarity between the key entity segmentations of the entity to be linked and the candidate entity, and determine the second similarity between the entity to be linked and the candidate entity.

[0166] In this embodiment, a first vector is generated based on the key entity segmentation of the entity to be linked, a second vector is generated based on the key entity segmentation of the candidate entity, and the similarity between the first vector and the second vector is calculated as the first similarity between the key entity segmentations of the entity to be linked and the candidate entity.

[0167] Among them, the calculation method of the first vector of the key entity word segmentation of the entity to be linked is as follows: calculate the word vectors of each character in the key entity word segmentation of the entity to be linked, and perform an averaging operation on all the word vectors of the key entity word segmentation of the entity to be linked to obtain the above-mentioned first vector. Similarly, the calculation method of the second vector of the key entity word segmentation of the candidate entity is as follows: calculate the word vectors of each character in the key entity word segmentation of the candidate entity, and perform an averaging operation on all the word vectors of the key entity word segmentation of the candidate entity to obtain the above-mentioned second vector.

[0168] Generate a third vector based on the entity to be linked, generate a fourth vector based on the candidate entity, and calculate the similarity between the third vector and the fourth vector as the second similarity between the entity to be linked and the candidate entity.

[0169] Among them, the calculation method of the third vector of the entity to be linked is as follows: calculate the word vectors of each character in the entity to be linked, and perform an averaging operation on all the word vectors of the entity to be linked to obtain the above-mentioned third vector. Similarly, the calculation method of the fourth vector of the candidate entity is as follows: calculate the word vectors of each character in the candidate entity, and perform an averaging operation on all the word vectors of the candidate entity to obtain the above-mentioned fourth vector.

[0170] Exemplarily, calculate the cosine similarity between the first vector and the second vector as the similarity between the first vector and the second vector, that is, obtain the first similarity; calculate the cosine similarity between the third vector and the fourth vector as the similarity between the third vector and the fourth vector, that is, obtain the second similarity.

[0171] S503. Take the maximum similarity between the first similarity and the second similarity as the semantic similarity between the entity to be linked and the candidate entity.

[0172] Compare the magnitudes of the first similarity and the second similarity, and take the maximum similarity between the first similarity and the second similarity as the semantic similarity between the entity to be linked and the candidate entity.

[0173] The first similarity in this embodiment is determined based on the key entity word segmentations of the entity to be linked and the candidate entity, which can effectively avoid the interference of non-key entity word segmentations on semantics and improve the accuracy of semantic similarity calculation. And taking the maximum similarity between the first similarity and the second similarity as the semantic similarity between the entity to be linked and the candidate entity can avoid the problem of too small semantic similarity caused by incorrect extraction of key entity word segmentations, and further improve the accuracy of semantic similarity calculation.

[0174] In another embodiment of the present application, the steps of the above embodiment respectively determine the key entity word segmentations of the entity to be linked and the candidate entity, which can be specifically implemented through the following steps:

[0175] Randomly cover the target word segments with a deletion degree greater than the second preset deletion degree threshold among the entities to be linked and the candidate entities to obtain the key entity word segments.

[0176] The calculation method of the deletion degree of the target word segments in the entities to be linked and the candidate entities is the same as that of the deletion degree in the above embodiments. Those skilled in the art can refer to the records of the above embodiments and will not be elaborated here.

[0177] Specifically, among the entities to be linked and the candidate entities, the target word segments with a deletion degree greater than the first preset deletion degree threshold have been deleted during entity cleaning. In this embodiment, the target word segments with a deletion degree greater than the second preset deletion degree threshold in the entities to be linked and the candidate entities are randomly covered. If there are target word segments with a deletion degree greater than the first preset deletion degree threshold that are missed in this step, then the target word segments with a deletion degree greater than the first preset deletion degree threshold are covered. The target word segments with a deletion degree less than the second preset deletion degree threshold and the target word segments for which the deletion degree is not calculated in the entities to be linked and the candidate entities are not covered, thereby obtaining the key entity word segments.

[0178] Exemplarily, as Figure 6 shown, if the entity to be linked after cleaning is "XX", "YY", "Stock", among which the deletion degree corresponding to "Stock" is greater than the second preset deletion degree threshold, then "Stock" needs to be randomly covered. If the deletion degrees corresponding to "XX" and "YY" are less than the second preset deletion degree threshold, "XX" and "YY" cannot be covered. Then the key entity word segments extracted from this entity to be linked are: "XX", "YY", "MASK". Here, MASK indicates that the target word segment is covered.

[0179] This embodiment determines the coverage situation of the target word segments according to the deletion degrees of the target word segments in the entities to be linked and the candidate entities, further avoiding the interference of non-key entity word segments on the semantics and improving the accuracy of semantic similarity calculation.

[0180] In another embodiment of the present application, the steps of the above embodiments determine the first similarity between the key entity word segments of the entities to be linked and the candidate entities, and determine the second similarity between the entities to be linked and the candidate entities. Specifically, it can be implemented through the following steps:

[0181] Input the entity to be linked, the candidate entity, and the key entity word segments of the entity to be linked and the candidate entity into a pre-trained similarity calculation model to obtain the first similarity and the second similarity.

[0182] Specifically, the above similarity calculation model is obtained through similar entity differential training and similarity calculation training. Among them, the similar entity differential training is used to train the similarity calculation model's ability to identify similar entities and dissimilar entities; the similarity calculation training is used to train the similarity calculation model's ability to calculate the similarity of the key entity word segments of the entity to be linked and the candidate entity, and to calculate the similarity between the entity to be linked and the candidate entity.

[0183] Further, as Figure 7 shown, the similar entity differential training in the above embodiment includes:

[0184] S701. Input the first sample into the similarity calculation model to obtain the first result output by the similarity calculation model.

[0185] Specifically, the above first sample includes at least a pair of similar entities. Exemplarily, there are only a few characters different in the target word segments that have a greater impact on the semantic features in Sentence_A and Sentence_B, and Sentence_A and Sentence_B are a pair of similar entities. For example, there is only one character different between "XY Co., Ltd." and "XZ Co., Ltd.", which is a pair of similar entities.

[0186] Sentence_A and Sentence_B can be replaced and split to expand the first sample. Figure 8 It is a schematic diagram of the construction of the similar entity differential training sample provided by the embodiment of the present application. As Figure 8 shown, Enity_A is the target word segment extracted from Sentence_A that has a greater impact on the semantic features of Sentence_A. For example, "XZ" is extracted from "XZ Co., Ltd." that has a greater impact on the semantic features of "XZ Co., Ltd."; Enity_B is the target word segment extracted from Sentence_B that has a greater impact on the semantic features of Sentence_B. For example, "XY" is extracted from "XY Co., Ltd." that has a greater impact on the semantic features of "XY Co., Ltd.". Enity_A+Sentence_B means replacing Enity_B in Sentence_B with Enity_A; Enity_B+Sentence_A means replacing Enity_A in Sentence_A with Enity_B. Among them, Enity_A+Sentence_A and Enity_A+Sentence_B are similar, and Enity_B+Sentence_B and Enity_A+Sentence_B are not similar.

[0187] The purpose of differential training of similar entities is to train the ability of the similarity calculation model to identify similar entities and dissimilar entities. The contrastive learning framework provided in this embodiment learns the feature representation of samples by comparing data with positive and negative samples in the feature space, and pulls the distance between positive and negative samples farther apart. Specifically, in order to improve the ability of the similarity calculation model to identify similar entities and dissimilar entities, the gap between the number of similar entities and the number of dissimilar entities in the first sample in a training batch should be maximized. Exemplarily, in a training batch, there can be only a pair of similar entities, Enity_A+Sentence_A and Enity_A+Sentence_B, and Enity_A+Sentence_A is not similar to other entities.

[0188] The above first result is the result of calculating the semantic similarity between every two entities in the first sample.

[0189] S702. Determine the first loss value of the similarity calculation model according to the result of calculating the semantic similarity between every two entities in the first sample and the first label.

[0190] The above first label is the semantic similarity label between every two entities in the first sample. Based on the result of calculating the semantic similarity between every two entities in the first sample and the first label, the first loss value of the similarity calculation model can be determined.

[0191] S703. Adjust the operation parameters of the similarity calculation model according to the first loss value.

[0192] In this embodiment, the operation parameters of the similarity calculation model are adjusted with the aim of reducing the first loss value.

[0193] S704. Repeat the above process until the first loss value of the similarity calculation model is less than the first set value.

[0194] After adjusting the operation parameters of the similarity calculation model, continue to input the first sample to train the similarity calculation model, and then obtain the first result again. According to this first result and the first label, determine the first loss value of the similarity calculation model again, and adjust the operation parameters of the similarity calculation model again with the aim of reducing the first loss value. Repeat this process until the first loss value of the similarity calculation model is less than the first set value.

[0195] Among them, the above first set value can be set according to the actual situation, and this embodiment does not make any limitations.

[0196] In this embodiment, by setting at least a pair of similar entities in the first sample and performing differential training on the similarity calculation model for similar entities, the ability of the similarity calculation model to identify similar entities and dissimilar entities can be effectively improved.

[0197] Further, as Figure 9 shown, the similarity calculation training of the above embodiment includes:

[0198] S901. Input the second sample into the similarity calculation model to obtain a second result output by the similarity calculation model.

[0199] The above second sample includes multiple entity groups, each entity group includes a pair of similar entities, and key entity word segments extracted from the similar entities. The steps of extracting key entity word segments from the similar entities are the same as those recorded when extracting key entity word segments from the entities to be linked or candidate entities in the above embodiment. Those skilled in the art can refer to the records of the above embodiment and will not be elaborated here.

[0200] The above second result is the first similarity calculation result between the similar entities in each entity group, and the second similarity calculation result between the key entity word segments in each entity group.

[0201] Among them, the process of the similarity calculation model calculating the first similarity calculation result between the similar entities in each entity group and the second similarity calculation result between the key entity word segments in each entity group is as follows:

[0202] Determine the word vectors of each entity in each entity group, perform an averaging operation on the word vectors of each entity to obtain the vector of each entity, calculate the similarity between the vectors of the two similar entities in each entity group, such as cosine similarity, to obtain the first similarity calculation result between the similar entities in each entity group; determine the word vectors of each key entity word segment in each entity group, perform an averaging operation on the word vectors of each key entity word segment to obtain the vector of each key entity word segment, calculate the similarity between the vectors of the two key entities in each entity group, such as cosine similarity, to obtain the second similarity calculation result between the key entity word segments in each entity group.

[0203] S902. Take the maximum similarity in the first similarity calculation result and the second similarity calculation result corresponding to each entity group as the semantic similarity of the corresponding entity group.

[0204] Compare the magnitudes of the first similarity calculation result and the second similarity calculation result corresponding to each entity group, and take the maximum similarity in each entity group as the semantic similarity of the corresponding entity group.

[0205] S903. Determine the second loss value of the similarity calculation model based on the semantic similarity of all entity groups in the second sample and the second label.

[0206] The above-mentioned second label is the semantic similarity label of all entity groups in the second sample. Based on the semantic similarity of all entity groups in the second sample and the second label, the second loss value of the similarity calculation model can be determined.

[0207] S904. Adjust the operation parameters of the similarity calculation model according to the second loss value.

[0208] In this embodiment, the operation parameters of the similarity calculation model are adjusted with the aim of reducing the second loss value.

[0209] S905. Repeat the above process until the second loss value of the similarity calculation model is less than the second set value.

[0210] After adjusting the operation parameters of the similarity calculation model, continue to input the second sample to train the similarity calculation model, and then obtain the second result again. According to this second result and the second label, determine the second loss value of the similarity calculation model again, and adjust the operation parameters of the similarity calculation model again with the aim of reducing the second loss value, and so on, until the second loss value of the similarity calculation model is less than the second set value.

[0211] Among them, the above-mentioned second set value can be set according to the actual situation, and this embodiment does not make any limitations.

[0212] In this embodiment, performing similarity calculation training on the similarity calculation model can effectively improve the ability of the similarity calculation model to calculate the similarity of the key entity segmentations between the entity to be linked and the candidate entity, and to calculate the similarity between the entity to be linked and the candidate entity.

[0213] In another embodiment of the present application, a BERT model pre-trained based on Whole Word Masking (WWM) is used as the basic model of the similarity calculation model. The corpus required for the pre-training process of the BERT model is unlabeled Chinese corpus, which can be extracted and cleaned from Chinese websites.

[0214] In another embodiment of the present application, as Figure 10 shown, the similarity calculation model of the above embodiment includes an adapter component after the first feed-forward network and the second feed-forward network.

[0215] The internal structure of the above-mentioned adapter component is as Figure 11As shown, assume that the dimension of the input input is d. First, the dimension is changed to m through a feed-forward neural network for dimensionality reduction, where m is much smaller than d. Then, the output output is obtained through a feed-forward neural network for dimensionality increase, and the dimension of the output becomes d again. Finally, a residual connection is performed, that is, input + output is used as the output result of the adapter component.

[0216] Furthermore, the steps of adjusting the operation parameters of the similarity calculation model in the above embodiments can be specifically implemented through the following steps:

[0217] Adjust the parameters of the adapter component.

[0218] Specifically, in order to obtain faster calculation efficiency and a lighter similarity calculation model, an adapter component is set after the feed-forward network of the similarity calculation model. During the differential training of similar entities and the similarity calculation training process, in the step of adjusting the operation parameters of the similarity calculation model, only the parameters corresponding to the adapter component need to be adjusted, and there is no need to adjust the operation parameters of the similarity calculation model. And the parameters corresponding to the adapter component are much smaller than the operation parameters of the similarity calculation model, thereby effectively improving the parameter adjustment efficiency.

[0219] Furthermore, the steps of determining the entity similarity between the entity to be linked and each candidate entity according to the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the above embodiments can be specifically implemented through the following steps:

[0220] Calculate the weighted sum of the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity, and use the weighted sum as the entity similarity between the entity to be linked and each candidate entity.

[0221] Specifically, the entity similarity H between the entity to be linked and each candidate entity can be calculated through the following formula:

[0222] H = H Z × S Z + H C × S C + H Y × S Y

[0223] In the above formula, H Z represents the character similarity between the entity to be linked and each candidate entity, H C represents the word similarity between the entity to be linked and each candidate entity, H Y represents the semantic similarity between the entity to be linked and each candidate entity; S Z represents the weight value of the character similarity, S C represents the weight value of the word similarity, S YThe weight value representing semantic similarity.

[0224] Among them, the weight value of character similarity, the weight value of word similarity, and the weight value of semantic similarity can be adjusted and set according to the actual situation, and this embodiment does not make any limitations.

[0225] In this embodiment, the similarity between the entity to be linked and the candidate entity is determined from three dimensions: character similarity, word similarity, and semantic similarity, effectively improving the accuracy of similarity calculation.

[0226] Furthermore, before the step of linking the entity to be linked with the target candidate entity in the above embodiment, the following steps may further be included:

[0227] Detect whether there is an entity in the entity library that matches the entity to be linked; if there is an entity in the entity library that matches the entity to be linked, then use the entity in the entity library that matches the entity to be linked as the linked entity.

[0228] Specifically, this embodiment can detect whether there is an entity in the entity library that matches the entity to be linked. If there is an entity in the entity library that matches the entity to be linked, it is not necessary to calculate the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity. Instead, directly use the entity in the entity library that matches the entity to be linked as the linked entity, and link the entity to be linked with the linked entity.

[0229] The above-mentioned entity that matches the entity to be linked includes a first type of entity and / or a second type of entity. The first type of entity includes the entity to be linked, and the entity to be linked is located at a set position in the first type of entity; the second type of entity is included by the entity to be linked, and the second type of entity is located at a set position in the entity to be linked.

[0230] The above-mentioned set position can be set according to the type of the entity to be linked. Exemplarily, if the entity to be linked is an organization name, in this embodiment, the last word segment of the entity to be linked is used as the starting point for sorting in order. According to the arrangement order of this embodiment, the set position should be a position closer to the end in the entity to be linked. For example, if the entity to be linked is "XX", and an entity in the entity library is "XX Co., Ltd.", "XX Co., Ltd." includes "XX", and "XX Co., Ltd." belongs to the first type of entity; on the contrary, if the entity to be linked is "XX Co., Ltd.", and an entity in the entity library is "XX", "XX Co., Ltd." includes "XX", and "XX" belongs to the second type of entity; furthermore, if the entity to be linked is "XX", and an entity in the entity library is "XX", and the entity to be linked is exactly the same as an entity in the entity library, then this entity in the entity library belongs to both the first type of entity and the second type of entity at the same time.

[0231] Exemplarily, it is possible to first detect whether there is an entity that belongs to both the first type of entity and the second type of entity in the entity library, that is, to detect whether there is an entity in the entity library that is exactly the same as the entity to be linked. If there is an entity in the entity library that is exactly the same as the entity to be linked, then this entity in the entity library is used as the linked entity. If there is no entity in the entity library that is exactly the same as the entity to be linked, then it is further detected whether there is a first type of entity or a second type of entity in the entity library. In order to reduce the calculation amount and save time, it is possible to first confirm the candidate entities of the entity to be linked from the entity library, and then detect whether there is a first type of entity or a second type of entity among the candidate entities. If there is a first type of entity or a second type of entity among the candidate entities, then this entity in the candidate entities is used as the linked entity. If there is no first type of entity and second type of entity among the candidate entities, then it is necessary to calculate the entity similarity between the entity to be linked and each candidate entity one by one according to the description of the above embodiments, so as to determine the linked entity from the candidate entities according to the entity similarity between the entity to be linked and the candidate entities.

[0232] It should be noted that when detecting whether there is an entity in the entity library that matches the entity to be linked, the entity to be linked and the entities in the entity library need to be in the same language environment. For example, in the same Chinese environment, to ensure the accuracy of the entity extracted from the entity library that matches the entity to be linked.

[0233] In this embodiment, first detect whether there is an entity in the entity library that matches the entity to be linked. If there is no entity in the entity library that matches the entity to be linked, then calculate the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library, which can effectively reduce the calculation amount in the entity linking process.

[0234] Corresponding to the above method for entity linking, an embodiment of the present application also discloses an apparatus for entity linking. Refer to Figure 12 As shown, the apparatus includes:

[0235] A calculation module 100, configured to calculate the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library respectively;

[0236] A first determination module 110, configured to determine the entity similarity between the entity to be linked and each candidate entity according to the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity;

[0237] A second determination module 120, configured to determine the candidate entity with the highest entity similarity to the entity to be linked as the linked entity corresponding to the entity to be linked.

[0238] For the device for entity linking in this application, the calculation module 100 calculates the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library respectively; the first determination module 110 determines the entity similarity between the entity to be linked and each candidate entity according to the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity; the second determination module 120 determines the candidate entity with the highest entity similarity to the entity to be linked as the linked entity corresponding to the entity to be linked. In the solution of this application, the similarity between the entity to be linked and the candidate entity is determined from three dimensions of character similarity, word similarity, and semantic similarity, effectively improving the accuracy of similarity calculation. For short texts, even if the short texts have characteristics such as short length, scarce features, and lack of rich context, the solution of this application can also determine the entity similarity between the entity to be linked and the candidate entity from three dimensions of character similarity, word similarity, and semantic similarity, effectively improving the accuracy of similarity calculation for short texts and meeting the entity linking requirements of short texts.

[0239] Optionally, in another embodiment of this application, the entity library is an entity library that has been pre-cleaned according to a preset entity cleaning method; the device for entity linking further includes:

[0240] A cleaning module, configured to clean the entity to be linked according to a preset entity cleaning method.

[0241] Optionally, in another embodiment of this application, the cleaning module includes:

[0242] A word segmentation unit, configured to perform word segmentation processing on the entity to be linked to obtain multiple target word segments of the entity to be linked;

[0243] A determination unit, configured to determine the elimination degree of each target word segment;

[0244] An elimination unit, configured to eliminate the target word segments with an elimination degree greater than the first preset elimination degree threshold from the entity to be linked to obtain the cleaned entity to be linked.

[0245] Optionally, in another embodiment of this application, the determination unit includes:

[0246] A statistics subunit, configured to statistically count the entity order frequency of each target word segment based on the target entities in the entity library; where the target entity is an entity containing the target word segment, and the entity order frequency is the frequency of the order in which the target word segment is located in the target entity in the entity library;

[0247] A first determination subunit, configured to determine the order weight of the order in which each target word segment is located in the target entity;

[0248] A calculation subunit, configured to use the order weight of the position of each target token in the target entity as the weight of the corresponding entity order frequency, and calculate the weighted sum of the respective entity order frequencies corresponding to the target token;

[0249] A correction subunit, configured to use the token order weight of each target token in the entity to be linked to correct the weighted sum corresponding to the target token, and obtain the elimination degree corresponding to the target token.

[0250] Optionally, in another embodiment of the present application, the calculation module 100 includes:

[0251] A third determination module, configured to respectively determine the key entity tokens of the entity to be linked and the candidate entity;

[0252] A fourth determination module, configured to determine the first similarity between the key entity tokens of the entity to be linked and the candidate entity, and determine the second similarity between the entity to be linked and the candidate entity;

[0253] A fifth determination module, configured to take the maximum similarity between the first similarity and the second similarity as the semantic similarity between the entity to be linked and the candidate entity.

[0254] Optionally, in another embodiment of the present application, when the third determination module respectively determines the key entity tokens of the entity to be linked and the candidate entity, specifically:

[0255] Randomly cover the target tokens in the entity to be linked and the candidate entity whose elimination degree is greater than the second preset elimination degree threshold to obtain the key entity tokens; the elimination degree is obtained according to the entity order frequency of each target token and the order weight of the position of each target token in the target entity; the target tokens are obtained by tokenizing the entity to be linked and the candidate entity, the target entity is the entity containing the target tokens, and the entity order frequency is the frequency of the position where the target token is located in the target entity in the entity library.

[0256] Optionally, in another embodiment of the present application, when the fourth determination module determines the first similarity between the key entity tokens of the entity to be linked and the candidate entity, and determines the second similarity between the entity to be linked and the candidate entity, specifically:

[0257] Input the entity to be linked, the candidate entity, and the key entity tokens of the entity to be linked and the candidate entity into a pre-trained similarity calculation model to obtain the first similarity and the second similarity; wherein, the similarity calculation model is obtained through similar entity differentiation training and similarity calculation training; the similar entity differentiation training is used to train the ability of the similarity calculation model to identify similar entities and dissimilar entities.

[0258] Optionally, in another embodiment of the present application, the similarity calculation model includes an adapter component disposed after the feedforward network; during the process of the first training module and the second training module performing similar entity differentiation training and similarity calculation training on the similarity calculation model, the parameters of the adapter component are adjusted.

[0259] Specifically, for the specific working content of each unit of the above-mentioned entity linking device, please refer to the content of the above method embodiment, which will not be elaborated here.

[0260] Another embodiment of the present application also proposes an electronic device, as shown in Figure 13 shown, the device includes:

[0261] A memory 200 and a processor 210;

[0262] Wherein, the memory 200 is connected to the processor 210 and is used to store programs;

[0263] The processor 210 is configured to implement the entity linking method disclosed in any of the above embodiments by running the programs stored in the memory 200.

[0264] Specifically, the above-mentioned electronic device may further include: a bus, a communication interface 220, an input device 230, and an output device 240.

[0265] The processor 210, the memory 200, the communication interface 220, the input device 230, and the output device 240 are interconnected through the bus. Among them:

[0266] The bus may include a path for transmitting information between various components of the computer system.

[0267] The processor 210 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application solution. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0268] The processor 210 may include a main processor and may also include a baseband chip, a modem, etc.

[0269] The program for implementing the technical solution of this application is stored in the memory 200, and the operating system and other key services may also be stored. Specifically, the program may include program codes, and the program codes include computer operation instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, and so on.

[0270] The input device 230 may include devices for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.

[0271] The output device 240 may include devices for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.

[0272] The communication interface 220 may include devices of any transceiver type for communicating with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0273] The processor 210 executes the program stored in the memory 200 and calls other devices, and can be used to implement each step of the method for entity linking provided in the above embodiments of this application.

[0274] Another embodiment of this application also provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, each step of the method for entity linking provided in any of the above embodiments is implemented.

[0275] Specifically, for the specific processing content of the above-mentioned electronic device and the computer program on the above-mentioned storage medium when run by a processor, reference can be made to the content of each embodiment of the above-mentioned method for entity linking, which will not be elaborated here.

[0276] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0277] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the similarities and common parts among the embodiments, reference can be made to each other. For device embodiments, since they are basically similar to method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the descriptions in the method embodiments.

[0278] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs. The technical features recorded in each embodiment can be replaced or combined.

[0279] The modules and sub-modules in the devices and terminals in the embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0280] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are only illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.

[0281] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0282] In addition, the functional modules or sub-modules in each embodiment of the present application can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware or in the form of software functional modules or sub-modules.

[0283] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0284] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0285] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0286] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for entity linking, characterized in that, Including: Calculate the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library respectively; Determine the entity similarity between the entity to be linked and each candidate entity according to the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity; wherein, determine the key entity segmentation words of the entity to be linked and the candidate entity respectively; determine the first similarity between the key entity segmentation words of the entity to be linked and the candidate entity, and the second similarity between the entity to be linked and the candidate entity; take the maximum similarity between the first similarity and the second similarity as the semantic similarity; Determine the candidate entity with the highest entity similarity to the entity to be linked as the linked entity corresponding to the entity to be linked.

2. The method according to claim 1, wherein The entity library is an entity library that has been pre-cleaned according to a preset entity cleaning method; Before calculating the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library respectively, it further includes: cleaning the entity to be linked according to a preset entity cleaning method.

3. The method according to claim 2, wherein The cleaning the entity to be linked according to a preset entity cleaning method includes: Perform word segmentation processing on the entity to be linked to obtain multiple target segmentation words of the entity to be linked; Determine the elimination degree of each target segmentation word; Eliminate the target segmentation words with an elimination degree greater than the first preset elimination degree threshold from the entity to be linked to obtain the cleaned entity to be linked.

4. The method according to claim 3, characterized in that, The determining the elimination degree of each target segmentation word includes: Based on the target entities in the entity library, count the entity order frequencies of each target segmentation word; wherein, the target entity is an entity containing the target segmentation word, and the entity order frequency is the frequency of the order in which the target segmentation word is located in the target entity in the entity library; Determine the order weight of the order in which each target segmentation word is located in the target entity; Take the order weight of the order in which each target segmentation word is located in the target entity as the weight of the corresponding entity order frequency, and calculate the weighted sum of the entity order frequencies corresponding to the target segmentation word; Use the word segmentation order weight of each target segmentation word in the entity to be linked to correct the weighted sum corresponding to the target segmentation word to obtain the elimination degree corresponding to the target segmentation word.

5. The method according to claim 1, characterized in that, The respectively determining the key entity segmentation words of the entity to be linked and the candidate entity includes: Randomly cover the target segmentation words with an elimination degree greater than the second preset elimination degree threshold in the entity to be linked and the candidate entity to obtain the key entity segmentation words; The elimination degree is obtained according to the entity order frequency of each target segmentation word and the order weight of the order in which each target segmentation word is located in the target entity; the target segmentation word is obtained by performing word segmentation processing on the entity to be linked and the candidate entity, the target entity is an entity containing the target segmentation word, and the entity order frequency is the frequency of the order in which the target segmentation word is located in the target entity in the entity library.

6. The method according to claim 1, wherein The determining the first similarity between the key entity segmentation words of the entity to be linked and the candidate entity, and the determining the second similarity between the entity to be linked and the candidate entity include: The entity to be linked, the candidate entity, and the key entity word segmentations of the entity to be linked and the candidate entity are all input into a pre-trained similarity calculation model to obtain the first similarity and the second similarity; wherein, the similarity calculation model is obtained through similar entity differential training and similarity calculation training; the similar entity differential training is used to train the ability of the similarity calculation model to identify similar entities and dissimilar entities.

7. The method according to claim 6, characterized in that The similarity calculation model includes an adapter component arranged after a feedforward network; During the process of performing similar entity differential training and similarity calculation training on the similarity calculation model, the parameters of the adapter component are adjusted.

8. An apparatus for entity linking, characterized in that, It includes: A calculation module, configured to calculate the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity in the entity library respectively; A first determination module, configured to determine the entity similarity between the entity to be linked and each candidate entity according to the character similarity, word similarity, and semantic similarity between the entity to be linked and each candidate entity; wherein, the key entity word segmentations of the entity to be linked and the candidate entity are determined respectively; the first similarity between the key entity word segmentations of the entity to be linked and the candidate entity and the second similarity between the entity to be linked and the candidate entity are determined; the maximum similarity between the first similarity and the second similarity is taken as the semantic similarity; A second determination module, configured to determine the candidate entity with the highest entity similarity to the entity to be linked as the linked entity corresponding to the entity to be linked.

9. An electronic device, characterized in that, It includes: A memory and a processor; Wherein, the memory is used to store a program; The processor is configured to implement the method for entity linking according to any one of claims 1 to 7 by running the program in the memory.

10. A storage medium, characterized in that, It includes: A computer program is stored on the storage medium, and when the computer program is executed by a processor, each step of the method for entity linking according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Device and method for short text similarity calculation

    CN106484678A

  • Named entity linking method and device, equipment and readable storage medium

    CN112836513A