Entity linking method, related device and storage medium

By adopting search-enhanced learning and k-nearest neighbor search mechanisms in medical entity links, using knowledge of similar instances in the training set for reasoning, the linking problem of long-tail entities is solved, and data consistency and information retrieval ability are improved.

CN120277221APending Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410011209.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing medical entity linking technology is difficult to effectively solve the problem of long-tail entities, resulting in data sparsity, difficulty in labeling and insufficient generalization capabilities, making it difficult to accurately link low-frequency medical entities.

Method used

The search-enhanced learning method is adopted, and the knowledge of similar instances in the training set is used to reason through the k-nearest neighbor search mechanism, combining dynamic and difficult negative sample sampling and pre-trained model to improve the link accuracy of the model in long-tail entities.

Benefits of technology

It improves the accuracy of medical entity links and data consistency, enhances information retrieval and reasoning capabilities, and can more accurately link entities from different sources to standardized entities in the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277221A_ABST
    Figure CN120277221A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of knowledge maps, and provides an entity linking method, a related device and a storage medium, the method comprises the following steps: obtaining an initial search text input by a user, the initial search text comprising at least one target word; calling an entity link model to obtain at least one candidate pair matched with the initial search text, wherein the candidate pair comprises a standard word associated with at least one target word on at least one of semantic features and font features; and setting a candidate entity corresponding to the standard word in the candidate pair as a target entity corresponding to the target word in the initial search text according to a corresponding relationship between the standard word and the entity. According to the scheme, reasoning can be carried out by utilizing knowledge of similar instances in a training set during entity prediction, so that the prediction result of the model is corrected, and a correct entity is easier to output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of knowledge graphs, and in particular, to an entity linking method, a related device, and a storage medium. Background Art

[0002] Biomedical Entity Linking (BioEL) aims to map words (such as diseases, drugs) to standard entities in the biomedical ontology, such as the Unified Medical Language System (UMLS). BioEL is crucial for various biomedical tasks, such as automatic diagnosis and medical question answering. Existing biomedical entity linking techniques can generally be divided into three categories:

[0003] (1) The bi-encoder-based method: mainly uses a bi-encoder to encode words and all candidate entities into the same vector space, and then calculates the cosine similarity between the word vector and the entity, and takes the entity with the highest similarity as the linking target of the word.

[0004] (2) The two-stage-based method: uses a bi-encoder to recall some candidate entities, and further uses a cross-encoder to model the fine-grained word-entity interaction to rank the recalled candidate entities.

[0005] (3) The generation-based method: directly generates associated entities using a generative model, thus avoiding the need for negative sample mining.

[0006] In the process of researching and practicing the prior art, the inventors of the embodiments of the present application found that since long-tail entities refer to entities with a low frequency of occurrence in a given domain or dataset, and their frequency of occurrence is low, the number of samples of long-tail entities in the dataset is often small, which may lead to data sparsity, difficult annotation, and generalization ability in model training and application, and ultimately make it difficult to solve the long-tail entity problem in the medical field. Summary of the Invention

[0007] The embodiments of the present application provide an entity linking method, a related device, and a storage medium, which can improve the use of the knowledge of similar instances in the training set for inference when predicting entities, so as to correct the prediction results of the model and make it easier to output the correct entity.

[0008] In a first aspect, the embodiments of the present application provide an entity linking method, which includes:

[0009] Obtain an initial search text input by a user; wherein, at least one target word is included in the initial search text;

[0010] Invoke the entity linking model to obtain at least one candidate pair that matches the initial search text. The candidate pair includes a standard word associated with at least one of the target words in at least one of semantic features and glyph features, and includes the corresponding relationship between the standard word and the entity;

[0011] According to the corresponding relationship between the standard word and the entity, set the candidate entity corresponding to the standard word in the candidate pair as the target entity corresponding to the target word in the initial search text.

[0012] In some embodiments, any candidate negative sample satisfies at least one of the following items:

[0013] The candidate entity that has a similarity higher than the preset similarity with the target word and is a wrong entity of the target word forms a candidate negative sample;

[0014] Alternatively, pair the target word with the candidate entities corresponding to other words in the negative sample set except the target word respectively to obtain the candidate negative sample of the target word.

[0015] In a second aspect, an embodiment of the present application provides an entity linking device, which has the function of implementing the entity linking method provided in the first aspect above. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. In some embodiments, the entity linking device includes:

[0016] An input / output module, configured to obtain an initial search text input by a user; wherein, at least one target word is included in the initial search text;

[0017] A processing module, configured to invoke an entity linking model to obtain at least one candidate pair that matches the initial search text obtained by the input / output module. The candidate pair includes a standard word associated with at least one of the target words in at least one of semantic features and glyph features, and includes the corresponding relationship between the standard word and the entity;

[0018] The processing module is further configured to set the candidate entity corresponding to the standard word in the candidate pair as the target entity corresponding to the target word in the initial search text according to the corresponding relationship between the standard word and the entity.

[0019] In some embodiments, before the input / output module obtains the initial search text input by the user, the processing module is further configured to:

[0020] Sample the initial training set to obtain global negative samples of the words in each training pair; the initial training set includes multiple training pairs in the medical field, each training pair includes a word and its corresponding relationship with the correct entity, and the global negative samples include the first type of negative samples and the second type of negative samples of the words in each training pair;

[0021] Train an entity linking model based on the initial training set and the global negative samples, such that the feature distance between the target word in the initial training set and the negative samples is greater than a first distance, and the feature distance between the target word and the positive samples is less than a second distance, where the first distance is greater than the second distance;

[0022] Among them, the first type of negative samples are composed of candidate entities with a similarity higher than a preset similarity to the target word and being wrong entities of the target word; the target word is the word in any training pair of the initial training set;

[0023] The second type of negative samples are obtained by pairing the target word with the candidate entities corresponding to other words in the initial training set except the target word respectively.

[0024] In some embodiments, the processing module is specifically configured to:

[0025] Obtain the first vector representation of the target word and the vector representations of all entity indices in the initial training set;

[0026] Obtain the similarity between the first vector representation and the vector representations of each entity index;

[0027] Form candidate negative samples by pairing the target word with candidate entities with a similarity higher than a preset similarity to the target word and being wrong entities of the target word, and update the candidate negative samples to the first type of negative samples; until the first type of negative samples of all words are obtained.

[0028] In some embodiments, the processing module is further configured to:

[0029] Obtain a first set, where the first set includes multiple training pairs, and each training pair includes a first word and a first entity of the first word;

[0030] Obtain a second set based on the first set, where the second set includes multiple key-value pairs, and each key-value pair is composed of the vector representation of the first word and the first entity.

[0031] In some embodiments, after obtaining the second set based on the first set, the processing module is further configured to:

[0032] Obtain the target vector representation of a second word;

[0033] Traverse the second set based on the target vector representation to obtain a third set, where the third set includes multiple key-value pairs recalled by the second word, and the probability that the second word links to the entity in each key-value pair in the third set is non-zero;

[0034] According to the classification result between the second word and the entity in each key-value pair in the third set and the first hyperparameter, obtain the predicted distribution of the second word on each entity in the third set;

[0035] Aggregate the probabilities of the same entity in the third set to form a nearest neighbor distribution;

[0036] Linearly interpolate the nearest neighbor distribution and the predicted distribution to obtain a target probability distribution, where the target probability distribution is used to predict the probability that the second word links to the first set.

[0037] In some embodiments, the processing module is specifically configured to:

[0038] Determine a nearest neighbor set according to the similarity between the target vector representation and the key in each key-value pair, where the nearest neighbor set includes a preset number of nearest neighbor entities of the second word;

[0039] According to the similarity between the target vector representation and the key in each key-value pair, obtain the candidate distribution of the first entity on the nearest neighbor set, and aggregate the probabilities with the same label between key-value pairs to form the nearest neighbor distribution.

[0040] In some embodiments, the third set is included in the second set, and the third set includes a first key-value pair and a second key-value pair, where the first key-value pair and the second key-value pair correspond to the same second entity; the processing module is specifically configured to:

[0041] Obtain a first similarity and a second similarity;

[0042] According to the first similarity, the second similarity, and the second hyperparameter, obtain the probability that the second word links to the second entity;

[0043] Wherein, the first similarity is the similarity between the target vector representation and the vector representation in the first key-value pair, and the second similarity is the similarity between the target vector representation and the vector representation in the second key-value pair.

[0044] In some embodiments, the third set further includes a third key-value pair, where the second vector representation and the third entity included in the third key-value pair are different from the first key-value pair and the second mean pair, and the probability that the second word links to the third entity is non-zero; the processing module is further configured to:

[0045] Obtain a third similarity, where the third similarity is the similarity between the target vector representation and the second vector representation;

[0046] According to the third similarity and the second hyperparameter, obtain the probability that the second word links to the third entity.

[0047] In some embodiments, before the input / output module obtains the initial search text input by the user, the processing module is further configured to:

[0048] Perform adversarial perturbation detection on the initial search text;

[0049] If adversarial perturbations that meet preset conditions are detected, at least one of the following items is performed on the initial search text:

[0050] Remove the adversarial perturbations;

[0051] Alternatively, return an interception message.

[0052] In a third aspect, an embodiment of the present application provides an entity linking device, where the face entity linking device includes: at least one processor and a memory; wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the steps in the entity linking method provided in the first aspect, any implementation manner of the first aspect, or any implementation manner provided in the second aspect.

[0053] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which has a function of implementing the entity linking method corresponding to the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. Specifically, the computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the entity linking method provided in the first aspect or any implementation manner of the first aspect in the embodiments of the present application.

[0054] Compared with the prior art, in the solution provided by the embodiment of the present application, after obtaining the initial search text input by the user, an entity linking model is first called to obtain at least one candidate pair that matches the initial search text. Through this matching, it can be seen that because at least one target word in the initial search text is associated with the standard word in each candidate pair in at least one of the semantic features and glyph features (that is, the candidate pair includes a standard word that is associated with at least one of the target words in the initial search text in at least one of the semantic features and glyph features), and the candidate pair includes the corresponding relationship between the standard word and the entity. Therefore, by providing candidate pairs that are associated with the initial search text in at least one of the semantic features and glyph features for reference in the reasoning stage, the candidate entity corresponding to the standard word in the candidate pair can finally be set as the target entity corresponding to the target word in the initial search text. It can be seen that this solution can assist in correcting the retrieval results without additional knowledge, and can link entities from different sources to the standardized entities in the knowledge graph, thereby improving the consistency and comparability of data and effectively improving the information retrieval and reasoning capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a schematic logical framework diagram for implementing the entity linking method in the embodiment of the present application;

[0056] Figure 2a It is a schematic overall flow diagram for implementing the entity linking method in the embodiment of the present application;

[0057] Figure 2b It is a schematic flow diagram for model training in the embodiment of the present application;

[0058] Figure 3 It is a schematic flow diagram for model reasoning in the embodiment of the present application;

[0059] Figure 4 It is a schematic flow diagram for the entity linking method in the embodiment of the present application;

[0060] Figure 5 It is a schematic diagram for comparing the effectiveness of the entity linking model in the embodiment of the present application with 4 types of entity linking models in the related art;

[0061] Figure 6 It is a schematic diagram of the prediction results after two long-tail entities in the embodiment of the present application adopt the entity linking method of the present application;

[0062] Figure 7 It is a schematic structural diagram of the entity linking device in the embodiment of the present application;

[0063] Figure 8 It is a schematic structural diagram of the entity device for implementing the entity linking method in the embodiment of the present application;

[0064] Figure 9 A structural schematic diagram of a mobile phone for implementing the entity linking method in the embodiments of the present application;

[0065] Figure 10 A structural schematic diagram of a server for implementing the entity linking method in the embodiments of the present application. Detailed implementation manners

[0066] It should be noted that the principle of the embodiments of the present application is illustrated by implementing in a suitable computing environment. The following description is based on the specific embodiments of the present application exemplified, and it should not be regarded as limiting other specific embodiments not detailed herein.

[0067] In the following description of the embodiments of the present application, the terms "some embodiments" and "some implementation manners" are involved, which describe subsets of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0068] The terms "first", "second", etc. in the description, claims and above-mentioned drawings of the embodiments of the present application are used to distinguish similar objects (for example, the first delegation queue and the second delegation queue in the embodiments of the present application respectively represent different delegation queues with the same attributes), and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or modules does not have to be limited to those clearly listed steps or modules, but may include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. The division of modules in the embodiments of the present application is only a logical division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed, which are not limited in the embodiments of the present application. And the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed to multiple circuit modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present application.

[0069] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0070] It should be noted that the solution provided in the embodiments of this application involves technologies such as Artificial Intelligence (AI), Natural Language Processing (NLP), Machine Learning (ML), and pre-trained models. Specific illustrations are given through the following embodiments:

[0071] Among them, AI is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0072] AI technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include Machine Learning (ML) technology. Among them, Deep Learning (DL) is a new research direction in machine learning, which is introduced into machine learning to make it closer to the original goal, that is, artificial intelligence. Currently, deep learning is mainly applied in fields such as machine vision and natural language processing.

[0073] Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Using deep learning technology and the corresponding training set, network models that can achieve different functions can be trained. For example, taking natural language processing as an example, a question-and-answer model for question answering can be trained based on a training set, and a viewpoint extraction model for text viewpoint extraction can be trained based on another training set, etc.

[0074] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0075] A pre-training model (PTM), also known as a foundation model or a large model, refers to a deep neural network (DNN) with a large number of parameters. It is trained on a large amount of unlabeled data, and uses the function approximation ability of the large-parameter DNN to enable the PTM to extract common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, it is applicable to downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be classified into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multi-modal models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modalities they process. Among them, multi-modal models refer to models that establish feature representations of two or more data modalities. The pre-training model is an important tool for outputting artificial intelligence-generated content (AIGC), and can also be used as a general interface connecting multiple specific task models.

[0076] For the NLP direction in the field of artificial intelligence, the embodiments of this application can be used to perform semantic understanding on the words in each text, match and search for the target entities of each keyword in the text, so as to realize linking words from different sources to the entities in the knowledge graph.

[0077] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application. The embodiments of the present application provide an entity linking method, related device and storage medium, which can be used in various knowledge bases, retrieval platforms, intelligent question answering systems, etc., and can be implemented on a server or a terminal device. Specifically, it can be used to optimize entity linking, so as to link entities from different sources to the standardized entities in the knowledge graph, thereby improving the consistency and comparability of data, enriching the content and semantics of the knowledge graph, and enhancing information retrieval and reasoning capabilities. In this way, after the entities in the search text are linked to the knowledge graph, the structure and relationships in the knowledge graph can be used for more accurate querying and reasoning. The embodiments of the present application link medical entities to a medical-related knowledge graph, and other fields can also refer to it. The embodiments of the present application do not limit the knowledge field to which the long-tail entities belong.

[0078] In the related technology, due to the low occurrence frequency of long-tail entities and the small number of samples in the dataset, it may lead to data sparsity, difficult annotation and generalization ability in model training and application, and ultimately make it difficult to solve the long-tail entity problem in the medical field. The embodiments of the present application mainly adopt the following technical solutions:

[0079] Embed retrieval-enhanced learning (such as the k-Nearest Neighbor (kNN) algorithm) into the entity linking model to utilize the knowledge of similar instances in the training set for reasoning when predicting the entity corresponding to the target word.

[0080] Taking the medical field as an example, design a model architecture: kNN-BioEL. Specifically, incorporate retrieval-enhanced learning into the training and reasoning processes of the BioEL model. First, construct a database that stores the word vectors generated by the BioEL model and the corresponding entity labels of these word vectors. During the training process, use dynamic hard negative sampling to train the model; during the reasoning process, the model retrieves the top-k nearest neighbor instances from the database as clues based on the word vector similarity, and then performs reasoning by linearly interpolating the prediction distribution of the model with its nearest neighbor distribution, that is, utilize the knowledge of similar instances in the training set for reasoning when predicting the entity corresponding to the target word. By using dynamic hard negative sampling to train the model, it is possible to accurately mine the fine-grained semantic differences between different medical entities (for example, type 1 diabetes and type 2 diabetes), and improve the quality of the retrieved neighbor entities. In some embodiments, other types of methods (such as two-stage methods, generative methods) can also be used as the BioEL model, and the present application does not limit this.

[0081] In some embodiments, the entity linking method of the embodiments of the present application can be based on the logic framework as Figure 1 shown, incorporating retrieval reinforcement learning into the training and inference processes of the BioEL model. The logic framework includes an encoder for the model training phase, a database for the model inference phase, and an encoder. Given a word m to be recognized and an ontology ξ consisting of N candidate entities, the goal of BioEL is to map the word m to the correct entity e ∈ ξ of the word m. Among them, the entity linking device of the entity linking method in the embodiments of the present application can be a server or a terminal. In the embodiments of the present application, a server is taken as an example. The server: can be a retrieval platform related to biology and medicine. In the embodiments of the present application, taking the medical entity linking to the medical-related knowledge graph as an example. It can be understood that the entity linking method in the embodiments of the present application can also be used in combination with the entity linking methods in related technologies (such as the method based on bi-encoder, the method based on bi-encoder + cross-encoder, the method based on generative model, etc.), and can be specifically selected according to actual business needs. The embodiments of the present application do not limit this.

[0082] The following combines Figures 2a - 10 to exemplarily illustrate the technical solutions of the embodiments of the present application.

[0083] The overall flowchart is as Figure 2a shown. Since the entity linking method of the embodiments of the present application performs entity linking on the search term based on a pre-trained entity linking model, before introducing the entity linking method, the training phase of the entity linking model is introduced first (as Figure 2b shown), and the model inference phase (as Figure 3 shown) and the model application phase (as Figure 4 shown) will be introduced in sequence later.

[0084] I. Model Training Phase

[0085] As Figure 2a shown, Figure 2b is a schematic diagram of a training process of an entity linking model. The training process includes:

[0086] 101. Obtain an initial training set.

[0087] Among them, the initial training set may include multiple training pairs in the medical field. Each training pair includes a word and the corresponding relationship with the correct entity. For example, the training pair can be expressed as: word - entity pair (m, e).

[0088] For example, taking the initial data set B as an example, the initial data set B includes multiple training pairs (m, e), where m is a word and e is an entity, and the entity e is the entity label corresponding to the word m. For example, when m is "back molar", e is "structure of molar tooth".

[0089] 102. Sample the initial training set to obtain global negative samples of the words in each training pair.

[0090] Among them, the global negative samples include the first type of negative samples and the second type of negative samples of the words in each training pair.

[0091] Regarding the first type of negative samples, they are composed of candidate entities that are target words with a similarity higher than a preset similarity and are incorrect entities of the target words; the target word is the word in any training pair of the initial training set. For example, according to the cosine similarity between the vector representation of the word m and the vector representation of the entity, top-p (p is a hyperparameter) candidate entities with high similarity but incorrect are retrieved from all candidate entities, and these entities e″ are paired with the word m to construct the first type of negative samples of the word m.

[0092] Regarding the second type of negative samples, they are obtained by pairing the target word with candidate entities corresponding to other words in the initial training set except the target word. For example, it can be obtained by sampling the initial data set B, that is, pairing the word m with the entities in other training pairs in the initial data set B to construct the second type of negative samples of the word m.

[0093] It should be noted that the first type of negative samples are negative examples that can be obtained based on the dynamic hard negative sampling (DHNS) in the embodiments of the present application, and the second type of negative samples are negative examples directly sampled from the initial data set B during the training process. Considering that negative sample sampling within the initial data set B is restricted by the size of the data volume in the initial data set B, it increases the difficulty of obtaining fine-grained semantic differences between entities. Since biological and medical entities usually have several different expression forms. On the one hand, these expression forms can show significant differences. For example, for "motrin" and "ibuprofen", although they are quite different literally but refer to the same chemical entity. On the other hand, entities with high literal similarity can also have different meanings, such as the disease entities "Type 1 Diabetes" and "Type 2 Diabetes", or the chemical entities "xyloglucan en dotransglycosylase" and "xyloglucan endoglucanase". Therefore, in order to more accurately obtain the fine-grained semantic relationships in medical entities, the embodiments of the present application combine the first type of negative samples and the second type of negative samples, and on the basis of the second type of negative samples, add the first type of negative samples for auxiliary inference to make the negative samples of specific attributives more comprehensive. Especially for entities in the biological and medical fields, the probability of long-tail entities appearing is higher. A long-tail entity refers to an entity with a relatively low occurrence frequency in a given domain or data set.

[0094] In some embodiments, DHNS retrieves from all entities in the initial training set to construct global negative samples. DHNS pre-computes the vector representations of all entities and constructs an entity index. The first type of negative samples can be updated iteratively. In each iteration, DHNS calculates the similarity between the vector representation of a single word in the initial training set and the vector representations of all entity indexes respectively, and uses the retrieved entities with high similarity but incorrect as the first type of negative samples. Since these negative samples will be dynamically updated in each iteration process, it makes the first type of negative samples more challenging in each iteration. The first type of negative samples in the global negative samples can be obtained through the following steps (1)-(3):

[0095] (1) Obtain the first vector representation of the target word and the vector representations of all entity indexes in the initial training set.

[0096] In some embodiments, a bi-encoder can be used as the BioEL model, and SapBERT[3] can be used as the encoder to generate corresponding vector representations for words and each candidate entity. The vector representation f(m) of word m can be expressed by the following formula (1):

[0097] f(m) = SapBERT(m)[CLS], (1)

[0098] where [CLS] represents the hidden state of the special token in the last layer of SapBERT, aiming to obtain a fixed-size feature vector for each input. The calculation method of the vector representation f(e) of entity e is similar to that of word m.

[0099] (2) Obtain the similarity between the first vector representation and the vector representations of each entity index.

[0100] In some embodiments, the similarity score of the word-entity pair (m, e) is expressed as the following formula (2):

[0101] s(m, e) = g(f(m), f(e)), (2)

[0102] where g represents the cosine similarity.

[0103] (3) Construct candidate negative samples from the candidate entities whose similarity to the target word is higher than the preset similarity and are wrong entities of the target word, and update the candidate negative samples to the first type of negative samples; until the first type of negative samples for all words are obtained; the global negative samples are used to obtain the semantic relationships in the entities of the preset domain.

[0104] Specifically, in the model training stage, the embodiments of the present application combine the first type of negative samples and the second type of negative samples to accurately infer the fine-grained semantic relationships in entities (such as medical entities).

[0105] 103. Train the entity linking model based on the initial training set and the global negative samples, so that the feature distance between the target word in the initial training set and the negative samples is greater than the first distance, and the feature distance from the positive samples is less than the second distance.

[0106] Wherein, the first distance is greater than the second distance. The smaller the feature distance between the target word and the positive sample, the more similar they are; the greater the feature distance between the target word and the negative sample, the less similar they are.

[0107] In some embodiments, the objective function during model training can adopt the following formula (3):

[0108]

[0109] Wherein, s(m,e) is calculated by formula (2), where τ is a temperature hyperparameter. This objective function can encourage the model to bring positive samples closer and push negative samples farther away, thereby improving the quality of text representation.

[0110] In the embodiments of the present application, on the one hand, in order to accurately capture the fine-grained semantic differences between different medical entities (e.g., type 1 diabetes and type 2 diabetes) and improve the quality of retrieved neighbors, during training, dynamic hard negative sampling is adopted to train the model, that is, through dynamic iteration, continuously obtain the first type of negative samples that can effectively assist the model in reasoning, and then combine with batch generation of incorrect word-entity pairs (i.e., training pairs) between words in the initial training set and entities of other words. Even in the face of long-tail entities with very low frequencies of occurrence in the domain or dataset, the model of the embodiments of the present application is trained based on global negative samples, and during the training process, the feature distances of positive samples can be made closer, and the feature distances of negative samples can be made farther. Therefore, when there are words from different sources and different expressions to be linked to the knowledge graph, the knowledge of similar training pairs in the training set can be referred to to infer entities with higher accuracy, rather than being unable to link to entities or linking to incorrect entities. Thus, it can be seen that this solution can improve the consistency and comparability of data; on the other hand, this solution is general and can be directly applied to most existing BioEL models without the need to redesign a new solution to replace it, without high costs. The search platforms using similar models can be quickly implemented with only some optimizations. Thus, through the solution of the embodiments of the present application, during the biological and medical entity linking, the consistency and comparability of data can be improved, the content and semantics of the knowledge graph related to biological and medical entities can be enriched, and the information retrieval and reasoning capabilities can be enhanced, which further helps to improve the accuracy and efficiency of medical decision-making. On the other hand, the embodiments of the present application can link these words to accurate entities without using any auxiliary information (e.g., word context or entity description) and by referring to approximate instances.

[0111] II. Model Inference Stage

[0112] In some embodiments of the present application, in order to improve the ability of the entity linking model for long-tail entities, a k-nearest neighbor retrieval mechanism is also proposed. This k-nearest neighbor retrieval mechanism provides valuable clues during the reasoning process using existing instances. The k-nearest neighbor retrieval mechanism includes two nodes: constructing a database of training instances and predicting with a kNN distribution interpolation model.

[0113] Node 1: Constructing a Database of Training Instances

[0114] The database stores word vectors generated by the BioEL model, and the entities corresponding to these word vectors. Specifically, first, a first set is obtained, the first set includes multiple training pairs, each of which includes a first word and a first entity of the first word. Then, based on the first set, a second set is obtained, the second set includes multiple key-value pairs, each key-value pair consists of a vector representation of the first word and the first entity. It can be understood that the first word is only a general reference and can be a word in any training pair in the first set. Correspondingly, the first entity is a candidate entity for its pairing.

[0115] For example, given the i-th instance (m i ,e i ), define the key-value pair (f(m i ),e i ), where f(m i ) is calculated by formula (1) to get word m i The vector representation of e i It is the word m i Entity tags. Therefore, the database is the set of all key-value pairs constructed from all training instances in the training set D, namely:

[0116]

[0117] By constructing this database, during the reasoning process, the model can retrieve the top-k nearest neighbor instances from the database as clues based on the similarity between the word vector and the candidate entity, and then perform reasoning by linearly interpolating the model's predicted distribution with its nearest neighbor distribution.

[0118] Node 2, Nearest Neighbor Distribution Interpolation

[0119] like Figure 3 As shown, after obtaining the second set based on the first set, the method further includes:

[0120] 201. Obtain a target vector representation of the second word.

[0121] For example, given a second word x, the model outputs the vector representation f(x) of word x through formula (1).

[0122] 202. Traverse the second set based on the target vector representation to obtain a third set.

[0123] The third set includes multiple key-value pairs recalled by the second word, and the probability that the second word is linked to an entity in each key-value pair in the third set is non-zero.

[0124] For example, the entity linking model uses f(x) to query the database constructed in node 1 According to f(x) and the database in retrieve the k (k is a hyperparameter) nearest neighbors of the word x according to the cosine similarity g(·,·) in

[0125] 203. Obtain the predicted distribution of the second word on each entity in the third set according to the classification result between the second word and the entities in each key-value pair in the third set and the first hyperparameter.

[0126] For example, given the word x, the model outputs the vector representation f(x) of the word x through formula (1) and generates the predicted distribution p BioEL (y|x). In some embodiments, the predicted distribution p BioEL (y|x) can refer to the following formula (4):

[0127]

[0128] where β1 is the temperature hyperparameter, which can make the distribution flatter or sharper.

[0129] 204. Aggregate the probabilities of the same entities in the third set to form the nearest neighbor distribution.

[0130] In some embodiments, aggregating the probabilities of the same entities in the third set to form the nearest neighbor distribution includes:

[0131] a. Determine the nearest neighbor set according to the similarity between the target vector representation and the keys in each key-value pair.

[0132] where the nearest neighbor set includes a preset number of nearest neighbor entities of the second word.

[0133] b. Obtain the candidate distribution of the first entity on the nearest neighbor set according to the similarity between the target vector representation and the keys in each key-value pair.

[0134] c. Aggregate the probabilities of the same entities between key-value pairs to form the nearest neighbor distribution.

[0135] In some embodiments, in the process of calculating the participation probability, the probability calculation of linking to a single entity is involved. Since there are at least two states of the entities participating in the probability aggregation process: from key-value pairs with different vector representations but the same entity; from key-value pairs with different vector representations and entities. Therefore, in some embodiments, the probabilities of the same entities between key-value pairs can be aggregated in the following two ways:

[0136] Method 1: Calculate the probability of a single entity for key-value pairs with different vector representations but the same entity:

[0137] In this embodiment, the third set is included in the second set. The third set includes a first key-value pair and a second key-value pair, and the first key-value pair and the second key-value pair correspond to the same second entity. Among them, the probability that the aggregated key-value pairs have the same entity includes:

[0138] c1: Obtain a first similarity and a second similarity.

[0139] Among them, the first similarity is the similarity between the target vector representation (such as f(x)) and the vector representation in the first key-value pair (such as (f(m1), e1)). The second similarity is the similarity between the target vector representation and the vector representation in the second key-value pair (such as (f(m2), e1)).

[0140] c2: According to the first similarity, the second similarity, and the second hyperparameter, obtain the probability that the second word links to the second entity.

[0141] Method 2: Calculate the probability of a single entity for key-value pairs with different vector representations and different entities:

[0142] In this embodiment, the third set further includes a third key-value pair. The second vector representation and the third entity included in the third key-value pair are both different from the first key-value pair and the second mean pair, and the probability that the second word links to the third entity is non-zero. The specific calculation process includes:

[0143] c1': Obtain a third similarity, where the third similarity is the similarity between the target vector representation and the second vector representation;

[0144] c2': According to the third similarity and the second hyperparameter, obtain the probability that the second word links to the third entity.

[0145] For example, the entity linking model uses f(x) to query the database constructed in node 1 According to f(x) and the database in retrieve the k (k is a hyperparameter) nearest neighbors of the word x through the cosine similarity g(·,·) After that, by applying the softmax function to the cosine similarity, calculate the distribution on the nearest neighbors and aggregate (the max operation can be used) the probability values with the same label in the instances to finally form the kNN distribution. In some embodiments, the kNN distribution can refer to the following formula (5):

[0146]

[0147] Among them, β2 is the temperature hyperparameter. For entities that do not appear in the nearest neighbors , the probability is defaulted to 0. For example, if the word x retrieves 3 retrieval instances from the database B, i.e., {(f(m1), e1), (f(m2), e1), (f(m3), e2)}, where the entity labels corresponding to the first instance and the second instance are the same, both being e1.

[0148] Then, the probability value for the label e1 is:

[0149]

[0150] The probability value for the label e2 is:

[0151]

[0152] For entities that do not appear in the retrieval instances, that is, the p kNN probability values of all entities except e1 and e2 are 0. e1 and e2 can be approximate entities. In the medical field, taking e1 as type 1 diabetes and e2 as type 2 diabetes as an example. During the inference process, the reference instances in the database of this solution are adopted, that is, the entity linking model retrieves 3 nearest neighbor instances (i.e., (f(m1), e1), (f(m2), e1), (f(m3), e2)) from the database according to the similarity g(·,·) between the target vector representation f(x) and the database in as clues, and then the inference is carried out by linearly interpolating the predicted distribution of the model and its nearest neighbor distribution, so as to capture the fine-grained semantic differences between type 1 diabetes and type 2 diabetes, thereby improving the quality of the retrieved neighbor entities and making it easier for the entity linking model to output the correct entity.

[0153] 205. Linearly interpolate the nearest neighbor distribution and the predicted distribution to obtain the target probability distribution.

[0154] Among them, the target probability distribution is used to predict the probability that the second word links to the first set.

[0155] For example, the hyperparameter λ can be used to linearly interpolate the kNN distribution p kNN (y|x) and the predicted distribution p BioEL (y|x) to obtain the final target distribution p(y|x). In some embodiments, the target distribution p(y|x) can refer to the following formula (6):

[0156] p(y|x) = λp kNN (y|x) + (1 - λ)pBioEL (y|x). (6)

[0157] It can be seen that in the embodiments of the present application, by constructing the above database it is convenient to traverse the above database without passing through the vector representation of the second word during the model training phase so as to obtain the nearest neighbor entities including multiple second words, which are used to subsequently aggregate the probabilities of having the same entity to form a nearest neighbor distribution, and finally linearly interpolate the nearest neighbor distribution and the prediction distribution to obtain the target probability distribution. When the obtained target probability distribution is used to predict any word (such as the second word), the probability of linking the word to the first set can be effectively improved, thereby providing a better and more accurate search experience for the end user and reducing the situation of incorrect linking.

[0158] After the above entity linking model is trained, any word in the search text can be linked to the entity in the knowledge graph based on the entity linking model. Specifically, the embodiments of the present application further include the following Figure 4 embodiments shown, including the following steps:

[0159] 301. Obtain the initial search text input by the user.

[0160] Wherein, the initial search text includes at least one target word. The target words can be mapped to the same entity or different entities, and the embodiments of the present application do not limit this.

[0161] For example, the initial search text includes target words such as "lung cancer", "pulmonary malignant tumor", "type 1 diabetes", "type 2 diabetes", etc. Among them, "lung cancer" and "pulmonary malignant tumor" should be matched to the same entity, while "type 1 diabetes" and "type 2 diabetes" have different meanings and should be matched to different entities.

[0162] It can be understood that the initial search text input by the user can be the text directly input by the user, or can be extracted or converted from any form such as audio, picture or video input by the user. The embodiments of the present application do not limit this. The embodiments of the present application only take directly obtaining the user's input as the text as an example.

[0163] In some embodiments, since this solution is applied to a retrieval system, it may involve text processing of large language models. The initial search text input into the retrieval system covers a wide range of topics and may pose AI security risks, including but not limited to: using prompt injection to make the entity linking model finally output a specified entity or a randomly answered entity, adding characters to the text through an adversarial attack algorithm to form the initial search text and bypass the risk detection system, causing the entity linking model to generate opposite or completely irrelevant entity linking results, etc. Therefore, before obtaining the initial search text input by the user, the initial search text can be detected, specifically including:

[0164] Performing adversarial perturbation detection on the initial search text;

[0165] If an adversarial perturbation that meets the preset conditions is detected, at least one of the following items is performed on the initial search text:

[0166] Removing the adversarial perturbation;

[0167] Or, returning an interception message.

[0168] It can be seen that by performing adversarial perturbation detection in advance and performing corresponding operations, such as removing the adversarial perturbation in the initial search text and then inputting it into the entity linking model, or directly intercepting the search request of the initial search text, this mechanism can effectively avoid such adversarial attacks, improve the security of the entity linking model, thereby ensuring that the output of the entity linking model is more accurate and also improving the user experience.

[0169] It can be understood that before the entity linking model related to the embodiments of this application is launched, the entity linking model can also be subjected to adversarial training. By discovering the algorithm risks of the entity linking model in advance and introducing a value evaluation model, a coach model, etc. to strengthen and enhance the training of the entity linking model, the entity linking model is made less vulnerable to attacks. During the adversarial training process, a generative model can be combined to generate an adversarial sample set, and these adversarial samples can be integrated into the training set for adversarial training. The specific implementation means are not elaborated in this application. Optionally, an adversarial model can be used to generate multiple rounds of adversarial questions and input them into the entity linking model to be tested (when implemented based on a large language model) for multiple rounds of question-answer iteration training. The output is scored by a trained prophet evaluation model, and the scoring result is fed back to the model to be tested, enabling it to continuously learn the key points and differences between good and bad outputs through reinforcement learning until the question-answer ability is gradually iterated to the optimal. In this way, the efficiency of testing the entity linking model is significantly improved, which is suitable for industrial-level mass requirements.

[0170] In some other embodiments, error correction processing may also be performed on the initial search text, including but not limited to correcting misspelled words in the initial search text and semantic errors caused by interfering characters after word segmentation. It may also prompt the user with "any errors in the initial search text" and the like. The embodiments of the present application do not limit this.

[0171] 302. Invoke the entity linking model to obtain at least one candidate pair that matches the initial search text.

[0172] Among them, the candidate pair includes the corresponding relationship between a standard word and an entity. For example, the candidate pair obtained by matching is (m i , e i ), the word m i is the standard word obtained by matching, and e i is the entity label corresponding to the standard word m i . The corresponding relationship between the standard word and the entity in each candidate pair is preset and comes from the above-mentioned entity linking model training stage, which will not be elaborated here.

[0173] In the initial search text, at least one of the target words is associated with the standard word in at least one of the semantic features and glyph features in at least one of the candidate pairs, that is, the candidate pairs obtained by matching include standard words that are associated with at least one of the target words in at least one of the semantic features and glyph features. It can be understood that the same target word can match at least one candidate pair, and different target words can match the same or different candidate pairs.

[0174] For example, given a target word "back molar", a candidate pair obtained by matching is (molar, structure of molar tooth), then the standard word that matches the target word "back molar" is molar.

[0175] 303. According to the corresponding relationship between the standard word and the entity, set the candidate entity corresponding to the standard word in the candidate pair as the target entity corresponding to the target word in the initial search text. Finally, multiple target entities can be obtained, and the multiple target entities can be sorted.

[0176] Correspondingly, for a given target word "back molar", since the standard word matching the target word "back molar" is "molar", then, according to the correspondence between "molar" and "structure of molar tooth" in the candidate pairs, the entity "structure of molar tooth" can be used as one of the candidate entities for the target word "back molar", and it is sorted with the candidate entities in other matched candidate pairs. The specific sorting strategy for each candidate entity is not limited in this application.

[0177] Compared with the prior art, in the solution provided by the embodiments of this application, after obtaining the initial search text input by the user, the foregoing entity linking model is first called to obtain at least one candidate pair that matches the initial search text. From this match, since at least one target word in the initial search text is associated with the standard word in each candidate pair in at least one of the semantic features and glyph features (that is, the candidate pair includes a standard word that is associated with at least one of the target words in the initial search text in at least one of the semantic features and glyph features), and the candidate pair includes the correspondence between the standard word and the entity, therefore, by providing candidate pairs that are associated with the initial search text in at least one of the semantic features and glyph features for reference in the reasoning stage, the candidate entity corresponding to the standard word in the candidate pair can finally be set as the target entity corresponding to the target word in the initial search text. It can be seen that this solution can assist in correcting the retrieval results without additional knowledge, and can link entities from different sources to the standardized entities in the knowledge graph, thereby improving the consistency and comparability of data and effectively improving the information retrieval and reasoning capabilities.

[0178] For ease of understanding, experiments are conducted on the publicly available medical entity linking dataset below to verify the effectiveness of the entity linking method. Specifically, experiments are conducted on 4 publicly available medical entity linking datasets. The experimental results are as Figure 5 shown. In this experiment, the evaluation metrics used are Acc@1 and Acc@5.

[0179] Acc@1 means that in the prediction of a given sample by the model, if the correct answer is in the first prediction of the model, it is counted as 1, otherwise it is counted as 0. The final Acc@1 score is the average of the accuracies of all samples.

[0180] Acc@5 means that in the prediction of a given sample by the model, if the correct answer is in the first five predictions of the model, it is counted as 1, otherwise it is counted as 0. The final Acc@5 score is the average of the accuracies of all samples.

[0181] Regarding the effect of model reasoning, fromFigure 5 As can be seen, the entity linking method (e.g., kNN-BioEL) proposed in the embodiments of this application exceeds the previous best baseline models in terms of Acc@1 on NCBI, COMETA, and AAP, with relative improvements of 0.2%, 2.4%, and 1.0% respectively, and achieves comparable results to Prompt-BioEL on BC5CDR.

[0182] In addition, in the related art, the two models GenBioEL and Prompt-BioEL rely to a large extent on additional synonym knowledge to pre-train the model, but these synonyms are difficult to obtain in real scenarios, thus limiting their application scope. However, kNN-BioEL adopted in the embodiments of this application uses existing training instances to enhance the model and does not require additional synonym knowledge. It can be seen that, compared with the mainstream method of fine-tuning using pre-trained language models, kNN-BioEL adopted in the embodiments of this application proves that the proposed retrieval-enhanced learning is more effective.

[0183] Figure 6 For given target words (ulcerated, novacaine), a comparison of the test results of two cases with and without kNN retrieval based on the database COMETA is made. Taking two long-tail entities "ulcer lesion" and "procaine" as examples, it shows that kNN retrieval can help correct the prediction results of the target words (ulcerated, novacaine). Without kNN retrieval, the BioEL model tends to predict the word as a literally similar but incorrect entity. For example, it predicts "ulcerated" as "ulcerated mass". However, after using kNN retrieval, that is, adopting kNN-BioEL, the kNN-BioEL is more likely to output a more correct entity. For example, the model can refer to the retrieved instances (ulcers, ulcer lesion) and predict the result as the correct entity "ulcer lesion".

[0184] Figures 1 to 6 Any technical feature mentioned in the embodiments corresponding to any one of them also applies to the embodiments in this application Figures 7 to 10 corresponding to the embodiments, and the subsequent similar parts will not be elaborated.

[0185] The above describes an entity linking method in the embodiments of this application. Next, an entity linking device for executing the above entity linking method will be introduced.

[0186] Refer to Figure 7 as Figure 7Schematic diagram of the structure of an entity linking device 40 shown, which can be applied to link any word in a search text to a target entity in a knowledge graph. For example, terms in a medical text are associated with standardized entities in a medical knowledge base, so as to provide more accurate and consistent entities. The entity linking device 40 in the embodiments of the present application can implement the steps in the entity linking method executed by the entity linking device 40 corresponding to any of the above Figures 1 - 6 The steps in the entity linking method executed by the entity linking device 40 corresponding to any one of the embodiments. The functions implemented by the entity linking device 40 can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. The entity linking device 40 may include an input / output module 401 and a processing module 402. The implementation of the functions of the input / output module 401 and the processing module 402 can refer to Figures 1 - 6 The operations executed in any corresponding embodiment, which will not be elaborated here.

[0187] In some embodiments, the input / output module 401 can be used to obtain an initial search text input by a user, where the initial search text includes at least one target word;

[0188] The processing module 402 can be used to call an entity linking model to obtain at least one candidate pair that matches the initial search text obtained by the input / output module. The candidate pair includes a standard word associated with at least one of the target words in at least one of semantic features and glyph features, and includes the corresponding relationship between the standard word and the entity;

[0189] The processing module 402 is further used to set the candidate entity corresponding to the standard word in the candidate pair as the target entity corresponding to the target word in the initial search text according to the corresponding relationship between the standard word and the entity.

[0190] Other embodiments can refer to the introduction of the method-side embodiments, which will not be elaborated here.

[0191] The entity linking device 40 for executing the entity linking method in the embodiments of the present application has been described above from the perspective of modular functional entities. Next, the entity linking device 40 for executing the entity linking method in the embodiments of the present application will be described from the perspective of hardware processing. It should be noted that in the embodiments of the present application Figure 7 The entity device corresponding to the input / output module 401 in the shown embodiment can be an input / output unit, a transceiver, a radio frequency circuit, a communication module, an output interface, etc., and the entity device corresponding to the processing module 402 can be a processor. Figure 7 The shown entity linking device 40 can have a structure as Figure 8 shown, when Figure 7When the entity link device 40 shown has a structure as shown in Figure 8 , the processor and transceiver in Figure 8 can implement the same or similar functions as the input / output module 401 and the processing module 402 provided by the device embodiment corresponding to the entity link device 40 described above. Figure 8 The memory in stores a computer program that the processor needs to call when executing the above entity link method.

[0192] An embodiment of the present application also provides another entity device for implementing the entity link method. As shown in Figure 9 , for the sake of convenience of description, only the parts related to the embodiment of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The entity link device can be any entity link device including a mobile phone, a tablet computer, a personal digital assistant (English full name: Personal Digital Assistant, English abbreviation: PDA), a point of sales (English full name: Point of Sales, English abbreviation: POS), an in-vehicle computer, etc. Taking the entity link device as a mobile phone as an example:

[0193] Figure 9 Shown is a block diagram of a part of the structure of a mobile phone related to the entity link device provided by the embodiment of the present application. Referring to Figure 9 , the mobile phone includes: a radio frequency (English full name: Radio Frequency, English abbreviation: RF) circuit 710, a memory 720, an input unit 730, a display unit 740, a sensor 750, an audio circuit 760, a wireless-fidelity (English full name: wireless-fidelity, English abbreviation: Wi-Fi) module 770, a processor 780, and a power supply 790, etc. Those skilled in the art can understand that Figure 9 the structure of the mobile phone shown in does not limit the mobile phone, and it may include more or fewer components than shown in the figure, or combine some components, or arrange different components.

[0194] Next, the various components of the mobile phone will be specifically introduced in conjunction with Figure 9 :

[0195] The RF circuit 710 can be used for receiving and transmitting information or signals during communication. Specifically, it receives the downlink information from the base station and processes it with the processor 780. Additionally, it sends the uplink data to the base station. Generally, the RF circuit 710 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (Full English name: Low Noise Amplifier, English abbreviation: LNA), a duplexer, etc. Moreover, the RF circuit 710 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (Full English name: Global System of Mobile communication, English abbreviation: GSM), General Packet Radio Service (Full English name: General Packet Radio Service, English abbreviation: GPRS), Code Division Multiple Access (Full English name: Code Division Multiple Access, English abbreviation: CDMA), Wideband Code Division Multiple Access (Full English name: Wideband Code Division Multiple Access, English abbreviation: WCDMA), Long Term Evolution (Full English name: Long Term Evolution, English abbreviation: LTE), email, Short Messaging Service (Full English name: Short Messaging Service, English abbreviation: SMS), etc.

[0196] The memory 720 can be used to store software programs and modules. The processor 780 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 720. The memory 720 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 720 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0197] The input unit 730 can be used to receive input numerical or character information and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 730 can include a touch panel 731 and other input devices 732. The touch panel 731, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel 731), and drive corresponding connection devices according to a pre-set program. Optionally, the touch panel 731 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, then sends it to the processor 780, and can receive and execute the commands sent by the processor 780. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 731. In addition to the touch panel 731, the input unit 730 can also include other input devices 732. Specifically, the other input devices 732 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0198] The display unit 740 can be used to display the information input by the user or the information provided to the user and various menus of the mobile phone. The display unit 740 can include a display panel 741. Optionally, the display panel 741 can be configured in forms such as a liquid crystal display (English full name: Liquid Crystal Display, English abbreviation: LCD), an organic light-emitting diode (English full name: Organic Light-Emitting Diode, English abbreviation: OLED), etc. Further, the touch panel 731 can cover the display panel 741. When the touch panel 731 detects a touch operation thereon or nearby, it transmits it to the processor 780 to determine the type of touch event. Subsequently, the processor 780 provides corresponding visual output on the display panel 741 according to the type of touch event. Although in Figure 9 the touch panel 731 and the display panel 741 are implemented as two independent components to realize the input and output functions of the mobile phone, in some embodiments, the touch panel 731 and the display panel 741 can be integrated to realize the input and output functions of the mobile phone.

[0199] The mobile phone may further include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 741 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 741 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer attitude calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the mobile phone can also be configured with, they will not be elaborated here.

[0200] The audio circuit 760, the speaker 761, and the microphone 762 can provide an audio interface between the user and the mobile phone. The audio circuit 760 can transmit the electrical signal converted from the received audio data to the speaker 761, and the speaker 761 converts it into a sound signal for output; on the other hand, the microphone 762 converts the collected sound signal into an electrical signal, which is received by the audio circuit 760 and then converted into audio data. After the audio data is output to the processor 780 for processing, it is sent through the RF circuit 710 to, for example, another mobile phone, or the audio data is output to the memory 720 for further processing.

[0201] Wi-Fi belongs to short-distance wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the Wi-Fi module 770, which provides users with wireless broadband Internet access. Although Figure 9 the Wi-Fi module 770 is shown, it can be understood that it does not belong to an essential component of the mobile phone and can be omitted entirely within the scope of not changing the essence of the application according to needs.

[0202] The processor 780 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 720, and by calling the data stored in the memory 720, it executes various functions of the mobile phone and processes data, thereby performing an overall detection of the mobile phone. Optionally, the processor 780 may include one or more processing units; the processor 780 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 780 either.

[0203] The mobile phone further includes a power supply 790 (such as a battery) for powering each component. The power supply can be logically connected to the processor 780 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.

[0204] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0205] In the embodiment of the present application, the processor 780 included in the mobile phone further has the function of controlling the execution of the method flow performed by the entity linking device 40 shown above. The steps performed by the entity linking device in the above embodiments may be based on the mobile phone structure shown above. For example, the processor 780 performs the following operations by calling instructions in the memory 720: Figure 9 Obtain an initial search text input by a user through the input unit 730, where the initial search text includes at least one target word; Figure 9 Call an entity linking model to obtain at least one candidate pair that matches the initial search text, where the candidate pair includes a standard word associated with at least one of the target words in at least one of semantic features and glyph features;

[0206] According to the correspondence between the standard word and the entity, set the candidate entity corresponding to the standard word in the candidate pair as the target entity corresponding to the target word in the initial search text.

[0207] Call an entity linking model to obtain at least one candidate pair that matches the initial search text, where the candidate pair includes a standard word associated with at least one of the target words in at least one of semantic features and glyph features;

[0208] According to the correspondence between the standard word and the entity, set the candidate entity corresponding to the standard word in the candidate pair as the target entity corresponding to the target word in the initial search text.

[0209] Details of others are not elaborated.

[0210] The embodiment of the present application further provides another entity linking device for implementing the above entity linking method, or a search device for implementing the above entity linking method, as shown in Figure 10 shown, Figure 10It is a schematic diagram of a server structure provided by an embodiment of the present application. The server 1020 may vary significantly due to configuration or performance differences, and may include one or more central processing units (full English name: central processing units, English abbreviation: CPU) 1022 (for example, one or more processors) and a memory 1032, and one or more storage media 1030 (for example, one or more mass storage devices) for storing application programs 1042 or data 1044. Among them, the memory 1032 and the storage medium 1030 can be transient storage or persistent storage. The program stored in the storage medium 1030 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1022 may be configured to communicate with the storage medium 1030 and execute a series of instruction operations in the storage medium 1030 on the server 1020.

[0211] The server 1020 may further include one or more power supplies 1026, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1058, and / or one or more operating systems 1041, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.

[0212] In the above embodiment, the Figure 7 steps performed by the entity linking device 40 shown may be based on the Figure 10 structure of the server 100 shown. For example, in the above embodiment, the Figure 7 steps performed by the entity linking device 40 shown may be based on the Figure 10 server structure shown. For example, the central processing unit 1022 performs the following operations by calling instructions in the memory 1032:

[0213] Obtain the initial search text input by the user through the input / output interface 1058, and the initial search text includes at least one target word;

[0214] Obtain at least one candidate pair that matches the initial search text, and the candidate pair includes a standard word associated with at least one of the target words in at least one of semantic features and glyph features;

[0215] According to the corresponding relationship between the standard word and the entity, set the candidate entity corresponding to the standard word in the candidate pair as the target entity corresponding to the target word in the initial search text.

[0216] Other details are not elaborated.

[0217] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0218] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0219] In several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.

[0220] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0221] In addition, in each embodiment of the embodiments of the present application, the various functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0222] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0223] The computer program product includes one or more computer instructions. When the computer program is loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0224] The technical solutions provided in the embodiments of the present application have been introduced in detail above. Specific examples are used in the embodiments of the present application to illustrate the principles and implementation manners of the embodiments of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the embodiments of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the embodiments of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the embodiments of the present application.

Claims

1. An entity linking method, characterized in that, Including: Obtain an initial search text input by a user; wherein, at least one target word is included in the initial search text; Call an entity linking model to obtain at least one candidate pair that matches the initial search text, where the candidate pair includes a standard word associated with at least one of the target words in at least one of semantic features and glyph features, and includes the corresponding relationship between the standard word and the entity; According to the corresponding relationship between the standard word and the entity, set the candidate entity corresponding to the standard word in the candidate pair as the target entity corresponding to the target word in the initial search text.

2. The method according to claim 1, wherein Before obtaining the initial search text input by the user, the method further includes: Sample an initial training set to obtain global negative samples of words in each training pair; the initial training set includes multiple training pairs in the medical field, each training pair includes a word and the corresponding relationship with the correct entity, and the global negative samples include the first type of negative samples and the second type of negative samples of words in each training pair; Train the entity linking model based on the initial training set and the global negative samples, so that the feature distance between the target word in the initial training set and the negative sample is greater than a first distance, and the feature distance from the positive sample is less than a second distance, and the first distance is greater than the second distance; Wherein, the first type of negative sample is composed of a candidate entity whose similarity with the target word is higher than a preset similarity and is a wrong entity of the target word; the target word is a word in any training pair of the initial training set; The second type of negative sample is obtained by pairing the target word with candidate entities corresponding to other words in the initial training set except the target word respectively.

3. The method according to claim 2, wherein The sampling of the initial training set to obtain global negative samples of words in each training pair includes: Obtain the first vector representation of the target word and the vector representations of all entity indexes in the initial training set; Obtain the similarity between the first vector representation and the vector representations of each entity index; Form a candidate negative sample with the target word and a candidate entity whose similarity is higher than the preset similarity and is a wrong entity of the target word, and update the candidate negative sample to the first type of negative sample; until the first type of negative samples of all words are obtained.

4. The method according to any one of claims 1 to 3, characterized in that The method further includes: Obtain a first set, the first set includes multiple training pairs, and each training pair includes a first word and a first entity of the first word; Obtain a second set based on the first set, the second set includes multiple key-value pairs, and each key-value pair is composed of the vector representation of the first word and the first entity.

5. The method according to claim 4, wherein After obtaining the second set based on the first set, the method further includes: Obtain the target vector representation of the second word; Traverse the second set based on the target vector representation to obtain a third set, the third set includes multiple key-value pairs recalled by the second word, and the probability that the second word links to the entity in each key-value pair in the third set is non-zero; Obtain the predicted distribution of the second word on each entity in the third set according to the classification result between the second word and the entities in each key-value pair in the third set and a first hyperparameter; Aggregate the probabilities of the same entities in the third set to form a nearest neighbor distribution; Linearly interpolate the nearest neighbor distribution and the predicted distribution to obtain a target probability distribution, which is used to predict the probability that the second word links to the first set.

6. The method according to claim 5, characterized in that, The aggregating the probabilities of the same entities in the third set to form a nearest neighbor distribution includes: Determine a nearest neighbor set according to the similarity between the target vector representation and the keys in each key-value pair, where the nearest neighbor set includes a preset number of nearest neighbor entities of the second word; According to the similarity between the target vector representation and the keys in each key-value pair, obtain a candidate distribution of the first entity on the nearest neighbor set, and aggregate the probabilities of the same entities between key-value pairs to form the nearest neighbor distribution.

7. The method according to claim 6, wherein The third set is included in the second set, the third set includes a first key-value pair and a second key-value pair, and the first key-value pair and the second key-value pair correspond to the same second entity; the aggregating the probabilities of the same entities between key-value pairs includes: Obtain a first similarity and a second similarity; According to the first similarity, the second similarity, and the second hyperparameter, obtain the probability that the second word links to the second entity; Wherein, the first similarity is the similarity between the target vector representation and the vector representation in the first key-value pair, and the second similarity is the similarity between the target vector representation and the vector representation in the second key-value pair.

8. The method according to claim 6 or 7, characterized in that The third set further includes a third key-value pair, where the second vector representation and the third entity included in the third key-value pair are different from the first key-value pair and the second mean pair, and the probability that the second word links to the third entity is non-zero; the method further includes: Obtain a third similarity, where the third similarity is the similarity between the target vector representation and the second vector representation; According to the third similarity and the second hyperparameter, obtain the probability that the second word links to the third entity.

9. The method according to claim 1, wherein Before obtaining the initial search text input by the user, the method further includes: Perform adversarial perturbation detection on the initial search text; If adversarial perturbations that meet preset conditions are detected, at least one of the following items is performed on the initial search text: Remove the adversarial perturbations; Or, return an interception message.

10. An entity linking device, characterized in that, The entity linking device includes: An input / output module, configured to obtain an initial search text input by a user; wherein, at least one target word is included in the initial search text; A processing module, configured to call an entity linking model to obtain at least one candidate pair that matches the initial search text obtained by the input / output module, where the candidate pair includes a standard word associated with at least one of the target words in at least one of semantic features and glyph features; The processing module is further configured to set the candidate entity corresponding to the standard word in the candidate pair as the target entity corresponding to the target word in the initial search text according to the corresponding relationship between the standard word and the entity.

11. A computer device, characterized in that, The computer device includes: At least one processor and a memory; Wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, It includes instructions that, when run on a computer, cause the computer to execute the method according to any one of claims 1-9.

13. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps in the entity linking method according to any one of claims 1 to 9 are implemented.