Training method for entity linking model, entity linking method and device

By using multimodal algorithms and self-attention mechanisms in the entity link model, the entity link model is trained, which solves the problem of low model reliability in multimodal scenarios in the existing technology, and achieves more efficient information fusion and denoising training effects.

CN114896421BActive Publication Date: 2025-06-03ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210622106.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-01
Publication Date
2025-06-03
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

The prior art models used for entity linking in multimodal scenarios are less reliable, and it is difficult to effectively extract and fuse image and text information, resulting in noise information interference and poor training effects.

Method used

The multimodal algorithm is used combined with the self-attention mechanism to train the entity link model. By obtaining training samples containing images and text, the self-attention mechanism is used to determine the correlation between samples, and encoding is performed through image encoder ResNET and text encoder BERT, and finally obtaining the entity link model through backpropagation and derivative training.

Benefits of technology

It improves the reliability of multimodal information fusion, removes noise information, enhances the effectiveness and accuracy of training, and improves the performance of entity link models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114896421B_ABST
    Figure CN114896421B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for training an entity linking model, a method and device for entity linking, including: obtaining a training sample set, where the training sample set includes mention sample data and entity sample data, the mention sample data includes a sample image and a sample text of a mention, and the entity sample data includes a sample image and a sample text of an entity; determining respective binary classification results corresponding to each sample image and each sample text according to a self-attention mechanism, and training an entity linking model according to each binary classification result, where the self-attention mechanism is used to determine the correlation between each sample image and each sample text, and the entity linking model is used to determine an entity corresponding to a mention to be recognized. The correlation between each modality is fully considered, thereby improving the reliability of modality fusion, and noise information is removed to avoid interference of noise information on training, so as to perform training based on effective information, thereby improving the effectiveness and reliability of training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of multimodal machine learning technology, and in particular, to a method for training an entity linking model, a method for entity linking, and an apparatus therefor. Background Art

[0002] Entity linking technology can be applied to search scenarios, which can include fields such as information extraction, information retrieval, content analysis, automatic question answering, knowledge base expansion, etc.

[0003] In the related art, an entity linking model can be constructed through a unimodal algorithm to implement entity linking based on the entity linking model. Summary of the Invention

[0004] The present disclosure provides a method for training an entity linking model, a method for entity linking, and an apparatus therefor, which are used to improve the reliability of the entity linking model.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for training an entity linking model, including:

[0006] Obtain a training sample set, where the training sample set includes mention sample data and entity sample data, the mention sample data includes a sample image and a sample text of the mention, and the entity sample data includes a sample image and a sample text of the entity corresponding to the mention;

[0007] Determine a binary classification result corresponding to each sample image and each sample text according to the self-attention mechanism, and train an entity linking model according to each binary classification result, where the self-attention mechanism is used to determine the correlation between each sample image and each sample text, and the entity linking model is used to determine the entity corresponding to the mention to be recognized.

[0008] In an embodiment of the present disclosure, after obtaining the training sample set, the method further includes:

[0009] Fuse each sample text to obtain a fused text;

[0010] And, determining the binary classification result corresponding to each sample image and each sample text according to the self-attention mechanism includes: determining the binary classification result corresponding to each sample image, each sample text, and the fused text according to the self-attention mechanism.

[0011] In an embodiment of the present disclosure, after fusing each sample text to obtain a fused text, the method further includes:

[0012] Perform encoding processing on each sample image, each sample text, and the fused text respectively to obtain respective corresponding encoding vectors;

[0013] Moreover, determining the binary classification results corresponding to each sample image, each sample text, and the fused text according to the self-attention mechanism includes: determining the binary classification results corresponding to each encoding vector according to the self-attention mechanism.

[0014] In an embodiment of the present disclosure, encoding each sample image, each sample text, and the fused text respectively to obtain the corresponding encoding vectors includes:

[0015] Encoding each sample image according to a preset image encoder ResNET to obtain the encoding vectors corresponding to each sample image;

[0016] Encoding each text according to a preset text encoder BERT to obtain the encoding vectors corresponding to each text, where each text includes each sample text and the fused text.

[0017] In an embodiment of the present disclosure, training the entity linking model according to each binary classification result includes:

[0018] Performing backpropagation derivation according to each binary classification result to obtain the entity linking model.

[0019] In an embodiment of the present disclosure, there is a one-to-one correspondence between the sample image and the image encoder, and a one-to-one correspondence between the text and the text encoder.

[0020] In a second aspect, an embodiment of the present disclosure provides a method for entity linking, which is applied to a knowledge graph and includes:

[0021] Obtaining a mention to be recognized;

[0022] Determining an entity corresponding to the mention to be recognized according to a pre-trained entity linking model;

[0023] Wherein, the entity linking model is trained based on the method as described in the first aspect.

[0024] In an embodiment of the present disclosure, after obtaining the mention to be recognized, the method further includes:

[0025] Obtaining multiple candidate entities with the highest similarity to the mention to be recognized from a preset knowledge graph, where each candidate entity has a candidate image and a candidate text;

[0026] And, determining an entity corresponding to the to-be-recognized mention according to a pre-trained entity linking model, including: inputting the to-be-recognized mention, each candidate image, and each candidate text into the entity linking model, obtaining a matching result between each candidate entity and the to-be-recognized mention, and determining an entity corresponding to the to-be-recognized mention from each candidate entity according to each matching result;

[0027] Wherein, the matching result is used to characterize the similarity between the to-be-recognized mention and the candidate entity.

[0028] In an embodiment of the present disclosure, the to-be-recognized mention is determined based on a search request, and the entity corresponding to the to-be-recognized mention is a search result corresponding to the search request; wherein,

[0029] The search request includes the to-be-recognized mention; or,

[0030] The search request includes a to-be-recognized image and / or a to-be-recognized text.

[0031] In an embodiment of the present disclosure, determining an entity corresponding to the to-be-recognized mention from each candidate entity according to each matching result includes:

[0032] Determining a matching result used to characterize the maximum similarity from each matching result, and determining the candidate entity corresponding to the matching result used to characterize the maximum similarity as the entity corresponding to the to-be-recognized mention.

[0033] In a third aspect, an embodiment of the present disclosure provides a training device for an entity linking model, including:

[0034] A first acquisition unit, configured to acquire a training sample set, wherein the training sample set includes mention sample data and entity sample data, the mention sample data includes sample images and sample texts of mentions, and the entity sample data includes sample images and sample texts of entities corresponding to the mentions;

[0035] A first determination unit, configured to determine binary classification results corresponding to each sample image and each sample text according to a self-attention mechanism;

[0036] A training unit, configured to train an entity linking model according to each binary classification result, wherein the self-attention mechanism is used to determine the correlation between each sample image and each sample text, and the entity linking model is used to determine an entity corresponding to a to-be-recognized mention.

[0037] In an embodiment of the present disclosure, the device further includes:

[0038] A fusion unit, configured to perform a fusion process on each sample text to obtain a fused text;

[0039] Moreover, the first determination unit is configured to determine the binary classification results corresponding to each sample image, each sample text, and the fused text according to the self-attention mechanism.

[0040] In an embodiment of the present disclosure, the apparatus further includes:

[0041] An encoding unit, configured to perform encoding processing on each sample image, each sample text, and the fused text respectively to obtain the encoding vectors corresponding to each of them;

[0042] Moreover, the first determination unit is configured to determine the binary classification results corresponding to each encoding vector according to the self-attention mechanism.

[0043] In an embodiment of the present disclosure, the encoding unit includes:

[0044] A first encoding unit, configured to perform encoding processing on each sample image according to a preset image encoder ResNET to obtain the encoding vector corresponding to each sample image;

[0045] A second encoding unit, configured to perform encoding processing on each text according to a preset text encoder BERT to obtain the encoding vector corresponding to each text, where each text includes each sample text and the fused text.

[0046] In an embodiment of the present disclosure, the training unit is configured to perform backpropagation derivation according to each binary classification result to obtain the entity linking model.

[0047] In an embodiment of the present disclosure, there is a one-to-one correspondence between the sample images and the image encoder, and a one-to-one correspondence between the texts and the text encoder.

[0048] Fourthly, an embodiment of the present disclosure provides an entity linking apparatus, which is applied to a knowledge graph and includes:

[0049] A second acquisition unit, configured to acquire a mention to be recognized;

[0050] A second determination unit, configured to determine an entity corresponding to the mention to be recognized according to a pre-trained entity linking model;

[0051] Wherein, the entity linking model is trained based on the method as described in the first aspect.

[0052] In an embodiment of the present disclosure, the apparatus further includes:

[0053] A third acquisition unit, configured to acquire a plurality of candidate entities with the highest similarity to the mention to be recognized from a preset knowledge graph, where each candidate entity has a candidate image and a candidate text;

[0054] And, the second determination unit includes:

[0055] An input subunit, configured to input the mention to be recognized, each candidate image, and each candidate text into the entity linking model, so as to obtain a matching result between each candidate entity and the mention to be recognized;

[0056] A determination subunit, configured to determine an entity corresponding to the mention to be recognized from each candidate entity according to each matching result;

[0057] Wherein, the matching result is used to characterize the similarity between the mention to be recognized and the candidate entity.

[0058] In an embodiment of the present disclosure, the mention to be recognized is determined based on a search request, and the entity corresponding to the mention to be recognized is a search result corresponding to the search request; wherein,

[0059] The search request includes the mention to be recognized; or,

[0060] The search request includes a to-be-recognized image and / or a to-be-recognized text.

[0061] In an embodiment of the present disclosure, the determination subunit includes:

[0062] A first determination module, configured to determine, from each matching result, a matching result used to characterize the maximum similarity;

[0063] A second determination module, configured to determine the candidate entity corresponding to the matching result characterizing the maximum similarity as the entity corresponding to the mention to be recognized.

[0064] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including:

[0065] At least one processor; and

[0066] A memory communicatively connected to the at least one processor; wherein,

[0067] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the electronic device is enabled to execute the method described in any one of the first aspect or the second aspect of the present disclosure.

[0068] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the first aspect or the second aspect of the present disclosure is implemented.

[0069] In a seventh aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any one of the first aspect of the present disclosure is implemented.

[0070] The present disclosure is trained by combining two dimensions of images and texts, and by combining the self-attention mechanism. The technical features fully consider the correlation between multimodals, thereby improving the reliability of modal fusion, removing noise information, avoiding the interference of noise information on training, and training based on effective information, so as to improve the effectiveness and reliability of training. Brief Description of the Drawings

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0072] Figure 1 Schematic diagram of the training method of the entity linking model according to an embodiment of the present disclosure;

[0073] Figure 2 Flowchart of the training method of the entity linking model according to another embodiment of the present disclosure;

[0074] Figure 3 Flowchart of the training method of the entity linking model according to another embodiment of the present disclosure;

[0075] Figure 4 Principle diagram of the training method of the entity linking model according to an embodiment of the present disclosure;

[0076] Figure 5 Schematic diagram of the entity linking method according to an embodiment of the present disclosure;

[0077] Figure 6 Schematic diagram of the entity linking method according to another embodiment of the present disclosure;

[0078] Figure 7 Schematic diagram of the training device of the entity linking model according to an embodiment of the present disclosure;

[0079] Figure 8 Schematic diagram of the training device of the entity linking model according to another embodiment of the present disclosure;

[0080] Figure 9 Schematic diagram of the entity linking device according to an embodiment of the present disclosure;

[0081] Figure 10 Schematic diagram of the entity linking device according to another embodiment of the present disclosure;

[0082] Figure 11Schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present disclosure.

[0083] Through the above-mentioned drawings, specific embodiments of the present disclosure have been shown, and more detailed descriptions will be given later. These drawings and textual descriptions are not intended to limit the scope of the concept of the present disclosure in any way, but to illustrate the concept of the present disclosure to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0084] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0085] The terms "first", "second", "third", etc. in the specification, claims, and drawings of the present disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein.

[0086] In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0087] For the convenience of readers' understanding of the present disclosure, some terms are explained as follows:

[0088] Multi-modal machine learning (MMML) aims to achieve the ability to process and understand multi-source modal information through machine learning methods, such as multi-modal learning between images, videos, audio, and semantics.

[0089] Entity Linking, also known as entity coreference resolution, is a task that requires us to identify the words representing entities in unstructured data (i.e., so-called mentions, referring terms for a certain entity), and find the entity represented by the mention from a knowledge base (such as a domain thesaurus or a knowledge graph, etc.).

[0090] A mention, also known as an entity reference, refers to the words representing entities in unstructured data. Among them, unstructured data refers to data with irregular or incomplete data structures and no predefined data models, including office documents, texts, pictures, HTML, various reports, images, and audio / video information, etc.

[0091] A Knowledge Graph, known as knowledge domain visualization or knowledge domain mapping map in the library and information science field, is a series of various graphs showing the development process and structural relationships of knowledge. It uses visualization technology to describe knowledge resources and their carriers, mine, analyze, construct, draw, and display knowledge and the interconnections between them. That is, a Knowledge Graph is composed of entities and the relationships between entities, and is presented in the form of a graph.

[0092] The self-attention mechanism is a variant of the attention mechanism, which reduces the dependence on external information and is better at capturing the internal correlations of data or features. The application of the self-attention mechanism in text mainly solves the long-distance dependence problem by calculating the mutual influence between words.

[0093] Exemplarily, entity linking can extract the mentions (or entity references) in a piece of text and map these mentions to the unique entities in a specified knowledge base. Entity linking can help find the important semantic information in a sentence, judge the different meanings of words in different context, and is indispensable for helping computers understand natural language.

[0094] Entity linking technology can be applied to search scenarios, which can include fields such as information extraction, information retrieval, content analysis, automatic question answering, and knowledge base expansion.

[0095] In some embodiments, entity linking can be implemented through a network model. For example, an entity linking model can be trained based on an algorithm of natural language processing (NLP) single modality to determine the entity corresponding to the mention based on the entity linking model. For example, an entity linking model can be trained based on the text including the mention.

[0096] However, with the continuous application of visual information such as images and videos in the Knowledge Graph, it is difficult for the above-mentioned entity linking model trained by a single modality algorithm to be applicable to multimodal scenarios. Therefore, the reliability of the entity corresponding to the mention determined by the entity linking model trained based on the above method is relatively low.

[0097] In some other embodiments, an entity linking model can also be obtained by combining multimodal algorithm training. Exemplarily, an unsupervised encoder can be used to encode the entity, the text of the mention, and the image respectively, and a fully connected layer can be used for modality fusion, thereby training an entity linking model.

[0098] As Figure 1 shown, the training method of the entity linking model includes:

[0099] S101: Obtain a training sample set.

[0100] Among them, the training sample set includes mention sample data and entity sample data. The mention sample data includes the sample image and sample text of the mention, and the entity sample data includes the sample image and sample text of the entity.

[0101] The mention sample data and the entity sample data are relative concepts. The mention sample data refers to the sample data of the mention, including: the image used for visual representation of the mention (i.e., the sample image of the mention), and the descriptive text containing the mention (i.e., the sample text of the mention).

[0102] Correspondingly, the entity sample data refers to the sample data of the entity, including: the image used for visual representation of the entity (i.e., the sample image of the mention), and the descriptive text containing the entity (i.e., the sample text of the entity). The entity here can be understood as a candidate entity, and the number of candidate entities can be multiple, so the number of entity sample data is also multiple, and one candidate entity corresponds to one entity sample data.

[0103] Among them, the number of candidate entities can be determined based on requirements, historical records, and experiments, etc., and this embodiment does not make a limitation. For example, the number of candidate entities is ten.

[0104] For example, for a training scenario with relatively high reliability, the number of candidate entities can be relatively large. On the contrary, for a training scenario with relatively low reliability, the number of candidate entities can be relatively small.

[0105] S102: Input the mention sample data and the entity sample data into a sentence vector model (sent2vec), and output the corresponding encoding vectors respectively.

[0106] Among them, the sentence vector model can also be called sentence embedding. Exemplarily, input the mention sample data into the sentence vector model, and output the encoding vector corresponding to the mention sample data; input the entity sample data into the sentence vector model, and output the encoding vector corresponding to the entity sample data.

[0107] S103: Extract features from each encoding vector to obtain the data features corresponding to each encoding vector.

[0108] It should be understood that, for the convenience of distinguishing the encoding vector corresponding to the mention sample data from the encoding vector corresponding to the entity sample data, the encoding vector corresponding to the mention sample data is called the mention encoding vector, and the encoding vector corresponding to the entity sample data is called the entity encoding vector.

[0109] Correspondingly, feature extraction is performed on the mention encoding vector to obtain data features corresponding to the mention encoding vector; feature extraction is performed on the entity encoding vector to obtain data features corresponding to the entity encoding vector.

[0110] Among them, feature extraction can be performed on each encoding vector by means of deep learning (Feature extraction), and the specific implementation principle will not be elaborated here.

[0111] S104: Input each data feature into its corresponding first fully connected layer (fully connected layers, FC), activation function (Rectified Linear Units, ReLu), text connection function (Concatenate), and second fully connected layer in sequence, and output triplet loss information.

[0112] It should be understood that the number of fully connected layers can be two layers or more layers. In this embodiment, two layers are taken as an example for illustration, and it should not be construed as a limitation on the number of fully connected layers.

[0113] S105: Combine the output results of each second fully connected layer (Comb) to obtain combined information.

[0114] S106: Input the combined information and the heat value corresponding to the preset entity sample data into a multi-layer perceptron (Multilayer Perceptron, MLP) network, and output binary cross entropy loss information.

[0115] S107: Generate an entity link model according to the triplet loss information and the binary cross entropy loss information.

[0116] However, when using the above method to train an entity link model, the encoding ability is relatively weak, and the fully connected layers corresponding to each data feature are in a "parallel" manner, and the effect of modality fusion is relatively poor, making it difficult to extract effective information and remove noise information.

[0117] To avoid the above drawbacks, the inventors of the present disclosure have obtained the inventive concept of the present disclosure through creative labor: training an entity link model based on a multi-modal algorithm and combining a self-attention mechanism.

[0118] Next, the technical solutions of the present disclosure will be described in detail through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0119] Please refer to Figure 2 , Figure 2 which is a flowchart of a method for training an entity linking model according to another embodiment of the present disclosure. As Figure 2 shown, the method includes:

[0120] S201: Obtain a training sample set. Among them, the training sample set includes mention sample data and entity sample data. The mention sample data includes the mentioned sample images and sample texts, and the entity sample data includes the entity's sample images and sample texts.

[0121] Exemplarily, the execution subject of this embodiment can be a training device for an entity linking model (hereinafter simply referred to as the training device). The training device can be a computer, a server (such as a cloud server or a local server), a terminal device, a processor, a chip, etc. This embodiment is not limited.

[0122] It should be understood that in order to avoid cumbersome statements, the same technical features of this embodiment as those of the above embodiment will not be described again in this embodiment.

[0123] S202: Determine the binary classification results corresponding to each sample image and each sample text according to the self-attention mechanism, and train an entity linking model according to each binary classification result.

[0124] Among them, the self-attention mechanism is used to determine the correlation between each sample image and each sample text. The entity linking model is used to determine the entity corresponding to the mention to be recognized. The binary classification result is used to characterize the association degree between the mention and the entity.

[0125] The fully connected layer usually does not consider that for the same word, it may have different meanings at different positions in the sentence. Therefore, for the mention in the sample text, the position of the mention in the sample text may cause the meaning of the mention to be different, especially when the length of the sample text is long. Based on the Figure 1 embodiment shown, the performance of the entity linking model trained may seriously decline.

[0126] The self-attention mechanism enables the training process to pay attention to and learn the correlations between various input information. In this embodiment, the various input information includes various sample images and various sample texts, and specifically includes: the mentioned sample images, the mentioned sample texts, the sample images of entities, and the sample texts of entities. Therefore, the entity linking model trained by combining the self-attention mechanism can learn the correlations between the mentioned sample images, the mentioned sample texts, the sample images of entities, and the sample texts of entities, thereby improving the effect of modality fusion, and extracting effective information from each input information (the mentioned sample images, the mentioned sample texts, the sample images of entities, and the sample texts of entities), removing noise information, and further improving the effectiveness and reliability of training.

[0127] Correspondingly, during the inference stage, that is, when applying the entity linking model trained by the method described in the embodiment as Figure 2 shown, the accuracy and reliability of determining the entity corresponding to the mention to be recognized can be improved.

[0128] Based on the above analysis, the present disclosure provides a method for training an entity linking model, including: obtaining a training sample set, where the training sample set includes mention sample data and entity sample data, the mention sample data includes the mentioned sample images and sample texts, the entity sample data includes the sample images and sample texts of entities, determining the binary classification results corresponding to each sample image and each sample text according to the self-attention mechanism, and training an entity linking model according to each binary classification result, where the self-attention mechanism is used to determine the correlations between each sample image and each sample text, and the entity linking model is used to determine the entity corresponding to the mention to be recognized. In this embodiment, by combining the two dimensions of image and text for training, and the technical feature of combining the self-attention mechanism for training, the correlations between various modalities are fully considered, thereby improving the reliability of modality fusion, removing noise information, avoiding the interference of noise information on training, and training based on effective information, thereby improving the effectiveness and reliability of training.

[0129] To enable readers to more deeply understand the implementation principle of the present disclosure, the embodiments of the present disclosure will now be described in more detail in combination with Figure 3 this.

[0130] As Figure 3 shown, the method for training an entity linking model includes:

[0131] S301: Obtain the sample images representing the mention and the sample texts including the mention.

[0132] Similarly, to avoid cumbersome statements, for the technical features that are the same as those in the above embodiments in this embodiment, this embodiment will not be elaborated again.

[0133] S302: Retrieve the entity corresponding to the mention from a preset database, and obtain the sample image for characterizing the entity and the sample text including the entity.

[0134] Among them, the preset knowledge base can be a domain word library or a knowledge graph, etc., which is not limited in this embodiment.

[0135] Regarding the method for retrieving the entity corresponding to the mention, this embodiment is not limited either. For example, the entity corresponding to the mention can be retrieved from the preset database by calculating the similarity. From the combined analysis, the number of entities can be determined based on requirements, historical records, and experiments, etc.

[0136] Exemplarily, if the number of entities is N (N is a positive integer greater than or equal to 1), then the first N entities with the greatest similarity, that is, the TopN entities, can be determined from the preset database, and the sample image and sample text corresponding to each entity are obtained.

[0137] That is to say, through S301 and S302, a training sample set can be obtained. The mention sample data includes the sample image and sample text of the mention obtained through S301, and the entity sample data includes the sample image and sample text of the entity obtained through S302.

[0138] S303: Perform a fusion process on the sample text of the mention and the sample text of the entity to obtain a fusion text.

[0139] Among them, the fusion process can be understood as combining two texts into one text, that is, the fusion text includes both the content of the sample text of the mention and the content of the sample text of the entity.

[0140] Exemplarily, the fusion process can be a splicing process, that is, splicing the sample text of the mention and the sample text of the entity into one text (i.e., the fusion text). The order of splicing is not limited in this embodiment. In the fusion text, the sample text of the mention can be the front part of the fusion text or the back part of the fusion text.

[0141] The fusion process can also be an insertion process. For example, taking the sample text of the mention as a whole and inserting it into a certain position in the sample text of the entity, or randomly inserting the sample text of the mention into multiple positions in the sample text of the entity. Of course, it can also be inserting the sample text of the entity into the sample text of the mention, which will not be elaborated here.

[0142] It should be noted that in this embodiment, by performing a fusion process on the sample text of the mention and the sample text of the entity to obtain a fusion text for training an entity linking model, the information of the mention and the information of the entity can be combined for training, thereby improving the accuracy and reliability of the training.

[0143] S304: Input the mentioned sample image and the sample image of the entity into their respective corresponding image encoders ResNET (Convolutional Neural Network (CNN)), and output their respective corresponding encoded vectors.

[0144] Regarding the encoding principle of the image encoder, reference can be made to the related technology, which will not be elaborated here.

[0145] For a clearer understanding of the embodiments of the present disclosure, the following will be elaborated in conjunction with the schematic diagram shown as Figure 4 shown. As Figure 4 shown:

[0146] Input the mentioned sample image into the first image encoder to output the first encoded vector. Input the sample image of the entity into the second image encoder to output the second encoded vector.

[0147] Among them, the first image encoded vector is used to characterize the features of the mentioned visual dimension. The second encoded vector is used to characterize the features of the visual dimension of the entity.

[0148] S305: Input the mentioned sample text, the sample text of the entity, and the fused text into their respective corresponding text encoders BERT, and output their respective corresponding encoded vectors.

[0149] Similarly, regarding the encoding principle of the text encoder, reference can be made to the related technology, which will not be elaborated here.

[0150] As Figure 4 shown, input the mentioned sample text into the first text encoder to output the third encoded vector. Input the sample text of the entity into the second text encoder to output the fourth encoded vector. Input the fused text into the third text encoder to output the fifth encoded vector.

[0151] Among them, the third encoded vector is used to characterize the features of the mentioned text (such as semantics) dimension. The fourth encoded vector is used to characterize the features of the text (such as semantics) dimension of the entity. The fifth encoded vector is used to characterize the features of the text (such as semantics) dimension of "mention + entity".

[0152] It should be noted that in this embodiment, by adopting different encoding methods to encode and process the sample images and texts, the flexibility and diversity of the encoding process are improved.

[0153] S306: Input each encoded vector into the self-attention network, and output the binary classification result (logits) corresponding to each encoded vector.

[0154] Among them, the binary classification result, i.e., the classification cross-entropy loss, is a probability value between 0 and 1, which is used to characterize the degree of association between a mention and an entity. Relatively speaking, the larger the probability value, the greater the degree of association; the smaller the probability value, the smaller the degree of association.

[0155] Exemplarily, in combination with the above analysis and Figure 4 , input the first encoding vector, the second encoding vector, the third encoding vector, the fourth encoding vector, and the fifth encoding vector into the self-attention network, output the binary classification result C1 corresponding to the first encoding vector, output the binary classification result C2 corresponding to the second encoding vector, output the binary classification result C3 corresponding to the third encoding vector, output the four-classification result C4 corresponding to the second encoding vector, and output the binary classification result C5 corresponding to the fifth encoding vector.

[0156] S307: Perform backpropagation and derivation according to each binary classification result to obtain an entity linking model.

[0157] Exemplarily, according to each binary classification result and a preset label, determine the difference information between each prediction result (i.e., each binary classification result) and the true result (i.e., the preset label), and adjust the model parameters of each encoder (such as Figure 4 each encoder shown therein) and the self-attention network, so as to obtain an entity linking model.

[0158] In combination with the above analysis and Figure 4 , the number of binary classification results is 5, which are C1 to C5 respectively. Then, for each binary classification result, the difference information between the binary classification result and the preset label can be determined, so as to obtain 5 pieces of difference information.

[0159] Correspondingly, backpropagation and derivation can be performed based on the 5 pieces of difference information to obtain an entity linking model. And backpropagation and derivation can be performed based on the average difference information of the 5 pieces of difference information to obtain an entity linking model, or backpropagation and derivation can be performed based on the weighted average difference information of the 5 pieces of difference information to obtain an entity linking model, etc. This embodiment does not make a limitation.

[0160] After obtaining the entity linking model through the above embodiment training, the entity linking model can be applied to determine the entity of the mention to be recognized through the entity linking model.

[0161] Please refer to Figure 5 , Figure 5 is a schematic diagram of a method for entity linking according to an embodiment of the present disclosure. This method can be applied to a knowledge graph, such as Figure 5 shown, and this method includes:

[0162] S501: Obtain the mention to be recognized.

[0163] Exemplarily, the execution subject of this embodiment may be a device for entity linking. This device may be the same as the training device or different from the training device, which is not limited in this embodiment.

[0164] S502: Determine the entity corresponding to the mention to be recognized according to the pre-trained entity linking model.

[0165] Among them, the entity linking model is trained based on the method described in any of the above embodiments.

[0166] To enable readers to more deeply understand the implementation principle of the entity linking method, the following will be combined with Figure 6 for a more detailed elaboration.

[0167] As Figure 6 shown, the entity linking method includes:

[0168] S601: Obtain the mention to be recognized.

[0169] Similarly, to avoid cumbersome statements, the same technical features of this embodiment and the above embodiments will not be elaborated in this embodiment.

[0170] Among them, the method for obtaining the mention to be recognized is not limited in this embodiment. Exemplarily, the device for entity linking may be a search device with a search function. The user can establish a communication connection between the user device and the search device by means of touch or voice, etc., and can initiate a search request to the search device based on touch or voice, etc. The search request may carry the mention to be recognized. Correspondingly, the search device obtains the mention to be recognized.

[0171] For example, if the search request is "XX item", the search device can determine that the mention to be recognized carried in this request is "XX". Among them, the item may be a virtual item or a physical item.

[0172] In some other embodiments, the search request may include the text to be recognized. Correspondingly, the search device recognizes the text to be recognized to obtain the mention to be recognized in the text to be recognized.

[0173] In still some other embodiments, the search request may include the image to be recognized. Correspondingly, the search device can recognize the image to be recognized to obtain the mention to be recognized represented by the image to be recognized.

[0174] In some other embodiments, the search request may include the text to be recognized and the image to be recognized. Correspondingly, the search device may recognize the text to be recognized to obtain the mentions to be recognized in the text to be recognized, and may recognize the image to be recognized to obtain the mentions to be recognized represented by the image to be recognized, and determine the mentions to be recognized of the final search request according to the mentions to be recognized in the text to be recognized and the mentions to be recognized represented by the image to be recognized.

[0175] For example, if the mentions to be recognized in the text to be recognized are the same as the mentions to be recognized represented by the image to be recognized, then the same mentions to be recognized are determined as the mentions to be recognized of the final search request. If the mentions to be recognized in the text to be recognized are different from the mentions to be recognized represented by the image to be recognized, then the mentions to be recognized in the text to be recognized may be corrected based on the mentions to be recognized represented by the image to be recognized, so as to obtain the mentions to be recognized of the final search request.

[0176] S602: Obtain multiple candidate entities with the highest similarity to the mentions to be recognized from a preset knowledge graph.

[0177] Exemplarily, the knowledge graph includes multiple entities. Calculate the similarity between the mentions to be recognized and each entity respectively, and select M (M is a positive integer greater than 1) entities with the largest similarities as candidate entities.

[0178] Similarly, M can be determined based on requirements, historical records, and experiments, etc., and this embodiment does not make any limitations.

[0179] S603: For each candidate entity, obtain the candidate image and candidate text of the candidate entity from the knowledge graph.

[0180] S604: Perform fusion processing on the candidate samples and the mentions to be recognized to obtain a fused text.

[0181] S605: Input the candidate image into an image encoder to output the encoded vector of the candidate image. Among them, the entity linking model includes the image encoder.

[0182] S606: Input the candidate text and the fused text into their respective text encoders to output their respective encoded vectors. Among them, the entity linking model includes each text encoder.

[0183] Combined with the above analysis, in one example, the search request may include the text to be recognized, and the text to be recognized includes the mentions to be recognized. Then S604 can be replaced with: Perform fusion processing on the candidate sample text and the text to be recognized to obtain a fused text.

[0184] In another example, the search request may include an image to be recognized, and the image to be recognized is used to characterize the mention to be recognized. Then, S604 can be replaced with: input the image to be recognized into an image encoder to output an encoded vector of the image to be recognized.

[0185] Alternatively, by recognizing the image to be recognized, text including the mention to be recognized can be obtained. Then, S604 can be replaced with: perform a fusion process on the candidate sample text and the text including the mention to be recognized to obtain a fused text. And add the step of "inputting the image to be recognized into an image encoder to output an encoded vector of the image to be recognized".

[0186] In yet another example, the search request may include an image to be recognized and text to be recognized. The image to be recognized is an image characterizing the mention to be recognized, and the text to be recognized is text including the mention to be recognized.

[0187] Correspondingly, S604 can be replaced with: perform a fusion process on the candidate sample text and the text including the mention to be recognized to obtain a fused text. And add the step of "inputting the image to be recognized into an image encoder to output an encoded vector of the image to be recognized".

[0188] S607: Input each encoded vector into a self-attention model to obtain a binary classification result corresponding to each encoded vector. Among them, the entity linking model includes a self-attention model. The binary classification result is used to characterize the degree of association between the mention to be recognized and the candidate entity.

[0189] S608: Determine the maximum degree of association according to each binary classification result.

[0190] S609: By analogy, obtain the maximum degree of association corresponding to each candidate entity, and extract the candidate entity corresponding to the maximum degree of association from each maximum degree of association as the entity corresponding to the mention to be recognized.

[0191] Please refer to Figure 7 , Figure 7 which is a schematic diagram of a training device for an entity linking model according to an embodiment of the present disclosure. As Figure 7 shown, the device 700 includes:

[0192] A first acquisition unit 701, configured to acquire a training sample set, where the training sample set includes mention sample data and entity sample data. The mention sample data includes a sample image and sample text of the mention, and the entity sample data includes a sample image and sample text of the entity corresponding to the mention.

[0193] A first determination unit 702, configured to determine a binary classification result corresponding to each sample image and each sample text according to a self-attention mechanism.

[0194] A training unit 703 is configured to train an entity linking model based on the binary classification results. The self-attention mechanism is used to determine the correlations between the sample images and the sample texts, and the entity linking model is used to determine the entity corresponding to the mention to be recognized.

[0195] Please refer to Figure 8 , Figure 8 which is a schematic diagram of a training device for an entity linking model according to another embodiment of the present disclosure. As Figure 8 shown, the device 800 includes:

[0196] A first acquisition unit 801 is configured to acquire a training sample set, where the training sample set includes mention sample data and entity sample data. The mention sample data includes the sample images and sample texts of the mentions, and the entity sample data includes the sample images and sample texts of the entities corresponding to the mentions.

[0197] A fusion unit 802 is configured to perform a fusion process on the sample texts to obtain a fused text.

[0198] An encoding unit 803 is configured to perform encoding processes on the sample images, the sample texts, and the fused text respectively to obtain their corresponding encoding vectors.

[0199] In some embodiments, as Figure 8 shown, the encoding unit 803 includes:

[0200] A first encoding unit 8031 is configured to perform an encoding process on the sample images according to a preset image encoder ResNET to obtain the encoding vectors corresponding to the sample images respectively.

[0201] A second encoding unit 8032 is configured to perform an encoding process on the texts according to a preset text encoder BERT to obtain the encoding vectors corresponding to the texts respectively, where the texts include the sample texts and the fused text.

[0202] In an embodiment of the present disclosure, there is a one-to-one correspondence between the sample images and the image encoder, and a one-to-one correspondence between the texts and the text encoder.

[0203] A first determination unit 804 is configured to determine the binary classification results corresponding to the encoding vectors according to the self-attention mechanism, where the self-attention mechanism is used to determine the correlations between the sample images and the sample texts.

[0204] A training unit 805 is configured to perform backpropagation and derivation according to the binary classification results to obtain the entity linking model, where the entity linking model is used to determine the entity corresponding to the mention to be recognized.

[0205] Please refer toFigure 9 , Figure 9 Schematic diagram of an entity linking device according to an embodiment of the present disclosure. The device can be applied to a knowledge graph, such as Figure 9 shown. The device 900 includes:

[0206] A second acquisition unit 901, configured to acquire a mention to be recognized.

[0207] A second determination unit 902, configured to determine an entity corresponding to the mention to be recognized according to a pre-trained entity linking model.

[0208] Wherein, the entity linking model is trained based on the method described in any one of the foregoing embodiments.

[0209] Please refer to Figure 10 , Figure 10 Schematic diagram of an entity linking device according to another embodiment of the present disclosure. The device can be applied to a knowledge graph, such as Figure 10 shown. The device 1000 includes:

[0210] A second acquisition unit 1001, configured to acquire a mention to be recognized.

[0211] In some embodiments, the mention to be recognized is determined based on a search request, and the entity corresponding to the mention to be recognized is a search result corresponding to the search request; wherein,

[0212] the search request includes the mention to be recognized; or,

[0213] the search request includes a to-be-recognized image and / or to-be-recognized text.

[0214] A third acquisition unit 1002, configured to acquire a plurality of candidate entities with the highest similarity to the mention to be recognized from a preset knowledge graph, where each candidate entity has a candidate image and a candidate text.

[0215] A second determination unit 1003, configured to determine an entity corresponding to the mention to be recognized according to a pre-trained entity linking model.

[0216] Wherein, the entity linking model is trained based on the method described in the first aspect.

[0217] Combined with Figure 10 it can be seen that in some embodiments, the second determination unit 1003 includes:

[0218] An input subunit 10031, configured to input the mention to be recognized, each candidate image, and each candidate text into the entity linking model to obtain a matching result of each candidate entity and the mention to be recognized.

[0219] Determine subunit 10032, configured to determine, according to each matching result, an entity corresponding to the to-be-recognized mention from each candidate entity.

[0220] Wherein, the matching result is used to characterize the similarity between the to-be-recognized mention and the candidate entity.

[0221] Combined Figure 10 It can be seen that, in some embodiments, the determining subunit 10032 includes:

[0222] A first determining module, configured to determine, from each matching result, a matching result used to characterize the maximum similarity.

[0223] A second determining module, configured to determine the candidate entity corresponding to the matching result characterizing the maximum similarity as the entity corresponding to the to-be-recognized mention.

[0224] Figure 11 This is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present disclosure. As Figure 11 shown, the electronic device 1100 of the embodiments of the present disclosure may include: at least one processor 1101 ( Figure 11 only one processor is shown in the figure); and a memory 1102 communicatively connected to at least one processor. Wherein, the memory 1102 stores instructions executable by at least one processor 1101, and the instructions are executed by at least one processor 1101 so that the electronic device 1100 can execute the technical solutions in any of the foregoing method embodiments.

[0225] Optionally, the memory 1102 may be either independent or integrated with the processor 1101.

[0226] When the memory 1102 is a device independent of the processor 1101, the electronic device 1100 further includes: a bus 1103, configured to connect the memory 1102 and the processor 1101.

[0227] The electronic device provided by the embodiments of the present disclosure can execute the technical solutions in any of the foregoing method embodiments, and the implementation principles and technical effects are similar, which will not be elaborated herein.

[0228] The embodiments of the present disclosure further provide a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it is used to implement the technical solutions in any of the foregoing method embodiments.

[0229] The embodiments of the present disclosure provide a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the technical solutions in any of the foregoing method embodiments.

[0230] An embodiment of the present disclosure further provides a chip, including: a processing module and a communication interface, and the processing module can execute the technical solutions in the foregoing method embodiments.

[0231] Further, the chip further includes a storage module (such as a memory), the storage module is used to store instructions, the processing module is used to execute the instructions stored in the storage module, and the execution of the instructions stored in the storage module enables the processing module to execute the technical solutions in the foregoing method embodiments.

[0232] It should be understood that the foregoing processor may be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), or may also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0233] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0234] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the bus in the drawings of the present disclosure is not limited to only one bus or one type of bus.

[0235] The foregoing storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0236] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device.

[0237] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present disclosure.

Claims

1. A method for training an entity linking model, comprising: Obtaining a training sample set, wherein the training sample set includes mention sample data and entity sample data, the mention sample data includes sample images and sample texts of mentions, and the entity sample data includes sample images and sample texts of entities corresponding to the mentions; Determining binary classification results corresponding to each sample image and each sample text according to the self-attention mechanism, and training an entity linking model based on each binary classification result, wherein the self-attention mechanism is used to determine the correlation between each sample image and each sample text, the binary classification result is a probability value between 0 and 1 and is used to characterize the association degree between a mention and an entity, and the entity linking model is used to determine an entity corresponding to a mention to be recognized; After obtaining the training sample set, the method further includes: Fusing the sample text of a mention and the sample text of an entity to obtain a fused text; And, determining binary classification results corresponding to each sample image and each sample text according to the self-attention mechanism includes: determining binary classification results corresponding to each sample image, each sample text, and the fused text according to the self-attention mechanism.

2. The method according to claim 1, wherein, After fusing the sample text of a mention and the sample text of an entity to obtain a fused text, the method further includes: Performing encoding processing on each sample image, each sample text, and the fused text respectively to obtain respective corresponding encoding vectors; And, determining binary classification results corresponding to each sample image, each sample text, and the fused text according to the self-attention mechanism includes: determining binary classification results corresponding to each encoding vector according to the self-attention mechanism.

3. The method according to claim 2, wherein, Performing encoding processing on each sample image, each sample text, and the fused text respectively to obtain respective corresponding encoding vectors includes: Performing encoding processing on each sample image according to a preset image encoder ResNET to obtain an encoding vector corresponding to each sample image; Performing encoding processing on each text according to a preset text encoder BERT to obtain an encoding vector corresponding to each text, wherein each text includes each sample text and the fused text.

4. The method according to any one of claims 1-3, wherein, Training an entity linking model based on each binary classification result includes: Performing backpropagation and derivation according to each binary classification result to obtain the entity linking model.

5. A method for entity linking, which is applied to a knowledge graph, comprising: Obtaining a mention to be recognized; Determining an entity corresponding to the mention to be recognized according to a pre-trained entity linking model; wherein the entity linking model is trained based on the method according to any one of claims 1-4.

6. The method according to claim 5, wherein, After obtaining the mention to be recognized, the method further includes: Obtaining multiple candidate entities with the highest similarity to the mention to be recognized from a preset knowledge graph, wherein each candidate entity has a candidate image and a candidate text; Moreover, determining an entity corresponding to the mention to be recognized according to a pre-trained entity linking model includes: inputting the mention to be recognized, each candidate image, and each candidate text into the entity linking model to obtain a matching result between each candidate entity and the mention to be recognized, and determining an entity corresponding to the mention to be recognized from each candidate entity according to each matching result; wherein, the matching result is used to characterize the similarity between the mention to be recognized and the candidate entity.

7. The method according to claim 5 or 6, wherein, the mention to be recognized is determined based on a search request, and the entity corresponding to the mention to be recognized is a search result corresponding to the search request; wherein, the search request includes the mention to be recognized; or, the search request includes an image to be recognized and / or text to be recognized.

8. An entity linking device, which is applied to a knowledge graph, comprising: a second obtaining unit, configured to obtain a mention to be recognized; a second determining unit, configured to determine an entity corresponding to the mention to be recognized according to a pre-trained entity linking model; wherein, the entity linking model is trained based on the method according to any one of claims 1-4.

9. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the electronic device can execute the method according to any one of claims 1-4; or, so that the electronic device can execute the method according to any one of claims 5-7.

10. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method according to any one of claims 1-4 is implemented; or, when the computer program is executed by a processor, the method according to any one of claims 5-7 is implemented.

11. A computer program product, comprising a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-4 is implemented; or, when the computer program is executed by a processor, the method according to any one of claims 5-7 is implemented.

Citation Information

Patent Citations

  • Entity recognition method and device, electronic equipment and storage medium

    CN113657100A

  • Information extraction method and device, electronic equipment and storage medium

    CN113806552A