Text-based Entity Linking, Recognition Method, Electronic Device and Storage Medium

By integrating semantic similarity and prior knowledge characteristics in the entity link model, the problem of low entity link accuracy in the prior art is solved, and higher accuracy and generalization ability are achieved.

CN115017325BActive Publication Date: 2025-07-29ALIBABA (CHINA) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210480047.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-05
Publication Date
2025-07-29
Estimated Expiration
2042-05-05

AI Technical Summary

Technical Problem

The existing entity linking algorithm only considers the relevant information of the entity itself in the text, resulting in low accuracy of entity linking.

Method used

By inputting text data into the entity link model, injecting semantic similarity matrix and prior knowledge features into the translation module, fusing the semantic similarity of mention words, entity names and entity attributes, constructing a semantic similarity matrix, and splicing with the prior knowledge features in the knowledge graph to judge entity linkage.

Benefits of technology

Improve the accuracy and generalization ability of entity links, reduce the dependence on training data sets, and can predict entity links more accurately in the absence of training data or obvious tendencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115017325B_ABST
    Figure CN115017325B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a text-based entity linking and recognition method, an electronic device, and a storage medium. The method includes: determining text data as training data; inputting the training data into an entity linking model; processing the training data by a translation module of the entity linking model to obtain corresponding output features; splicing the output features with prior knowledge features to obtain fused features; and making a judgment on entity linking for the fused features to determine the linked entity as the output result of the entity linking model. By integrating prior knowledge in a knowledge graph into the model and comprehensively predicting entity linking, the accuracy of the entity linking model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to a text-based entity linking method, a text recognition method, a speech recognition method, an electronic device, and a storage medium. Background Art

[0002] Entity linking refers to identifying a mention word from natural language and linking it to the corresponding entity. Currently, the entity linking algorithm usually first matches the corresponding mention word from the text, then obtains the candidate entities corresponding to the mention word according to the knowledge graph, and then selects the entity with the highest correlation with the mention word appearing in the text as the linked entity based on the context relationship in the text and the relevant information of the candidate entity itself.

[0003] However, the above entity linking algorithm only considers the relevant information of the entity itself in the text, resulting in some entity linkings being ambiguous and having low accuracy. Summary of the Invention

[0004] Embodiments of this application provide a text-based entity linking method to improve processing efficiency.

[0005] Correspondingly, embodiments of this application also provide a text recognition method, a speech recognition method, an electronic device, and a storage medium to ensure the implementation and application of the above system.

[0006] To solve the above problems, embodiments of this application disclose a text-based entity linking method, and the method includes:

[0007] Determine text data as training data;

[0008] Input the training data into an entity linking model;

[0009] In a translation module of the entity linking model, process the training data to obtain corresponding output features;

[0010] Concatenate the output features with prior knowledge features to obtain fused features;

[0011] Perform entity linking judgment on the fused features, and determine the linked entity as the output result of the entity linking model.

[0012] Optionally, the determining text data as training data includes:

[0013] Concatenate the mention word, the context text containing the mention word, and the entity description information of the candidate entity as sample data;

[0014] Use multiple sample data to form training data.

[0015] Optionally, when the translation module of the entity linking model processes the training data to obtain corresponding output features, it includes:

[0016] During the process of the translation module of the entity linking model processing the corresponding features of the training data, inject a semantic similarity matrix and process it to obtain corresponding output features.

[0017] Optionally, it further includes:

[0018] Obtain a set of words, where the types of words in the set of words at least include: mention words, entity names, and entity attributes;

[0019] Determine the semantic similarity of the words in the set of words and construct a semantic similarity matrix of the words.

[0020] Optionally, when the translation module of the entity linking model processes the corresponding features of the training data, inject a semantic similarity matrix and process it to obtain corresponding output features, including:

[0021] Input the input vector corresponding to the sample data into the translation module of the entity linking model;

[0022] Fuse the output vector of the specified intermediate component with the semantic similarity matrix;

[0023] Input the fused vector into the next component for processing to obtain the output features of the translation module.

[0024] Optionally, splicing the output features with prior knowledge features to obtain fused features, including:

[0025] Obtain the probability feature of the mention word linking to the entity and the popularity feature of the candidate entity;

[0026] Normalize the probability feature of the mention word linking to the entity and the popularity feature of the candidate entity respectively to obtain the normalized probability feature and the normalized popularity feature;

[0027] Splice the output features with the normalized probability feature and the normalized popularity feature to obtain corresponding fused features.

[0028] Optionally, inputting the training data into the entity linking model includes:

[0029] Determine the text features of the sample data in the training data and input the text features into the entity linking model.

[0030] An embodiment of the present application also discloses a text recognition method, and the method includes:

[0031] Obtain the text data to be recognized;

[0032] Input the text data into the entity linking model;

[0033] Process the text data in the translation module of the entity linking model to obtain corresponding output features;

[0034] Concatenate the output features with the prior knowledge features to obtain fused features;

[0035] Judge entity linking for the fused features, and determine the linked entity as the output result of the entity linking model;

[0036] Use the linked entity as a keyword and perform recognition processing based on the keyword.

[0037] Optionally, the processing the text data in the translation module of the entity linking model to obtain corresponding output features includes:

[0038] Inject a semantic similarity matrix and perform processing during the process of processing the corresponding features of the text data in the translation module of the entity linking model to obtain corresponding output features.

[0039] An embodiment of the present application also discloses a speech recognition method, and the method includes:

[0040] Obtain the speech data to be recognized, and recognize the corresponding text data through the speech data;

[0041] Input the text data into the entity linking model;

[0042] Process the text data in the translation module of the entity linking model to obtain corresponding output features;

[0043] Concatenate the output features with the prior knowledge features to obtain fused features;

[0044] Judge entity linking for the fused features, and determine the linked entity as the output result of the entity linking model;

[0045] Use the linked entity as a keyword and perform recognition processing based on the keyword.

[0046] Optionally, the processing the text data in the translation module of the entity linking model to obtain corresponding output features includes:

[0047] Inject a semantic similarity matrix and perform processing during the process of processing the corresponding features of the text data in the translation module of the entity linking model to obtain corresponding output features.

[0048] An embodiment of the present application also discloses a method for identifying multimodal data, the method comprising:

[0049] Obtain multimodal data, the multimodal data including at least two of the following: text data, audio data, image data, and video data;

[0050] Identify the multimodal data to determine the corresponding text data;

[0051] Input the text data into an entity linking model;

[0052] Process the text data in a translation module of the entity linking model to obtain corresponding output features;

[0053] Concatenate the output features with prior knowledge features and multimodal features of the multimodal data to obtain fused features;

[0054] Judge entity linking for the fused features to determine the linked entity as the output result of the entity linking model;

[0055] Use the linked entity as a keyword and perform corresponding processing based on the keyword.

[0056] An embodiment of the present application also discloses an electronic device, comprising: a processor; and a memory storing executable code thereon, which when executed by the processor, executes the method as described in the embodiment of the present application.

[0057] An embodiment of the present application also discloses one or more machine-readable media storing executable code thereon, which when executed by the processor, executes the method as described in the embodiment of the present application.

[0058] Compared with the prior art, the embodiment of the present application has the following advantages:

[0059] In the embodiment of the present application, text data can be determined as training data and input into an entity linking model. Then, the training data is processed in a translation module of the entity linking model to obtain corresponding output features. The output features are concatenated with prior knowledge features to obtain fused features, integrating prior knowledge in the knowledge graph into the model and comprehensively predicting entity linking. Judge entity linking for the fused features to determine the linked entity as the output result of the entity linking model, which can improve the accuracy of the entity linking model. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a schematic diagram of an entity linking example of an embodiment of the present application;

[0061] Figure 2 It is a flowchart of the steps of an embodiment of a text-based entity linking method of the present application;

[0062] Figure 3 It is a schematic diagram of an example of a translation module of an embodiment of the present application;

[0063] Figure 4 It is a flowchart of the steps of another embodiment of a text-based entity linking method of the present application;

[0064] Figure 5 It is a schematic diagram of an example of a BERT model of an embodiment of the present application;

[0065] Figure 6 It is a flowchart of the steps of an embodiment of a text recognition method of the present application;

[0066] Figure 7 It is a flowchart of the steps of an embodiment of a speech recognition method of the present application;

[0067] Figure 8 It is a flowchart of the steps of an embodiment of a recognition method for multi-modal data of the present application;

[0068] Figure 9 It is a schematic diagram of the structure of an exemplary device provided by an embodiment of the present application. Detailed implementation manners

[0069] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0070] Embodiments of the present application can be applied to various scenarios related to entity linking. For example, in scenarios such as text recognition and speech recognition, corresponding content is linked to entities for subsequent processing. Among them, entity linking refers to identifying mention words from natural language and linking them to corresponding entities. A mention word refers to a language fragment in natural language that expresses an entity. An entity refers to a single and unique semantics or object in the real world, such as a single person, a single product, a single organization, etc.

[0071] The embodiments of the present application can effectively integrate the prior knowledge of the knowledge graph in the entity linking model, thereby reducing the ambiguity of entity linking and improving the accuracy. Among them, in the knowledge graph, there is a lot of prior knowledge, such as the link probability of the mention word to the candidate entity, the popularity of the entity, the mutual relationship between entities, the semantic similarity of the mention word, entity name, etc. as words. These prior knowledge are crucial for entity linking. For example, for the word "White Crane", there is an 80% probability that it refers to a crane species called the White Crane, and a 20% probability that it refers to the football player White Crane. If such prior knowledge is added to the entity linking model, it can effectively enhance the accuracy of entity linking and reduce the dependence on the training data set.

[0072] The embodiments of the present application can be applied to various entity linking algorithm models, such as applying to the training of the entity linking model in the BERT (Bidirectional Encoder Representation from Transformers) model.

[0073] Such as Figure 1 The schematic diagram of an entity linking example shown. The embodiments of the present application can improve the entity linking model. By injecting the knowledge of word similarity in the Transformer module, when the entity linking model extracts the semantic features of the training data text, it can increase the learning of the word semantic similarity features, enabling the entire model to comprehensively consider the text features where the mention word is located, the entity itself features, and the semantic similarity features between the mention word and the entity name, and make more accurate feature extraction and correlation prediction. And after the entity linking model extracts the text features and word similarity features, by combining the probability features of the mention word linking to the entity, the popularity features of the entity, etc., the prior knowledge of the relationship between the mention word and the entity in the knowledge graph is integrated into the model, and the entity linking prediction is comprehensively performed, which is particularly significant for some mention words and entities lacking training data or entity linking predictions with obvious tendencies.

[0074] Referring to Figure 2 , the flowchart of the steps of an embodiment of the text-based entity linking method of the present application is shown.

[0075] Step 202, determine the text data as the training data.

[0076] Embodiments of the present application can pre-acquire text data to construct training data, and the text data can be obtained from various websites, such as e-commerce websites, social websites, news websites, science popularization websites, etc. Among them, the data of each website can be obtained first, and then sample data can be constructed. Among them, the data types of entities generally include mention words, entity names, and entity attributes. Among them, entity attributes refer to the data describing entities. Therefore, mention words, the context text containing the mention words, and the entity description information of candidate entities can be analyzed from the basic data. Among them, the context text containing the mention words can be a sentence containing the mention words, such as a sentence extracted from the text obtained from the website. The entity description information of the candidate entity can be the description data composed of entity attributes. Therefore, the entity description information includes one or more entity attributes. In the embodiments of the present application, the mention words, the context text containing the mention words, and the entity description information of the candidate entity can be concatenated into sample data, and multiple sample data are used to construct training data.

[0077] For example, in one example, the mention word is "White Crane", the context text containing the mention word is "Wild White Crane... protected animal", and the corresponding candidate entities are a species of crane called White Crane and the football player Bai He. The entity description information of the crane species White Crane is "White Crane is a large wading bird...", and the entity description information of the football player Bai He is "Bai He, a Chinese football player, was born in...". In the context "Wild White Crane... protected animal" containing the mention word, "White Crane" refers to the crane species White Crane. Therefore, when its entity description information is concatenated with the mention word and the context text containing the mention word, the sample 1 "[CLS] White Crane|Wild White Crane... protected animal[SEP] White Crane is a large wading bird..." is a correct match, and this sample is a positive sample; when the entity description information of the football player Bai He is concatenated with the mention word and the context text containing the mention word, the sample 2 "[CLS] White Crane|Wild White Crane... protected animal[SEP] Bai He, a Chinese football player, was born in..." is an incorrect match, and this sample is a negative sample. Among them, the [CLS] flag is the sentence start flag, placed at the beginning of the first sentence, and the representation vector C obtained by processing through an entity linking model such as BERT can be used for subsequent classification tasks. The [SEP] flag is used to separate two input sentences. For example, for input sentences A and B, the [SEP] flag needs to be added after sentence A and before sentence B for separation.

[0078] Step 204, input the training data into the entity linking model.

[0079] Each sample data in the training data can be sequentially input into the entity linking model for processing. Among them, inputting the training data into the entity linking model includes: determining the text features of the sample data in the training data, and inputting the text features into the entity linking model. The text features of the sample data can be determined, for example, by converting each text into a feature vector and inputting the converted feature vector into the entity linking model.

[0080] Step 206, process the training data in the translation module of the entity linking model to obtain corresponding output features.

[0081] The entity linking model includes a translation module Transformer, and a structural example is shown as Figure 3 shown. In the embodiments of the present application, knowledge of word similarity can be injected into the translation (Transformer) module, enabling the entity linking model to increase the learning of word semantic similarity features when extracting the text semantic features of the training data. The process of processing the training data in the translation module of the entity linking model to obtain corresponding output features includes: injecting a semantic similarity matrix and performing processing during the process of processing the corresponding features of the training data in the translation module of the entity linking model to obtain corresponding output features.

[0082] Among them, since a mentioned word may have multiple semantics, the link from the mentioned word to the entity needs to be determined based on the semantics of the mentioned word to identify the entity it refers to, achieving entity disambiguation. Therefore, the similarity between the mentioned word and the entity name can provide reference features for entity disambiguation. Thus, semantic similarity is very important for the link from the mentioned word to the entity. Both the mentioned word and the entity name are individual words, and there is an inherent semantic relationship between the words themselves, that is, the semantic similarity of the words. The semantic similarity of words can represent the semantic similarity relationship between words. Among them, if a word has multiple meanings, its semantic similarity with other words is a weighted average of the semantic similarities of different meanings. Therefore, a semantic similarity matrix can be injected during the process of processing the corresponding features of the training data in the translation module to increase the learning of word semantic similarity features, enabling the entire model to comprehensively consider the text features where the mentioned word is located, the features of the entity itself, and the semantic similarity features between the mentioned word and the entity name, and make more accurate feature extraction and correlation prediction.

[0083] In an alternative embodiment, during the process of the translation module of the entity linking model processing the corresponding features of the training data, a semantic similarity matrix is injected and processed to obtain corresponding output features, including: inputting the input vector corresponding to the sample data into the translation module of the entity linking model; fusing the output vector of the specified intermediate component with the semantic similarity matrix; inputting the fused vector into the next component for processing to obtain the output features of the translation module. The translation module may include multiple components. The specified intermediate component can be determined, the output vector of the specified intermediate component is fused with the semantic similarity matrix, and then the fused vector is input into the next component for further processing, and finally the output features of the translation module are obtained. A set of words can be obtained. The types of words in the set of words at least include: mention words, entity names, and entity attributes; the semantic similarity of the words in the set of words is determined, and a semantic similarity matrix of the words is constructed. The set of words can be determined based on the knowledge graph. Various types of words such as mention words, entity names, and entity attributes are obtained from the knowledge graph, and then the semantic similarity between any two words can be calculated. The semantic similarity can be calculated in various ways, such as calculating the semantic similarity between words according to word libraries such as WordNet (Word Network) and Chinese Open Wordnet, to obtain the similarity matrix of the words.

[0084] As Figure 3 In the example shown, when the input vector corresponding to the sample data is used as the query vector (Query, Q), the Transformer module also corresponds to key-value pairs, namely the key (Key, K) and the value (Value, V). The Transformer module includes a first component (MatMul), a second component (Scale), a third component (Softmax), and a fourth component (MatMul). Inputting the query vector Q and the key vector K into the first component MatMul, the query vector Q and the key K can be multiplied in matrix, and then multiplied by a constant for Scale. Taking the second component Scale as the specified intermediate component, multiplying the output of the second component Scale by the semantic similarity matrix makes the Transformer module pay more attention to the feature information between similar words. Then, it is input into the third component for Softmax operation and multiplied by V in matrix to obtain the output features.

[0085] Among them, the formula for injecting the semantic similarity matrix of words into the Transformer module is as follows:

[0086]

[0087] Among them, S is the semantic similarity matrix of words, Q, K, and V are the parameters Query, Key, and Value of the Transformer module, dk is the normalization parameter, and Attention(Q, K, V) is the attention result of the Transformer module. Through the above method, the semantic similarity information between all words in the vocabulary can be injected into the Transformer module, enabling the Transformer module to simultaneously focus on the semantic features of the input text and the word similarity features, realizing data-driven based on the training text and knowledge-driven based on word similarity, improving the utilization effect of the data features and knowledge features of the entire model, and optimizing the algorithm model for entity linking.

[0088] The entity linking model of the embodiment of the present application can be based on the attention mechanism. For example, the multi-head attention mechanism can be adopted. Therefore, it can have multiple Transformer modules, and the output features of each Transformer module can be spliced together for subsequent operations.

[0089] Step 208: Splice the output features with the prior knowledge features to obtain the fusion features.

[0090] The entity linking model of the embodiment of the present application can be based on the attention mechanism. For example, the multi-head attention mechanism can be adopted. Therefore, it can have multiple Transformer modules, and the output features of each Transformer module can be spliced together for subsequent operations. Among them, the output features of each Transformer module can be spliced to obtain the spliced output features.

[0091] In the embodiment of the present application, the prior knowledge features are the features obtained from the prior knowledge, which may include: the probability features of the mention word linking to the entity, the popularity features of the entity, etc., and may also include various prior knowledge features such as the edit distance features. Among them, the probability feature of the mention word linking to the entity refers to the linking probability of the mention word to all candidate entities. For example, for the mention word "White Crane", there is an 80% probability that it refers to the white crane of the genus Grus, and a 20% probability that it refers to the football player White Crane. Then the probability feature of the mention word linking to the entity is [0.8, 0.2]. The popularity feature of the entity refers to the number of times the entity appears in a data set or a scenario. For example, the number of times the entity white crane of the genus Grus appears in the task scenario data of this entity linking is 200, then the popularity feature of this entity is 200. Since different prior knowledge features have different units, the prior knowledge features can also be normalized.

[0092] In an alternative embodiment, the output feature is concatenated with the prior knowledge feature to obtain a fused feature, including: obtaining the probability feature of the mention word linking to an entity and the popularity feature of the candidate entity; normalizing the probability feature of the mention word linking to an entity and the popularity feature of the candidate entity respectively to obtain a normalized probability feature and a normalized popularity feature; concatenating the output feature with the normalized probability feature and the normalized popularity feature to obtain the corresponding fused feature. The probability feature of the mention word linking to an entity and the popularity feature of the candidate entity can be obtained, then zero-padding is performed on the probability feature of the mention word linking to an entity to reach a specified matrix size, and then normalization is performed to obtain the normalized probability feature. Zero-padding is performed on the popularity feature of the candidate entity to reach a specified matrix size, and then normalization is performed to obtain the normalized popularity feature. Among them, zero-padding is to fill 0s after the data with insufficient length in the matrix so that the matrix size reaches a fixed size for convenient calculation.

[0093] The output features are concatenated as the output feature of the multi-head attention mechanism, and then this output feature is concatenated with the normalized probability feature and the normalized popularity feature, that is, the data-driven text semantic feature and the knowledge-driven knowledge graph prior knowledge feature are concatenated together to obtain the corresponding fused feature.

[0094] Step 210, determine whether the fused feature is an entity link, and determine the linked entity as the output result of the entity link model.

[0095] It is possible to determine whether the fused feature is an entity link. For example, the fused feature is input into the Softmax layer to obtain a comprehensive judgment score result. This score result can take values between [0,1]. Based on the score result, the linked entity is judged among the candidate entities. For example, the candidate entity with the highest score can be used as the linked entity to obtain the output result of this entity link model.

[0096] Subsequently, the loss function can be determined based on the output result and the label data corresponding to the sample data, etc. The parameters of the entity link model are adjusted based on the loss function and iteratively processed until an entity link model that meets the conditions is obtained, which can be applied to subsequent text recognition scenarios. For example, according to the loss function, the parameters in the model can be adjusted by methods such as SGD (Stochastic Gradient Descent) or Adam (Adaptive Momentum Estimation) so that the loss function meets certain conditions, and the parameters after the model converges are the model parameters obtained by training.

[0097] In summary, it can be determined that the text data is used as training data and input into the entity linking model. Then, the training data is processed by the translation module of the entity linking model to obtain corresponding output features. The output features are concatenated with the prior knowledge features to obtain fused features, and the prior knowledge in the knowledge graph is fused into the model. The entity linking prediction is comprehensively performed, and the fused features are judged for entity linking to determine the linked entity as the output result of the entity linking model. For the mention words lacking data in the training set, accurate entity linking can also be performed, reducing the excessive dependence on the data set, which is important for improving the accuracy and generalization of the entity linking algorithm.

[0098] The embodiments of the present application can be applied to various scenarios for entity linking, such as text recognition, speech recognition, etc. For example, when applied to scenarios such as customer service, corresponding entities can be recognized based on the text, so as to provide services related to the entities. For example, if the user's text or speech includes "apple", based on the processing of this entity linking model, it can be determined whether the user's "apple" refers to the fruit apple or the electronic product brand, so as to provide corresponding services. It can also be applied to speech recognition scenarios such as smart speakers and TV boxes to assist in recognizing the entities corresponding to the words spoken by the user and provide better services to the user.

[0099] Based on the above embodiments, the embodiments of the present application also provide a text-based entity linking method, which can add prior knowledge such as word semantic similarity, the probability feature of mention word linking entity, and entity popularity, and jointly train an entity linking model that integrates prior knowledge and the text semantic features of training data, improving the accuracy of the entity linking algorithm. As Figure 4 shown:

[0100] Step 402, concatenate the mention word, the context text containing the mention word, and the entity description information of the candidate entity as sample data; multiple sample data are used to form training data.

[0101] Taking the training of the entity linking model based on the BERT model as an example. The mention word, the context text containing the mention word, and the entity description information of the candidate entity can be obtained from various word banks, knowledge graphs, websites, etc., and the sample data can be concatenated.

[0102] For example, the mentioned word is "White Crane", the context text containing the mentioned word is "Wild White Crane... protected animal", and the corresponding candidate entities are a species of crane called the White Crane and the football player Bai He. The entity description information of the White Crane of the genus Grus is "The White Crane is a large wading bird...", and the entity description information of the football player Bai He is "Bai He, a Chinese football player, was born in...". In the context "Wild White Crane... protected animal" containing the mentioned word, "White Crane" refers to the White Crane of the genus Grus. Therefore, when splicing its entity description information with the mentioned word and the context containing the mentioned word, sample 1 "[CLS] White Crane | Wild White Crane... protected animal [SEP] The White Crane is a large wading bird..." is a correct match, and this sample is a positive sample; when splicing the entity description information of the football player Bai He with the mentioned word and the context containing the mentioned word, sample 2 "[CLS] White Crane | Wild White Crane... protected animal [SEP] Bai He, a Chinese football player, was born in..." is an incorrect match, and this sample is a negative sample.

[0103] Step 404, determine the text features of the sample data in the training data, and input the text features into the entity linking model.

[0104] The text features can be determined and transformed based on the sample data to obtain corresponding text vectors, and the text vectors are input into the entity linking model.

[0105] As Figure 5 In the shown example, the text vector of the sample "[CLS] White Crane | Wild White Crane... protected animal [SEP] The White Crane is a large wading bird..." is determined and input into the entity linking model. Figure 5 It is an example based on the BERT model, and some structures in the model are omitted, such as the Transformer module, etc.

[0106] Step 406, obtain a set of words, and the types of words in the set of words at least include: the mentioned word, the entity name, and the entity attribute.

[0107] Step 408, determine the semantic similarity of the words in the set of words, and construct a semantic similarity matrix of the words.

[0108] The mentioned word, the entity name, the entity attribute and other words can be obtained from various word banks, knowledge graphs, websites, etc. Then, the semantic similarity between any two words can be calculated. The semantic similarity can be calculated in various ways, such as calculating the semantic similarity between words according to word banks such as WordNet and Chinese Open Wordnet to obtain the similarity matrix of the words.

[0109] Step 410, during the process of processing the corresponding features of the training data by the translation module of the entity linking model, inject the semantic similarity matrix and perform processing to obtain the corresponding output features.

[0110] Among them, during the process of processing the corresponding features of the training data by the translation module of the entity linking model, injecting the semantic similarity matrix and performing processing to obtain the corresponding output features includes: inputting the input vector corresponding to the sample data into the translation module of the entity linking model; fusing the output vector of the specified intermediate component with the semantic similarity matrix; inputting the fused vector into the next component for processing to obtain the output features of the translation module. When taking the input vector corresponding to the sample data as the query vector (Query, Q), this Transformer module also corresponds to key-value pairs (Key-Value), that is, the keyword (Key, K) and the key value (Value, V). This Transformer module includes a first component (MatMul), a second component (Scale), a third component (Softmax), and a fourth component (MatMul). Inputting the query vector Q and the keyword vector K into the first component MatMul can perform matrix multiplication on the query vector Q and the keyword K, and then multiply by a constant for Scale. Taking the second component Scale as the specified intermediate component, multiplying the output of the second component Scale by the semantic similarity matrix makes the Transformer module pay more attention to the feature information between similar words. Then, input it into the third component for Softmax operation and perform matrix multiplication with V to obtain the output features.

[0111] Step 412, obtain the probability feature of the mention word linking to the entity and the popularity feature of the candidate entity.

[0112] Step 414, perform normalization processing on the probability feature of the mention word linking to the entity and the popularity feature of the candidate entity respectively to obtain the normalized probability feature and the normalized popularity feature.

[0113] Step 416, splice the output features with the normalized probability feature and the normalized popularity feature to obtain the corresponding fused features.

[0114] In the embodiments of the present application, prior knowledge features can be extended, such as more prior knowledge features like edit distance. After operations such as normalization, they can continue to be spliced at this position, and then an overall feature vector that combines the text semantic features output at the [CLS] position and various prior knowledge features in the knowledge graph can be obtained. Then, entity linking judgment is made based on this feature vector.

[0115] Step 418: Determine entity linking for the fusion features, and identify the linked entity as the output result of the entity linking model.

[0116] Subsequently, a loss function can be determined based on the output result and the label data corresponding to the sample data. The parameters of the entity linking model can be adjusted based on the loss function, and iterative processing can be performed until an entity linking model that meets the conditions is obtained, which can be applied to subsequent text recognition scenarios.

[0117] In the embodiments of the present application, on the basis of training the entity linking model using BERT, prior knowledge features of the knowledge graph such as word semantic similarity, the probability feature of the mention word linking to the entity, and entity popularity features are fully integrated. The advantages of data-driven and knowledge-driven are comprehensively utilized, which is of great significance for improving the utilization rate of knowledge graph features and thus improving the effect of entity linking.

[0118] Based on the above embodiments, the embodiments of the present application also provide a text recognition method, as Figure 6 shown:

[0119] Step 602: Obtain the text data to be recognized.

[0120] The text data to be recognized can be obtained. Among them, the sample data to be recognized can be determined based on the application scenario. For example, in the customer service scenario, the text data to be recognized can be the text information sent by the user. In the translation scenario, the text data to be recognized can be the text data to be translated, etc.

[0121] Step 604: Input the text data into the entity linking model.

[0122] The inputting of the training data into the entity linking model includes: determining the text features of the sample data in the training data, and inputting the text features into the entity linking model.

[0123] Step 606: Process the text data in the translation module of the entity linking model to obtain corresponding output features.

[0124] The processing of the training data in the translation module of the entity linking model to obtain corresponding output features includes: injecting a semantic similarity matrix and performing processing during the processing of the corresponding features of the training data in the translation module of the entity linking model to obtain corresponding output features.

[0125] Among them, a set of words can be obtained, and the types of words in the set of words at least include: mention words, entity names, and entity attributes; determine the semantic similarity of the words in the set of words, and construct a semantic similarity matrix of the words. During the process of the translation module of the entity linking model processing the corresponding features of the training data, inject the semantic similarity matrix and process it to obtain the corresponding output features, including: input the input vector corresponding to the sample data into the translation module of the entity linking model; fuse the output vector of the specified intermediate component with the semantic similarity matrix; input the fused vector into the next component for processing to obtain the output features of the translation module.

[0126] Step 608, splice the output features with the prior knowledge features to obtain fused features.

[0127] Among them, splicing the output features with the prior knowledge features to obtain fused features includes: obtaining the probability features of the mention word linking to the entity and the popularity features of the candidate entity; respectively performing normalization processing on the probability features of the mention word linking to the entity and the popularity features of the candidate entity to obtain normalized probability features and normalized popularity features; splice the output features with the normalized probability features and the normalized popularity features to obtain the corresponding fused features.

[0128] Step 610, perform entity linking judgment on the fused features, and determine the linked entity as the output result of the entity linking model.

[0129] Step 612, use the linked entity as a keyword and perform recognition processing based on the keyword.

[0130] Based on the entity linking model, the entity corresponding to the link of this text can be determined, and then the entity name of the entity can be used as a keyword to process the keyword. For example, in the customer service scenario, the recognized entity name can be used as a keyword to perform corresponding customer service processing, such as matching the corresponding service template, providing corresponding customer service scripts, etc. Another example is that in the translation scenario, based on the recognized entity name as a keyword, the corresponding related field can be determined, thereby assisting subsequent translation processing.

[0131] Based on the above embodiments, an embodiment of the present application further provides a speech recognition method, as Figure 7 shown:

[0132] Step 702, obtain the speech data to be recognized, and recognize the corresponding text data through the speech data.

[0133] The speech data to be recognized can be obtained. The speech data to be recognized can be determined based on the application scenario. For example, in the customer service scenario, the speech data to be recognized can be the conversation speech between the user and the customer service. Another example is in the speech recognition scenarios based on devices such as speakers and set-top boxes, the speech data to be recognized can be the speech commands issued by the user.

[0134] First, speech recognition can be performed on the speech data to obtain the corresponding text data.

[0135] Step 704: Input the text data into the entity linking model.

[0136] The inputting the training data into the entity linking model includes: determining the text features of the sample data in the training data and inputting the text features into the entity linking model.

[0137] Step 706: Process the text data in the translation module of the entity linking model to obtain the corresponding output features.

[0138] The processing the training data in the translation module of the entity linking model to obtain the corresponding output features includes: injecting and processing the semantic similarity matrix during the process of processing the corresponding features of the training data in the translation module of the entity linking model to obtain the corresponding output features.

[0139] Among them, a set of words can be obtained. The types of words in the set of words at least include: mention words, entity names, and entity attributes. Determine the semantic similarity of the words in the set of words and construct a semantic similarity matrix of the words. The processing the training data in the translation module of the entity linking model to obtain the corresponding output features includes: inputting the input vector corresponding to the sample data into the translation module of the entity linking model; fusing the output vector of the specified intermediate component with the semantic similarity matrix; inputting the fused vector into the next component for processing to obtain the output features of the translation module.

[0140] Step 708: Concatenate the output features with the prior knowledge features to obtain fused features.

[0141] Among them, the concatenating the output features with the prior knowledge features to obtain fused features includes: obtaining the probability features of the mention word linking to the entity and the popularity features of the candidate entities; respectively normalizing the probability features of the mention word linking to the entity and the popularity features of the candidate entities to obtain the normalized probability features and the normalized popularity features; concatenating the output features with the normalized probability features and the normalized popularity features to obtain the corresponding fused features.

[0142] Step 710: Determine whether to perform entity linking on the fused feature, and identify the linked entity as the output result of the entity linking model.

[0143] Step 712: Use the linked entity as a keyword and perform recognition processing based on the keyword.

[0144] Based on the entity linking model, the entity linked to the text corresponding to the voice can be determined. Then, the entity name of the entity can be used as a keyword to process the keyword. For example, in a customer service scenario, the identified entity name can be used as a keyword to perform corresponding customer service processing, such as matching the corresponding service template and providing corresponding customer service scripts. Another example is in the voice recognition scenario of a device. Based on the identified entity name as a keyword or wake-up word, corresponding voice processing is performed.

[0145] Based on the above embodiments, an embodiment of the present application further provides a method for recognizing multi-modal data, which can recognize multi-modal data, organically combine text, audio, image, video data, etc., link them to corresponding entities, and thus perform required processing based on the entity.

[0146] Refer to Figure 8 , which shows a flowchart of the steps of an embodiment of the method for recognizing multi-modal data of the present application.

[0147] Step 802: Obtain multi-modal data.

[0148] Step 804: Recognize the multi-modal data to determine the corresponding text data.

[0149] Among them, the multi-modal data includes at least two of the following: text data, audio data, image data, and video data. Among them, for audio data, the text data can be determined through speech recognition. For image data, if it contains text, the text data can be recognized through Optical Character Recognition (OCR). For video data, which is composed of an audio stream and an image stream, for the audio data in the audio stream, the text data can be obtained through speech recognition, and for the image data in the image stream, the text data can be recognized through OCR.

[0150] In addition, for image data, it can also be processed through an image model to obtain the feature vector of the image. Similarly, for the image stream in video data, the feature vector is obtained through image model processing. The feature vector is used for subsequent vector splicing and fusion.

[0151] Step 808: Input the text data into the entity linking model.

[0152] Inputting the training data into the entity linking model includes: determining the text features of the sample data in the training data and inputting the text features into the entity linking model.

[0153] Step 810, processing the text data by the translation module of the entity linking model to obtain corresponding output features.

[0154] Processing the training data by the translation module of the entity linking model to obtain corresponding output features includes: injecting and processing a semantic similarity matrix during the process of processing the corresponding features of the training data by the translation module of the entity linking model to obtain corresponding output features.

[0155] Among them, a set of words can be obtained, and the types of words in the set of words at least include: mention words, entity names, and entity attributes; determining the semantic similarity of the words in the set of words and constructing a semantic similarity matrix of the words. Processing the training data by the translation module of the entity linking model to obtain corresponding output features includes: inputting the input vector corresponding to the sample data into the translation module of the entity linking model; fusing the output vector of the specified intermediate component with the semantic similarity matrix; inputting the fused vector into the next component for processing to obtain the output features of the translation module.

[0156] Step 812, concatenating the output features with prior knowledge features and multi-modal features of multi-modal data to obtain fused features.

[0157] Among them, concatenating the output features with prior knowledge features and multi-modal features of multi-modal data to obtain fused features includes: obtaining the probability features of mention word linked entities, the popularity features of candidate entities, and the multi-modal features of multi-modal data; respectively normalizing the probability features of mention word linked entities, the popularity features of candidate entities, and the multi-modal features of multi-modal data to obtain normalized probability features, normalized popularity features, and normalized multi-modal features; concatenating the output features with the normalized probability features, normalized popularity features, and normalized multi-modal features to obtain corresponding fused features.

[0158] Step 814, judging entity linking for the fused features and determining the linked entity as the output result of the entity linking model.

[0159] Step 816, using the linked entity as a keyword and performing corresponding processing according to the keyword.

[0160] Therefore, it can be applied to the processing scenarios of multi-modal data, such as the recognition of multi-modal data on websites. For example, in a video website, entities are recognized based on videos, titles, etc. and corresponding processing is performed, such as auditing.

[0161] In the existing entity linking algorithms, they simply rely on the text semantic features in the training dataset, which is a data-driven algorithm. They cannot effectively integrate and absorb prior knowledge such as the linking probability of a mention word pair to a candidate entity, that is, they cannot make full use of the effective information in the knowledge graph. Compared with the existing entity linking algorithms, the entity linking algorithm model that fuses prior knowledge in the embodiments of the present application can make full use of the prior knowledge in the knowledge graph for mention words and entities, and can train these prior knowledge together with the text data features to obtain an algorithm that combines knowledge-driven and data-driven, which can greatly improve the effect of entity linking and reduce the dependence on the training data, especially for few-shot data and data with uneven distributions.

[0162] The embodiments of the present application can fuse semantic similarity features into the entity linking model based on BERT. A basic word library can be constructed according to all the words involved in the knowledge graph (including mention words, entity names, entity attributes, etc.), and then the semantic similarity between different words is calculated through WordNet, Chinese Open Wordnet, etc. to construct a semantic similarity matrix of words. Then, this semantic similarity matrix is injected into the Transformer module of the BERT model, enabling it to simultaneously pay attention to the semantic information of the input text and the similarity information between words during the training process, and improving the feature extraction ability for synonyms in the knowledge graph. Since there are a large number of related relationships of synonyms between mention words and entity names in the knowledge graph, this method is of great significance for enhancing the extraction of word semantic features by the entity linking model and improving the effect of the entity linking model.

[0163] The vector features of the input text are obtained from the [CLS] position. Then, after performing operations such as zero-padding and normalization on other prior knowledge features (such as the link probability feature of the mentioned word pair to the entity, the entity popularity feature, etc.), they are concatenated with the text semantic features output from the [CLS] position to obtain a feature vector that combines prior knowledge features and text semantic features. Then, entity linking is judged. This part can also be extended according to requirements. If there are other prior knowledge features, such as the edit distance feature, etc., after performing zero-padding and normalization operations, they can be continuously concatenated at this position to expand the feature vector. Finally, entity linking is judged based on all the text semantic features and prior knowledge features. It can make full use of the prior knowledge features in the knowledge graph, greatly increasing the extraction ability of training data and knowledge graph features, and can effectively improve the accuracy of the entity linking model, as well as its robustness and generalization ability.

[0164] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present application.

[0165] This embodiment also provides a text-based entity linking device, which is applied to an electronic device on the server side.

[0166] The data determination module is used to determine text data as training data;

[0167] The input module is used to input the training data into the entity linking model;

[0168] The translation processing module is used to process the training data in the translation module of the entity linking model to obtain corresponding output features;

[0169] The knowledge fusion module is used to concatenate the output features with prior knowledge features to obtain fusion features;

[0170] The output module is used to judge entity linking for the fusion features and determine the linked entity as the output result of the entity linking model.

[0171] In summary, it is possible to determine the text data as training data, input it into the entity linking model, and then process the training data in the translation module of the entity linking model to obtain corresponding output features. The output features are concatenated with the prior knowledge features to obtain fused features, integrating the prior knowledge in the knowledge graph into the model, comprehensively performing entity linking prediction, judging the entity linking of the fused features, determining the linked entity as the output result of the entity linking model, and being able to accurately perform entity linking for the mention words lacking data in the training set, reducing the over-reliance on the data set, which is important for improving the accuracy and generalization of the entity linking algorithm.

[0172] The data determination module is used to concatenate the mention word, the context text containing the mention word, and the entity description information of the candidate entity as sample data; and multiple sample data are used to form training data.

[0173] The translation processing module is used to inject and process the semantic similarity matrix during the process of processing the corresponding features of the training data in the translation module of the entity linking model to obtain corresponding output features.

[0174] It further includes: a similarity determination module, which is used to obtain a set of words, and the types of words in the set of words at least include: mention words, entity names, and entity attributes; determine the semantic similarity of the words in the set of words, and construct a semantic similarity matrix of the words.

[0175] The translation processing module is used to input the input vector corresponding to the sample data into the translation module of the entity linking model; fuse the output vector of the specified intermediate component with the semantic similarity matrix; input the fused vector into the next component for processing to obtain the output features of the translation module.

[0176] The knowledge fusion module is used to obtain the probability feature of the mention word linking to an entity and the popularity feature of the candidate entity; respectively normalize the probability feature of the mention word linking to an entity and the popularity feature of the candidate entity to obtain the normalized probability feature and the normalized popularity feature; concatenate the output features with the normalized probability feature and the normalized popularity feature to obtain corresponding fused features.

[0177] The input module is used to determine the text features of the sample data in the training data and input the text features into the entity linking model.

[0178] In the embodiments of the present application, on the basis of training an entity linking model using BERT, prior knowledge features of a knowledge graph such as word semantic similarity, the probability feature of a mentioned word linking to an entity, and entity popularity features are fully integrated. The advantages of data-driven and knowledge-driven are comprehensively utilized, which is of great significance for improving the utilization rate of knowledge graph features and further improving the effect of entity linking.

[0179] This embodiment also provides an identification device, which is applied to an electronic device on a server side.

[0180] In a text recognition scenario, for example, the identification device includes:

[0181] An acquisition module, configured to acquire text data to be recognized;

[0182] A linking module, configured to input the text data into an entity linking model; in a translation module of the entity linking model, process the text data to obtain corresponding output features; splice the output features with prior knowledge features to obtain a fusion feature; perform entity linking judgment on the fusion feature to determine the linked entity as the output result of the entity linking model;

[0183] An identification processing module, configured to use the linked entity as a keyword and perform identification processing based on the keyword.

[0184] The linking module is configured to inject a semantic similarity matrix and perform processing during the process of the translation module of the entity linking model processing the corresponding features of the text data to obtain corresponding output features.

[0185] In a voice recognition scenario, for example, the identification device includes:

[0186] An acquisition module, configured to acquire voice data to be recognized and recognize corresponding text data through the voice data;

[0187] A linking module, configured to input the text data into an entity linking model; in a translation module of the entity linking model, process the text data to obtain corresponding output features; splice the output features with prior knowledge features to obtain a fusion feature; perform entity linking judgment on the fusion feature to determine the linked entity as the output result of the entity linking model;

[0188] An identification processing module, configured to use the linked entity as a keyword and perform identification processing based on the keyword.

[0189] The linking module is configured to inject a semantic similarity matrix and perform processing during the process of the translation module of the entity linking model processing the corresponding features of the text data to obtain corresponding output features.

[0190] In a recognition scenario of multimodal data, the recognition device includes:

[0191] An acquisition module, configured to acquire multimodal data, where the multimodal data includes at least two of the following: text data, audio data, image data, and video data; recognize the multimodal data to determine corresponding text data;

[0192] A linking module, configured to input the text data into an entity linking model; in the translation module of the entity linking model, process the text data to obtain corresponding output features; splice the output features with prior knowledge features and multimodal features of the multimodal data to obtain fused features; perform entity linking judgment on the fused features to determine the linked entity as the output result of the entity linking model;

[0193] A recognition processing module, configured to use the linked entity as a keyword and perform recognition processing based on the keyword.

[0194] The linking module is configured to inject and process a semantic similarity matrix during the process of processing the corresponding features of the text data by the translation module of the entity linking model to obtain corresponding output features.

[0195] An embodiment of the present application also provides a non-volatile readable storage medium, in which one or more modules (programs) are stored. When the one or more modules are applied to a device, the device can execute instructions (instructions) of each method step in the embodiment of the present application.

[0196] An embodiment of the present application provides one or more machine-readable media, on which instructions are stored. When executed by one or more processors, the instructions cause an electronic device to execute one or more of the methods as described in the above embodiments. In an embodiment of the present application, the electronic device includes devices such as a server and a terminal device.

[0197] Embodiments of the present disclosure can be implemented as a device configured with any suitable hardware, firmware, software, or any combination thereof. The device may include electronic devices such as a server (cluster) and a terminal. Figure 9 Exemplary device 900 that can be used to implement the various embodiments described in the present application is schematically shown.

[0198] For one embodiment, Figure 9An exemplary device 900 is shown, which has one or more processors 902, a control module (chipset) 904 coupled to at least one of the one or more processors 902, a memory 906 coupled to the control module 904, a non-volatile memory (NVM) / storage device 908 coupled to the control module 904, one or more input / output devices 910 coupled to the control module 904, and a network interface 912 coupled to the control module 904.

[0199] The processor 902 may include one or more single-core or multi-core processors, and the processor 902 may include any combination of general-purpose processors or dedicated processors (such as graphics processors, application processors, baseband processors, etc.). In some embodiments, the device 900 can act as devices such as the server, terminal, etc. described in the embodiments of the present application.

[0200] In some embodiments, the device 900 may include one or more computer-readable media (e.g., the memory 906 or the NVM / storage device 908) having instructions 914, and one or more processors 902 combined with the one or more computer-readable media and configured to execute the instructions 914 to implement modules so as to perform the actions described in the present disclosure.

[0201] For one embodiment, the control module 904 may include any suitable interface controller to provide any suitable interface to at least one of the one or more processors 902 and / or any suitable device or component communicating with the control module 904.

[0202] The control module 904 may include a memory controller module to provide an interface to the memory 906. The memory controller module can be a hardware module, a software module, and / or a firmware module.

[0203] The memory 906 can be used, for example, to load and store data and / or instructions 914 for the device 900. For one embodiment, the memory 906 may include any suitable volatile memory, e.g., suitable DRAM. In some embodiments, the memory 906 may include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).

[0204] For one embodiment, the control module 904 may include one or more input / output controllers to provide an interface to the NVM / storage device 908 and the one or more input / output devices 910.

[0205] For example, the NVM / storage device 908 can be used to store data and / or instructions 914. The NVM / storage device 908 can include any suitable non-volatile memory (e.g., flash memory) and / or can include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more compact disc (CD) drives, and / or one or more digital versatile disc (DVD) drives).

[0206] The NVM / storage device 908 can include storage resources that are part of a device installed as device 900, or it can be accessible to the device without necessarily being part of the device. For example, the NVM / storage device 908 can be accessed via a network through (one or more) input / output devices 910.

[0207] (One or more) input / output devices 910 can provide an interface for device 900 to communicate with any other suitable devices. The input / output devices 910 can include communication components, audio components, sensor components, etc. The network interface 912 can provide an interface for device 900 to communicate through one or more networks. Device 900 can wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing a wireless network based on communication standards like WiFi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.

[0208] For one embodiment, at least one of (one or more) processors 902 can be logically encapsulated with one or more controllers of the control module 904 (e.g., a memory controller module). For one embodiment, at least one of (one or more) processors 902 can be logically encapsulated with one or more controllers of the control module 904 to form a system-in-package (SiP). For one embodiment, at least one of (one or more) processors 902 can be logically integrated with one or more controllers of the control module 904 on the same die. For one embodiment, at least one of (one or more) processors 902 can be logically integrated with one or more controllers of the control module 904 on the same die to form a system-on-chip (SoC).

[0209] In various embodiments, the device 900 can be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the device 900 can have more or fewer components and / or a different architecture. For example, in some embodiments, the device 900 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit (ASIC), and a speaker.

[0210] Among them, a main control chip can be used as a processor or a control module in the detection device, sensor data, location information, etc. are stored in a memory or an NVM / storage device, the sensor group can be used as an input / output device, and the communication interface can include a network interface.

[0211] Embodiments of the present application also provide an electronic device, including: a processor; and a memory storing executable code thereon, when the executable code is executed, causing the processor to execute one or more of the methods as in the embodiments of the present application. In the embodiments of the present application, various data can be stored in the memory, such as target files, file-application association data, and other various data, and can also include user behavior data, etc., so as to provide a data basis for various processes.

[0212] Embodiments of the present application also provide one or more machine-readable media storing executable code thereon, when the executable code is executed, causing a processor to execute one or more of the methods as in the embodiments of the present application.

[0213] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0214] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other.

[0215] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate for implementation in the process Figure 1one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks

[0216] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the processes Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks

[0217] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal device, such that a series of operational steps are performed on the computer or other programmable terminal device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one or more processes and / or blocks Figure 1 one or more blocks

[0218] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0219] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0220] The above has introduced in detail a text-based entity linking method, a text recognition method, a speech recognition method, an electronic device, and a storage medium provided by the present application. Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A text-based entity linking method, characterized in that The method includes: Determine text data as training data; Input the training data into an entity linking model; Input the input vector corresponding to the training data into the translation module of the entity linking model; Fuse the output vector of a specified intermediate component with a semantic similarity matrix; Input the fused vector into the next component for processing to obtain the output features of the translation module; Concatenate the output features with prior knowledge features to obtain fused features; Perform entity linking judgment on the fused features to determine the linked entity as the output result of the entity linking model.

2. The method according to claim 1, characterized in that, The determination of text data as training data includes: Concatenate the mention word, the context text containing the mention word, and the entity description information of the candidate entity as sample data; Use multiple sample data to form training data.

3. The method according to claim 1, wherein It also includes: Obtain a set of words, where the types of words in the set of words at least include: mention words, entity names, and entity attributes; Determine the semantic similarity of the words in the set of words and construct a semantic similarity matrix of the words.

4. The method according to claim 1, wherein The concatenation of the output features with prior knowledge features to obtain fused features includes: Obtain the probability features of the mention word linked to an entity and the popularity features of the candidate entity; Normalize the probability features of the mention word linked to an entity and the popularity features of the candidate entity respectively to obtain normalized probability features and normalized popularity features; Concatenate the output features with the normalized probability features and the normalized popularity features to obtain corresponding fused features.

5. The method according to claim 1, wherein The input of the training data into the entity linking model includes: Determine the text features of the sample data in the training data and input the text features into the entity linking model.

6. A text recognition method, characterized in that, The method includes: Obtain text data to be recognized; Input the text data into an entity linking model; Input the input vector corresponding to the text data into the translation module of the entity linking model; Fuse the output vector of a specified intermediate component with a semantic similarity matrix; Input the fused vector into the next component for processing to obtain the output features of the translation module; Concatenate the output features with prior knowledge features to obtain fused features; Perform entity linking judgment on the fused features to determine the linked entity as the output result of the entity linking model; Use the linked entity as a keyword and perform recognition processing based on the keyword.

7. A voice recognition method, characterized in that, The method includes: Obtain speech data to be recognized and recognize the corresponding text data through the speech data; Input the text data into an entity linking model; Input the input vector corresponding to the text data into the translation module of the entity linking model; Fuse the output vector of a specified intermediate component with a semantic similarity matrix; Input the fused vector into the next component for processing to obtain the output features of the translation module; Concatenate the output features with prior knowledge features to obtain fused features; Perform entity linking judgment on the fused features to determine the linked entity as the output result of the entity linking model; Use the linked entity as a keyword and perform recognition processing based on the keyword.

8. A recognition method for multimodal data, characterized in that, The method includes: Obtain multimodal data, where the multimodal data includes at least two of the following: text data, audio data, image data, and video data; Identify the multimodal data to determine the corresponding text data; Input the text data into an entity linking model; Input the input vector corresponding to the text data into the translation module of the entity linking model; Fuse the output vector of the specified intermediate component with the semantic similarity matrix; Input the fused vector into the next component for processing to obtain the output features of the translation module; Concatenate the output features with the prior knowledge features and the multimodal features of the multimodal data to obtain fused features; Judge entity linking for the fused features to determine the linked entity as the output result of the entity linking model; Use the linked entity as a keyword and perform corresponding processing based on the keyword.

9. An electronic device, comprising: A processor; And a memory storing executable code that, when executed by the processor, performs the method according to any one of claims 1-8.

10. One or more machine-readable media storing executable code that, when executed by the processor, performs the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Knowledge question and answer retrieval method and device based on tourism domain knowledge graph

    CN111353030A

  • Knowledge base question and answer entity linking method and system based on similarity

    CN112100356A

  • Text vector representation method and device and electronic equipment

    CN114117062A