Method and System for Training Entity Recognition Model, and Entity Recognition Method and System
Through the K-BERT model combined with knowledge graphs and graph neural networks, efficient identification of explicit and metaphorical entities in text is achieved, and the problem of inability to identify metaphorical entities in the existing technology is solved, and the accuracy and efficiency of naming entity recognition is improved.
Patent Information
- Application Number
- CN202210078338.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-01-24
AI Technical Summary
Existing named entity recognition technology is difficult to identify entities that do not appear in text, especially metaphorical entities, which makes it impossible to accurately identify entities pointed to by user feedback in scenarios such as smart customer service.
A K-BERT-based model is used to perform multi-task learning of sequence annotation and entity matching tasks. Combined with the knowledge graph, entity vectors are generated through graph neural networks, and metaphorical entities are identified using vector search.
It can accurately identify explicit entities of interest and metaphorical entities, improves the comprehensiveness and efficiency of entity recognition, and is suitable for scenarios such as intelligent customer service.
Smart Images

Figure CN114492443B_ABST
Abstract
Description
Technical Field
[0001] This application relates to entity recognition in sentences, and in particular to a method for training an entity recognition model, an entity recognition method, and related systems, devices, and media. Background Art
[0002] Named entity recognition technology has been widely applied. In named entity recognition technology, machine learning models have been used to identify named entities in text (such as sentences). Named entity recognition technology can also be called "proper name recognition", and can identify person names, place names, organization names, proper nouns, etc. in text. Named entity recognition technology can have many applications, such as for text analysis, question and answer dialogues, machine translation, and so on.
[0003] However, in many cases, the text (such as a sentence) may not include the entity itself. For example, for the sentence "The meal I ordered was delivered too long overdue", the entity it refers to may be a certain online food ordering service, but the name of this online food ordering service is not included in the sentence itself. Existing entity recognition solutions in this field have not been able to solve this problem. In fact, the existence of this problem may never have been recognized in the prior art.
[0004] Therefore, there is a need for a solution that can accurately and more comprehensively identify entities in text. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, one or more embodiments of this specification use a K-BERT-based model to perform two tasks: sequence labeling and entity matching, so as to be able to comprehensively and efficiently identify entities in text, including explicitly interested entities and metaphorical entities.
[0006] One or more embodiments of this specification achieve the above object through the following technical solutions.
[0007] In one aspect, there is provided a method for training an entity recognition model, including:
[0008] Constructing a training set, where the training set includes a plurality of training samples d = {S input , X NER , X TER}, where S input is a sentence, X NER is the sequence labeling label of this sentence, X TER is the metaphorical entity label of this sentence, and the metaphorical entity label is used to represent the metaphorical entity of this sentence, where the metaphorical entity is the entity that this sentence actually refers to but does not appear in this sentence;
[0009] Training the entity recognition model using the training set, where the entity recognition model is based on a pre-trained K-BERT model, and training the entity recognition model includes:
[0010] Inputting the training samples in the training set into the entity recognition model to obtain the sequence annotation prediction output and the entity matching prediction output of the sentences in the training samples,
[0011] Determining the sequence annotation loss Loss_sequence of the sentence based on the sequence annotation prediction output of the sentence and the sequence annotation label of the sentence;
[0012] Determining the entity matching loss Loss_match of the sentence at least partially based on the entity matching prediction output of the sentence and the metaphorical entity label of the sentence;
[0013] Determining the total loss Loss_total of the entity recognition model, where the total loss is the weighted sum of the sequence annotation loss and the entity matching loss, i.e., Loss_total = Loss_sequence + α * Loss_match, where α indicates the weight of the entity matching loss; and
[0014] Iteratively performing training to minimize the total loss of the entity recognition model, thereby obtaining the trained entity recognition model.
[0015] Preferably, the knowledge graph associated with the sentence is also input into the entity recognition model.
[0016] Preferably, determining the entity matching loss Loss_match of the sentence includes:
[0017] Using a graph neural network associated with the knowledge graph to generate the metaphorical entity vector of the sentence;
[0018] Determining the vector distance between the entity matching prediction output of the sentence and the metaphorical entity vector; and
[0019] Determining the entity matching loss Loss_match of the sentence, where the entity matching loss Loss_match of the sentence is proportional to the vector distance between the entity matching prediction output and the metaphorical entity vector.
[0020] Preferably, determining the entity matching loss Loss_match of the sentence includes:
[0021] Determining the random entity of the sentence based at least on the knowledge graph and the metaphorical entity label of the sentence, where the random entity is a randomly obtained entity;
[0022] Using the graph neural network to generate a random entity vector for the sentence; and
[0023] Determining an entity matching loss Loss_match for the sentence at least partially based on the entity matching prediction output of the sentence and the metaphor entity label and random entity label of the sentence, wherein the entity matching loss Loss_match of the sentence is inversely proportional to the vector distance between the entity matching prediction output and the random entity vector.
[0024] Preferably, determining the entity matching loss Loss_match for the sentence includes:
[0025] Determining at least based on the knowledge graph and the metaphor entity label of the sentence a knowledge-embedded entity of the sentence, the knowledge-embedded entity being an entity different from the metaphor entity embedded into the sentence based on the knowledge graph;
[0026] Using the graph neural network to generate a knowledge-embedded entity vector for the sentence;
[0027] Determining at least partially based on the entity matching prediction output of the sentence and the metaphor entity label and knowledge-embedded entity label of the sentence an entity matching loss Loss_match for the sentence, wherein the entity matching loss Loss_match of the sentence is inversely proportional to the vector distance between the entity matching prediction output and the knowledge-embedded entity vector.
[0028] Preferably, determining the entity matching loss Loss_match for the sentence includes:
[0029] Determining at least based on the knowledge graph and the metaphor entity label of the sentence a knowledge-embedded entity and a random entity of the sentence, the knowledge-embedded entity being an entity different from the metaphor entity embedded into the sentence based on the knowledge graph, the random entity being a randomly obtained entity;
[0030] Using the graph neural network to generate a random entity vector and a knowledge-embedded entity vector for the sentence;
[0031] Determining at least partially based on the entity matching prediction output of the sentence and the metaphor entity, knowledge-embedded entity and random entity of the sentence an entity matching loss Loss_match for the sentence, wherein the entity matching loss Loss_match of the sentence is directly proportional to the vector distance between the entity matching prediction output and the metaphor entity vector, inversely proportional to the vector distance between the entity matching prediction output and the knowledge-embedded entity vector, inversely proportional to the vector distance between the entity matching prediction output and the random entity vector, and inversely proportional to the vector distance between the knowledge-embedded entity vector and the random entity vector.
[0032] Preferably, the graph neural network is a graph convolutional network.
[0033] Preferably, when initializing the graph neural network, the word embeddings of the K-BERT model of the entity recognition model are used to generate the initial embedding representations of each node in the knowledge graph.
[0034] On the other hand, a method for entity recognition is disclosed, including:
[0035] Obtain the sentence to be processed;
[0036] Use the entity recognition model trained by the method as described herein to process the sentence to be processed. If an entity is obtained in the sequence annotation prediction output of the entity recognition model, output the obtained entity as the recognized entity;
[0037] If the entity recognition model does not recognize an entity, output the entity matching prediction output of the entity recognition model;
[0038] Use the entity matching prediction output to perform retrieval in the entity vector library to retrieve an entity vector that matches the entity matching prediction output in the entity vector library; and
[0039] Take the entity corresponding to the retrieved entity vector as the recognized entity.
[0040] Preferably, the entity vector library is obtained by performing vectorization on the knowledge graph associated with the input sentence.
[0041] Preferably, the entity vector library is obtained by performing vectorization on the input sentence through a graph neural model, and the graph neural model is iteratively updated while training the entity recognition model.
[0042] Preferably, performing retrieval in the entity vector library using the entity matching prediction output is implemented through the FAISS library.
[0043] On the other hand, a system for training an entity recognition model is disclosed, including:
[0044] A training set construction module for constructing a training set, where the training set includes a plurality of training samples d = {S input , X NER , X TER}, where S input is a sentence, X NER is the sequence annotation label of the sentence, and X TER is the metaphor entity label of the sentence, and the metaphor entity label is used to represent the metaphor entity of the sentence, where the metaphor entity is the entity that the sentence actually refers to but does not appear in the sentence; and
[0045] An entity recognition model training module for training the entity recognition model using the training set, the entity recognition model being based on a pre-trained K-BERT model, wherein the entity recognition model training module includes:
[0046] A prediction module for inputting a training sample in the training set into the entity recognition model to obtain a sequence annotation prediction output and an entity matching prediction output of the sentence in the training sample,
[0047] A loss calculation module for determining a sequence annotation loss Loss_sequence of the sentence based on the sequence annotation prediction output of the sentence and the sequence annotation label of the sentence; determining an entity matching loss Loss_match of the sentence at least partially based on the entity matching prediction output of the sentence and the metaphor entity label of the sentence; and determining a total loss Loss_total of the entity recognition model, the total loss being a weighted sum of the sequence annotation loss and the entity matching loss, i.e., Loss_total = Loss_sequence + α * Loss_match, where α indicates the weight of the entity matching loss; and
[0048] An iterative training module for iteratively performing training to minimize the total loss of the entity recognition model, thereby obtaining a trained entity recognition model.
[0049] Preferably, a knowledge graph associated with the sentence is also input into the entity recognition model.
[0050] Preferably, the loss calculation module includes an entity vectorization module for vectorizing an entity into an entity vector using a graph neural network.
[0051] Preferably, when initializing the graph neural network, the word embeddings of the K-BERT model of the entity recognition model are used to generate an initial embedding representation of each node in the knowledge graph.
[0052] In another aspect, a system for entity recognition is disclosed, including:
[0053] A sentence acquisition module for acquiring a sentence to be processed;
[0054] An entity recognition model for processing the sentence to be processed, if an entity is obtained in the sequence annotation prediction output of the entity recognition model, outputting the obtained entity as the recognized entity, and if the entity recognition model does not recognize an entity, outputting the entity matching prediction output of the entity recognition model;
[0055] A retrieval module, configured to perform retrieval in an entity vector library by using the entity matching prediction output, so as to retrieve an entity vector matching the entity matching prediction output in the entity vector library, and use the entity corresponding to the retrieved entity vector as the identified entity.
[0056] In another aspect, a device for training an entity recognition model is disclosed, including:
[0057] A memory; and
[0058] A processor, configured to execute the method for training an entity recognition model as described above.
[0059] In another aspect, a device for performing entity recognition is disclosed, including:
[0060] A memory; and
[0061] A processor, configured to execute the method for performing entity recognition.
[0062] In yet another aspect, a computer-readable storage medium storing instructions is provided, which when executed by a computer, causes the computer to execute the above method.
[0063] Compared with the prior art, one or more embodiments of this specification can achieve one or more of the following technical effects:
[0064] It can not only identify explicit entities of interest, but also identify metaphorical entities;
[0065] It can perform end-to-end training. When metaphorical entities are included, instead of re-predicting, it can directly use the obtained matching entity vectors to perform vector retrieval, thus greatly improving the efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The above invention content and the following specific implementation manners will be better understood when read in conjunction with the accompanying drawings. It should be noted that the drawings are only examples of the claimed invention. In the drawings, the same reference numerals represent the same or similar elements.
[0067] Figure 1 A very general schematic block diagram showing the process of multi-task learning for performing an entity recognition model according to an embodiment of this specification.
[0068] Figure 2 A part showing a specific example of a knowledge graph.
[0069] Figure 3 A schematic diagram showing an example of using a knowledge graph to expand a sentence according to an embodiment of this specification.
[0070] Figure 4 A schematic diagram showing the process of a sequence labeling task according to an embodiment of the present specification.
[0071] Figure 5 A schematic diagram showing the process of an entity matching task according to an embodiment of the present specification.
[0072] Figure 6 A schematic diagram showing the loss of the overall model according to an embodiment of the present specification.
[0073] Figure 7 A schematic flowchart showing an example method for training an entity recognition model according to an embodiment of the present specification.
[0074] Figure 8 A flowchart showing a method for identifying an entity using the entity recognition model according to an embodiment of the present specification.
[0075] Figure 9 A schematic diagram showing an example system for training an entity recognition model according to an embodiment of the present specification.
[0076] Figure 10 A schematic diagram showing an example system for entity recognition according to an embodiment of the present specification.
[0077] Figure 11 A schematic block diagram showing an apparatus for implementing a system according to one or more embodiments of the present specification. Detailed implementation manners
[0078] The content of the following detailed implementation manners is sufficient for any person skilled in the art to understand the technical content of one or more embodiments of the present specification and to implement it accordingly. And according to the specification, claims and drawings disclosed in the present specification, those skilled in the art can easily understand the objectives and advantages related to one or more embodiments of the present specification.
[0079] As described above, named entity recognition technology has been applied to various scenarios such as information extraction, relation extraction, syntactic analysis, information retrieval, question answering systems, machine translation, and so on. However, current named entity recognition technology usually can only recognize entities (or their synonyms) that appear in the text.
[0080] In practical applications, in many cases, the text (such as a sentence) may not include the entity itself. For example, in scenarios such as intelligent customer service and automated question answering, the customer's speech may not contain the entity itself. For example, in the intelligent customer service scenario, the user may not accurately describe the specific service / product with problems, but only describe the appearance and problems that occur. Therefore, it is necessary to infer and recognize the specific function / product feedback by the user according to the user's description.
[0081] A specific example is a certain online food ordering service. A customer may make a comment like "The delivery of the meal I ordered was too long overdue" in the feedback. This comment points to the online food ordering service, but the name of the online food ordering service does not exist in this comment, and there is not even a near-synonym of the online food ordering service. In this case, it may be difficult to accurately identify the entity pointed to by this comment using traditional named entity recognition techniques.
[0082] For ease of description, three types of entities that may be used in the embodiments of this specification are introduced below: metaphorical entities, knowledge-embedded entities, and random entities.
[0083] An explicit entity refers to an entity that appears in a sentence. However, it should be understood that there may be multiple entities in a sentence, some of which may be of interest to the user, while some other entities may not be of interest to the user; and even in some cases, all explicit entities may not be of interest to the user. Hereinafter, the explicit entity that the user is interested in may be referred to as an "explicitly interested entity", and the entity that the user is not interested in may be referred to as an "explicit other entity".
[0084] A metaphorical entity refers to an entity that the sentence actually points to but does not appear in the sentence. In most cases, the metaphorical entity exists in the entities embedded in the sentence based on the knowledge graph. It should be understood that the "entity that the sentence actually points to" here refers to the entity related to the sentence and of interest to the user. Which specific entity it is can be determined by the user through annotation according to actual needs. Therefore, the metaphorical entity can also be referred to as an "implicitly interested entity".
[0085] A knowledge-embedded entity refers to an entity embedded in the sentence based on the knowledge graph and different from the metaphorical entity.
[0086] A random entity refers to an entity obtained randomly. In most cases, a random entity refers to an entity that is not embedded in the sentence, that is, an entity that does not exist in the sentence after knowledge integration. In many cases, this random entity can be randomly selected from the knowledge graph, which is an entity other than the entities embedded in the sentence. In some other cases, this random entity can be generated independently of the knowledge graph.
[0087] Taking the sentence "The delivery of the meal I ordered was too long overdue, but fortunately the taste was good" as an example, there are explicit entities such as "I", "meal", "delivery", "overdue", "the taste was good", etc. Depending on the specific application (or depending on the user's interest), one or more of these explicit entities may be entities of interest to the user (i.e., explicitly interested entities), while some other entities may be entities that the user is not interested in (i.e., explicit other entities), or it is also possible that all of them are entities that the user is not interested in. In the example application of this specification, these explicit entities are not explicitly interested entities but explicit other entities.
[0088] After expanding the sentence through knowledge incorporation and obtaining the sentence with knowledge incorporation, the sentence becomes "The meal I ordered was delivered overdue. The Ele.me courier took too long. Fortunately, the taste is good. The Ele.me reputation." Through the expansion, three entities are incorporated into the sentence, namely "Ele.me", "courier", and "reputation". In the example application of this specification, only "Ele.me" is the entity of interest. Therefore, "Ele.me" is an implicitly interesting entity, that is, a metaphorical entity; while "courier" and "reputation" are knowledge-embedded entities. Any other entity can be a random entity.
[0089] This specification provides a solution for comprehensively and efficiently performing entity recognition. Specifically, this solution trains a K-BERT model incorporated with a knowledge graph in a multi-task learning manner that trains a fusion sequence matching task and an entity matching task, and performs vector retrieval when necessary, which can accurately and efficiently identify the entity of interest and can also identify the metaphorical entity.
[0090] See Figure 1 , which shows a very general schematic block diagram of the process for performing multi-task learning of an entity recognition model according to an embodiment of this specification.
[0091] As Figure 1 shown, the input sentence 102 together with the knowledge graph 104 used is input into the K-BERT model 106. Generally, the K-BERT model 106 can be a pre-trained K-BERT model. For example, a large-scale open corpus can be used to pre-train the K-BERT model to obtain a pre-trained K-BERT model. Examples of such large-scale open corpora can include WikiZh, WebtextZh, etc.
[0092] Subsequently, the sequence annotation task 110 can be performed using the K-BERT model 106.
[0093] In addition, the entity matching task 112 can also be performed using the K-BERT model and using the associated entity.
[0094] The sequence annotation loss obtained through the sequence annotation task 110 and the entity matching loss obtained through the entity matching task 112 are combined (for example, through weighted summation) to obtain the total loss of the entity recognition model. By iterating the sequence to minimize this total loss, a trained (or fine-tuned) entity recognition model can be obtained.
[0095] It should be understood that Figure 1 is only for illustrating the general framework of the solution of the embodiment of this specification. It shows a very rough and even inaccurate block diagram, and the specific details should refer to the following description.
[0096] Figure 1 The specific details will be described in further detail below with reference to Figure 4 and Figure 5 further described in detail.
[0097] Figure 1 The knowledge graph in Figure 2 shows a part of a specific example of a Knowledge Graph. A knowledge graph is a knowledge base that uses a graphical structured data model or topological structure to aggregate knowledge. As Figure 2 shown, the knowledge graph may include nodes and edges, where the nodes represent various entities or concepts, and the edges may represent the associations between these entities or concepts. In some examples ( Figure 2 not shown), the edges may have specific meanings (e.g., they may indicate the membership relationship between entities, etc.). Examples of knowledge graphs may include CN-DBpedia, MedicalKG, HowNet, etc.
[0098] Preferably, a knowledge graph for a specific application domain can be used. For example, in the application scenario of the Ele.me online food ordering service, the knowledge graph may include entities or concepts associated with this online food ordering service. For example, express delivery may be associated with refund and delivery timeliness, both of which are associated with Ele.me. In addition, credit limit adjustment may be associated with Huabei, and Huabei may cause page lags, etc.
[0099] Next, a brief introduction to the K-BERT model used in this application will be given first. The K-BERT model integrates the knowledge graph into the BERT model, thereby enabling the model to understand language knowledge more accurately. Specifically, the K-BERT model adopts the form of a sentence tree to incorporate knowledge into the input sentence and obtains a vector representation in the representation space of the pre-trained model. The K-BERT model mainly includes four components: the Knowledge layer, the Embedding layer, the Seeing layer, and the Mask-Transformer Encoder. For more details about the K-BERT model, reference can be made to the paper "K-BERT: Enabling Language Representation with Knowledge Graph" published by Weijie Liu et al. (hereinafter referred to as the "K-BERT paper"). By citing, the content of this paper is incorporated herein in its entirety, and more details about K-BERT will not be described in detail here.
[0100] SeeFigure 3 , which shows a schematic diagram of an example of using a knowledge graph to expand a sentence according to an embodiment of this specification. As described in the K-BERT paper mentioned above, in the knowledge layer of the K-BERT model, a knowledge graph can be used to expand the input sentence S input = {s0, s1, … s i , …, s n} (where s i is the i-th character), and finally obtain the sentence S knowledge = {s0, s1, … si{(r i0 , s i0 ), …, (r ik , s ik )}, …, s n}, where s ik represents the k-th entity connected to s i (that is, the k-th node connected to the node of entity s i ) in the knowledge graph, and r ik represents the relationship between entity s ik and entity s i (that is, the edge between node s i and node s ik ) in the knowledge graph).
[0101] In Figure 3 's example, assume that the input sentence "The meal I ordered was delivered too late, but fortunately the taste was good." is being processed. After integrating the knowledge graph as shown in Figure 2 , since "delivered too late (i.e., delivery timeliness)" is connected to "Ele.me" and "express delivery", and "the taste was good (i.e., taste evaluation)" is connected to "Ele.me" and "word-of-mouth", the sentence is converted into a knowledge-incorporated sentence tree "The meal I ordered was delivered too late Ele.me express delivery for too long, but fortunately the taste was good Ele.me word-of-mouth.". It can be seen that by embedding the connected entities into the sentence, a knowledge-incorporated sentence is obtained. This knowledge-incorporated sentence tree can be referred to as the "knowledge-incorporated sentence" hereinafter. Note that although only the incorporated content is shown in Figure 3 and it is not presented in the form of a sentence tree, in actual operation, it can be processed into a sentence tree structure. The process of integrating knowledge into the input sentence is usually completed by the knowledge layer of the K-BERT model. The specific details of the above method can be referred to the implementation of the knowledge embedding layer of the K-BERT model.
[0102] Subsequently, the sentence integrated with knowledge can be processed by subsequent modules of the K-BERT model. Specifically, the sentence integrated with knowledge will be input into the embedding layer and the viewing layer, and the outputs of the embedding layer and the viewing layer are further input into the masked transformer encoder to output the sentence representation of the sentence integrated with knowledge, which can be used to perform various tasks such as classification, sequence labeling, including entity recognition as described in the embodiments of this specification, and so on.
[0103] According to the principles discussed in the K-BERT paper, different ways from the above can be adopted to use the knowledge graph to augment the sentence to obtain the sentence integrated with knowledge, and the specific details are not elaborated here.
[0104] See Figure 4 , which shows a schematic diagram of the process of the sequence labeling task according to the embodiments of this specification. This process can be viewed in combination with the left half of Figure 1 therein.
[0105] Through the description in the K-BERT paper mentioned above, it can be known how to use the K-BERT model to obtain the sequence labeling prediction output of the input sentence, and further perform the sequence labeling task through the obtained vector. In the following, this process will be briefly introduced. During the training process of the entity recognition model, this sequence labeling task is used to obtain the sequence labeling task loss, which constitutes a part of the total loss of the entity recognition model. During the process of performing the entity recognition task with the entity recognition model, this sequence labeling task is used to perform named entity recognition, so as to obtain the explicitly interested entities of the sentence to be processed.
[0106] In the following, the training and use of the entity recognition model according to the embodiments of this specification will be specifically described. This entity recognition model is a K-BERT-based model (such as the K-BERT-based model 406 and the K-BERT-based model 506 described below). This K-BERT model can be a pre-trained K-BERT model. Specifically, this entity recognition model is based on the K-BERT model and is implemented by combining other layers for performing specific tasks (such as sequence labeling, entity matching, etc.).
[0107] In the following, the entity recognition model and the K-BERT-based model can be used interchangeably.
[0108] As Figure 4 shown, the input sentence 402 and the knowledge graph 404 are input into the K-BERT-based model 406.
[0109] The K-BERT-based model is based on K-BERT and may include other processing layers as well to perform specific tasks, such as Sigmoid layer, Softmax layer, LTSM layer, CRF layer, etc. The model 406 can process the input sentence 402 based on the knowledge graph 404 and generate a sequence labeling prediction output 408, which is the predicted sequence labeling representation of the input sentence. For example, the predicted vector representation of the input sentence can be obtained through the K-BERT model, and the sequence labeling representation of the input sentence can be obtained through the LTSM layer or CRF layer, etc., as the sequence labeling prediction output 408 of the input sentence.
[0110] To perform training for the sequence labeling task (specifically the named entity recognition (NER) task in this embodiment), the samples in the training set can be first sequence-labeled. Any suitable labeling strategy or format can be used to perform sequence labeling for named entity recognition on the input sentence. In this way, the labeled sequence labeling samples d NER ={S input ,X NER}, where S input is the input sentence {s0, s1, … s i ,…, s n}, and X NER is the label {x0, x1, … x i ,…, x n} after named entity annotation of the knowledge-incorporated sentence, where x i is the named entity annotation label of the i-th character.
[0111] For example, referring to the example shown in Figure 3 , for the sentence "Ordered a chicken cutlet rice on Ele.me", its sequence after expansion and addition of the start symbol (token) "[CLS]" and end symbol "[SEP]" is "[CLS] Ordered a chicken cutlet rice on Ele.me [SEP]", and its label X NER after named entity annotation is {O O B-T I-T I-T O O B-T I-T I-T O}. Although the BIO format is used for annotation in this example, other applicable annotation formats can also be used.
[0112] Any method known to those skilled in the art can be used to perform sequence labeling on the sentences in the sample for named entity recognition to label the explicitly interested entities. The specific details of performing annotation on the input sentence for named entity recognition are well known in the art and will not be elaborated here. Note that not all named entities in the sentence need to be labeled, but only the interested entities. For example, in the example of "ordered a chicken cutlet rice on Ele.me", only "Ele.me" can be labeled, and "chicken cutlet rice" does not need to be labeled, that is, the label X after named entity annotation NER is {O O B-T I-T I-T O O O O O O}. In this way, only the interested entities can be recognized.
[0113] In addition, if all the entities in the input sentence are not interested entities, that is, they are all "explicit other entities", then the label X after named entity annotation NER can have all elements as O. For example, in the example of "The meal I ordered was delivered too late, but fortunately it tasted good.", assuming there are no interested entities in this sentence (the actually interested one is the metaphorical entity "Ele.me"), then the labels of this sentence can all be O, that is, X NER= {O O O……}
[0114] Subsequently, based on the sequence annotation prediction output 408 of the input sentence and the sequence annotation label 412 of the input sentence, the sequence annotation loss 410 of the sentence can be determined. This sequence annotation loss can be expressed as Loss_sequence, for example. In one example, the sequence annotation loss Loss_sequence can be defined as follows:
[0115] Loss_sequence = cross_entropy_loss(label, target)
[0116] where label is the sequence annotation label 412, and target is the sequence annotation prediction output 408 of the input sentence output by the BERT model. Among them, cross_entropy_loss is the cross-entropy loss function, which is known to those skilled in the art. For example, it can be found in multiple open-source libraries (such as Pytorch) (CrossEntropyLoss() function) and can be directly called, which will not be elaborated here.
[0117] It can be understood that the above cross-entropy loss function is only an example and not a limitation. Any other applicable loss function conceivable by those skilled in the art can be used to calculate the sequence loss annotation.
[0118] In the traditional single-task training process, only by minimizing the sequence annotation loss Loss_sequence can a trained K-BERT model for the sequence annotation task be obtained.
[0119] However, when performing multi-task learning, it is not simply to minimize the sequence annotation loss Loss_sequence alone, but to minimize the total loss of multiple tasks, and its specific process will be described below.
[0120] It should be noted that the above description of the sequence annotation task is only for ease of understanding and not to limit the scope of the present invention. Any appropriate method different from the above can be used to utilize the K-BERT model to perform the sequence annotation task and generate the corresponding sequence annotation loss.
[0121] See Figure 5 , which shows a schematic diagram of the process of the entity matching task according to an embodiment of the present specification. This process can be combined with the right half in Figure 1 .
[0122] In the training process of the entity recognition model, this sequence annotation task is used to obtain the sequence annotation task loss, and this sequence annotation task loss constitutes a part of the total loss of the entity recognition model. In the process of using the entity recognition model to perform the entity recognition task, this sequence annotation task is used to perform named entity recognition, so as to obtain the explicit entities of interest in the sentence to be processed.
[0123] Next, the entity matching task will be introduced through multiple examples.
[0124] First example
[0125] The input sentence 502 and the knowledge graph 504 are input into the K-BERT-based model 506 (for example, input into the K-BERT model). Similarly, the entity matching prediction output 508 is obtained. Specifically, this entity matching prediction output 508 can be the vector representation of this sentence output by the K-BERT model. In actual implementation, the vector representation of this sentence can usually take the vector representation of the first symbol of this sentence, that is, the vector representation of "[CLS]". It should be understood that any method conceivable by those skilled in the art can be used to obtain the output vector of the input sentence 502 as the entity matching prediction output 508.
[0126] As Figure 5 shown, in order to perform training for the entity recognition task, the samples in the training set can be first annotated with metaphorical entities to obtain their metaphorical entity labels.
[0127] The annotation of metaphorical entities can be carried out in various ways.
[0128] In one implementation, metaphorical entity annotation can be performed manually on the input sentence. For example, metaphorical entities can be manually embedded into the input sentence, and the embedded metaphorical entities can be annotated. Usually, this manual annotation can be carried out by specialized personnel. Usually, this specialized personnel knows the knowledge graph association between the metaphorical entity and the entity in the input sentence. Therefore, when performing manual annotation, the determined metaphorical entity can be an entity related to the entity in the input sentence in the knowledge graph. For example, in the example of "The meal I ordered was delivered too late, but fortunately the taste was good.", this specialized personnel knows the association between "delivered too late" and "Ele.me" and "express delivery" in the knowledge graph, and makes annotations accordingly. In some examples, this specialized personnel can use references to the knowledge graph (such as querying the knowledge graph) to assist in the annotation.
[0129] In another implementation, metaphorical entity annotation can be performed on the input sentence based on rules. For example, it can be performed based on the associations in the knowledge graph and according to a specific entity library of interest. For example, in the example of "The meal I ordered was delivered too late, but fortunately the taste was good.", by automatically querying the knowledge graph, entities such as "Ele.me", "express delivery", and "Word of Mouth" can be obtained. Subsequently, a search can be performed in the entity library of interest. For example, it can be retrieved that "Ele.me" is in the entity library of interest among the above entities. At this time, "Ele.me" can be automatically annotated as the metaphorical entity of this sentence.
[0130] By performing metaphorical entity annotation on the input sentence, the metaphorical entity 512 can be obtained.
[0131] In this way, an entity matching sample d annotated with metaphorical entities can be obtained ER ={S input ,X TER}, where S input is the input sentence {s0, s1, … s i ,…, s n}, and X TER is the metaphorical entity label of this input sentence.
[0132] The metaphorical entity label can take any appropriate form. For example, this metaphorical entity label can be the name of the metaphorical entity 510 of this input sentence. For example, referring to the example shown in Figure 3 , for the sentence "The meal I ordered was delivered too late...", its sentence after being expanded and added with the start symbol (token) "[CLS]" and the end symbol "[SEP]" ( Figure 3 not shown) is "[CLS] The meal I ordered was delivered too late Ele.me express delivery too long...", and its metaphorical entity label X TER after metaphorical entity annotation is "Ele.me", that is, X TER = "Ele.me".
[0133] The metaphor entity labels can also take other forms conceivable by those skilled in the art. For example, a form similar to the labels commonly used in sequence annotation can be adopted, except that the sentences incorporated with knowledge are labeled. For example, for Figure 3 the example of "[CLS] The meal I ordered was delivered overdue. The Ele.me courier took too long..." in TER its metaphor entity label can be "O O OO O O O O O B-T I-T I-T O O O O O...", that is, the sequence representation of labeling "Ele.me" as the entity of interest. That is, namely X
[0134] It can be understood that regardless of which label form is adopted, through the metaphor entity label 510, the metaphor entity 512 can be obtained.
[0135] Subsequently, the metaphor entity 512 is processed by the graph neural network 518 to vectorize the metaphor entity 512, so as to obtain the metaphor entity vector 520 of the metaphor entity label. The specific process of using the knowledge graph GCN to vectorize the entity will be described in detail below.
[0136] It should be understood that after obtaining the entity matching prediction output 508 of the input sentence 502 through the K-BERT-based model 506 and getting the metaphor entity vector 520 of the metaphor entity 512 of the input sentence 502, the entity matching loss 526 of the entity matching task can be calculated. The entity matching loss can be represented as Loss_match for example. In one example, the entity matching loss Loss_match can be calculated in the following manner:
[0137] Loss_match = distance(e_i, e_t) (Formula 1)
[0138] where e_i is the entity matching prediction output 508 of the sentence output by the K-BERT model (which is the vector representation of the sentence output by the K-BERT model), and e_t is the metaphor entity vector 520 of the sentence. Among them, distance is a vector distance function used to compare the vector distances between vectors. The larger the vector distance, the larger the value of this function. That is to say, in this first example, the entity matching loss Loss_match of the sentence is proportional to the vector distance between the entity matching prediction output 508 and the metaphor entity vector.
[0139] Various ways can be adopted to implement the above vector distance function distance. For example, the cosine distance, Euclidean distance, Pearson correlation coefficient, Jaccard similarity coefficient, etc. between two vectors can be calculated.
[0140] It can be seen that the calculation method of the above loss function makes the predicted output of entity matching in the sentence as similar as possible to the metaphor entity vector, so that the predicted entity is as similar as possible to the metaphor entity.
[0141] Second example
[0142] In a preferred example, when determining the entity matching loss, in addition to the metaphor entity 512, a knowledge-embedded entity 516 can also be introduced. A knowledge-embedded entity refers to an entity different from the metaphor entity that is embedded in the sentence based on a knowledge graph.
[0143] Different from the metaphor entity being labeled by a metaphor entity label, there is no need to label the knowledge-embedded entity. Instead, the entity recognition model can automatically determine the knowledge-embedded entity based on an algorithm.
[0144] For example, after determining the entities embedded in the sentence, excluding the labeled metaphor entity 512 from the entities embedded in the sentence, the knowledge-embedded entity 516 can be obtained. Preferably, the knowledge layer of the K-BERT model can be used to determine which entities are embedded in the sentence, as described above.
[0145] After determining the knowledge-embedded entity, any one of the knowledge-embedded entities can be selected and put into a sample as the knowledge-embedded entity 516. In this way, multiple samples can be generated.
[0146] Subsequently, similarly, the knowledge-embedded entity 516 can also be processed by the graph neural network 518 to vectorize the knowledge-embedded entity 516, so as to obtain the knowledge-embedded entity vector 524 of the knowledge-embedded entity 516. The specific process of using the knowledge graph GCN to vectorize the entity will be described in detail below.
[0147] In this preferred example, the entity matching loss Loss_match can be calculated in the following way:
[0148] Loss_match = distance(e_i, e_t) – β * distance(e_i, e_s), (Formula 2)
[0149] Where e_s is the knowledge-embedded entity vector 524, and β is a weight parameter. The value of β can be implemented in any known way in the art.
[0150] It can be seen that the above loss function aims to make the entity matching prediction output of the sentence as similar as possible to the metaphor entity vector, and make the entity matching prediction output as dissimilar as possible to the knowledge-embedded entity vector, so that the predicted entity is as similar as possible to the metaphor entity and dissimilar to the knowledge-embedded entity. That is to say, in this second example, the entity matching loss Loss_match of the sentence is proportional to the vector distance between the entity matching prediction output 508 and the metaphor entity vector, and inversely proportional to the vector distance between the entity matching prediction output 508 and the knowledge-embedded entity vector 524.
[0151] Introducing the knowledge-embedded entity increases the amount of information, enabling the trained model to distinguish between metaphor entities and knowledge-embedded entities in the knowledge graph.
[0152] The third example
[0153] When determining the entity matching loss, in addition to the metaphor entity 516, a random entity 514 can also be introduced.
[0154] The knowledge-embedded entity refers to an entity other than the metaphor entity that is embedded into the sentence according to the knowledge graph.
[0155] However, there is no need to label the knowledge-embedded entity. Instead, the model can automatically determine the random entity 514 based on the algorithm.
[0156] As the name implies, the random entity is a randomly obtained entity. For example, the random entity can be, for example, an entity in the knowledge graph other than the metaphor entity and the knowledge-embedded entity. For example, in Figure 2 and Figure 3 's example, it has been determined that "Ele.me" is the metaphor entity. At this time, "express delivery" and "Word of Mouth" can be determined as knowledge-embedded entities, and "Huabei", "Ride Code", etc. are random entities.
[0157] Alternatively, the random entity can be an entity randomly constructed in other ways. For example, it can be an entity not included in the knowledge graph. For example, the random entity can be an entity randomly selected from a wider database, or can be an entity randomly generated in any way.
[0158] Subsequently, similarly, the random entity 514 can also be processed through the graph neural network 518 to vectorize the random entity 514, so as to obtain the random entity vector 522 of the random entity 514. The specific process of using the knowledge graph GCN to vectorize the entity will be described in detail below.
[0159] In this preferred example, the entity matching loss Loss_match can be calculated in the following way:
[0160] Loss_match = distance(e_i, e_t) – β * distance(e_i, e_r), (Equation 3)
[0161] where e_r is the random entity vector 522, and β is the weight parameter. The value of β can be implemented in any manner known in the art. This β can be the same as or different from the β in the second example above.
[0162] It can be seen that the above loss function is intended to make the entity matching prediction output of the sentence as similar as possible to the metaphor entity vector, and make the entity matching prediction output as dissimilar as possible to the random entity vector, so that the predicted entity is as similar as possible to the metaphor entity and dissimilar to the random entity. That is to say, in this third example, the entity matching loss Loss_match of the sentence is proportional to the vector distance between the entity matching prediction output 508 and the metaphor entity vector, and inversely proportional to the vector distance between the entity matching prediction output 508 and the random entity vector 522.
[0163] Introducing the knowledge-embedded entity increases the amount of information, enabling the trained model to distinguish between metaphor entities and random entities in the knowledge graph. That is to say, the trained model can be made to know that the metaphor entity is an entity that appears in the knowledge graph rather than a random entity.
[0164] Fourth Example
[0165] In a more preferred example, the metaphor entity, the knowledge-embedded entity, and the random entity can be considered simultaneously. At this time, the entity matching loss Loss_match can be calculated as follows:
[0166] Loss_match = distance(e_i, e_t) – β * (distance(e_i, e_r) + distance(e_i, e_s) + distance(e_r, e_s)), (Equation 4)
[0167] It can be seen that the above loss function is intended to make the entity matching prediction output of the sentence as similar as possible to the metaphor entity vector, and make the entity matching prediction output, the knowledge-embedded entity vector, and the random entity vector dissimilar to each other pairwise. That is to say, in this fourth example, the entity matching loss Loss_match of the sentence is proportional to the vector distance between the entity matching prediction output and the metaphor entity vector, inversely proportional to the vector distance between the entity matching prediction output and the knowledge-embedded entity vector, inversely proportional to the vector distance between the entity matching prediction output and the random entity vector, and inversely proportional to the vector distance between the knowledge-embedded entity vector and the random entity vector.
[0168] It can be understood that the above formula 4 is merely a special example, and any example satisfying the above direct and inverse proportion relationships can be adopted. For example:
[0169] Loss_match = distance(e_i, e_t) – (β1 * distance(e_i, e_r) + β2 * distance(e_i, e_s) + β3 * distance(e_r, e_s)), (Formula 5)
[0170] Introducing knowledge-embedded entities increases the amount of information, enabling the trained model to distinguish metaphorical entities, knowledge-embedded entities, and random entities.
[0171] As described above, in the case of multiple knowledge-embedded entities or multiple random entities, the metaphorical entity, knowledge-embedded entity, and random entity can be combined so that each sample includes one metaphorical entity, zero or one knowledge-embedded entity, and zero or one random entity. In this way, for the same sentence, multiple input examples can be obtained, and these input examples can include different embedded entities and / or random entities.
[0172] As previously mentioned, one or more of the metaphorical entity 512, random entity 514, and knowledge-embedded entity 516 need to be vectorized to obtain the corresponding metaphorical entity vector 520, random entity vector 522, and knowledge-embedded entity vector 524. It can be understood that the vectorization of the above entities should be performed in the context of the knowledge graph. In the embodiments of this specification, a graph neural network associated with the knowledge graph is used to vectorize the entities. Specifically, the graph neural network is used to vectorize the knowledge graph to obtain the vectors of the entities in the knowledge graph.
[0173] In a preferred example, a graph convolutional network associated with the knowledge graph (abbreviated as "knowledge graph GCN") can be used to vectorize the entities.
[0174] Specifically, in one example, when initializing the knowledge graph GCN, the word embeddings of the K-BERT model can be used to generate the initial embedding representations of each node in the knowledge graph. Subsequently, the initial embedding representations of each node can be further processed (such as pooling, etc.) through the knowledge graph GCN to obtain the final layer embedding representations of each node. Subsequently, during the training process of the entity recognition model, the knowledge graph GCN can be iteratively updated to achieve the optimal vectorization of the nodes in the knowledge graph. The specific implementation details of the graph convolutional network are known to those skilled in the art and will not be elaborated here.
[0175] It should be understood that the graph convolutional network is merely an example of the graph neural network, and any suitable graph neural network conceivable by those skilled in the art can be adopted.
[0176] After separately calculating the loss of the sequence labeling task and the loss of the entity matching task, the total loss of the entire model can be calculated. Refer to Figure 6 , which shows a schematic diagram of the loss of the total model according to an embodiment of the present specification.
[0177] By combining the training sets for the above two tasks, the total training samples of the entity recognition model according to an embodiment of the present specification can be obtained. For example, by combining the training samples {S input , X NER} for the sequence labeling task and the training samples {S input , X TER} for the entity matching task, the training sample d = {S input , X NER , X TER} for training the overall entity recognition model can be obtained.
[0178] As Figure 6 shown, by combining the sequence labeling loss and the entity matching loss, the total loss of the model can be obtained.
[0179] Subsequently, training is performed in the manner described above, and the sequence labeling loss Loss_sequence and the entity matching loss Loss_match are obtained respectively.
[0180] Subsequently, the total loss Loss_total of the model can be calculated by the following formula:
[0181] Loss_total = Loss_sequence + α * Loss_match
[0182] where α is a hyperparameter, which indicates the weight of the entity matching loss and can be used to adjust the influence degree of the sequence labeling loss and the entity matching loss on the total loss. The value of α can be determined according to experience or based on experimental data.
[0183] According to the description above with reference to Figures 1-6 , the composition and training process of the entity recognition model of the present application can be understood.
[0184] After obtaining the functional representation of the total loss Loss_total, training can be iteratively performed to minimize the total loss of the entity recognition model, thereby obtaining the trained entity recognition model. In the case where the samples and the loss function have been described, those skilled in the art know how to obtain the trained entity recognition model through iterative training, and details are not described herein again.
[0185] Based on the above description, the training method of the entity recognition model is generally described below. Refer to Figure 7, which shows a schematic flowchart of an example method 700 for training an entity recognition model according to an embodiment of this specification. For specific details of the operations of this method, reference may be made to the above description.
[0186] As Figure 7 shown, method 700 may include: at operation 702, a training set may be constructed. As described above, the training set may include a plurality of training samples d = {S, X NER , X TER}, where S is a sentence, X NER is the sequence annotation label of the sentence, and X TER is the metaphorical entity label of the sentence. As described above, the sequence annotation label may annotate the explicit entities of interest (if any) in the sentence. The metaphorical entity label may be used to represent the metaphorical entity of the sentence. As described above, a metaphorical entity is an entity that the sentence actually refers to but does not appear in the sentence.
[0187] After constructing the training set, the constructed one may be used to train the entity recognition model. As described above, the entity recognition model may be based on a pre-trained K-BERT model. Specifically, the underlying layer of the entity recognition model may be a pre-trained K-BERT model, and the subsequent layers may be other layers for performing specific tasks. The specific implementation of the other layers may be selected by those skilled in the art according to their tasks, and the embodiments of this specification are not limited in this regard. It can be understood that in addition to the training sample set, the knowledge graph associated with the sentence is also input into the entity recognition model. The knowledge graph may be, for example, a knowledge graph for a specific application domain.
[0188] Specifically, the training of the entity recognition model may be performed in the following manner.
[0189] As Figure 7 shown, method 700 may include: at operation 704, the training samples in the training set may be input into the entity recognition model to obtain the sequence annotation prediction output and the entity matching prediction output of the sentence in the training sample. As described above, the sequence annotation prediction output of the sentence may be the sequence annotation representation of the sentence, which may be implemented by the process described above in combination with Figure 4 description. The entity matching prediction output of the sentence may be the vector representation of the sentence, which may be implemented by the process described above in combination with Figure 5 description. The vector representation may be, for example, the vector representation of the first character (i.e., "[CLS]") of the sentence.
[0190] Method 700 may further include: at operation 706, the sequence annotation loss Loss_sequence of the sentence may be determined based on the sequence annotation prediction output of the sentence and the sequence annotation label of the sentence.
[0191] The method 700 may further include: at operation 708, determining an entity matching loss Loss_match for the sentence, at least partially based on the entity matching prediction output of the sentence and the metaphor entity label of the sentence.
[0192] As shown in the first example above, the entity matching loss of the sentence can be determined only based on the metaphor entity label.
[0193] In this case, operation 708 may include the following steps:
[0194] First, a graph neural network associated with the knowledge graph can be used to generate a metaphor entity vector for the sentence. The metaphor entity can be obtained from the metaphor entity label, and a metaphor entity vector corresponding to the metaphor entity can be obtained using a graph neural network associated with the knowledge graph (such as a knowledge spectral graph GCN).
[0195] Subsequently, the vector distance between the entity matching prediction output of the sentence and the metaphor entity vector can be determined. The calculation of the vector distance can refer to the description above.
[0196] Then, the entity matching loss Loss_match for the sentence can be determined, where the entity matching loss Loss_match for the sentence is proportional to the vector distance between the entity matching prediction output and the metaphor entity vector. Formula 1 above can be referred to.
[0197] As shown in the second example above, the entity matching loss of the sentence can also be determined based on the metaphor entity label and the determined knowledge-embedded entity.
[0198] In this case, operation 708 may include the following steps:
[0199] First, the knowledge-embedded entity of the sentence can be determined at least based on the knowledge graph and the metaphor entity label of the sentence. As described above, the knowledge-embedded entity can refer to an entity different from the metaphor entity embedded into the sentence based on the knowledge graph. Through the metaphor entity label, the metaphor entity can be determined. One or more entities to be embedded into the sentence can be obtained based on the knowledge graph (such as using the knowledge layer of K-BERT), and after excluding the metaphor entity from the one or more entities, the knowledge-embedded entity of the sentence can be obtained.
[0200] Subsequently, the graph neural network can be used to generate a knowledge-embedded entity vector for the sentence.
[0201] Then, the vector distance between the entity matching prediction output and the knowledge-embedded entity vector can be determined. The vector distance between the entity matching prediction output and the metaphor entity vector can also be determined.
[0202] Then, the entity matching loss Loss_match of the sentence can be determined at least partially based on the entity matching prediction output of the sentence and the metaphor entity label and knowledge embedding entity label of the sentence, where the entity matching loss Loss_match of the sentence is inversely proportional to the vector distance between the entity matching prediction output and the knowledge embedding entity vector. That is to say, the entity matching loss can be proportional to the vector distance between the entity matching prediction output and the metaphor entity vector, and inversely proportional to the vector distance between the entity matching prediction output and the knowledge embedding entity vector. Reference can be made to Equation 2 above.
[0203] As shown in the third example above, the entity matching loss of a sentence can also be determined based on the metaphor entity label and a random entity.
[0204] In this case, operation 708 may include the following steps:
[0205] First, the random entity of the sentence can be determined at least based on the knowledge graph and the metaphor entity label of the sentence, where the random entity is a randomly obtained entity. The process of obtaining the random entity can be referred to the description above. Preferably, the random entity is different from the metaphor entity and the knowledge embedding entity, but this is not necessary.
[0206] Subsequently, the graph neural network can be used to generate the random entity vector of the sentence.
[0207] Then, the vector distance between the entity matching prediction output and the random entity vector can be determined. At the same time, the vector distance between the entity matching prediction output and the metaphor entity vector can also be determined
[0208] Then, the entity matching loss Loss_match of the sentence can be determined at least partially based on the entity matching prediction output of the sentence and the metaphor entity label and random entity label of the sentence, where the entity matching loss Loss_match of the sentence is inversely proportional to the vector distance between the entity matching prediction output and the random entity vector. That is to say, the entity matching loss can be proportional to the vector distance between the entity matching prediction output and the metaphor entity vector, and inversely proportional to the vector distance between the entity matching prediction output and the random entity vector. Reference can be made to Equation 3 above.
[0209] As shown in the fourth example above, the entity matching loss of a sentence can also be determined based on the metaphor entity label and a combination of the knowledge embedding entity and the random entity.
[0210] First, the knowledge embedding entity and the random entity of the sentence can be determined at least based on the knowledge graph and the metaphor entity label of the sentence.
[0211] Subsequently, the graph neural network can be used to generate a random entity vector and a knowledge-embedded entity vector for the sentence.
[0212] Then, the vector distance between the entity matching prediction output of the sentence and the metaphor entity vector can be determined. The entity matching prediction output of the sentence and the knowledge-embedded entity vector can be determined. The vector distance between the entity matching prediction output of the sentence and the random entity vector can be determined. The vector distance between the entity matching prediction output of the sentence and the random entity vector can also be determined.
[0213] Then, the entity matching loss Loss_match of the sentence can be determined based at least in part on the entity matching prediction output of the sentence and the metaphor entity, knowledge-embedded entity, and random entity of the sentence, where the entity matching loss Loss_match of the sentence is proportional to the vector distance between the entity matching prediction output and the metaphor entity vector, inversely proportional to the vector distance between the entity matching prediction output and the knowledge-embedded entity vector, inversely proportional to the vector distance between the entity matching prediction output and the random entity vector, and inversely proportional to the vector distance between the knowledge-embedded entity vector and the random entity vector. See Formulas 4 and 5 above.
[0214] The graph neural network used to convert the metaphor entity, random entity, or knowledge-embedded entity into the corresponding metaphor entity vector, random entity vector, or knowledge-embedded entity vector respectively can adopt various applicable models. Preferably, the graph neural network is a graph convolutional network. When initializing the graph neural network, preferably, the word embeddings of the K-BERT model of the entity recognition model can be used to generate the initial embedding representation of each node in the knowledge graph.
[0215] It can be appreciated that although in Figure 7 operation 708 is shown outside operation 706, this is not a limitation. In actual applications, operations 706 and 708 can be executed in any order or in parallel.
[0216] Method 700 may further include: at operation 710, the total loss Loss_total of the entity recognition model can be determined, and the total loss is the weighted sum of the sequence annotation loss and the entity matching loss, that is: Loss_total = Loss_sequence + α * Loss_match. Where α indicates the weight of the entity matching loss.
[0217] Method 700 may further include: at operation 712, training can be iteratively executed to minimize the total loss of the entity recognition model, thereby obtaining a trained entity recognition model. For example, the above operations can be iteratively executed using a large number of training samples in the training set to minimize the total loss.
[0218] It can be appreciated that the total loss reflects two tasks: on the one hand, it enables the predicted output of the sequence annotation of the output sequence to match the explicitly interested entity as much as possible, and on the other hand, it makes the predicted output of the entity match as similar as possible to the metaphorical entity vector (and / or dissimilar to the knowledge-embedded entity and the random entity). Through such training, the following effects can be achieved: either the explicitly interested entity can be directly predicted through sequence annotation, or the entity vector of the entity similar to the metaphorical entity can be output, and the entity vector can then be used to perform retrieval in the entity vector space.
[0219] The following describes the specific process of using the entity recognition model to recognize entities. Refer to Figure 8 , which shows a flowchart of an example method 800 for recognizing entities using the entity recognition model according to an embodiment of the present specification.
[0220] As Figure 8 shown, method 800 may include: at operation 802, a sentence to be processed may be obtained. The sentence to be processed refers to the sentence for which entity recognition is to be performed.
[0221] Method 800 may further include: at operation 804, the trained entity recognition model as described above may be used to process the sentence to be processed, and if an entity is obtained in the sequence annotation predicted output of the entity recognition model, the obtained entity may be output as the recognized entity. That is, if an explicitly interested entity is recognized in the sentence to be processed, the recognized entity may be directly output as the result of entity recognition at operation 806.
[0222] Method 800 may further include: at operation 808, if the entity recognition model does not recognize an entity, the entity match predicted output of the entity recognition model may be output. The entity match predicted output is an entity vector similar to the metaphorical entity, which can be used to directly perform entity vector retrieval subsequently.
[0223] Method 800 may further include: at operation 810, retrieving may be performed in the entity vector library using the entity matching prediction output to retrieve an entity vector that matches the entity matching prediction output in the entity vector library. The entity vector library may be obtained, for example, by performing vectorization on a knowledge graph associated with the input sentence. The vectorization may be performed, for example, by a graph neural model. The graph neural model may preferably be a knowledge graph GCN model. The graph neural model used when performing prediction is the same as the entity recognition model used when training the entity recognition model. As described above, the graph neural model is iteratively updated while training the entity recognition model. Preferably, retrieving in the entity vector library using the entity matching prediction output is implemented through the FAISS library. The FAISS library is a library developed by Facebook for efficiently performing vector retrieval. The specific process of performing vector retrieval using the FAISS library will not be elaborated here.
[0224] Method 800 may further include: at operation 812, the entity corresponding to the retrieved entity vector may be used as the recognized entity. For example, the corresponding entity of the entity vector may be obtained through the graph neural network.
[0225] See Figure 9 , which shows a schematic diagram of an example system 900 for training an entity recognition model according to an embodiment of the present specification. The system 900 may be used to perform the training of the entity recognition model.
[0226] As Figure 9 shown, the system 900 may include a training set construction module 902 for constructing a training set, where the training set includes a plurality of training samples d = {Sinput, XNER, XTER}, where Sinput is a sentence, XNER is the sequence annotation label of the sentence, and XTER is the metaphorical entity label of the sentence. The metaphorical entity label is used to represent the metaphorical entity of the sentence, where the metaphorical entity is the entity that the sentence actually refers to but does not appear in the sentence.
[0227] The system 900 may further include an entity recognition model training module 904 for training the entity recognition model using the training set. The entity recognition model is based on a pre-trained K-BERT model. The entity recognition model training module 904 further includes:
[0228] A prediction module 906 for inputting the training samples in the training set into the entity recognition model to obtain a sequence annotation prediction output and an entity matching prediction output for the sentence in the training sample. Specifically, it may include a sequence annotation prediction module and an entity matching prediction module (not shown in the figure).
[0229] A loss calculation module 908 is configured to determine a sequence annotation loss Loss_sequence of the sentence based on the sequence annotation prediction output of the sentence and the sequence annotation label of the sentence, and determine an entity matching loss Loss_match of the sentence at least partially based on the entity matching prediction output of the sentence and the metaphor entity label of the sentence. Specifically, it may include a sequence annotation loss calculation module and an entity matching loss calculation module.
[0230] The loss calculation module 908 is further configured to determine a total loss Loss_total of the entity recognition model, where the total loss is a weighted sum of the sequence annotation loss and the entity matching loss, that is: Loss_total = Loss_sequence + α * Loss_match, where α indicates the weight of the entity matching loss.
[0231] An iterative training module 910 is configured to iteratively perform training to minimize the total loss of the entity recognition model, thereby obtaining a trained entity recognition model.
[0232] For specific details of related operations, reference may be made to the specific description of method 700 above.
[0233] See Figure 10 , which shows a schematic diagram of an example system 1000 for entity recognition according to an embodiment of the present specification. The system can be used to perform entity recognition using a trained entity recognition model.
[0234] As Figure 10 shown, the system 1000 may include a sentence acquisition module 1002, which can be used to acquire a sentence to be processed.
[0235] The system 1000 may further include an entity recognition model 1004, which can be used to process the sentence to be processed. If an entity is obtained in the sequence annotation prediction output of the entity recognition model, the obtained entity is output as the recognized entity, and if the entity recognition model does not recognize an entity, the entity matching prediction output of the entity recognition model is output.
[0236] The system 1000 may further include: a retrieval module 1006, which can be used to perform a retrieval in an entity vector library using the entity matching prediction output to retrieve an entity vector that matches the entity matching prediction output in the entity vector library, and use the entity corresponding to the retrieved entity vector as the recognized entity.
[0237] For specific details of related operations, reference may be made to the specific description of method 800 above.
[0238] Figure 11FIG. 1100 is a schematic block diagram of an apparatus for implementing a system (such as system 900 or system 1000 above) according to one or more embodiments of the present specification. The apparatus may include a processor 1110 and a memory 1115, the processor being configured to execute any of the methods described above, such as Figure 2 , Figure 4 , Figure 5 , Figure 6 , Figure 7 and Figure 8 shown in the methods and so on. The memory may store, for example, a training set, sentence inputs to be processed, various intermediate data, and associated algorithms and so on.
[0239] The apparatus 1100 may include a network connection element 1125, which may include, for example, a network connection device that connects to other devices via a wired connection or a wireless connection. The wireless connection may be, for example, a WiFi connection, a Bluetooth connection, a 3G / 4G / 5G network connection, etc. For example, a module for obtaining data or outputting data may obtain data from various data sources and output the data to other devices via the network connection element. Inputs made by a user from other devices may also be received via the network connection element or data may be transmitted to other devices for display.
[0240] The apparatus may also optionally include other peripheral elements 1120, such as an input device (such as a keyboard, mouse), an output device (such as a display), etc. For example, in a method based on user input, the user may perform an input operation via the input device. Corresponding information may also be output to the user via the output device.
[0241] Each of these modules may communicate directly or indirectly with each other, for example, via one or more buses (such as bus 1105).
[0242] Moreover, the present application also discloses a computer-readable storage medium storing computer-executable instructions thereon, the computer-executable instructions, when executed by a processor, causing the processor to execute the methods of the various embodiments described herein.
[0243] In addition, the present application also discloses an apparatus including a processor and a memory storing computer-executable instructions, the computer-executable instructions, when executed by the processor, causing the processor to execute the methods of the various embodiments described herein.
[0244] In addition, the present application also discloses a system including an apparatus for implementing the methods of the various embodiments described herein.
[0245] It will be understood that the methods according to one or more embodiments of the present specification may be implemented in software, firmware, or a combination thereof.
[0246] It should be understood that the various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the description of the method embodiments.
[0247] It should be understood that the above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0248] It should be understood that an element described in the singular form herein or shown as only one in the drawings does not represent limiting the quantity of that element to one. Additionally, modules or elements described or shown as separate herein may be combined into a single module or element, and a module or element described or shown as a single herein may be split into multiple modules or elements.
[0249] It should also be understood that the terms and expressions used herein are only for description, and one or more embodiments of this specification should not be limited to these terms and expressions. Using these terms and expressions does not mean excluding any equivalent features of the illustration and description (or parts thereof). It should be recognized that various modifications that may exist should also be included within the scope of the claims. Other modifications, variations, and substitutions may also exist. Accordingly, the claims should be regarded as covering all such equivalents.
[0250] Similarly, it should be noted that although reference has been made to the current specific embodiments for description, those of ordinary skill in the art in this technical field should recognize that the above embodiments are only used to illustrate one or more embodiments of this specification, and various equivalent changes or substitutions can be made without departing from the spirit of the invention. Therefore, as long as the changes and modifications to the above embodiments are within the scope of the spirit of the invention, they will fall within the scope of the claims of this application.
Claims
1. A method for training an entity recognition model, comprising: Construct a training set, where the training set includes a plurality of training samples d = {S input , X NER , X TER}, where S input is a sentence, X NER is the sequence annotation label of the sentence, X TER is the metaphor entity label of the sentence, and the metaphor entity label is used to represent the metaphor entity of the sentence, where the metaphor entity is the entity that the sentence actually refers to but does not appear in the sentence; Performing training on the entity recognition model using the training set, where the entity recognition model is based on a pre-trained K-BERT model, and performing training on the entity recognition model includes: Inputting a training sample in the training set into the entity recognition model to obtain a sequence annotation prediction output and an entity matching prediction output of the sentence in the training sample; Determining a sequence annotation loss Loss_sequence of the sentence based on the sequence annotation prediction output of the sentence and the sequence annotation label of the sentence; Determining an entity matching loss Loss_match of the sentence at least partially based on the entity matching prediction output of the sentence and the metaphorical entity label of the sentence; Determining a total loss Loss_total of the entity recognition model, where the total loss is a weighted sum of the sequence annotation loss and the entity matching loss, that is: Loss_total = Loss_sequence + α * Loss_match, where α indicates the weight of the entity matching loss; and Iteratively performing training to minimize the total loss of the entity recognition model, thereby obtaining a trained entity recognition model.
2. The method according to claim 1, wherein a knowledge graph associated with the sentence is also input into the entity recognition model.
3. The method according to claim 2, wherein determining the entity matching loss Loss_match of the sentence includes: Using a graph neural network associated with the knowledge graph to generate a metaphorical entity vector of the sentence; Determining a vector distance between the entity matching prediction output of the sentence and the metaphorical entity vector; and Determining the entity matching loss Loss_match of the sentence, where the entity matching loss Loss_match of the sentence is proportional to the vector distance between the entity matching prediction output and the metaphorical entity vector.
4. The method according to claim 3, wherein determining the entity matching loss Loss_match of the sentence includes: Determining a random entity of the sentence at least based on the knowledge graph and the metaphorical entity label of the sentence, where the random entity is a randomly obtained entity; Using the graph neural network to generate a random entity vector of the sentence; and Determining the entity matching loss Loss_match of the sentence at least partially based on the entity matching prediction output of the sentence and the metaphorical entity label and random entity label of the sentence, where the entity matching loss Loss_match of the sentence is inversely proportional to the vector distance between the entity matching prediction output and the random entity vector.
5. The method according to claim 3, wherein determining the entity matching loss Loss_match of the sentence includes: Determining a knowledge-embedded entity of the sentence at least based on the knowledge graph and the metaphorical entity label of the sentence, where the knowledge-embedded entity is an entity different from the metaphorical entity embedded into the sentence based on the knowledge graph; Using the graph neural network to generate a knowledge-embedded entity vector of the sentence; Determine the entity matching loss Loss_match of the sentence at least partially based on the entity matching prediction output of the sentence, the metaphor entity label of the sentence, and the knowledge-embedded entity label of the sentence, where the entity matching loss Loss_match of the sentence is inversely proportional to the vector distance between the entity matching prediction output and the knowledge-embedded entity vector.
6. The method according to claim 3, wherein determining the entity matching loss Loss_match of the sentence includes: Determine at least based on the knowledge graph and the metaphor entity label of the sentence the knowledge-embedded entity and the random entity of the sentence, where the knowledge-embedded entity is an entity different from the metaphor entity embedded into the sentence based on the knowledge graph, and the random entity is a randomly obtained entity; Use the graph neural network to generate the random entity vector and the knowledge-embedded entity vector of the sentence; Determine the entity matching loss Loss_match of the sentence at least partially based on the entity matching prediction output of the sentence, the metaphor entity, the knowledge-embedded entity, and the random entity of the sentence, where the entity matching loss Loss_match of the sentence is directly proportional to the vector distance between the entity matching prediction output and the metaphor entity vector, inversely proportional to the vector distance between the entity matching prediction output and the knowledge-embedded entity vector, inversely proportional to the vector distance between the entity matching prediction output and the random entity vector, and inversely proportional to the vector distance between the knowledge-embedded entity vector and the random entity vector.
7. The method according to any one of claims 3-6, wherein the graph neural network is a graph convolutional network.
8. The method according to any one of claims 3-6, wherein when initializing the graph neural network, use the word embedding of the K-BERT model of the entity recognition model to generate the initial embedding representation of each node in the knowledge graph.
9. A method for entity recognition, comprising: Obtain a sentence to be processed; Use the entity recognition model trained by the method according to any one of claims 1-8 to process the sentence to be processed, and if an entity is obtained in the sequence annotation prediction output of the entity recognition model, output the obtained entity as the recognized entity; If the entity recognition model does not recognize an entity, output the entity matching prediction output of the entity recognition model; Perform a retrieval in the entity vector library using the entity matching prediction output to retrieve an entity vector that matches the entity matching prediction output in the entity vector library; and Use the entity corresponding to the retrieved entity vector as the recognized entity.
10. The method according to claim 9, wherein the entity vector library is obtained by performing vectorization on the knowledge graph associated with the input sentence.
11. The method according to claim 10, wherein the entity vector library is obtained by performing vectorization on the input sentence by a graph neural model, and the graph neural model is iteratively updated while training the entity recognition model.
12. The method according to claim 9, wherein performing retrieval in the entity vector library using the entity matching prediction output is implemented by the FAISS library.
13. A system for training an entity recognition model, comprising: A training set construction module for constructing a training set, where the training set includes a plurality of training samples d = {S input , X NER , X TER}, where S input is a sentence, X NER is the sequence annotation label of the sentence, X TER is the metaphor entity label of the sentence, and the metaphor entity label is used to represent the metaphor entity of the sentence, where the metaphor entity is the entity that the sentence actually refers to but does not appear in the sentence; and an entity recognition model training module for training the entity recognition model using the training set, the entity recognition model being based on a pre-trained K-BERT model, wherein the entity recognition model training module includes: a prediction module for inputting a training sample in the training set into the entity recognition model to obtain a sequence annotation prediction output and an entity matching prediction output for the sentence in the training sample, a loss calculation module for determining a sequence annotation loss Loss_sequence for the sentence based on the sequence annotation prediction output and the sequence annotation label of the sentence; determining an entity matching loss Loss_match for the sentence at least partially based on the entity matching prediction output and the metaphor entity label of the sentence; and determining a total loss Loss_total of the entity recognition model, the total loss being a weighted sum of the sequence annotation loss and the entity matching loss, i.e., Loss_total = Loss_sequence + α * Loss_match, where α indicates the weight of the entity matching loss; and an iterative training module for iteratively performing training to minimize the total loss of the entity recognition model, thereby obtaining a trained entity recognition model.
14. The system according to claim 13, wherein a knowledge graph associated with the sentence is also input into the entity recognition model.
15. The system according to claim 13, wherein the loss calculation module includes an entity vectorization module for vectorizing entities into entity vectors using a graph neural network.
16. The system according to claim 15, wherein when initializing the graph neural network, the word embeddings of the K-BERT model of the entity recognition model are used to generate an initial embedding representation for each node in the knowledge graph.
17. A system for entity recognition, comprising: a sentence acquisition module for acquiring a sentence to be processed; an entity recognition model for processing the sentence to be processed, the entity recognition model being trained using the method according to any one of claims 1-8, and if an entity is obtained from the sequence annotation prediction output of the entity recognition model, outputting the obtained entity as the recognized entity, and if the entity recognition model does not recognize an entity, outputting the entity matching prediction output of the entity recognition model; a retrieval module for performing retrieval in the entity vector library using the entity matching prediction output to retrieve an entity vector in the entity vector library that matches the entity matching prediction output, and taking the entity corresponding to the retrieved entity vector as the recognized entity.
18. An apparatus for training an entity recognition model, comprising: a memory; and a processor configured to execute the method according to any one of claims 1-8.
19. A device for performing entity recognition, comprising: a memory; and a processor configured to execute the method according to any one of claims 9-12.
20. A computer-readable storage medium storing instructions that, when executed by a computer, cause the computer to execute the method according to any one of claims 1-12.
Citation Information
Patent Citations
Chinese statement metaphor recognition system
CN111859934A
Metaphor calculation and device based on knowledge graph representation learning
CN113157932A