Hierarchical Category Named Entity Recognition Model Design Method Based on Multi-task Learning

By designing a hierarchical category named entity recognition model based on multi-task learning, the problem of entity inconsistency and category relationship conflicts in the existing model when it is difficult to identify named entities and faces complex scenarios of multi-level fine-grained categories, and a more efficient naming entity recognition effect is achieved.

CN114881032BActive Publication Date: 2025-05-06BEIJING INST OF COMP TECH & APPL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210462583.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-28
Publication Date
2025-05-06
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

It is difficult for existing named entity recognition models to identify multiple categories of named entities at the same time, and when faced with complex scenarios of multi-level fine-grained categories, it is easy to have entity inconsistencies and category relationship conflicts.

Method used

A hierarchical category named entity recognition model based on multi-task learning is designed. By treating named entity recognition at different levels as multiple tasks, using one model to train multiple tasks, using a multi-task learning mechanism to simultaneously predict named entity recognition between multiple levels, and two information transmission mechanisms are designed to transmit identification information between tasks at different levels.

Benefits of technology

This model can effectively identify multiple categories of named entities, solve the problem of hierarchical category naming entity recognition, improve the recognition effect of the model, and avoid conflicts in output results between different levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114881032B_ABST
    Figure CN114881032B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for designing a hierarchical category named entity recognition model based on multi-task learning, and belongs to the technical field of natural language processing. The present invention enables the model to simultaneously recognize multiple categories of named entities by adding modeling of category relationships to the named entity recognition model. At the same time, the present invention proposes a model based on multi-task learning to solve the problem of named entity recognition with hierarchical categories. The model uses a multi-task learning mechanism to simultaneously learn multiple levels of named entity recognition tasks, and these tasks share the same encoding layer, so that the encoding vectors learned by the encoding layer can simultaneously adapt to multiple levels of named entity recognition instead of overfitting to a single level. Finally, two information transmission mechanisms are designed to transmit recognition information between different levels to improve the recognition effect of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a method for designing a hierarchical category named entity recognition model based on multi-task learning. Background Art

[0002] Named entity recognition is one of the basic tasks in the field of natural language processing. Its goal is to identify meaningful named entities such as names of people and places in sentences. In the existing research on named entity recognition, most of them are only aimed at coarse-grained named entities. The number of categories specified in the data set is mostly less than 10, and the relationship between categories is not considered. However, in reality, only coarse-grained classification of named entities is far from meeting actual needs. Named entities are usually composed of multiple categories with different granularities, and a lot of key information is contained in the fine-grained dimensions. The more levels of categories of named entities and the finer the granularity, the richer the information given by the named entity recognition results. Therefore, it is of great practical significance to study the design method of named entity recognition models for hierarchical categories.

[0003] Named entity recognition for simple scenes cannot adapt to complex scenarios with multi-level fine-grained categories. If multiple named entity recognition models for simple scenes are used to identify categories at different levels, it will inevitably lead to two phenomena: inconsistency between entities at different levels and conflict between parent-child categories of entities. At the same time, multiple tasks work independently of each other, and there is no analysis of correlation between different model categories. If the named entity recognition model for simple scenes is used to directly identify the finest-grained categories, and the identified fine-grained categories are output as coarse-grained categories of entities, there may be a problem of insufficient training due to insufficient fine-grained entity data. At the same time, this method does not utilize coarse-grained category information and does not model the relationship between categories. At present, the mainstream method in the field of multi-level named entity recognition is a two-stage pipeline method. The first stage identifies the boundaries of the entity, and the second stage determines the categories of each level of the entity. When classifying, the idea of ​​classification from coarse to fine is often adopted. This method has two disadvantages. First, the pipeline method will have the problem of error accumulation. Errors in the previous task will lead to errors in subsequent tasks. Second, the pipeline method does not fully utilize the information in the data set, which will lead to performance loss. Because entity types also help segment entity boundaries, and fine-grained category information also helps coarse-grained entity classification. In summary, the core challenge of hierarchical named entity recognition is how to simultaneously use multi-level information to identify and classify named entities, and avoid conflicts between output results at different levels. Summary of the invention

[0004] 1. Technical issues to be resolved

[0005] The technical problem to be solved by the present invention is: how to design a named entity recognition model so that the model can simultaneously recognize multiple categories of named entities, solve the problem of named entity recognition with hierarchical categories, and improve the recognition effect of the model.

[0006] (II) Technical solution

[0007] In order to solve the above technical problems, the present invention provides a method for designing a hierarchical category named entity recognition model based on multi-task learning. In the method, the designed hierarchical category named entity recognition model based on multi-task learning is named MTBP. When designing the model, named entity recognition at different levels is regarded as multiple tasks, one model is used to train multiple tasks, and a multi-task learning mechanism is used to simultaneously perform named entity recognition prediction between multiple levels. Encoders are shared between multiple tasks, wherein two different information transfer mechanisms are designed to transfer recognition information between tasks at different levels. The first one adopts a top-down information transfer order, first predicts the top-level class, and then passes the top-level information to the next layer for prediction, which is called MTBP-T, and the second one is a bottom-up transfer order, which is called MTBP-B.

[0008] Preferably, in this method, the design principle of the MTBP-T model is: the model output of the coarse-grained category is passed as information to the next layer to assist fine-grained named entity recognition; the MTBP-T model uses BERT as an encoder, and the input characters are passed through the encoder to obtain a preliminary word vector, and the low-level representation vector is spliced ​​by the BERT output result and the label prediction result of the previous layer.

[0009] Preferably, in the method, the MTBP-T model is designed as an MBTP-T model structure for named entity recognition tasks with a three-layer category structure:

[0010] The first layer uses the output of BERT as the embedding vector, and the calculation process is shown in the following formula:

[0011] E0=BERT(X)

[0012] After the second layer, the concatenation of the embedding of the previous layer and the recognition result of the previous layer is used as the embedding vector:

[0013] E k =Concat(E k-1 , label k-1 )

[0014] Where E0 represents the BERT output, which has a shape of m×l, l is the number of characters in the input sequence, and m is the size of the BERT word vector; E kRepresents the input character vector used in each layer, 0<k≤n, n is the number of levels of categories; label k-1 It is the extraction result output by the previous model;

[0015] After obtaining each layer of word vectors, a probability matrix is ​​obtained as a prediction matrix through a linear layer and a sigmoid activation layer. Each column in the probability matrix maps a word in the input sequence, and every two rows in the probability matrix map a category. The first row of the two rows corresponds to the probability that the word is the beginning of the entity of this category, and the second row is the probability of the end. The specific calculation process is shown in the following formula:

[0016] pred j =sigmoid(W j E j )

[0017] Among them, E j represents the vector representation of the jth word, pred j That is, the probability that the predicted characters are the start and end positions of the entity, Among them C j Represents the number of categories in the j-th layer.

[0018] Preferably, in this method, the MTBP-B model is designed as: a named entity recognition model based on multi-task learning that transmits information from the bottom up, and its design principle is: due to the subordinate relationship between categories, predicting a subclass entity in entity prediction has actually predicted the parent class entity, and the low-level entity output predicted by the model contains information about the parent class distribution, so the predicted distribution of the parent class can be obtained from the predicted distribution of the subclass.

[0019] Preferably, in the method, the MTBP-B model is designed as an entity-oriented MTBP-B model with a three-layer structure;

[0020] The MTBP-B model also uses BERT as an encoder to encode the input sequence into a character vector, as shown in the following formula:

[0021] E=Bert(X)

[0022] E is the vector of input characters, where the MTBP-B model directly uses the character vector for the finest-grained named entity prediction. The prediction process is still to pass the character vector through two fully connected layers and a sigmoid activation layer to obtain a matrix indicating whether the character is the beginning and end of a class of entities. The calculation process is shown in the following formula:

[0023] pred n =sigmoid(W n E)

[0024] Among them, W n is the parameter of the fully connected layer. The MTBP-B model uses the low-level prediction results to obtain the high-level prediction results. It aggregates the prediction data of the subclasses of the same parent class to obtain the prediction data of the parent class. For the starting matrix, the specific transformation process is: the two prediction matrices of the subclasses are divided by category to form several small matrices, and the type of row mapping in each matrix has the same parent class; take the maximum value of the column of each small matrix to form a new row, and then splice these rows to obtain a new matrix. This new matrix is ​​the prediction matrix of the parent class. This transformation process is called levelmax operation. The overall process is shown in the following formula:

[0025] pred j =levelmax(pred j+1 )

[0026] Where 0≤j<n.

[0027] Preferably, MTBP-B and MTBP-T use a single model to simultaneously perform entity recognition at multiple levels, which requires the use of a multi-task learning paradigm. Therefore, a multi-task loss function is introduced into the loss to perform learning of multiple tasks. The multi-task loss function is designed as follows:

[0028] Each single task of named entity recognition at each level can be decomposed into multiple binary classification problems. The cross entropy loss function is used as the loss function for the binary classification problem. The loss function is:

[0029] loss 二分类 = -tlogp-(1-t)log(1-p)

[0030] Where t∈{0,1} is the label and p is an item in the matrix output by the model. In this way, the loss function of named entity recognition at a single level is:

[0031] loss 单任务 =∑loss 二分类

[0032] When adding tasks at multiple levels, because the number of categories at the lower level is higher than that at the higher level, the corresponding loss function value is greater than the loss function value of the higher level task, so a hyperparameter 0≤λ is set for each task i ≤1, 1≤i≤n, to adjust the importance of the task, and limit the sum of all hyperparameters to 1. The total multi-task loss function is shown in the following formula:

[0033]

[0034]

[0035] Preferably, in this method, the recognition result is constructed by predicting the matrix, taking a threshold z, 0<z<1, setting the values ​​in the prediction matrix greater than the threshold to 1, and setting the values ​​less than the threshold to 0, so as to obtain a label matrix label with the same shape j , as shown in the following formula:

[0036]

[0037] labelj is the label matrix predicted by the jth layer. The starting and ending positions of the predicted entities can be obtained through the label values ​​of the label matrix, so as to extract the named entities of the hierarchical category as the final output result of the hierarchical category named entity recognition model based on multi-task learning.

[0038] Preferably, in the training phase, a teacher-supervised learning method is used to directly use the fine-grained category information in the training set to construct the correct label matrix for guidance, that is, the character label data label used in training j The correct label in the training set rather than the output of the previous layer is used to accelerate the convergence of the hierarchical category named entity recognition model based on multi-task learning.

[0039] Preferably, the output of a high-level category among multiple outputs of the multi-task based hierarchical category named entity recognition model is taken as the true output result.

[0040] The present invention also provides an application of the method in the technical field of natural language processing.

[0041] (III) Beneficial effects

[0042] The present invention adds modeling of category relationships to the named entity recognition model, so that the model can simultaneously recognize multiple categories of named entities. At the same time, the present invention proposes a model based on multi-task learning to solve the problem of named entity recognition with hierarchical categories. The model uses a multi-task learning mechanism to simultaneously learn multiple levels of named entity recognition tasks. These tasks share the same encoding layer, so that the encoding vectors learned by the encoding layer can simultaneously adapt to multiple levels of named entity recognition instead of overfitting to a single level. Finally, two information transmission mechanisms are designed to transmit recognition information between different levels to improve the recognition effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is the MTBP-T model architecture diagram of the present invention;

[0044] Figure 2 This is the MTBP-B model architecture diagram of the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, content, and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below in conjunction with the accompanying drawings and examples.

[0046] In the present invention, the named entity recognition model based on the hierarchical category of multi-task learning is named MTBP (Multi-Task-BERT-Pointer). The basic idea is to regard named entity recognition at different levels as multiple tasks, use one model to train multiple tasks, use a multi-task learning mechanism to simultaneously perform named entity recognition prediction between multiple levels, and share the encoding layer (BERT encoder) between multiple tasks. There is a great correlation between named entity recognition tasks at different levels. Multi-task learning can avoid overfitting to a certain task and reduce the probability of falling into a local minimum. Since tasks at different levels have a strong correlation, multi-task learning can help entity recognition at each level have better performance. The present invention also designs two different information transfer mechanisms to transfer recognition information between tasks at different levels. The first structure adopts a top-down (Top-down) information transfer order, first predicts the top-level class, and then passes the top-level information to the next layer for prediction. It is called MTBP-T in the present invention, and the second structure is a bottom-up (Bottom-up) transfer order. It is called MTBP-B in the present invention. One of them is used when used, and is introduced below.

[0047] 1. MTBP-T

[0048] The main motivation of the MTBP-T model is that fine-grained entity recognition will be more accurate after obtaining coarse-grained entity recognition information, so the model output of the coarse-grained category is passed as information to the next layer to assist fine-grained named entity recognition. The overall architecture of the MTBP-T model is as follows Figure 1 shown.

[0049] The MTBP-T model uses BERT as an encoder, and the input characters are passed through the encoder to obtain preliminary word vectors. The low-level representation vector is concatenated by the BERT output result and the label prediction result of the previous layer. Figure 1 The MBTP-T model structure for named entity recognition tasks with a three-layer category structure is shown:

[0050] The first layer uses the output of BERT as the embedding vector, and the calculation process is shown in the following formula:

[0051] E0=BERT(X)\*MERGEFORMAT (1)

[0052] After the second layer, the concatenation of the embedding of the previous layer and the recognition result of the previous layer is used as the embedding vector:

[0053] E k =Concat(E k-1 , label k-1 )\*MERGEFORMAT (2)

[0054] Among them, E0 represents the BERT output, whose shape is m×l, l is the number of characters in the input sequence, and m is the size of the BERT word vector, which is usually 768; E k Represents the input character vector used in each layer, 0<k≤n, n is the number of levels of categories; label k-1 It is the extraction result output by the previous model, and the calculation method is listed in the formula.

[0055] After obtaining each layer of word vectors, a probability matrix is ​​obtained through a linear layer and a sigmoid activation layer. Each column in the probability matrix maps a word in the input sequence, and every two rows in the probability matrix map a category. The first row of the two rows corresponds to the probability that the word is the beginning of the entity of this category, and the second row is the probability of the end. The specific calculation process is shown in the following formula:

[0056] pred j =sigmoid(W j E j )\*MERGEFORMAT (3)

[0057] Among them, E j represents the vector representation of the jth word, pred j That is, the probability that the predicted characters are the start and end positions of the entity, Among them C j Represents the number of categories in the j-th layer.

[0058] 2. MTBP-B

[0059] The MTBP-B model is a named entity recognition model based on multi-task learning that transfers information from bottom to top. Its motivation is that due to the subordinate relationship between categories, predicting a sub-category entity in entity prediction actually predicts the parent-category entity. The low-level entity output predicted by the model contains information about the parent-category distribution, so the predicted distribution of the parent-category can be obtained from the predicted distribution of the sub-category. Figure 2 A MTBP-B model with a three-layer structure for entity categories is presented.

[0060] Similar to the MTBP-T model, the MTBP-B model also uses BERT as an encoder to encode the input sequence into a character vector. As shown in the following formula:

[0061] E=Bert(X)\*MERGEFORMAT (4)

[0062] E is the vector of input characters. The difference is that the MTBP-B model directly uses character vectors for the finest-grained named entity prediction. The prediction process is still to pass the character vector through two fully connected layers and a sigmoid activation layer to obtain a matrix indicating whether the character is the beginning and end of a certain type of entity. The calculation process is shown in the following formula:

[0063] pred n =sigmoid(W n E)\*MERGEFORMAT (5)

[0064] Among them, W n is the parameter of the fully connected layer. n represents the nth layer of named entity recognition. The MTBP-B model uses low-level prediction results to obtain high-level prediction results. The specific idea is to aggregate the prediction data of subclasses of the same parent class to obtain the prediction data of the parent class. Taking the starting matrix as an example, the specific transformation process is: divide the two prediction matrices of the subclasses by category to form several small matrices, and the types of row mappings in each matrix have the same parent class; take the maximum value of the column of each small matrix to form a new row, and then splice these rows to obtain a new matrix. This new matrix is ​​the prediction matrix of the parent class. The above process is called the levelmax operation, and the overall process is shown in the following formula:

[0065] pred j =levelmax(pred j+1 )

[0066] Where 0≤j<n.

[0067] The multi-task loss function is introduced below:

[0068] MTBP-B and MTBP-T use a single model to perform entity recognition at multiple levels at the same time. It is necessary to use the multi-task learning paradigm, that is, to introduce multi-task loss functions in the loss to learn multiple tasks. Each single task of named entity recognition at each level can be decomposed into multiple binary classification problems. The cross entropy loss function is used as the loss function for the binary classification problem. The loss function is:

[0069] loss 二分类 =-tlogp-(1-t)log(1-p)\*MERGEFORMAT (6)

[0070] Where t∈{0,1} is the label and p is an item in the matrix output by the model. In this way, the loss function of named entity recognition at a single level is:

[0071] loss 单任务 =∑loss 二分类 \*MERGEFORMAT (7)

[0072] When multiple levels of tasks are added together, because the number of categories in the lower level is higher than that in the higher level, the corresponding loss function value is greater than the loss function value of the higher level task. Therefore, a hyperparameter 0≤λ is set for each task. i ≤1, 1≤i≤n, to adjust the importance of the task and limit the sum of all hyperparameters to 1. The overall multi-task loss function is shown in the following formula:

[0073]

[0074]

[0075] Finally, the model prediction output is introduced:

[0076] Finally, the recognition result can be constructed through the prediction matrix. Take a threshold z, 0<z<1, set the values ​​in the matrix greater than the threshold to 1, and the values ​​less than the threshold to 0, and you can get a label matrix label with the same shape j , as shown in the following formula:

[0077]

[0078] label j That is, the label matrix predicted by the jth layer. The starting and ending positions of the predicted entities can be obtained through the label values ​​of the label matrix, so as to extract the named entities of the hierarchical category as the final output result of the hierarchical category named entity recognition model based on multi-task learning designed by the present invention.

[0079] During the training phase, when the model is not fully trained, the label output given by the upper layer may contain a lot of errors and cannot play a guiding role. Therefore, the teacher supervised learning method is used during training to directly use the fine-grained category information in the training set to construct the correct label matrix for guidance, that is, the character label data label used during training j The correct labels in the training set are not the output of the previous layer. This can help accelerate the convergence of the model.

[0080] The multi-task named entity recognition model will have multiple outputs, and each level will have corresponding outputs. However, there is a lack of hard constraints between the outputs of different levels, and the output results of different levels may have the following inconsistencies: 1. Entity inconsistency. Specifically, the output entities of different levels are not completely consistent, and the entity sets of different level outputs are not completely consistent. 2. Inconsistency in the parent-child relationship of entity categories. There is no parent-child relationship between the classification results of the same entity at different levels. At the same time, since each low-level category has and only has one father. So when the model recognizes a low-level entity, it actually gives an output of a high-level category, and this output has no entity inconsistency and parent-child relationship conflict with the low-level classification result. In other words, the latter (the output of the high-level category) is more suitable as the true output of the hierarchical category named entity recognition model based on multi-task learning.

[0081] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for designing a hierarchical category named entity recognition model based on multi-task learning, characterized in that: In this method, the designed named entity recognition model based on multi-task learning hierarchical categories is named MTBP. When designing this model, named entity recognition at different levels is regarded as multiple tasks, and one model is used to train multiple tasks. A multi-task learning mechanism is used to predict named entity recognition between multiple levels, and encoders are shared between multiple tasks. Two different information transfer mechanisms are designed to transfer recognition information between tasks at different levels. The first one adopts a top-down information transfer order, first predicting the top-level class, and then passing the top-level information to the next layer for prediction. It is called MTBP-T, and the second one is a bottom-up transfer order, which is called MTBP-B. In this method, the design principle of the MTBP-T model is: the model output of the coarse-grained category is passed as information to the next layer to assist the fine-grained named entity recognition; The MTBP-T model uses BERT as an encoder. The input characters are passed through the encoder to obtain preliminary word vectors. The low-level representation vectors are concatenated by the BERT output results and the label prediction results of the previous layer. In this method, the MTBP-T model is designed as an MBTP-T model structure for named entity recognition tasks with a three-layer category structure: The first layer uses the output of BERT as the embedding vector, and the calculation process is shown in the following formula: E0=BERT(X) After the second layer, the concatenation of the embedding of the previous layer and the recognition result of the previous layer is used as the embedding vector: AND k =Concat(E k-1 ,label k-1 ) Among them, E0 represents the BERT output, whose shape is m×l, where l is the number of characters in the input sequence and m is the dimension of the word vectors of BERT; E k represents the input character vectors used in each layer, where 0 < k ≤ N and N is the number of levels of categories; label k-1 is the extraction result output by the previous layer of the model; After obtaining each layer of word vectors, a probability matrix is ​​obtained as a prediction matrix through a linear layer and a sigmoid activation layer. Each column in the probability matrix maps a word in the input sequence, and every two rows in the probability matrix map a category. The first row of the two rows corresponds to the probability that the word is the beginning of the entity of this category, and the second row is the probability of the end. The specific calculation process is shown in the following formula: pred j =sigmoid(W j E j ) Among them, E j represents the vector representation of the jth word, pred j That is, the probability that the predicted characters are the start and end positions of the entity, Among them C j represents the number of categories in the jth layer; In this method, the MTBP-B model is designed as a named entity recognition model based on multi-task learning that transfers information from bottom to top. The design principle is that due to the subordinate relationship between categories, predicting a sub-category entity in entity prediction actually predicts the parent-category entity. The low-level entity output predicted by the model contains information about the parent-category distribution, so the predicted distribution of the parent-category can be obtained from the predicted distribution of the sub-category. In this method, the MTBP-B model is designed as an entity-oriented MTBP-B model with a three-layer structure; The MTBP-B model also uses BERT as an encoder to encode the input sequence into a character vector, as shown in the following formula: E=Bert(X) E is the vector of input characters, where the MTBP-B model directly uses the character vector for the finest-grained named entity prediction. The prediction process is still to pass the character vector through two fully connected layers and a sigmoid activation layer to obtain a matrix indicating whether the character is the beginning and end of a class of entities. The calculation process is shown in the following formula: pred n =sigmoid(W n E) Among them, W n is the parameter of the fully connected layer, n represents the nth layer of named entity recognition, and the MTBP-B model uses the low-level prediction results to obtain the high-level prediction results. It aggregates the prediction data of the subclasses of the same parent class to obtain the prediction data of the parent class. For the starting matrix, the specific transformation process is: the two prediction matrices of the subclasses are divided by category to form several small matrices, and the types of row mappings in each matrix have the same parent class; the maximum value of the column of each small matrix is ​​taken to form a new row, and then these rows are spliced ​​to obtain a new matrix. This new matrix is ​​the prediction matrix of the parent class. This transformation process is called the levelmax operation, and the overall process is shown in the following formula: pred j =levelmax(pred j+1 ) Where 0≤j<n.

2. The method according to claim 1, characterized in that MTBP-B and MTBP-T use a single model to simultaneously perform entity recognition at multiple levels, which requires the use of a multi-task learning paradigm. Therefore, a multi-task loss function is introduced into the loss to learn multiple tasks. The multi-task loss function is designed as follows: Each single task of named entity recognition at each level can be decomposed into multiple binary classification problems. The cross entropy loss function is used as the loss function for the binary classification problem. The loss function is: loss 二分类 =-tlogp-(1-t)log(1-p) Where t∈{0,1} is the label and p is an item in the matrix output by the model. In this way, the loss function of named entity recognition at a single level is: loss 单任务 =∑loss 二分类 When adding tasks at multiple levels, because the number of categories at the lower level is higher than that at the higher level, the corresponding loss function value is greater than the loss function value of the higher level task, so a hyperparameter 0≤λ is set for each task i ≤1, 1≤i≤n, to adjust the importance of the task, and limit the sum of all hyperparameters to 1. The total multi-task loss function is shown in the following formula:

3. The method according to claim 2, characterized in that In this method, the recognition result is constructed by predicting the matrix, taking a threshold z, 0<z<1, setting the values ​​in the prediction matrix greater than the threshold to 1, and the values ​​less than the threshold to 0, and the label matrix label with the same shape can be obtained. j , as shown in the following formula: label j That is, it is the label matrix predicted by the jth layer. The starting and ending positions of the predicted entities can be obtained through the label values ​​of the label matrix, so as to extract the named entities of the hierarchical category as the final output result of the hierarchical category named entity recognition model based on multi-task learning.

4. The method according to claim 3, characterized in that In the training phase, the teacher supervised learning method is used to directly use the fine-grained category information in the training set to construct the correct label matrix for guidance, that is, the character label data label used in training j The correct label in the training set rather than the output of the previous layer is used to accelerate the convergence of the hierarchical category named entity recognition model based on multi-task learning.

5. The method according to any one of claims 1 to 4, characterized in that The output of the high-level category among multiple outputs of the multi-task based hierarchical category named entity recognition model is taken as the true output result.

Citation Information

Patent Citations

  • Hierarchical multi-label text classification method and system

    CN113590815A

  • Clustering method and device based on comprehensive similarity

    CN114118310A

  • Entity identification method and device, server and storage medium

    CN114218945A