Named entity recognition method based on prompt learning and label perception

By adopting a method based on prompt learning and tag perception in the named entity recognition task, combining the BERT model and the bidirectional LSTM network, the interaction between prompt slots and sentences is achieved, which solves the problems of large time overhead and template sensitivity in the named entity recognition task of the existing prompt learning model, and improves the entity extraction ability of the model, especially in a small sample environment.

CN119990132AActive Publication Date: 2025-05-13CHONGQING UNIV OF POSTS & TELECOMM
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510084440.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The existing prompt learning model has problems such as high time overhead, sensitive templates, independent prompts, poor performance in the environment of few samples, and single task in named entity recognition tasks.

Method used

The named entity recognition method based on prompt learning and tag perception is adopted. By obtaining the text to be identified and preprocessed, the trained named entity recognition model is used for processing. Combining the BERT model and the bidirectional LSTM network, the self-attention mechanism is used to achieve interaction between the prompt slots and sentences, and the entity extraction ability of the model is improved through an encouraging nearest neighbor template filling mechanism.

Benefits of technology

While increasing the model inference speed, this method reduces the size of the prompt template in the input, enhances the interaction between different prompt slots and within the same prompt slot, and improves the model's entity extraction ability, especially in a small sample environment, deeper semantics can be learned.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990132A_ABST
    Figure CN119990132A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of natural language processing, and particularly relates to a named entity recognition method based on prompt learning and tag perception. Comprising the following steps: splicing a plurality of prompt templates in front of each sentence in pre-processed named entity identification data; adding a prompt mask to the processed sentence and inputting the sentence into a BERT model to obtain semantic representation of a prompt slot and the sentence; inputting the entity label into the BERT model to obtain a label semantic representation; inputting the semantic representation of the sentence into a bidirectional LSTM and processing the semantic representation of the prompt slot by using a self-attention mechanism to further obtain a sentence context semantic matrix and a prompt slot deep semantic matrix; calculating classification slot probability distribution according to the prompt slot deep semantic matrix and the label semantic representation to realize entity classification; calculating positioning groove probability distribution according to the sentence context semantic matrix and the prompt groove deep semantic matrix to realize entity positioning; according to the method, the entity extraction capability of the model can be remarkably improved, and time and space overhead is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a named entity recognition method based on prompt learning and label perception. Background Art

[0002] Named Entity Recognition (NER) is one of the most basic tasks in NLP. This task aims to identify and classify nouns with special meanings from texts of predefined semantic types, such as names of people, places, and organization names.

[0003] In the early days, NER was mainly implemented by linguists manually constructing rule templates based on language knowledge characteristics. Supervised learning is the most widely used method in statistical machine learning, which is mainly manifested in selecting appropriate features based on text information as the basis for entity classification. In recent years, NER models based on deep learning have become mainstream models. The release of models such as Transformer and BERT has opened a new era of pre-trained language models. The pre-training + fine-tuning method has achieved excellent performance. Recently, prompt learning as the fourth paradigm of NLP has also attracted the interest of many researchers, and the pre-training + prompt + prediction method has gradually emerged. The classic solution to apply the idea of ​​prompt learning to NER is to iteratively obtain all candidate entity fragments through the N-gram method, and then splice them with the designed manual template, and use the BART model to score each fragment to predict the entity category. This method is significantly better than the traditional sequence labeling method and the distance-based few-sample NER method in cross-domain and few-sample scenarios. A corresponding method is to construct a prompt for each entity type and then guide the model to locate a specific type of entity. This solution avoids the disadvantage of too many iterations of TemplateNER and is suitable for nested NER. One method uses a soft template to transform the original sequence labeling task into a sequence-to-sequence generation task. By incorporating prompt information into the self-attention mechanism, the model achieves better results. During training, the parameters of the pre-trained model are fixed, and only the prompt-related parameters are optimized, making the model more flexible and lightweight. The PromptNER method uses a dual-slot multi-prompt template and a dynamic template filling mechanism to fill all prompt templates in parallel, achieving the effect of full entity output in one round of prompts, while ensuring performance and improving operating efficiency.

[0004] However, most of the above-mentioned prompt learning models have problems such as high time cost, template sensitivity, independence between prompts, poor performance in a few-sample environment, and single-task targeting. For example, TemplateNER, which iteratively generates spans, requires N(N+1) / 2 rounds of prompts (N is the total number of tokens in a sentence). Each round of prompts in the above models is independent of each other, ignoring the potential relationship between prompts. For example, the PromptNER method reduces the time complexity by giving prompts first and then matching, but the positioning and classification slots designed are functionally separated from each other, lacking cross-task interaction, and are prone to losing a lot of useful prior information. In addition, the entity repetition mechanism is relatively rough, which can easily cause template confusion problems. The use of double slots also implicitly increases the limit on the length of the text. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention proposes a named entity recognition method based on prompt learning and label perception, the method comprising: obtaining a text to be recognized and preprocessing it, inputting the preprocessed text into a trained named entity recognition model for processing, and obtaining a named entity recognition result;

[0006] The training process of the named entity recognition model includes:

[0007] S1: Obtain a named entity recognition dataset and preprocess it to obtain preprocessed named entity recognition data;

[0008] S2: concatenate multiple prompt templates before each sentence in the named entity recognition data, and unify the length of the sentences after concatenating the prompt templates;

[0009] S3: Add a prompt mask to the sentence after the prompt template is concatenated and input into the BERT model to obtain the prompt slot semantic representation and sentence semantic representation;

[0010] S4: Input the entity label into the BERT model to obtain the label semantic representation;

[0011] S5: Input the sentence semantic representation into the bidirectional LSTM network to obtain the sentence context semantic matrix; use the self-attention mechanism to process the cue slot semantic representation according to the sentence context semantic matrix to obtain the cue slot deep semantic matrix;

[0012] S6: Map the deep semantic matrix of the prompt slot into the classification slot encoding matrix through a linear layer, and calculate the probability distribution of the classification slot according to the classification slot encoding matrix and the label semantic representation;

[0013] S7: Calculate the probability distribution of the positioning slot based on the deep semantic matrix of the prompt slot and the sentence context semantic matrix;

[0014] S8: Calculate the minimum matching loss between the prompt slot and the entity according to the probability distribution of the classification slot and the probability distribution of the positioning slot, and obtain the optimal matching solution;

[0015] S9: Calculate the total model loss under the optimal matching scheme, adjust the model parameters according to the total model loss, and obtain a trained named entity recognition model.

[0016] Preferably, the semantic representation of the prompt slot is processed using a self-attention mechanism and represented as:

[0017]

[0018] Among them, H E represents the original hint slot semantic representation, H X Represents the sentence context semantic matrix, H S represents the semantic representation of the hint slot after the interaction between slots, H D It means that the semantic representation of the cue slot takes into account both the interaction between slots and the interaction between slots and sentences, that is, the deep semantic representation of the cue slot; and denote the key, query and value matrices of the hint slot respectively, and W k , W q and W v denote the first, second and third weight matrices respectively; and Represents the key and value matrix of the sentence; L represents the slot identifier embedding matrix, h represents the dimension of the hidden vector, and Softmax represents the activation function.

[0019] Preferably, the process of calculating the probability distribution of classification slots includes: mapping the deep semantic matrix of the prompt slots to the classification slot encoding matrix through linear mapping; performing a dot product operation on the classification slot encoding matrix and the label semantic representation to obtain a matching matrix between the prompt slots and the entity categories; and decoding the matching matrix between the prompt slots and the entity categories through a linear layer to obtain the probability distribution of the classification slots.

[0020] Preferably, the process of calculating the probability distribution of the positioning slots includes:

[0021] The deep semantic matrix of the cue slot is linearly mapped to obtain the positioning slot encoding matrix;

[0022] Add the positioning slot encoding matrix to the sentence context semantic matrix after linear mapping to obtain the first matrix;

[0023] Add the classification slot encoding matrix to the sentence context semantic matrix after linear mapping to obtain the second matrix;

[0024] The first matrix is ​​added to the second matrix and decoded again through a linear layer to obtain the probability distribution of the positioning slots.

[0025] Preferably, the process of calculating the matching loss between the prompt slot and the entity includes:

[0026] Calculate the sampling weight of each entity according to the classification slot probability distribution and the positioning slot probability distribution;

[0027] The entities are repeatedly sampled according to the sampling weights to fill all the prompt slots, and the matching schemes between all the prompt slots and the entities are obtained;

[0028] Calculate the minimum matching loss among all matching schemes between prompt slots and entities, and select the matching scheme corresponding to the minimum matching loss as the optimal matching scheme.

[0029] Preferably, the formula for calculating the sampling weight of each entity is:

[0030]

[0031] Among them, MC j Represents the total number of prompt slots matching the jth real entity in the sentence, Weight j represents the sampling weight corresponding to the jth real entity in the sentence, M is the number of prompt slots, and A is the total number of entities originally in a sentence; Indicates the matching classification, left boundary, and right boundary value obtained when the hint slot k matches the jth real entity; max k It represents the real entity index that best matches the hint slot k, and equal is the marking function.

[0032] Preferably, the formula for calculating the minimum matching loss among all matching schemes between hint slots and entities is:

[0033]

[0034] Among them, σ * represents the optimal matching solution, M represents the number of prompt slots, represents the indicator function, represents the confidence probability value corresponding to the true category output by the prompt slot i under the σth arrangement, It represents the confidence probability value corresponding to the true left and right boundaries output by the prompt slot i under the σ-th arrangement, and argmin represents the minimum value function.

[0035] Preferably, the total model loss is the weighted sum of the positioning loss and the classification loss, where:

[0036] Classification loss L type It is expressed as:

[0037]

[0038] Positioning loss L pOsitionIt is expressed as:

[0039]

[0040] in, The optimal matching arrangement σ * The confidence probability value corresponding to the true category output by the next prompt slot i, represents the confidence probability value corresponding to the true left and right boundaries output by the prompt slot i under the optimal matching arrangement, represents the indicator function, and M represents the number of prompt slots.

[0041] The beneficial effects of the present invention are as follows: the present invention uses a single-slot dual-function template as a framework, which avoids the time overhead problem commonly existing in conventional prompt learning in terms of time, and reduces the size of the prompt template in the input in terms of space. The optimization of these two dimensions can improve the reasoning speed of the model; the present invention realizes the interaction between different prompt slots and within the same prompt slot through the interaction layer and cross-task representation fusion, integrates the semantics of the two tasks of positioning and classification, and realizes the mutual assistance of positioning and classification through prior knowledge to improve the entity extraction ability of the model; the present invention proposes an encouraging nearest neighbor template filling mechanism, which increases the number of entities extracted by the model while alleviating the template confusion problem, and can alleviate the problem of misjudgment of loss in the later stage of training to a certain extent; the present invention introduces the natural semantics of the label, on the one hand, implicitly introduces external knowledge, so that the model can learn deeper semantics in the scenario of few samples, and on the other hand, through contrast learning, the distributed semantic representation generated by the model can be as far away from non-entity categories as possible. In a few-sample environment, the accuracy of the model is improved with the help of external knowledge, which can alleviate the problem of insufficient learning of the model in a few-sample environment to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a training flow chart of the named entity recognition model in the present invention;

[0043] Figure 2 This is a structural diagram of the named entity recognition model in the present invention. DETAILED DESCRIPTION

[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] The present invention proposes a named entity recognition method based on prompt learning and label perception, and the method comprises the following contents:

[0046] The text to be recognized is obtained and preprocessed, and the preprocessed text is input into the trained named entity recognition model for processing to obtain the named entity recognition result.

[0047] like Figure 1 , Figure 2 As shown in Figure 1, the training process of the named entity recognition model includes:

[0048] S1: Obtain a named entity recognition dataset and preprocess it to obtain preprocessed named entity recognition data.

[0049] The process of preprocessing the data includes: removing stop words in the original data set, then modifying the traditional sequence annotation data set format to span representation, representing a sentence and the entities contained in it as an object, and representing the entire data set in the form of an object array. The sentence is represented by the "tokens" field, whose value is an array of multiple tokens obtained after the sentence is divided by words. The entity field "entities" is an array consisting of all entity objects contained in the sentence. The object consists of three attributes: start, end, and type. Start and end represent the index values ​​of the left and right boundaries of the entity respectively, and type represents the category of the entity.

[0050] S2: Concatenate multiple prompt templates before each sentence in the named entity recognition data, and unify the length of the sentences after concatenating the prompt templates.

[0051] Add multiple prompt templates in front. The prompt template consists of a prompt slot and a prompt context. The entire input is converted into the dictionary id representation of the BERT model. If it exceeds the maximum input length of the model, the excess part is truncated and the insufficient part is filled with special characters [PAD], so that each sentence in a training batch is processed into the same length.

[0052] The prompt template can use a hard template, that is, the prompt context uses a specific word designed manually for prompts. For example, if the hard template "Entity is[E]" is used, "Entity is" is the prompt context, and "[E]" represents the prompt slot. You can also use a soft template to let the model learn the prompt context by itself during the training process, that is, " <soft>[E]”, for <soft>Randomly initialize and update the semantic vector of this part during training. The prompt slot [E] will eventually be decoded into the location and category probability distribution of the entity. When using a hard template, the final input of the model is:

[0053] Input(X)=Entity is[E1]Entity is[E2]…Entity is[E n ][SEP]x1x2…x n [PAD]…

[0054] Among them, X represents the sentence, Entity is is the prompt context, and E n Indicates the nth prompt slot, x n It represents the nth word in the sentence. [SEP] is a special symbol in the vocabulary, which is used to separate the prompt template from the sentence. [PAD] represents the filler.

[0055] S3: Add a prompt mask to the sentence after the concatenation of the prompt template and input it into the BERT model to obtain the prompt slot semantic representation and sentence semantic representation.

[0056] Since the self-attention mechanism of the BERT model itself will make sentence X pay attention to the randomly initialized prompt template at the beginning of encoding, thus destroying the original semantics of the sentence, it needs to be adjusted. Here, the prompt mask is used to make the sentence no longer pay attention to the information of the prompt template, and then the input is BERT encoded based on this mask. The formulas involved in the above steps are:

[0057] H = vector(Input(X))

[0058]

[0059] PA(H)=αHW v

[0060] H′=BERT PA (H)

[0061] Where X represents the sentence, Vector(Input(X)) represents the input word embedding representation, and W q ∈R h×h represents the query vector parameter matrix, W k ∈R h×h represents the key vector parameter matrix, W v ∈R h×h Represents the value vector parameter matrix, I∈R (N+M)×(N+M) is the prompt mask matrix, the I matrix is ​​set to 0 for the elements that retain attention, and the masked elements, that is, the elements at the corresponding positions of the prompt template are set to negative infinity (M is the number of prompt slots, N is the length of the original sentence X), Softmax is an activation function that maps values ​​from 0 to 1 and makes the sum equal to 1, PA(H) represents the prompt-independent unidirectional self-attention mechanism, T represents the matrix transpose, h and D represent the dimensions of the vector, and H′ represents the semantic representation of the prompt slot H E And the sentence semantic representation obtained after BERT encoding the input based on the prompt mask.

[0062] S4: Input the entity label into the BERT model to obtain the label semantic representation.

[0063] The entity category set, i.e., entity label, needs to be mapped into a natural language description. For example, the entity category set = [organization category, person name category, place name category], which is mapped into a natural language description: set = [organization, person, location].

[0064] If the label description is an abbreviation, it needs to be manually annotated, and then the processed result is input into BERT for label semantic encoding to obtain label semantic representation.

[0065] S5: Input the sentence semantic representation into the bidirectional LSTM network to obtain the sentence context semantic matrix; according to the sentence context semantic matrix, the cue slot semantic representation is processed using the self-attention mechanism to obtain the cue slot deep semantic matrix.

[0066] The sentence encoding needs to pass through a bidirectional LSTM module to deepen the context perception, so as to serve the subsequent boundary judgment. After passing through the LSTM module, the sentence encoding X will change. Here, a bidirectional LSTM is used, and the formula of the forward LSTM is:

[0067] i t =σ(W i x t +W i′ h t-1 +b i )

[0068] f t =σ(W f x t +W f′ h t-1 +b f )

[0069] o t =σ(W o x t +W o′ h t-1 +b o )

[0070] g t =tanh(W g x t +W g′ h t-1 +b g )

[0071] c t =f t ·c t-1 +i t ·g t

[0072] h t =o t ·tanh(c t )

[0073] Among them, x t is the input sequence at time t, h t-1 is the hidden state of the previous time step, i t 、f t , o t is the output of the input gate, forget gate, and output gate. t is the cell state at the current time step, h t is the hidden state of the current time step, W i , W i′ , W f , W f′ , W o , W o′ , W g , W g′ is the weight matrix, b i 、b f 、b o 、b g is the bias vector, σ is the sigmoid function, and tanh is the hyperbolic tangent function.

[0074] The formula for backward LSTM is:

[0075] i′ t =σ(W′ i x t +W′i ′ h t+1 +b′ i )

[0076] f′ t =σ(W′ f x t +W′ f′ h t+1 +b′ f )

[0077] o′ t =σ(W′ o x t +W′ o′ h t+1 +b′ o )

[0078] g′ t =tanh(W′ g x t +W′ g′ h t+1 +b′ g )

[0079] c′ t =f′ t ·c′ t-1 +i′ t ·g′ t

[0080] h′ t =o′ t tanh(c′ c )

[0081] Among them, x t is the input sequence at time t, h t+1 is the hidden state of the next time step, i t 、f t , o t is the output of the input gate, forget gate, and output gate. t is the cell state at the current time step, h t is the hidden state of the current time step, W′ i , W′ i′ , W′ f , W′ f′ , W′ o , W′ o′ , W′ g , W′ g′ is the weight matrix, b′ i , b′ f , b′ o , b′ g is the bias vector, σ is the sigmoid function, and tanh is the hyperbolic tangent function.

[0082] After the above steps, we get the sentence context semantic matrix H 3 . X With H E Through the interaction layer, the self-attention mechanism of the Transformer encoder is used to realize the self-attention between the prompt slots and the cross-attention between the prompt slots and the sentence, thereby obtaining a deep semantic representation of the prompt slots, which can be expressed as:

[0083]

[0084] Among them, H E represents the original hint slot semantic representation, H X Represents the sentence context semantic matrix, H s represents the semantic representation of the hint slot after the interaction between slots, H D It means that the semantic representation of the cue slot takes into account both the interaction between slots and the interaction between slots and sentences, that is, the deep semantic representation of the cue slot; and denote the key, query and value matrices of the hint slot respectively, and W k , W q and W v denote the first, second and third weight matrices respectively; and Represents the key and value matrix of the sentence; L represents the slot identifier embedding matrix, and Softmax represents the activation function.

[0085] S6: Map the deep semantic matrix of the prompt slot to the classification slot encoding matrix through a linear layer, and calculate the classification slot probability distribution based on the classification slot encoding matrix and the label semantic representation.

[0086] After the above encoding process is completed, decoding is performed next, that is, the probability output is obtained based on the features extracted from the representation vector of the prompt slot. First, classification is performed, and the deep semantic matrix H of the prompt slot is transformed into D Mapped into a classification slot encoding matrix; dot product operation on the classification slot encoding matrix and the label semantic representation to obtain the matching matrix of the prompt slot and the entity category; the matching matrix of the prompt slot and the entity category is decoded through a linear layer to obtain the classification slot probability distribution to achieve entity classification. The formula is:

[0087] cls result =Softmax(W1(W2H D +b1)·H l )+b2)

[0088]

[0089] Among them, cls result ∈R M×K is the probability distribution associated with entity classification, W1∈R h×h , W2∈R K×K is a trainable weight matrix, b1 and b2 are biases, and H E ∈R M×h is the semantic representation of the prompt slot, H l ∈R K×h is the semantic representation of the label, represents the probability that the entity extracted by the i-th classification slot is of the j-th category, K is the number of entity categories, and h represents the dimension of the hidden vector.

[0090] S7: Calculate the positioning slot probability distribution based on the hint slot deep semantic matrix and the sentence context semantic matrix.

[0091] The deep semantic matrix of the cue slot is linearly mapped to obtain the positioning slot encoding matrix;

[0092] Add the positioning slot encoding matrix to the sentence context semantic matrix after linear mapping to obtain the first matrix;

[0093] Add the classification slot encoding matrix to the sentence context semantic matrix after linear mapping to obtain the second matrix;

[0094] The first matrix is ​​added to the second matrix, and then decoded through a linear layer to obtain the probability distribution of the positioning slot, thus realizing the entity positioning. The above operation can be expressed as:

[0095]

[0096] Among them, W l1 , W l2 , W l3 , W l4 , W l5 ∈R h×h is the trainable left boundary parameter matrix, W r1 , W r2 , W r3 , W r4 , W r5 ∈R h×h is the trainable right boundary parameter matrix, W p ∈R h×h and W t ∈R h×h are the mapping matrices that map the deep semantic representation of the prompt slot to the positioning slot encoding and the classification slot encoding, respectively. h represents the dimension of the hidden vector, and b p and b t is the bias, is the deep semantic representation of the cue slot corresponding to cue slot i; and is the probability distribution of the positioning slot, represents the probability that the entity extracted by the prompt slot has j as the left boundary index value, Represents the probability that the entity extracted by hint slot i has j as the right boundary index value.

[0097] S8: Calculate the minimum matching loss between the prompt slot and the entity according to the probability distribution of the classification slot and the probability distribution of the positioning slot to obtain the optimal matching solution.

[0098] Because the correspondence between the prompt template and the entity is unknown and it is impossible to assign labels to the slots in advance, slot filling is regarded as a linear assignment problem, where any entity can be filled into any prompt template, which will produce a loss value, and then the optimal match can be calculated based on the loss value.

[0099] Since there are more hint templates than real entities, many hints will match The training efficiency is reduced. In order to improve the utilization rate of the prompt template, one-to-many matching is achieved by repeatedly sampling real entities, that is, the dynamic template filling mechanism. The nearest neighbor template filling is used here. First, the most matching entity is calculated for each prompt slot, and the entity is repeatedly sampled based on this as the sampling weight. The purpose is to ensure that the template matches the nearest neighbor entity first, avoid template confusion, and maximize the recall rate. The sampling weight formula is:

[0100]

[0101] Among them, MC j represents the total number of prompt slots matching the jth real entity in a sentence, Weight j It represents the sampling weight corresponding to the jth real entity in a sentence, M is the number of prompt slots, and A is the total number of entities originally in a sentence. Represents the confidence score related to entity classification, left boundary, and right boundary obtained by matching the jth real entity in the hint slot k of a sentence. max k It indicates the real entity index that the prompt slot k best matches. Equal is a flag function. When the two parameters in equal are equal, the equal function takes 1, otherwise it takes 0. That is, if the real entity that prompt slot k best matches is the jth real entity, the equal function takes 1, otherwise it takes 0. If MC j is 0, that is, there is no hint slot matching the j-th entity, then the numerator of the sampling weight is 1 to ensure that each real entity always has a small probability of being repeatedly sampled.

[0102] The entities are sampled repeatedly according to the sampling weights until the number of real entities reaches 0.9*M, that is, 90% of the prompt slots can be matched one-to-one with the real entities, and only 10% of the prompt slots will match empty values. These repeated real entities are used to fill the prompt slots to obtain the matching scheme between all prompt slots and entities.

[0103] Calculate the minimum matching loss among all matching solutions between prompt slots and entities, and select the matching solution corresponding to the minimum matching loss as the optimal matching solution. The formula is:

[0104]

[0105] Among them, σ * represents the optimal matching solution, represents the indicator function, which takes 1 when the i-th slot has a real entity matching it, otherwise it takes 0. represents the confidence probability value corresponding to the true category output by the prompt slot i under the σ-th arrangement, and argmin represents the minimum value function; Represents the confidence probability value corresponding to the true left and right boundaries output by prompt slot i under the σth arrangement.

[0106] S9: Calculate the total model loss under the optimal matching scheme; adjust the model parameters according to the total model loss to obtain a trained named entity recognition model.

[0107] The total loss of the model is L total It is the weighted sum of positioning loss and classification loss:

[0108] L total =αL type +βL position

[0109] Among them, α represents the classification weight, β represents the positioning weight, and L type represents the classification loss, L position represents the positioning loss.

[0110] The classification loss is expressed as:

[0111]

[0112] The localization loss is expressed as:

[0113]

[0114] in, The optimal matching arrangement σ * The confidence probability value corresponding to the true category output by the next prompt slot i, Represents the confidence probability value corresponding to the true left and right boundaries output by prompt slot i under the optimal matching arrangement.

[0115] According to the total loss of the model, back propagation updates the model parameters, and the training is continuously iterated. When the loss function converges or reaches the maximum preset number of iterations, the training is stopped and the optimal model parameters are saved. The Adam optimizer is used to dynamically adjust the learning rate of the model, and the BERT model parameters are frozen in the first 5 epochs to prevent catastrophic forgetting. The formula of the Adam algorithm is:

[0116]

[0117] in, represents the corrected first-order moment estimate, represents the corrected second-order moment estimate, ∈, η are the parameters that need to be adjusted during the training process.

[0118] The Micro-F1 value evaluation model is used during the test. The Micro-F1 formula is as follows:

[0119]

[0120] Among them, TP i Represents the true number of examples of the i-th category, FP i Indicates the number of false positive cases in the i-th category, FN i represents the number of false negative examples in the i-th category, and K is the total number of categories.

[0121] After the model training is completed, the text to be recognized is obtained and preprocessed, and the preprocessed text is input into the trained named entity recognition model for processing, and the category and left and right boundaries corresponding to the maximum value in the classification slot probability distribution and the positioning slot probability distribution are output, and the named entity recognition result is obtained.

[0122] In summary, the present invention splices the prompt template before the sentence and obtains the semantic representation through BERT encoding; uses the self-attention mechanism to realize the interaction between prompt slots and between prompt slots and sentences; combines the positioning information with the classification information and the text semantics to decode the entity boundary, and classifies the entity according to the classification information and the semantic similarity of the natural label; based on the prompt learning framework, the present invention designs a single-slot multi-task prompt template to extract multiple entities in parallel, solving the high time overhead and computational cost problems of the prompt learning method. The interaction between different prompt slots and tasks is realized through the fusion of the interaction layer and cross-task representation, which can significantly improve the entity extraction ability of the model.

[0123] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.< / soft> < / soft>

Claims

1. A method for named entity recognition based on prompt learning and label perception, characterized in that: include: Obtain the text to be recognized and preprocess it, input the preprocessed text into the trained named entity recognition model for processing, and obtain the named entity recognition result; The training process of the named entity recognition model includes: S1: Obtain a named entity recognition dataset and preprocess it to obtain preprocessed named entity recognition data; S2: concatenate multiple prompt templates before each sentence in the named entity recognition data, and unify the length of the sentences after concatenating the prompt templates; S3: Add a prompt mask to the sentence after the prompt template is concatenated and input into the BERT model to obtain the prompt slot semantic representation and sentence semantic representation; S4: Input the entity label into the BERT model to obtain the label semantic representation; S5: Input the sentence semantic representation into the bidirectional LSTM network to obtain the sentence context semantic matrix; use the self-attention mechanism to process the cue slot semantic representation according to the sentence context semantic matrix to obtain the cue slot deep semantic matrix; S6: Map the deep semantic matrix of the prompt slot into the classification slot encoding matrix through a linear layer, and calculate the probability distribution of the classification slot according to the classification slot encoding matrix and the label semantic representation; S7: Calculate the probability distribution of the positioning slot based on the deep semantic matrix of the prompt slot and the sentence context semantic matrix; S8: Calculate the minimum matching loss between the prompt slot and the entity according to the probability distribution of the classification slot and the probability distribution of the positioning slot, and obtain the optimal matching solution; S9: Calculate the total model loss under the optimal matching scheme, adjust the model parameters according to the total model loss, and obtain a trained named entity recognition model.

2. According to claim 1, a method for named entity recognition based on prompt learning and label perception is characterized in that: The semantic representation of the prompt slot is processed using the self-attention mechanism as follows: Among them, H E represents the original hint slot semantic representation, H X Represents the sentence context semantic matrix, H S represents the semantic representation of the hint slot after the interaction between slots, H D It means that the semantic representation of the cue slot takes into account both the interaction between slots and the interaction between slots and sentences, that is, the deep semantic representation of the cue slot; and denote the key, query and value matrices of the hint slot respectively, and W k , W q and W v denote the first, second and third weight matrices respectively; and Represents the key and value matrix of the sentence; L represents the slot identifier embedding matrix, h represents the dimension of the hidden vector, and Softmax represents the activation function.

3. According to the method for named entity recognition based on prompt learning and label perception in claim 1, it is characterized in that: The process of calculating the probability distribution of classification slots includes: mapping the deep semantic matrix of the prompt slots to the classification slot encoding matrix through linear mapping; performing dot product operation on the classification slot encoding matrix and the label semantic representation to obtain the matching matrix of the prompt slots and entity categories; the matching matrix of the prompt slots and entity categories is decoded through a linear layer to obtain the probability distribution of classification slots.

4. The method for named entity recognition based on prompt learning and label perception according to claim 1, characterized in that: The process of calculating the probability distribution of the positioning slots includes: The deep semantic matrix of the cue slot is linearly mapped to obtain the positioning slot encoding matrix; Add the positioning slot encoding matrix to the sentence context semantic matrix after linear mapping to obtain the first matrix; Add the classification slot encoding matrix to the sentence context semantic matrix after linear mapping to obtain the second matrix; The first matrix is ​​added to the second matrix and decoded again through a linear layer to obtain the probability distribution of the positioning slots.

5. The method for named entity recognition based on prompt learning and label perception according to claim 1, characterized in that: The process of calculating the matching loss between the prompt slot and the entity includes: Calculate the sampling weight of each entity according to the classification slot probability distribution and the positioning slot probability distribution; The entities are repeatedly sampled according to the sampling weights to fill all the prompt slots, and the matching schemes between all the prompt slots and the entities are obtained; Calculate the minimum matching loss among all matching schemes between prompt slots and entities, and select the matching scheme corresponding to the minimum matching loss as the optimal matching scheme.

6. A method for named entity recognition based on prompt learning and label perception according to claim 5, characterized in that: The formula for calculating the sampling weight of each entity is: Among them, MC j Represents the total number of prompt slots matching the jth real entity in the sentence, Weight j represents the sampling weight corresponding to the jth real entity in the sentence, M is the number of prompt slots, and A is the total number of entities originally in a sentence; Indicates the matching classification, left boundary, and right boundary value obtained when the hint slot k matches the jth real entity; max k It represents the real entity index that best matches the hint slot k, and equal is the marking function.

7. The method for named entity recognition based on prompt learning and label perception according to claim 5, characterized in that: The formula for calculating the minimum matching loss among all matching schemes between hint slots and entities is: Among them, σ * represents the optimal matching solution, M represents the number of prompt slots, represents the indicator function, represents the confidence probability value corresponding to the true category output by the prompt slot i under the σth arrangement, It represents the confidence probability value corresponding to the true left and right boundaries output by the prompt slot i under the σ-th arrangement, and argmin represents the minimum value function.

8. The method for named entity recognition based on prompt learning and label perception according to claim 1, characterized in that: The total model loss is the weighted sum of the positioning loss and the classification loss, where: Classification loss L type It is expressed as: Positioning loss L position It is expressed as: in, The optimal matching arrangement σ * The confidence probability value corresponding to the true category output by the next prompt slot i, represents the confidence probability value corresponding to the true left and right boundaries output by the prompt slot i under the optimal matching arrangement, represents the indicator function, and M represents the number of prompt slots.

Citation Information

Patent Citations

  • Bidirectional LSTM named entity recognition method based on predicted position attention

    CN109933801A

  • Knowledge prompt fused legal text small sample named entity recognition method

    CN115062104A

  • BERT-based military field composite named entity recognition method

    CN115238690A

  • Small sample named entity recognition method based on multiple tasks and prompt learning

    CN116151256A

  • BERT-BiLSTM-CRF dangerous chemical named entity identification method fusing multiple features

    CN116432647A