Sample named entity identification method based on pointer network and prototype network

By combining the method of pointer network and prototype network, the semantic information utilization and adaptability of named entity recognition technology are achieved, the problems of insufficient use and poor adaptability of semantic information in the existing technology are solved, and efficient naming entity recognition is achieved.

CN120068867AActive Publication Date: 2025-05-30BEIJING UNIV OF POSTS & TELECOMM
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202411971726.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

Existing named entity recognition technology is difficult to effectively utilize the semantic information of words, and traditional methods require a large amount of data to be retrained when applied to new fields, with high coupling and poor adaptability.

Method used

The sample named entity recognition method based on pointer network and prototype network is adopted, entity range detection is performed through pointer network, and entity classification is performed by prototype network, and semantic information of words is clearly used to reduce the coupling between entity categories and models.

Benefits of technology

It realizes efficient naming entity recognition in the target field, avoids the limitation that entity annotation can only be based on characters, reduces interference from non-entity characters, and improves the adaptability and efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068867A_ABST
    Figure CN120068867A_ABST
Patent Text Reader

Abstract

The invention discloses a sample named entity identification method based on a pointer network and a prototype network. The method comprises the following steps: S1, acquiring a public named entity identification data set; s2, performing support set and query set division on the data set, wherein each sample of the divided data set is composed of a support set and a query set; s3, constructing a decomposition type model combining the pointer network and the prototype network, and storing model parameters with the best effect; s4, selecting samples of each type of entities in the target field to perform manual annotation, training the model by using the annotated data, and performing named entity recognition on the text in the target field by using the trained model; according to the method, traditional pattern named entity recognition is divided into two parts of entity range detection and entity classification, entities in texts are recalled through a pointer network, entity classification is performed through a prototype network, the limitation that entity labeling can only be based on characters is solved, and interference of non-entity characters during entity classification is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing information extraction, and specifically relates to a method for sample named entity recognition based on pointer network and prototype network. Background Art

[0002] Named entity recognition (NER), as a basic task in natural language processing, aims to identify entities in text and classify them into predefined types. In recent years, various deep learning-based methods have been used for the named entity recognition task and achieved remarkable results. Deep learning-based methods often rely on large-scale labeled data and require a large amount of manual annotation of the training set. In addition, traditional deep neural network models use CRF or a classifier as the decoder, and the entity categories are highly coupled with the model. When the model is applied to a new field, a large amount of data is required to retrain the classifier.

[0003] Deficiencies of the prior art:

[0004] Currently, metric learning can reduce the coupling degree between the model and entity categories. When the model is applied to a new field, the model training can be completed with a small amount of data. However, traditional metric learning methods perform sequence annotation on text based on characters, cannot explicitly utilize the semantic information of words, and are interfered by a large number of non-entity categories. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for sample named entity recognition based on pointer network and prototype network to solve the problems raised in the above background art.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for sample named entity recognition based on pointer network and prototype network, which specifically includes the following steps:

[0007] S1. Obtain a publicly available named entity recognition data set;

[0008] S2. Divide the data set into a support set and a query set. Each sample of the divided data set is composed of a support set and a query set;

[0009] S3. Construct a decomposable model combining a pointer network and a prototype network, and use the obtained data set to train the decomposable model. After each round of training, use the validation set to verify the model effect and save the model parameters with the best effect;

[0010] S4. Select samples of each type of entity in the target field for manual annotation, use the labeled data to train the model, and use the trained model to perform named entity recognition on the text in the target field.

[0011] Preferably, in step S1, to construct the categories of entities, an annotation tool is used to annotate the entities in the domain text, and the annotation results in the JavaScript Object Notation (JSON) format or the Comma-Separated Values (CSV) file format are obtained.

[0012] Preferably, the annotation tool is a Quick Annotation Tool or Label Studio.

[0013] Preferably, step S2 specifically includes the following steps:

[0014] a1. Divide the support set and the query set. For the named entity recognition dataset D, ∈=(S, Q) ∈ D is a sample in the dataset, where S and Q represent the support set and the query set respectively. ∈ follows the N-way K~2K shot setting. N-way means that the support set and the query set each contain N categories of entities, and K~2K shot means that there are K~2K instances for each category of entity.

[0015] a2. Construct each sample in the dataset in sequence to construct the support set and the query set of the sample.

[0016] a3. For the support set where the number of entity categories does not reach N, add the text containing new type entities to the support set where the number of entity categories does not reach N, and the newly added text will not violate the K~2K shot constraint. The query set does not need to satisfy the K~2K shot constraint.

[0017] Preferably, step S3 specifically includes the following steps:

[0018] b1. Using the model-agnostic meta-learning method, divide one training into two parts: an inner loop and an outer loop. One outer loop contains multiple inner loops, and each inner loop corresponds to a sample in the dataset.

[0019] b2. At the beginning of each inner loop, obtain the mirror image of the model, use the support set data to train the mirror image of the model, use the trained mirror image to perform entity recognition on the query set, and obtain the loss through the cross-entropy function. Differentiate the loss with respect to the parameters to obtain the gradient.

[0020] b3. In the outer loop, average all the gradients obtained in the inner loops to update the initial model parameters.

[0021] Preferably, the model in step S3 includes an entity range detection module and an entity classification module. For the input text, use a pointer network to generate the position of the entity in the text, extract the feature vector hidden vector of the corresponding position character, obtain the feature vector of the entity by taking the average of the feature vectors of each character of the entity, and use a prototype network for entity classification.

[0022] Preferably, the main steps of using a pointer network for entity scope detection include:

[0023] c1. Given an input text X = {x 1 , x 2 , …, x n}, use the encoder Encoder(·) to encode X into a representation vector H en in the hidden space:

[0024] H en = Encoder(X)

[0025] where d is the dimension of the hidden vector; the decoder Decoder(·) takes the encoding result H en of the encoder and the entity positions i 1 , i 2 , …, i t-1 already output by the decoder as inputs, and decodes i t , where i represents the position of the named entity in the sentence, 0 ≤ i m ≤ n;

[0026] c2. Obtain the corresponding characters through , Map(·) is the mapping from character position to character, is the character at position i m , represents the entity sequence that has been decoded, and use the following formula to calculate the probability at i t :

[0027] h t = Decoder(H en ; x)

[0028] q t = W q h t

[0029] K = W k H en

[0030]

[0031] where is the decoded hidden vector, W q , is the mapping matrix of the attention mechanism, and Softmax(·) is the activation function. Denote the attention of the decoder position t to the encoded sequence X, regarded as the probability distribution of each character position in the encoded sequence. In the training phase, use the cross-entropy function CrossEntropy(·) to measure the probability distribution p of the position t and the correct position i t The loss between them is:

[0032]

[0033] c3. In the prediction phase, use the trained model to sequentially generate the positions of entities in the sentence, and use the beam search algorithm for decoding to generate the sequence of entity positions.

[0034] Preferably, the main steps of using the prototype network for entity classification include:

[0035] d1. Given an input text X = {x 1 , x 2 , …, x n}, use the encoder Encoder(·) to calculate the context semantic representation H = {h 1 , h 2 , …, h n} of the input text:

[0036] H = Encoder(X)

[0037] Given the data ∈ = (S, Q), for the entity type t ∈ T, Denote the set of entity types, and use the semantic representations of entities of this type in the support set S to construct prototypes, Denote that the training set contains L instances, X u Denote the character sequence of the u-th instance, and let Denote the semantic representation of the entity of type t in the u-th instance. The prototype can be calculated using the following formula:

[0038]

[0039] d2. Construct prototypes for each type of entity in the support set, and classify the entities in the query set according to the type prototypes. Let w [j,k] Denote the entity obtained in the entity range detection stage, and the start position is i j , and the end position is i k . The semantic representation corresponding to this entity is obtained through the encoding result of :

[0040]

[0041] The probability that the type of this entity is t is:

[0042] p(t|e[j,k] ) = Softmax(-d(c t , e [j,k] ))

[0043] Among them, d(·, ·) represents the distance function, and let t [j,k] ∈T represent the actual type of e [j,k] . The cross-entropy loss function is as follows:

[0044]

[0045] Preferably, in the step S3 of saving the parameters of the model: insert a low-rank adapter into the model, and adjust the parameters of the inserted part for the projection matrix in the model is decomposed into the form of the product of two low-rank matrices:

[0046] W + ΔW = W + BA

[0047] Among them, and r << min(d, k). During the training process, W does not update its parameters, while A and B contain trainable parameters. For the input vector x, the output vector h = Wx is redefined as:

[0048] h = (W + ΔW)x = Wx + BAx

[0049] At the beginning of training, use random Gaussian initialization for matrix A, and the value of matrix B is all zero, then ΔW = BA is zero. During the training process, the update of the parameters of matrix A and B can be regarded as the update of the projection matrix W; the LoRA structure is inserted into the feed-forward neural network or the attention layer. In the attention layer, it includes the query mapping matrix W q , the key mapping matrix W k , the value mapping matrix W v and the output mapping matrix W o ; in the feed-forward neural network layer, it includes the input mapping matrix W i and the output mapping matrix W o .

[0050] Preferably, the beam search is as follows: during the process of generating the target sequence, the beam search algorithm outputs the value of the current position for each candidate sequence and calculates the overall probability of the candidate sequence, sorts the k candidate sequences according to the probability, some sequences output the termination symbol in advance, and some terminate after reaching the pre-set maximum output sequence length. After meeting the end condition, select the candidate sequence with the highest probability as the final solution.

[0051] Compared with the prior art, the beneficial effects of the present invention are:

[0052] A method for named entity recognition of samples based on pointer network and prototype network, the steps include: S1, obtaining a public named entity recognition data set; S2, dividing the data set into a support set and a query set, and each sample of the divided data set is composed of a support set and a query set; S3, constructing a decomposable model combining a pointer network and a prototype network, training the decomposable model using the obtained data set, verifying the model effect using a validation set at the end of each round of training, and saving the model parameters with the best effect; S4, selecting samples of each type of entity in the target domain for manual annotation, training the model using the annotated data, and using the trained model to perform named entity recognition on the text in the target domain; The present invention divides the named entity recognition of traditional patterns into two parts: entity range detection and entity classification. The pointer network is used to recall entities in the text, and the prototype network is used for entity classification, clearly utilizing the semantic information of words, solving the limitation that entity annotation can only be based on characters, and avoiding the interference of non-entity characters during entity classification. Description of the Drawings

[0053] Figure 1 It is a flowchart of the method provided by the present invention. Detailed Embodiments

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0055] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.

[0056] In the description of this patent, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "setting" should be understood in a broad sense. For example, it can be fixedly connected and set, or detachably connected and set, or integrally connected and set. For those of ordinary skill in the art, the specific meanings of the above terms in this patent can be understood according to specific circumstances.

[0057] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, the meaning of "a number of" is two or more, unless otherwise specifically defined.

[0058] Embodiment

[0059] Please refer to Figure 1 As shown, a technical solution of a sample named entity recognition method based on a pointer network and a prototype network provided by the present invention: This method specifically includes the following steps:

[0060] S1. Obtain a publicly available named entity recognition data set, construct entity categories, use an annotation tool to annotate entities in domain texts, obtain annotation results in JavaScript object representation format or comma-separated value file format, use an annotation tool to annotate entities in domain texts, and obtain annotation results in JavaScript object representation format or comma-separated value file format;

[0061] S2. Divide the data set into a support set and a query set. Each sample of the divided data set consists of a support set and a query set;

[0062] a1. Divide the support set and the query set. For the named entity recognition data set D, ∈=(S, Q)∈D is a sample in the data set, where S and Q represent the support set and the query set respectively. ∈ follows the N way K~2K shot setting. N way means that the support set and the query set respectively contain N types of entities, and K~2K shot means that there are K~2K instances for each type of entity;

[0063] a2. Construct each sample in the data set in sequence to construct the support set and the query set of the sample;

[0064] a3. For the support set whose number of entity categories does not reach N, add the text containing new type entities

[0065] to the support set whose number of entity categories does not reach N, and the newly added text will not violate the K~2K shot constraint. The query set does not need to satisfy the K~2K shot constraint;

[0066] S3. Construct a decomposed model combining a pointer network and a prototype network, use the obtained data set to train the decomposed model, verify the model effect using a validation set at the end of each round of training, and save the model parameters with the best effect;

[0067] b1. Use model-agnostic meta-learning method to divide one training into two parts: inner loop and outer loop. One outer loop contains multiple inner loops, and each inner loop corresponds to a sample in the dataset.

[0068] b2. At the beginning of each inner loop, obtain the mirror image of the model, use the support set data to train the mirror image of the model, use the trained mirror image to perform entity recognition on the query set, and obtain the loss through the cross-entropy function. Differentiate the loss with respect to the parameters to obtain the gradient.

[0069] b3. In the outer loop, average all the gradients obtained in the inner loops to update the initial model parameters.

[0070] S4. Select samples of each type of entity in the target domain for manual annotation, use the annotated data to train the model, and use the trained model to perform named entity recognition on the text in the target domain.

[0071] Furthermore, the model includes an entity scope detection module and an entity classification module. For the input text, use the pointer network to generate the position of the entity in the text, extract the feature vector hidden vector of the corresponding position character, obtain the feature vector of the entity by taking the average of the feature vectors of each character of the entity, and use the prototype network for entity classification. The main steps of using the pointer network for entity scope detection include:

[0072] c1. Given an input text X = {x 1 , x 2 , …, x n}, use the encoder Encoder(·) to encode X into the representation vector H en in the hidden space:

[0073] H en = Encoder(X)

[0074] where d is the dimension of the hidden vector; the decoder Decoder(·) takes the encoding result H en of the encoder and the entity positions i 1 , i 2 , …, i t-1 output by the decoder as inputs, and decodes i t , representing the position of the named entity in the sentence, 0 ≤ i m ≤ n;

[0075] c2. Obtain the corresponding character through , Map(·) is the mapping from character position to character, is the character at position i m , Denote the decoded entity sequence, and calculate the probability at position i using the following formula: t :

[0076] h t = Decoder(H en ; x)

[0077] q t = W q h t

[0078] K = W k H en

[0079]

[0080] where is the decoded hidden vector, W q , is the mapping matrix of the attention mechanism, Softmax(·) is the activation function, represents the attention of the decoder position t to the encoded sequence X, regarded as the probability distribution of each character position in the encoded sequence. In the training stage, the cross-entropy function CrossEntropy(·) is used to measure the loss between the probability distribution p t and the correct position i t :

[0081]

[0082] On the one hand, the number of entities generated by the pointer network is only related to the number of entities in the training data, and the imbalance between the number of entities and non-entities does not affect its prediction results; on the other hand, the pointer network does not pay attention to boundary information, thus avoiding the problem of entity range overlap caused by the intersection of the start and end boundaries of different entities;

[0083] c3. In the prediction stage, use the trained model to sequentially generate the positions of entities in the sentence, and use the beam search algorithm for decoding to make the overall effect of the generated entity position sequence better; the beam search algorithm is to output the value of the current position for each candidate sequence and calculate the overall probability of the candidate sequence during the process of generating the target sequence, sort the k candidate sequences according to the probability, some sequences output the termination symbol in advance, and some terminate after reaching the pre-set maximum output sequence length. After meeting the end condition, select the candidate sequence with the highest probability as the final solution.

[0084] Furthermore, the main steps of using the prototype network for entity classification include:

[0085] d1. Given an input text X = {x 1 , x2 ,…,x n}, use the encoder Encoder(·) to calculate the context semantic representation H = {h 1 , h 2 ,…, h n}:

[0086] H = Encoder(X)

[0087] Given the data ∈ = (S, Q), for the entity type t ∈ T, representing the set of entity types, construct prototypes using the semantic representations of entities of this type in the support set S, representing that the training set contains L instances, X u representing the character sequence of the u-th instance, let representing the semantic representation of the entity of type t in the u-th instance, the prototype can be calculated using the following formula:

[0088]

[0089] d2. Construct prototypes for each type of entity in the support set, and classify the entities in the query set according to the type prototypes. Let e [j,k] represent the entity obtained in the entity range detection stage, and the starting position is i j , and the ending position is i k , and the semantic representation corresponding to this entity is obtained through the encoding result of :

[0090]

[0091] The probability that the type of this entity is t is:

[0092] p(t|e [j,k] ) = Softmax(-d(c t , e [j,k] ))

[0093] where d(·,·) represents the distance function. Let t [j,k] ∈ T represent the actual type of e [j,k] , and the cross-entropy loss function is as follows:

[0094]

[0095] Furthermore, among the parameters of the saved model in step S3: Insert the LoRA component into the model and adjust the parameters of the inserted part for the projection matrix in the model to be decomposed into the form of the product of two low-rank matrices:

[0096] W + ΔW = W + BA

[0097] Among them, and r << min(d, k). During the training process, W does not update its parameters, while A and B contain trainable parameters. For the input vector x, the output vector h = Wx is redefined as:

[0098] h = (W + ΔW)x = Wx + BAx

[0099] At the beginning of training, the matrix A is initialized with random Gaussian values, and the values of the matrix B are all zero, so ΔW = BA is zero. During the training process, the update of the parameters of the matrices A and B can be regarded as the update of the projection matrix W; The LoRA structure is inserted into the feed-forward neural network or the attention layer. In the attention layer, it includes the query mapping matrix W q , the key mapping matrix W k , the value mapping matrix W v and the output mapping matrix W o ; In the feed-forward neural network layer, it includes the input mapping matrix W i and the output mapping matrix W o . Using the LoRA method, the model parameters can be fixed to avoid the problem of catastrophic forgetting; and only the parameters of the LoRA component part need to be stored for each target domain, and the parameters of the model main body only need to be stored once, saving storage space.

[0100] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A sample named entity recognition method based on pointer network and prototype network, characterized by: This method specifically comprises the following steps: S1. Obtain a public named entity recognition dataset; S2, divide the data set into support set and query set, and each sample of the divided data set consists of support set and query set; S3. Build a decomposition model combining the pointer network and the prototype network, use the obtained data set to train the decomposition model, use the validation set to verify the model effect at the end of each round of training, and save the model parameters with the best effect; S4. Select samples of each type of entity in the target field for manual labeling, use the labeled data to train the model, and use the trained model to perform named entity recognition on the text in the target field.

2. According to claim 1, a sample named entity recognition method based on pointer network and prototype network is characterized in that: In the step S1, the entity category is constructed, the entity in the domain text is annotated using an annotation tool, and the annotation result in JavaScript object representation format or comma separated value file format is obtained.

3. According to claim 2, a sample named entity recognition method based on pointer network and prototype network is characterized in that: The marking tool is a quick annotation tool or a label studio.

4. The sample named entity recognition method based on pointer network and prototype network according to claim 1, characterized in that: The step S2 specifically includes the following steps: a1. Divide the support set and query set. For the named entity recognition dataset D, ∈=(S,Q)∈D is a sample in the dataset, S and Q represent the support set and query set respectively, ∈ follows the N wayK~2K shot setting, N way means that the support set and query set contain N types of entities respectively, K~2K shot means that there are K~2K instances of each type of entity; a2. Construct each sample in the data set in sequence to construct the support set and query set of the sample; a3. For the support set whose number of entity categories does not reach N, add the text containing the new type of entity to the support set whose number of entity categories does not reach N, and the newly added text will not violate the K~2K shot constraint, and the query set does not need to satisfy the K~2K shot constraint.

5. The sample named entity recognition method based on pointer network and prototype network according to claim 1, characterized in that: The step S3 specifically includes the following steps: b1. Use a model-independent meta-learning method to divide a training session into two parts: an inner loop and an outer loop. An outer loop contains multiple inner loops, and each inner loop corresponds to a sample in the data set. b2. At the beginning of each inner loop, obtain the mirror image of the model, use the support set data to train the mirror image of the model, use the trained mirror image to perform entity recognition on the query set, and obtain the loss through the cross entropy function. The loss is differentiated to obtain the gradient of the parameters; b3. In the outer loop, average all the gradients obtained in the inner loop to update the initial model parameters.

6. The sample named entity recognition method based on pointer network and prototype network according to claim 1, characterized in that: The model in step S3 includes an entity range detection module and an entity classification module. For the input text, a pointer network is used to generate the position of the entity in the text, and the feature vector hidden vector of the character at the corresponding position is extracted. The feature vector of the entity is obtained by averaging the feature vectors of each character of the entity, and the prototype network is used to perform entity classification.

7. The sample named entity recognition method based on pointer network and prototype network according to claim 6 is characterized by: The main steps of using pointer networks for entity range detection include: c1. Given an input text X = {x1, x2, ..., x n }, use encoder Encoder(·) to encode X into a representation vector H in the latent space en : H en =Encoder(X) in, d is the dimension of the hidden vector; the decoder Decoder(·) takes the encoding result H of the encoder en and the entity positions i1,i2,…,i output by the decoder t-1 As input, decode i t , Indicates the position of the named entity in the sentence, 0≤i m ≤n; c2. Pass Get the character at the corresponding position. Map(·) is the mapping from character position to character. is the position i m Characters, Indicates the decoded entity sequence, and uses the following formula to calculate i t The probability of: h t =Decoder(H en ;x) q t =W q h t K=W k H en in, is the decoded hidden vector, is the mapping matrix of the attention mechanism, Softmax(·) is the activation function, represents the attention of decoder position t to the encoding sequence X, which is regarded as the probability distribution of each character position in the encoding sequence. During the training phase, the cross entropy function CrossEntropy(·) is used to measure the probability distribution of the position p t With the correct position t The loss between: c3. In the prediction stage, the trained model is used to sequentially generate the positions of entities in the sentence, and the beam search algorithm is used to decode the generated entity position sequence.

8. The method for sample named entity recognition based on pointer network and prototype network according to claim 6, characterized in that: The main steps of using prototype networks for entity classification include: d1. Given an input text X = {x1, x2, ..., x n }, use encoder Encoder(·) to calculate the contextual semantic representation of the input text H = {h1,h2,…,h n }: H=Encoder(X) Given data ∈=(S,Q), for entity type t∈T, Represents a set of entity types, and uses the semantic representation of entities of this type in the support set S to construct a prototype. Indicates that the training set contains L instances, X u Represents the character sequence of the uth instance, let Represents the semantic representation of entity type t in the u-th instance. The prototype can be calculated using the following formula: d2. Construct prototypes of each type of entity in the support set, and classify the entities in the query set according to the type prototype. Let e [j,k] Represents the entity obtained in the entity range detection phase, and the starting position is i j , the end position is i k , the semantic representation of this entity is expressed by The encoding result is: The probability that the entity is of type t is: p(t|e [j,k] )=Softmax(-d(c t ,e [j,k] )) Where d(·,·) represents the distance function, let t [j,k] ∈T represents e [j,k] The actual type of the cross entropy loss function is as follows:

9. The method for sample named entity recognition based on pointer network and prototype network according to claim 1, characterized in that: The step S3 saves the parameters of the model: inserting the low-rank adapter into the model, adjusting the parameters of the inserted part for the projection matrix in the model Decomposed into the form of multiplication of two low-rank matrices: W+ΔW=W+BA in, And r<<min(d,k), W does not perform parameter updates during training, while A and B contain trainable parameters. For the input vector x, the output vector h=Wx is redefined as: h=(W+ΔW)x=Wx+BAx At the beginning of training, random Gaussian initialization is used for matrix A, and the values ​​of matrix B are all zero, then ΔW=BA is zero. During the training process, the update of the parameters of matrices A and B can be regarded as the update of the projection matrix W; the LoRA structure is inserted into the feedforward neural network or the attention layer. In the attention layer, the query mapping matrix W q , key mapping matrix W k , value mapping matrix W v And the output mapping matrix W o ; In the feedforward neural network layer, including the input mapping matrix W i And the output mapping matrix W o .

10. The sample named entity recognition method based on pointer network and prototype network according to claim 7, characterized in that: The beam search is as follows: in the process of generating the target sequence, the beam search algorithm outputs the value of the current position for each candidate sequence and calculates the probability of the candidate sequence as a whole, sorts the k candidate sequences according to the probability, outputs the termination symbol in advance for some sequences, and terminates the other sequences after reaching the preset maximum output sequence length. After the termination condition is met, the candidate sequence with the largest probability is selected as the final solution.

Citation Information

Patent Citations

  • Nested named entity recognition method and system, electronic equipment and readable medium

    CN110956042A

  • Named entity identification model establishing method and named entity identification method

    CN112364655A

  • Few-sample named entity identification method and system for entity boundary category decoupling

    CN112541355A

  • Chinese electronic medical record named entity identification method based on active learning

    CN115440330A

  • Named entity recognition method based on global pointer and adversarial training

    CN115688782A