Named entity identification method and system based on multivariate information span score
By using the multivariate information span scoring method in named entity recognition, and using the sequence encoder and span scorer to improve feature extraction capabilities, the problem of poor recognition effect of traditional methods is solved, and a higher accuracy of naming entity recognition is achieved.
Patent Information
- Application Number
- CN202510027147.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional naming entity recognition methods fail to make full use of the relationship between the multivariate information contained in the text span itself and the adjacent span, resulting in limited feature extraction capabilities and poor recognition results.
A named entity recognition method based on multivariate information span score is adopted. A named entity recognition model is created through sequence encoder, span filter, span classifier and output module. A filter vector sequence and a classification vector sequence are used for scoring, and an entity score is output combined with filter scores and classification scores to achieve named entity recognition.
By making full use of the relationship between the multi-information of text span and the adjacent span, the feature extraction capability is significantly improved, and the accuracy of naming entity recognition is greatly improved.
Smart Images

Figure CN120068868A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to a named entity recognition method and system based on multi-information span scoring. Background Art
[0002] Named entity recognition (NER) is a basic task in the technical field of natural language processing. Its purpose is to extract entities with specific meanings from unstructured text sequences and determine them as predefined categories, such as person names, place names, organization names, etc. In engineering practice, named entity recognition is widely applied in fields such as news, medicine, law, and social media.
[0003] Named entity recognition research faces various challenges, such as semantic diversity and entity nesting. For the recognition of named entities, the traditionally mainstream methods are recognition methods based on deep learning, and most of them adopt the paradigm of supervised feature learning plus fine-tuning, mainly including sequence annotation methods and span-based methods. However, traditional methods fail to fully utilize the multi-information contained in the text span itself, and the relationship between adjacent spans is not taken seriously, resulting in limited feature extraction ability and poor recognition effect.
[0004] Therefore, how to provide a named entity recognition method and system based on multi-information span scoring to improve the accuracy of named entity recognition has become an urgent technical problem to be solved. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a named entity recognition method and system based on multi-information span scoring to improve the accuracy of named entity recognition.
[0006] The present invention is implemented as follows: A named entity recognition method based on multi-information span scoring includes the following steps:
[0007] Step S1, collect a large amount of text data, perform preprocessing on each of the text data including at least data cleaning and data verification, label the entities in each of the preprocessed text data in the form of triples to construct a data set;
[0008] Step S2, create a named entity recognition model based on a sequence encoder, a span filter, a span classifier, and an output module, and set the loss function and optimizer of the named entity recognition model;
[0009] The sequence encoder is used to encode the input text data to obtain a filtered vector sequence and a classification vector sequence;
[0010] The span filter is used to score using a sequence of filter vectors to obtain a filter score;
[0011] The span classifier is used to score using a sequence of classification vectors to obtain a classification score;
[0012] The output module is used to sum the filter score and the classification score to obtain an entity score, and output a named entity recognition result based on the entity score;
[0013] Step S3: Train the named entity recognition model with the dataset;
[0014] Step S4: Execute the named entity recognition task with the trained named entity recognition model.
[0015] Further, in step S1, the specific operation of annotating the entities in each of the preprocessed text data in the form of triples is as follows:
[0016] Annotate the entities in each of the preprocessed text data in the form of triples, and the annotation content is:
[0017] (i, j, t);
[0018] where i represents the index number of the first character of the entity; j represents the index number of the last character of the entity; t represents the entity type.
[0019] Further, in step S2, the sequence encoder is composed of a backbone encoder, a first linear encoder, and a second linear encoder;
[0020] The backbone encoder is used to encode the input text data into a sequence of token vectors, and the formula is:
[0021] H = PLM(X) = {h 1 , h 2 , …, h n};
[0022] where H represents the sequence of token vectors; PLM() represents the backbone encoder; X represents the text data, and X = {x 1 , x 2 , …, x n}, x n represents the nth token in the text data; h n represents the nth token vector in the sequence of token vectors;
[0023] The first linear encoder is used to linearly encode the sequence of token vectors to obtain a sequence of filter vectors:
[0024] H F = HWF +b F ={h F,1 ,h F,2 ,…,h F,n}};
[0025] Among them, H F represents the screening vector sequence; W F represents the first trainable parameter, W F ∈R D×d , R represents real numbers, D represents the dimension of the token vector, and d represents the size of the feature dimension; b F represents the second trainable parameter, b F ∈R d ; h F,n represents the nth screening vector in the screening vector sequence;
[0026] The second linear encoder is used to linearly encode the token vector sequence to obtain a classification vector sequence:
[0027] H C =HW C +b C ={h C,1 ,h C,2 ,...,h C,n}};
[0028] Among them, H C represents the classification vector sequence; W C represents the third trainable parameter, W C ∈R D×d ; b C represents the fourth trainable parameter, b C ∈R d ; h C,n represents the nth classification vector in the classification vector sequence.
[0029] Furthermore, in step S2, the span filter is composed of a third linear encoder, a fourth linear encoder, and a rotary position encoder;
[0030] Both the third linear encoder and the fourth linear encoder are used to encode the screening vector sequence to obtain a first feature vector h s,i and a second feature vector h e,j :
[0031]
[0032] Among them, W s represents the fifth trainable parameter, W s ∈R d×d ; T represents transpose; h F,iDenote the first screening vector of any span (i, j) in the screening vector sequence; b s Denote the sixth trainable parameter, b s ∈R d ; W e Denote the seventh trainable parameter, W e ∈R d ×d ; h F,j Denote the last screening vector of any span (i, j) in the screening vector sequence; b e Denote the eighth trainable parameter, b e ∈R d ;
[0033] The rotation position encoder is used to calculate a screening score according to the first feature vector h s,i and the second feature vector h e,j :
[0034]
[0035] where, S F(i,j) Denote the screening score of the screening vector corresponding to the span (i, j) in the screening vector sequence; Q i Denote the rotation position encoding of token i; Q j Denote the rotation position encoding of token j; Q j-1 Denote the rotation position encoding of token j - 1.
[0036] Furthermore, in the step S2, the span classifier is composed of a bi-affine attention module, a position attention module, a first fusion module, a linear classification module, an inter-span interaction module, and a second fusion module;
[0037] The bi-affine attention module is used to encode the classification vector sequence to obtain a cross-boundary information representation e B(i,j) :
[0038] e B(i,j) =[h C,i ; 1] T U[h C,j ; 1];
[0039] where, h C,i Denote the first classification vector of any span (i, j) in the classification vector sequence; h C,j Denote the last classification vector of any span (i, j) in the classification vector sequence; [.;.] represents the vector concatenation operation; U denotes the ninth trainable parameter, U ∈ R (d+1)×d×(d+1) ;
[0040] The position attention module is used to encode the classification vector sequence to obtain an intra-span information representation eI(i,j) :
[0041]
[0042] Among them, α u represents the attention weight corresponding to the u-th classification vector in the classification vector sequence; h C,u represents the u-th classification vector in the classification vector sequence; a u and a v respectively represent the attention scores corresponding to the u-th and v-th classification vectors in the classification vector sequence; W I represents the tenth trainable parameter, W I ∈R d×1 ; b I represents the eleventh trainable parameter, b I ∈R d ;
[0043] The first fusion module is constructed based on a gating network and is used to perform weighted fusion on the cross-boundary information representation e B(i,j) and the intra-span information representation to obtain an enhanced span information representation e (i,j) :
[0044] e (i,j) = g (i,j) e B(i,j) + (1 - g (i,j) ) e I(i,j) ;
[0045]
[0046] Among them, g (i,j) represents the output of the gating network; σ() represents the sigmoid function; W g represents the twelfth trainable parameter, W g ∈R 2d×1 ; b g represents the thirteenth trainable parameter, b g ∈R d ;
[0047] The linear classification module is used to calculate the original score s (i,j) with k dimensions based on the enhanced span information representation e c(i,j) :
[0048]
[0049] Among them, W c represents the fourteenth trainable parameter, W c ∈R d×k ; b c represents the fifteenth trainable parameter, b c ∈Rk ;
[0050] The span - to - span interaction module is used to enhance the span information representation e (i,j) Calculate the calibration score s of dimension k r(i,j) for the tensor S r :
[0051] S r = λ·Conv2d(W r , E);
[0052] where λ represents the scaling factor, 0 < λ ≤ 0.1; Conv2d() represents the two - dimensional convolution function; W r represents the sixteenth trainable parameter, W r ∈ R k×d×(2c+1)×(2c+1) , (2c + 1)×(2c + 1) represents the convolution kernel size, c represents the interaction size; E represents the tensor of the enhanced span information representation e (i,j) containing all spans, E ∈ R n×n×d ;
[0053] The second fusion module is used to sum the original score s c(i,j) and the calibration score s r(i,j) to obtain the fusion vector S C(i,j) :
[0054] S C(i,j) = S c(i,j) + S r(i,j) ;
[0055] For any span (i, j) and entity type, the fusion vector S C(i,j) contains the corresponding classification score S C(i,j,t) ;
[0056] The calculation formula of the output module is:
[0057] S (i,j,t) = S F(i,j) + S C(i,j,t) ;
[0058] where S (i,j,t) represents the entity score.
[0059] Furthermore, in step S2, the formula of the loss function is:
[0060]
[0061] where L t represents the loss value of the loss function; P t represents the positive sample set of entity type t; N t represents the negative sample set of entity type t;
[0062] The optimizer uses the AdamW optimizer.
[0063] Furthermore, step S3 is specifically as follows:
[0064] The data set is divided into a training set, a validation set, and a test set based on a preset ratio. The named entity recognition model is trained using the training set until the loss value of the loss function is less than a preset loss threshold.
[0065] The precision rate, recall rate, and micro-F1 value of the trained named entity recognition model are calculated using the validation set to verify the named entity recognition model. If the verification fails, the training set is expanded and training continues; if the verification passes, then:
[0066] The confidence level of the verified named entity recognition model is calculated using the test set to test the named entity recognition model.
[0067] The advantages of the present invention are as follows:
[0068] By collecting a large amount of text data, performing preprocessing on each text data including at least data cleaning and data verification, annotating entities in the preprocessed text data in the form of triples to construct a data set; then creating a named entity recognition model based on a sequence encoder, a span filter, a span classifier, and an output module, setting the loss function and optimizer of the named entity recognition model, then training the named entity recognition model using the data set, and finally performing a named entity recognition task using the trained named entity recognition model; the sequence encoder of the named entity recognition model is used to encode the input text data to obtain a filtered vector sequence and a classification vector sequence; the span filter is used to score using the filtered vector sequence to obtain a filtered score; the span classifier is used to score using the classification vector sequence to obtain a classification score; the output module is used to sum the filtered score and the classification score to obtain an entity score, and output a named entity recognition result based on the entity score; that is, by the sequence encoder, two groups of vector sequences with different features (filtered vector sequence, classification vector sequence) are respectively output to the span filter and the span classifier. The span filter is used to distinguish whether each span is an entity and output a filtered score, and the span classifier is used to classify and correct the entity and output a classification score, and then a named entity recognition result is output based on the filtered score and the classification score, making full use of the multiple information contained in the text span itself and the relationship between adjacent spans, greatly improving the feature extraction ability, and thus greatly improving the named entity recognition accuracy. Description of the Drawings
[0069] The present invention will be further described below with reference to the drawings in conjunction with embodiments.
[0070] Figure 1 The present invention is a flowchart of a method for named entity recognition based on multivariate information span scoring.
[0071] Figure 2 It is a structural diagram of the named entity recognition model of the present invention.
[0072] Figure 3 It is a schematic diagram of the span scoring matrix output by the present invention. DETAILED DESCRIPTION
[0073] The technical solution in the embodiment of the present application has the following overall idea: a named entity recognition model is created based on a sequence encoder, a span filter, a span classifier and an output module to perform a named entity recognition task, and two sets of vector sequences with different features are output to the span filter and the span classifier respectively through the sequence encoder, the span filter is used to distinguish whether each span is an entity and output a screening score, the span classifier is used to classify and correct the entity and output a classification score, and then the named entity recognition result is output based on the screening score and the classification score, that is, named entity recognition is performed based on the multivariate information span score, making full use of the multivariate information contained in the text span itself and the relationship between adjacent spans to improve the feature extraction capability, thereby improving the named entity recognition accuracy, and is simple and lightweight, has good parallelism, does not require any external auxiliary resources, has no domain knowledge requirements, and does not rely on auxiliary training methods.
[0074] Please refer to Figures 1 to 3 As shown, a preferred embodiment of a method for named entity recognition based on multivariate information span scoring of the present invention comprises the following steps:
[0075] Step S1, collecting a large amount of text data, performing preprocessing on each of the text data including at least data cleaning and data proofreading, and annotating entities (named entities) in each of the preprocessed text data in the form of triples to construct a data set; each entity is a continuous subsequence of the text data (text sequence) to which it belongs; the entity annotation list (which may be empty) obtained by annotation can be saved in a structured data format such as JSON or YAML;
[0076] Step S2, creating a named entity recognition model based on the sequence encoder, the span filter, the span classifier and the output module, setting the loss function and the optimizer of the named entity recognition model; the named entity recognition model is constructed based on the multivariate information span scoring mechanism;
[0077] The sequence encoder is used to encode the input text data to obtain a screening vector sequence and a classification vector sequence;
[0078] The span filter is used to score using a sequence of filter vectors to obtain a filter score, that is, to use span start and end token feature vectors that embed relative position information, calculate the dot product based on token pairs as the discrimination score (filter score) for entity / non-entity, to distinguish whether each span is a named entity, without entity classification;
[0079] The span classifier is used to score using a sequence of classification vectors to obtain a classification score, that is, to use a span information representation that fuses boundary and internal information, and calculate the span score (classification score) according to entity types;
[0080] The output module is used to sum the filter score and the classification score to obtain an entity score, and output the named entity recognition result based on the entity score;
[0081] Step S3: Train the named entity recognition model through the dataset;
[0082] Step S4: Execute the named entity recognition task through the trained named entity recognition model.
[0083] In the said step S1, the specific operation of annotating the entities in each of the preprocessed text data in the form of triples is as follows:
[0084] Annotate the entities in each of the preprocessed text data in the form of triples, and the content of the annotation is:
[0085] (i, j, t);
[0086] where i represents the index number of the starting character of the entity; j represents the index number of the ending character of the entity; t represents the entity type.
[0087] Suppose two entity spans are S 1 (i 1 , j 1 , t 1 ) and S 2 (i 2 , j 2 , t 2 ), and there are the following situations:
[0088] (1) If i 1 ≤ i 2 and j 2 ≤ j 1 , then S 2 is nested in S 1 ;
[0089] (2) If i 2 ≤ i 1 and j 1 ≤ j 2 , then S1 Nested in S 2 inside;
[0090] (3)i 1 <i 2 ≤j 1 <j 2 or i 2 <i 1 ≤j 2 <j 1 ,then S 1 and S 2 have boundary crossings, do not conform to the nesting rule, and are not marked.
[0091] In entity annotation, nesting is the only allowed case of overlapping spans. As Figure 3 shown, in the text sequence "severe lung lesions", "lung lesions" is an entity of the disease name, and within its span (3, 6), there is another nested anatomical site entity "lung", corresponding to the span (3, 3).
[0092] In the said step S2, the sequence encoder is composed of a backbone encoder, a first linear encoder, and a second linear encoder;
[0093] The backbone encoder is used to encode the input text data into a sequence of token vectors, with the formula:
[0094] H = PLM(X) = {h 1 , h 2 ,..., h n};
[0095] Among them, H represents the sequence of token vectors; PLM() represents the backbone encoder. Specifically, when implemented, BERT, RoBERTa, etc. can be selected; X represents the text data, and X = {x 1 , x 2 , …, x n}, x n represents the nth token in the text data; h n represents the nth token vector in the sequence of token vectors;
[0096] The first linear encoder is used to perform linear encoding on the sequence of token vectors to obtain a sequence of screening vectors:
[0097] H F = HW F + b F = {h F,1 , h F,2 , …, h F,n};
[0098] Among them, H F represents the sequence of screening vectors; WF Denote the first trainable parameter, \(W\) F \(\in\mathbb{R}\) D×d , where \(\mathbb{R}\) represents the set of real numbers, \(D\) represents the dimension of the token vectors, and \(d\) represents the size of the feature dimension; \(b\) F Denote the second trainable parameter, \(b\) F \(\in\mathbb{R}\) d ; \(h\) F,n Denote the \(n\)-th screening vector in the screening vector sequence;
[0099] The second linear encoder is used to linearly encode the token vector sequence to obtain a classification vector sequence:
[0100] \(H\) C \(=HW\) C \(+b\) C \(=\{h\) C,1 ,h\) C,2 ,\(\cdots,h\) C,n \(\}\);
[0101] where \(H\) C denotes the classification vector sequence; \(W\) C denotes the third trainable parameter, \(W\) C \(\in\mathbb{R}\) D×d ; \(b\) C denotes the fourth trainable parameter, \(b\) C \(\in\mathbb{R}\) d ; \(h\) C,n denotes the \(n\)-th classification vector in the classification vector sequence.
[0102] In step S2, the span filter consists of a third linear encoder, a fourth linear encoder, and a rotary position encoder; the span filter calculates a binary classification score (screening score) for each span, with the goal of determining whether each span is a named entity without making specific entity classifications;
[0103] Both the third linear encoder and the fourth linear encoder are used to encode the screening vector sequence to obtain a \(d\)-dimensional first feature vector \(h\) s,i and a \(d\)-dimensional second feature vector \(h\) e,j :
[0104]
[0105] where \(W\) s denotes the fifth trainable parameter, \(W\) s \(\in\mathbb{R}\) d×d ; \(T\) represents the transpose; \(h\) F,i denotes the first screening vector of any span \((i,j)\) in the screening vector sequence; \(b\) s denotes the sixth trainable parameter, \(b\) s \(\in\mathbb{R}\) d ; \(W\)e Denotes the seventh trainable parameter, W e ∈R d ×d ; h F,j Denotes the last screening vector of any span (i, j) in the screening vector sequence; b e Denotes the eighth trainable parameter, b e ∈R d ;
[0106] The rotation position encoder is used to calculate the screening score based on the first feature vector h s,i and the second feature vector h e,j Calculate the screening score:
[0107]
[0108] where, S F(i,j) Denotes the screening score of the screening vector corresponding to the span (i, j) in the screening vector sequence; Q i Denotes the rotation position encoding of token i; Q j Denotes the rotation position encoding of token j; Q j-1 Denotes the rotation position encoding of token j-1.
[0109] To make the span filter more sensitive to the length of the entity span, it is necessary to explicitly embed the relative position information using the rotation position encoding, that is, to satisfy
[0110] In the step S2, the span classifier consists of a bi-affine attention module, a position attention module, a first fusion module, a linear classification module, an inter-span interaction module, and a second fusion module; the span classifier regards the named entity recognition task as k independent and parallel binary classification tasks, and calculates k binary classification scores (classification scores) for each span, where k is the number of predefined entity types.
[0111] The bi-affine attention module is used to encode the classification vector sequence to obtain the cross-boundary information representation e B(i,j) :
[0112] e B(i,j) =[h C,i ; 1] T U[h C,j ; 1];
[0113] where, h C,i Denotes the first classification vector of any span (i, j) in the classification vector sequence; h C,j Denotes the last classification vector of any span (i, j) in the classification vector sequence; [.;.] denotes the vector concatenation operation; U denotes the ninth trainable parameter, U∈R (d+1)×d×(d+1) ;
[0114] The position attention module is used to encode the classification vector sequence to obtain the internal information representation e of the span I(i,j) :
[0115]
[0116] where α u represents the attention weight corresponding to the u-th classification vector in the classification vector sequence; h C,u represents the u-th classification vector in the classification vector sequence; a u and a v respectively represent the attention scores corresponding to the u-th and v-th classification vectors in the classification vector sequence; W I represents the tenth trainable parameter, W I ∈R d×1 ; b I represents the eleventh trainable parameter, b I ∈R d ;
[0117] The first fusion module is constructed based on a gated network and is used to perform weighted fusion on the cross-boundary information representation e B(i,j) and the internal information representation of the span to obtain the enhanced span information representation e (i,j) :
[0118] e (i,j) = g (i,j) e B(i,j) +(1 - g (i,j) )e I(i,j) ;
[0119]
[0120] where g (i,j) represents the output of the gated network; σ() represents the sigmoid function; W g represents the twelfth trainable parameter, W g ∈R 2d×1 ; b g represents the thirteenth trainable parameter, b g ∈R d ;
[0121] The linear classification module is used to calculate the original score s of k dimensions based on the enhanced span information representation e (i,j) : c(i,j) :
[0122]
[0123] where W c represents the fourteenth trainable parameter, W c∈R d×k ; b c represents the fifteenth trainable parameter, b c ∈R k ;
[0124] The span - to - span interaction module is used to enhance the span information representation e (i,j) Calculate the calibration score s of k - dimension, r(i,j) for the tensor S r :
[0125] S r = λ·Conv2d(W r , E);
[0126] where λ represents the scaling factor, 0 < λ ≤ 0.1; Conv2d() represents the two - dimensional convolution function, which is used to perform span - to - span interaction, capture the relationship information between adjacent spans, correct the original score, and effectively improve the accuracy of nested named entity recognition through the original score correction mechanism; zero padding is used in the convolution operation and no bias value is set; W r represents the sixteenth trainable parameter, W r ∈R k×d×(2c+1)×(2c+1) , (2c + 1)×(2c + 1) represents the convolution kernel size, c represents the interaction size, and the value of c is 1 or 2; E represents the tensor of the enhanced span information representation e (i,j) for all spans, E∈R n×n×d , for any invalid span (the head and tail indices are in reverse order: i > j), its e (i,j) is set to 0; the tensor S r has the calibration score s of k - dimension r(i,j) , representing its calibration scores under all k predefined entity types;
[0127] Spans that are spatially adjacent may carry information indicating each other's entity types, and such information can be used to learn the rules related to entity nesting; to make full use of the relationship information between adjacent spans, the two - dimensional convolution block Conv2d() is used as the span - to - span interactor;
[0128] The second fusion module is used to sum the original score s c(i,j) and the calibration score s r(i,j) to obtain the fusion vector S C(i,j) :
[0129] S C(i,j) = S c(i,j) + S r(i,j) ;
[0130] For any span (i, j) and entity type, the fusion vector S C(i,j) contains the corresponding classification score S C(i,j,t) ;
[0131] The calculation formula of the output module is as follows:
[0132] S (i,j,t) = S F(i,j) + S C(i,j,t) ;
[0133] Among them, S (i,j,t) represents the entity score.
[0134] S (i,j,t) reflects the confidence level of the named entity recognition model in determining the span (i, j) as a t-type entity; for all invalid spans, S (i,j,t) is set to negative infinity to facilitate the calculation of the loss value and the decoding of the named entity recognition result.
[0135] The entity score S (i,j,t) is also used to decode the named entity recognition result. A zero threshold needs to be set for decoding: Denote T' = T ∪ {0} = {t ∈ N | t ≤ k} as the extended entity type set, where T is the predefined entity type set, the type t = 0 represents "non-entity", and for any span (i, j), it is stipulated that s (i,j,0) = 0. In the text data of length n, the final entity scores of all spans form an n×n matrix, so the decoding process is parallel. The decoding rule is: if S (i,j,t) obtains the maximum positive value under a certain entity type (t ≠ 0), then the named entity recognition model recognizes the span as an entity of this type:
[0136]
[0137] Among them, represents the predicted type label of the span (i, j), that is, the named entity recognition result.
[0138] In step S2, the formula of the loss function is:
[0139]
[0140] Among them, L t represents the loss value of the loss function; P t represents the positive sample set of entity type t; N t represents the negative sample set of entity type t;
[0141] Since there are a total of spans in the text data of length n, and in the actual scenario, most spans are non-entities, there is a serious class imbalance problem. Therefore, the ZLPR loss function (Zero-bounded Log-sum-exp&Pairwise Rank-based Loss) is adopted to alleviate this problem.
[0142] In specific implementation, the training objective of the named entity recognition model is to minimize the mean of the ZLPR losses of all k types of entities:
[0143]
[0144] The optimizer adopts the AdamW optimizer.
[0145] The specific steps of step S3 are as follows:
[0146] Based on a preset ratio, the dataset is divided into a training set, a validation set, and a test set. The named entity recognition model is trained through the training set until the loss value of the loss function is less than a preset loss threshold. When dividing the dataset, it is necessary to ensure that the average sequence length, the proportion of each type of entity, and the average entity length among the subsets (training set, validation set, and test set) are basically the same;
[0147] The precision, recall, and micro-F1 score of the trained named entity recognition model are calculated through the validation set to verify the named entity recognition model. If the verification fails, the training set is expanded and training continues; if the verification passes, then:
[0148] The confidence of the verified named entity recognition model is calculated through the test set to test the named entity recognition model.
[0149]
[0150] Among them, P represents precision; R represents recall; F represents the micro-F1 score. The larger the value, the better the performance of the named entity recognition model.
[0151] In specific implementation, when passing text data into the named entity recognition model, it is necessary to set the length limit of the text data and truncate the overly long text data. An early stopping strategy can also be set for the training of the named entity recognition model, and combined with the change of the micro-F1 value of the named entity recognition model on the validation set, the best checkpoint in the training process is retained.
[0152] To sum up, the advantages of the present invention are as follows:
[0153] By collecting a large amount of text data, preprocessing each text data including at least data cleaning and data verification, annotating entities in each preprocessed text data in the form of triples to construct a dataset; then creating a named entity recognition model based on a sequence encoder, a span filter, a span classifier, and an output module, setting the loss function and optimizer of the named entity recognition model, then training the named entity recognition model through the dataset, and finally performing a named entity recognition task through the trained named entity recognition model; the sequence encoder of the named entity recognition model is used to encode the input text data to obtain a filtered vector sequence and a classification vector sequence; the span filter is used to score using the filtered vector sequence to obtain a filtered score; the span classifier is used to score using the classification vector sequence to obtain a classification score; the output module is used to sum the filtered score and the classification score to obtain an entity score, and output a named entity recognition result based on the entity score; that is, by the sequence encoder, two vector sequences with different features (filtered vector sequence, classification vector sequence) are respectively output to the span filter and the span classifier, the span filter is used to distinguish whether each span is an entity and output a filtered score, the span classifier is used to classify and correct the entity and output a classification score, and then a named entity recognition result is output based on the filtered score and the classification score, fully utilizing the multiple information contained in the text span itself and the relationship between adjacent spans, greatly improving the feature extraction ability, and thus greatly improving the named entity recognition accuracy.
[0154] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.
Claims
1. A method for named entity recognition based on multivariate information span scoring, characterized in that: The steps include: Step S1, collecting a large amount of text data, performing preprocessing on each of the text data including at least data cleaning and data proofreading, and annotating entities in each of the preprocessed text data in the form of triples to construct a data set; Step S2, creating a named entity recognition model based on the sequence encoder, the span filter, the span classifier and the output module, and setting the loss function and the optimizer of the named entity recognition model; The sequence encoder is used to encode the input text data to obtain a screening vector sequence and a classification vector sequence; The span filter is used to score using the screening vector sequence to obtain a screening score; The span classifier is used to perform scoring using the classification vector sequence to obtain a classification score; The output module is used to sum the screening score and the classification score to obtain an entity score, and output a named entity recognition result based on the entity score; Step S3, training a named entity recognition model using the data set; Step S4: performing a named entity recognition task using the trained named entity recognition model.
2. The method for named entity recognition based on multivariate information span scoring as claimed in claim 1, characterized in that: In the step S1, the entities in the preprocessed text data are labeled in the form of triples as follows: The entities in the preprocessed text data are annotated in the form of triples, and the annotated contents are: (i, j, t); Among them, i represents the index number of the first character of the entity; j represents the index number of the last character of the entity; and t represents the entity type.
3. The method for named entity recognition based on multivariate information span scoring as claimed in claim 1, characterized in that: In step S2, the sequence encoder is composed of a backbone encoder, a first linear encoder and a second linear encoder; The backbone encoder is used to encode the input text data into a word element vector sequence, and the formula is: H=PLM(X)={h1,h2,…,h n }; Where H represents the word vector sequence; PLM() represents the backbone encoder; X represents the text data, and X={x1,x2,…,x n }, x n Represents the nth word in the text data; h n Represents the nth word vector in the word vector sequence; The first linear encoder is used to linearly encode the word element vector sequence to obtain a screening vector sequence: H F =HW F +b F ={h F,1 ,h F,2 ,…,h F,n }; Among them, H F represents the screening vector sequence; W F represents the first trainable parameter, W F ∈R D×d , R represents a real number, D represents the dimension of the word element vector, and d represents the feature dimension size; b F represents the second trainable parameter, b F ∈R d ;h F,n represents the nth sifting vector in the sifting vector sequence; The second linear encoder is used to linearly encode the word element vector sequence to obtain a classification vector sequence: H C =HW C +b C ={h C,1 ,h C,2 ,…,h C,n }; Among them, H C represents the classification vector sequence; W C represents the third trainable parameter, W C ∈R D×d ; b C represents the fourth trainable parameter, b C ∈R d ;h C,n Represents the nth classification vector in a classification vector sequence.
4. A method for named entity recognition based on multivariate information span scoring as claimed in claim 3, characterized in that: In step S2, the span filter is composed of a third linear encoder, a fourth linear encoder and a rotary position encoder; The third linear encoder and the fourth linear encoder are both used to encode the screening vector sequence to obtain the first characterization vector h s,i and the second characterization vector h e,j : Among them, W s represents the fifth trainable parameter, W s ∈R d×d ; T represents transpose; h F,i represents the first screening vector of any span (i, j) in the screening vector sequence; b s represents the sixth trainable parameter, b s ∈R d ; W e represents the seventh trainable parameter, W e ∈R d×d ;h F,j represents the last screening vector of any span (i, j) in the screening vector sequence; b e represents the eighth trainable parameter, b e ∈R d ; The rotary position encoder is used to characterize the first vector h s,i and the second characterization vector h e,j Calculate filter score: Among them, S F(i,j) represents the screening score of the screening vector corresponding to span (i, j) in the screening vector sequence; Q i represents the rotational position encoding of word i; Q j represents the rotational position encoding of word j; Q j-1 Represents the rotational position encoding of word j-1.
5. The method for named entity recognition based on multivariate information span scoring as claimed in claim 4, characterized in that: In step S2, the span classifier is composed of a dual affine attention module, a position attention module, a first fusion module, a linear classification module, an inter-span interaction module and a second fusion module; The dual affine attention module is used to encode the classification vector sequence to obtain the cross-boundary information representation e B(i,j) : e B(i,j) =[h C,i ;1] T U[h C,j ;1]; Among them, h C,i Represents the first classification vector of any span (i, j) in the classification vector sequence; h C,j Represents the last classification vector of any span (i, j) in the classification vector sequence; [.;.] represents the vector concatenation operation; U represents the ninth trainable parameter, U∈R (d +1)×d×(d+1) ; The position attention module is used to encode the classification vector sequence to obtain the span internal information representation e I(i,j) : Among them, α u represents the attention weight corresponding to the u-th classification vector in the classification vector sequence; h C,u represents the u-th classification vector in the classification vector sequence; a u and a v Respectively represent the attention scores corresponding to the u-th and v-th classification vectors in the classification vector sequence; W I represents the tenth trainable parameter, W I ∈R d×1 ; b I represents the eleventh trainable parameter, b I ∈R d ; The first fusion module is constructed based on a gated network and is used to represent the cross-boundary information. B(i,j) And the span internal information representation is weighted fused to obtain the enhanced span information representation e (i,j) : yes (i,j) =g (i,j) yes B(i,j) +(1-g (i,j) )e I(i,j) ; Among them, g (i,j) represents the output of the gating network; σ() represents the sigmoid function; W g represents the twelfth trainable parameter, W g ∈R 2d×1 ; b g represents the thirteenth trainable parameter, b g ∈R d ; The linear classification module is used to represent e according to the enhanced span information. (i,j) Calculate the original score s of k dimensions c(i,j) : Among them, W c represents the fourteenth trainable parameter, W c ∈R d×k ; b c represents the fifteenth trainable parameter, b c ∈R k ; The span interaction module is used to enhance the span information representation. (i,j) Calculate the corrected score s with k dimensions r(i,j) The tensor S r : S r =λ·Conv2d(W r ,E); Where λ represents the scaling factor, 0<λ≤0.1; Conv2d() represents the two-dimensional convolution function; W r represents the sixteenth trainable parameter, W r ∈R k×d×(2c+1)×(2c+1) , (2c+1)×(2c+1) represents the convolution kernel size, c represents the interaction size; E represents the enhanced span information representation containing all spans. (i,j) Tensor, E∈R n×n×d ; The second fusion module is used to s c(i,j) and correction scores r(i,j) Sum and get the fusion vector S C(i,j) : S C(i,j) =S c(i,j) +S r(i,j) ; For any span (i, j) and entity type, the fusion vector S C(i,j) Contains the corresponding classification score S C(i,j,t) ; The calculation formula of the output module is: S (i,j,t) =S F(i,j) +S C(i,j,t) ; Among them, S (i,j,t) Represents an entity rating.
6. The method for named entity recognition based on multivariate information span scoring according to claim 1, characterized in that: In step S2, the formula of the loss function is: Among them, L t Represents the loss value of the loss function; P t Represents the positive sample set of entity type t; N t Represents the negative sample set of entity type t; The optimizer adopts AdamW optimizer.
7. The method for named entity recognition based on multivariate information span scoring according to claim 1, characterized in that: The step S3 is specifically as follows: Dividing the data set into a training set, a validation set, and a test set based on a preset ratio, and training the named entity recognition model through the training set until the loss value of the loss function is less than a preset loss threshold; The precision, recall and micro-F1 value of the trained named entity recognition model are calculated through the verification set to verify the named entity recognition model. If the verification fails, the training set is expanded to continue training; if the verification passes, then: The confidence of the verified named entity recognition model is calculated using the test set to test the named entity recognition model.