Sequence labeling optimization method and system, computer equipment and medium
By introducing entity boundary offset and smoothing processing technology, the position deviation problem of the sequence annotation model when identifying entity boundaries is solved, which significantly improves the accuracy and data quality of entity recognition.
Patent Information
- Application Number
- CN202510069995.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The sequence annotation model has position deviations from the real entity boundary when identifying entity boundaries, and the generation of offset sequences may be disturbed by noise, resulting in data fluctuations and deviations.
The concept of entity boundary offset is introduced, and noise interference is reduced through smoothing processing technology, data quality is improved, and candidate span is generated using the smoothed offset, and entity boundaries in the label sequence are adjusted.
It significantly improves the accuracy of entity recognition, reduces the impact of noise on data, and improves the accuracy of entity boundaries.
Smart Images

Figure CN119990139A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer application and natural language processing, and in particular to a sequence annotation optimization method, system, computer equipment and medium. Background Art
[0002] Named entity recognition (NER) is an important basic task in the field of natural language processing (NLP). The purpose of named entity recognition is to identify the type and boundary of entities in a given text. Entity types mainly include time, personal names, place names, institution names, etc. The performance of named entity recognition directly affects the performance and efficiency of downstream tasks such as knowledge graphs, search engines, and question-answering systems.
[0003] The sequence labeling model can output the labeling path with the highest probability, capture the global structural information of the sentence, and perform well in the named entity recognition task. However, the sequence labeling model only relies on the original input sentence, and does not introduce any additional information about the entity boundary when identifying the entity type and boundary. The related study "Sequence labeling optimization method combined with entity boundary offset" proposed a sequence labeling optimization method combined with entity boundary offset, aiming to solve the problem of position deviation between the entity boundary identified by the sequence labeling model and the real entity boundary. By introducing the offset of the entity boundary, the semantic dependency relationship between the entity boundary and its context is constructed, so as to obtain an entity boundary closer to the real boundary, optimize the entity recognition results output by the sequence labeling model, and improve the performance of entity recognition. However, in the process of generating the offset sequence, it may be interfered by noise, resulting in deviations caused by fluctuations in the offset sequence data.
[0004] To solve this problem, the present invention is based on the fact that the position offsets of adjacent tags are regular and symmetrical, which imply the spatial shape of named entities, which can be used to construct entity position dependencies in sentences. Inspired by the noise elimination in image smoothing, the sliding window is designed, and the value in it is obtained by a specified calculation method, and it is updated to the current value. The offset information is captured around the entity boundary, and the data with excessive deviation is smoothed to reduce the impact of noise and improve the accuracy of the entity boundary. Summary of the invention
[0005] The purpose of the present invention is to provide a sequence labeling optimization method, system, computer equipment and medium to solve the problem of position deviation between the sequence labeling model and the real entity boundary when identifying the entity boundary. To solve this problem, the present invention introduces the concept of offset to quantify this boundary deviation and associate the positional relationship between each word and the entity boundary, thereby providing richer entity boundary information. It is particularly critical that the present invention adopts a smoothing processing technology. This technology effectively reduces the interference of noise and improves data quality by smoothing the entity boundary offset sequence. On this basis, candidate spans are generated according to the smoothed entity boundary offsets, and these candidate spans are further used to adjust the entity boundaries in the label sequence, thereby significantly improving the accuracy of entity recognition.
[0006] To achieve the above object, the present invention provides a sequence annotation optimization method, comprising the following steps:
[0007] S1, preprocessing the text data set and constructing the preprocessed data set;
[0008] S2, obtain the word vectors in the preprocessed data set;
[0009] S3, input the obtained word vector into the label classifier and two different offset classifiers at the same time, and obtain the label sequence and offset sequence respectively;
[0010] S4, extracting entity span sets based on the acquired tag sequence;
[0011] S5. Based on the obtained offset sequence, a smoothing process is performed to smooth the noise in the offset sequence, and a candidate span set is extracted;
[0012] S6. Filter out low-quality candidate spans through the intersection-and-union strategy to obtain filtered candidate spans;
[0013] S7. Based on the filtered candidate spans, update the corresponding entity spans in the tag sequence.
[0014] Preferably, in step S2, word vectors in the preprocessed data are obtained, and the specific operations are as follows:
[0015] S21. For sentences in the preprocessed dataset, given an input sequence X, use BERT to obtain the deep semantic features in the input sequence X and obtain the word embedding representation of each word. in, Represents the word vector of the i-th word;
[0016] S22. Use BiLSTM to further process the output of BERT and capture the bidirectional semantic dependencies in the sentence to obtain the context representation H = {h 1,h 2 ,...,h n}, where h i ∈R d Represents the word vector, and d represents the dimension of the word vector;
[0017] The formal representation of step S21 and step S22 is:
[0018]
[0019] in, Represent the forward and backward hidden states of BiLSTM respectively; [;] represents the concatenation operation.
[0020] Preferably, in step S3, the obtained word vector is input into a label classifier to obtain a label sequence, and the specific operation is as follows:
[0021] The word vector h i In the input label classifier, the word vector h i Convert from high dimension to the dimension corresponding to the number of labels, and use the softmax layer to obtain the word vector h for each word i The probability matrix decodes the score of the label sequence and converts the word vector h of each word i Finally mapped to a label set:
[0022]
[0023] Among them, the label sequence is y, which is expressed as: Represents the i-th word x i Label: MLP 1 Represents a multilayer perceptron, i.e., a label classifier, consisting of two layers of Linear and GELU activation functions.
[0024] Preferably, in step S3, the obtained word vector is input into two different offset classifiers to obtain an offset sequence, and the specific operation is as follows:
[0025] The word vector h i Input into two different offset classifiers to generate the offset of the entity start position and the offset of the entity end position respectively, and use the softmax layer to obtain the word vector h of each word i The probability of the offset relative to the entity boundary position, and finally, the highest scoring offset sequence is decoded:
[0026]
[0027] in, and Respectively represent the offset of the entity start position and the offset of the entity end position; MLP 2 and MLP 3 Both represent multi-layer perceptrons, i.e., offset classifiers, and their structures are the same as MLP 1 Consistent.
[0028] Preferably, in step S4, the entity span set is extracted based on the acquired tag sequence, and the specific operations are as follows:
[0029] Extract entity span set S = {s 0 ,s 1 ,...,s m}, where m represents the number of spans, s q =(yb q ,ye q ) represents the span of the qth entity, yb q and ye q Respectively represent the start and end positions of the entity span.
[0030] Preferably, in step S5, the obtained offset sequence is smoothed to smooth the noise in the offset sequence, and a candidate span set is extracted. The specific operations are as follows:
[0031] S51, the entity boundary offset sequence and It is regarded as a one-dimensional digital signal for smoothing, and the entity boundary offset sequence is smoothed using the sliding window method;
[0032] S52, treat the smoothing process as a sliding window operation, using the smoothing check window (left 2 ,left 1 , mid) to decide whether the mid in the current window needs to be smoothed. When the smoothing check window finds abnormal data fluctuations, the smoothing calculation window (left 2 ,left 1 ,mid,right 1 ,right 2 ) Calculate the new value of mid at the outlier point;
[0033] Among them, mid is the smoothing value of the window, left 1 is the previous value of the smoothed value, left 2 left 1 The previous value of right 1 is the value after the smoothing value, right 2 for right 1 The next value of
[0034] S521, adding edge values to the beginning and end of the entity boundary offset sequence to prevent the window from crossing the boundary, the beginning edge value is the first value of the sequence, and the end edge value is the last value of the sequence;
[0035] S522, in the smoothing check window sliding, the window is smoothed from left to right, starting from the data with the entity boundary offset sequence index position 1. In this process, it is assumed that the data that has passed through the window is consistent with the real offset, and v is used. 1 =|left 1 -mid| and v 2 =|left 2 -mid|Two variables to determine whether the data fluctuates;
[0036] S523, when v 1 The value range is not [0,1] or v 2 If the value is not 0 or 2, it is considered as a numerical anomaly, and a new smoothing value is calculated using the smoothing window at the abnormal window.
[0037] Determine whether left1±1==right1±1 holds. If the equation holds, the smoothing value is the value that makes the equation hold.
[0038] If the equality does not hold, the smoothing value is:
[0039]
[0040] Entity boundary offset sequence O b and O e After smoothing, the new sequence entity boundary offset O is obtained. b′ and O e′ , from O b′ and O e′ Extract the entity start position and entity end position information and combine them into an offset span;
[0041] S53, let two sets B = {b 1 ,b 2 ,...,b c} and E={e 1 ,e 2 ,...,e r}, respectively O b′ and O e′ The position index set with offset 0, where 0≤c≤n, 0≤r≤n;
[0042] S54, by enumerating all candidate start positions and candidate end positions in sets B and E, a candidate span set D of offsets of entity boundaries is formed = {d 1 ,d 2 ,...,dk}, where d w =(b w ,e j ) represents the w-th candidate span of the enumeration, 1≤w≤c, 1≤j≤r, b w represents the wth candidate start position in the candidate span set, e j represents the jth candidate end position in the candidate span set, and k=c*r represents the number of candidate spans.
[0043] Preferably, in step S6, low-quality candidate spans are filtered out by using an intersection-over-union strategy to obtain filtered candidate spans. The specific operations are as follows:
[0044] Calculate the intersection-over-union ratio:
[0045]
[0046] Among them, min(e j ,ye q ) represents the candidate span d w and entity span s q The minimum value of the end position in the w ,yb q ) represents the candidate span d w and entity span s q The maximum value of the starting position in the j ,ye q ) represents the candidate span d w and entity span s q The maximum value of the end position in the w ,yb q ) represents the candidate span d w and entity span s q The minimum value of the starting position in ;
[0047] Based on the calculated intersection-and-union ratio, the candidate span with an intersection-and-union ratio greater than 0.8 is selected as the final result of the entity span.
[0048] The present invention also provides a sequence annotation optimization system, comprising:
[0049] The data preprocessing module is used to preprocess the text data set and construct the preprocessed data set;
[0050] The word vector acquisition module is used to obtain the word vectors in the preprocessed data set;
[0051] The sequence decoding and offset classification module inputs the word vector into the sequence decoding module to obtain the label sequence, and simultaneously inputs it into two offset classifiers to obtain the offset sequence;
[0052] The entity span and candidate span extraction module extracts the entity span set based on the label sequence, performs smoothing based on the obtained offset sequence, smoothes the noise in the offset sequence, and then extracts the candidate span set;
[0053] The candidate span filtering module filters out low-quality candidate spans through the intersection-and-union strategy to obtain filtered candidate spans;
[0054] The entity span update module updates the corresponding entity spans in the label sequence based on the filtered candidate spans.
[0055] The present invention also provides a computer device, comprising: a memory and a processor; the memory stores a computer program, and the processor implements the steps of the above-mentioned sequence labeling optimization method when executing the computer program.
[0056] The present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the above-mentioned sequence labeling optimization method are implemented.
[0057] Therefore, the present invention adopts the above-mentioned sequence labeling optimization method, system, computer device and medium, which have beneficial technical effects: the contextual information and semantic ambiguity of entity boundaries in the text cause the sequence labeling model to have a deviation between the recognized entity boundaries and the actual entity boundaries, and the concept of entity boundary offset is introduced to quantify the deviation; the sequence labeling model has limitations in perceiving entity boundaries, and the boundary offset is used as a connection to construct a semantic dependency relationship between the entity boundary and its context, and capture more boundary information to enhance the sequence labeling model's ability to recognize entity boundaries; the F1 values of this method on the CLUENER2020, Resume-zh and MSRA datasets reached 80.61, 96.53 and 94.96, respectively, verifying the versatility and effectiveness of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 A flowchart of a sequence annotation optimization method of the present invention;
[0059] Figure 2 Entity recognition model diagram;
[0060] Figure 3 For adjustment example diagram. DETAILED DESCRIPTION
[0061] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.
[0062] Unless otherwise defined, technical or scientific terms used in the present invention shall have the common meanings understood by one having ordinary skills in the field to which the present invention belongs.
[0063] Embodiment 1
[0064] like Figure 1 FIG. 1 is a flow chart of a sequence annotation optimization method of the present invention, comprising the following steps:
[0065] S1. Preprocess the text dataset to construct a preprocessed dataset; preprocessing includes word segmentation and named entity tagging.
[0066] S2. Get the word vectors in the preprocessed data set. The specific operations are as follows:
[0067] S21. For sentences in the preprocessed dataset, given an input sequence X, use BERT to obtain the deep semantic features in the input sequence X and obtain the word embedding representation of each word. in, Represents the word vector of the i-th word;
[0068] S22. Use BiLSTM to further process the output of BERT and capture the bidirectional semantic dependencies in the sentence to obtain the context representation H = {h 1 ,h 2 ,...,h n}, where h i ∈R d Represents the word vector, and d represents the dimension of the word vector;
[0069] The formal representation of step S21 and step S22 is:
[0070]
[0071] in, They represent the forward and backward hidden states of BiLSTM respectively, and [;] represents the concatenation operation.
[0072] S3, input the obtained word vector into the label classifier and two different offset classifiers at the same time, and obtain the label sequence and offset sequence respectively;
[0073] Get the tag sequence. The specific operations are as follows:
[0074] The word vector h i In the input label classifier, the word vector h i Convert from high dimension to the dimension corresponding to the number of labels, and use the softmax layer to obtain the word vector h for each word i The probability matrix decodes the score of the label sequence and converts the word vector h of each word iFinally mapped to a label set:
[0075]
[0076] Among them, the label sequence is y, which is expressed as: Represents the i-th word x i Label: MLP 1 Represents a multilayer perceptron, i.e., a label classifier, consisting of two layers of Linear and GELU activation functions.
[0077] Get the offset sequence. The specific operations are as follows:
[0078] The word vector h i Input into two different offset classifiers to generate the offset of the entity start position and the offset of the entity end position respectively, and use the softmax layer to obtain the word vector h of each word i The probability of the offset relative to the entity boundary position, and finally, the highest scoring offset sequence is decoded:
[0079]
[0080] in, and Respectively represent the offset of the entity start position and the offset of the entity end position; MLP 2 and MLP 3 Both represent multi-layer perceptrons, i.e., offset classifiers, and their structures are the same as MLP 1 Consistent.
[0081] S4. Extract entity span sets based on the acquired tag sequence. The specific operations are as follows:
[0082] Extract entity span set S = {s 0 ,s 1 ,...,s m}, where m represents the number of spans, s q =(yb q ,ye q ) represents the span of the qth entity, yb q and ye q Respectively represent the start and end positions of the entity span.
[0083] S5. Based on the obtained offset sequence, smoothing processing is performed to smooth the noise in the offset sequence and extract the candidate span set. The specific operations are as follows:
[0084] S51, the entity boundary offset sequence and It is regarded as a one-dimensional digital signal for smoothing, and the entity boundary offset sequence is smoothed using the sliding window method;
[0085] S52, treat the smoothing process as a sliding window operation, using the smoothing check window (left 2 ,left 1 , mid) to decide whether the mid in the current window needs to be smoothed. When the smoothing check window finds abnormal data fluctuations, the smoothing calculation window (left 2 ,left 1 ,mid,right 1 ,right 2 ) Calculate the new value of mid at the outlier point;
[0086] Among them, mid is the smoothing value of the window, left 1 is the previous value of the smoothed value, left 2 left 1 The previous value of right 1 is the next value of the smoothed value, right 2 for right 1 The next value of
[0087] S521, adding edge values to the beginning and end of the entity boundary offset sequence to prevent the window from crossing the boundary, the beginning edge value is the first value of the sequence, and the end edge value is the last value of the sequence;
[0088] S522, in the smoothing check window sliding, the window is smoothed from left to right, starting from the data with the entity boundary offset sequence index position 1. In this process, it is assumed that the data that has passed through the window is consistent with the real offset, and v is used. 1 =|left 1 -mid| and v 2 =|left 2 -mid|Two variables to determine whether the data fluctuates;
[0089] S523, when v 1 The value range is not [0,1] or v 2 If the value is not 0 or 2, it is considered as a numerical anomaly, and a new smoothing value is calculated using the smoothing window at the abnormal window.
[0090] Determine whether left1±1==right1±1 holds. If the equation holds, the smoothing value is the value that makes the equation hold.
[0091] If the equality does not hold, the smoothing value is:
[0092]
[0093] Entity boundary offset sequence O b and O e After smoothing, the new sequence entity boundary offset O is obtained. b′ and O e′ , from O b′ and O e′ Extract the entity start position and entity end position information and combine them into an offset span;
[0094] S53, let two sets B = {b 1 ,b 2 ,...,b c} and E={e 1 ,e 2 ,...,e r}, respectively O b′ and O e′ The position index set with offset 0, where 0≤c≤n, 0≤r≤n;
[0095] S54, by enumerating all candidate start positions and candidate end positions in sets B and E, a candidate span set D of offsets of entity boundaries is formed = {d 1 ,d 2 ,...,d k}, where d w =(b w ,e j ) represents the w-th candidate span of the enumeration, 1≤w≤c, 1≤j≤r, b w represents the wth candidate start position in the candidate span set, e j represents the jth candidate end position in the candidate span set, and k=c*r represents the number of candidate spans.
[0096] S6. Filter out low-quality candidate spans through the intersection-and-union strategy to obtain filtered candidate spans. The specific operations are as follows:
[0097] Calculate the intersection-over-union ratio:
[0098]
[0099] Among them, min(e j ,ye q ) represents the candidate span d w and entity span s q The minimum value of the end position in the w ,yb q ) represents the candidate span d w and entity span s q The maximum value of the starting position in the j ,yeq ) represents the candidate span d w and entity span s q The maximum value of the end position in the w ,yb q ) represents the candidate span d w and entity span s q The minimum value of the starting position in ;
[0100] Based on the calculated intersection-and-union ratio, the candidate span with an intersection-and-union ratio greater than 0.8 is selected as the final result of the entity span.
[0101] S7. Based on the filtered candidate spans, the corresponding entity spans in the tag sequence are updated. The specific operations are as follows:
[0102] S71. Create an empty new label sequence with the same length as the label sequence. For each candidate span, find its corresponding position in the new label sequence. Update the label of this position to the entity category label corresponding to the candidate span, and update all other positions to non-entity "O".
[0103] S72. Finally, output an updated tag sequence.
[0104] The present invention will be further described below through specific examples.
[0105] like Figure 2 Figure 1 is an entity recognition model diagram. First, the abstract semantic representation of each word in the sentence is obtained through the pre-trained language model, and the context encoder is used to obtain the sentence context semantic information to obtain the word vector. Then, these word vectors are used to generate the label sequence and offset sequence. Next, the offset sequence is smoothed and the boundary information in these offsets is combined to generate candidate spans, and the low-quality candidate spans are filtered out by the intersection-over-union method, and finally the candidate spans that are most likely to represent the entity boundary are left. Finally, these candidate spans are used to update the corresponding entity spans in the label sequence.
[0106] like Figure 3An example diagram for the adjustment of the invention. The input sequence "Nanjing Yangtze River Bridge" contains two entities, "Nanjing City" and "Yangtze River Bridge", and its true label should be {B, I, E, B, I, I, E}. However, there is a certain ambiguity in the boundaries of these two entities at the semantic level. The end boundary "City" of the entity "Nanjing City" and the start boundary "Yangtze" of the entity "Yangtze River Bridge" can combine to form a new semantics "Mayor". According to the context information in the sentence, the word "Nanjing" can also combine with "Mayor", further complicating the determination of the boundary. When the sequence labeling model predicts the entity boundary, due to the dual influence of semantic ambiguity relationships and complex context information, the model incorrectly labels "Nanjing Mayor" and "River Bridge" as entities {B, I, I, E, B, I, E}, resulting in a deviation Δ in the division of the entity boundary. The present invention introduces the concept of entity boundary offset to quantify this boundary deviation. Regarding the start boundary or end boundary of the entity as the starting reference point of the boundary offset, using the index values of these starting points in the sentence as the benchmark, calculate the index offset of each non-entity boundary word from the nearest starting point, and use this as the quantitative expression of the position characteristics of each word and the entity boundary, thereby establishing the semantic dependence relationship between the entity boundary and its context. "Nan" and "Chang" are the start boundaries of the entities "Nanjing City" and "Yangtze River Bridge", and their offsets are the starting points B. The offset of the non-boundary word is the minimum of the absolute index differences from the entity boundaries "Nan" or "Chang", thus obtaining the offset sequence {B, B - 1, B + 1, B, B - 1, B - 2, B - 3}, where the index differences between the non-entity boundary word "City" and the entity boundaries "Nan" and "Chang" are -2 and 1 respectively, and the minimum of the absolute index differences is 1, so the boundary offset representing its position is B + 1. Similarly, taking the entity end words "City" and "Bridge" as the starting points, the offset sequence for the entity end offset can be obtained {E + 2, E + 1, E, E - 1, E + 2, E + 1, E}. The offset sequence can provide more entity boundary features for the model, clarify the word boundaries, and avoid the fuzzy boundaries caused by semantic ambiguity.
[0107] This embodiment conducts experiments on three Chinese datasets, CLUENER2020, Resume-zh, and MSRA. The CLUENER2020 dataset is a Chinese dataset screened based on the open-source text classification dataset THUCNEWS. The entity types include 10 categories: game, organization, government, movie, name, book, company, scene, position, and address.
[0108] Resume-zh is a Chinese entity recognition dataset for resumes.
[0109] MSRA is a Chinese named entity recognition dataset for the news field. The entity types include three categories: place names (LOC), names (NAME), and organizations (ORG).
[0110] Since the datasets CLUENER2020 and MSRA do not have a validation set, 10% of the training set is randomly selected to form a validation set to adjust the model parameters.
[0111] The evaluation indicators used in the experiment include precision P, recall R and measurement value F1. Among them, P represents the ratio of correctly identified entities to all identified entities, R represents the ratio of correctly identified entities to entities that should be identified, and F1 is a comprehensive evaluation indicator that combines P and R. In order to verify the effectiveness of the model, the model proposed in the present invention is experimentally compared with other models on different data sets, and the comparison results are shown in Table 1.
[0112] Table 1 Comparison of experimental results of various models on different data sets Unit: %
[0113]
[0114] This invention is based on the paper "Sequence Labeling Optimization Method Combined with Entity Boundary Offset" and adds a smoothing module on the basis of the original text. The generated offset sequence may be disturbed by noise, resulting in data fluctuations and deviations. Because the position offsets of adjacent tags are regular and symmetrical, they imply the spatial shape of the named entity, which can be used to construct entity position dependencies in sentences. Therefore, an effective boundary offset update algorithm and a smoothing algorithm are designed relative to the true offset to reduce the impact of noise and improve the regression accuracy of entity boundaries.
[0115] Embodiment 2
[0116] A sequence labeling optimization system, comprising:
[0117] The data preprocessing module is used to preprocess the text data set and construct the preprocessed data set;
[0118] The word vector acquisition module is used to obtain the word vectors in the preprocessed data set;
[0119] The sequence decoding and offset classification module inputs the word vector into the sequence decoding module to obtain the label sequence, and simultaneously inputs it into two offset classifiers to obtain the offset sequence;
[0120] The entity span and candidate span extraction module extracts the entity span set based on the label sequence, performs smoothing based on the obtained offset sequence, smoothes the noise in the offset sequence, and then extracts the candidate span set;
[0121] The candidate span filtering module filters out low-quality candidate spans through the intersection-and-union strategy to obtain filtered candidate spans;
[0122] The entity span update module updates the corresponding entity spans in the label sequence based on the filtered candidate spans.
[0123] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.
[0124] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0125] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0126] It is worth noting that the contents not elaborated in detail in the present invention are all prior art and are well known to those skilled in the art.
[0127] Therefore, the present invention adopts the above-mentioned sequence annotation optimization method, system, computer device and medium to improve the accuracy of named entity recognition.
[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A sequence annotation optimization method, characterized in that: The following steps are involved: S1, preprocessing the text data set and constructing the preprocessed data set; S2, obtain the word vectors in the preprocessed data set; S3, input the obtained word vector into the label classifier and two different offset classifiers at the same time, and obtain the label sequence and offset sequence respectively; S4, extracting entity span sets based on the acquired tag sequence; S5. Based on the obtained offset sequence, a smoothing process is performed to smooth the noise in the offset sequence, and a candidate span set is extracted; S6. Filter out low-quality candidate spans through the intersection-and-union strategy to obtain filtered candidate spans; S7. Based on the filtered candidate spans, update the corresponding entity spans in the tag sequence.
2. A sequence annotation optimization method according to claim 1, characterized in that: In step S2, the word vectors in the preprocessed data are obtained. The specific operations are as follows: S21. For sentences in the preprocessed dataset, given an input sequence X, use BERT to obtain the deep semantic features in the input sequence X and obtain the word embedding representation of each word. in, Represents the word vector of the i-th word; S22. Use BiLSTM to further process the output of BERT and capture the bidirectional semantic dependencies in the sentence to obtain the context representation H = {h1,h2,...,h n }, where h i ∈R d Represents the word vector, and d represents the dimension of the word vector; The formal representation of step S21 and step S22 is: in, Represent the forward and backward hidden states of BiLSTM respectively; [;] represents the concatenation operation.
3. A sequence annotation optimization method according to claim 2, characterized in that: In step S3, the obtained word vector is input into the label classifier to obtain the label sequence. The specific operations are as follows: The word vector h i In the input label classifier, the word vector h i Convert from high dimension to the dimension corresponding to the number of labels, and use the softmax layer to obtain the word vector h for each word i The probability matrix decodes the score of the label sequence and converts the word vector h of each word i Finally mapped to a label set: Among them, the label sequence is y, which is expressed as: Represents the i-th word x i Label; MLP1 represents a multi-layer perceptron, that is, a label classifier, which consists of two layers of Linear and GELU activation functions.
4. A sequence annotation optimization method according to claim 3, characterized in that: In step S3, the obtained word vector is input into two different offset classifiers to obtain an offset sequence. The specific operations are as follows: The word vector h i Input into two different offset classifiers to generate the offset of the entity start position and the offset of the entity end position respectively, and use the softmax layer to obtain the word vector h of each word i The probability of the offset relative to the entity boundary position, and finally, the highest scoring offset sequence is decoded: in, and They represent the offset of the entity's start position and the offset of the entity's end position respectively; MLP2 and MLP3 both represent multi-layer perceptrons, i.e., offset classifiers, and their structures are consistent with MLP1.
5. A sequence annotation optimization method according to claim 4, characterized in that: In step S4, the entity span set is extracted based on the acquired tag sequence. The specific operations are as follows: Extract entity span set S = {s0, s1, ..., s m }, where m represents the number of spans, s q =(yb q ,ye q ) represents the span of the qth entity, yb q and ye q Respectively represent the start and end positions of the entity span.
6. A sequence annotation optimization method according to claim 5, characterized in that: In step S5, the obtained offset sequence is smoothed to smooth the noise in the offset sequence, and a candidate span set is extracted. The specific operations are as follows: S51, the entity boundary offset sequence and It is regarded as a one-dimensional digital signal for smoothing, and the entity boundary offset sequence is smoothed using the sliding window method; S52, the smoothing process is regarded as a sliding window operation, and the smoothing check window (left2, left1, mid) is used to determine whether mid in the current window needs to be smoothed. When the smoothing check window finds abnormal data fluctuations, the smoothing calculation window (left2, left1, mid, right1, right2) is used to calculate the new value of mid at the abnormal point; Among them, mid is the smoothing value of the window, left1 is the previous value of the smoothing value, left2 is the previous value of left1, right1 is the next value of the smoothing value, and right2 is the next value of right1; S521, adding edge values to the beginning and end of the entity boundary offset sequence to prevent the window from crossing the boundary, the beginning edge value is the first value of the sequence, and the end edge value is the last value of the sequence; S522, in the smoothing check window sliding, the window is smoothed from left to right, starting from the data with the entity boundary offset sequence index position 1. In this process, it is assumed that the data that has passed through the window is consistent with the real offset, and two variables v1=|left1-mid| and v2=|left2-mid| are used to determine whether the data fluctuates; S523, when the value range of v1 is not [0,1] or the value of v2 is not 0 or 2, it is judged as a numerical anomaly, and a new smoothing value at the abnormal window is calculated using a smoothing window; Determine whether left1±1==right1±1 holds. If the equation holds, the smoothing value is the value that makes the equation hold. If the equality does not hold, the smoothing value is: Entity boundary offset sequence O b and O e After smoothing, the new sequence entity boundary offset O is obtained. b′ and O e′ , from O b′ and O e′ Extract the entity start position and entity end position information and combine them into an offset span; S53, let two sets B = {b1, b2, ..., b c } and E={e1,e2,...,e r }, respectively O b′ and O e′ The position index set with offset 0, where 0≤c≤n, 0≤r≤n; S54, by enumerating all candidate start positions and candidate end positions in sets B and E, a candidate span set D of offsets of entity boundaries is formed, namely, {d1, d2, ..., d k }, where d w =(b w ,e j ) represents the w-th candidate span of the enumeration, 1≤w≤c, 1≤j≤r, b w represents the wth candidate start position in the candidate span set, e j represents the jth candidate end position in the candidate span set, and k=c*r represents the number of candidate spans.
7. A sequence annotation optimization method according to claim 6, characterized in that: In step S6, low-quality candidate spans are filtered out by using the intersection-over-union strategy to obtain filtered candidate spans. The specific operations are as follows: Calculate the intersection-over-union ratio: Among them, min(e j ,ye q ) represents the candidate span d w and entity span s q The minimum value of the end position in the w ,yb q ) represents the candidate span d w and entity span s q The maximum value of the starting position in the j ,ye q ) represents the candidate span d w and entity span s q The maximum value of the end position in the w ,yb q ) represents the candidate span d w and entity span s q The minimum value of the starting position in ; Based on the calculated intersection-and-union ratio, the candidate span with an intersection-and-union ratio greater than 0.8 is selected as the final result of the entity span.
8. A sequence annotation optimization system, characterized in that: include: The data preprocessing module is used to preprocess the text data set and construct the preprocessed data set; The word vector acquisition module is used to obtain the word vectors in the preprocessed data set; The sequence decoding and offset classification module inputs the word vector into the sequence decoding module to obtain the label sequence, and simultaneously inputs it into two offset classifiers to obtain the offset sequence; The entity span and candidate span extraction module extracts the entity span set based on the label sequence, performs smoothing based on the obtained offset sequence, smoothes the noise in the offset sequence, and then extracts the candidate span set; The candidate span filtering module filters out low-quality candidate spans through the intersection-and-union strategy to obtain filtered candidate spans; The entity span update module updates the corresponding entity spans in the label sequence based on the filtered candidate spans.
9. A computer device comprising: Memory and processor; The memory stores a computer program, wherein the processor implements the steps of the sequence annotation optimization method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the sequence annotation optimization method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Social media multi-modal named entity identification method based on label information guidance
CN117172253A
Method for semantic retrieval, device and storage medium
US20220027569A1