A Specific Entity Recognition Method for Scenarios with Scarce or Imbalanced Label Distribution

By introducing the adaptive resampling strategy of pseudo-label distribution-aware and the deconfused marginal loss function in the self-training framework, the rare entity recognition problem of the entity recognition model in the scenarios of scarcity and distribution imbalance is solved, and the generalization performance and recognition accuracy of the model are significantly improved.

CN115345165BActive Publication Date: 2025-05-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210990180.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-18
Publication Date
2025-05-27
Estimated Expiration
2042-08-18

AI Technical Summary

Technical Problem

When existing entity recognition models face scenarios with scarcity of labels and unbalanced distribution, it is difficult to effectively identify rare entities, and existing self-training methods cannot effectively solve this problem.

Method used

The self-training framework is adopted, combining the adaptive resampling strategy of pseudo-label distribution perception and the deconfused marginal loss function, and dynamically perceive the label distribution of new data during each round of self-training, optimize the training objectives, and improve the model's recognition accuracy of rare entities.

Benefits of technology

It significantly improves the generalization performance of entity recognition models in scenarios of scarce labels or unbalanced distribution, improves the accuracy, recall and F1 values ​​of rare categories, and provides a more robust in-domain solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115345165B_ABST
    Figure CN115345165B_ABST
Patent Text Reader

Abstract

The present invention discloses a specific entity recognition method for scenarios with scarce or unevenly distributed labels, proposes a pseudo-label distribution-aware adaptive resampling strategy and a deconfounding margin loss function, has a high tolerance for the label data distribution in the training set, solves the problem of uneven entity category distribution in the scenario of scarce in-domain labels, significantly improves the generalization performance of the entity recognition model in difficult scenarios with scarce or unevenly distributed labels, significantly improves evaluation metrics such as the precision, recall, and F1 value of rare categories, and is applicable to specific entity recognition tasks with fewer label samples or a higher degree of imbalance in the training set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a specific entity recognition method for scenarios with scarce or imbalanced label distribution. Background Art

[0002] Entity recognition aims to automatically label entities with specific meanings in text, mainly including person names, place names, organization names, proper nouns, etc. As a branch of sequence labeling tasks, entity recognition is an important basic tool in application fields such as information extraction, question answering systems, syntactic analysis, and machine translation, and occupies an important position in the process of the practical application of natural language processing technology. For example, in the construction of a domain-specific knowledge graph, named entity recognition is usually used to automatically capture domain-specific proper nouns and their attribute words to construct triples.

[0003] Scarcity of in-domain labels is a major challenge faced by entity recognition tasks. Since training samples require annotators to perform fine-grained token-level annotations, and entity annotations in professional fields often require domain experts to contribute knowledge, it is expensive and time-consuming to obtain fine-grained in-domain labeled data. On the contrary, in-domain unlabeled data is often massive and easy to obtain. Existing entity recognition models often use the self-training framework to solve the problem of scarce in-domain labels, and use massive in-domain unlabeled data to iteratively produce models. As a classic semi-supervised training framework, the self-training method is widely used in low-resource scenarios. Different from data augmentation-based methods such as consistency regularization, self-training does not require modifying the backbone network or preprocessing data. In the self-training framework, first, a small amount of labeled in-domain data is collected to form a labeled dataset, and the remaining massive in-domain unlabeled data forms an unlabeled dataset. The model is first trained on the labeled dataset and will be marked as the teacher model after convergence. Then, the teacher model is used to make predictions on the unlabeled dataset, and the prediction results with high confidence are set as the pseudo-labels of the data points. The pseudo-labeled data will be incorporated into the labeled set for iterative training of a new model, where the newly trained model is marked as the student model. The above process is iterated until the prediction metrics of the teacher model converge.

[0004] In addition to the scarcity of in-domain labels, the imbalance of entity label distribution is also a technical difficulty in entity recognition tasks. Since a single input text will contain multiple entities at the same time, the natural co-occurrence of entities in the training corpus leads to a generally unbalanced distribution of different entity types. Usually, named entities can be divided into common entities and rare entities according to their appearance frequencies in the in-domain dataset. Some rare entities are very important in actual application scenarios, such as organization names, contact information, etc. However, training an entity recognition model in a corpus with imbalanced entity distribution will cause the decision boundary of entity types to shift towards rare entities, resulting in misjudgment of rare entities.

[0005] The performance of semi-supervised methods represented by self-training is greatly affected in the setting of distribution imbalance. Since new labels are continuously added during the iterative process of the teacher-student model, the imbalance of the label class distribution often intensifies during this iterative process.

[0006] In recent years, this problem has received extensive attention in the field of images. "CReST: A Class-Rebalancing Self-Training Framework for Imbalanced Semi-Supervised Learning" (Chen Wei, Kihyuk Sohn, Clayton Mellina, Alan Yuille, and Fan Yang. 2021a. Crest: A class-rebalancing self-training framework for imbalanced semi-supervised learning. 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR).) proposed using weighted sampling to encourage the model to add more pseudo-labels from the minority classes during the iterative process, and achieved good performance on the long-tailed image classification dataset. "Rethinking Re-Sampling in Imbalanced Semi-Supervised Learning" (Ju He, Adam Kortylewski, Shaokang Yang, Shuai Liu, Cheng Yang, Changhu Wang, and Alan Yuille. 2021. Rethinking re-sampling in imbal-anced semi-supervised learning. arXiv preprint arXiv:2106.00209.) proposed decoupling the re-sampling and representation learning processes, and adopting different re-sampling methods at different stages of self-training to solve the class imbalance problem.

[0007] However, for the dual challenges of "scarce labels + imbalanced distribution", there are two defects in existing self-training methods: First, the imbalanced distribution of the classification dataset is at the instance level, while the entity category imbalance in the sequence labeling dataset is a token-level imbalanced distribution. An input corpus may contain both common entities and rare entities at the same time. Existing resampling methods are not applicable to this complex token-level distribution, and directly applying them to the sequence labeling task cannot well balance the dataset. Second, existing methods mostly make improvements from the data level and do not discuss solving the category imbalance problem from the perspective of the training objective of the student model. Summary of the Invention

[0008] In view of the above two defects, the present invention designs a specific entity recognition method for scenarios with scarce labels or imbalanced distribution with self-training as the framework, from the perspectives of resampling scheme design and training objective optimization, aiming to learn a more reasonable decision boundary and improve the overall recognition accuracy, especially the recognition accuracy of rare entities. Thus, it provides a more robust in-domain solution for entity recognition applications for specific tasks.

[0009] To achieve the above object, the present invention provides the following technical solutions:

[0010] The present invention provides a specific entity recognition method for scenarios with scarce labels or imbalanced distribution, including the following steps:

[0011] S1. Using the labeled data, train the model with the goal of minimizing the deconfounding margin loss function;

[0012] S2. Use the trained model to predict the unlabeled data, and assign a pseudo-label to each sample according to the class confidence predicted by the model;

[0013] S3. According to the pseudo-label distribution-aware adaptive resampling strategy, based on the class confidence obtained in step S2, assign weights according to the label distribution of the pseudo-labeled data newly added in the previous round of self-training, calculate the weighted confidence score for each pseudo-labeled sample, and then use the smoothing threshold function and Bernoulli sampling to finally determine whether the sample is selected;

[0014] S4. Use the pseudo-label of the sampled pseudo-labeled sample as its true label, delete this part of the data from the unlabeled dataset, and merge it with the training set in the original labeled data as the training set for the next iteration;

[0015] S5. Repeat steps S1 to S4 multiple times until the model converges;

[0016] S6. Input the text to be recognized into the trained model for prediction.

[0017] Furthermore, the model uses the Bert+BiLSTM+CRF model as the backbone network; when the text sequence is input into the network, the Bert pre-trained model is first used to pre-encode the text to obtain the word vectors of each character; then the BiLSTM network is used to further perform downstream encoding on the vectors to model the context information; finally, the CRF is used as the decoder to decode the encoding results, thereby obtaining the entity label sequence.

[0018] Furthermore, the deconfounded margin loss function in step S1 consists of the conditional random field loss the label distribution-aware margin loss the class-suppressed confusion loss and is composed of three parts, and the formula is as follows:

[0019]

[0020] Among them, λ 1 and λ 2 are hyperparameters representing the weights of different losses;

[0021] The conditional random field loss function is the original loss function of the Bert+BiLSTM+CRF model;

[0022] The label distribution-aware margin loss function is as follows:

[0023]

[0024] Among them N j represents the number of entities of the j-th class, H is a hyperparameter, and z j represents the output score of the model for classifying the word s as an entity of the j-th class;

[0025] The class-suppressed confusion loss function is as follows:

[0026]

[0027] Among them, ξ is a score threshold parameter, and σ(·) represents the Sigmoid function.

[0028] Furthermore, the specific selection method in step S3 is as follows:

[0029] S301. According to the distribution of the number of entity labels in the newly added pseudo-labeled data, sort the entities in descending order of the number, N 1 ≥N 2 ≥…≥N l ≥…≥N L , assign the weight μ to the entity ss , calculate the pseudo-labeled text S i weighted confidence score C i :

[0030]

[0031]

[0032] l is the index of the entity, and δ, γ, ρ are hyperparameters;

[0033] S302. Design a smoothing threshold function to calculate the text S i The probability of being selected is:

[0034]

[0035] C min is a score threshold, and α, β are hyperparameters, where α > 0 and β ≥ 1;

[0036] S303. Perform Bernoulli sampling on the candidates. The sampling probability p of the Bernoulli distribution is weighted by the entity weights, and its formula is:

[0037]

[0038] Furthermore, in step S5, the Viterbi algorithm is used for decoding in the CRF layer, and the entity label sequence with the highest score is selected as the recognition result, and the post-processing outputs the structured recognition result.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] 1. Based on the self-training method of semi-supervised learning, the present invention reduces the cost of manual annotation, makes full use of a large amount of unlabeled data to expand the scarce labeled data set, reduces the requirement for the scale of labeled data, and alleviates the problem of low recognition accuracy of deep learning models on data with fewer labeled samples and higher imbalance.

[0041] 2. The present invention designs a new pseudo-label distribution-aware adaptive resampling strategy for specific entity recognition tasks. It can dynamically perceive the label distribution of the newly added data in each round of self-training, and adaptively sample pseudo-labeled data to be added to the training set of the next iteration (when the confidence is high enough, the more entity categories with fewer samplings in the previous round are included in the text, the greater the probability of being selected), which helps to balance the entity number distribution of the training set and improve the recognition performance of the model on rare classes.

[0042] 3. The present invention proposes a deconfounding label distribution-aware margin loss function, which acts on each round of training and learning process of the model, aiming to correct the deviation of the classification decision surface between entities and eliminate potential confusion caused by semantic similarity or quantity distribution differences between entities, making the recognition results more confident and improving the precision of the model in all categories.

[0043] In summary, the specific entity recognition method for label-scarce or distribution-imbalanced scenarios proposed by the present invention proposes a pseudo-label distribution-aware adaptive resampling strategy and a deconfounding margin loss function, which has a high tolerance for the label data distribution in the training set, solves the problem of imbalanced entity category distribution in the in-domain label-scarce scenario, significantly improves the generalization performance of the entity recognition model in difficult scenarios of label scarcity or distribution imbalance, significantly improves evaluation metrics such as precision, recall, and F1 value of rare categories, is applicable to specific entity recognition tasks with fewer label samples or higher imbalance degree in the training set, and helps to alleviate the problem of low recognition accuracy of rare entity categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0045] Figure 1 It is a flowchart of a specific entity recognition method for label-scarce or distribution-imbalanced scenarios provided by an embodiment of the present invention.

[0046] Figure 2 It is a text data annotation mode provided by an embodiment of the present invention.

[0047] Figure 3 It is a model training flowchart provided by an embodiment of the present invention.

[0048] Figure 4 It is a model inference flowchart provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] In order to better understand the technical solution, the method of the present invention will be described in detail below with reference to the drawings.

[0050] The specific entity recognition method for label-scarce or distribution-imbalanced scenarios of the present invention, as Figure 1 shown, includes the following steps:

[0051] Step 1: Pre-define specific entity types in the target domain.

[0052] Domain experts define specific meaningful concepts to be recognized from the text as entity types, i.e., entity labels for model learning. For example, the CMID (Chinese Medical Intention) dataset in the medical field defines entity types such as "Diseases and Diagnoses", "Imaging Examinations", "Anatomical Sites", "Drugs", "Surgeries", etc.

[0053] Step 2: Prepare the dataset.

[0054] Collect text data within the target domain, remove illegal characters, and label some of the data. Divide the dataset into a labeled dataset and an unlabeled dataset based on whether the sentences are labeled, and further divide the labeled dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1. Each data subset needs to cover all entity types. The text data annotation mode adopts the BIO method, where B (Begin) represents the start position of the entity, I (Inner) represents the internal position of the entity, and O is used to mark irrelevant characters. An example is as Figure 2 shown.

[0055] Step 3: Build a named entity recognition model.

[0056] In the embodiment of the present invention, the commonly used Bert+BiLSTM+CRF model for named entity recognition is used as the backbone network. When the text sequence is input into the network, first use the Bert pre-trained model to pre-encode the text to obtain the word vectors of each character; then use the BiLSTM network to further perform downstream encoding on the vectors to model the context information; finally, CRF is used as the decoder to decode the encoded result, thereby obtaining the entity label sequence.

[0057] Step 4: Train the model.

[0058] For the difficulty of scarce labels, few-shot learning methods represented by meta-learning can be adopted. The core idea of meta-learning is to let the model train on multiple decomposed tasks with a large amount of labeled data, learn the classifier training experience on the relevant data, and obtain better initial parameters, so as to have the ability to generalize to a new task with only a small amount of labeled data. However, this method still requires a large amount of labeled data in the relevant field, and the training is complex and time-consuming.

[0059] Since the labeled data in the dataset is scarce and it is relatively easy to obtain unlabeled data in the domain, the present invention adopts the self-training method in semi-supervised learning to use a large amount of unlabeled data to expand the labeled dataset, so that the model learns a better feature extractor and enhances the generalization ability of the model.

[0060] The overall process is as Figure 3 shown, including the following steps:

[0061] S1. Use the labeled data to train the model with the goal of minimizing the deconfounding marginal loss function.

[0062] In step S1, training the model on a dataset with an imbalanced label distribution will bring about biases in the decision boundary and label confusion, resulting in rare entities being easily confused by the model as common entities or other entities with similar semantics. Commonly used loss functions for imbalanced distributions include Focal Loss, Dice Loss, etc. However, these loss functions essentially focus more on the imbalance problem between easy and difficult samples and neglect the penalty for quantity relationships and confusion. The present invention proposes a training objective of minimizing the deconfounding marginal loss function, aiming to correct the boundary size between entities and mitigate potential confusion phenomena during model training. The specific approach is as follows:

[0063] Deconfounding marginal loss function Consists of a conditional random field loss Label distribution-aware marginal loss Class-suppressed confusion loss And is composed of three parts. The total loss function formula is as follows:

[0064]

[0065] Where λ 1 And λ 2 Are hyperparameters representing the weights of different losses;

[0066] The conditional random field loss function Is the original loss function of the Bert+BiLSTM+CRF model;

[0067] The label distribution-aware marginal loss function Is as follows:

[0068]

[0069] Where N j Represents the number of entities of the j-th class, H is a hyperparameter, and z j Represents the output score of the model for classifying the word s as an entity of the j-th class; it can be seen from the formula that this loss function is related to the number of entity classes. The fewer the number, the more the model is forced to have a high output score for this entity, thereby encouraging the model to widen the boundary distance of rare entities to correct the bias problem of the decision boundary caused by the imbalanced distribution.

[0070] During discrimination, the model is prone to entity confusion, among which the confusion between rare entities and common entities or other entities with similar semantics is very serious. The model lacks discrimination between rare classes and common classes with similar semantics. To address this issue, the present invention adopts a class-suppressed confusion loss function as follows:

[0071]

[0072] where ξ is a fractional threshold parameter, and the losses on non-genuine categories with scores greater than ξ are included in the loss function to suppress other categories that are easily confused with them, promoting the protection of the accuracy of all entities, especially rare entities, when the parameters are updated, and σ(·) represents the Sigmoid function.

[0073] S2. Use the trained model to predict the unlabeled data, and assign a pseudo-label to each sample according to the confidence level;

[0074] S3. According to the pseudo-label distribution-aware adaptive resampling strategy, based on the confidence levels obtained in step S2, assign weights according to the label distribution of the pseudo-labeled data newly added in the previous round of self-training, and calculate the weighted confidence score for each pseudo-labeled sample. Then, use the smoothing threshold function and Bernoulli sampling to finally determine whether the sample is selected;

[0075] The most crucial step in self-training lies in step S3 - the selection of pseudo-labeled data. Generally speaking, pseudo-labels are highly noisy. The traditional approach is to sum and average the confidence levels of all entities, and then use a fractional threshold to filter out some low-confidence predictions. However, due to the unbalanced distribution of entities in a sentence, simply calculating the confidence score by summing and averaging will ignore the quality of the pseudo-labels of rare entities. In addition, since common entity types are likely to obtain high scores, the traditional approach tends to select sentences containing more common entities, further exacerbating the unbalanced distribution in the training set.

[0076] Therefore, the present invention proposes a pseudo-label distribution-aware adaptive resampling strategy to sample a high-confidence subset from the pseudo-labeled data. The specific approach is as follows:

[0077] S301. According to the distribution of the number of entity labels in the newly added pseudo-labeled data (the number distribution in the original labeled dataset is used in the first round), sort the entities in descending order of the number, N 1 ≥N 2 ≥…≥N l ≥…≥N L , assign the weight μ to the entity s s , and calculate the weighted confidence score C of the pseudo-labeled text S i : i :

[0078]

[0079]

[0080] l is the index of the entity, and δ, γ, ρ are hyperparameters; from μs It can be seen from the formula that the weight and the quantity are negatively correlated. The more entities there are, the smaller the weight, and the greater the contribution to the confidence score in the text.

[0081] S302. Design a smooth threshold function to calculate the text S i The probability of being selected is:

[0082]

[0083] C min is a score threshold, α and β are hyperparameters, α > 0, β ≥ 1; different from the classical step function, the smooth threshold function adopts a smooth transformation for the score.

[0084] S303. Conduct Bernoulli sampling on the candidates. The sampling probability p of the Bernoulli distribution is weighted by the entity weight, and its formula is:

[0085]

[0086] Select the pseudo-labeled statements that can finally be added to the next round of training set.

[0087] As the self-training progresses, under the adaptive sampling strategy of pseudo-label distribution perception, those statements with more rare entities and relatively high scores on these rare entities are more likely to be selected into the training corpus, which helps to alleviate the highly unbalanced distribution among entities in the training set.

[0088] An example of a set of hyperparameter values is as follows:

[0089]

[0090] γ = 2, ρ = 1, α = 10, β = 1, C min = 0.95

[0091] S4. Use the pseudo-labels of the sampled pseudo-labeled data as their true labels, delete this part of the data from the unlabeled dataset, and merge it with the training set in the original labeled data as the training set for the next iteration;

[0092] Repeat S1 - S4 multiple times until the model converges.

[0093] Step Five: Model inference.

[0094] Input the text to be recognized into the trained model for prediction. Use the Viterbi algorithm for decoding in the CRF layer, select the entity label sequence with the highest score as the recognition result, and post-process to output the structured recognition result. As Figure 4 shown.

[0095] On the dataset implemented in the present invention, the F1 value of rare entity categories can be increased by 6% - 9%. For example, the F1 value is increased by 8.7% for rare categories in the 10-shot SNIPS dataset; the F1 value is increased by 6.4% for rare categories in the 10-shot Few-NERD dataset.

[0096] In summary, the specific entity recognition method for label-scarce or distribution-imbalanced scenarios proposed by the present invention proposes a pseudo-label distribution-aware adaptive resampling strategy and a deconfounding margin loss function, has a high tolerance for the distribution of label data in the training set, solves the problem of entity category distribution imbalance in the in-domain label-scarce scenario, significantly improves the generalization performance of the entity recognition model in difficult scenarios of label scarcity or distribution imbalance, significantly improves evaluation indicators such as the precision, recall, and F1 value of rare categories, is applicable to specific entity recognition tasks with fewer label samples or higher imbalance degrees in the training set, and helps to alleviate the problem of low recognition accuracy of rare entity categories.

[0097] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features, but these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A specific entity recognition method for scenarios with scarce or imbalanced label distribution, characterized in that, it includes the following steps: S1. Use the labeled data to train a model with the goal of minimizing the deconfounding margin loss function; S2. Use the trained model to predict the unlabeled data, and assign a pseudo-label to each sample according to the class confidence predicted by the model; S3. According to the pseudo-label distribution-aware adaptive resampling strategy, based on the class confidence obtained in step S2, assign weights according to the label distribution of the pseudo-labeled data newly added in the previous round of self-training, calculate the weighted confidence score for each pseudo-labeled sample, and finally determine whether the sample is selected by using the smoothing threshold function and Bernoulli sampling; S4. Use the pseudo-labels of the sampled pseudo-labeled samples as their true labels, delete this part of the data from the unlabeled dataset, and merge it with the training set in the original labeled data as the training set for the next iteration; S5. Repeat steps S1 to S4 multiple times until the model converges; S6. Input the text to be recognized into the trained model for prediction.

2. The specific entity recognition method for scenarios with scarce or imbalanced label distribution according to claim 1, characterized in that, the model uses the Bert+BiLSTM+CRF model as the backbone network; when the text sequence is input into the network, first use the Bert pre-trained model to pre-encode the text to obtain the word vectors of each character; then use the BiLSTM network to further perform downstream encoding on the vectors to model the context information; finally, CRF is used as the decoder to decode the encoding result to obtain the entity label sequence.

3. The specific entity recognition method for scenarios with scarce or imbalanced label distribution according to claim 1, characterized in that, The de - obfuscation marginal loss function in step S1 consists of the conditional random field loss the label - distribution - aware marginal loss and the class - suppression obfuscation loss and is composed of three parts, with the formula as follows: Among them, λ 1 and λ 2 are hyperparameters representing the weights of different losses; Conditional Random Field Loss Function is the original loss function of the Bert+BiLSTM+CRF model; Label Distribution Aware Margin Loss Function As follows: Among them N j represents the number of entities of the j-th class, H is a hyperparameter, and z j represents the output score of the model for classifying the word s as an entity of the j-th class; Class-inhibited confusion loss function As follows: where ξ is a fractional threshold parameter, and σ(·) represents the Sigmoid function.

4. The specific entity recognition method for scenarios with scarce or imbalanced label distribution according to claim 1, characterized in that, the specific selection method in step S3 is: S301. Sort the entities in descending order of the number according to the entity label number distribution in the newly added pseudo-labeled data, N 1 ≥N 2 ≥…≥N l ≥…≥N L , assign the weight μ to the entity s s , calculate the pseudo-labeled text S i weighted confidence score C i : l is the index of the entity, and δ, γ, ρ are hyperparameters; S302. Design a smooth threshold function to calculate text S i The probability of being selected is: C min is a score threshold, α and β are hyperparameters, where α > 0 and β ≥ 1; S303. Perform Bernoulli sampling on the candidates, and the sampling probability p of the Bernoulli distribution is weighted by the entity weight, and its formula is:

5. The specific entity recognition method for scenarios with scarce or imbalanced label distribution according to claim 1, characterized in that, in step S5, the Viterbi algorithm is used for decoding in the CRF layer, and the entity label sequence with the highest score is selected as the recognition result, and the post-processing outputs the structured recognition result.