Multi-domain entity identification method and device under data scarcity condition and readable medium
By generating pseudo-labels and fine-tuning of pre-trained language models, the problem of insufficient recognition accuracy and generalization capabilities of named entity recognition technology in the case of scarcity of data and inconsistent entity categories is solved, and efficient multi-domain entity recognition is achieved.
Patent Information
- Application Number
- CN202510387793.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-08-12
AI Technical Summary
In the face of scarcity of data and inconsistent entity categories, existing named entity recognition technology is difficult to effectively identify new entity categories in multiple fields, resulting in insufficient recognition accuracy and generalization capabilities.
By generating pseudo-labels and scoring and filtering methods, combining pseudo-label generation and fine-tuning of pre-trained language models, a multi-domain entity recognition model is built, which supplements the missing entity categories and optimizes the inference efficiency of the identification model.
It improves the accuracy and stability of named entity recognition in multi-domain tasks, simplifies the inference process, reduces the cost of manual labeling, and improves the generalization ability and recognition effect of the model.
Smart Images

Figure CN120471056A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of named entity recognition, and in particular to a method, device and readable medium for multi-domain entity recognition under data scarcity conditions. Background Art
[0002] Named Entity Recognition (NER) is a core task in natural language processing. Its goal is to identify and classify entities with specific semantics from text. However, current practical applications often face the contradiction between the diverse entity requirements and insufficient dataset coverage, which has become a major bottleneck in building high-quality entity recognition systems.
[0003] (1) Diversity of entity needs
[0004] Entity detection tasks often require flexible definition of target entity categories based on specific scenarios and requirements. For example, in the medical field, certain tasks may require identifying entities such as diseases, medical test items, and drug dosages. In the financial field, entities such as company names, financial transaction amounts, and transaction times may be of interest. In general domains, entities such as age, height, and weight may need to be identified.
[0005] However, the diversity of practical needs leads to significant discrepancies between entity categories. A dataset may cover some entity categories (such as diseases and medical test items), but when the task is expanded, new entity categories (such as age, height, and weight) are often required. In this case, the disconnect between the dataset and practical needs limits the improvement of task performance.
[0006] (2) Limitations of dataset coverage
[0007] Existing NER datasets usually have the following problems:
[0008] 1. Limited label range
[0009] Many datasets define only a small number of entity categories for general or specific domains. For example, a general NER dataset may only contain common entity categories such as "name," "address," and "organization," but in the medical field, it lacks annotations for specialized entities such as "symptoms" and "cause of disease."
[0010] 2. Insufficient samples
[0011] Even if the target entity class exists in the dataset, its sample number is often insufficient. For example, a medical dataset may contain the “drug” entity class, but lack sufficient samples on drug subtypes (such as dosage, frequency) to meet the needs of fine-grained tasks.
[0012] 3. New demands are difficult to adapt
[0013] In real-world applications, task requirements often change dynamically, and users may need to identify newly added entity categories. For example, in a medical consultation system, initially only "disease" and "clinical manifestation" recognition is required, but later the ability to detect "drugs" may need to be added. In this case, existing datasets often do not cover the full range.
[0014] (3) The contradiction between dataset and model capabilities
[0015] During model training, the definition of entity categories is usually determined by the dataset. When the model is trained on a dataset with limited coverage:
[0016] ① Loss of unlabeled entities
[0017] Target entity categories not covered in the dataset are considered noise in the model prediction. Even if the model can correctly detect the entity, it cannot generate effective labels for it. If the dataset contains entities and entity category sets that lack some target entity categories required by the task, the labeling workload for these entities will increase significantly.
[0018] ② Limited model generalization ability
[0019] When new entity categories are introduced, the model has difficulty adapting to new task requirements in the absence of relevant training data, resulting in a significant drop in performance.
[0020] (4) Impact of actual tasks
[0021] Due to insufficient dataset coverage, the model's performance may be severely affected in the following scenarios:
[0022] ① Multi-tasking applications in the medical field
[0023] In medical NER tasks, the lack of annotation of fine-grained entities such as "drug dosage" and "image description" leads to a decrease in the comprehensiveness and accuracy of electronic medical record information extraction.
[0024] ②Dynamic demand scenarios
[0025] In real-world applications, task requirements may change over time. For example, during an epidemic, the need for identifying infectious disease-related terms may increase. In this case, datasets and models that lack flexible scalability cannot adapt quickly to changes.
[0026] In summary, the diversity of entity recognition tasks requires models with flexible generalization and scalability. However, the shortcomings of existing NER datasets in terms of coverage, sample distribution, and dynamic adaptability severely restrict the improvement of task performance. This indicates that a new solution is needed to effectively bridge the gap between datasets and task requirements.
[0027] Existing named entity recognition (NER) technologies mainly rely on machine learning and deep learning methods, and use large-scale annotated datasets to improve the accuracy of entity recognition. Existing technologies mainly include the following categories:
[0028] (1) Traditional machine learning methods
[0029] Common algorithms used in traditional NER include conditional random fields (CRFs) and maximum entropy models (MaxEnt). These methods extract useful information from text through feature engineering and train it using machine learning algorithms to identify entities. CRFs effectively capture contextual information in sequential data, while maximum entropy models leverage global information for discriminative modeling.
[0030] (2) Deep Learning Methods
[0031] In recent years, deep learning-based methods have made significant progress and become the mainstream technology in NER tasks. Deep learning methods usually use neural networks to automatically extract features, avoiding the tedious process of manually designing features in traditional methods. They mainly include the following methods:
[0032] Recurrent Neural Networks (RNNs): RNNs can process sequential data and capture long-range contextual dependencies. RNN-based variants, such as LSTM (Long Short-Term Memory) and GRU (Gated Recurrent Unit), further improve the ability to model long sequences and are widely used in NER tasks.
[0033] BERT and its variants: BERT (Bidirectional Encoder Representations from Transformers) uses a bidirectional encoder to model contextual information and improves model performance on various tasks through large-scale pre-training. BERT-based models can achieve excellent entity recognition results in multiple domains, especially for general NER tasks, outperforming traditional methods.
[0034] (3) Data enhancement and transfer learning
[0035] Data augmentation techniques enhance the diversity of training sets by synthesizing new training samples (such as synonym replacement and sentence rearrangement), thereby improving the model's generalization ability. Transfer learning uses pre-trained language models, such as BERT, to migrate models to specific domains, avoiding the training issues of small datasets and thus improving performance on specific tasks.
[0036] Although existing NER technologies have solved the problem of entity recognition in text to a certain extent, they still face some challenges and limitations in practical applications. The following are the main problems of existing technologies:
[0037] (1) Insufficient dataset coverage
[0038] Most existing NER models rely on fixed annotated datasets that typically cover common entity categories (such as names of people, places, and organizations). However, in some specific fields or tasks, the required entity categories may not be adequately annotated. Especially in professional fields (such as medicine and law), the entity categories of existing datasets cannot meet the needs of practical applications. For example, NER tasks in the medical field require the identification of fine-grained entities such as drug dosages and imaging sites, which are usually not adequately annotated in existing datasets.
[0039] When faced with new entities or domain-specific entities, existing models cannot effectively identify and distinguish them, which affects recognition accuracy.
[0040] (2) The definition of entity categories is inconsistent with task requirements
[0041] Existing NER datasets typically predefine entity categories, which are fixed during training. However, real-world application requirements can be more diverse. For example, in the medical field, as new diseases or drugs emerge, existing datasets may not be able to adapt to these changes in a timely manner, resulting in the model being unable to recognize new entity categories.
[0042] Due to the inflexible definition of entity categories, the model cannot meet actual needs when processing certain specific tasks, resulting in unsatisfactory recognition results.
[0043] (3) Lack of labeled data
[0044] Although large-scale labeled datasets have a significant effect on NER tasks, labeled data is still severely scarce in some fields or tasks. For example, in emerging fields or specific tasks, it is very difficult to obtain sufficient labeled data, which directly affects the training effect and accuracy of the model.
[0045] In the absence of data, it is difficult to obtain sufficient diversity in model training, which affects the model's generalization ability and performance in practical applications.
[0046] Existing NER technologies have achieved good results in many fields, but they still have significant limitations when faced with problems such as insufficient dataset coverage, inconsistent entity categories, and poor domain adaptability. In order to improve the accuracy and applicability of entity recognition, it is necessary to address these issues and develop more flexible, scalable, and domain-adaptive NER methods. Summary of the Invention
[0047] The purpose of this application is to propose a multi-domain entity recognition method, device and readable medium under data scarcity conditions to address the above-mentioned technical problems.
[0048] In a first aspect, the present invention provides a method for multi-domain entity recognition under data scarcity conditions, comprising the following steps:
[0049] Determine a target entity category set required for a target named entity recognition task, where the target entity category set includes several target entity categories;
[0050] Obtain a set of original data sets, where each original data set in the set of original data sets includes a text data set and a set of entities and entity categories. Under data scarcity conditions, the entity and entity category set of each original data set stores entities that annotate the text of the text data set and their corresponding entity categories that do not completely cover the target entity category set; based on each original data set and the target entity category set, annotate pseudo labels for the remaining unannotated entities in the text of the text data set of each original data set;
[0051] Using the pre-trained first language model and the target entity category set, the pseudo labels of the remaining unlabeled entities in the text of each original dataset are scored and filtered, and the pseudo labels with high confidence and their corresponding texts are retained and combined with the corresponding original datasets to generate the final dataset, forming the final dataset set;
[0052] The final dataset set is used to fine-tune the pre-trained second largest language model to obtain the entity recognition model corresponding to the target named entity recognition task; the text to be recognized is obtained and input into the entity recognition model corresponding to the target named entity recognition task to identify the entity and its corresponding entity category.
[0053] Preferably, based on each original data set and the target entity category set, pseudo labels of the remaining unlabeled entities in the text of the text data set of each original data set are labeled, specifically including:
[0054] S21, traversing one of the original data sets in the original data set set and taking it as the data set to be labeled, and forming a missing entity category set of the data set to be labeled with some target entity categories in the target entity category set that do not exist in the entity and entity category set of the data set to be labeled;
[0055] S22, traversing one of the remaining original data sets in the original data set set except the data set to be labeled and using it as a supplementary data set, taking the intersection of the missing entity category set of the data set to be labeled and the entity and entity category set of the supplementary data set, obtaining the entity categories that the supplementary data set can supplement for the data set to be labeled, and in response to determining that the number of entity categories that the supplementary data set can supplement for the data set to be labeled is greater than or equal to 1, performing entity recognition on the text in the text data set of the data set to be labeled using the text data set of the supplementary data set and the entity categories that the supplementary data set can supplement for the data set to be labeled, to obtain a first pseudo label predicted for the data set to be labeled using the supplementary data set;
[0056] S23, repeating steps S22-S23 until all the original data sets except the data set to be labeled are traversed, and a first pseudo label predicted for the data set to be labeled is obtained using each supplementary data set;
[0057] S24, taking some target entity categories in the target entity category set that do not exist in all entities and entity category sets of the dataset to be labeled as the entity categories to be predicted, and performing entity recognition based on the entity categories to be predicted on the text in the text data set of the dataset to be labeled, to obtain a second pseudo label for the dataset to be labeled;
[0058] S25, taking the union of the first pseudo labels predicted by all supplementary datasets for the dataset to be labeled and the second pseudo labels of the dataset to be labeled, to obtain the remaining unlabeled entities in the text of the text dataset of the dataset to be labeled and their corresponding pseudo labels;
[0059] S26, repeating steps S21-S26 until all original data sets in the original data set set are traversed, and obtaining the remaining unlabeled entities in the text of the text data set of each original data set and their corresponding pseudo labels.
[0060] Preferably, entity recognition is performed on text in the text data set of the to-be-annotated dataset using the text data set of the supplementary dataset and the entity categories that the supplementary dataset can supplement for the to-be-annotated dataset, specifically including:
[0061] The pre-trained first entity recognition model is fine-tuned using the text data set of the supplementary data set and the entity categories that the supplementary data set can supplement for the data set to be labeled, thereby obtaining a fine-tuned first entity recognition model; and the fine-tuned first entity recognition model is used to perform entity recognition on the text in the text data set of the data set to be labeled.
[0062] Preferably, entity recognition based on the entity category to be predicted is performed on the text in the text data set of the labeled data set, specifically including:
[0063] A pre-trained second entity recognition model is used to perform entity recognition on text in the text data set of the to-be-annotated dataset, and the scope of the recognized entity categories is limited to cover the entity categories to be predicted.
[0064] Preferably, the pre-trained first language model and the target entity category set are used to score and filter the pseudo labels of the remaining unlabeled entities in the text of the text data set of each original data set, specifically including:
[0065] Setting a prompt word, which is used to guide the pre-trained first language model to give a confidence score of the pseudo-label of each unlabeled entity and an overall score of the pseudo-labels of all unlabeled entities in each text based on each text of the text data set of each original data set input and the unlabeled entities and their pseudo-labels. In the scoring process, it is necessary to consider whether the entity recognition is correct, whether the pseudo-label is within the range of the target entity category set, and whether the pseudo-label is semantically appropriate in the text;
[0066] In response to determining that the confidence score of the pseudo-label of each unlabeled entity in the text is greater than or equal to a first threshold, and the overall score of the pseudo-labels of all unlabeled entities in the text is greater than or equal to a second threshold, determining that the pseudo-labels of all unlabeled entities in the text are high-confidence pseudo-labels, and retaining the corresponding text;
[0067] In response to determining that the confidence score of the pseudo-label of at least one unlabeled entity in the text is less than a first threshold, or the overall score of the pseudo-labels of all unlabeled entities in the text is less than a second threshold, the pseudo-labels of all unlabeled entities in the text are determined to be low-confidence pseudo-labels, and the corresponding text is discarded.
[0068] Preferably, the pre-trained first language model includes the Qwen2.5-72B-Instruct model, the pre-trained second language model includes the Qwen2.5-0.5B-Instruct model, and during the fine-tuning process of the pre-trained second language model, the entity category of each text is limited to be within the range of the target entity category set through the Schema instruction.
[0069] In a second aspect, the present invention provides a method and apparatus for multi-domain entity recognition under data scarcity conditions, comprising:
[0070] a target entity category determination module, configured to determine a target entity category set required for a target named entity recognition task, wherein the target entity category set includes a plurality of target entity categories;
[0071] The pseudo-label prediction module is configured to obtain a set of original data sets, each of which contains a text data set and a set of entities and entity categories. Under data scarcity conditions, the entity and entity category set of each original data set stores entities that annotate the text of the text data set and their corresponding entity categories that do not completely cover the target entity category set; based on each original data set and the target entity category set, pseudo-labels of the remaining unannotated entities in the text of the text data set of each original data set are annotated;
[0072] a pseudo-label filtering module configured to score and filter the pseudo-labels of the remaining unlabeled entities in the text of each original data set using the pre-trained first language model and the target entity category set, retain the pseudo-labels with high confidence and their corresponding texts, and combine them with the corresponding original data set to generate a final data set to constitute a final data set set;
[0073] The entity recognition module is configured to use the final data set to fine-tune the pre-trained second largest language model to obtain an entity recognition model corresponding to the target named entity recognition task; obtain the text to be recognized and input it into the entity recognition model corresponding to the target named entity recognition task to recognize the entity and its corresponding entity category.
[0074] In a third aspect, the present invention provides an electronic device comprising one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0075] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0076] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which implements the method described in any implementation manner in the first aspect when the computer program is executed by a processor.
[0077] Compared with the prior art, the present invention has the following beneficial effects:
[0078] (1) The multi-domain entity recognition method proposed in this paper uses a pre-trained entity recognition model to generate pseudo-labels and supplement the entity categories missing in the original dataset. Combined with the pseudo-label scoring and filtering process, it completes the construction of a high-quality multi-domain dataset. By using pseudo-label generation to supplement the missing target entity categories and screening high-confidence pseudo-labels through a scoring mechanism, it solves the problem of insufficient generalization ability of entity recognition models caused by data scarcity and incomplete labels in related technologies, and improves the accuracy and stability of named entity recognition in multi-domain tasks.
[0079] (2) The multi-domain entity recognition method proposed in this paper, under data-scarce conditions, further optimizes the inference efficiency of the entity recognition model corresponding to the target named entity recognition task by fine-tuning the pre-trained second language model, combining Schema instruction construction and dynamic annotation output. Because the optimized instructions clarify the entity categories of the input text, the inference process is simplified, solving the problems of slow inference speed and high resource consumption of entity recognition models in related technologies, and significantly improving the inference efficiency of entity recognition models.
[0080] (3) The proposed multi-domain entity recognition method under data-scarce conditions achieves efficient training of entity recognition models under data-scarce conditions through the synergistic effect of multi-domain pseudo-label expansion and fine-tuning techniques. By adopting a combined strategy of pseudo-label generation and scoring screening, it supplements entity annotation under data-scarce conditions, solving the problems of high manual annotation costs and poor cross-domain generalization capabilities in related technologies. This significantly improves the recognition effect and generalization capabilities of the model for the target named entity recognition task under data-scarce conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0082] Figure 1 A flowchart of a multi-domain entity recognition method under data scarcity conditions according to an embodiment of the present application;
[0083] Figure 2 A schematic diagram of a multi-domain entity recognition device under data scarcity conditions according to an embodiment of the present application;
[0084] Figure 3 A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0085] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only some, not all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.
[0086] Figure 1 The embodiment of the present application provides a method for multi-domain entity recognition under data scarcity conditions, comprising the following steps:
[0087] S1, determine the target entity category set required for the target named entity recognition task, and the target entity category set contains several target entity categories.
[0088] Specifically, in the named entity recognition (NER) task, we first need to clarify what the entity category is to be recognized. The scope of the target entity category is usually determined by the background and application scenario of the target named entity recognition task. Assume that the target entity category set is E target ={e1,e2,e3,...,e k}, where e k Represents the target entity category, such as disease name, drug name, etc. The goal is to identify the entity in the text T that belongs to E target For example, in the medical field, possible entity categories include disease names, drug names, symptoms, etc.; in the news field, they may include names of people, places, and organizations. The choice of target entity category will directly affect subsequent data annotation and model design.
[0089] S2, obtaining a set of original data sets, where each original data set in the original data set contains a text data set and a set of entities and entity categories. Under the condition of data scarcity, the entity and entity category set of each original data set stores the entities that annotate the text of the text data set and their corresponding entity categories that do not completely cover the target entity category set; based on each original data set and the target entity category set, annotate the pseudo labels of the remaining unannotated entities in the text of the text data set of each original data set.
[0090] Specifically, in the target named entity recognition task, the original dataset collected usually has labels for some target entity categories, while labels for other target entity categories are missing. To solve this problem of incomplete labels, the following strategies can be adopted:
[0091] Assume there are i original data sets D i ={T i ,E i}, where i∈{a,b,c,d...}, where each original dataset contains a text dataset T and an entity and entity category set E. It is necessary to consider the applicability and quality of the original dataset and check whether the entity categories in the dataset are consistent with the target entity categories E. target If they are not consistent, the dataset needs to be adjusted or supplemented.
[0092] The target named entity recognition task has multiple target entity categories. The original dataset set contains at least two original datasets. In the case of data scarcity, multiple original datasets respectively mark some target entity categories, but do not completely cover all target entity categories in the target entity category set. a and the original dataset D b For example, the original data set D a Some entity categories E are marked a ={e1,e2,e3,...,e k1}, original dataset D b Some entity categories E are marked b ={e1,e2,e3,...,e k2 In the case of data scarcity, in the required target named entity recognition task, the target entity category set is E target ={e1,e2,e3,...,e k}, E a ={e1,e2,e3,...,e k1} and E b ={e1,e2,e3,...,e k2} can only target some entities in the target named entity recognition task, and does not cover all target entity categories in the target entity category set required in the target named entity recognition task. Therefore, for the target named entity recognition task, the original dataset D a and the original dataset D b The entity and entity category set of the original dataset stores the entities and entity categories corresponding to the text in the text dataset. The entity categories cover some of the target entity categories in the target entity category set, but do not completely cover all the target entity categories in the target entity category set.
[0093] In a specific embodiment, based on each original data set and the target entity category set, pseudo labels of the remaining unlabeled entities in the text of the text data set of each original data set are labeled, specifically including:
[0094] S21, traversing one of the original data sets in the original data set set and taking it as the data set to be labeled, and forming a missing entity category set of the data set to be labeled with some target entity categories in the target entity category set that do not exist in the entity and entity category set of the data set to be labeled;
[0095] S22, traversing one of the remaining original data sets in the original data set set except the data set to be labeled and using it as a supplementary data set, taking the intersection of the missing entity category set of the data set to be labeled and the entity and entity category set of the supplementary data set, obtaining the entity categories that the supplementary data set can supplement for the data set to be labeled, and in response to determining that the number of entity categories that the supplementary data set can supplement for the data set to be labeled is greater than or equal to 1, performing entity recognition on the text in the text data set of the data set to be labeled using the text data set of the supplementary data set and the entity categories that the supplementary data set can supplement for the data set to be labeled, to obtain a first pseudo label predicted for the data set to be labeled using the supplementary data set;
[0096] S23, repeating steps S22-S23 until all the original data sets except the data set to be labeled are traversed, and a first pseudo label predicted for the data set to be labeled is obtained using each supplementary data set;
[0097] S24, taking some target entity categories in the target entity category set that do not exist in all entities and entity category sets of the dataset to be labeled as the entity categories to be predicted, and performing entity recognition based on the entity categories to be predicted on the text in the text data set of the dataset to be labeled, to obtain a second pseudo label for the dataset to be labeled;
[0098] S25, taking the union of the first pseudo labels predicted by all supplementary datasets for the dataset to be labeled and the second pseudo labels of the dataset to be labeled, to obtain the remaining unlabeled entities in the text of the text dataset of the dataset to be labeled and their corresponding pseudo labels;
[0099] S26, repeating steps S21-S26 until all original data sets in the original data set set are traversed, and obtaining the remaining unlabeled entities in the text of the text data set of each original data set and their corresponding pseudo labels.
[0100] In a specific embodiment, entity recognition is performed on text in the text dataset to be annotated using the text dataset of the supplementary dataset and the entity categories that the supplementary dataset can supplement for the dataset to be annotated, specifically including:
[0101] The pre-trained first entity recognition model is fine-tuned using the text data set of the supplementary data set and the entity categories that the supplementary data set can supplement for the data set to be labeled, thereby obtaining a fine-tuned first entity recognition model; and the fine-tuned first entity recognition model is used to perform entity recognition on the text in the text data set of the data set to be labeled.
[0102] In a specific embodiment, entity recognition based on the entity category to be predicted is performed on the text in the text data set of the to-be-annotated dataset, specifically including:
[0103] A pre-trained second entity recognition model is used to perform entity recognition on text in the text data set of the to-be-annotated dataset, and the scope of the recognized entity categories is limited to cover the entity categories to be predicted.
[0104] Specifically, take the original data set D a As the dataset to be labeled, the original dataset D b As a supplementary dataset, let’s explain it. Assume that the original dataset D a (i.e. the dataset to be labeled) lacks some target entity categories in the target entity category set, that is, some target entity categories in the target entity category set do not exist in the entity and entity category set of the dataset to be labeled, then this part of the target entity category constitutes the missing entity category set E of the dataset to be labeled lack . The original dataset D b (i.e., supplementary dataset) contains the original dataset D a The missing target entity categories in the dataset are intersected with the missing entity category set of the dataset to be labeled and the entity and entity category set of the supplementary dataset, that is, E B_need =E lack ∩E b , the obtained supplementary dataset can be used to supplement the entity category E of the dataset to be labeled B_need The number of entities is greater than or equal to 1. At this time, the supplementary dataset can be used to fine-tune the pre-trained first entity recognition model, that is, M B =FineTuning(M pretrained1 ,T B ,E B_need ), where M pretrained1 represents the pre-trained first entity recognition model, FineTuning represents fine-tuning, M B Represents the fine-tuned first entity recognition model, which enhances the recognition ability of the first entity recognition model for the existing target entity category. After obtaining the fine-tuned first entity recognition model, the fine-tuned first entity recognition model can be used to classify the original data set D a The text data set T aEntity recognition is performed on the text in the original dataset D a The text data set T a The text in the dataset can identify the entity and mark the corresponding first pseudo label. That is to say, in the dataset to be marked, the entity categories that the supplementary dataset can supplement for the dataset to be marked can be pseudo-labeled. c , original dataset D d As a supplementary data set, pseudo-label the target entity category that does not exist in the dataset to be labeled but exists in the supplementary data set, and obtain the first pseudo label E predicted by all supplementary data sets to the dataset to be labeled a_add .
[0105] If there is a target entity category e in the target entity category set n If it does not exist in all the original data sets, the entity category to be predicted can be obtained, and the pre-trained second entity recognition model M is used pretrained2 The pre-training ability of the original dataset D i The text data set T i The text in M is directly pseudo-labeled with the entity category to be predicted, that is, pretrained (T i ), that is, directly input the text in the text data set of the to-be-annotated dataset into the pre-trained second entity recognition model M pretrained2 By limiting the scope of entity categories to be identified to the entity categories to be predicted, the second pseudo-label of the dataset to be labeled can be obtained. Finally, the first pseudo-labels predicted by all supplementary datasets for the dataset to be labeled and the second pseudo-labels of the dataset to be labeled are merged to obtain the remaining unlabeled entities in the text of the text dataset of the dataset to be labeled and their corresponding pseudo-labels.
[0106] In one embodiment, the pre-trained first entity recognition model and the pre-trained second entity recognition model can use the same pre-trained entity recognition model, or different pre-trained entity recognition models can be selected. As an example, the pre-trained first entity recognition model and the pre-trained second entity recognition model can use the OneKE model. The OneKE model has been pre-trained in multiple fields and has cross-domain knowledge extraction capabilities, so it can predict these missing entity categories. In addition, the OneKE model uses a Schema-based polling instruction construction technology to extract entities within the specified entity category range.
[0107] By using the above strategies, we can continuously expand the target entity categories that are missing in the dataset to be labeled, thereby gradually enriching the scope of entity categories and ensuring that the deficiencies in the entity labeling of the original dataset are supplemented.
[0108] S3, using the pre-trained first language model and the target entity category set to score and filter the pseudo labels of the remaining unlabeled entities in the text of each original dataset, retain the pseudo labels with high confidence and their corresponding texts and combine them with the corresponding original dataset to generate the final dataset, forming the final dataset set.
[0109] In a specific embodiment, the pre-trained first language model and the target entity category set are used to score and filter the pseudo labels of the remaining unlabeled entities in the text of each original data set, specifically including:
[0110] Setting a prompt word, which is used to guide the pre-trained first language model to give a confidence score of the pseudo-label of each unlabeled entity and an overall score of the pseudo-labels of all unlabeled entities in each text based on each text of the text data set of each original data set input and the unlabeled entities and their pseudo-labels. In the scoring process, it is necessary to consider whether the entity recognition is correct, whether the pseudo-label is within the range of the target entity category set, and whether the pseudo-label is semantically appropriate in the text;
[0111] In response to determining that the confidence score of the pseudo-label of each unlabeled entity in the text is greater than or equal to a first threshold, and the overall score of the pseudo-labels of all unlabeled entities in the text is greater than or equal to a second threshold, determining that the pseudo-labels of all unlabeled entities in the text are high-confidence pseudo-labels, and retaining the corresponding text;
[0112] In response to determining that the confidence score of the pseudo-label of at least one unlabeled entity in the text is less than a first threshold, or the overall score of the pseudo-labels of all unlabeled entities in the text is less than a second threshold, the pseudo-labels of all unlabeled entities in the text are determined to be low-confidence pseudo-labels, and the corresponding text is discarded.
[0113] Specifically, the generated pseudo labels may have low confidence, especially when the entity recognition model is uncertain about the prediction of some entities. In order to improve data quality, it is necessary to score the pseudo labels and set a threshold based on the score to filter out pseudo labels with low confidence.
[0114] In one embodiment, a pre-trained first language model with strong instruction compliance and stable format output, such as Qwen2.5-72B-Instruct, is used, and appropriate prompt words are set to limit the range of pseudo-labels in the target entity category set, and each pseudo-label generated is scored. The score ranges from 0 to 10, and the score represents the credibility of the pseudo-label. At the same time, an overall scoring mechanism is introduced to appropriately reduce the overall score when the pseudo-label is lacking, so as to prevent the omission of pseudo-labels. The text in the text data set of the original data set and the entities identified in the text and their corresponding pseudo-labels are input into the pre-trained first language model with prompt words, and each pseudo-label can be output. The rating is s i , the overall score is S, and the threshold θ is set score (such as 8 points), when the score s of each pseudo label in a text i ≥θ score And the overall score S ≥ θ score When the score of at least one pseudo label in a text is s, the text and its pseudo label are retained. i ≥θ score Or the overall score S<θ score , the text is directly discarded. Since the original dataset retains the original entities and entity categories in the entity category set, the high-confidence pseudo-labels and their corresponding texts filtered by the above filtering and scoring strategies are combined with the original dataset to generate a complete final dataset. The final dataset retains the entity categories and high-confidence pseudo-labels in the original entity category set. All final datasets constitute the final dataset set.
[0115] As an example, the prompt might be: Objective: As a seasoned information extraction expert, you are given a score for a JSON dataset containing input text, entities, and predicted pseudo-labels. Your task is to evaluate the accuracy of the predicted pseudo-labels for these entities. Scoring should follow the following rules: 1. Assign a confidence score to each predicted pseudo-label and then score the overall prediction on a scale of 0 to 10, with 10 indicating a completely correct prediction and 0 indicating a completely incorrect prediction. 2. Scoring should consider whether the entity recognition is correct, whether the entity category meets the requirements, and whether its semantics are appropriate within the context. 3. Entity categories are limited to the following range: ["medical procedure", "disease", "clinical manifestation", "medical test item", "body part", "cause", "drug"]; Requirements: Output directly in a JSON format similar to the "scoring result" in "Example", without explanation. Example: Input text, recognized text, and predicted pseudo-labels: {example text}; Scoring result: {example output}.
[0116] {example text} = '{"text":"31, degree 500","entity":[{"entity":"degree 500","entity_type":"clinical manifestations"}]}';
[0117] example output = '{"text":"31,degree 500","entity":[{"entity":"degree 500","entity type":"clinical manifestation","score":"0"}],"total score":"0"}'.
[0118] S4, using the final dataset set to fine-tune the pre-trained second largest language model to obtain the entity recognition model corresponding to the target named entity recognition task; obtain the text to be recognized and input it into the entity recognition model corresponding to the target named entity recognition task to identify the entity and its corresponding entity category.
[0119] In a specific embodiment, the pre-trained first language model includes the Qwen2.5-72B-Instruct model, the pre-trained second language model includes the Qwen2.5-0.5B-Instruct model, and during the fine-tuning process of the pre-trained second language model, the entity category of each text is limited to be within the range of the target entity category set through the Schema instruction.
[0120] Specifically, in order to further improve the performance and efficiency of the entity recognition model, the embodiment of the present application adopts a strategy of fine-tuning the small model to optimize the reasoning speed and accuracy. The fine-tuning process is based on a pseudo-label dataset and combined with the Schema polling instruction construction technology. In one embodiment, an appropriate small model (such as Qwen2.5-0.5B-Instruct) is selected as the second largest language model after pre-training, and fine-tuned through the following steps:
[0121] (1) Training based on the final dataset. Through the Schema instruction construction, specify the entity category of each input text and annotate the input text.
[0122] (2) Optimize the model’s inference speed and accuracy. By simplifying the instruction format and adjusting the output strategy, the inference time is reduced and the accuracy of the model is improved.
[0123] The format of the prompt words used in the fine-tuning process is as follows:
[0124]
[0125] Among them, Instruction": "You are an expert in entity extraction. Please extract the entity categories that conform to the definition in the schema from the input, accurately identify all entities that meet the requirements, and output them in JSON data format: return {\"entity type\": [\"entity\"]} if there is an entity, and return {} if there is no entity.
[0126] Note: 1. Please strictly follow the entity categories defined in the schema for extraction; 2. The output must be in the correct JSON data format."; "schema": ["medical procedure","disease","clinical manifestation","medical test items","site","cause","drug","organization","name","address","price","age","time","weight","height"]; "input" is the input text.
[0127] During fine-tuning, text from the final dataset is input into a pre-trained second-largest language model. The above prompts guide the pre-trained second-largest language model to perform entity recognition on the input text, limiting the output entity categories to the target entity category set. The entities and entity categories in the final dataset serve as labels, enabling the pre-trained second-largest language model to learn to perform entity recognition on text and complete the target named entity recognition task, resulting in a fine-tuned second-largest language model. The fine-tuned second-largest language model is the entity recognition model corresponding to the target named entity recognition task.
[0128] The entity recognition model for the target named entity recognition task dynamically generates entity annotations that meet the requirements based on this instruction and outputs them in the corresponding JSON format. Through multiple optimization experiments and using an improved prompt word data format, the entity recognition model for the target named entity recognition task is able to consistently output JSON data that meets the requirements and significantly reduce inference costs. Experimental data shows that, when tested on a manually annotated data set containing 300 test data items, the entity recognition model for the target named entity recognition task improved its accuracy from 50% to 90%.
[0129] Further references Figure 2 As an implementation of the methods shown in the above figures, this application provides an embodiment of a multi-domain entity recognition device under data scarcity conditions. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0130] The present invention provides a method and apparatus for multi-domain entity recognition under data scarcity conditions, including:
[0131] A target entity category determination module 1 is configured to determine a target entity category set required for a target named entity recognition task, where the target entity category set includes a plurality of target entity categories;
[0132] The pseudo-label prediction module 3 is configured to obtain a set of original data sets, each of which contains a text data set and a set of entities and entity categories. Under the condition of data scarcity, the entity and entity category set of each original data set stores entities that annotate the text of the text data set and their corresponding entity categories that do not completely cover the target entity category set; based on each original data set and the target entity category set, the pseudo-labels of the remaining unannotated entities in the text of the text data set of each original data set are annotated;
[0133] The pseudo-label filtering module 3 is configured to score and filter the pseudo-labels of the remaining unlabeled entities in the text of each original data set using the pre-trained first language model and the target entity category set, retain the pseudo-labels with high confidence and their corresponding texts, and combine them with the corresponding original data set to generate a final data set to constitute a final data set set;
[0134] The entity recognition module 4 is configured to use the final data set set to fine-tune the pre-trained second largest language model to obtain an entity recognition model corresponding to the target named entity recognition task; obtain the text to be recognized and input it into the entity recognition model corresponding to the target named entity recognition task to recognize the entity and its corresponding entity category.
[0135] Figure 3 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Figure 3 As shown, the electronic device of this embodiment includes: a processor 301 and a memory 302; wherein the memory 302 is used to store computer-executable instructions; and the processor 301 is used to execute the computer-executable instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant description of the above method embodiment.
[0136] Optionally, the memory 302 may be independent or integrated with the processor 301 .
[0137] When the memory 302 is independently provided, the electronic device further includes a bus 303 for connecting the memory 302 and the processor 301 .
[0138] An embodiment of the present invention further provides a computer storage medium, in which computer execution instructions are stored. When the processor 301 executes the computer execution instructions, the above method is implemented.
[0139] An embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 301, the above method is implemented.
[0140] In the embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical functional division. In actual implementation, other division methods may be used. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not implemented. In addition, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or module, which may be electrical, mechanical or other forms.
[0141] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of these modules may be selected to implement the solution of this embodiment based on actual needs.
[0142] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each module may exist physically separately, or two or more modules may be integrated into a single unit. The units formed by the above modules may be implemented in the form of hardware or hardware plus software functional units.
[0143] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or processor 301 to perform some steps of the methods of various embodiments of the present application.
[0144] It should be understood that the processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), or application-specific integrated circuits (ASIC). A general-purpose processor may be a microprocessor, or the processor 301 may be any conventional processor 301. The steps of the method disclosed in the present invention may be directly implemented by the hardware processor 301 or implemented by a combination of hardware and software modules in the processor 301.
[0145] The memory 302 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disk.
[0146] Bus 303 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Bus 303 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, the bus 303 in the drawings of this application is not limited to a single bus 303 or a single type of bus 303.
[0147] The storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0148] An exemplary storage medium is coupled to the processor 301, so that the processor 301 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor 301. The processor 301 and the storage medium can be located in an application-specific integrated circuit (ASIC). Of course, the processor 301 and the storage medium can also exist as discrete components in an electronic device or a host control device.
[0149] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-domain entity recognition method under data scarcity conditions, characterized by: The following steps are involved: Determine a target entity category set required for a target named entity recognition task, wherein the target entity category set includes a plurality of target entity categories; Obtain a set of original data sets, each original data set in the original data set includes a text data set and an entity and entity category set. Under data scarcity conditions, the entity and entity category set of each original data set stores entities that annotate the text of the text data set and their corresponding entity categories that do not completely cover the target entity category set; based on each original data set and the target entity category set, annotate pseudo labels of remaining unannotated entities in the text of the text data set of each original data set; Using the pre-trained first language model and the target entity category set, the pseudo labels of the remaining unlabeled entities in the text of the text data set of each original data set are scored and filtered, and the pseudo labels with high confidence and their corresponding texts are retained and combined with the corresponding original data sets to generate a final data set to form a final data set set; The final dataset set is used to fine-tune the pre-trained second largest language model to obtain an entity recognition model corresponding to the target named entity recognition task; the text to be recognized is obtained and input into the entity recognition model corresponding to the target named entity recognition task to identify the entity and its corresponding entity category.
2. The multi-domain entity recognition method under data scarcity conditions according to claim 1 is characterized in that: Based on each original data set and the target entity category set, pseudo labels of the remaining unlabeled entities in the text of the text data set of each original data set are labeled, specifically including: S21, traversing one of the original data sets in the original data set set and taking it as the data set to be labeled, and forming a missing entity category set of the data set to be labeled with some target entity categories in the target entity category set that do not exist in the entity and entity category set of the data set to be labeled; S22, traversing one of the remaining original data sets in the original data set set except the data set to be labeled and using it as a supplementary data set, taking the intersection of the missing entity category set of the data set to be labeled and the entity and entity category set of the supplementary data set, obtaining the entity categories that the supplementary data set can supplement for the data set to be labeled, and in response to determining that the number of entity categories that the supplementary data set can supplement for the data set to be labeled is greater than or equal to 1, performing entity recognition on the text in the text data set of the data set to be labeled by using the text data set of the supplementary data set and the entity categories that the supplementary data set can supplement for the data set to be labeled, and obtaining a first pseudo label predicted for the data set to be labeled using the supplementary data set; S23, repeating steps S22-S23 until all the original data sets except the data set to be labeled are traversed, and a first pseudo label predicted for the data set to be labeled using each supplementary data set is obtained; S24, taking some target entity categories in the target entity category set that do not exist in all entities and entity category sets of the dataset to be labeled as entity categories to be predicted, and performing entity recognition based on the entity categories to be predicted on the text in the text dataset of the dataset to be labeled to obtain a second pseudo label for the dataset to be labeled; S25, taking the union of the first pseudo labels predicted by all supplementary datasets for the dataset to be labeled and the second pseudo labels of the dataset to be labeled, to obtain the remaining unlabeled entities in the text of the text dataset of the dataset to be labeled and their corresponding pseudo labels; S26, repeating steps S21-S26 until all original data sets in the original data set set are traversed, and obtaining the remaining unlabeled entities in the text of the text data set of each original data set and their corresponding pseudo labels.
3. The multi-domain entity recognition method under data scarcity conditions according to claim 2 is characterized in that: Performing entity recognition on text in the text dataset of the dataset to be labeled using the text dataset of the supplementary dataset and the entity categories that the supplementary dataset can supplement for the dataset to be labeled, specifically includes: The pre-trained first entity recognition model is fine-tuned using the text data set of the supplementary data set and the entity categories that the supplementary data set can supplement for the data set to be labeled, thereby obtaining a fine-tuned first entity recognition model; and the fine-tuned first entity recognition model is used to perform entity recognition on the text in the text data set of the data set to be labeled.
4. The multi-domain entity recognition method under data scarcity conditions according to claim 2 is characterized in that: Performing entity recognition based on the entity category to be predicted on the text in the text data set of the to-be-annotated data set, specifically comprising: A pre-trained second entity recognition model is used to perform entity recognition on the text in the text data set of the data set to be labeled, and the range of the recognized entity categories is limited to cover the entity categories to be predicted.
5. The multi-domain entity recognition method under data scarcity conditions according to claim 1 is characterized in that: The pre-trained first language model and the target entity category set are used to score and filter the pseudo labels of the remaining unlabeled entities in the text of each original data set, specifically including: Setting a prompt word, wherein the prompt word is used to guide the pre-trained first large language model to give a confidence score of the pseudo-label of each unlabeled entity and an overall score of the pseudo-labels of all unlabeled entities in each text according to each text of the text data set of each original data set input and the unlabeled entities and their pseudo-labels therein, and in the scoring process, it is necessary to consider whether the entity recognition is correct, whether the pseudo-label is within the range of the target entity category set, and whether the semantics of the pseudo-label in the text are appropriate; In response to determining that the confidence score of the pseudo-label of each unlabeled entity in the text is greater than or equal to a first threshold, and the overall score of the pseudo-labels of all unlabeled entities in the text is greater than or equal to a second threshold, determining that the pseudo-labels of all unlabeled entities in the text are high-confidence pseudo-labels, and retaining the corresponding text; In response to determining that the confidence score of the pseudo-label of at least one unlabeled entity in the text is less than a first threshold, or the overall score of the pseudo-labels of all unlabeled entities in the text is less than a second threshold, the pseudo-labels of all unlabeled entities in the text are determined to be low-confidence pseudo-labels, and the corresponding text is discarded.
6. The multi-domain entity recognition method under data scarcity conditions according to claim 1 is characterized in that: The pre-trained first large language model includes a Qwen2.5-72B-Instruct model, and the pre-trained second large language model includes a Qwen2.5-0.5B-Instruct model. In the fine-tuning process of the pre-trained second large language model, the entity category of each text is limited to be within the scope of the target entity category set through the Schema instruction.
7. A method and device for multi-domain entity recognition under data scarcity conditions, characterized in that: include: a target entity category determination module, configured to determine a target entity category set required for a target named entity recognition task, wherein the target entity category set includes a plurality of target entity categories; The pseudo-label prediction module is configured to obtain a set of original data sets, each of which contains a text data set and a set of entities and entity categories. Under data scarcity conditions, the entity and entity category set of each original data set stores entities that annotate the text of the text data set and their corresponding entity categories that do not completely cover the target entity category set; based on each original data set and the target entity category set, annotate the pseudo-labels of the remaining unannotated entities in the text of the text data set of each original data set; a pseudo-label filtering module configured to score and filter the pseudo-labels of the remaining unlabeled entities in the text of each original data set using the pre-trained first language model and the target entity category set, retain the pseudo-labels with high confidence and their corresponding texts, and combine them with the corresponding original data set to generate a final data set to constitute a final data set set; The entity recognition module is configured to use the final data set to fine-tune the pre-trained second largest language model to obtain an entity recognition model corresponding to the target named entity recognition task; obtain the text to be recognized and input it into the entity recognition model corresponding to the target named entity recognition task, and recognize the entity and its corresponding entity category.
8. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.