A text data annotation method and system based on weakly supervised learning

Through a weakly supervised learning method, pseudo-label generation and deep learning and adversarial training algorithms are used to build a text data annotation model, solving the problems of high cost of text data annotation and uneven labeling quality in the existing technology, achieving efficient and accurate text data annotation, and reducing dependence on label data.

CN119669477BActive Publication Date: 2025-05-13YUNHAI SPACETIME (BEIJING) TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510168698.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-13
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

In the prior art, text data has high cost of labeling, uneven labeling quality, and high dependence on label data, making it difficult to guarantee efficiency and accuracy.

Method used

Using a method based on weakly supervised learning, a text data annotation model is constructed through pseudo-label generation and deep learning and adversarial training algorithms, and a continuous learning algorithm is used to adjust it to reduce manual annotation work and improve labeling efficiency and accuracy.

Benefits of technology

It significantly reduces the labeling cost, improves the overall efficiency and accuracy of data labeling, reduces dependence on labeling data, and improves the generalization ability and stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669477B_ABST
    Figure CN119669477B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of natural language processing, and discloses a text data annotation method and system based on weakly supervised learning. The method comprises the following steps: using a weakly supervised learning algorithm to generate pseudo labels for a number of historical text data, and obtaining a number of source domain data with real labels and a number of target domain data with pseudo labels; using a deep learning and adversarial training algorithm to construct a text data annotation model, and using a continuous learning algorithm to adjust the text data annotation model; using the adjusted text data annotation model to annotate real-time text data, and obtaining annotated real-time text data, and using a continuous learning algorithm to update the adjusted text data annotation model. The present invention solves the problems of high manual annotation cost, uneven annotation quality, and high dependence on label data in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a text data annotation method and system based on weakly supervised learning. Background Art

[0002] In the field of natural language processing (NLP), text data annotation is one of the important steps in building machine learning models. Annotation usually involves assigning category labels or annotations to text data so that the model can learn how to classify or interpret unlabeled data. However, text data annotation is a time-consuming and costly process, especially when it requires a lot of expertise and human resources.

[0003] The existing text data annotation technology has the following defects:

[0004] 1) High cost of manual annotation: Traditional text data annotation methods rely on a lot of manual operations, which is not only time-consuming and labor-intensive, but also costly. In scenarios where large-scale annotated data is required, this cost becomes a limiting factor;

[0005] 2) Uneven annotation quality: Due to the subjectivity of manual annotation, the annotation standards and quality may vary among different annotators, making it difficult to ensure the consistency and accuracy of the annotation data;

[0006] 3) High dependence on labeled data: In existing technologies, a large amount of labeled data is required for model training. The dependence on labeled data is high, and unlabeled data is often ignored. Summary of the invention

[0007] In order to solve the problems of high manual annotation cost, uneven annotation quality and high dependence on label data in the prior art, the present invention aims to provide a text data annotation method and system based on weakly supervised learning.

[0008] The technical solution adopted by the present invention is:

[0009] A text data labeling method based on weakly supervised learning includes the following steps:

[0010] Collect some historical text data, and use a weakly supervised learning algorithm to generate pseudo labels for some historical text data, so as to obtain some source domain data with real labels and some target domain data with pseudo labels;

[0011] According to a number of source domain data and a number of target domain data, a text data annotation model is constructed using a deep learning and adversarial training algorithm, and the text data annotation model is adjusted using a continuous learning algorithm to obtain an adjusted text data annotation model;

[0012] Real-time text data is collected, and the real-time text data is annotated using the adjusted text data annotation model to obtain annotated real-time text data, and the adjusted text data annotation model is updated using a continuous learning algorithm to obtain an updated text data annotation model.

[0013] Furthermore, a number of historical text data are collected, and a weakly supervised learning algorithm is used to generate pseudo labels for the historical text data, so as to obtain a number of source domain data with real labels and a number of target domain data with pseudo labels, including the following steps:

[0014] Collecting a number of historical text data, and preprocessing the number of historical text data to obtain a number of preprocessed historical text data;

[0015] The preprocessed historical text data with real labels are divided into source domain data, and the preprocessed historical text data without real labels are divided into original data;

[0016] Based on several source domain data with real labels, a pseudo-label generation model is constructed using a weakly supervised learning algorithm.

[0017] The pseudo-label generation model is used to generate pseudo-labels for the original data, and a number of target domain data with pseudo-labels are obtained.

[0018] Furthermore, the pseudo-label generation model is constructed based on the LSTM-cGAN algorithm, and the pseudo-label generation model includes a sequence feature extraction module constructed based on the LSTM algorithm and a pseudo-label generation module constructed based on the cGAN algorithm, which are connected in sequence. The pseudo-label generation module includes a generator and a discriminator that are connected in sequence and are both constructed based on the RNN algorithm.

[0019] Furthermore, based on a number of source domain data with real labels, a weakly supervised learning algorithm is used to construct a pseudo label generation model, including the following steps:

[0020] Using a weakly supervised learning algorithm, an initial pseudo-label generation model is constructed; the initial pseudo-label generation model includes an initial sequence feature extraction module and an initial pseudo-label generation module;

[0021] Integrate the first loss function of the generator of the initial pseudo-label generation module and the second loss function of the discriminator to obtain a first comprehensive loss function of the initial pseudo-label generation module;

[0022] According to a number of source domain data with real labels, the initial sequence feature extraction module is optimized and trained to obtain a final sequence feature extraction module, and a number of historical sequence features are obtained;

[0023] According to the historical sequence characteristics of the source domain data, the generator of the initial pseudo-label generation module is used to generate pseudo-labels for the source domain data;

[0024] Use the discriminator of the initial pseudo-label generation module to perform label authenticity analysis based on the pseudo-labels and true labels of the source domain data to obtain the historical label authenticity analysis results;

[0025] According to the authenticity analysis results of the pseudo-labels and historical labels, the first comprehensive loss function of the initial pseudo-label generation module is used to obtain the corresponding first historical loss value;

[0026] Traverse all source domain data and perform the optimization training step of the pseudo label generation module. If the first historical loss value is less than the first loss value threshold, output the final pseudo label generation module.

[0027] The final sequence feature extraction module and the final pseudo-label generation module are integrated to obtain the final pseudo-label generation model.

[0028] Furthermore, a pseudo-label generation model is used to generate pseudo-labels for the original data to obtain a number of target domain data with pseudo-labels, including the following steps:

[0029] Use the sequence feature extraction module of the pseudo-label generation model to extract the second historical sequence features of the original data without setting the real label;

[0030] Using a pseudo-label generation module of a pseudo-label generation model, generating pseudo-labels according to the second historical sequence characteristics, to obtain pseudo-labels for the original data;

[0031] Traverse all the original data without real labels, perform the above pseudo-label generation step, and obtain a number of target domain data with pseudo labels.

[0032] Further, according to the plurality of source domain data and the plurality of target domain data, a text data annotation model is constructed using a deep learning and adversarial training algorithm, and the text data annotation model is adjusted using a continuous learning algorithm to obtain an adjusted text data annotation model, including the following steps:

[0033] Use deep learning and adversarial training algorithms to build an initial text data annotation model;

[0034] According to a number of source domain data and a number of target domain data, the initial text data annotation model is optimized and trained to obtain an optimized text data annotation model, and a number of historical text data annotation experiences are obtained;

[0035] Use the experience replay mechanism of the continuous learning algorithm to build an experience replay pool, and use the elastic weight connection mechanism of the continuous learning algorithm to build an elastic loss function;

[0036] According to the experience replay pool and the elastic loss function, the text data annotation model is adjusted to obtain an adjusted text data annotation model, and several historical text data annotation experiences are stored in the experience replay pool.

[0037] Furthermore, the text data annotation model is constructed based on the DBN-DANN-Elman algorithm, and the text data annotation model includes a text feature extraction model constructed based on the DBN algorithm, a domain adversarial training module constructed based on the DANN algorithm, and a text data annotation module constructed based on the Elman algorithm, which are connected in sequence. The domain adversarial training module includes a label predictor and a domain classifier connected in sequence.

[0038] Furthermore, according to the plurality of source domain data and the plurality of target domain data, the initial text data annotation model is optimized and trained to obtain an optimized text data annotation model, and a plurality of historical text data annotation experiences are obtained, including the following steps:

[0039] Pre-training an initial text feature extraction model and an initial text data annotation module of an initial text data annotation model according to a number of target domain data to obtain a pre-trained text feature extraction model and a pre-trained text data annotation module;

[0040] According to a number of source domain data and a number of target domain data, a pre-trained text feature extraction model and a pre-trained text data annotation module are optimized and trained to obtain a final feature extraction model and a final text data annotation module, and obtain a number of historical text features;

[0041] Integrate the third loss function of the label predictor of the initial text data annotation model and the fourth loss function of the domain classifier to obtain the second comprehensive loss function of the initial domain adversarial training module;

[0042] According to the historical text features, a label predictor is used to generate corresponding historical label predictions, and a domain classifier is used to generate corresponding historical domain classification results;

[0043] According to the historical label prediction and the historical domain classification results, the second comprehensive loss function of the initial domain adversarial training module is used to obtain the corresponding second historical loss value;

[0044] The historical text features and historical model parameters of the text data annotation model generated in each optimization training are retained to obtain historical text data annotation experience;

[0045] Traversing all source domain data and target domain data, performing the optimization training step of the above domain adversarial training module, and outputting the final domain adversarial training module if the second historical loss value is less than the second loss value threshold;

[0046] Integrate the final text feature extraction model, the final domain adversarial training module, and the final text data annotation module to obtain an optimized text data annotation model and some historical text data annotation experiences.

[0047] Further, real-time text data is collected, and the real-time text data is annotated using the adjusted text data annotation model to obtain annotated real-time text data, and the adjusted text data annotation model is updated using a continuous learning algorithm to obtain an updated text data annotation model, including the following steps:

[0048] Collecting real-time text data, and preprocessing the real-time text data to obtain preprocessed real-time text data;

[0049] Extracting real-time text features of the preprocessed real-time text data using a text feature extraction model of the adjusted text data annotation model;

[0050] Using the text data annotation module of the adjusted text data annotation model, annotating the real-time text data according to the real-time text features to obtain annotated real-time text data;

[0051] The real-time text features and the real-time model parameters of the adjusted text data annotation model are retained to obtain real-time text data annotation experience;

[0052] Randomly extract a number of historical text data annotation experiences from the experience replay pool, and mix them with the real-time text data annotation experiences to obtain a number of mixed text data annotation experiences;

[0053] Based on the mixed text data annotation experience, the adjusted text data annotation model is continuously trained, and the elastic loss function is used to obtain the elastic loss value of the continuous training;

[0054] Traverse all mixed text data annotation experiences and perform the above continuous training steps. If the elastic loss value is less than the third loss value threshold, output the updated text data annotation model.

[0055] A text data annotation system based on weakly supervised learning is used to implement a text data annotation method. The system comprises a weakly supervised learning unit, a model building unit, and an annotation and updating unit which are connected in sequence.

[0056] The beneficial effects of the present invention are:

[0057] The invention provides a text data annotation method and system based on weak supervised learning. By using a weak supervised learning algorithm to automatically generate pseudo labels, the data annotation work with manual participation is greatly reduced, thereby significantly reducing the annotation cost. By using an automated text data annotation model, a large amount of unlabeled text data can be quickly processed, thereby improving the overall efficiency of data annotation, especially in scenarios requiring rapid response. By combining deep learning and adversarial training algorithms, the accuracy and consistency of annotation can be improved, human errors can be reduced, and a small amount of labeled data and a large amount of unlabeled data can be used to automatically generate high-quality annotation results, thereby reducing the dependence on labeled data. By generating pseudo labels, unlabeled data can be fully utilized to extract valuable information from it, thereby improving the utilization efficiency of unlabeled data. The model is adjusted by a continuous learning algorithm, so that the model can better adapt to new data and changes, thereby improving the generalization ability of the model, and enabling the model to continuously learn from new data to maintain the timeliness and accuracy of the model. By using an experience replay pool and an elastic loss function, model training can be more effectively performed, thereby reducing the risk of overfitting and improving the stability and performance of the model.

[0058] Other beneficial effects of the present invention will be further described in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a flowchart of the text data labeling method based on weakly supervised learning in the present invention.

[0060] Figure 2 It is a structural block diagram of the text data annotation system based on weakly supervised learning in the present invention. DETAILED DESCRIPTION

[0061] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments.

[0062] Embodiment 1:

[0063] like Figure 1 As shown, this embodiment provides a text data labeling method based on weakly supervised learning, comprising the following steps:

[0064] S1: Collect some historical text data, and use a weakly supervised learning algorithm to generate pseudo labels for some historical text data, so as to obtain some source domain data with real labels and some target domain data with pseudo labels, including the following steps:

[0065] S1-1: collecting some historical text data, and preprocessing the some historical text data to obtain some preprocessed historical text data;

[0066] Preprocessing includes text cleaning, word segmentation, stop word removal, stemming, etc., to improve data quality and provide data support for subsequent model training;

[0067] S1-2: dividing the preprocessed historical text data with real labels into source domain data, and dividing the preprocessed historical text data without real labels into original data;

[0068] Clear data division helps to train models and generate pseudo labels in a targeted manner, improving the generalization ability of the model and the accuracy of annotation;

[0069] S1-3: Based on several source domain data with real labels, a pseudo-label generation model is constructed using a weakly supervised learning algorithm. The pseudo-label generation model can effectively utilize unlabeled data, reduce dependence on a large amount of manually labeled data, and reduce costs.

[0070] The pseudo-label generation model is constructed based on the Long Short-Term Memory (LSTM)-Conditional Generative Adversarial Network (cGAN) algorithm, and the pseudo-label generation model includes a sequence feature extraction module constructed based on the LSTM algorithm and a pseudo-label generation module constructed based on the cGAN algorithm, which are connected in sequence. The pseudo-label generation module includes a generator and a discriminator that are connected in sequence and are both constructed based on the Recurrent Neural Network (RNN) algorithm.

[0071] The sequence feature extraction module is used to capture the word order information and contextual relationships in the sentence and extract the deep features of the input data. The generator attempts to generate pseudo labels that can deceive the discriminator, while the discriminator attempts to better identify real and false labels. The pseudo label generation module performs well in generating labels, especially when there is not enough new data. The generator can help generate labeled data for training, and achieves the generation of sufficient training samples in the case of a small amount of new data, ensuring the diversity of training samples and enhancing the generalization ability of the pseudo label generation module, so as to better process real new data.

[0072] Based on several source domain data with real labels, a pseudo label generation model is constructed using a weakly supervised learning algorithm, including the following steps:

[0073] S1-3-1: Use a weakly supervised learning algorithm to build an initial pseudo-label generation model; the initial pseudo-label generation model includes an initial sequence feature extraction module and an initial pseudo-label generation module;

[0074] S1-3-2: Integrate the first loss function of the generator of the initial pseudo-label generation module and the second loss function of the discriminator to obtain the first comprehensive loss function of the initial pseudo-label generation module; this helps to balance the performance of the generator and the discriminator during the training process;

[0075] S1-3-3: According to a number of source domain data with real labels, the initial sequence feature extraction module is optimized and trained to obtain a final sequence feature extraction module, and a number of historical sequence features are obtained;

[0076] S1-3-4: Based on the historical sequence characteristics of the source domain data, the generator of the initial pseudo-label generation module is used to generate pseudo-labels for the source domain data. Pseudo-label generation allows unlabeled data to be used for training, thereby expanding the scale of the training data set and helping to improve the robustness and accuracy of the model.

[0077] S1-3-5: Use the discriminator of the initial pseudo-label generation module to perform label authenticity analysis based on the pseudo-labels and true labels of the source domain data to obtain the historical label authenticity analysis results;

[0078] S1-3-6: According to the authenticity analysis results of the pseudo-labels and historical labels, the first comprehensive loss function of the initial pseudo-label generation module is used to obtain the corresponding first historical loss value;

[0079] S1-3-7: traverse all source domain data and perform the optimization training step of the pseudo label generation module. If the first historical loss value is less than the first loss value threshold, output the final pseudo label generation module;

[0080] S1-3-8: Integrate the final sequence feature extraction module and the final pseudo-label generation module to obtain the final pseudo-label generation model;

[0081] S1-4: Use the pseudo-label generation model to generate pseudo-labels for the original data to obtain a number of target domain data with pseudo-labels, including the following steps:

[0082] S1-4-1: Use the sequence feature extraction module of the pseudo-label generation model to extract the second historical sequence features of the original data without setting the real label; the sequence feature extraction module can capture the sequence dependency and context information in the text data and convert the original text into a set of meaningful feature vectors. By extracting effective sequence features, the model can better understand the text content and provide high-quality feature input for subsequent pseudo-label generation;

[0083] S1-4-2: Use the pseudo-label generation module of the pseudo-label generation model to generate pseudo-labels according to the second historical sequence characteristics to obtain pseudo-labels of the original data; by generating pseudo-labels, the scale of the labeled data set can be expanded to provide more learning signals for model training;

[0084] S1-4-3: Traverse all the original data without real labels, perform the above pseudo-label generation steps, and obtain a number of target domain data with pseudo-labels; by using the pseudo-label generation model to automatically generate labels for a large amount of unlabeled text data, the efficiency and scale of data labeling are significantly improved;

[0085] S2: Based on a number of source domain data and a number of target domain data, a text data annotation model is constructed using a deep learning and adversarial training algorithm, and the text data annotation model is adjusted using a continuous learning algorithm to obtain an adjusted text data annotation model, including the following steps:

[0086] S2-1: Use deep learning and adversarial training algorithms to build an initial text data annotation model;

[0087] The text data annotation model is constructed based on a deep belief network (DBN)-domain adversarial network (DANN)-Elman algorithm, and the text data annotation model includes a text feature extraction model constructed based on the DBN algorithm, a domain adversarial training module constructed based on the DANN algorithm, and a text data annotation module constructed based on the Elman algorithm, which are connected in sequence. The domain adversarial training module includes a label predictor and a domain classifier connected in sequence.

[0088] S2-2: Based on a number of source domain data and a number of target domain data, the initial text data annotation model is optimized and trained to obtain an optimized text data annotation model, and a number of historical text data annotation experiences are obtained, including the following steps:

[0089] S2-2-1: pre-training an initial text feature extraction model and an initial text data annotation module of an initial text data annotation model according to a number of target domain data, to obtain a pre-trained text feature extraction model and a pre-trained text data annotation module;

[0090] S2-2-2: According to a number of source domain data and a number of target domain data, the pre-trained text feature extraction model and the pre-trained text data annotation module are optimized and trained to obtain a final feature extraction model and a final text data annotation module, and obtain a number of historical text features;

[0091] S2-2-3: Integrate the third loss function of the label predictor of the initial text data annotation model and the fourth loss function of the domain classifier to obtain the second comprehensive loss function of the initial domain adversarial training module;

[0092] S2-2-4: Based on the historical text features, use the label predictor to generate the corresponding historical label prediction, and use the domain classifier to generate the corresponding historical domain classification results;

[0093] S2-2-5: According to the historical label prediction and the historical domain classification results, the second comprehensive loss function of the initial domain adversarial training module is used to obtain the corresponding second historical loss value;

[0094] S2-2-6: retain the historical text features and historical model parameters of the text data annotation model generated in each optimization training to obtain historical text data annotation experience;

[0095] S2-2-7: Traverse all source domain data and target domain data, perform the optimization training step of the above domain adversarial training module, and if the second historical loss value is less than the second loss value threshold, output the final domain adversarial training module;

[0096] S2-2-8: Integrate the final text feature extraction model, the final domain adversarial training module, and the final text data annotation module to obtain an optimized text data annotation model and obtain some historical text data annotation experience;

[0097] S2-3: Use the experience replay mechanism of the continuous learning algorithm to build an experience replay pool, and use the elastic weight connection mechanism of the continuous learning algorithm to build an elastic loss function;

[0098] The experience replay pool is used to store the interactive experience of the model in the environment. These experiences are usually stored in the form of (s, a, r, s'), where s is the current state, including the current text features and model parameters, a is the model parameter adjustment action taken, r is the reward obtained, which is used to characterize the impact of the action on the state, and s' is the next state, that is, the state after the model parameters are adjusted during the training process; the experience replay pool is a database that stores previous task learning experience, including parameter settings, model status, learning strategies, etc. Joining the experience pool allows the model to store and reuse past experience to improve learning efficiency;

[0099] Since the data in the experience replay pool is randomly drawn, it helps to break the correlation between consecutive experiences, thereby reducing the variance in model training. By reusing experience, the model can learn more from limited experience, which is especially useful in new tasks or when samples are scarce. Experience replay helps stabilize the learning process and reduce fluctuations during training. Regularly check the size of the experience replay pool. If it exceeds the preset capacity, remove some experience according to a certain strategy (such as priority sampling, longest unvisited).

[0100] When training a new task, in order to prevent the model from forgetting the previously learned experience, an additional penalty term can be added to the loss function using the Elastic Weight Consolidation (EWC) method. This penalty term is proportional to the importance measure of the weight and inversely proportional to the change in the weight on the new task. This penalty term ensures that when training on a new task, the weights that are important for the old task do not change too much, thereby reducing the risk of catastrophic forgetting. The larger the importance measure, the more restricted the change of the corresponding weight on the new task. In this way, the model can retain the experience of the old task while learning the new task;

[0101] S2-4: According to the experience replay pool and the elastic loss function, the text data annotation model is adjusted to obtain an adjusted text data annotation model, and several historical text data annotation experiences are stored in the experience replay pool;

[0102] S3: collecting real-time text data, using the adjusted text data annotation model to annotate the real-time text data, obtaining annotated real-time text data, and using a continuous learning algorithm to update the adjusted text data annotation model to obtain an updated text data annotation model, including the following steps:

[0103] S3-1: Collect real-time text data and preprocess the real-time text data to obtain preprocessed real-time text data; the preprocessing steps include cleaning, word segmentation, and vectorization to prepare standard data, which helps to improve the accuracy of annotation and make it suitable for model processing;

[0104] S3-2: using the text feature extraction model of the adjusted text data annotation model to extract real-time text features of the preprocessed real-time text data;

[0105] S3-3: using the text data annotation module of the adjusted text data annotation model to annotate the real-time text data according to the real-time text features, thereby obtaining annotated real-time text data;

[0106] S3-4: retaining the real-time text features and the real-time model parameters of the adjusted text data annotation model to obtain real-time text data annotation experience;

[0107] S3-5: Randomly extract some historical text data annotation experiences from the experience playback pool and mix them with the real-time text data annotation experiences to obtain some mixed text data annotation experiences; prevent the model from forgetting key information in historical data and improve the generalization ability of the model;

[0108] S3-6: Based on the mixed text data annotation experience, the adjusted text data annotation model is continuously trained, and the elastic loss function is used to obtain the elastic loss value of the continuous training; through continuous training, the model can adapt to the new data distribution while retaining the ability to process old data;

[0109] S3-7: Traverse all mixed text data annotation experiences and perform the above-mentioned continuous training steps. If the elasticity loss value is less than the third loss value threshold, output the updated text data annotation model.

[0110] Embodiment 2:

[0111] like Figure 2 As shown, this embodiment provides a text data annotation system based on weakly supervised learning, which is used to implement a text data annotation method. The system includes a weakly supervised learning unit, a model building unit, and an annotation and update unit connected in sequence;

[0112] A weakly supervised learning unit is used to collect a number of historical text data, and use a weakly supervised learning algorithm to generate pseudo labels for the number of historical text data, so as to obtain a number of source domain data with real labels and a number of target domain data with pseudo labels;

[0113] A model building unit is used to build a text data annotation model based on a plurality of source domain data and a plurality of target domain data using a deep learning and adversarial training algorithm, and to adjust the text data annotation model using a continuous learning algorithm to obtain an adjusted text data annotation model;

[0114] The annotation and updating unit is used to collect real-time text data, use the adjusted text data annotation model to annotate the real-time text data to obtain annotated real-time text data, and use a continuous learning algorithm to update the adjusted text data annotation model to obtain an updated text data annotation model.

[0115] The invention provides a text data annotation method and system based on weak supervised learning. By using a weak supervised learning algorithm to automatically generate pseudo labels, the data annotation work with manual participation is greatly reduced, thereby significantly reducing the annotation cost. By using an automated text data annotation model, a large amount of unlabeled text data can be quickly processed, thereby improving the overall efficiency of data annotation, especially in scenarios requiring rapid response. By combining deep learning and adversarial training algorithms, the accuracy and consistency of annotation can be improved, human errors can be reduced, and a small amount of labeled data and a large amount of unlabeled data can be used to automatically generate high-quality annotation results, thereby reducing the dependence on labeled data. By generating pseudo labels, unlabeled data can be fully utilized to extract valuable information from it, thereby improving the utilization efficiency of unlabeled data. The model is adjusted by a continuous learning algorithm, so that the model can better adapt to new data and changes, thereby improving the generalization ability of the model, and enabling the model to continuously learn from new data to maintain the timeliness and accuracy of the model. By using an experience replay pool and an elastic loss function, model training can be more effectively performed, thereby reducing the risk of overfitting and improving the stability and performance of the model.

[0116] The present invention is not limited to the above optional implementations, and anyone can derive other various forms of products under the enlightenment of the present invention. The above specific implementations should not be understood as limiting the scope of protection of the present invention. The scope of protection of the present invention should be based on the definition in the claims, and the description can be used to interpret the claims.

Claims

1. A text data annotation method based on weakly supervised learning, characterized in that: The steps include: Collect some historical text data, and use a weakly supervised learning algorithm to generate pseudo labels for some historical text data, so as to obtain some source domain data with real labels and some target domain data with pseudo labels; According to a number of source domain data and a number of target domain data, a text data annotation model is constructed using a deep learning and adversarial training algorithm, and the text data annotation model is adjusted using a continuous learning algorithm to obtain an adjusted text data annotation model, including the following steps: Use deep learning and adversarial training algorithms to build an initial text data annotation model; According to a number of source domain data and a number of target domain data, the initial text data annotation model is optimized and trained to obtain an optimized text data annotation model, and a number of historical text data annotation experiences are obtained; Use the experience replay mechanism of the continuous learning algorithm to build an experience replay pool, and use the elastic weight connection mechanism of the continuous learning algorithm to build an elastic loss function; According to the experience replay pool and the elastic loss function, the text data annotation model is adjusted to obtain an adjusted text data annotation model, and several historical text data annotation experiences are stored in the experience replay pool; The text data annotation model is constructed based on the DBN-DANN-Elman algorithm, and the text data annotation model includes a text feature extraction model constructed based on the DBN algorithm, a domain adversarial training module constructed based on the DANN algorithm, and a text data annotation module constructed based on the Elman algorithm, which are connected in sequence. The domain adversarial training module includes a label predictor and a domain classifier connected in sequence. The initial text data annotation model is optimized and trained according to a number of source domain data and a number of target domain data to obtain an optimized text data annotation model, and a number of historical text data annotation experiences are obtained, including the following steps: Pre-training an initial text feature extraction model and an initial text data annotation module of an initial text data annotation model according to a number of target domain data to obtain a pre-trained text feature extraction model and a pre-trained text data annotation module; According to a number of source domain data and a number of target domain data, a pre-trained text feature extraction model and a pre-trained text data annotation module are optimized and trained to obtain a final feature extraction model and a final text data annotation module, and obtain a number of historical text features; Integrate the third loss function of the label predictor of the initial text data annotation model and the fourth loss function of the domain classifier to obtain the second comprehensive loss function of the initial domain adversarial training module; According to the historical text features, a label predictor is used to generate corresponding historical label predictions, and a domain classifier is used to generate corresponding historical domain classification results; According to the historical label prediction and the historical domain classification results, the second comprehensive loss function of the initial domain adversarial training module is used to obtain the corresponding second historical loss value; The historical text features and historical model parameters of the text data annotation model generated in each optimization training are retained to obtain historical text data annotation experience; Traversing all source domain data and target domain data, performing the optimization training step of the above domain adversarial training module, and outputting the final domain adversarial training module if the second historical loss value is less than the second loss value threshold; Integrate the final text feature extraction model, the final domain adversarial training module, and the final text data annotation module to obtain an optimized text data annotation model and some historical text data annotation experience; Real-time text data is collected, and the real-time text data is annotated using the adjusted text data annotation model to obtain annotated real-time text data, and the adjusted text data annotation model is updated using a continuous learning algorithm to obtain an updated text data annotation model.

2. A text data labeling method based on weakly supervised learning according to claim 1, characterized in that: Collect some historical text data, and use weakly supervised learning algorithm to generate pseudo labels for some historical text data, so as to obtain some source domain data with real labels and some target domain data with pseudo labels, including the following steps: Collecting a number of historical text data, and preprocessing the number of historical text data to obtain a number of preprocessed historical text data; The preprocessed historical text data with real labels are divided into source domain data, and the preprocessed historical text data without real labels are divided into original data; Based on several source domain data with real labels, a pseudo-label generation model is constructed using a weakly supervised learning algorithm. The pseudo-label generation model is used to generate pseudo-labels for the original data, and a number of target domain data with pseudo-labels are obtained.

3. A text data labeling method based on weakly supervised learning according to claim 2, characterized in that: The pseudo-label generation model is constructed based on the LSTM-cGAN algorithm, and the pseudo-label generation model includes a sequence feature extraction module constructed based on the LSTM algorithm and a pseudo-label generation module constructed based on the cGAN algorithm, which are connected in sequence. The pseudo-label generation module includes a generator and a discriminator that are connected in sequence and are both constructed based on the RNN algorithm.

4. A text data labeling method based on weakly supervised learning according to claim 3, characterized in that: Based on several source domain data with real labels, a pseudo label generation model is constructed using a weakly supervised learning algorithm, including the following steps: Using a weakly supervised learning algorithm, an initial pseudo-label generation model is constructed; the initial pseudo-label generation model includes an initial sequence feature extraction module and an initial pseudo-label generation module; Integrate the first loss function of the generator of the initial pseudo-label generation module and the second loss function of the discriminator to obtain a first comprehensive loss function of the initial pseudo-label generation module; According to a number of source domain data with real labels, the initial sequence feature extraction module is optimized and trained to obtain a final sequence feature extraction module, and a number of historical sequence features are obtained; According to the historical sequence characteristics of the source domain data, the generator of the initial pseudo-label generation module is used to generate pseudo-labels for the source domain data; Use the discriminator of the initial pseudo-label generation module to perform label authenticity analysis based on the pseudo-labels and true labels of the source domain data to obtain the historical label authenticity analysis results; According to the authenticity analysis results of the pseudo-labels and historical labels, the first comprehensive loss function of the initial pseudo-label generation module is used to obtain the corresponding first historical loss value; Traverse all source domain data and perform the optimization training step of the pseudo label generation module. If the first historical loss value is less than the first loss value threshold, output the final pseudo label generation module. The final sequence feature extraction module and the final pseudo-label generation module are integrated to obtain the final pseudo-label generation model.

5. A text data labeling method based on weakly supervised learning according to claim 4, characterized in that: Using the pseudo-label generation model, pseudo-label generation is performed on the original data to obtain a number of target domain data with pseudo-labels, including the following steps: Use the sequence feature extraction module of the pseudo-label generation model to extract the second historical sequence features of the original data without setting the real label; Using a pseudo-label generation module of a pseudo-label generation model, generating pseudo-labels according to the second historical sequence characteristics, to obtain pseudo-labels for the original data; Traverse all the original data without real labels, perform the above pseudo-label generation step, and obtain a number of target domain data with pseudo labels.

6. The text data labeling method based on weakly supervised learning according to claim 1, characterized in that: Collecting real-time text data, using the adjusted text data annotation model to annotate the real-time text data to obtain annotated real-time text data, and using a continuous learning algorithm to update the adjusted text data annotation model to obtain an updated text data annotation model, including the following steps: Collecting real-time text data, and preprocessing the real-time text data to obtain preprocessed real-time text data; Extracting real-time text features of the preprocessed real-time text data using a text feature extraction model of the adjusted text data annotation model; Using the text data annotation module of the adjusted text data annotation model, annotating the real-time text data according to the real-time text features to obtain annotated real-time text data; The real-time text features and the real-time model parameters of the adjusted text data annotation model are retained to obtain real-time text data annotation experience; Randomly extract a number of historical text data annotation experiences from the experience replay pool, and mix them with the real-time text data annotation experiences to obtain a number of mixed text data annotation experiences; Based on the mixed text data annotation experience, the adjusted text data annotation model is continuously trained, and the elastic loss function is used to obtain the elastic loss value of the continuous training; Traverse all mixed text data annotation experiences and perform the above continuous training steps. If the elastic loss value is less than the third loss value threshold, output the updated text data annotation model.

7. A text data annotation system based on weakly supervised learning, used to implement the text data annotation method according to any one of claims 1 to 6, characterized in that: The system comprises a weakly supervised learning unit, a model building unit and a labeling and updating unit which are connected in sequence.

Citation Information

Patent Citations

  • Entity recognition model generation method and device and computer readable storage medium

    CN112420205A

  • Enterprise multi-type data labeling method and system based on feature engineering

    CN119128612A

  • SKU intelligent classification and label generation method fusing visual pre-training model

    CN119202826A