Semantic tag generation method based on semi-supervised learning
By using semi-supervised learning methods in semantic tag generation, data is collected and processed to generate pseudo-labels, and improving the accuracy of label generation through iterative training, the problem of insufficient label generation accuracy in the prior art is solved, and higher semantic tag generation accuracy and recall rate are achieved.
Patent Information
- Application Number
- CN202510157833.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
AI Technical Summary
Existing semi-supervised learning faces the influence of labeled data and unlabeled data in semantic label generation, resulting in insufficient accuracy of label generation.
By collecting supervised and unlabeled data, feature extraction and preliminary label generation are performed, supervised learning models are trained and generated pseudo-labels, and iterative training is carried out round by round to improve the accuracy of label generation.
The accuracy and recall of semantic label generation are improved, and the generalization ability and performance of the model are improved through iterative update of pseudo-labels and optimization of the model.
Smart Images

Figure CN120086593A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of semantic label generation, and specifically to a semantic label generation method based on semi-supervised learning. Background Art
[0002] With the development of big data and deep learning technologies, semantic label generation has been increasingly widely applied in fields such as natural language processing, image recognition, and video analysis. Semantic labels can help computers understand the meaning of data and provide valuable information for subsequent tasks. Traditional semantic label generation methods rely on a large amount of manually labeled data, but the cost of manual labeling is high, time-consuming, and it is difficult to meet the needs of large-scale datasets. To solve this problem, semi-supervised learning methods have emerged.
[0003] Semi-supervised learning is a learning paradigm between supervised learning and unsupervised learning, which uses a large amount of unlabeled data and a small amount of labeled data to jointly train a model. It can, in the case of scarce labeled data, improve the generalization ability and accuracy of the model by mining the potential structure in the unlabeled data. Semantic label generation methods based on semi-supervised learning mainly achieve efficient label generation by designing algorithms that can effectively utilize unlabeled data, such as generative adversarial networks, autoencoders, and graph neural networks.
[0004] However, the application of semi-supervised learning in semantic label generation still faces some challenges, such as the influence of labeled data and unlabeled data. Therefore, it is necessary to design a semantic label generation method based on semi-supervised learning that can improve the accuracy of label generation. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] Aiming at the deficiencies of the prior art, the present invention provides a semantic label generation method based on semi-supervised learning, which has the advantage of improving the accuracy of label generation and solves the problems in the above background art.
[0007] (2) Technical Solutions
[0008] To achieve the above object of improving the accuracy of label generation, the present invention provides the following technical solutions: A semantic label generation method based on semi-supervised learning, including the following steps:
[0009] S1: Collect supervised data with labels and unlabeled data without labels, perform denoising, removing irrelevant information, and missing value processing on the original data, and extract features for semantic label generation from the original data;
[0010] Preferably, S1 further includes collecting labeled data sets in known fields or tasks, removing features irrelevant to the target task in feature selection, processing missing data by interpolation, mean filling or deleting missing value samples, feature extraction including text data feature extraction, image data feature extraction and audio data feature extraction, selecting features related to label generation through algorithms, training a preliminary classification model based on existing supervised data, predicting unlabeled data, and generating preliminary labels.
[0011] S2: Use the labeled data to train a supervised learning model, use the trained model to predict the unlabeled data and generate preliminary labels;
[0012] Preferably, S2 further includes that for each unlabeled sample, the model will output a predicted probability of a category, set a confidence threshold, select high-confidence predictions as pseudo-labels, and low-confidence predictions need to be discarded or re-evaluated. In the generated pseudo-labels, the label consistency with the existing label data is checked, and the similarity between the labeled data and the pseudo-labels is calculated to evaluate the label quality. The generated high-confidence pseudo-labels are added to the training set to form a new mixed training set, and the model is retrained using this new training set. The training process is iterated repeatedly, each time the model is trained using the labeled data and pseudo-labels, and new pseudo-labels are generated and added to the training set again for updating.
[0013] S3: The preliminary labels generated by the model are added to the unlabeled data as pseudo labels, and the model is trained jointly using the labeled data and the unlabeled data with pseudo labels.
[0014] Preferably, S3 further includes combining the original labeled data with the unlabeled data with pseudo labels to form an extended training set, using the mixed data for training, and using the trained model to predict the unlabeled data again to generate new pseudo labels.
[0015] S4: In each iteration, the labeled data and the unlabeled data with high-confidence pseudo-labels are trained together, and the high-confidence labels are used for label update, and the low-confidence pseudo-labels are removed;
[0016] Preferably, S4 further includes setting a confidence threshold, treating all pseudo-labels with confidence greater than the threshold as high-confidence labels, considering these labels to be credible, evaluating the confidence of each pseudo-label output by the model, dynamically adjusting the threshold according to the performance of the current model, and in each round of iteration, using high-confidence pseudo-labels to update the original labels, merging the unlabeled data and labeled data of these high-confidence labels together to form a new training set, and conducting the next round of training. In each round of training, the labeled data and the screened high-confidence pseudo-label data are merged for joint training.
[0017] S5: Use the semi - supervised learning model completed by training to predict all data, generate the final semantic labels, evaluate the accuracy and recall rate of the model according to the comparison between the true labels and the generated labels, and optimize the parameters of the model.
[0018] Preferably, S5 further includes using the trained semi - supervised learning model to predict all data, generate the final semantic labels, compare the labels generated by the model with the true labels, calculate the accuracy and recall rate of the model on the test set or validation set, optimize the model performance by adjusting the hyperparameters according to the model evaluation results, retrain the model according to the evaluation results and the settings after hyperparameter optimization, train with all data, and continue to adjust the model parameters according to the feedback.
[0019] (III) Beneficial Effects
[0020] Compared with the prior art, the present invention provides a semantic label generation method based on semi - supervised learning, having the following beneficial effects:
[0021] In the present invention, the preliminary labels generated by the model are used as pseudo - labels and added to the unlabeled data, and the model is jointly trained using the labeled data and the unlabeled data with pseudo - labels; in each iteration, the labeled data and the unlabeled data with high - confidence pseudo - labels are trained together, the high - confidence labels are used for label update, and the low - confidence pseudo - labels are removed; the semi - supervised learning model completed by training is used to predict all data, generate the final semantic labels, evaluate the accuracy and recall rate of the model according to the comparison between the true labels and the generated labels, and optimize the parameters of the model. It has the advantage of improving the accuracy of label generation. Brief Description of the Drawings
[0022] Figure 1 It is a schematic diagram of the method of the present invention. Detailed Embodiments
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0024] Embodiment 1
[0025] The present invention provides a technical solution: a semantic label generation method based on semi - supervised learning, including the following steps:
[0026] S1: Collect supervised data with labels and unlabeled data without labels, denoise the original data, remove irrelevant information and missing values, and extract features generated from semantic labels from the original data.
[0027] Collect labeled datasets from known domains or tasks. This can be obtained from public datasets, company internal data, manual annotation, etc. This part of the data has no labels and usually comes from a large amount of user-generated content, automatically collected data, social media, log data, etc. Statistically analyze the distribution of each category in the labeled data to understand the skewness of the data. For unlabeled data, estimate the data quality through pre-statistical analysis or expert evaluation.
[0028] Remove noise through data cleaning techniques. For text data, reduce noise by removing stop words, punctuation marks, low-frequency words, etc. For image data, apply image filters, remove backgrounds, etc. Use the noise ratio to measure the degree of noise in the data, such as the number of noise points in an image, the proportion of irrelevant words in text, etc. During feature selection or dimensionality reduction, remove features or information irrelevant to the target task, and use feature importance metrics to evaluate the relevance of features. Handle missing data by interpolation, mean filling, or deleting missing value samples. The proportion of missing values can be used to evaluate the impact of missing data.
[0029] Feature extraction for text data:
[0030] Bag of Words model: Statistically analyze the occurrence frequency of each word in the text and convert it into a vector.
[0031] TF-IDF: A feature extraction method based on term frequency and inverse document frequency, used to measure the importance of each word to semantics.
[0032] Word Embedding: Map words to a continuous vector space to capture semantic relationships between words.
[0033] Use the numerical values of TF-IDF to calculate the sparsity of the feature dimension and the distribution of the feature vector, such as the Euclidean distance and cosine similarity of the vector.
[0034] Feature extraction for image data:
[0035] Convolutional Neural Network: Use the CNN model to extract visual features of images, including edges, textures, shapes, etc.
[0036] Image descriptors: These methods describe image content by extracting local features.
[0037] Evaluate the distribution and diversity of features by calculating statistical metrics such as the mean, variance, and maximum value of the feature vector.
[0038] Audio data feature extraction:
[0039] MFCC: Extract frequency and time-domain features from audio signals to describe the speech or pitch of audio.
[0040] Use statistics such as spectral entropy and spectrum entropy of the signal to quantify the complexity or information content of audio features.
[0041] Feature vector dimension: Analyze the dimension of the vector after feature extraction to understand the degree of information compression.
[0042] Feature correlation measurement: Use Pearson correlation coefficient or mutual information to quantify the correlation between different features.
[0043] Select the features most relevant to label generation through algorithms such as recursive feature elimination, L1 regularization, PCA, etc. Evaluate the effect of the selected features through cross-validation, such as classification accuracy, recall rate, etc., and calculate the correlation between features and labels, such as using statistics such as Pearson correlation coefficient and Spearman rank correlation coefficient.
[0044] S2: Train a supervised learning model using labeled data, and use the trained model to predict unlabeled data to generate preliminary labels;
[0045] For each unlabeled sample, the model will output a prediction probability for a class. By setting a confidence threshold, select high confidence, for example, a prediction with a probability greater than 0.9 as a pseudo-label. Low-confidence predictions, for example, those with a probability less than 0.6, usually need to be discarded or re-evaluated. Calculate the proportion of high-confidence labels and their contribution to the training set.
[0046] Add the generated high-confidence pseudo-labels to the training set to form a new mixed training set, which is labeled data + high-confidence pseudo-labels. Retrain the model using this new training set, calculate the difference in model performance between the original training set and the training set with newly added pseudo-labels. If the performance improves, it indicates that the pseudo-labels have a positive effect on training. Iteratively repeat the training process. Each time, train the model using labeled data and pseudo-labels, generate new pseudo-labels, and add them to the training set for update again. Monitor the performance changes after each round of training, evaluate the convergence of the model. If the improvement in the model's performance tends to be stable, it means convergence.
[0047] Through certain manual inspections or sample samplings, conduct quality verification on the generated pseudo-labels. Check whether some pseudo-labels conform to the actual situation through random sampling, or use expert annotations to compare pseudo-labels with true labels, and calculate the consistency between manual verification and pseudo-labels. For example, measure the matching degree between manual labels and pseudo-labels through the consistency rate. According to the verification results, correct the incorrect pseudo-labels. For suspected incorrect labels, use other methods for correction, such as adding additional constraints, combining other data sources, etc. After correction, re-evaluate the performance changes of the model to ensure that the corrected pseudo-labels can improve the model accuracy.
[0048] S3: Use the preliminary labels generated by the model as pseudo-labels, add them to the unlabeled data, and jointly train the model using the labeled data and the unlabeled data with pseudo-labels;
[0049] Combine the original labeled data with the unlabeled data with pseudo-labels to form an extended training set. The training set consists of labeled data and pseudo-labeled data, so that the model can be trained on more data. Calculate the label ratios in the extended training set, such as the ratio of labeled data to the total data and the ratio of pseudo-labeled data, to evaluate the contribution of pseudo-labels to the training data. Use the mixed data for training. By merging the labeled data and the pseudo-labeled data, the model can learn from more samples. During the training process, the model needs to be able to handle the potential noise brought by pseudo-labels, monitor the performance changes of the model after each round of training, such as accuracy, loss value, F1 score, etc., especially the performance fluctuations after the introduction of pseudo-labels, to evaluate the effectiveness of pseudo-labels.
[0050] Evaluate the quality of each pseudo-label by checking its confidence. Pseudo-labels with low confidence may introduce noise and affect the accuracy of model training. Set a threshold and only use pseudo-labels with a confidence higher than a certain standard. Evaluate the overall quality of pseudo-labels by calculating statistical quantities such as the average confidence and standard deviation of pseudo-labels. Conduct a consistency check with the labeled data, compare the similarity between pseudo-labels and labeled data, and ensure that the generated pseudo-labels have a high consistency with the target labels. Use similarity metrics to calculate the consistency between pseudo-labels and labeled data. Correct or discard low-confidence or inconsistent pseudo-labels. Incorrect pseudo-labels can be adjusted through manual verification or other algorithms. For the corrected pseudo-labels, recalculate the degree of performance improvement of the model to evaluate the effect of correction.
[0051] Use the trained model to predict the unlabeled data again and generate new pseudo-labels. Then, add these new pseudo-labels together with the original pseudo-labels to the training set for the next round of training. Through multiple iterations of optimization, the model gradually improves the generation of pseudo-labels, making the quality of the pseudo-labels higher and the model performance continuously improved. By monitoring the loss value, accuracy, etc. of each round of training, evaluate the model convergence during the self-training process.
[0052] S4: In each iteration, train the labeled data and the unlabeled data with high-confidence pseudo-labels together, use the high-confidence labels for label updating, and eliminate the low-confidence pseudo-labels;
[0053] When using the trained model to predict the unlabeled data, the output is not only the class label but also the corresponding confidence. Set a confidence threshold, and regard all pseudo-labels with a confidence greater than this threshold as high-confidence labels, believing that these labels are credible. Evaluate the confidence of each pseudo-label output by the model. For example, calculate the average confidence and standard deviation of all pseudo-labels to ensure that the high-confidence labels do have relatively high reliability. Dynamically adjust the threshold according to the performance of the current model. As the model is gradually improved in the iteration, gradually increase the screening criteria for high-confidence pseudo-labels. After each round of training, check the impact of different confidence thresholds on the final model performance. Obtain the optimal confidence threshold through experiments to make the model stable and effective.
[0054] In each iteration, update the original labels with the high-confidence pseudo-labels. That is, use the labeled data and the unlabeled data with high-confidence pseudo-labels together for training, and update the model weights through backpropagation. Compare the change in the loss value before and after adding the high-confidence labels during the training process to check whether the addition of the high-confidence labels effectively improves the training effect. Merge the unlabeled data with these high-confidence labels and the labeled data together to form a new training set and conduct the next round of training. Statistically analyze the frequency and proportion of the training set update, and monitor the impact of the increase in the data set on the model training.
[0055] For the pseudo-labels below the set confidence threshold, they need to be eliminated to prevent low-quality pseudo-labels from introducing noise and affecting the learning process of the model. Calculate the proportion of low-confidence labels and monitor their negative impact on the model performance during the training process. Low-confidence pseudo-labels can not only be eliminated after each round of training, but also be dynamically evaluated during the training process, and high-quality labels are selected for retention and labels with greater noise are eliminated according to the quality of label updating.
[0056] S5: Use the trained semi-supervised learning model to predict all data and generate the final semantic labels. According to the comparison between the true labels and the generated labels, evaluate the accuracy and recall rate of the model and optimize the model parameters.
[0057] Use the trained semi - supervised learning model to predict all data, including labeled data and unlabeled data, and generate the final semantic labels. These labels include the results learned from the training data and inferred through pseudo - labels. Calculate the prediction confidence of each data point, and generate labels and confidence scores to ensure that the final labels have a high level of credibility.
[0058] Compare the labels generated by the model with the true labels, calculate the accuracy and recall rate of the model on the test set or validation set, measure the proportion of the labels predicted by the model that are exactly the same as the true labels, and measure the proportion of the model's ability to correctly identify the positive class labels. Especially for the problem of data imbalance, recall rate is particularly important. Calculate the accuracy, recall rate, and F1 - score for each category to ensure a quantitative evaluation of the model's performance for each category.
[0059] View the prediction situation of the model in each category through the confusion matrix, understand which categories are misclassified and the types of misclassification. Extract the precision, recall rate, and F1 - score for each category from the confusion matrix to evaluate the performance of different categories. Plot the ROC curve and calculate the AUC to evaluate the classification ability of the model. Especially for binary classification tasks, AUC can measure the overall classification performance of the model. Calculate the ROC curve for each category and obtain the AUC value to evaluate the classification ability of the model.
[0060] According to the model evaluation results, optimize the model performance by adjusting hyperparameters. Commonly used hyperparameter optimization methods include grid search and random search. Train different combinations of hyperparameters, evaluate their performance on the validation set, and select the best hyperparameter configuration. According to the evaluation results and the settings after hyperparameter optimization, retrain the model, use all the data for training, and continue to adjust the model parameters according to the feedback. After each retraining, calculate the changes in the training error and validation error of the model to ensure that the error gradually decreases during the training process. Deploy the trained model to the actual application for real - time monitoring to ensure that the model can run stably in the actual scenario and process new data in a timely manner. Monitor the real - time performance of the model in the production environment, collect real - time prediction data, and regularly evaluate the prediction accuracy of the model.
[0061] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non - exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0062] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A semantic label generation method based on semi-supervised learning, characterized in that: The following steps are involved: S1: Collect labeled supervised data and unlabeled data, denoise the original data, remove irrelevant information and missing values, and extract features for semantic label generation from the original data; S2: Use the labeled data to train a supervised learning model, use the trained model to predict the unlabeled data and generate preliminary labels; S3: The preliminary labels generated by the model are added to the unlabeled data as pseudo labels, and the model is trained jointly using the labeled data and the unlabeled data with pseudo labels. S4: In each iteration, the labeled data and the unlabeled data with high-confidence pseudo-labels are trained together, and the high-confidence labels are used for label update, and the low-confidence pseudo-labels are removed; S5: Use the trained semi-supervised learning model to predict all data and generate final semantic labels. Compare the real labels with the generated labels to evaluate the accuracy and recall of the model and optimize the model parameters.
2. The method for generating semantic tags based on semi-supervised learning according to claim 1, characterized in that: The S1 further includes collecting labeled data sets from known fields or tasks, removing features irrelevant to the target task in feature selection, processing missing data by interpolation, mean filling or deleting missing value samples, feature extraction including text data feature extraction, image data feature extraction and audio data feature extraction, selecting features related to label generation through algorithms, training a preliminary classification model based on existing supervised data, predicting unlabeled data, and generating preliminary labels.
3. The method for generating semantic tags based on semi-supervised learning according to claim 1, characterized in that: The S2 further includes that for each unlabeled sample, the model will output a predicted probability of a category, set a confidence threshold, select high-confidence predictions as pseudo-labels, and low-confidence predictions need to be discarded or re-evaluated. In the generated pseudo-labels, the label consistency with the existing label data is checked, and the similarity between the labeled data and the pseudo-labels is calculated to evaluate the label quality. The generated high-confidence pseudo-labels are added to the training set to form a new mixed training set, and the model is retrained using this new training set. The training process is iterated repeatedly, each time the model is trained using the labeled data and pseudo-labels, and new pseudo-labels are generated and added to the training set again for updating.
4. The method for generating semantic tags based on semi-supervised learning according to claim 1, characterized in that: The S3 further includes combining the original labeled data with the unlabeled data with pseudo labels to form an extended training set, using the mixed data for training, and using the trained model to predict the unlabeled data again to generate new pseudo labels.
5. The method for generating semantic tags based on semi-supervised learning according to claim 1, characterized in that: The S4 further includes setting a confidence threshold, treating all pseudo-labels with confidence greater than the threshold as high-confidence labels, considering these labels to be credible, evaluating the confidence of each pseudo-label output by the model, dynamically adjusting the threshold according to the performance of the current model, and in each round of iteration, using high-confidence pseudo-labels to update the original labels, merging the unlabeled data and labeled data of these high-confidence labels together to form a new training set, and conducting the next round of training. In each round of training, the labeled data and the screened high-confidence pseudo-label data are merged for joint training.
6. The method for generating semantic tags based on semi-supervised learning according to claim 1, characterized in that: The S5 further includes using the trained semi-supervised learning model to predict all data, generate final semantic labels, compare the labels generated by the model with the true labels, calculate the accuracy and recall of the model on the test set or validation set, and optimize the model performance by adjusting the hyperparameters based on the model evaluation results. According to the evaluation results and the settings after hyperparameter optimization, retrain the model, use the full amount of data for training, and continue to adjust the model parameters based on feedback.
Citation Information
Cited By
News content core-oriented labeling method, equipment and medium
CN120470127A
Road structure health monitoring data key information real-time extraction method based on edge calculation
CN120724150A
Image abstract generation method and system based on adaptive pseudo supervision
CN121545157A
Data labeling method, device and system based on incremental learning and confidence learning
CN122347716A
Data labeling method, device and system based on incremental learning and confidence learning
CN122347716B