Abnormal tag identification method and system and related equipment
By combining the abnormal label recognition method with multiple abnormal recognition models and judgment rules, text labels can be automatically identified and corrected, solving the problems of low efficiency and poor accuracy of manual inspection in the existing technology, and realizing efficient and low-cost text label quality inspection and error correction.
Patent Information
- Application Number
- CN202410288001.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-09-16
AI Technical Summary
The quality inspection of text labels in the existing technology relies on manual methods, which leads to low efficiency, poor accuracy, high cost, and requires a lot of human resources.
Adopting the abnormal label recognition method, the text labels are automatically identified and corrected through multiple abnormal recognition models and judgment rules, and the data-driven and knowledge-driven models are combined to realize automated label quality inspection and error correction.
It significantly improves the efficiency and accuracy of text label quality inspection, reduces labor costs, supports quality inspection and classification of single-text and multi-text labels, and ensures the reliability and interpretability of corrected labels.
Smart Images

Figure CN120653942A_ABST
Abstract
Claims
1. A method for identifying abnormal labels, characterized in that: The method comprises: Get the text and the actual label of said text; In the case where the actual label is an abnormal label, one of the labels in the predicted label set and the inferred label set is used as the corrected label for the text, wherein the predicted label set includes one or more predicted labels, one of which is determined based on the text using a first processing method, and the inferred label set includes one or more inferred labels, one of which is determined based on the text using a second processing method, and the first processing method and the second processing method are different.
2. The method according to claim 1, characterized in that The first processing mode is a target anomaly recognition model, and the target anomaly recognition model belongs to a plurality of anomaly recognition models. Before using one of the labels in the predicted label set and the inferred label set as the corrected label of the text, the method further includes: Inputting the text into the multiple anomaly recognition models respectively to obtain the prediction label set including multiple prediction labels; Comparing the multiple predicted labels with the actual labels to determine whether they are the same, and obtaining an inconsistency rate; When the inconsistency rate is greater than or equal to a first threshold, the actual label is determined to be an abnormal label.
3. The method according to claim 2, characterized in that The method further comprises: The target anomaly recognition model is updated using the text and the correction label.
4. The method according to any one of claims 1 to 3, characterized in that The second processing manner is a first determination rule, which belongs to a plurality of determination rules. Before using one of the predicted label set and the inferred label set as the corrected label for the text, the method further includes: The text is matched with the multiple decision rules respectively to obtain the inference label set including multiple inference labels.
5. The method according to claim 4, characterized in that The step of using one of the predicted label set and the inferred label set as a corrected label for the text includes: If the inference label set is credible, selecting the inference label set from the prediction label set and the inference label set; The inference label with the largest proportion in the inference label set is selected as the correction label of the text.
6. The method according to claim 5, characterized in that When the confidence of the inference label set is greater than or equal to a second threshold, the inference label set is credible, wherein the confidence of the inference label set is the ratio between the number of labels in the intersection between the inference label set and the prediction label set and the number of the multiple determination rules.
7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: In the case where the actual tag is a normal tag, determining a second determination rule for text matching; In a case where the second determination rule is not credible, the second determination rule is updated using the text and the actual label.
8. The method according to claim 7, characterized in that When the confidence of the second determination rule is less than or equal to a third threshold, the second determination rule is unreliable, wherein the confidence of the second determination rule is the proportion of texts in the texts matching the second determination rule whose actual labels are also the semantic labels of the second determination rule.
9. An abnormal tag identification system, characterized in that: Including acquisition unit and correction unit, The acquiring unit is used to acquire the text and the actual label of the text; The correction unit is used to use one of the labels in the predicted label set and the inferred label set as the correction label of the text when the actual label is an abnormal label, wherein the predicted label set includes one or more predicted labels, one of which is determined based on the text using a first processing method, and the inferred label set includes one or more inferred labels, one of which is determined based on the text using a second processing method, and the first processing method and the second processing method are different.
10. The system according to claim 9, characterized in that The first processing mode is a target anomaly recognition model, the target anomaly recognition model belongs to multiple anomaly recognition models, and the system further includes an identification unit, The recognition unit is configured to input the text into the multiple anomaly recognition models respectively before using one of the predicted label set and the inferred label set as a corrected label for the text, to obtain the predicted label set including the multiple predicted labels; The recognition unit is further configured to compare the plurality of predicted labels with the actual labels to determine whether they are the same, and obtain an inconsistency rate; The identification unit is further configured to determine that the actual label is an abnormal label when the inconsistency rate is greater than or equal to a first threshold.
11. The system according to claim 10, wherein: The system further comprises a training unit, The training unit is used to update the target anomaly recognition model using the text and the correction label.
12. The system according to any one of claims 9 to 11, characterized in that: The second processing method is a first determination rule, the first determination rule belongs to a plurality of determination rules, and the system further includes an inference unit, The inference unit is configured to match the text with the multiple decision rules respectively before using one of the predicted label set and the inference label set as the correction label for the text, to obtain the inference label set including multiple inference labels.
13. The system according to claim 12, wherein: The correction unit is specifically configured to select the inference label set from the prediction label set and the inference label set when the inference label set is credible; and select the inference label with the largest proportion in the inference label set as the correction label for the text.
14. The system according to claim 13, wherein: When the confidence of the inference label set is greater than or equal to a second threshold, the inference label set is credible, wherein the confidence of the inference label set is the ratio between the number of labels in the intersection between the inference label set and the prediction label set and the number of the multiple determination rules.
15. The system according to any one of claims 9 to 14, characterized in that: The system further comprises a generating unit, The inference unit is further configured to determine a second determination rule for text matching when the actual label is a normal label; The generating unit is configured to update the second determination rule using the text and the actual label when the second determination rule is not credible.
16. The system according to claim 15, wherein: When the confidence of the second determination rule is less than or equal to a third threshold, the second determination rule is unreliable, wherein the confidence of the second determination rule is the proportion of texts in the texts matching the second determination rule whose actual labels are also the semantic labels of the second determination rule.
17. A computing device, characterized in that The method comprises a processor and a memory, wherein the memory is used to store instructions, and the processor is used to execute the instructions. When the processor executes the instructions, the method according to any one of claims 1 to 8 is implemented.
18. A computing device cluster, characterized in that: comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 8.
19. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device, the computing device is caused to perform the method according to any one of claims 1 to 8.
20. A computer-readable storage medium, characterized in that The method comprises computer program instructions, and when the computer program instructions are executed by a computing device, the computing device performs the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
A fusion reasoning system and method for intelligent tags of news programs
CN109635171A
Text data processing method, neural network training method and related equipment
CN113807089A
Text data processing method and device, computer equipment and storage medium
CN113807096A
Entity recognition model training method and device, equipment, storage medium and product
CN116956915A