Automatic data labeling method based on multi-modal fusion and iterative optimization
Through the data automatic labeling method of multimodal fusion and iterative optimization, the problems of instability and inefficiency in multimodal data labeling are solved, and efficient and low-cost data labeling and model generalization are achieved, with strong adaptability.
Patent Information
- Application Number
- CN202510486851.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-29
AI Technical Summary
In the process of labeling multimodal, multi-time phase, and multi-organization level data, the prior art has problems such as unstable labeling quality, low efficiency, high expert labeling cost and poor adaptability.
The data automatic labeling method of multimodal fusion and iterative optimization is adopted. Through the closed-loop process of "machine labeling-screening exceptions-manual verification-model retraining", combined with multimodal enhancement, standardized processing and cross-domain knowledge fusion, a self-evolution system of "machine main mark + manual retraining" is formed.
It effectively improves the labeling quality and efficiency of multimodal data, enhances the generalization ability of the model, has good adaptability and promotion, and reduces costs.
Smart Images

Figure CN120387135A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and data processing, and relates to a data automatic annotation method based on multimodal fusion and iterative optimization. Background Art
[0002] In the current rapid development of artificial intelligence, data annotation, as a key link, plays a decisive role in the training effect of models. Especially when dealing with multi-modal, multi-temporal, and multi-organizational level data, the importance of annotation work is even more prominent. The traditional annotation mode has the problem of unstable annotation quality when dealing with complex data. The existing artificial intelligence and data processing fields face two main challenges, resulting in unstable or inefficient annotation quality: (i) improving the coordination of annotation efficiency and quality; (ii) enhancing the adaptability and accuracy of intelligent annotation systems. Summary of the Invention
[0003] To solve these problems, the present invention provides a data automatic annotation method based on multimodal fusion and iterative optimization, which combines multimodal enhancement, standardization processing, and cross-domain knowledge fusion, and forms a closed-loop process of "machine annotation - screening anomalies - manual verification - model retraining" through an iterative mechanism, breaking through the limitations of the existing expert-dependent annotation with low efficiency, high cost, and difficult quality control, being applicable to multi-modal, multi-temporal, and multi-organizational level data, effectively improving data quality, annotation efficiency, and model generalization ability, and having good adaptability and popularization.
[0004] The technical solution of the present invention is as follows:
[0005] A data automatic annotation method based on multimodal fusion and iterative optimization, which combines multimodal enhancement, standardization processing, and cross-domain knowledge fusion, and forms a closed-loop process of "machine annotation + manual refinement" that can continuously learn and self-evolve through an iterative mechanism of "machine annotation - screening anomalies - manual verification - model retraining".
[0006] This method breaks through the limitations of the existing expert-dependent annotation with low efficiency, high cost, and difficult quality control, is applicable to multi-modal, multi-temporal, and multi-organizational level data, effectively improves data quality, annotation efficiency, and model generalization ability, and has good adaptability and popularization.
[0007] Specifically, it includes the following steps:
[0008] (1) Obtain the original multi-modal data Including forms such as text, image, audio, video, sensors, etc.;
[0009] (2) Perform format conversion and noise removal processing on each type of data. Fill in missing fields (such as unentered fields in text, missing frames in images, and NaN values in time series), detect and remove outliers, and merge and unify duplicate images, similar texts, and repeatedly recorded audio segments, etc.; Process heterogeneous data using unified encoding and size standards (such as text UTF-8 encoding and image size standardized to 224×224);
[0010] (3) Perform data augmentation, that is, artificially expand the number and diversity of training samples while maintaining the original semantics, improve the generalization ability of the model, and alleviate problems such as data scarcity and class imbalance. Enhance by applying techniques such as rotation, mirroring, and blurring to images, and using methods such as synonym replacement and word order scrambling for texts. The formula is as follows:
[0011] D * =f aug (D)
[0012] where D * is the augmented data sample, and f aug is the augmentation function;
[0013] (4) Establish an expert-annotator team. The expert group consists of medical imaging experts, manufacturing industry engineers, linguists, etc. with many years of experience in relevant fields. The annotator team can be composed of professionally trained students, field assistant engineers, or part-time outsourced personnel; The experts formulate annotation rules and supervise the annotators to execute the annotation tasks, and the two complete the high-quality annotation of multimodal samples through cooperation;
[0014] (5) Set annotation dimensions for each modality (such as target type and bounding box in images; entity categories and relationships in texts; intonation and emotion in audio, etc.), and formulate clear and structured feature expression methods and annotation dimensions, whose role is equivalent to "formulating a dictionary" and "defining fields" for each data modality, so that subsequent manual or automatic systems know "which dimensions to pay attention to" and "how to express these dimensions";
[0015] (6) After the annotators complete the first round of annotation, it is reviewed and corrected by the experts. Because the annotators lack domain knowledge, it is easy to produce mislabeling, missing labeling, and the complexity of multimodal data itself is high. It is necessary to unify the "cognitive standard", that is, form a golden standard sample set to achieve the unity of annotation standards and the controllability of data quality. The formula is as follows:
[0016]
[0017] where x i is the multimodal input (such as images, texts, audio), and y iStructured tags annotated by experts (such as categories, borders, attributes), L expert is a high-quality annotation set;
[0018] (7) Calculate the annotation consistency score and eliminate low-consistency data. After multiple annotators complete the annotation of the same data sample, measure the consistency of their annotation results through statistical methods, automatically screen out data with serious disagreements, and only retain data with high consistency into the training set or the gold standard set. If the consistency score is greater than or equal to the threshold, the sample is determined to be valid; otherwise, the sample is eliminated.
[0019] (8) Construct a multi-modal annotation model, an artificial intelligence model that can simultaneously receive and understand data input from multiple modalities (such as images, text, audio, video, sensors) and automatically output the corresponding annotation results. It can learn the annotation rules based on the existing expert-annotated samples and achieve fast, automatic, and scalable annotation of new data. Use a fusion neural network to model data from different modalities. The formula is as follows:
[0020] y = arg max F θ (x text , x image , x audio )
[0021] where F θ represents the parameters of the multi-modal fusion network, x text is the text data, x image is the image data, x audio is the audio data;
[0022] (9) Train the model with the expert-annotated data L expert to conduct the first complete training of the multi-modal annotation model, enabling the model to have preliminary automatic annotation capabilities and laying a foundation for subsequent large-scale inference and iterative training. The loss function is the cross-entropy loss. The formula is as follows:
[0023]
[0024] where F θ (x i ) is the predicted output of the model for x i , log(F θ (x i )) is to take the logarithm of the model's predicted value to form the cross-entropy term with the true label;
[0025] (10) For the multi-modal annotation model after preliminary training, perform inference and prediction on the original samples that have not been annotated, automatically generate the corresponding annotation labels to replace the process of manually generating preliminary labels. Perform inference and prediction on the unlabeled sample x ∈ U to obtain the predicted label y'.
[0026] (11) Calculate whether manual review is required using entropy or model confidence to determine whether the model is "confident". If the class distribution predicted by the model is unclear (high entropy), the sample is sent to the review pool to ensure the credibility of the output of the automatic annotation system. The formula is as follows:
[0027]
[0028] where \(H(y')\) is the entropy of the predicted label and \(P(c)\) is the predicted probability;
[0029] (12) In some cases, the automatic annotation system "seems to have marked" but is actually very uncertain. Set a threshold \(\delta\). If \(H(y') > \delta\), the sample is marked as uncertain, and the expert team reviews and optimizes the highly uncertain samples to obtain a correction set, which can avoid directly using "low-confidence annotations" for training and prevent mislearning and model degradation. The formula is as follows:
[0030]
[0031] where \(L\) check is the correction set, \(x\) j is the original input of the highly uncertain sample, is the expert-corrected label;
[0032] (13) The system continuously feeds back the corrected samples reviewed by experts, etc. as new training data for incremental training or fine-tuning of the existing model, making the model more accurate when facing similar scenarios in the future and gradually improving the quality of automatic annotation. Incorporate \(L\) check into the training set and continue to fine-tune the model parameters. If the model performs poorly on some specific samples, the expert correction results can be used as "feedback samples". The formula is as follows:
[0033] \(\Gamma\) total =\(\Gamma\) cls +\(\lambda_1\Gamma\) bbox +\(\lambda_2\Gamma\) attr
[0034]
[0035] where \(\Gamma\) total is the joint loss, \(\Gamma\) cls is the cross-entropy loss, \(\Gamma\) bbox is the bounding box regression loss, \(\Gamma\) attr is the attribute binary classification loss, \(\lambda_1\) and \(\lambda_2\) are weight coefficients, \(\theta\) is the learning rate, is the gradient operator;
[0036] (14) After the model accuracy reaches the upper limit standard, the system enters the "machine main label - manual fine-tuning" stage, continuously expands the training samples and knowledge base, realizes self-evolution, and constructs a data annotation system that continuously learns, dynamically optimizes, and becomes more accurate with use.
[0037] Advantages of the present invention:
[0038] The present invention combines multi-modal enhancement, standardized processing, and cross-domain knowledge fusion, and forms a closed-loop process of "machine annotation - screening anomalies - manual verification - model retraining" through an iterative mechanism, resulting in a "machine main label + manual fine-tuning" process that can continuously learn and self-evolve.
[0039] This method breaks through the limitations of the existing expert annotation method, such as low efficiency, high cost, and difficult quality control. It is applicable to multi-modal, multi-temporal, and multi-organizational level data, effectively improves data quality, annotation efficiency, and model generalization ability, and has good adaptability and scalability. Description of the Drawings
[0040] Figure 1 It is a flowchart of the implementation of the present invention. Detailed Embodiments
[0041] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the protection scope of the present invention.
[0042] As Figure 1 shown, a data automatic annotation method based on multi-modal fusion and iterative optimization covers core links such as model automatic annotation, uncertainty identification, expert review and correction, and model continuous optimization. It combines multi-modal enhancement, standardized processing, and cross-domain knowledge fusion, and forms a closed-loop process of "machine annotation - screening anomalies - manual verification - model retraining" through an iterative mechanism, resulting in a "machine main label + manual fine-tuning" process that can continuously learn and self-evolve.
[0043] This method breaks through the limitations of the existing expert annotation method, such as low efficiency, high cost, and difficult quality control. It is applicable to multi-modal, multi-temporal, and multi-organizational level data, effectively improves data quality, annotation efficiency, and model generalization ability, and has good adaptability and scalability.
[0044] Specifically, it includes the following steps:
[0045] (1) Obtain original multi-modal data Including forms such as text, image, audio, video, sensors, etc.;
[0046] (2) Perform format conversion and noise removal processing on each type of data. Fill in missing fields (such as unentered fields in text, missing frames in images, and NaN values in time series), detect and remove outliers, and merge and unify duplicate images, similar texts, and multiple-recorded audio segments, etc. Process heterogeneous data using a unified encoding and size standard (such as UTF-8 encoding for text and standardizing image size to 224×224).
[0047] (3) Perform data augmentation, that is, under the premise of keeping the original semantics unchanged, artificially expand the number and diversity of training samples, improve the generalization ability of the model, and alleviate problems such as data scarcity and class imbalance. Enhance images by applying techniques such as rotation, mirroring, and blurring, and enhance text by using synonymous replacement, word order scrambling, etc. The formula is as follows:
[0048] D * =f aug (D)
[0049] where D * is the enhanced data sample, and f aug is the enhancement function;
[0050] (4) Establish an expert-annotator team. The expert group consists of medical imaging experts, manufacturing industry engineers, linguists, etc. with many years of experience in related fields. The annotator team can be composed of trained professional students, field assistant engineers, or part-time outsourced personnel. The experts formulate annotation rules and supervise the annotators to execute the annotation tasks, and the two complete the high-quality annotation of multi-modal samples through collaboration;
[0051] (5) Set annotation dimensions for each modality (such as target type and bounding box in images; entity categories and relationships in text; intonation and emotion in audio, etc.), and formulate clear and structured feature expression methods and annotation dimensions, whose role is equivalent to "formulating a dictionary" and "defining fields" for each data modality, so that subsequent manual or automatic systems know "which dimensions to focus on" and "how to express these dimensions";
[0052] (6) After the annotators complete the first round of annotation, it is reviewed and corrected by the experts. Since the annotators lack domain knowledge, they are prone to mislabeling and missing labels, and the multi-modal data itself is highly complex. It is necessary to unify the "cognitive standard", that is, form a gold standard sample set to achieve the unity of annotation standards and the controllability of data quality. The formula is as follows:
[0053]
[0054] where x i is the multi-modal input (such as images, text, audio), and y iStructured tags annotated by experts (such as categories, borders, attributes), L expert is a high-quality annotation set;
[0055] (7) Calculate the annotation consistency score and eliminate low-consistency data. After multiple annotators complete the annotation of the same data sample, measure the consistency of their annotation results through statistical methods, automatically screen out data with serious disagreements, and only retain highly consistent data for the training set or the gold standard set. If the consistency score is greater than or equal to the threshold, the sample is determined to be valid; otherwise, the sample is eliminated.
[0056] (8) Construct a multi-modal annotation model, which is an artificial intelligence model that can simultaneously receive and understand data input from multiple modalities (such as images, text, audio, video, sensors) and automatically output the corresponding annotation results. It can learn the annotation rules based on the existing expert-annotated samples and achieve fast, automatic, and scalable annotation of new data. Use a fusion neural network to model data from different modalities. The formula is as follows:
[0057] y = arg max F θ (x text , x image , x audio )
[0058] where F θ represents the parameters of the multi-modal fusion network, x text is the text data, x image is the image data, x audio is the audio data;
[0059] (9) Train the model with the expert-annotated data L expert to conduct the first complete training of the multi-modal annotation model, enabling the model to have preliminary automatic annotation capabilities and laying a foundation for subsequent large-scale inference and iterative training. The loss function is the cross-entropy loss. The formula is as follows:
[0060]
[0061] where F θ (x i ) is the predicted output of the model for x i , log(F θ (x i )) is to take the logarithm of the model's predicted value to form the cross-entropy term with the true label;
[0062] (10) For the multi-modal annotation model after preliminary training, perform inference and prediction on the original samples that have not been annotated, automatically generate the corresponding annotation labels to replace the process of manually generating preliminary labels, and perform inference and prediction on the unlabeled sample x ∈ U to obtain the predicted label y'.
[0063] (11) Calculate whether manual review is required using entropy or model confidence to determine whether the model is "confident". If the class distribution predicted by the model is unclear (high entropy), the sample is put into the review pool to ensure the credibility of the output of the automatic annotation system. The formula is as follows:
[0064]
[0065] where \(H(y')\) is the entropy of the predicted label and \(P(c)\) is the predicted probability;
[0066] (12) In some cases, the automatic annotation system "seems to have marked" but is actually very uncertain. Set a threshold \(\delta\). If \(H(y') > \delta\), the sample is marked as uncertain, and the expert team reviews and optimizes the highly uncertain samples to obtain a correction set, which can avoid directly using "low-confidence annotations" for training and prevent incorrect learning and model degradation. The formula is as follows:
[0067]
[0068] where \(L\) check is the correction set, \(x\) j is the original input of the highly uncertain sample, is the expert-corrected label;
[0069] (13) The system will continuously feedback the corrected samples reviewed by experts, etc. as new training data for incremental training or fine-tuning of the existing model, making the model more accurate when facing similar scenarios in the future and gradually improving the quality of automatic annotation. Incorporate \(L\) check into the training set and continue to fine-tune the model parameters. If the model performs poorly on some specific samples, the expert correction results can be used as "feedback samples". The formula is as follows:
[0070] \(\Gamma\) total =\(\Gamma\) cls +\(\lambda_1\Gamma\) bbox +\(\lambda_2\Gamma\) attr
[0071]
[0072] where \(\Gamma\) total is the joint loss, \(\Gamma\) cls is the cross-entropy loss, \(\Gamma\) bbox is the bounding box regression loss, \(\Gamma\) attr is the attribute binary classification loss, \(\lambda_1\) and \(\lambda_2\) are weight coefficients, \(\theta\) is the learning rate, is the gradient operator;
[0073] (14) After the model accuracy reaches the upper line standard, the system enters the "machine main label - manual fine-tuning" stage, continuously expands the training samples and knowledge base, realizes self-evolution, and constructs a data annotation system that continuously learns, dynamically optimizes, and becomes more accurate with use.
[0074] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A data automatic annotation method based on multimodal fusion and iterative optimization, characterized in that The steps include: (1) Obtain the original multi-modal data including text, image, audio, video, and sensor forms; (2) Perform format conversion and noise removal on each type of data, and use unified encoding and size standards to handle heterogeneous data; (3) Perform data enhancement, that is, while keeping the original semantics unchanged, artificially expand the number and diversity of training samples to improve the generalization ability of the model and alleviate the problems of data scarcity and category imbalance. This is done by applying rotation, mirroring, and blurring techniques to images, and using synonym replacement and word order disruption to text. The formula is as follows: D * = f aug (D) Among them, D * is the enhanced data sample, and f aug is the enhancement function; (4) Establish an expert-annotator team, where experts formulate annotation rules and supervise annotators to perform annotation tasks. The two collaborate to complete high-quality annotation of multimodal samples; (5) Set the annotation dimension for each modality; (6) After the annotators complete the first round of annotation, the experts review and revise the annotations to form the gold standard sample set. The formula is as follows: where x i is the multimodal input, y i is the structured label annotated by experts, and L expert is the high-quality annotation set; (7) Calculate the annotation consistency score and eliminate low consistency data. After multiple annotators complete the annotation of the same data sample, use statistical methods to measure the consistency of their annotation results. (8) Use fusion neural network to model different modal data. The formula is as follows: y = arg max F θ (x text , x image , x audio ) Among them, F θ represents the parameters of the multimodal fusion network, x text is the text data, x image is the image data, x audio is the audio data; (9) Use the expert-annotated data L expert Train the model with the cross-entropy loss as the loss function. The formula is as follows: Among them, F θ (x i ) is the predicted output of the model for x i , log(F θ (x i )) is to take the logarithm of the model's predicted value to form the cross-entropy term with the true label; (10) After completing the initial training, the multimodal annotation model performs inference prediction on the unlabeled sample x∈U to obtain the predicted label y′; (11) Use entropy or model confidence to calculate whether manual review is needed to determine whether the model is confident. The formula is as follows: Where H(y′) is the entropy of the predicted label and P(c) is the predicted probability; (12) Set a threshold δ. If H(y′)>δ, the sample is marked as uncertain. Then the expert team reviews and optimizes the high uncertainty samples to obtain a correction set. The formula is as follows: where L check is the correction set, x j is the original input of the highly uncertain samples, and is the expert correction label; (13), Incorporate L check into the training set and continue to fine-tune the model parameters. The formula is as follows: Γ total = Γ cls + λ1Γ bbox + λ2Γ attr where Γ total is the combined loss, Γ cls is the cross-entropy loss, Γ bbox is the bounding box regression loss, Γ attr is the binary classification loss for attributes, λ1 and λ2 are weight coefficients, and θ is the learning rate, is the gradient operator; (14) After the model accuracy reaches the online standard, the system enters the machine main labeling-manual refinement stage, continuously expanding the training samples and knowledge base to achieve self-evolution and build a data labeling system that continuously learns, dynamically optimizes, and becomes more accurate with use.
Citation Information
Cited By
Multi-modal sample data synthesis and labeling integration method and device, equipment and storage medium
CN121071627A
Large model Prompt dynamic generation method and device and medium
CN121214462A
Rice phenotype extraction method integrating deep learning and expert system
CN121366191A
Placental lesion tissue detection method and placental lesion tissue detection system based on artificial intelligence
CN121563999A
Intelligent annotation and optimization system for multi-modal data based on active learning
CN122758159A