Intelligent data augmentation and cleaning system and method based on multi-modal consistency detection

By introducing multimodal consistency detection and intelligent auxiliary labeling technology into the data processing system, the problem that traditional methods are difficult to deal with multimodal data is solved, high-quality and automated data augmentation and cleaning are achieved, and the consistency and reliability of the data set are improved.

CN120179996APending Publication Date: 2025-06-20INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510253033.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional data processing methods are difficult to effectively process large-scale multimodal data sets, especially in terms of data generation, cleaning and consistency maintenance. The existing technology is difficult to meet the requirements of multimodal data consistency, resulting in a decline in data quality.

Method used

An intelligent data augmentation and cleaning system based on multimodal consistency detection is adopted to generate semantically consistent multimodal data by generating adversarial networks (GAN), variational autoencoder (VAE) and cross-modal feature fusion technology. Combining the similarity score and context consistency detection of the big model, the cleaning and labeling process is automated to ensure the consistency and quality of the data.

Benefits of technology

It significantly improves the quality and consistency of multimodal data sets, reduces manual processing costs, improves the efficiency and accuracy of data annotation, and continuously improves the data processing effect through a closed-loop feedback optimization mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179996A_ABST
    Figure CN120179996A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an intelligent data augmentation and cleaning system and method based on multi-modal consistency detection, and the method comprises the following steps: generating multi-modal data; carrying out multi-mode consistency detection; performing data cleaning and anomaly detection; performing intelligent auxiliary labeling; quality control and feedback optimization; compliance and data traceability management; the method has the beneficial effects that multi-modal data (such as an image-text pair and an image-audio pair) with consistent semantics is generated through a generative adversarial network (GAN), a variational auto-encoder (VAE) and a cross-modal feature fusion technology, and content matching of the generated data among different modals is ensured. According to the consistency guarantee, the quality of a multi-modal data set is remarkably improved, richer and more real training data is provided for a large model, and model performance reduction caused by modal mismatching is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and specifically to an intelligent data augmentation and cleaning system and method based on multi-modal consistency detection. Background Art

[0002] In the context of the rapid development of deep learning and large models, data quality and diversity have become key factors in improving model performance. Especially in the training of multi-modal data, ensuring high-quality, semantic consistency, and diversity of data has an important impact on the accuracy, robustness, and generalization ability of the model. However, traditional data processing methods face many technical challenges when dealing with the generation, cleaning, and consistency maintenance of large-scale multi-modal data sets.

[0003] Current data augmentation methods mainly rely on single-modal processing technologies such as image enhancement and text synonym replacement, and are difficult to meet the multi-modal data consistency requirements. For example, in multi-modal pairs such as image-text or audio-image, augmenting only one modality may lead to semantic inconsistency between modalities, thereby reducing the quality of the data set. In addition, although existing generative models such as generative adversarial networks (GANs) and variational autoencoders (VAEs) can generate diverse data, it is difficult to effectively ensure the consistency and correlation between different modalities.

[0004] In terms of data cleaning, traditional anomaly detection methods usually rely on algorithms such as Isolation Forest or One-Class SVM, and can only detect outliers in single-modal data, and are difficult to identify cross-modal anomalies in multi-modal data. For example, traditional methods have limited effectiveness in detecting mismatches between pictures and text in image-text data sets. Especially in multi-modal data, the lack of semantic consistency and modality matching makes it necessary to spend a lot of manual effort on cleaning and verification after the data set is generated, significantly increasing the cost and time investment. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent data augmentation and cleaning system and method based on multi-modal consistency detection, which realizes the automatic processing of multi-modal data through means such as multi-modal generation, automatic cleaning, and consistency detection, reduces the labor cost of data processing, and improves the quality and diversity of the data set to better support the training and application of large models in multi-modal tasks, so as to solve the problems raised in the above background art.

[0006] To achieve the above purpose, the present invention provides the following technical solution: An intelligent data augmentation and cleaning system based on multi-modal consistency detection, the system includes:

[0007] Multimodal data generation module, including: Generative Adversarial Network (GAN) and Variational Autoencoder (VAE): used to generate multimodal data and ensure semantic consistency between different modalities; Cross-modal feature fusion: through cross-modal feature fusion technology, ensure the relevance of data in different modalities, so that the generated multimodal data matches in semantics and expression;

[0008] Multimodal consistency detection module, including: Inter-modal similarity scoring: use a large model to score the similarity between data in different modalities to ensure data consistency; Context consistency detection: combine context features to detect outliers in multimodal data and ensure data quality;

[0009] Data cleaning and anomaly detection module, including: Context anomaly detection: identify outliers and noise in data through temporal and spatial information, and automatically clean unqualified data; Dynamic modal matching: detect and clean inconsistent data between modalities in multimodal data pairs to ensure data quality and accuracy;

[0010] Intelligent assisted annotation module, including: Automatic annotation and correction: use a large model to generate initial labels and correct the annotation content through the consistency detection module to improve the efficiency and accuracy of annotation; Annotation quality control: perform quality detection on the annotation content through multimodal consistency scoring to ensure data reliability;

[0011] Feedback optimization and quality control module, including: Quality feedback mechanism: continuously optimize the generation and cleaning parameters through user feedback and system automatic scoring to ensure the high quality of augmented and cleaned data; Augmented data evaluation and optimization: based on the quality detection results of the generated data, adjust the generation and cleaning strategies in real time to achieve closed-loop optimization of data quality.

[0012] Preferably, the multimodal data generation module uses the Generative Adversarial Network (GAN) to generate multimodal data pairs. The Generative Adversarial Network (GAN) generates high-quality data through the adversarial training of two neural networks: a generator and a discriminator. The generator generates new image-text pairings, and the discriminator discriminates the authenticity of the generated data, gradually improving the authenticity of the generated data; The Variational Autoencoder (VAE) is used to generate diverse multimodal data. The Variational Autoencoder (VAE) encodes and decodes through a latent space, making the generated data diverse and conforming to the distribution of real data; Through cross-modal feature fusion, text features are embedded into the image generation process to achieve the fusion of different modal features. When generating images, the semantic information described in the text is incorporated into the content and details of the generated image to ensure semantic consistency of the generated data across different modalities.

[0013] Preferably, the multimodal consistency detection module calculates the similarity score between image and text multimodal data based on the feature extraction ability of the large model. By training the multimodal encoder, it can capture the semantic relevance between image-text or image-audio and generate a similarity score. Samples with low scores will be screened out to ensure the multimodal consistency of the data. It uses temporal, spatial information, and feature distribution to detect the context consistency between modalities and ensure the rationality of the data in the same task or scenario.

[0014] Preferably, the data cleaning and anomaly detection module detects data anomalies based on the Isolation Forest and One-Class SVM algorithms, combined with temporal and spatial information. In time-series data, it detects and filters points with inconsistent contexts to ensure that anomaly detection is not only based on a single feature but also comprehensively considers the context of the data. It detects and cleans data samples with inconsistent modalities through deep learning and contrast learning techniques.

[0015] Preferably, the intelligent auxiliary annotation module uses a pre-trained large model to preliminarily annotate the data; automatically corrects the annotation content based on the consistency detection module; the system automatically scores after annotation is completed, and combines the inter-modal similarity detection to screen out low-quality annotations, improving the accuracy and consistency of data annotation;

[0016] The feedback optimization and quality control module continuously optimizes the generation parameters and cleaning strategies by collecting user feedback through the system and combining the generated data quality scores. The user's evaluation of data quality is used to adjust the weights and parameters in the augmentation and cleaning algorithms to improve the quality of the generated data. The system automatically evaluates the distribution, features, and modal consistency of the augmented data, screens out data that does not meet the quality standards, and dynamically adjusts the parameters of the generation model and cleaning algorithm according to the detection results. Through feedback optimization, it realizes the adaptive improvement of data generation and cleaning.

[0017] The system also includes compliance and data copyright protection. The system records the process of each data augmentation, cleaning, and annotation, and uses distributed storage technology to save the processing logs to ensure the transparency and traceability of the data processing process. The system encrypts and stores the generated and processed data and provides access rights management to ensure the secure use of sensitive data in a compliant environment.

[0018] An intelligent data augmentation and cleaning method based on multimodal consistency detection uses an intelligent data augmentation and cleaning system based on multimodal consistency detection. The method includes the following steps:

[0019] Multimodal data generation;

[0020] Multimodal consistency detection;

[0021] Data cleaning and anomaly detection;

[0022] Intelligent assisted annotation;

[0023] Quality control and feedback optimization;

[0024] Compliance and data traceability management.

[0025] Preferably, the specific operations of multi-modal data generation include: the system extracts samples from the original data and generates multi-modal data pairs based on the Generative Adversarial Network (GAN) and Variational Autoencoder (VAE); using cross-modal feature fusion technology, the multi-modal features are embedded into the generation process to ensure semantic consistency and content matching across different modalities of the generated data;

[0026] The multi-modal consistency detection specifically includes: the system performs similarity scoring on the generated multi-modal data pairs through a large model to ensure semantic consistency between different modal data. Based on the context consistency detection algorithm, the system cleans the noise and outliers in the multi-modal data, filtering out samples that do not conform to the context or semantic consistency.

[0027] Preferably, the data cleaning and anomaly detection specifically include: the system uses the context anomaly detection algorithm to identify and clean the noise and outliers in the dataset, especially the inter-modal inconsistent data in multi-modal data. Through the dynamic modal matching technology, it further screens the samples that do not conform to the modal consistency, ensuring the high quality and consistency of the multi-modal dataset.

[0028] Preferably, the specific operations of intelligent assisted annotation include: the system performs intelligent annotation on the generated and cleaned data, automatically generates labels through a large model, and the system corrects the annotation content in combination with the consistency detection module after the annotation is generated, correcting the inconsistencies or errors in the annotation to improve the accuracy of the annotation.

[0029] Preferably, the specific operations of quality control and feedback optimization include: the system performs quality scoring on the generated and cleaned data, and performs real-time optimization in combination with user feedback. If the quality of the generated data or the annotated data does not meet the standard, the system automatically adjusts the generation parameters, cleaning rules, and annotation strategies to optimize the effect of data processing, forming a closed-loop optimization mechanism;

[0030] The specific operations of compliance and data traceability management include: the system stores all operation records during the data generation, cleaning, and annotation processes, generates data processing logs, and ensures the traceability and compliance of the data through distributed storage. The processed data is encrypted and stored with access permissions set to ensure the compliance and security of data usage.

[0031] Compared with the prior art, the beneficial effects of the present invention are:

[0032] The intelligent data augmentation and cleaning system and method based on multi-modal consistency detection proposed by the present invention generate multi-modal data with consistent semantics (such as image-text pairs, image-audio pairs) through generative adversarial networks (GANs), variational autoencoders (VAEs), and cross-modal feature fusion techniques, ensuring content matching between different modalities for the generated data. This consistency guarantee significantly improves the quality of multi-modal datasets, provides richer and more realistic training data for large models, and reduces the performance degradation of models caused by modality mismatches.

[0033] Using the context anomaly detection algorithm combined with temporal, spatial, and feature correlation information, automatically identify and clean outliers and noisy data in multi-modal datasets. Compared with traditional single-modal cleaning methods, the cleaning module of the present invention not only improves the cleaning efficiency but also significantly enhances the overall quality and reliability of multi-modal datasets, ensuring the correlation consistency of data across different modalities.

[0034] Automatically generate initial labels through a large model and combine the consistency detection module to correct the annotation content, ensuring the semantic accuracy and modal consistency of the annotation content. This intelligent assisted annotation and automatic correction mechanism reduces the manual annotation burden, improves the annotation efficiency and accuracy at the same time, and provides a guarantee for the high quality of the dataset.

[0035] Adaptive optimization of the generation and cleaning processes through user feedback and system scoring. This closed-loop feedback significantly improves the effects of data augmentation and cleaning, ensures the high quality and diversity of the generated data, supports the continuous improvement of model training data, and enhances the adaptability of the platform in a dynamic data environment.

[0036] By recording the operation logs of the data processing process, ensure the traceability of each augmentation, cleaning, and annotation, and at the same time provide encryption storage and access permission control functions to achieve the secure management and compliant use of data. This compliance guarantee is particularly important in the processing of sensitive data (such as medical and financial data), effectively reducing the risk of data management and providing support for compliance and transparency in data circulation. Description of the Drawings

[0037] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments

[0038] In order to clearly and completely describe the objectives, technical solutions of the present invention, and make the advantages more clearly understood, the following further elaborates on the embodiments of the present invention in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, rather than all of the embodiments, and are merely used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0039] Embodiment 1, the present invention provides a technical solution: an intelligent data augmentation and cleaning system based on multi-modal consistency detection, the system includes:

[0040] 1) Multi-modal data generation module

[0041] Generative Adversarial Network (GAN): Use GAN to generate multi-modal data pairs, such as image-text pairs. GAN generates high-quality data through the adversarial training of two neural networks (generator and discriminator). The generator generates new image-text pairings, and the discriminator discriminates the authenticity of the generated data, thereby gradually improving the authenticity of the generated data.

[0042] Variational Autoencoder (VAE): Use VAE to generate diverse multi-modal data. VAE encodes and decodes through the latent space, making the generated data diverse and conforming to the distribution of real data.

[0043] Cross-modal feature fusion: Through cross-modal feature fusion (such as the Cross-Attention mechanism based on Transformer), embed text features into the image generation process to achieve the fusion of different modal features. When generating images, integrate the semantic information described in the text into the content and details of the generated images to ensure the semantic consistency of the generated data across different modalities.

[0044] 2) Multi-modal consistency detection module

[0045] Inter-modal similarity scoring: Based on the feature extraction ability of large models (such as CLIP), calculate the similarity scores between multi-modal data such as images and texts. By training a multi-modal encoder, enable it to capture the semantic relevance between image-text or image-audio, and generate similarity scores. Samples with low scores will be screened out to ensure the multi-modal consistency of the data.

[0046] Context consistency detection: Use temporal, spatial information, and feature distribution to detect the context consistency between modalities, ensuring the rationality of the data in the same task or scenario. For example, in video-audio pairings, by detecting the consistency between the frame sequence of the video and the audio changes, automatically clean the noise and abnormal data that do not conform to the context.

[0047] 3) Data cleaning and anomaly detection module

[0048] Context anomaly detection: Based on algorithms such as Isolation Forest and One-Class SVM, combined with temporal and spatial information to detect data anomalies. In time series data, detect and filter out points with inconsistent contexts, ensuring that anomaly detection is not only based on a single feature but also comprehensively considers the context of the data.

[0049] Dynamic modality matching: Through deep learning and contrastive learning techniques, detect and clean data samples with inconsistent modalities. For example, for image-text data pairs, the system cleans out mismatched samples by analyzing the consistency between the image content and the text description, ensuring that the final generated data has modality matching and consistency.

[0050] 4) Intelligent assisted annotation module

[0051] Automatic annotation generation: Use pre-trained large models to perform preliminary annotation on the data. For example, use an image classification model to generate labels for images and then use a natural language processing model to generate labels for text.

[0052] Annotation correction: Automatically correct the annotation content based on the consistency detection module. For example, for the annotation content in image-text pairs, by calculating the similarity score between the image and the text, ensure the consistency between the image content and the text annotation, and automatically correct the mismatched annotation content, thereby reducing manual annotation errors.

[0053] Annotation quality control: The system automatically scores after annotation is completed, and combines inter-modal similarity detection to screen out low-quality annotations, improving the accuracy and consistency of data annotation.

[0054] 5) Feedback optimization and quality control module:

[0055] Quality feedback mechanism: The system collects user feedback and combines it with the generated data quality score to continuously optimize the generation parameters and cleaning strategies. The user's evaluation of data quality is used to adjust the weights and parameters in the augmentation and cleaning algorithms to improve the quality of the generated data.

[0056] Augmented data evaluation and optimization: The system automatically evaluates the distribution, features, and modality consistency of the augmented data, screens out data that does not meet the quality standards, and dynamically adjusts the parameters of the generation model and cleaning algorithm according to the detection results. Through feedback optimization, achieve adaptive improvement of data generation and cleaning.

[0057] 6) Compliance and data copyright protection:

[0058] Data Traceability and Compliance Management: The system records the process of each data augmentation, cleaning, and annotation, and uses distributed storage technology to save the processing logs to ensure the transparency and traceability of the data processing process.

[0059] Data Copyright and Access Control: The system encrypts and stores the generated and processed data, and provides access permission management to ensure the secure use of sensitive data in a compliant environment.

[0060] Example 2, referring to the appendix Figure 1 , based on Example 1, a method for intelligent data augmentation and cleaning based on multimodal consistency detection is proposed, and an intelligent data augmentation and cleaning system based on multimodal consistency detection is adopted. The method includes the following steps:

[0061] Multimodal Data Generation:

[0062] The system extracts samples from the original data and generates multimodal data pairs based on the Generative Adversarial Network (GAN) and Variational Autoencoder (VAE), such as image-text or image-audio data. Using cross-modal feature fusion technology, multimodal features are embedded in the generation process to ensure semantic consistency and content matching across different modalities of the generated data.

[0063] Multimodal Consistency Detection:

[0064] The system performs similarity scoring on the generated multimodal data pairs through a large model to ensure semantic consistency between different modality data. Based on the context consistency detection algorithm, the system cleans the noise and outliers in the multimodal data, filtering out samples that do not conform to the context or semantic consistency.

[0065] Data Cleaning and Anomaly Detection:

[0066] The system uses the context anomaly detection algorithm to identify and clean the noise and outliers in the dataset, especially the inter-modal inconsistent data in the multimodal data. Through the dynamic modal matching technology, samples that do not conform to the modal consistency are further screened to ensure the high quality and consistency of the multimodal dataset.

[0067] Intelligent Assistant Annotation:

[0068] The system performs intelligent annotation on the generated and cleaned data, and automatically generates labels through a large model. After the annotation is generated, the system combines the consistency detection module to correct the annotation content, correcting the inconsistencies or errors in the annotation to improve the accuracy of the annotation.

[0069] Quality Control and Feedback Optimization:

[0070] The system performs quality scoring on the generated and cleaned data and conducts real-time optimization in combination with user feedback. If the quality of the generated data or labeled data fails to meet the standards, the system automatically adjusts the generation parameters, cleaning rules, and labeling strategies to optimize the effect of data processing, forming a closed-loop optimization mechanism.

[0071] Compliance and data traceability management:

[0072] The system stores all operation records during the data generation, cleaning, and labeling processes, generates data processing logs, and ensures the traceability and compliance of the data through distributed storage. The processed data is encrypted and stored with access permissions set to ensure the compliance and security of data usage.

[0073] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent data augmentation and cleaning system based on multimodal consistency detection, characterized by: The system comprises: Multimodal data generation module, including: Generative Adversarial Network (GAN) and Variational Autoencoder (VAE): used to generate multimodal data and ensure the semantic consistency of generated data between different modalities; Cross-modal feature fusion: through cross-modal feature fusion technology, ensure the relevance of data in different modalities, so that the generated multimodal data matches in semantics and expression; The multimodal consistency detection module includes: inter-modal similarity scoring: using a large model to score the similarity between different modal data to ensure data consistency; contextual consistency detection: combining contextual features to detect outliers in multimodal data to ensure data quality; Data cleaning and anomaly detection module, including: context anomaly detection: identifying outliers and noise in data through time series and spatial information, and automatically removing unqualified data; dynamic mode matching: detecting and cleaning inconsistent data between modes in multi-modal data pairs to ensure data quality and accuracy; Intelligent auxiliary annotation module, including: automatic annotation and correction: using the large model to generate initial labels, and correcting the annotation content through the consistency detection module to improve the efficiency and accuracy of annotation; annotation quality control: through multi-modal consistency scoring, the quality of the annotation content is checked to ensure the reliability of the data; Feedback optimization and quality control module, including: Quality feedback mechanism: Continuously optimize generation and cleaning parameters through user feedback and system automatic scoring to ensure the high quality of augmented and cleaned data; Augmented data evaluation and optimization: Based on the quality detection results of generated data, adjust the generation and cleaning strategies in real time to achieve closed-loop optimization of data quality.

2. The intelligent data augmentation and cleaning system based on multimodal consistency detection according to claim 1, characterized in that: The multimodal data generation module uses the adversarial network GAN to generate multimodal data pairs. The adversarial network GAN generates high-quality data through adversarial training of two neural networks: the generator and the discriminator. The generator generates new image-text pairs, and the discriminator judges the authenticity of the generated data, gradually improving the authenticity of the generated data. The variational autoencoder VAE is used to generate diverse multimodal data. The variational autoencoder VAE encodes and decodes through the latent space, so that the generated data is diverse and conforms to the distribution of real data. Through cross-modal feature fusion, text features are embedded into the image generation process to achieve the fusion of different modal features. When generating images, the semantic information described in the text is integrated into the content and details of the generated image to ensure the semantic consistency of the generated data in different modalities.

3. The intelligent data augmentation and cleaning system based on multimodal consistency detection according to claim 2, characterized in that: The multimodal consistency detection module calculates the similarity score between image and text multimodal data based on the feature extraction capability of the large model. By training the multimodal encoder, it is able to capture the semantic correlation between image and text or image and audio, and generate a similarity score. Low-scoring samples will be screened out to ensure the multimodal consistency of the data. The module uses temporal, spatial information and feature distribution to detect contextual consistency between modalities to ensure the rationality of data in the same task or scenario.

4. The intelligent data augmentation and cleaning system based on multimodal consistency detection according to claim 1, characterized in that: The data cleaning and anomaly detection module is based on the Isolation Forest and One-Class SVM algorithms, combining time series and spatial information to detect data anomalies. In time series data, it detects and filters points with inconsistent contexts to ensure that anomaly detection is not only based on a single feature, but also comprehensively considers the context of the data; through deep learning and contrastive learning techniques, it detects and cleans inconsistent data samples between modalities.

5. The intelligent data augmentation and cleaning system based on multimodal consistency detection according to claim 1, characterized in that: The intelligent auxiliary annotation module uses a pre-trained large model to perform preliminary annotation on the data; the annotation content is automatically corrected based on the consistency detection module; the system automatically scores after the annotation is completed, and combines the inter-modal similarity detection to filter out low-quality annotations, thereby improving the accuracy and consistency of data annotation; Feedback optimization and quality control module, through the system to collect user feedback and combine the generated data quality score, continuously optimize the generation parameters and cleaning strategy, the user's evaluation of data quality is used to adjust the weights and parameters in the augmentation and cleaning algorithms, and improve the quality of generated data; the system automatically evaluates the distribution, characteristics and modal consistency of augmented data, screens out data that does not meet the quality standards, and dynamically adjusts the parameters of the generation model and cleaning algorithm according to the detection results, and realizes adaptive improvement of data generation and cleaning through feedback optimization; The system also includes compliance and data copyright protection. The system records each process of data augmentation, cleaning and labeling, and uses distributed storage technology to save processing logs to ensure the transparency and traceability of the data processing process. The system encrypts and stores the generated and processed data, and provides access permission management to ensure the safe use of sensitive data in a compliant environment.

6. An intelligent data augmentation and cleaning method based on multimodal consistency detection, using an intelligent data augmentation and cleaning system based on multimodal consistency detection as described in any one of claims 1 to 5 above, characterized in that: The method comprises the following steps: Multimodal data generation; Multimodal consistency detection; Data cleaning and anomaly detection; Intelligent auxiliary annotation; Quality control and feedback optimization; Compliance and data traceability management.

7. The intelligent data augmentation and cleaning method based on multimodal consistency detection according to claim 6 is characterized in that: The specific operations of multimodal data generation include: the system extracts samples from the original data and generates multimodal data pairs based on the generative adversarial network (GAN) and the variational autoencoder (VAE); using cross-modal feature fusion technology, multimodal features are embedded in the generation process to ensure the semantic consistency and content matching of the generated data in different modalities; Multimodal consistency detection specifically includes: the system uses a large model to perform similarity scoring on the generated multimodal data pairs to ensure semantic consistency between data of different modalities. Based on the contextual consistency detection algorithm, the system cleans the noise and outliers in the multimodal data and filters out samples that do not conform to the context or semantic consistency.

8. The intelligent data augmentation and cleaning method based on multimodal consistency detection according to claim 6, characterized in that: Data cleaning and anomaly detection specifically include: the system uses contextual anomaly detection algorithms to identify and clean noise and outliers in the data set, especially inconsistent data between modalities in multimodal data, and further screens samples that do not meet modal consistency through dynamic modal matching technology to ensure the high quality and consistency of multimodal data sets.

9. The intelligent data augmentation and cleaning method based on multimodal consistency detection according to claim 6, characterized in that: The specific operations of intelligent assisted labeling include: the system intelligently labels the generated and cleaned data, automatically generates labels through a large model, and after the label is generated, the system combines the consistency detection module to correct the label content, correct inconsistencies or errors in the label, and improve the accuracy of the label.

10. The intelligent data augmentation and cleaning method based on multimodal consistency detection according to claim 6, characterized in that: The specific operations of quality control and feedback optimization include: the system scores the quality of generated and cleaned data, and performs real-time optimization based on user feedback. If the quality of generated data or labeled data does not meet the standards, the system automatically adjusts the generation parameters, cleaning rules and labeling strategies to optimize the effect of data processing and form a closed-loop optimization mechanism; The specific operations of compliance and data traceability management include: the system stores all operation records during data generation, cleaning and annotation, generates data processing logs, and ensures data traceability and compliance through distributed storage, encrypts and stores processed data and sets access rights to ensure compliance and security of data use.

Citation Information

Cited By

  • Data cleaning method, device and equipment and storage medium

    CN121071310A

  • Intelligent content consistency detection method and system for multi-modal report material

    CN121301615A

  • Intelligent detection method and system for content consistency of multi-modal report materials

    CN121301615B

  • Large-model multi-modal and multi-dimensional data augmentation method and device based on man-machine interaction

    CN121996338A