Multi-modal large model data cleaning treatment method and system
By performing format filtering and semantic-level alignment evaluation on the multimodal data set, high-quality, semantic-aligned image-text data pairs are selected, which solves the problem of insufficient cross-modal semantic association evaluation in multimodal large models, and improves the accuracy and robustness of the model.
Patent Information
- Application Number
- CN202510820032.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The prior art fails to effectively evaluate the cross-modal semantic association between images and text in multimodal data cleaning, resulting in cognitive bias in multimodal large models during feature extraction and cross-modal alignment, affecting accuracy and robustness.
After format filtering of multimodal data sets, a single-modal quality evaluation mechanism is used to quantify image clarity and text fluency, and a semantic-level alignment evaluation mechanism is introduced to conduct semantic-level interaction response analysis of image-text data pairs, and highly semantic-aligned data pairs are selected.
Ensure that multimodal training samples are highly aligned at the cross-modal semantic level, improving the accuracy and robustness of multimodal large models in cross-modal understanding and generation tasks.
Smart Images

Figure CN120336725A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data cleaning, and more specifically, to a method and system for cleaning and governing multi-modal large model data. Background Art
[0002] With the rapid development of artificial intelligence technology towards the direction of multi-modal integration, large-scale multi-modal pre-trained models have become the core driving force for promoting cross-modal understanding and generation tasks. The performance of such models highly depends on the quality of training data. However, in actual application scenarios, the original multi-modal data sets often have significant quality defects: image data may contain visual information with low resolution, blur, or distortion; text descriptions may have grammar errors, information redundancy, or semantic ambiguity. More seriously, there are often problems of semantic misalignment or insufficient correlation between image-text data pairs. These data defects will lead to cognitive biases in the model during feature extraction and cross-modal alignment processes, thereby affecting the accuracy and robustness of downstream tasks. Therefore, constructing an effective multi-modal large model data cleaning and governing solution to screen high-quality and highly relevant training samples from massive original data is crucial for improving the performance of multi-modal large models.
[0003] Currently, for the cleaning and governance of multi-modal data, existing technologies are mostly limited to the independent processing of single modalities. In the image dimension, simple size filtering or format verification is usually adopted, and in the text dimension, it relies on rule-based keyword matching or length truncation. Although these methods can remove obviously invalid data, they fail to deeply evaluate the quality characteristics within each modality and lack refined measurement of cross-modal semantic associations. Especially when dealing with image-text paired data, traditional technologies often ignore the deep correspondence between visual features and semantic descriptions, resulting in low-quality data samples that do not trigger basic filtering rules to continuously contaminate the training set, ultimately causing feature response deviation or semantic understanding deviation in the model's multi-modal alignment task.
[0004] Therefore, an optimized method and system for cleaning and governing multi-modal large model data are expected. Summary of the Invention
[0005] To solve the above technical problems, the present application is proposed. Embodiments of the present application provide a method and system for cleaning and governing multi-modal large model data. After performing basic format filtering on the original multi-modal data set, the clarity of images and the fluency of texts in the multi-modal data set are respectively quantitatively evaluated through a single-modal quality assessment mechanism to screen out qualified image and text data samples. Furthermore, by introducing a semantic-level alignment assessment mechanism, semantic-level interaction response analysis is performed on each corresponding image sample and image text description in the data set to quantitatively evaluate the semantic alignment degree between the image sample and the text description, thereby further screening out highly semantically aligned image-text data pairs. By performing multi-level cleaning and governance on the multi-modal data set, this method can ensure that the multi-modal training samples not only meet the quality standards but also achieve a high degree of alignment at the cross-modal semantic level, thus effectively improving the accuracy and robustness of the multi-modal large model in cross-modal understanding and generation tasks.
[0006] According to one aspect of the present application, there is provided a method for cleaning and governing multi-modal large model data, which includes: Obtain an original multi-modal data set; After performing initial data cleaning on the original multi-modal data set, extract multi-modal data samples to be selected, where the multi-modal data samples to be selected include image data to be selected and corresponding text descriptions to be selected for the image data to be selected; Extract visual features from the image data to be selected to obtain a visual feature encoding vector of the image to be selected; Extract semantic features from the text description to be selected to obtain a semantic feature encoding vector of the text description to be selected; Perform semantic-level fine-grained alignment encoding on the visual feature encoding vector of the image to be selected and the semantic feature encoding vector of the text description to be selected to obtain a fine-grained interaction response encoding vector of the image-text semantic level to be selected; Based on the fine-grained interaction response encoding vector of the image-text semantic level to be selected, determine whether to filter the multi-modal data samples to be selected.
[0007] According to another aspect of the present application, there is provided a system for cleaning and governing multi-modal large model data, which includes: A multi-modal data acquisition module for obtaining an original multi-modal data set; A multi-modal data preprocessing module for performing initial data cleaning on the original multi-modal data set and extracting multi-modal data samples to be selected therefrom, where the multi-modal data samples to be selected include image data to be selected and corresponding text descriptions to be selected for the image data to be selected; A visual feature extraction module for extracting visual features from the image data to be selected to obtain a visual feature encoding vector of the image to be selected; A semantic feature extraction module, configured to extract semantic features from the text description to be selected to obtain a semantic feature encoding vector of the text description to be selected; A fine-grained alignment encoding module, configured to perform semantic-level fine-grained alignment encoding on the visual feature encoding vector of the image to be selected and the semantic feature encoding vector of the text description to be selected to obtain a fine-grained interaction response encoding vector of the image-text semantic level of the image to be selected; A sample screening module, configured to determine whether to filter the multi-modal data sample to be selected based on the fine-grained interaction response encoding vector of the image-text semantic level of the image to be selected.
[0008] Compared with the prior art, the multi-modal large model data cleaning and governance method and system provided by this application, after performing basic format filtering on the original multi-modal data set, respectively quantify and evaluate the image clarity and text fluency in the multi-modal data set through a single-modal quality assessment mechanism to screen out qualified image and text data samples. Furthermore, by introducing a semantic-level alignment assessment mechanism, perform semantic-level interaction response analysis on each corresponding image sample and image text description in the data set to quantify and evaluate the semantic alignment degree between the image sample and the text description, so as to further screen out highly semantically aligned image-text data pairs. This method can ensure that the multi-modal training samples not only meet the quality standards, but also achieve a high degree of alignment at the cross-modal semantic level through multi-level cleaning and governance of the multi-modal data set, thereby effectively improving the accuracy and robustness of the multi-modal large model in cross-modal understanding and generation tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0010] Figure 1 FIG. is a flowchart of a multi-modal large model data cleaning and governance method according to an embodiment of the present application.
[0011] Figure 2 FIG. is a data flow diagram of a multi-modal large model data cleaning and governance method according to an embodiment of the present application.
[0012] Figure 3 FIG. is a flowchart of sub-step S2 of a multi-modal large model data cleaning and governance method according to an embodiment of the present application.
[0013] Figure 4It is a flowchart of sub-step S22 of the multimodal large model data cleaning and governance method according to an embodiment of the present application.
[0014] Figure 5 It is a flowchart of sub-step S5 of the multimodal large model data cleaning and governance method according to an embodiment of the present application.
[0015] Figure 6 It is a flowchart of sub-step S6 of the multimodal large model data cleaning and governance method according to an embodiment of the present application.
[0016] Figure 7 It is a block diagram of the multimodal large model data cleaning and governance system according to an embodiment of the present application. Detailed implementation manners
[0017] As shown in the present application and the claims, unless the context clearly indicates an exception, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0018] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or server. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0019] In the present application, flowcharts are used to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations before or below do not necessarily need to be executed precisely in order. On the contrary, as needed, various steps can be processed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.
[0020] Next, example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described here.
[0021] It is worth noting that in the present application, all actions of obtaining data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining the authorization given by the owner of the corresponding device.
[0022] In view of the technical problems described in the above background art, the present application proposes a method for cleaning and governing multi-modal large model data. After performing basic format filtering on the original multi-modal data set, it respectively quantifies and evaluates the image clarity and text fluency in the multi-modal data set through a single-modal quality assessment mechanism to screen out qualified image and text data samples. Furthermore, by introducing a semantic-level alignment assessment mechanism, semantic-level interactive response analysis is performed on each corresponding image sample and image text description in the data set to quantify and evaluate the semantic alignment degree between the image sample and the text description, thereby further screening out highly semantically aligned image-text data pairs. By performing multi-level cleaning and governance on the multi-modal data set, this method can ensure that the multi-modal training samples not only meet the quality standards but also achieve a high degree of alignment at the cross-modal semantic level, thus effectively improving the accuracy and robustness of the multi-modal large model in cross-modal understanding and generation tasks.
[0023] Figure 1 It is a flowchart of the method for cleaning and governing multi-modal large model data according to an embodiment of the present application. Figure 2 It is a data flow diagram of the method for cleaning and governing multi-modal large model data according to an embodiment of the present application. As Figure 1 and Figure 2 shown, the method for cleaning and governing multi-modal large model data includes the steps: S1, obtaining an original multi-modal data set; S2, performing initial data cleaning on the original multi-modal data set and extracting multi-modal data samples to be selected therefrom, where the multi-modal data samples to be selected include image data to be selected and corresponding text descriptions to be selected for the image data to be selected; S3, performing visual feature extraction on the image data to be selected to obtain an encoded vector of visual features of the image to be selected; S4, performing semantic feature extraction on the text description to be selected to obtain an encoded vector of semantic features of the text description to be selected; S5, performing semantic-level fine-grained alignment encoding on the encoded vector of visual features of the image to be selected and the encoded vector of semantic features of the text description to be selected to obtain an encoded vector of semantic-level fine-grained interactive response of the image-text to be selected; S6, based on the encoded vector of semantic-level fine-grained interactive response of the image-text to be selected, determining whether to filter the multi-modal data samples to be selected.
[0024] In the above multi-modal large model data cleaning and governance method, in step S1, the original multi-modal data set is obtained. It should be understood that the training of multi-modal large models relies on a large amount of heterogeneous data. To construct an initial multi-modal data set, this application is based on a distributed data collection framework and aggregates multi-modal data containing image-text pairs from public or private data sources through methods such as customized crawler tools (such as Scrapy combined with Selenium to simulate browser behavior), API interfaces (such as Flickr API, social media open platforms), or database exports (such as MySQL relational databases). Specifically, for image data, standard format files such as JPEG and PNG and associated text descriptions (including user comments, automatically generated Alt-Text, or manually annotated captions) are collected, and metadata (such as image resolution, text encoding format) is recorded at the same time; for text data, natural language descriptions, tags, or annotation information corresponding to the images are extracted, and character set conflicts are avoided through UTF-8 encoding normalization processing. In this way, it helps to ensure the diversity and scale of the data source, provides the original input for subsequent data set cleaning and governance, and at the same time retains the potential semantic relevance of the original data.
[0025] In the specific implementation process, the acquisition of the original multi-modal data set usually depends on a distributed heterogeneous data collection framework. This framework aggregates through multiple data source channels, including but not limited to web crawler technology, open platform API interface calls, and database exports. These methods can efficiently capture multi-modal data resources containing image-text pairs from public or private data sources. For example, using a customized web crawler tool (such as Scrapy combined with Selenium to simulate browser behavior), the embedded images and their corresponding text description information can be automatically extracted from the web page structure; at the same time, with the help of standardized interfaces provided by Flickr API, social media open platforms, etc., image-text paired data with tags, annotations, and metadata can be systematically obtained. In addition, for the existing historical data assets within an enterprise, the graphically related data stored in a relational database (such as MySQL) can be imported into a unified data set through database export. These diverse collection means together constitute the technical foundation for the construction of the multi-modal data set, which helps to ensure the breadth and depth of the data obtained.
[0026] In the aspect of image data collection, the focus is mainly on obtaining standard format files, such as common image formats like JPEG and PNG, and simultaneously collecting the related natural language text descriptions. These text descriptions may come from user comments, automatically generated Alt-Text tags, or manually annotated captions, etc. They provide context information for the image at the semantic level. At the same time, during the data collection process, metadata information of the image is recorded, such as resolution, aspect ratio, camera model, etc. Although these information do not directly participate in the subsequent semantic analysis, they have important reference value in evaluating image clarity and judging whether it meets the training standards. In addition, to avoid text parsing errors caused by character encoding issues, all text descriptions will undergo UTF-8 encoding normalization to eliminate compatibility barriers between different character sets, thus ensuring the stability and consistency of the text content in the subsequent feature extraction stage.
[0027] It is worth noting that the construction of the original multimodal dataset is not simply stitching images and texts together, but the potential semantic relevance between the two needs to be retained. This semantic association may be explicit, such as having clear text descriptions below the image, or implicit, such as there being a certain thematic relevance between the pictures on social platforms and the text dynamics posted by users. Therefore, special attention should be paid to the pairing logic between images and texts during the data collection process, and data pairs with clear semantic connections should be selected as much as possible to improve the effectiveness of subsequent semantic-level alignment evaluation. For example, on social media platforms, a travel photo is often accompanied by the user's sharing of travel feelings or location descriptions, and such picture-text combinations naturally have a high degree of semantic consistency and are suitable as training samples for multimodal models.
[0028] In addition, considering the high demand for data scale in the training of multimodal large models, the construction of the original dataset also needs to balance the acquisition efficiency and data capacity. For this reason, a distributed data collection architecture is widely adopted. By multiple nodes executing data scraping tasks in parallel, the speed and stability of data acquisition are significantly improved. At the same time, a caching mechanism and a deduplication strategy are introduced to avoid repeated collection of the same or highly similar data samples, thus controlling data redundancy while ensuring data diversity. On this basis, an incremental update mechanism can also be combined to regularly scrape new content from the data source to keep the original dataset up-to-date and continuously expandable.
[0029] In the above multimodal large model data cleaning and governance method, in step S2, after initial data cleaning of the original multimodal dataset, the multi-modal data samples to be selected are extracted from it. The multi-modal data samples to be selected include the image data to be selected and the corresponding text descriptions to be selected for the image data to be selected. Among them, Figure 3It is a flowchart of sub-step S2 of the multi-modal large model data cleaning and governance method according to an embodiment of the present application. As Figure 3 shown, the step S2 includes steps: S21, performing basic filtering on the original multi-modal data set to obtain the original multi-modal data set after basic filtering; S22, performing single-modal quality assessment on the original multi-modal data set after basic filtering to obtain the original multi-modal data set after secondary filtering; S23, extracting the multi-modal data samples to be selected from the original multi-modal data set after secondary filtering.
[0030] Specifically, in a specific example of the present application, the step S21 includes: removing the obvious invalid data samples in the original multi-modal data set, and the obvious invalid data samples include unloadable images, too short / too long texts, and specific format errors. It should be understood that during the acquisition, storage, or transmission process, invalid data such as unloadable images, format errors, or text with missing content will inevitably be mixed in. It not only cannot provide effective information for the training of the multi-modal large model, but also occupies computing resources, increases the training time cost, and may even cause errors in model training. Therefore, in order to initially purify the original multi-modal data set, the present application, based on the principle of data validity determination, obtains the original multi-modal data set after basic filtering by removing the obvious invalid data samples in the original multi-modal data set. Specifically, in the image processing scenario, use an image loading library (such as the Pillow library in Python) to try to load each image in the data set. If an error occurs during the loading process or the basic information of the image (such as size, color mode) cannot be obtained, then determine that the image is an unloadable invalid sample and delete it; in the text processing scenario, for text data, set a reasonable length range, for example, stipulate that the text length needs to be between 10 characters and 500 characters, and screen out text samples that are too short (such as only containing 1-2 meaningless characters) or too long (more than 1000 characters and the content is redundant and chaotic); for specific format errors, such as the existence of garbled characters (including special character encodings that cannot be recognized) in text data or the image format does not conform to common standards (such as non-JPEG, PNG, BMP, etc. formats), directly remove such data samples. In this way, it is possible to quickly and effectively remove the obvious invalid data samples in the data set, making the original multi-modal data set after basic filtering more concise and effective, thereby reducing the complexity of subsequent data processing.
[0031] Specifically, in step S22, the original multi-modal data set after basic filtering is subjected to single-modal quality assessment to obtain the original multi-modal data set after secondary filtering. Specifically, since the above basic filtering can only remove obviously invalid data and cannot detect the quality of the data content, there may still be potential low-quality data samples such as blurred images, redundant texts, or unclear semantics in the original multi-modal data set after basic filtering. Therefore, in order to further improve the overall quality of the data set, the present application introduces a single-modal quality assessment mechanism on the basis of basic filtering to perform quality assessment on the image data and text data in the original multi-modal data set after basic filtering respectively, so as to screen out qualified image and text data samples. Among them, Figure 4 is a flowchart of sub-step S22 of the multi-modal large model data cleaning and governance method according to an embodiment of the present application. As Figure 4 shown, step S22 includes the steps of: S221, using a pre-trained model to evaluate the image clarity of each image data and the text fluency of each text description in the original multi-modal data set after basic filtering; S222, based on the comparison between the image clarity and a preset clarity threshold, filtering the image data in the original multi-modal data set whose quality does not meet the requirements; S223, based on the comparison between the text fluency and a preset fluency threshold, filtering the text descriptions in the original multi-modal data set after basic filtering whose quality does not meet the requirements.
[0032] Specifically, first, for the image data, using image sharpness as the quality evaluation index, the pre-trained Noise2Noise denoising network is employed to denoise each image data in the original multi-modal dataset after the basic filtering. By calculating the structural similarity index (SSIM) between the input image data and the image data output by the model after denoising, the noise level of the image data is revealed, thereby evaluating the image sharpness. Furthermore, taking the structural similarity index as the quantitative score of the image sharpness, the image data with a value lower than the preset sharpness threshold is regarded as a low-quality image sample and excluded, while the image data with a value higher than the preset sharpness threshold is regarded as a high-quality image sample and retained. For the text data, using text fluency as the quality evaluation index, a language fluency discriminator is constructed based on the RoBERTa-large model. The input text description is tokenized and sub-word processed (Byte-Pair Encoding, vocabulary size 50,265), and the predicted probability distribution of the next position for each sub-word unit is obtained through the model. Based on the cross-entropy loss (performing the natural exponential function operation on the cross-entropy loss function value), the perplexity of the text description is calculated; the lower the perplexity, the better the fluency of the text description, that is, the language is more natural, coherent, and easy to understand. Furthermore, taking the reciprocal of the perplexity as the quantitative score of the text fluency, the text data with a value lower than the preset fluency threshold is regarded as a low-quality text sample and excluded, while the text data with a value higher than the preset fluency threshold is regarded as a high-quality text sample and retained. In this way, potential low-quality data samples can be further excluded, ensuring that the retained image and text data samples have a high quality level within their respective modalities and meeting the requirements of high-quality training.
[0033] Specifically, in step S23, the multi-modal data samples to be selected are extracted from the original multi-modal dataset after the secondary filtering. Specifically, although single-modal filtering has been performed, the cross-modal semantic mismatch problem in the multi-modal dataset (such as the image content being irrelevant to the text description) may still contaminate the dataset (such as the text illustration of "a dog running" being an actual static cat image), resulting in semantic understanding biases during the training process of the multi-modal large model and affecting the accuracy and generalization ability of the model. Therefore, in order to accurately locate and exclude the multi-modal data samples with semantic mismatches, based on the image-text pairing relationship in the original multi-modal dataset after the secondary filtering, each corresponding pair of image data and text data is regarded as a multi-modal data sample to be selected, focusing on the possible cross-modal alignment problem between the two for semantic consistency verification.
[0034] In the above multi-modal large model data cleaning and governance method, in step S3, visual features of the to-be-selected image data are extracted to obtain a to-be-selected image visual feature encoding vector. In a specific example of the present application, step S3 includes: inputting the to-be-selected image data into a visual encoder based on the ViT model for visual feature extraction to obtain the to-be-selected image visual feature encoding vector. It should be understood that since image data itself is a pixel matrix and cannot be directly used for semantic consistency comparison with text descriptions. Therefore, in order to effectively extract the image content semantics of the to-be-selected image data, the present application is based on the Vision Transformer (ViT) model (ViT-B / 16 architecture, pre-trained on ImageNet-21k). By performing visual feature extraction on the to-be-selected image data, key information in the image, such as objects, scenes, actions, etc., is captured to convert it into a high-dimensional to-be-selected image visual feature encoding vector, providing a basis for subsequent semantic consistency comparison with text descriptions. Specifically, as an image feature extraction model based on the Transformer architecture, the ViT model divides the to-be-selected image data into a series of non-overlapping small patches, linearly embeds each image patch into a vector of a fixed size as the input sequence of the model, and uses the self-attention mechanism of the Transformer architecture to process the sequence of image patch embedding vectors to capture the global dependency relationships between the image patches, thereby generating a deep feature representation of the to-be-selected image data and obtaining the to-be-selected image visual feature encoding vector. Based on this, the to-be-selected image visual feature encoding vector, as a vectorized representation of the semantic connotation of the to-be-selected image data, can accurately describe the core semantic content in the image data, providing strong support for subsequent semantic consistency verification with text descriptions.
[0035] In the above-mentioned multimodal large model data cleaning and management method, the step S4 extracts semantic features from the text description to be selected to obtain a semantic feature encoding vector of the text description to be selected. Similarly, the text description to be selected is a natural language text. In order to convert it into a semantic feature representation that is understandable and operable by a computer, so as to align it with the visual feature space, the present application uses a pre-trained BERT model to perform semantic analysis on the text description to be selected, and extracts the contextual semantic embedding representation of the text description to be selected, so as to obtain a semantic feature encoding vector of the text description to be selected. Specifically, the BERT model performs word segmentation on the input text description, and converts the word segmentation result into a corresponding word embedding vector sequence through a word embedding layer, and uses the self-attention mechanism of the Transformer architecture to perform deep context modeling on the word embedding vector sequence, capture the semantic relationship between words, and mine the entity relationship (such as the action-object association in "boys playing football"), action logic and other fine-grained semantic information in the text description to be selected, thereby generating a semantic feature encoding vector of the text description to be selected. In the above-mentioned multimodal large model data cleaning and management method, the step S5 performs semantic-level fine-grained alignment encoding on the visual feature coding vector of the image to be selected and the semantic feature coding vector of the text description to be selected to obtain the semantic-level fine-grained interactive response coding vector of the image to be selected and the text. It should be understood that the present application takes into account that simple cosine similarity calculation cannot model the fine-grained semantic association pattern of cross-modal features (such as local semantic matching or implicit relationship). Therefore, in order to refine the measurement of the degree of alignment between images and texts (such as the correspondence between the "red car" in the text description and the local red area in the image), the present application proposes a cross-modal semantic-level fine-grained alignment coding method, which performs fine-grained semantic alignment interaction on the visual feature coding vector of the image to be selected and the semantic feature coding vector of the text description to be selected in the local space to mine the local semantic matching relationship and implicit semantic association between the two, and integrates the information of the global semantic space through the inference aggregation mechanism to obtain the semantic-level fine-grained interactive response coding vector of the image to be selected and the text. Among them, Figure 5 FIG. 5 is a flowchart of sub-step S5 of the multimodal large model data cleaning and management method according to an embodiment of the present application. Figure 5As shown, step S5 includes steps: S51, locally segmenting the visual feature encoding vector of the to-be-selected image and the semantic feature encoding vector of the to-be-selected text description to obtain a sequence of ordered encoding vectors of local visual features of the to-be-selected image and a sequence of ordered encoding vectors of local semantic features of the to-be-selected text description; S52, inputting each group of corresponding ordered encoding vectors of local visual features of the to-be-selected image and ordered encoding vectors of local semantic features of the to-be-selected text description in the sequence of ordered encoding vectors of local visual features of the to-be-selected image and the sequence of ordered encoding vectors of local semantic features of the to-be-selected text description into a semantic-level transfer interaction response inference unit to obtain a sequence of local semantic interaction response encoding matrices of the to-be-selected image-text description; S53, performing semantic transfer encoding on the sequence of local semantic interaction response encoding matrices of the to-be-selected image-text description to obtain the fine-grained interaction response encoding vector of the to-be-selected image-text semantic level.
[0036] Specifically, in a specific example of the present application, step S51 includes: First, perform an ordered arrangement based on the eigenvalue size on the visual feature encoding vector of the to-be-selected image and the semantic feature encoding vector of the to-be-selected text description to obtain an ordered arrangement encoding vector of visual features of the to-be-selected image and an ordered arrangement encoding vector of semantic features of the to-be-selected text description, which is expressed by the formula:
[0037]
[0038] Wherein, represents the visual feature encoding vector of the to-be-selected image, represents the semantic feature encoding vector of the to-be-selected text description, represents the sorting operation on vector elements, represents the ordered arrangement encoding vector of visual features of the to-be-selected image, represents the ordered arrangement encoding vector of semantic features of the to-be-selected text description.
[0039] That is, by reconstructing the sequence distribution of feature elements, the implicit interference introduced by the difference in the arrangement order of the visual feature encoding vector of the to-be-selected image and the semantic feature encoding vector of the to-be-selected text description in the original space is eliminated, thereby providing a feature expression basis with numerical distribution consistency for cross-modal alignment. Specifically, through the ordered arrangement, the visual feature encoding vector of the to-be-selected image and the semantic feature encoding vector of the to-be-selected text description are re-sorted according to the intensity of element values, so that a comparable standardized structure is formed in the numerical distribution dimension. The generated ordered arrangement encoding vector of visual features of the to-be-selected image and the ordered arrangement encoding vector of semantic features of the to-be-selected text description reduce the redundant noise of the feature sequence and enhance the robustness of cross-modal local feature matching at the same time.
[0040] Then, perform equal - granularity feature segmentation on the ordered encoding vector of the visual features of the image to be selected and the ordered encoding vector of the semantic features of the text description of the image to be selected, so as to obtain a sequence of ordered encoding vectors of local visual features of the image to be selected and a sequence of ordered encoding vectors of local semantic features of the text description of the image to be selected, which is expressed by the formula:
[0041]
[0042] Among them, 、 、 and respectively represent the 1st, 2nd, the th, and the th ordered encoding vectors of local visual features of the image to be selected in the sequence of ordered encoding vectors of local visual features of the image to be selected, is the number of ordered encoding vectors of local visual features of the image to be selected, 、 、 and respectively represent the 1st, 2nd, the th, and the th ordered encoding vectors of local semantic features of the text description of the image to be selected in the sequence of ordered encoding vectors of local semantic features of the text description of the image to be selected, represents the feature segmentation function.
[0043] That is, by aligning the feature segmentation granularity of the ordered encoding vector of the visual features of the image to be selected and the ordered encoding vector of the semantic features of the text description of the image to be selected, the subsequent cross - modal interaction response inference can capture the potential correlation between the visual local region and the semantic segment at the same structural scale. Specifically, cutting the globally ordered encoding vector of the visual features of the image to be selected and the ordered encoding vector of the semantic features of the text description of the image to be selected into sequences of ordered encoding vectors of local visual features of the image to be selected and sequences of ordered encoding vectors of local semantic features of the text description of the image to be selected with consistent dimensions can eliminate the interference of different modal feature dimension differences on fine - grained alignment, provide fragmented feature inputs with structurally matchable for cross - modal local semantic interaction modeling, and enhance the sensitivity to microscopic semantic coupling relationships.
[0044] Specifically, step S52 is expressed by the formula:
[0045] Among them, represents the linear transformation matrix, denotes the ReLU activation function, denotes matrix multiplication, denotes and the local semantic interaction response encoding matrix of the to-be-selected image-text description.
[0046] That is, by establishing a deep non-linear interaction modeling mechanism for cross-modal local features, the potential coupling relationships and conflicts between visual segments and semantic segments are revealed. The implicit semantic response patterns between the sequence of ordered encoding vectors of the local visual features of the to-be-selected image and the sequence of ordered encoding vectors of the local semantic features of the to-be-selected text description are dynamically modeled using a neural network model, thereby quantifying the internal consistency of the image-text data pair at the fine-grained semantic level. The generated local semantic interaction response encoding matrix of the to-be-selected image-text description can represent the explicit association strength between cross-modal segments, identify imperceptible local semantic deviations, and thus improve the parsing depth and discrimination accuracy of subsequent cross-modal alignment evaluation.
[0047] Specifically, considering that the adjustment of the segmentation granularity of the equal-granularity feature segmentation not only affects the regional segmentation granularity dimension of the local semantic interaction response encoding matrix of the to-be-selected image-text description, but also directly affects the mutual construction expression framework between the sequence of ordered encoding vectors of the local visual features of the to-be-selected image and the sequence of ordered encoding vectors of the local semantic features of the to-be-selected text description. Therefore, in a preferred example of the present application, the step S423 includes: First, based on the semantic-level interaction response connection density between each group of corresponding ordered encoding vectors of the local visual features of the to-be-selected image and the ordered encoding vectors of the local semantic features of the to-be-selected text description, sparsity constraints are imposed on each local semantic interaction response encoding matrix in the sequence of the local semantic interaction response encoding matrix of the to-be-selected image-text description to obtain a sequence of optimized local semantic interaction response encoding matrices of the to-be-selected image-text description.
[0048] Specifically, since the local semantic interaction response encoding matrix of the to-be-selected image-text description is a carrier for spatially mutual construction quantization representation, the norm of its row vectors represents the proximity constraint parameter. If the proximity constraint parameter, that is, the norm of the row vectors, is used as the intensity parameter for low-rank spatial feature quantization, then the low-rank spatial feature quantization of the local semantic interaction response encoding matrix of the to-be-selected image-text description, that is, the F norm should follow the Poisson distribution relationship:
[0049] where denotes factorial, denotes the Frobenius norm, denotes The norm of the row vector represents the exponential function with the natural constant as the base represents the Poisson distribution parameter, which characterizes the expected frequency of the semantic interaction of the image - text description to be selected. Thus, the Poisson distribution parameter can be solved .
[0050] In this way, when the cross - domain collaborative co - construction feedback generates a dynamic process of Poisson distribution with the local interaction constraint strength of and the mean expectation of , the boundary - associated feature expression of the interaction space metric descriptor is described by and , and then the semantic - level interaction response connection density between two local areas is further determined as:
[0051] where represents and the semantic - level interaction response connection density between represents the absolute - value operation represents the two - norm
[0052] Then, the adjacent - scale adjustment parameter is iteratively optimized and adjusted with the semantic - level interaction response connection density:
[0053] where represents the after optimization and adjustment
[0054] Finally, the sparsity of the local semantic interaction response coding matrix of the image - text description to be selected is re - constrained with the optimized and adjusted :
[0055] where represents the corresponding optimized local semantic interaction response coding matrix of the image - text description to be selected
[0056] That is, while maintaining the expected connection degree at , the global co - construction space quantization parameter is dynamically corrected by the regularization constraint driven by probability fluctuations, so that the probability fluctuations of the structured information interaction amount in the cross - domain co - construction system are limited, thereby suppressing the risk of neighborhood overfitting and improving the overall feature expression performance
[0057] Then, input the sequence of the optimized image-text description local semantic interaction response encoding matrix to the LSTM transfer encoding module with an attention mechanism to obtain the image-text semantic-level fine-grained interaction response encoding vector to be selected, which is expressed by the formula:
[0058]
[0059]
[0060] Wherein, represents the matrix flattening operation, represents the flattened image-text description local semantic interaction response encoding vector to be selected, represents the exponential function operation with base e, is the feature modulation function based on the attention mechanism, represents and the optimized image-text description local semantic interaction response encoding matrix between, represents and the optimized image-text description local semantic interaction response encoding matrix between, represents the LSTM model, represents the image-text semantic-level fine-grained interaction response encoding vector to be selected.
[0061] That is, by using the long-range dependence modeling ability of LSTM, capture the evolution law of the cross-region semantic association of the optimized image-text description local semantic interaction response encoding matrix, and at the same time, with the help of the attention mechanism, dynamically strengthen the weight assignment of key local interactions, so as to reconstruct the composite feature expression of multi-granularity semantic alignment at the global level. In this way, the generated image-text semantic-level fine-grained interaction response encoding vector not only fuses the semantic association topological structure across local regions, but also highlights the discriminative interaction features through attention-guided context encoding, providing a quantitative basis with both integrity and sensitivity for the filtering decision in data cleaning.
[0062] In the above multi-modal large model data cleaning and governance method, in step S6, based on the image-text semantic-level fine-grained interaction response encoding vector to be selected, determine whether to filter the multi-modal data sample to be selected. Wherein, Figure 6 is the flow chart of sub-step S6 of the multi-modal large model data cleaning and governance method according to the embodiment of the present application. As Figure 6As shown, step S6 includes steps: S61, performing feature decoding on the encoded vector of the fine-grained interaction response at the semantic level of the image-text to be selected to obtain a semantic-level alignment coefficient; S62, based on the comparison between the semantic-level alignment coefficient and a preset semantic alignment threshold, determining whether to filter the multi-modal data sample to be selected.
[0063] Specifically, in step S61, feature decoding is performed on the encoded vector of the fine-grained interaction response at the semantic level of the image-text to be selected to obtain a semantic-level alignment coefficient. It should be understood that the encoded vector of the fine-grained interaction response at the semantic level of the image-text to be selected, as a deep modeling representation of the implicit semantic relationship between the image data to be selected and the text description to be selected, contains the matching degree and interaction information between the two at the fine-grained semantic level. In order to further transform the high-dimensional semantic interaction features into interpretable alignment scores to support subsequent filtering decisions, the present application, based on a regression task framework, realizes the mapping from features to scalars through a multi-layer perceptron (MLP). Specifically, in the training stage of the multi-layer perceptron model, a contrastive learning strategy is used: positive samples (manually annotated high-quality image-text pairs) and negative samples (noise pairs with randomly replaced text, with a ratio of 1:3) are constructed, and the model is optimized through triplet loss (interval α = 0.2) to make the output for positive samples close to 1 and negative samples close to 0. At the same time, label smoothing (ε = 0.1) is introduced to alleviate overfitting. In practical applications, the encoded vector of the fine-grained interaction response at the semantic level of the image-text to be selected is input into an MLP model with two hidden layers (using the GeLU activation function). Through the non-linear transformation of the hidden layers, the high-dimensional encoded vector of the fine-grained interaction response at the semantic level of the image-text to be selected is gradually reduced in dimension, and the sigmoid function is used to constrain the final output to a semantic alignment coefficient with a value range of [0,1], so as to reflect the degree of semantic consistency between images and texts.
[0064] Specifically, in step S62, based on the comparison between the semantic-level alignment coefficient and a preset semantic alignment threshold, determining whether to filter the multi-modal data sample to be selected. That is, the semantic-level alignment coefficient is compared with a preset semantic alignment threshold (such as 0.7). If the semantic-level alignment coefficient is greater than or equal to the threshold, it is considered that the semantic alignment degree between the image sample and the text description sample in the multi-modal data sample to be selected is relatively high, and it is retained; otherwise, if the semantic-level alignment coefficient is less than the threshold, this group of data samples is filtered out. In this way, multi-modal data samples with a high degree of semantic alignment can be effectively screened out, reducing the interference of low-quality and semantically mismatched data on model training, thereby improving the accuracy and stability of the model in multi-modal alignment tasks, enabling the model to better learn cross-modal semantic associations, and enhancing the semantic understanding ability of the model.
[0065] In summary, the multi-modal large model data cleaning and governance method based on the embodiments of the present application is elucidated. After performing basic format filtering on the original multi-modal data set, the image clarity and text fluency in the multi-modal data set are respectively quantitatively evaluated through a single-modal quality assessment mechanism to screen out qualified image and text data samples. Furthermore, by introducing a semantic-level alignment assessment mechanism, semantic-level interactive response analysis is performed on each corresponding image sample and image text description in the data set to quantitatively evaluate the semantic alignment degree between the image sample and the text description, thereby further screening out highly semantically aligned image-text data pairs. Through multi-level cleaning and governance of the multi-modal data set, this method can ensure that the multi-modal training samples not only meet the quality standards but also achieve a high degree of alignment at the cross-modal semantic level, thus effectively improving the accuracy and robustness of the multi-modal large model in cross-modal understanding and generation tasks.
[0066] Furthermore, a multi-modal large model data cleaning and governance system is also provided.
[0067] Figure 7 The block diagram of the multi-modal large model data cleaning and governance system according to the embodiments of the present application is as follows. As Figure 7 shown, the multi-modal large model data cleaning and governance system 100 according to the embodiments of the present application includes: a multi-modal data acquisition module 110 for acquiring an original multi-modal data set; a multi-modal data preprocessing module 120 for performing initial data cleaning on the original multi-modal data set and extracting multi-modal data samples to be refined therefrom, where the multi-modal data samples to be refined include image data to be refined and corresponding text descriptions to be refined of the image data to be refined; a visual feature extraction module 130 for extracting visual features from the image data to be refined to obtain a visual feature encoding vector of the image data to be refined; a semantic feature extraction module 140 for extracting semantic features from the text descriptions to be refined to obtain a semantic feature encoding vector of the text descriptions to be refined; a fine-grained alignment encoding module 150 for performing semantic-level fine-grained alignment encoding on the visual feature encoding vector of the image data to be refined and the semantic feature encoding vector of the text descriptions to be refined to obtain a fine-grained interactive response encoding vector at the semantic level of the image-text to be refined; and a sample screening module 160 for determining whether to filter the multi-modal data samples to be refined based on the fine-grained interactive response encoding vector at the semantic level of the image-text to be refined.
[0068] Here, those skilled in the art can understand that the specific operations of each module in the above multi-modal large model data cleaning and governance system have been introduced in detail in the description of the multi-modal large model data cleaning and governance method above with reference to Figures 1 to 6 Therefore, the repeated description thereof will be omitted.
[0069] The basic principles of the present invention have been described above in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present invention. Additionally, the specific details of the above embodiments are only for the purposes of illustration and facilitating understanding, rather than limitations. The above details do not limit the present invention to necessarily adopting the above specific details for implementation.
[0070] In the above embodiments, the descriptions of each embodiment have their own focuses. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the unit division is only a logical function division, and there can be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0071] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.
[0072] In addition, obviously the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units stated in the system claims can also be implemented by one unit through software or hardware.
[0073] Finally, it should be noted that the above description has been given for the purposes of illustration and description. In addition, the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the technical solutions are modified or equivalently replaced with reference to the preferred embodiments, the spirit and scope of the technical solutions of the present invention are not departed from.
Claims
1. A method for cleaning and governing multi-modal large model data, characterized in that, Including: Obtain the original multi-modal dataset; After initially cleaning the original multi-modal dataset, extract the multi-modal data samples to be refined therefrom, where the multi-modal data samples to be refined include the image data to be refined and the corresponding text description to be refined for the image data to be refined; Extract visual features from the image data to be refined to obtain the visual feature encoding vector of the image to be refined; Extract semantic features from the text description to be refined to obtain the semantic feature encoding vector of the text description to be refined; Perform semantic-level fine-grained alignment encoding on the visual feature encoding vector of the image to be refined and the semantic feature encoding vector of the text description to be refined to obtain the semantic-level fine-grained interaction response encoding vector of the image-text to be refined; Based on the semantic-level fine-grained interaction response encoding vector of the image-text to be refined, determine whether to filter the multi-modal data samples to be refined.
2. The multimodal large model data cleaning and governance method according to claim 1, wherein, After initially cleaning the original multi-modal dataset, extract the multi-modal data samples to be refined, including: Perform basic filtering on the original multi-modal dataset to obtain the original multi-modal dataset after basic filtering; Perform single-modal quality assessment on the original multi-modal dataset after basic filtering to obtain the original multi-modal dataset after secondary filtering; Extract the multi-modal data samples to be refined from the original multi-modal dataset after secondary filtering.
3. The multimodal large model data cleaning and governance method according to claim 2, wherein Perform basic filtering on the original multi-modal dataset to obtain the original multi-modal dataset after basic filtering, including: Remove the obviously invalid data samples in the original multi-modal dataset, where the obviously invalid data samples include unloadable images, texts that are too short / too long, and specific format errors.
4. The multimodal large model data cleaning and governance method according to claim 3, characterized in that, Perform single-modal quality assessment on the original multi-modal dataset after basic filtering to obtain the original multi-modal dataset after secondary filtering, including: Use a pre-trained model to evaluate the image clarity of each image data and the text fluency of each text description in the original multi-modal dataset after basic filtering; Based on the comparison between the image clarity and the preset clarity threshold, filter the image data with unqualified quality in the original multi-modal dataset; Based on the comparison between the text fluency and the preset fluency threshold, filter the text descriptions with unqualified quality in the original multi-modal dataset after basic filtering.
5. The multimodal large model data cleaning and governance method according to claim 4, wherein Extract visual features from the image data to be refined to obtain the visual feature encoding vector of the image to be refined, including: Input the image data to be refined into a visual encoder based on the ViT model for visual feature extraction to obtain the visual feature encoding vector of the image to be refined.
6. The multimodal large model data cleaning and governance method according to claim 5, wherein, Perform semantic-level fine-grained alignment encoding on the visual feature encoding vector of the image to be refined and the semantic feature encoding vector of the text description to be refined to obtain the semantic-level fine-grained interaction response encoding vector of the image-text to be refined, including: Perform local segmentation on the visual feature encoding vector of the image to be refined and the semantic feature encoding vector of the text description to be refined to obtain a sequence of local visual feature ordered encoding vectors of the image to be refined and a sequence of local semantic feature ordered encoding vectors of the text description to be refined; Input the sequence of ordered encoded vectors of local visual features of the to-be-selected images and the sequence of ordered encoded vectors of local semantic features of the to-be-selected text descriptions, where each corresponding pair of ordered encoded vectors of local visual features of the to-be-selected images and ordered encoded vectors of local semantic features of the to-be-selected text descriptions are input into the semantic-level transfer interaction response inference unit to obtain a sequence of local semantic interaction response encoding matrices of the to-be-selected image-text descriptions; Perform semantic transfer encoding on the sequence of local semantic interaction response encoding matrices of the to-be-selected image-text descriptions to obtain the fine-grained interaction response encoding vectors at the semantic level of the to-be-selected image-text.
7. The multimodal large model data cleaning and governance method according to claim 6, wherein, Perform local segmentation on the encoded vectors of the visual features of the to-be-selected images and the encoded vectors of the semantic features of the to-be-selected text descriptions to obtain a sequence of ordered encoded vectors of local visual features of the to-be-selected images and a sequence of ordered encoded vectors of local semantic features of the to-be-selected text descriptions, including: Perform an ordered arrangement based on the eigenvalue magnitudes on the encoded vectors of the visual features of the to-be-selected images and the encoded vectors of the semantic features of the to-be-selected text descriptions to obtain an ordered arrangement encoded vector of the visual features of the to-be-selected images and an ordered arrangement encoded vector of the semantic features of the to-be-selected text descriptions; Perform equal-grained feature segmentation on the ordered arrangement encoded vector of the visual features of the to-be-selected images and the ordered arrangement encoded vector of the semantic features of the to-be-selected text descriptions to obtain a sequence of ordered encoded vectors of local visual features of the to-be-selected images and a sequence of ordered encoded vectors of local semantic features of the to-be-selected text descriptions.
8. The multi-modal large model data cleaning and governance method according to claim 7, wherein Perform semantic transfer encoding on the sequence of local semantic interaction response encoding matrices of the to-be-selected image-text descriptions to obtain the fine-grained interaction response encoding vectors at the semantic level of the to-be-selected image-text, including: Based on the semantic-level interaction response connection density between each corresponding pair of ordered encoded vectors of local visual features of the to-be-selected images and ordered encoded vectors of local semantic features of the to-be-selected text descriptions, perform sparsity constraints on each local semantic interaction response encoding matrix in the sequence of local semantic interaction response encoding matrices of the to-be-selected image-text descriptions to obtain a sequence of optimized local semantic interaction response encoding matrices of the to-be-selected image-text descriptions; Input the sequence of optimized local semantic interaction response encoding matrices of the to-be-selected image-text descriptions into the LSTM transfer encoding module with an attention mechanism to obtain the fine-grained interaction response encoding vectors at the semantic level of the to-be-selected image-text.
9. The multimodal large model data cleaning and governance method according to claim 8, characterized in that Based on the fine-grained interaction response encoding vectors at the semantic level of the to-be-selected image-text, determine whether to filter the to-be-selected multimodal data samples, including: Perform feature decoding on the fine-grained interaction response encoding vectors at the semantic level of the to-be-selected image-text to obtain semantic-level alignment coefficients; Based on the comparison between the semantic-level alignment coefficients and a preset semantic alignment threshold, determine whether to filter the to-be-selected multimodal data samples.
10. A multi-modal large model data cleaning and governance system, characterized in that, Including: A multimodal data acquisition module for acquiring an original multimodal data set; A multi-modal data preprocessing module, which is used to perform initial data cleaning on the original multi-modal data set and extract multi-modal data samples to be selected therefrom. The multi-modal data samples to be selected include image data to be selected and corresponding text descriptions to be selected for the image data to be selected; A visual feature extraction module, which is used to extract visual features from the image data to be selected to obtain a visual feature encoding vector of the image to be selected; A semantic feature extraction module, which is used to extract semantic features from the text description to be selected to obtain a semantic feature encoding vector of the text description to be selected; A fine-grained alignment encoding module, which is used to perform semantic-level fine-grained alignment encoding on the visual feature encoding vector of the image to be selected and the semantic feature encoding vector of the text description to be selected to obtain a fine-grained interaction response encoding vector of the image-text semantic level to be selected; A sample screening module, which is used to determine whether to filter the multi-modal data samples to be selected based on the fine-grained interaction response encoding vector of the image-text semantic level to be selected.
Citation Information
Patent Citations
Large model multi-modal data semantic representation alignment method
CN119380341A
Multi-modal model training fine tuning optimization method based on artificial intelligence
CN119442128A
Single stream multi-level alignment for vision-language pretraining
US20230281963A1
Cited By
Model training method, computer equipment and storage medium
CN120780820A
Model training method, computer device and storage medium
CN120780820B
Multi-modal data quality evaluation method based on deep learning
CN120804084A