Multi-modal large model data cleaning and feature enhancement preprocessing method and system

By classifying and cleaning and feature enhancement of multimodal large model data, combining field information and backtracking threshold judgment, the problem of insufficient dynamic adaptability in the existing technology is solved, efficient and reliable data processing is achieved, and data quality and user experience are improved.

CN120277502AActive Publication Date: 2025-07-08BEIJING YUNLIAN JINHUI DIGITAL TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510765051.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

In the process of data cleaning and feature enhancement of multimodal large models, it is difficult for the prior art to dynamically adapt to complex related scenarios, resulting in incomplete cross-modal coupled noise processing, generating samples that deviate from the real scene, and lack controllability, which reduces the efficiency and quality of data processing.

Method used

Through the methods of information acquisition, classification, data processing, judgment backtracking and feedback generation, combined with the domain information of the data, the data of the multimodal large model is classified and cleaned and feature enhanced, and a backtracking threshold is set to judge the data processing status, and feedback information is generated to provide a channel for reclassification and processing.

Benefits of technology

It improves the dynamic adaptability of data cleaning and feature enhancement, reduces the chance of enhanced deviation and semantic cleavage, improves the efficiency and quality of data processing, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277502A_ABST
    Figure CN120277502A_ABST
Patent Text Reader

Abstract

The invention discloses a preprocessing method and system for data cleaning and feature enhancement of a multi-modal large model, and relates to the technical field of data processing. Comprising the following steps: information acquisition: acquiring data required to be used by a multi-modal large model to obtain to-be-processed data; according to the method, the data to be processed is subjected to corresponding cleaning processing according to the classification result through the set data processing method, the completed data is subjected to corresponding enhancement processing according to the classification result through the data enhancement method, and in the data cleaning and feature enhancement process, the data is subjected to corresponding cleaning and enhancement in combination with the field information of the data, so that the data cleaning and feature enhancement efficiency is improved. The method is advantaged in that the method can adapt to data of different scenes, dynamic adaptability of data cleaning and feature enhancement is improved, probability of enhancement deviation and semantic segmentation is reduced, an effect of improving data processing efficiency and quality is realized, the range of classification results can be determined according to user demands, and user experience is improved. It can be determined whether to pursue precision or efficiency in the data cleaning and feature enhancement process, so that the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and specifically to a preprocessing method and system for multi-modal large model data cleaning and feature enhancement. Background Art

[0002] A multi-modal large model refers to an artificial intelligence model that can simultaneously process and fuse multiple data types. Before training a multi-modal large model, it is necessary to clean and enhance the features of the data used for training the multi-modal large model to improve the efficiency and quality of multi-modal large model training.

[0003] A method for pre-training data cleaning and balancing of a multi-modal large model with the patent publication number CN117932219A includes the following steps: S01, extracting features from the images and texts in the image-text pairs; S02, calculating the cosine similarity of the same image-text pair, and deleting the image-text pairs with a cosine similarity less than the similarity threshold; S03, segmenting the texts in the remaining image-text pairs to obtain text segments; S04, counting the number of times each entry in the metadata appears in the text segments as the word frequency quantity, and recording the positions of the text segments in the metadata; S05, calculating the word frequency probabilities of each entry; S06, generating a random probability p, indexing the corresponding word frequency probabilities in turn according to the positions of each text segment in the metadata in each pair of image-text pairs to be balanced. When all the corresponding word frequency probabilities are greater than p, the image-text pair is retained, otherwise the image-text pair is removed. The finally retained data is the pre-training data after cleaning and balancing; The methods based on the above and similar principles rely on artificial annotation rules or traditional algorithms, and it is difficult to dynamically adapt to the complex association scenarios of multi-modal data. Traditional cleaning methods (such as filtering and redundant filtering) do not thoroughly handle cross-modal coupling noise, and the rules overly rely on specific data sets and are difficult to migrate to new scenarios. As a result, the methods based on the above and similar principles have the disadvantages of insufficient dynamic adaptability and limited generation reliability. In the process of feature enhancement, the existing technology often uses enhancement based on GAN or diffusion models. This enhancement may generate samples that deviate from the real scenario, leading to model overfitting. At the same time, cross-modal generation lacks controllable constraints and is prone to semantic fragmentation, and it is necessary to rely on post-processing rules for correction, which will reduce the efficiency of the feature enhancement process and is not conducive to improving the overall data processing efficiency. Therefore, the present invention is proposed. Summary of the Invention

[0004] The purpose of the present invention is to provide a preprocessing method and system for multi-modal large model data cleaning and feature enhancement to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions: A preprocessing method for multi-modal large model data cleaning and feature enhancement includes: Information acquisition: Obtain the data required for the multi-modal large model to get the data to be processed, obtain the user requirements and feedback address information, and the feedback address information is the address information for sending messages to the user; Information classification: Use the classification method to divide the data to be processed into domain categories to obtain the classification result; Classification processing: The classification result includes several sub-results. Establish the association relationship between the sub-results and the data to be processed. Based on the association relationship, extract the data to be processed corresponding to the sub-results to get the sub-data. Establish a sub-library named after the sub-result to store the sub-data, and establish a storage library to store the sub-libraries; It is characterized by including: Data processing: Perform data cleaning and feature enhancement on the sub-data. Use the data processing method to perform corresponding cleaning processing on the data with different classification results to obtain the cleaning result. The cleaning result includes completed data and uncompleted data. Based on the association relationship, extract the classification result of the completed data to get the target result. Based on the target result and the completed data, use the data enhancement method to perform corresponding enhancement processing on the data with different classification results to obtain the enhancement result. The enhancement result includes enhanced data and unenhanced data; Judgment and backtracking: Preset the backtracking threshold. Based on the uncompleted data, unenhanced data, and classification results, use the backtracking method to reclassify and perform data cleaning and feature enhancement to obtain the reclassification result. Record the number of times the same data is backtracked in the backtracking method to get the backtracking times; Feedback generation: Judge the relationship between the backtracking times and the backtracking threshold. When the backtracking times exceed the backtracking threshold, extract the data at the backtracking times to get the feedback information, and send the feedback information to the user based on the feedback address information.

[0006] Furthermore, the data processing method includes: Preset the processing format, perform normalization processing on the sub-data based on the processing format to obtain the preliminary data, split the preliminary data to obtain several data groups, judge the number of individual data in the data group to get the target number, extract the data group with the target number as a single data to get the first data group, extract the data group with the target number as multiple data to get the second data group, perform monitoring on the single data group in the first data group using the detection method to obtain the detection result, extract the single data group with the detection result feedback as successful detection to get the first completed data, extract the single data group with the detection result feedback as failed detection to get the first uncompleted data, perform verification on the single data group in the second data group using the cross-validation method to obtain the verification result, extract the single data group with the verification result feedback as successful verification to get the second completed data, extract the single data group with the verification result feedback as failed verification to get the second uncompleted data, integrate the first completed data and the second completed data to get the completed data, and integrate the first uncompleted data and the second uncompleted data to get the uncompleted data.

[0007] Furthermore, the detection method includes: obtaining a sub-result of a first data group to obtain a first result, obtaining a single data group from the first data group to obtain a data group to be detected, splitting the data group to be detected to obtain individual pieces of information to be detected, obtaining domain information of the first result to obtain reference information, performing anomaly detection and correction on the information to be detected based on the reference information to obtain a detection result and the processed information to be detected, obtaining first completed data and first uncompleted data based on the detection result, when the detection result feedback indicates that the processed information to be detected has no anomaly, extracting the processed information to be detected to obtain the first completed data, and when the detection result feedback indicates that the processed information to be detected has an anomaly, extracting the processed information to be detected to obtain the first uncompleted data.

[0008] Furthermore, the process of performing anomaly detection and correction on the information to be detected to obtain a detection result and a correction result is as follows: performing paragraphing on the information to be detected to obtain a number of sub-detection information and processing positions, comparing the basic formats of the sub-detection information in combination with the reference information and extracting the sub-detection information with incorrect basic formats as anomaly values, correcting the anomaly values based on the reference information to obtain corrected sub-detection information, integrating the sub-detection information based on the processing positions to obtain the processed information to be detected, extracting the features in the processed information to obtain specific features, judging the relationship between the specific features and the reference information to obtain a target relationship, obtaining a detection result based on the target relationship, when the target relationship feedback indicates that the specific features belong to the reference information, the detection result is that the processed information to be detected has no anomaly, and extracting the processed information to be detected into the first completed data, and when the target relationship feedback indicates that the specific features do not belong to the reference information, the detection result is that the processed information to be detected has an anomaly, and extracting the processed information to be detected to obtain the first uncompleted data.

[0009] Furthermore, the cross-validation method includes: obtaining a sub-result of a second data group to obtain a second result, obtaining a single data group from the second data group to obtain a data group to be verified, splitting the data group to be verified to obtain a number of sub-information, obtaining the features of the number of sub-information to obtain target features, judging whether the number of target features is consistent to obtain a judgment result, when the judgment result feedback indicates that the number of target features is consistent and belongs to the second result, extracting the verification data group to obtain the second completed data, and when the judgment result feedback does not indicate that the number of target features is consistent and belongs to the second result, extracting the verification data group to obtain the second uncompleted data.

[0010] Furthermore, the data augmentation method includes: recording the features identified during the data cleaning process of the completed data to obtain a number of associated features, splitting the completed data to obtain sub-completed data, establishing the corresponding relationship between the associated features and the sub-completed data, performing feature derivation based on the associated features and the target result to obtain derived features, adding the derived features to the sub-completed data corresponding to the associated features based on the corresponding relationship to obtain augmented data, and extracting the sub-completed data that has not been added with derived features to obtain un-augmented data.

[0011] Furthermore, the classification method includes: obtaining the categories of the past data to be processed to obtain past categories, further obtaining category information based on the past categories to obtain filled categories, integrating the past categories and the filled categories to obtain a category set, determining the scope size based on the user's needs, determining the classified categories in the category set based on the scope size to obtain category criteria, splitting the data to be processed to obtain sub-processed data, matching the sub-processed data based on the category criteria to obtain sub-results, and integrating all the sub-results to obtain a classification result.

[0012] Furthermore, the backtracking method includes: reclassifying the incomplete data through the classification method to obtain a reclassification result, where the reclassification result is different from the classification result, obtaining a re-cleaning result through the data processing method based on the reclassification result and the incomplete data until the re-cleaning result does not contain incomplete data, and obtaining a re-augmentation result through the data augmentation method for the completed data in the re-cleaning result until the re-augmentation result does not contain un-augmented data.

[0013] The preprocessing system for multi-modal large model data cleaning and feature augmentation uses the above-mentioned preprocessing method for multi-modal large model data cleaning and feature augmentation.

[0014] Compared with the prior art, the beneficial effects of the present invention are: The preprocessing method and system for multi-modal large model data cleaning and feature augmentation perform corresponding cleaning processing on the data to be processed according to the classification result through the set data processing method, and perform corresponding augmentation processing on the completed data according to the classification result through the set data augmentation method. That is, during the data cleaning and feature augmentation process, combined with the domain information of the data, corresponding cleaning and augmentation are performed on the data to adapt to data in different scenarios, improve the dynamic adaptability of data cleaning and feature augmentation, and reduce the probability of augmentation deviation and semantic fragmentation, achieving the effect of improving the data processing efficiency and quality. Moreover, the scope size of the classification result can be determined according to the user's needs, and it can be decided whether to pursue accuracy or efficiency during the data cleaning and feature augmentation process to improve the user experience.

[0015] Meanwhile, it is determined whether the data is in a processable state by judging the relationship between the number of backtracks and the backtrack threshold. When the number of backtracks exceeds the backtrack threshold, it is determined that the data is in an unprocessable state, and then feedback information is generated and sent to the user for processing. Through the set backtracking method, a channel for reclassifying and cleaning and enhancing the data that has not been completely cleaned and enhanced is provided. Specifically, the process of re-cleaning and enhancing is the same as before, which is beneficial to improving the efficiency of data processing and reducing the number of times staff intervene in data processing.

[0016] Meanwhile, the information to be detected is divided into paragraphs to shorten the length of a single piece of data and improve the efficiency of anomaly detection and correction. The specific length of the paragraph division is formulated according to the actual usage situation. By analyzing whether the specific features belong to the reference information, it is determined whether there are anomalies in the information to be detected, and the information to be detected is classified and processed according to whether there are anomalies. The information to be detected with anomalies is reclassified and data cleaning processing is performed to improve the accuracy of data cleaning and data enhancement by the data processing method and the data enhancement method. Brief Description of the Drawings

[0017] Figure 1 It is a schematic diagram of the main process structure of the present invention; Figure 2 It is a schematic diagram of the feedback information generation structure of the present invention; Figure 3 It is a schematic diagram of the data processing method structure of the present invention; Figure 4 It is a schematic diagram of the third embodiment structure of the present invention; Figure 5 It is a schematic diagram of the classification method structure of the present invention. Detailed Embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] The training data types of the multi-modal large model include various different data types, and the various data types include text, images, audio, video, etc. The core of the multi-modal large model lies in realizing the in-depth understanding and generation of cross-modal information by jointly learning the features of different modal data. For example, generating images according to text, answering questions by combining speech and images, etc. By performing data cleaning and feature enhancement on the training data, the quality of model training is improved. The high-quality data after cleaning can reduce training bias, and the enhanced features can improve the model's adaptability to complex tasks.

[0020] As Figures 1 - 5 shown, the present invention provides a technical solution: a preprocessing method for multi-modal large model data cleaning and feature enhancement, including: Information acquisition: Obtain the data required to be used by the multi-modal large model to obtain the data to be processed, obtain the user requirements and feedback address information, and the feedback address information is the address information for sending information to the user; It should be noted that the data to be processed is the training data used by the multi-modal large model during training, the user requirements are the demand processing information of the user, which can be specifically obtained by the user himself, and the feedback address information is the address information, which can specifically be the user's email, phone number, IP address, etc.

[0021] Information classification: Use a classification method to classify the data to be processed to obtain a classification result; It should be noted that before performing data cleaning and feature enhancement on the data to be processed, classify the data to be processed and perform corresponding processing according to the different categories of the data to be processed to improve the accuracy and efficiency of processing.

[0022] Classification processing: The classification result includes several sub-results. Establish an association relationship between the sub-results and the data to be processed, extract the data to be processed corresponding to the sub-results based on the association relationship to obtain sub-data, establish a sub-library named after the sub-results to store the sub-data, and establish a storage library to store the sub-library; It should be noted that the sub-results in the classification result are the category information of the sub-data. By establishing an association relationship, it is convenient to determine the corresponding category of the sub-data through the association relationship. By establishing a sub-library named after the sub-results to store the sub-data, it is convenient to store data of the same category, so as to facilitate subsequent data cleaning and feature enhancement of the corresponding category.

[0023] Embodiment 1: The core object of multi-modal large model preprocessing is training data, including multi-modal raw data such as text, images, audio, and video. These data usually come from various sources (such as web pages, books, sensors, etc.) and have complex formats (such as HTML, PDF, EPUB, audio waveforms, etc.). Before use, they need to be uniformly structured, an information sending and receiving interface is established, the user requirements are obtained by the user sending demand information to the information sending and receiving interface, the feedback address information is obtained according to the address information of the user sending the information, and the size of the classification category interval is determined. Specifically, all category intervals need to be in the same standard. For example, the sub-results are traffic category, medical category, etc., and the specific sub-results can also be road category, medical imaging category, etc. The data to be processed is classified according to the determined categories to obtain a classification result.

[0024] It is characterized in that it includes: Data processing: Sub-data is subjected to data cleaning and feature enhancement. Through data processing methods, corresponding cleaning is performed on data with different classification results to obtain a cleaning result, which includes completed data and uncompleted data. Based on the association relationship, the classification results of the completed data are extracted to obtain a target result. Based on the target result and in cooperation with the completed data, corresponding enhancement processing is performed on data with different classification results through data enhancement methods to obtain an enhancement result, which includes enhanced data and unenhanced data; It should be noted that through the set data processing method, corresponding cleaning is performed on the data to be processed according to the classification results, and through the set data enhancement method, corresponding enhancement is performed on the completed data according to the classification results. In this way, during the process of data cleaning and feature enhancement, combined with the domain information of the data, corresponding cleaning and enhancement of the data can be carried out to adapt to data in different scenarios, improve the dynamic adaptability of data cleaning and feature enhancement, and reduce the probability of enhancement deviation and semantic fragmentation, achieving the effect of improving the efficiency and quality of data processing.

[0025] Judgment backtracking: A backtracking threshold is preset. Based on the uncompleted data, unenhanced data, and classification results, reclassification is performed through the backtracking method, and data cleaning and feature enhancement are carried out to obtain a reclassification result. The number of times the same data is backtracked in the backtracking method is recorded to obtain the backtracking times; It should be noted that the backtracking threshold is specific quantitative information. Through the set backtracking method, a backtracking function is provided to reclassify the data that has not been successfully subjected to data cleaning and feature enhancement and then perform data cleaning and feature enhancement, so as to reduce the amount of data processed manually and improve the efficiency of data processing.

[0026] Feedback generation: The relationship between the backtracking times and the backtracking threshold is judged. When the backtracking times exceed the backtracking threshold, the data at the backtracking times is extracted to obtain feedback information, and the feedback information is sent to the user based on the feedback address information.

[0027] It should be noted that by judging the relationship between the backtracking times and the backtracking threshold, it is judged whether the data is in a processable state. When the backtracking times exceed the backtracking threshold, it is determined that the data is in an unprocessable state, and then feedback information is generated and sent to the user for processing.

[0028] Example 2: Set the backtracking threshold to 3. When the number of backtracking times is 4, the number of backtracking times exceeds the backtracking threshold, and this data will be extracted to generate feedback information and fed back to the user. When the number of backtracking times is 1, the data will be re-classified, data-cleaned, and feature-enhanced. During the data cleaning process, the corresponding data cleaning model and data processing method in the existing technology can be used to clean the data to improve the efficiency of data cleaning. Similarly, during the feature enhancement process, the corresponding feature enhancement model and data enhancement method in the existing technology can be used to enhance the data to improve the efficiency of data enhancement. Specifically, after data cleaning and feature enhancement are completed, the sampling quantity and sampling interval can be preset. The sampling quantity can be set to 10 pieces of data, and the sampling interval can be set to every 30 minutes. Sampling the processed data according to the sampling quantity to obtain sampling data, and judging whether the sampling data meets the required status to obtain a judgment criterion. According to the judgment criterion, the corresponding data cleaning model and data enhancement model are adjusted and optimized accordingly to complete improving the accuracy and efficiency of data processing.

[0029] As Figure 1 and Figure 3 shown, the data processing method includes: presetting a processing format, normalizing sub-data based on the processing format to obtain preliminary data, splitting the preliminary data into several data groups, judging the number of individual data in the data groups to obtain a target number, extracting the data groups with the target number being individual to obtain the first data group, extracting the data groups with the target number being multiple to obtain the second data group, monitoring the first data group based on the individual data groups in the first data group by a detection method to obtain a detection result, extracting the individual data groups with the detection result being feedback as detection success to obtain the first completed data, extracting the individual data groups with the detection result being feedback as detection failure to obtain the first uncompleted data, verifying the individual data groups in the second data group by a cross-validation method to obtain a verification result, extracting the individual data groups with the verification result being feedback as verification success to obtain the second completed data, extracting the individual data groups with the verification result being feedback as verification failure to obtain the second uncompleted data, integrating the first completed data and the second completed data to obtain the completed data, and integrating the first uncompleted data and the second uncompleted data to obtain the uncompleted data.

[0030] It should be noted that the processing format is the standard format for data training of multi-modal large models, which can be obtained from the administrator specifically. Normalize the sub-data according to the processing format to improve the efficiency of subsequent processing methods. The multi-modal large model data has different characteristics, which can be specifically manifested as that a set of data may contain two or more pieces of data such as images and texts, or it can also be manifested as that a set of data only contains a single image or text data. By analyzing the number of data in the data group and performing data cleaning separately according to the number of data, the efficiency of data processing can be improved. For the data group with multiple pieces of data, judge whether the data expressions are consistent through cross-validation. For the data group with a single piece of data, detect whether the format information of the data conforms to its field and correct it. Specifically, in the actual use process, the detection method can also be added to the cross-validation method to detect and correct outliers for the data group with multiple pieces of data, so as to improve the operation efficiency of the cross-validation method. Figure 3 where a is the processing flow of the first data group. Figure 3 where b is the processing flow of the second data group.

[0031] As Figure 3 shown, the detection method includes: obtaining the first result by obtaining the sub-result of the first data group, obtaining the single data group in the first data group to get the data group to be detected, splitting the data group to be detected to get the single information to be detected, obtaining the domain information of the first result to get the reference information, performing outlier detection and correction on the information to be detected based on the reference information to get the detection result and the processed information to be detected, obtaining the first completed data and the first uncompleted data based on the detection result. When the detection result feedbacks that there is no abnormality in the processed information to be detected, extract the processed information to be detected to get the first completed data. When the detection result feedbacks that there is an abnormality in the processed information to be detected, extract the processed information to be detected to get the first uncompleted data.

[0032] It should be noted that the process of obtaining the domain information of the first result to get the reference information, that is, obtaining the domain information of the classification result. Specifically, when the first result is of the road category, the reference information is the domain information of the road category, which can be specifically the common sense information and specific information of the road category. By performing outlier detection and correction on the information to be detected according to the reference information, the accuracy of data processing can be improved to improve the dynamic adaptability of data processing.

[0033] The process of performing anomaly detection and correction on the information to be detected to obtain the detection result and the correction result is as follows: The information to be detected is segmented to obtain a number of sub-detection information and processing positions. Based on the sub-detection information and in combination with the reference information, a basic format comparison is performed, and the sub-detection information with incorrect basic format is extracted to obtain outliers. Based on the reference information, the outliers are corrected to obtain the corrected sub-detection information. Based on the processing positions, the sub-detection information is integrated to obtain the processed information to be detected. The features in the processed information are extracted to obtain specific features. The relationship between the specific features and the reference information is judged to obtain the target relationship. Based on the target relationship, the detection result is obtained. When the target relationship feedback is that the specific feature belongs to the reference information, the detection result is that the processed information to be detected has no anomaly, and the processed information to be detected is extracted to the first completed data. When the target relationship feedback is that the specific feature does not belong to the reference information, the detection result is that the processed information to be detected has an anomaly, and the processed information to be detected is extracted to obtain the first incomplete data.

[0034] It should be noted that the process of performing anomaly detection and correction on the information to be detected is carried out in combination with the reference information of the category of the information to be detected. Specifically, anomaly detection and correction can be performed by establishing an anomaly detection and correction model. When initially performing anomaly detection and correction on the information to be detected, the information to be detected is segmented to shorten the length of a single piece of data and improve the efficiency of anomaly detection and correction. The specific length of the segmentation is formulated according to the actual usage situation. Whether the information to be detected has an anomaly is obtained by analyzing whether the specific feature belongs to the reference information, and the information to be detected is classified and processed according to whether there is an anomaly. The information to be detected with an anomaly is reclassified and data cleaning processing is performed to improve the accuracy of data cleaning and data enhancement by using data processing methods and data enhancement methods.

[0035] Example 3: Anomaly detection and anomaly correction are performed on the information to be detected by establishing an anomaly detection and correction model for the category of information to be detected. Specifically, the process of establishing the anomaly detection and correction model can be as follows: Obtain training data, which includes but is not limited to data with format errors and data with semantic defects in this category. Specifically, mark the specific format errors and semantic defects in the training data. At the same time, the training data also includes data after correcting format errors and semantic errors, etc. Select a deep learning model as the model matrix, and input the training data into the model matrix for the model matrix to learn. After learning, an initial model is obtained. Select data that does not exist in the training data as verification data. The verification data is consistent with the training data, and there is no annotation for format errors or semantic defects in the verification data. Input the verification data into the initial model to obtain the anomaly detection result. When the anomaly detection result indicates that the verification data has errors, a correction result is generated. When the anomaly detection result indicates that the verification data has no errors, no correction result is generated. Manually judge whether the verification data has errors to obtain the verification result. Compare the verification result and the anomaly detection result, and adjust and optimize the initial model according to the differences between the verification result and the anomaly detection result to obtain the anomaly detection and correction model. The specific process is as Figure 4 shown.

[0036] As Figure 3 shown, the cross-validation method includes: obtaining the sub-results of the second data group to obtain the second result, obtaining a single data group in the second data group to obtain the data group to be verified, splitting the data group to be verified into several sub-information, obtaining the features of several sub-information to obtain the target features, judging whether several target features are consistent to obtain the judgment result. When the judgment result indicates that several target features are consistent and belong to the second result, extract the verification data group to obtain the second completed data. When the judgment result indicates that several target features are not consistent and belong to the second result, extract the verification data group to obtain the second uncompleted data.

[0037] It should be noted that the process of obtaining the features of several sub-information to obtain the target features can be performed by training a dedicated feature recognition model for feature recognition, or can be manually recognized by staff. The specific recognition method is determined according to the actual usage situation. Through the cross-validation method, data with different expressions of the same meaning in the data group to be verified can be cross-validated and combined with the category information of the data group to improve the reliability of data cleaning. The specific data cleaning can also be combined with deeper cleaning techniques and data processing methods in the existing technology to improve the data cleaning accuracy.

[0038] As Figure 1As shown in the figure, the data augmentation method includes: recording the features identified during the data cleaning process of the completed data to obtain a number of associated features, splitting the completed data to obtain sub-completed data, establishing the corresponding relationship between the associated features and the sub-completed data, performing feature derivation based on the associated features in cooperation with the target result to obtain derived features, adding the derived features to the sub-completed data corresponding to the associated features based on the corresponding relationship to obtain enhanced data, and extracting the sub-completed data that has not been added with the derived features to obtain unenhanced data.

[0039] It should be noted that the associated features can be obtained from the system's operation logs during the actual operation process. The specific associated features include target features and specific features. By recording the features of the completed data to obtain the associated features, the amount of data processing can be reduced and the work efficiency can be improved. The process of performing feature derivation based on the associated features in cooperation with the target result to obtain derived features can be achieved by recording the associated features and retrieving them in combination with the target result, or by training a dedicated feature derivation model to obtain the corresponding derived features by inputting the associated features and the target result. At the same time, manual derivation can also be performed by the staff, which is specifically determined according to the actual usage situation. Adding the obtained derived features to the sub-completed data can complete the data augmentation. The unenhanced data is the data for which the derived features cannot be found, and it is specifically fed back to the staff for further determination. At the same time, a verification agency can be added, and the enhanced data is sampled by random sampling and handed over to the staff for judgment to improve the data processing accuracy.

[0040] As Figure 1 and Figure 5 As shown in the figure, the classification method includes: obtaining the categories of the past data to be processed to obtain past categories, further obtaining category information based on the past categories to obtain filled categories, integrating the past categories and the filled categories to obtain a category set, determining the scope size based on the user's needs, determining the classified categories in the category set based on the scope size to obtain category criteria, splitting the data to be processed to obtain sub-processed data, matching the sub-processed data based on the category criteria to obtain sub-results, and integrating all the sub-results to obtain a classification result.

[0041] It should be noted that the process of obtaining the past categories by getting the categories of past data to be processed can be achieved by manual classification by staff according to the past data to be processed, or by classifying according to the known types of data to be imported into the multi-modal large model. The process of further obtaining category information based on the past categories to get the filled categories, that is, on the basis of the original past categories, further subdividing the categories to get the filled categories. The process of determining the scope size based on user needs, that is, according to user needs, determining the category scope to be used, specifically, it can be understood whether to use fine-grained categories or coarse-grained categories. For example, the coarse-grained category can be transportation, and the fine-grained category can be the extension of transportation, specifically road category, marking category, signal light category, etc. By selecting different depths of classified categories according to user needs, the user can decide whether the data cleaning and data enhancement processes pursue accuracy or efficiency, so as to improve the user experience. Figure 5 The extended path in it represents the inclusion relationship between the filled category and the past category.

[0042] Such as Figure 1 and Figure 2 As shown, the backtracking method includes: obtaining the reclassification result through the classification method based on the incomplete data. The reclassification result is different from the classification result. Based on the reclassification result and the incomplete data, obtaining the re-cleaning result through the data processing method until the re-cleaning result does not contain incomplete data. Obtaining the re-enhancement result through the data enhancement method for the completed data in the re-cleaning result until the re-enhancement result does not contain un-enhanced data.

[0043] It should be noted that through the set backtracking method, a channel for reclassifying and cleaning and enhancing the data that has not been completely cleaned and enhanced is provided. The specific process of re-cleaning and enhancing is the same as the original, which is conducive to improving the efficiency of data processing and reducing the number of times staff intervene in data processing.

[0044] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended embodiments and their equivalents.

Claims

1. A preprocessing method for data cleaning and feature enhancement of multi-modal large models, including: Information acquisition: Obtain the data required by the multi-modal large model to get the data to be processed, obtain user requirements and feedback address information, and the feedback address information is the address information for sending information to the user; Information classification: Use a classification method to divide the data to be processed into domain categories to obtain a classification result; Classification processing: The classification result includes several sub-results. Establish an association relationship between the sub-results and the data to be processed, extract the data to be processed corresponding to the sub-results based on the association relationship to get sub-data, establish a sub-library named after the sub-results to store the sub-data, and establish a storage library to store the sub-libraries; It is characterized in that it includes: Data processing: Perform data cleaning and feature enhancement on the sub-data. Use a data processing method to perform corresponding cleaning processing on the data with different classification results to obtain a cleaning result. The cleaning result includes completed data and uncompleted data. Extract the classification result of the completed data based on the association relationship to get the target result. Based on the target result and the completed data, use a data enhancement method to perform corresponding enhancement processing on the data with different classification results to obtain an enhancement result. The enhancement result includes enhanced data and unenhanced data; Judgment and backtracking: Preset a backtracking threshold. Based on the uncompleted data, unenhanced data, and classification results, use a backtracking method to reclassify and perform data cleaning and feature enhancement to obtain a reclassification result, and record the number of times the same data is backtracked in the backtracking method to get the backtracking times; Feedback generation: Judge the relationship between the backtracking times and the backtracking threshold. When the backtracking times exceed the backtracking threshold, extract the data at the backtracking times to get feedback information, and send the feedback information to the user based on the feedback address information.

2. The preprocessing method for multimodal large model data cleaning and feature enhancement according to claim 1, wherein: The data processing method includes: Preset a processing format, perform normalization processing on the sub-data based on the processing format to obtain preliminary data, split the preliminary data to get several data groups, judge the number of individual data in the data groups to get the target number, extract the data groups with the target number of individual data to get the first data group, extract the data groups with the target number of multiple data to get the second data group, use a detection method to monitor the individual data groups in the first data group to obtain a detection result, extract the individual data groups with the detection result feedback as detection success to get the first completed data, extract the individual data groups with the detection result feedback as detection failure to get the first uncompleted data, use a cross-validation method to verify the individual data groups in the second data group to obtain a verification result, extract the individual data groups with the verification result feedback as verification success to get the second completed data, extract the individual data groups with the verification result feedback as verification failure to get the second uncompleted data, integrate the first completed data and the second completed data to get the completed data, and integrate the first uncompleted data and the second uncompleted data to get the uncompleted data.

3. The preprocessing method for multimodal large model data cleaning and feature enhancement according to claim 2, wherein: The described detection method includes: obtaining a first result by acquiring sub-results of a first data group, obtaining a data group to be detected by acquiring a single data group in the first data group, splitting the data group to be detected to obtain individual information to be detected, obtaining reference information by acquiring domain information of the first result, performing anomaly detection and correction on the information to be detected based on the reference information to obtain a detection result and the processed information to be detected, obtaining first completed data and first uncompleted data based on the detection result, when the detection result feedback is that the processed information to be detected has no anomaly, extracting the processed information to be detected to obtain the first completed data, and when the detection result feedback is that the processed information to be detected has an anomaly, extracting the processed information to be detected to obtain the first uncompleted data.

4. The preprocessing method for multimodal large model data cleaning and feature enhancement according to claim 3, wherein: The process of performing anomaly detection and correction on the information to be detected to obtain a detection result and a correction result is as follows: performing paragraphing on the information to be detected to obtain a number of sub-detection information and processing positions, comparing the basic formats of the sub-detection information in combination with the reference information and extracting the sub-detection information with incorrect basic formats to obtain anomaly values, correcting the anomaly values based on the reference information to obtain corrected sub-detection information, integrating the sub-detection information based on the processing positions to obtain the processed information to be detected, extracting the features in the processed information to obtain specific features, judging the relationship between the specific features and the reference information to obtain a target relationship, obtaining a detection result based on the target relationship, when the target relationship feedback is that the specific features belong to the reference information, the detection result is that the processed information to be detected has no anomaly, extracting the processed information to be detected to the first completed data, and when the target relationship feedback is that the specific features do not belong to the reference information, the detection result is that the processed information to be detected has an anomaly, extracting the processed information to be detected to obtain the first uncompleted data.

5. The preprocessing method for multimodal large model data cleaning and feature enhancement according to claim 2, wherein: The described cross-validation method includes: obtaining a second result by acquiring sub-results of a second data group, obtaining a data group to be verified by acquiring a single data group in the second data group, splitting the data group to be verified to obtain a number of sub-informations, obtaining target features by acquiring the features of the number of sub-informations, judging whether the number of target features is consistent to obtain a judgment result, when the judgment result feedback is that the number of target features is consistent and belongs to the second result, extracting the verification data group to obtain the second completed data, and when the judgment result feedback is that the number of target features is not consistent and belongs to the second result, extracting the verification data group to obtain the second uncompleted data.

6. The preprocessing method for multimodal large model data cleaning and feature enhancement according to claim 1, wherein: The described data augmentation method includes: recording the features identified during data cleaning processing of the completed data to obtain a number of associated features, splitting the completed data to obtain sub-completed data, establishing a correspondence between the associated features and the sub-completed data, performing feature derivation on the associated features in cooperation with the target result to obtain derived features, adding the derived features to the sub-completed data corresponding to the associated features based on the correspondence to obtain augmented data, and extracting the sub-completed data without added derived features to obtain un-augmented data.

7. The preprocessing method for multi-modal large model data cleaning and feature enhancement according to claim 1, wherein: The classification method includes: obtaining the past categories of the past data to be processed to obtain the past categories, further obtaining category information based on the past categories to obtain the filled categories, integrating the past categories and the filled categories to obtain a category set, determining the scope size based on user requirements, determining the classified categories in the category set based on the scope size to obtain the category criteria, splitting the data to be processed to obtain sub-processed data, matching the sub-processed data based on the category criteria to obtain sub-results, and integrating all sub-results to obtain the classification result.

8. The preprocessing method for multimodal large model data cleaning and feature enhancement according to claim 1, characterized in that: The backtracking method includes: reclassifying the incomplete data through the classification method to obtain the reclassification result, where the reclassification result is different from the classification result, obtaining the re-cleaning result through the data processing method based on the reclassification result and the incomplete data until the re-cleaning result does not contain incomplete data, and obtaining the re-enhanced result through the data enhancement method for the completed data in the re-cleaning result until the re-enhanced result does not contain un-enhanced data.

9. A preprocessing system for data cleaning and feature enhancement of multi-modal large models, characterized in that: The preprocessing method for multi-modal large model data cleaning and feature enhancement according to any one of claims 1-8 is used.

Citation Information

Patent Citations

  • Multi-modal large model pre-training data cleaning and balancing method

    CN117932219A

  • Data processing method and device, storage medium and electronic equipment

    CN115145902A

  • Big data analysis method based on decision tree

    CN117056834A

  • Medical data cleaning method, system and equipment based on machine learning optimization and medium

    CN119003999A

  • Metadata driven combined real-time and batch data ingestion framework with real-time multi-view generation

    US20200372074A1