Data processing method and device, electronic equipment and storage medium

Through the method of classifying multimodal data and data cleaning, the problem of poor multimodal data cleaning in the existing technology is solved, and more efficient data cleaning effect is achieved, and the performance and generalization ability of the model are improved.

CN120217018APending Publication Date: 2025-06-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510338855.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the category redundancy and low quality problems in multimodal data, resulting in poor data cleaning results.

Method used

By classifying multiple multimodal data based on at least one data mode, data cleaning is performed to obtain the processed data set. The specific steps include: classifying the multiple multimodal data in the data set to obtain a plurality of first data sets; data cleaning of multiple multimodal data in the first data set to obtain a second data set; and obtaining the processed data set based on the multimodal data in the multiple second data sets.

Benefits of technology

Through category division and data cleaning, the data cleaning effect is significantly improved, and the category redundancy and low quality problems in multimodal data can be more effectively handled, providing cleaner data, thereby improving the accuracy and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217018A_ABST
    Figure CN120217018A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, natural language processing, deep learning, large models and the like. According to the specific implementation scheme, based on at least one data mode, multiple pieces of multi-mode data included in a data set are subjected to category division, and multiple first data sets are obtained; performing data cleaning on the multiple pieces of multi-modal data included in the first data set to obtain a second data set; and obtaining a processed data set based on the multi-modal data included in each of the plurality of second data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, particularly to technical fields such as computer vision, natural language processing, deep learning, large models, etc. More specifically, a data processing method, apparatus, electronic device, and storage medium are disclosed. Background Art

[0002] Data cleaning is a key step in data preprocessing, aiming to identify, correct, or delete incomplete, incorrect, duplicate, redundant, irrelevant, or inconsistently formatted data in a dataset to ensure the high quality and consistency of the data. In the field of artificial intelligence, the effect of data cleaning directly determines the performance and generalization ability of the model. Summary of the Invention

[0003] The present disclosure provides a data processing method, apparatus, electronic device, and storage medium.

[0004] According to one aspect of the present disclosure, a data processing method is provided, including: classifying multiple multimodal data included in a dataset based on at least one data modality to obtain multiple first data sets; performing data cleaning on the multiple multimodal data included in the first data sets to obtain second data sets; and obtaining a processed dataset based on the multimodal data included in each of the multiple second data sets.

[0005] According to another aspect of the present disclosure, a data processing apparatus is provided, including: a classification module for classifying multiple multimodal data included in a dataset based on at least one data modality to obtain multiple first data sets; a first processing module for performing data cleaning on the multiple multimodal data included in the first data sets to obtain second data sets; and a second processing module for obtaining a processed dataset based on the multimodal data included in each of the multiple second data sets.

[0006] According to another aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described above.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program which, when executed by a processor, implements the method as described above.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 Schematically shows an exemplary system architecture to which the data processing method and apparatus according to an embodiment of the present disclosure can be applied.

[0012] Figure 2 Schematically shows a flowchart of the data processing method according to an embodiment of the present disclosure.

[0013] Figure 3 Schematically shows a flowchart of the data processing method according to another embodiment of the present disclosure.

[0014] Figure 4 Schematically shows a schematic diagram of the data processing method according to an embodiment of the present disclosure.

[0015] Figure 5 Schematically shows a block diagram of the data processing apparatus according to an embodiment of the present disclosure.

[0016] Figure 6 Shows a schematic block diagram of an example electronic device that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0018] With the comprehensive development of technologies such as large models and multi-modal models, the demand for high-quality data is growing explosively. The data cleaning methods in related technologies can be divided into two categories according to the technical route, namely, multi-modal data cleaning methods based on hash series algorithms, and cleaning methods assisted by neural network deep learning, multi-modal large models, etc.

[0019] For example, the Simhash-based duplicate detection and cleaning method can use the Simhash algorithm to calculate the hash value of text or images, detect duplicate or highly similar data by comparing the hash values, and delete them. Another example is that the semantic consistency detection method based on multimodal large models can use multimodal large models to detect semantic inconsistencies between image content and text labels, thereby identifying and cleaning dirty samples, and effectively detecting noise labels and toxic samples. Another example is that the cleaning method based on feature fusion and deep learning can merge the feature vectors of multimodal data through feature-level fusion, and then use deep learning algorithms for anomaly detection and cleaning to effectively handle the inconsistencies between different modalities in multimodal data.

[0020] However, the data cleaning methods in related technologies mainly target single-modal data and are difficult to effectively handle problems such as category redundancy and low quality existing in multimodal data.

[0021] In view of this, embodiments of the present disclosure provide a data processing method. By classifying data at the category level and performing data cleaning according to the correlation between data in multiple modalities, the data cleaning effect can be effectively improved. Specifically, the data processing method includes: classifying multiple multimodal data included in a data set based on at least one data modality to obtain multiple first data sets; performing data cleaning on the multiple multimodal data included in the first data sets to obtain second data sets; and obtaining a processed data set based on the multimodal data included in each of the multiple second data sets.

[0022] Figure 1 An exemplary system architecture to which the data processing method and apparatus according to embodiments of the present disclosure can be applied is schematically shown.

[0023] It should be noted that Figure 1 The illustration is only an example of the system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the data processing method and apparatus can be applied may include a terminal device, but the terminal device can implement the data processing method and apparatus provided by embodiments of the present disclosure without interacting with a server.

[0024] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0025] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).

[0026] Terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0027] Server 105 can be a server that provides various services, such as a background management server that provides support for the content browsed by users using terminal devices 101, 102, and 103 (for example only). The background management server can analyze and process data such as received user requests, etc., and feedback the processing results (such as web pages, information, or data, etc. obtained or generated according to user requests) to the terminal device.

[0028] It should be noted that the data processing method provided by the embodiments of the present disclosure can generally be executed by terminal devices 101, 102, or 103. Correspondingly, the data processing device provided by the embodiments of the present disclosure can also be set in terminal devices 101, 102, or 103.

[0029] Alternatively, the data processing method provided by the embodiments of the present disclosure can generally also be executed by server 105. Correspondingly, the data processing device provided by the embodiments of the present disclosure can generally be set in server 105. The data processing method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103, and / or server 105. Correspondingly, the data processing device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103, and / or server 105.

[0030] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0031] are merely illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers.

[0032] In the technical solutions of the present disclosure, the authorization or consent of the user is obtained before acquiring or collecting the user's personal information.

[0033] Figure 2 Schematically shows a flowchart of a data processing method according to an embodiment of the present disclosure.

[0034] As Figure 2 shown, the method 200 includes operations S210 to S230.

[0035] In operation S210, based on at least one data modality, multiple multimodal data included in the data set are classified to obtain multiple first data sets.

[0036] In operation S220, the multiple multimodal data included in the first data set are data-cleaned to obtain a second data set.

[0037] In operation S230, a processed data set is obtained based on the multimodal data included in each of the multiple second data sets.

[0038] The data set may include multiple multimodal data for training, testing, etc. of a multimodal model. The multimodal data may be composed of sub-data of multiple data modalities. The multiple data modalities may include a text modality, an image modality, an audio modality, etc., which are not limited herein. Correspondingly, the multimodal data may include text sub-data corresponding to the text modality, image sub-data corresponding to the image modality, audio sub-data corresponding to the audio modality, and so on. For example, the multimodal data may be data that simultaneously has text content and image content.

[0039] The sub-data of a single data modality may be used as a classification benchmark to classify multiple multimodal data. For example, the text sub-data of multiple multimodal data may be classified, and according to the category to which the text sub-data belongs, the category to which the corresponding multimodal data belongs is determined. Alternatively, the sub-data of multiple data modalities may be used as a classification benchmark to classify multiple multimodal data. For example, classification may be performed based on the sub-data of each data modality respectively. For a multimodal data, the classification results of the sub-data of each data modality of the multimodal data may all represent the category to which the multimodal data belongs. If the categories represented by the multiple classification results are the same, it may be considered that the multimodal data belongs to this category; if the categories represented by the multiple classification results are different, means such as weight assignment and confidence analysis may be used to determine the category to which the multimodal data belongs, which is not limited herein.

[0040] Similar to data classification, when performing in-class data cleaning, data cleaning can be performed based on sub-data of a single data modality, or sub-data of multiple data modalities can be combined for data cleaning, which is not limited herein. When performing data cleaning, multimodal data that does not meet the requirements in the original multiple multimodal data in the first data set can be removed, and the remaining multimodal data can be used as the second data set.

[0041] Based on the multimodal data remaining after cleaning in the multiple second data sets, a data set after data cleaning can be determined, that is, the processed data set. Therefore, the multimodal data included in the processed data set are all multimodal data remaining after in-class data cleaning.

[0042] According to an embodiment of the present disclosure, multiple multimodal data can be classified based on at least one data modality, and then in-class data cleaning can be performed on the classified data set to achieve in-class granularity multimodal data cleaning, so as to fully explore the correlation between each modality, improve the data cleaning effect, and further provide cleaner data for multimodal tasks, improving the accuracy and generalization ability of the trained model.

[0043] The following further illustrates the Figure 2 method shown with specific embodiments. In the following embodiments, the multimodal data may include text sub-data corresponding to the text modality and image sub-data corresponding to the image modality. Optionally, for other data modalities, such as the audio modality, data conversion can be performed on the sub-data of the other data modality to obtain text sub-data or image sub-data. For example, for sub-data of the audio modality, it can be converted into audio data in the time domain, and the audio data in the time domain can be represented as a data sequence, then the converted audio data in the time domain can be regarded as text sub-data. For another example, for sub-data of the audio modality, it can be converted into audio data in the frequency domain, and the audio data in the frequency domain can be represented as a spectrogram, then the converted audio data in the frequency domain can be regarded as image sub-data.

[0044] As an example of classifying multiple multimodal data based on a single data modality, based on at least one data modality, classifying the multiple multimodal data included in the data set to obtain multiple first data sets may include the following operations:

[0045] Determine text sub-data corresponding to the text modality from the multimodal data; perform category name matching based on the text sub-data of the multiple multimodal data to obtain a matching result; and classify the multiple multimodal data into multiple first data sets based on the matching result.

[0046] The category name matching can be expressed as text matching with the category name as the keyword.

[0047] According to an embodiment of the present disclosure, performing category name matching based on the text sub-data of each of multiple multimodal data to obtain a matching result may include the following operations:

[0048] Using the category names included in each of the multiple text sub-data as keywords, performing text matching on the multiple text sub-data to obtain the similarity between each of the multiple text sub-data; and obtaining a matching result based on the similarity between each of the multiple text sub-data.

[0049] Optionally, the similarity between each of the multiple text sub-data may be expressed as the degree of overlap between the category names of each of the multiple text sub-data. For example, if the category name of text sub-data A1 is white crabapple flower, the category name of text sub-data A2 is pink Chinese rose, and the category name of text sub-data A3 is red Chinese rose, then the similarity between text sub-data A1 and text sub-data A2 or text sub-data A3 may be 1 / 4, and the similarity between text sub-data A2 and text sub-data A3 may be 3 / 4.

[0050] Alternatively, optionally, the similarity between each of the multiple text sub-data may also be expressed as the distance between the mappings of the category names of each of the multiple text sub-data in the vector space. For example, text encoding may be performed on the category names of the text sub-data to obtain text feature vectors, and the text feature vectors may be mapped to a point in the vector space. Similarly, the text feature vectors of each of the multiple text sub-data may be mapped to multiple points in the vector space. The similarity between two text sub-data may be expressed as the distance between the two points mapped in the vector space.

[0051] Optionally, the matching result may be determined based on the similarity between each of the multiple text sub-data by means of unsupervised clustering. For example, the similarity between each of the multiple text sub-data may be traversed. If the similarity between two text sub-data is higher than a threshold, the two text sub-data are classified into one category. If the similarity between each of a text sub-data and other text sub-data is less than a threshold, the text sub-data is classified into a separate category.

[0052] Alternatively, the matching result can also be determined based on the similarity between each of multiple text sub-data and the cluster center through supervised clustering. For example, some root category names can be determined by combining the tracing of existing knowledge based on the existing category names. For example, the root category name determined based on the category name of text sub-data A1 can be "Chinese flowering crabapple", and the root category name determined jointly based on the category names of text sub-data A2 and text sub-data A3 can be "Chinese rose". These root category names can be used as the cluster centers when performing supervised clustering. The similarity between the text sub-data and the cluster center can be calculated using the calculation method of the similarity between text sub-data described above, which will not be elaborated here. For each text sub-data, the category to which the text sub-data belongs can be determined based on the maximum value of the similarity between the text sub-data and each cluster center. For example, if the similarity between text sub-data A1 and cluster center α1 is the largest, it can be determined that the category to which text sub-data A1 belongs is the category corresponding to cluster center α1.

[0053] The matching result obtained through category name matching can be represented as the categories to which each text sub-data belongs. Based on the categories to which each text sub-data belongs, the multi-modal data corresponding to the text sub-data belonging to the same category can be added to a data set to obtain a first data set.

[0054] As another example of classifying multiple multi-modal data based on a single data modality, based on at least one data modality, classifying the multiple multi-modal data included in the data set to obtain multiple first data sets may include the following operations:

[0055] Determine the image sub-data corresponding to the image modality from the multi-modal data; perform image clustering based on the image sub-data of each of the multiple multi-modal data to obtain a clustering result; and classify the multiple multi-modal data into multiple first data sets based on the clustering result.

[0056] The two-dimensional image sub-data can be converted into a one-dimensional feature representation, and image clustering of multiple image sub-data can be performed based on the one-dimensional feature representation of the image sub-data.

[0057] According to an embodiment of the present disclosure, performing image clustering based on the image sub-data of each of the multiple multi-modal data to obtain a clustering result may include the following operations:

[0058] Perform feature encoding on each of the multiple image sub-data to obtain multiple image features; and perform clustering processing on the multiple image features to obtain a clustering result.

[0059] The feature encoding of the image sub-data can be implemented using various encoders and encoding algorithms, which are not limited herein. The image features obtained through feature encoding can be represented as a feature vector.

[0060] The clustering process for multiple image features can be implemented using various supervised clustering or unsupervised clustering methods, which are not limited herein.

[0061] Optionally, the clustering of the image sub-data can also be achieved in other ways. For example, through the method of generating text from images, the image sub-data can be converted into text descriptions, and then using the text clustering method, the multiple text descriptions obtained by conversion can be clustered to obtain the clustering results of the multiple image sub-data.

[0062] The clustering results obtained through image clustering can be represented as the categories to which each image sub-data belongs. Based on the categories to which each image sub-data belongs, the multi-modal data corresponding to the image sub-data belonging to the same category can be added to a data set to obtain a first data set.

[0063] As another alternative implementation, multiple multi-modal data can also be classified by combining multiple data modalities.

[0064] According to the embodiments of the present disclosure, based on at least one data modality, performing category division on multiple multi-modal data included in a data set to obtain multiple first data sets may include the following operations:

[0065] Determine the text sub-data corresponding to the text modality and the image sub-data corresponding to the image modality from the multi-modal data; perform category name matching based on the text sub-data of each of the multiple multi-modal data to obtain a matching result; perform image clustering on the image sub-data of each of the multiple multi-modal data to obtain a clustering result; and classify the multiple multi-modal data into multiple first data sets based on the matching result and the clustering result.

[0066] Optionally, both the matching result and the clustering result can be obtained through the method of supervised clustering. The matching result can include the probabilities that each of the multiple text sub-data belongs to multiple categories, and the clustering result can include the probabilities that each of the multiple image sub-data belongs to multiple categories. The multiple categories can be the categories preset during supervised clustering. For a multi-modal data, based on the product of the probability that the text sub-data of the multi-modal data belongs to multiple categories and the probability that the image sub-data of the multi-modal data belongs to multiple categories, the probability that the multi-modal data belongs to multiple categories can be determined, so that the multi-modal data included in each category can be determined, and the multi-modal data included in one category can be used as a first data set.

[0067] Alternatively, at least one of the matching result and the clustering result can be obtained by unsupervised clustering. The matching result can include the probabilities of multiple text sub-data belonging to multiple text categories respectively, and the clustering result can include the probabilities of multiple image sub-data belonging to multiple image categories respectively. The multiple categories corresponding to the multiple first data sets can be represented as the union of the multiple text categories and the multiple image categories. Accordingly, if the text category B1 and the image category B2 can be represented as the same category, the multi-modal data corresponding to this category can be the union of the multi-modal data corresponding to the text category B1 and the multi-modal data corresponding to the image category B2.

[0068] Alternatively, the multi-modal data can be roughly classified based on one of the matching result and the clustering result first, and then finely classified based on the other result to complete the classification of the multiple multi-modal data.

[0069] For example, based on the matching result and the clustering result, classifying the multiple multi-modal data into multiple first data sets may include the following operations:

[0070] Classifying the multiple multi-modal data into multiple third data sets based on the matching result; calculating the confidence of the multi-modal data belonging to the corresponding third data set based on the clustering result to obtain a confidence calculation result; and performing inter-class adjustment on the multi-modal data included in each of the multiple third data sets based on the confidence calculation result to obtain multiple first data sets.

[0071] Optionally, the correlation between the multiple image categories and the multiple text categories represented by the multiple third data sets can be calculated to obtain the probability that the multi-modal data under one image category belongs to each of the third data sets. The probability that the multi-modal data under one image category belongs to a third data set can be expressed as the confidence of the multi-modal data belonging to the corresponding third data set. The calculation of this correlation can be expressed as a text similarity calculation with the category name as the keyword, which will not be elaborated here. Accordingly, the confidence calculation result can include the probabilities that the multi-modal data included in each of the multiple image categories belongs to the multiple third data sets respectively. Multiplying this probability by the probability that the image sub-data represented by the clustering result belongs to each of the multiple image categories respectively, the final probability that the multi-modal data belongs to each of the third data sets respectively can be obtained. For a multi-modal data, if this multi-modal data originally belongs to the third data set C1, but the calculated probability that this multi-modal data belongs to the third data set C1 is less than the probability of belonging to another third data set, for example, the third data set C2, then this multi-modal data can be classified into the third data set C2 by means of inter-class adjustment.

[0072] According to an embodiment of the present disclosure, by classifying multi-modal data based on sub-data of multiple data modalities, the correlation between multiple data modalities can be combined to achieve data classification and improve the accuracy of data classification.

[0073] Optionally, before cleaning the data, the obtained multiple first data sets can also be merged by category to aggregate data sets with text or semantic relationships into one category.

[0074] Figure 3 The flowchart of a data processing method according to another embodiment of the present disclosure is schematically shown.

[0075] As Figure 3 shown, the method 300 includes operations S310 to S340.

[0076] In operation S310, based on at least one data modality, multiple multi-modal data included in the data set are classified by category to obtain multiple first data sets.

[0077] In operation S320, the multiple first data sets are merged by category to obtain multiple fourth data sets.

[0078] In operation S330, the multiple multi-modal data included in the fourth data set are cleaned to obtain a second data set.

[0079] In operation S340, based on the multi-modal data included in each of the multiple second data sets, a processed data set is obtained.

[0080] Multiple first data sets with similar text or semantics can be merged into one fourth data set. If a first data set has no text or semantic similarity with other first data sets, then this first data set can be used as one fourth data set.

[0081] Optionally, each first data set can be represented as a category, and this category can have a clustering center. The category merging of the multiple first data sets can be completed based on the clustering centers of the multiple first data sets respectively.

[0082] According to an embodiment of the present disclosure, the operation of merging the multiple first data sets by category to obtain multiple fourth data sets can include the following operations:

[0083] Determine the clustering centers of the multiple first data sets in the vector space respectively; and based on the distances between the multiple clustering centers, merge the multiple first data sets by category to obtain multiple fourth data sets.

[0084] For example, for each first data set, if the distance between the cluster center of the first data set and the cluster center of another first data set is less than the distance threshold, the first data set and the other first data set can be aggregated into one data set. If the distance between the cluster center of the first data set and the cluster centers of all other first data sets is greater than the distance threshold, the first data set can be separately used as a fourth data set.

[0085] According to an embodiment of the present disclosure, data cleaning is performed on multiple multimodal data included in the first data set to obtain a second data set, which may include the following operations:

[0086] Feature encoding is performed on sub-data of at least one modality of each of the multiple multimodal data included in the first data set to obtain multiple feature vectors; clustering processing is performed on the multiple feature vectors to determine a preset number of target feature vectors far from the cluster center; and the multimodal data corresponding to the preset number of target feature vectors is removed from the first data set to obtain the second data set.

[0087] For example, for multiple multimodal data included in a first data set, a text encoder can be used to perform feature encoding on the text sub-data of each of the multiple multimodal data to obtain multiple feature vectors. For another example, an image encoder can be used to perform feature encoding on the image sub-data of each of the multiple multimodal data to obtain multiple feature vectors. For still another example, multiple text sub-features can be obtained by performing feature encoding on the text sub-data of each of the multiple multimodal data using a text encoder, and multiple image sub-features can be obtained by performing feature encoding on the image sub-data of each of the multiple multimodal data using an image encoder. The corresponding text sub-features and image sub-features can be combined by means of splicing, weighted summation, etc. to obtain the corresponding feature vectors.

[0088] The clustering method used for clustering the multiple feature vectors can be various supervised or unsupervised clustering methods, which are not limited herein.

[0089] The preset number can be represented as a preset fixed value, for example, it can be set to 5, 6, etc. Alternatively, the preset number can also be represented as the number of a preset proportion of the number of multimodal data included in the first data set. For example, it can be set to 10%, 15%, etc. of the number of multimodal data included in the first data set.

[0090] According to an embodiment of the present disclosure, by performing multi-level cleaning on multimodal data, including category cleaning, category merging, and intra-category data cleaning, cleaner data can be provided for multimodal tasks, thereby improving the performance of the model.

[0091] Optionally, after completing the in-class data cleaning, the self-training process of the model can also be utilized to further clean the multi-modal data.

[0092] Figure 4 A schematic diagram of a data processing method according to an embodiment of the present disclosure is schematically shown.

[0093] As Figure 4 shown, the data set 401 includes data samples for model training and testing under multi-modal tasks, and these data samples are the multi-modal data 402.

[0094] The multi-modal data 402 may include text sub-data corresponding to the text modality and image sub-data corresponding to the image modality. Multiple multi-modal data 402 can be classified based on the text modality and the image modality to obtain multiple first data sets 403.

[0095] For each first data set 403, data cleaning can be performed based on at least one data modality of the multi-modal data included in the first data set 403. This data cleaning can be to clean the abnormal multi-modal data in the first data set 403 to obtain a second data set 404. Taking the at least one data modality as the image modality as an example, feature encoding can be performed on the image sub-data of the multi-modal data included in the first data set 403 to obtain feature vectors. By clustering the feature vectors, 10% of the image sub-data that is farthest from the cluster center can be screened out. Compared with the image sub-data close to the cluster center, it can be considered that the 10% of the image sub-data that is farthest from the cluster center is inferior in terms of image size, image quality, category error rate, etc. Therefore, the 10% of the image sub-data that is farthest from the cluster center can be removed from the first data set to obtain a second data set.

[0096] Aggregating the multi-modal data included in each of the multiple second data sets 404 can obtain a processed data set 405.

[0097] Optionally, data cleaning of the processed data set 405 can be performed by self-verifying the processed data set 405 using the classification model 406.

[0098] The multiple multi-modal data included in the processed data set 405 can be divided into multiple training data 407 and multiple test data 408. For example, the multiple multi-modal data included in the processed data set 405 can be divided according to a certain ratio to obtain multiple training data 407 and multiple test data 408.

[0099] The initial model 409 can be trained using multiple training data 407 to obtain a classification model 406. By processing multiple test data 408 using the classification model 406, the classification results 410 of the multiple test data 408 can be obtained. The classification result 410 can be represented as the second data set to which the test data 408 belongs inferred by the classification model 406.

[0100] Each test data 408 can be configured with a corresponding label, which represents the second data set 404 to which the test data 408 originally belongs.

[0101] Based on the classification results 410 of the multiple test data 408 and the second data sets 404 to which the multiple test data 408 belong, the classification result of each test data can be compared with the second data set to which the test data originally belongs. If the two do not match, the test data can be considered a bad example sample. Based on all the detected bad example samples, it can be determined which categories belong to the easily confused categories, and thus at least one target data set can be determined from the multiple second data sets 404. The category corresponding to the target data set is the easily confused category.

[0102] After determining the target data set, data cleaning can continue to be performed on the multimodal data included in the target data set to further clean the processed data set 405.

[0103] The method used for data cleaning of the multimodal data included in the target data set is not limited herein. For example, a large model can be called for data cleaning, or data cleaning can be performed based on the methods of large model, feature extraction, and clustering, or data cleaning can be performed through manual annotation, etc.

[0104] According to an embodiment of the present disclosure, by using the self - learning process of the model to further clean the sample data, the cost of data cleaning can be reduced, and the cleaning effect can be further improved.

[0105] Figure 5 A block diagram of a data processing device according to an embodiment of the present disclosure is schematically shown.

[0106] As Figure 5 shown, the data processing device 500 includes a classification module 510, a first processing module 520, and a second processing module 530.

[0107] The classification module 510 is configured to perform category division on multiple multimodal data included in a data set based on at least one data modality to obtain multiple first data sets.

[0108] The first processing module 520 is configured to perform data cleaning on multiple multimodal data included in the first data set to obtain a second data set.

[0109] The second processing module 530 is configured to obtain a processed data set based on the multimodal data included in each of the multiple second data sets.

[0110] According to an embodiment of the present disclosure, the classification module 510 includes a first classification unit, a second classification unit, and a third classification unit.

[0111] The first classification unit is configured to determine text sub-data corresponding to the text modality from the multimodal data.

[0112] The second classification unit is configured to perform category name matching based on the text sub-data of each of the multiple multimodal data to obtain a matching result.

[0113] The third classification unit is configured to classify the multiple multimodal data into multiple first data sets based on the matching result.

[0114] According to an embodiment of the present disclosure, the classification module 510 includes a fourth classification unit, a fifth classification unit, and a sixth classification unit.

[0115] The fourth classification unit is configured to determine image sub-data corresponding to the image modality from the multimodal data.

[0116] The fifth classification unit is configured to perform image clustering based on the image sub-data of each of the multiple multimodal data to obtain a clustering result.

[0117] The sixth classification unit is configured to classify the multiple multimodal data into multiple first data sets based on the clustering result.

[0118] According to an embodiment of the present disclosure, the classification module 510 includes a seventh classification unit, an eighth classification unit, a ninth classification unit, and a tenth classification unit.

[0119] The seventh classification unit is configured to determine text sub-data corresponding to the text modality and image sub-data corresponding to the image modality from the multimodal data.

[0120] The eighth classification unit is configured to perform category name matching based on the text sub-data of each of the multiple multimodal data to obtain a matching result.

[0121] The ninth classification unit is configured to perform image clustering based on the image sub-data of each of the multiple multimodal data to obtain a clustering result.

[0122] The tenth classification unit is configured to classify the multiple multimodal data into multiple first data sets based on the matching result and the clustering result.

[0123] According to an embodiment of the present disclosure, the tenth classification unit includes a first classification subunit, a second classification subunit, and a third classification subunit.

[0124] The first classification subunit is configured to classify a plurality of multimodal data into a plurality of third data sets based on a matching result.

[0125] The second classification subunit is configured to calculate a confidence level of the multimodal data belonging to the corresponding third data set based on a clustering result, to obtain a confidence level calculation result.

[0126] The third classification subunit is configured to perform inter-class adjustment on the multimodal data included in each of the plurality of third data sets based on the confidence level calculation result, to obtain a plurality of first data sets.

[0127] According to an embodiment of the present disclosure, the second classification unit or the eighth classification unit includes a fourth classification subunit and a fifth classification subunit.

[0128] The fourth classification subunit is configured to perform text matching on a plurality of text sub-data by using the category names included in the plurality of text sub-data as keywords, to obtain similarities between the plurality of text sub-data.

[0129] The fifth classification subunit is configured to obtain a matching result based on the similarities between the plurality of text sub-data.

[0130] According to an embodiment of the present disclosure, the fifth classification unit or the ninth classification unit includes a sixth classification subunit and a seventh classification subunit.

[0131] The sixth classification subunit is configured to perform feature encoding on a plurality of image sub-data respectively, to obtain a plurality of image features.

[0132] The seventh classification subunit is configured to perform clustering processing on the plurality of image features, to obtain a clustering result.

[0133] According to an embodiment of the present disclosure, the first processing module 520 includes a first processing unit, a second processing unit, and a third processing unit.

[0134] The first processing unit is configured to perform feature encoding on sub-data of at least one modality of each of the plurality of multimodal data included in the first data set, to obtain a plurality of feature vectors.

[0135] The second processing unit is configured to perform clustering processing on the plurality of feature vectors to determine a preset number of target feature vectors that are far from the clustering center.

[0136] The third processing unit is configured to remove the multimodal data corresponding to the preset number of target feature vectors from the first data set, to obtain a second data set.

[0137] According to an embodiment of the present disclosure, the data processing apparatus 500 further includes a merging module.

[0138] The merging module is configured to perform class merging on a plurality of first data sets to obtain a plurality of fourth data sets.

[0139] According to an embodiment of the present disclosure, the first processing module 520 includes a fourth processing unit.

[0140] The fourth processing unit is configured to perform data cleaning on a plurality of multimodal data included in the fourth data set to obtain a second data set.

[0141] According to an embodiment of the present disclosure, the merging module includes a first merging unit and a second merging unit.

[0142] The first merging unit is configured to determine the clustering centers of the plurality of first data sets in the vector space respectively.

[0143] The second merging unit is configured to perform class merging on the plurality of first data sets based on the distances between the plurality of clustering centers respectively to obtain a plurality of fourth data sets.

[0144] According to an embodiment of the present disclosure, the data processing apparatus 500 further includes a third processing module.

[0145] The third processing module is configured to perform data cleaning on the processed data set by self-verifying the processed data set using a classification model.

[0146] According to an embodiment of the present disclosure, the processed data set includes a plurality of training data and a plurality of test data.

[0147] According to an embodiment of the present disclosure, the third processing module includes a fifth processing unit, a sixth processing unit, a seventh processing unit, and an eighth processing unit.

[0148] The fifth processing unit is configured to train an initial model using a plurality of training data to obtain a classification model.

[0149] The sixth processing unit is configured to process a plurality of test data using the classification model to obtain the classification results of the plurality of test data respectively.

[0150] The seventh processing unit is configured to determine at least one target data set from the plurality of second data sets based on the classification results of the plurality of test data respectively and the second data sets to which the plurality of test data belong respectively.

[0151] The eighth processing unit is configured to perform data cleaning on the multimodal data included in the target data set to complete the data cleaning of the processed data set.

[0152] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0153] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method as described above.

[0154] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the method as described above.

[0155] According to an embodiment of the present disclosure, a computer program product includes a computer program which, when executed by a processor, implements the method as described above.

[0156] Figure 6 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0157] As Figure 6 shown, the device 600 includes a computing unit 601 which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0158] Multiple components in device 600 are connected to an input / output (I / O) interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0159] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the data processing method in any other suitable manner (e.g., by means of firmware).

[0160] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0161] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on the remote machine or server.

[0162] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0163] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0164] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0165] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0166] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0167] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A data processing method, comprising: Based on at least one data modality, classify the multiple multimodal data included in the data set into categories to obtain multiple first data sets; Performing data cleaning on a plurality of multimodal data included in the first data set to obtain a second data set; as well as A processed data set is obtained based on the multimodal data respectively included in the plurality of second data sets.

2. The method according to claim 1, wherein: The method of classifying the plurality of multimodal data included in the data set into categories based on at least one data modality to obtain a plurality of first data sets includes: Determining text sub-data corresponding to a text modality from the multimodal data; Performing category name matching based on the text sub-data of each of the plurality of multimodal data to obtain a matching result; and Based on the matching result, the plurality of multimodal data are classified into the plurality of first data sets.

3. The method according to claim 1, wherein: The method of classifying the plurality of multimodal data included in the data set into categories based on at least one data modality to obtain a plurality of first data sets includes: Determining image sub-data corresponding to an image modality from the multimodal data; Performing image clustering based on the image sub-data of each of the plurality of multimodal data to obtain a clustering result; and Based on the clustering result, the plurality of multimodal data are classified into the plurality of first data sets.

4. The method according to claim 1, wherein: The method of classifying the plurality of multimodal data included in the data set into categories based on at least one data modality to obtain a plurality of first data sets includes: Determining text sub-data corresponding to the text modality and image sub-data corresponding to the image modality from the multimodal data; Performing category name matching based on the text sub-data of each of the plurality of multimodal data to obtain a matching result; Performing image clustering based on the image sub-data of each of the plurality of multimodal data to obtain a clustering result; and Based on the matching result and the clustering result, the plurality of multimodal data are classified into the plurality of first data sets.

5. The method according to claim 4, wherein: The classifying the plurality of multimodal data into the plurality of first data sets based on the matching result and the clustering result includes: Based on the matching result, classifying the plurality of multimodal data into a plurality of third data sets; Based on the clustering result, calculating the confidence that the multimodal data belongs to the corresponding third data set to obtain a confidence calculation result; and Based on the confidence calculation result, inter-class adjustment is performed on the multimodal data included in each of the plurality of third data sets to obtain the plurality of first data sets.

6. The method according to claim 2 or 4, wherein: The performing category name matching based on the respective text sub-data of the plurality of multimodal data to obtain a matching result includes: Using the category names included in each of the plurality of text sub-data as keywords, performing text matching on the plurality of text sub-data to obtain the similarities between each of the plurality of text sub-data; and The matching result is obtained based on the similarities between the plurality of text sub-data.

7. The method according to claim 3 or 4, wherein: The performing image clustering based on the image sub-data of each of the plurality of multimodal data to obtain a clustering result includes: Performing feature encoding on the plurality of image sub-data respectively to obtain a plurality of image features; and Clustering is performed on the multiple image features to obtain the clustering result.

8. The method according to claim 1, wherein: The step of performing data cleaning on the plurality of multimodal data included in the first data set to obtain a second data set includes: Performing feature encoding on at least one modal sub-data of each of the plurality of multimodal data included in the first data set to obtain a plurality of feature vectors; Performing clustering processing on the plurality of feature vectors to determine a preset number of target feature vectors that are far away from the cluster center; and The multimodal data corresponding to the preset number of target feature vectors are removed from the first data set to obtain the second data set.

9. The method according to claim 1, further comprising: Merging the multiple first data sets by category to obtain multiple fourth data sets; The step of performing data cleaning on the plurality of multimodal data included in the first data set to obtain the second data set includes: Data cleaning is performed on the multiple multimodal data included in the fourth data set to obtain a second data set.

10. The method according to claim 9, wherein: The step of merging the multiple first data sets by categories to obtain multiple fourth data sets includes: determining a cluster center of each of the plurality of first data sets in a vector space; and Based on the distances between the plurality of cluster centers, the plurality of first data sets are merged into categories to obtain a plurality of fourth data sets.

11. The method according to claim 1, further comprising: The processed data set is cleaned by performing self-validation on the processed data set using a classification model.

12. The method according to claim 11, wherein: The processed data set includes a plurality of training data and a plurality of test data; Wherein, the self-verification of the processed data set by using the classification model to clean the processed data set includes: Using the plurality of training data to train an initial model to obtain the classification model; Processing the plurality of test data using the classification model to obtain classification results of the plurality of test data; Based on the classification results of each of the plurality of test data and the second data sets to which each of the plurality of test data belongs, determining at least one target data set from the plurality of second data sets; and Data cleaning is performed on the multimodal data included in the target data set to complete data cleaning of the processed data set.

13. A data processing device, comprising: A classification module, configured to classify a plurality of multimodal data included in the data set into categories based on at least one data modality, to obtain a plurality of first data sets; A first processing module, configured to perform data cleaning on the plurality of multimodal data included in the first data set to obtain a second data set; as well as The second processing module obtains a processed data set based on the multimodal data respectively included in the plurality of second data sets.

14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.

16. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12.