Data set cleaning method, device, electronic device and computer readable medium

By using training image samples and verification image samples during the image classification data set cleaning process, updating classification labels and iterative cleaning, the problems of low cleaning efficiency and data loss in the prior art are solved, and high-quality data sets and accurate image classification models are realized.

CN114676276BActive Publication Date: 2025-05-23MULTIPOINT (SHENZHEN) DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210224634.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-09
Publication Date
2025-05-23
Estimated Expiration
2042-03-09

AI Technical Summary

Technical Problem

The prior art has problems of missed detection, missed detection, low cleaning efficiency and loss of dirty and uncertain data during the cleaning process of image classification data sets.

Method used

By obtaining the initial set of training image samples and verifying the image sample set, the initial image classification model and classification accuracy are determined based on the initial set, the original model is trained using the training image sample set, the classification label of the training image sample is updated, and the uncertain data is added to the cleaning set, and iterative cleaning is performed to improve the quality of the data set.

Benefits of technology

Automatic data set cleaning is realized, cleaning efficiency and data set quality are improved, thereby improving the classification accuracy of image classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114676276B_ABST
    Figure CN114676276B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose a data set cleaning method, device, electronic device and computer-readable medium. A specific implementation of the method includes: obtaining an initial set of training image samples and a verification image sample set; using the initial set of training image samples to determine an initial image classification model, an initial classification accuracy and a training image sample set; determining the initial image classification model as a reference image classification model, and determining the value of the initial classification accuracy as the value of the reference classification accuracy, and determining the initial set of training image samples as a reference set of training image samples; using the training image sample set to perform iterative training, and determining a target image classification model and a target training image sample set. This implementation can automatically clean dirty data in the original image classification data set, improve cleaning efficiency and the quality of the data set, and thus improve the accuracy of the image classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a data set cleaning method, device, electronic device, and computer-readable medium. Background Art

[0002] Dataset cleaning is a technology that removes dirty data or corrects errors in a dataset to improve the quality of the dataset. The quality of the dataset largely determines the accuracy of the image classification model. Currently, when cleaning image classification datasets containing dirty data, the usual method is to remove or correct the dirty data through manual review.

[0003] However, when using the above method to clean the data set, the following technical problems often occur:

[0004] First, there is a large amount of data in the dataset, and manual review may result in missed detections and false detections, which results in a large proportion of dirty data in the dataset, leading to lower dataset quality.

[0005] Second, manual review and cleaning of dirty data sets is inefficient;

[0006] Third, dirty data in the data set is directly discarded without further processing, resulting in the loss of some data that is dirty but the true category of which can be known;

[0007] Fourth, the uncertain data is directly discarded without further processing the data in the data set that is uncertain whether it is dirty data, resulting in the loss of some data that is currently uncertain whether it is dirty data but its true category may be determined in the subsequent cleaning process. Summary of the invention

[0008] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0009] Some embodiments of the present disclosure propose a data set cleaning method, an apparatus, an electronic device, and a computer-readable medium to solve one or more of the technical problems mentioned in the above background technology section.

[0010] In a first aspect, some embodiments of the present disclosure provide a data set cleaning method, the method comprising: obtaining an initial set of training image samples and a verification image sample set, wherein the training image samples in the initial set of training image samples include training images and classification labels, and the verification image samples in the verification image sample set include verification images and classification labels; determining an initial image classification model, an initial classification accuracy and a training image sample set based on the initial set of training image samples; determining the initial image classification model as a reference image classification model, and determining the value of the initial classification accuracy as a reference classification accuracy value, and determining the initial set of training image samples as a reference set of training image samples; using the training image sample set, performing the following training steps: using the training image sample set to train the original image classification model to obtain an image classification model; determining the classification accuracy of the image classification model for the verification image sample set; in response to determining that the classification accuracy is less than or equal to the reference classification accuracy, determining the reference image classification model as a target image classification model, and determining the reference set of training image samples as a target training image sample set.

[0011] In some embodiments, the cleaning of the training image sample set using the image classification model further includes:

[0012] In response to determining that the classification label included in the target classification information is different from the classification label included in the training image sample, and the classification probability included in the target classification information is greater than or equal to a set first threshold, the classification label included in the training image sample is updated to the classification label included in the target classification information, and the updated training image sample is added to the cleaned training image sample set as a cleaned training image sample, wherein the first threshold is greater than or equal to 0.8.

[0013] In some embodiments, the cleaning of the training image sample set using the image classification model further includes:

[0014] In response to determining that the classification label included in the target classification information is different from the classification label included in the training image sample, and the classification probability included in the target classification information is less than the first threshold, and the classification probability corresponding to the classification label included in the training image sample is greater than or equal to a set second threshold, the training image sample is added to the cleaning training image sample set as a cleaning training image sample, wherein the second threshold is less than or equal to 0.1.

[0015] In a second aspect, some embodiments of the present disclosure provide a data set cleaning device, the device comprising: an acquisition unit, configured to acquire an initial set of training image samples and a verification image sample set, wherein the training image samples in the initial set of training image samples include training images and classification labels, and the verification image samples in the verification image sample set include verification images and classification labels; a determination unit, configured to determine an initial image classification model, an initial classification accuracy and a training image sample set based on the initial set of training image samples; a conversion unit, configured to determine the initial image classification model as a reference image classification model, and to determine the value of the initial classification accuracy as a reference classification accuracy value, and to determine the initial set of training image samples as a reference set of training image samples; a training unit, configured to use the training image sample set to perform the following training steps: use the training image sample set to train the original image classification model to obtain an image classification model; determine the classification accuracy of the image classification model for the verification image sample set; in response to determining that the classification accuracy is less than or equal to the reference classification accuracy, determine the reference image classification model as a target image classification model, and determine the reference set of training image samples as a target training image sample set.

[0016] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.

[0017] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above-mentioned first aspect is implemented.

[0018] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the data set cleaning method of some embodiments of the present disclosure, the dirty data in the data set can be automatically cleaned, the cleaning efficiency and the quality of the data set can be improved, so that a reliable image classification model can be trained using the cleaned data set. Specifically, the reasons for the low efficiency of data set cleaning and the high content of dirty data are: there are many data in the data set, the manual review speed is slow, and there will be missed detection and false detection. Based on this, the data set cleaning method of some embodiments of the present disclosure first obtains the initial set of training image samples and the verification image sample set, wherein the training image samples in the initial set of training image samples include training images and classification labels, and the verification image samples in the verification image sample set include verification images and classification labels. Then, based on the initial set of training image samples, the initial image classification model, the initial classification accuracy and the training image sample set are determined. Thus, the image classification model trained by the original data set with a high content of dirty data and the classification accuracy of the model are determined. Then, the initial image classification model is determined as the reference image classification model, and the value of the initial classification accuracy is determined as the value of the reference classification accuracy, and the initial set of training image samples is determined as the reference set of training image samples. Thus, it is convenient to compare the accuracy of the image classification model trained by the cleaned data set with the accuracy of the image classification model trained by the original data set. Finally, the following training steps are performed using the training image sample set: the original image classification model is trained using the training image sample set to obtain an image classification model; the classification accuracy of the image classification model for the above-mentioned verification image sample set is determined; in response to determining that the classification accuracy is less than or equal to the reference classification accuracy, the reference image classification model is determined as the target image classification model, and the training image sample reference set is determined as the target training image sample set. Thus, after iterative cleaning of the data set, the image classification model with the highest classification accuracy for the image can be quickly selected. Therefore, the above-mentioned various embodiments of the present disclosure can quickly and automatically clean dirty data in the image classification data set, improve the cleaning efficiency and the quality of the data set, and thus improve the classification accuracy of the image classification model for the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0020] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;

[0021] Figure 2is a flow chart of some embodiments of a data set cleaning method according to the present disclosure;

[0022] Figure 3 is a schematic diagram of the structure of some embodiments of the data set cleaning device according to the present disclosure;

[0023] Figure 4 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0024] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0025] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0026] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0027] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0028] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0029] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0030] Figure 1 An exemplary system architecture 100 is shown in which a method for data set cleaning and an apparatus for data set cleaning according to an embodiment of the present application can be applied.

[0031] like Figure 1As shown, the system architecture 100 may include terminals 101, 102, a network 103, a database server 104, and a server 105. The network 103 is used to provide a medium for communication links between the terminals 101, 102, the database server 104, and the server 105. The network 103 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.

[0032] The user 106 can use the terminals 101 and 102 to interact with the server 105 through the network 103 to receive or send messages, etc. Various client applications can be installed on the terminals 101 and 102, such as model training applications, web browsers, and instant messaging tools.

[0033] The terminals 101 and 102 here can be hardware or software. When the terminals 101 and 102 are hardware, they can be various electronic devices with display screens, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III, Dynamic Image Experts Compression Standard Audio Layer 3), laptops and desktop computers, etc. When the terminals 101 and 102 are software, they can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules (for example, to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0034] When the terminals 101 and 102 are hardware, an image acquisition device may be installed thereon. The image acquisition device may be any device capable of acquiring images, such as a camera, a sensor, etc. The user 106 may use the image acquisition device on the terminals 101 and 102 to acquire images of objects.

[0035] The database server 104 may be a database server that provides various services. For example, the database server may store an initial set of training image samples and a verification image sample set. The training image samples in the initial set of training image samples include training images and classification labels, and the verification image samples in the verification image sample set include verification images and classification labels. In this way, the user 106 may also select samples from the sample set stored in the database server 104 through the terminals 101 and 102.

[0036] The server 105 may also be a server that provides various services, such as a background server that provides support for various applications displayed on the terminals 101 and 102. The background server may use the training image samples in the initial set of training image samples sent by the terminals 101 and 102 to iteratively train the initial image classification model, and may send the training results (such as the category labels of the output training images) to the terminals 101 and 102. In this way, the user 106 may apply the category labels of the output training images.

[0037] The database server 104 and the server 105 here can also be hardware or software. When they are hardware, they can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When they are software, they can be implemented as multiple software or software modules (for example, for providing distributed services), or as a single software or software module. No specific limitation is made here.

[0038] It should be noted that the method for cleaning a data set provided in the embodiment of the present application is generally executed by the server 105. Accordingly, the device for cleaning a data set is generally also disposed in the server 105.

[0039] It should be pointed out that, in the case where the server 105 can implement the related functions of the database server 104 , the database server 104 may not be provided in the system architecture 100 .

[0040] It should be understood that Figure 1 The number of terminals, networks, database servers and servers in the embodiment is only for illustration. Any number of terminals, networks, database servers and servers may be provided according to the implementation requirements.

[0041] Continue to refer Figure 2 , shows a process 200 of some embodiments of the data set cleaning method according to the present disclosure. The process 200 of the data set cleaning method may include the following steps:

[0042] Step 201, obtaining an initial set of training image samples and a verification image sample set.

[0043] In some embodiments, the execution body of the data set cleaning method (such as Figure 1 The server 105 shown in the figure can obtain the initial set of training image samples and the verification image sample set in various ways. For example, the execution subject can obtain the initial set of training image samples and the verification image sample set from the database server (for example, Figure 1 The database server 104 shown in FIG. 104 is used to obtain the existing training image sample initial set and verification image sample set stored therein. For another example, the user can use a terminal (such as Figure 1In this way, the execution subject can receive the samples collected by the terminal and store these samples locally, thereby generating an initial set of training image samples and a verification image sample set.

[0044] Among them, the initial set of training image samples may include at least one training image sample. The training image samples may include training images and classification labels. The training images may display objects, and the classification labels may be the categories of the objects. For example, the objects in the training images may display apples, and the classification labels may be apples. The verification image sample set may include all classification labels, and each classification label may include at least one verification image. The verification image samples in the verification image sample set include verification images and classification labels. The verification images may display objects, and the classification labels may be the categories of the objects. For example, the objects in the verification images may display pears, and the classification labels may be pears.

[0045] Step 202: Determine an initial image classification model, initial classification accuracy, and a training image sample set based on an initial set of training image samples.

[0046] In some embodiments, the execution subject determines the initial image classification model, the initial classification accuracy and the training image sample set based on the initial set of training image samples, which may include the following steps:

[0047] In the first step, the original image classification model is trained using the initial set of training image samples to obtain an initial image classification model.

[0048] The original image classification model may be a model of various structures, such as ResNet (Residual Neural Network) and MobileNet (Efficient Convolutional Neural Networks for Mobile Vision Applications). This application does not make any specific limitation on this.

[0049] In some embodiments, there are many methods for training the original image classification model using the above-mentioned initial set of training image samples, such as SGD (Stochastic Gradient Descent) and Adam (Adaptive Momentum Estimation), etc. This application does not make specific limitations on this.

[0050] The second step is to determine the classification accuracy of the initial image classification model for the verification image sample set to obtain the initial classification accuracy.

[0051] In some embodiments, the verification image of the verification image sample in the verification image sample set can be input into the initial image classification model, and the classification accuracy of the initial image classification model for the verification image sample set can be determined according to the output result of the image classification model. For example, there are m output results of the initial image classification model, among which n (n≤m) are correct. Then the classification accuracy of the initial image classification model for the verification image sample set is

[0052] As an example, there are 50 output results of the above initial image classification model, of which 38 are correct. Therefore, the classification accuracy of the above initial image classification model for the above verification image sample set is 76%.

[0053] The third step is to use the initial image classification model to clean the initial set of training image samples to obtain a training image sample set.

[0054] In some embodiments, using the initial image classification model to clean the initial set of training image samples to obtain a training image sample set may include the following steps:

[0055] In the first step, the following cleaning steps are performed on each training image sample in the initial set of training image samples:

[0056] In the first sub-step, the training images in the training image samples are input into the initial image classification model to obtain a classification information group.

[0057] The classification information in the classification information group may include classification labels and classification probabilities. The sum of the classification probabilities included in each classification information in the classification information group is 1.

[0058] As an example, the classification information in the above classification information group may be "apple: 0.81, pear: 0.09, orange: 0.07, and persimmon: 0.03".

[0059] The second sub-step is to select the classified information that meets the preset conditions from the above-mentioned classified information group as the target classified information.

[0060] The preset condition may be that the classification probability included in the classification information is the largest classification probability among the classification probabilities included in the classification information group.

[0061] As an example, the classification information "apple: 0.81" may be selected as the target classification information from the classification information group "apple: 0.81, pear: 0.09, orange: 0.07, and persimmon: 0.03".

[0062] The third sub-step is, in response to determining that the classification label included in the target classification information is the same as the classification label included in the training image sample, adding the training image sample as a cleaning training image sample to the cleaning training image sample set.

[0063] As an example, if the target classification information includes a classification label of apple, and the training image sample also includes a classification label of apple, then the training image sample is clean data. The training image sample is added as a cleaned training image sample to the cleaned training image sample set.

[0064] The fourth sub-step, in response to determining that the classification label included in the above-mentioned target classification information is different from the classification label included in the above-mentioned training image sample, and the classification probability included in the above-mentioned target classification information is greater than or equal to the set first threshold, updates the classification label included in the above-mentioned training image sample to the classification label included in the above-mentioned target classification information, and adds the updated training image sample as a cleaned training image sample to the cleaned training image sample set.

[0065] The larger the first threshold is, the greater the possibility that updating the classification label included in the training image sample to the classification label included in the target classification information is correct. Generally, the first threshold can be set to a value greater than or equal to 0.8.

[0066] As an example, the target classification information is "apple: 0.81". The classification label included in the target classification information is apple, and the classification probability included in the target classification information is 0.81. The classification label included in the training image sample is pear. The first threshold is set to 0.8. The above training image sample is dirty data, and the true classification label is apple. The classification label included in the above training image sample is updated to apple, and the updated training image sample is added to the cleaned training image sample set as a cleaned training image sample.

[0067] The above-mentioned fourth sub-step and its related content, as an inventive point of an embodiment of the present disclosure, solve the third technical problem mentioned in the background technology, "directly discarding the dirty data in the data set without further processing the dirty data in the data set, resulting in the loss of some data that are dirty data but whose true categories are known". The factors that further lead to the loss of some data that are dirty data but whose true categories are known are as follows: no further processing is performed on this part of the data. If the above-mentioned factors are resolved, the effect of reducing data loss can be achieved. In order to achieve this effect, the present disclosure will update the classification labels included in the training image samples that are determined to be dirty data and whose true classification labels are known to be the classification labels included in the target classification information, so that the category of the dirty data can be changed to the true category. Therefore, data loss can be reduced.

[0068] The fifth sub-step, in response to determining that the classification label included in the above-mentioned target classification information is different from the classification label included in the above-mentioned training image sample, and the classification probability included in the above-mentioned target classification information is less than the above-mentioned first threshold, and the classification probability corresponding to the classification label included in the above-mentioned training image sample is greater than or equal to the set second threshold, adds the above-mentioned training image sample as a cleaning training image sample to the cleaning training image sample set.

[0069] The smaller the second threshold is, the more likely it is that the discarded training image samples are real dirty data. Generally, the second threshold can be set to be less than or equal to 0.1.

[0070] As an example, the target classification information is "Tomato: 0.56". The classification label included in the target classification information is Tomato, and the classification probability included in the target classification information is 0.56. The classification label included in the training image sample is Pear, and the classification probability is 0.61. The first threshold is set to 0.8. The second threshold is set to 0.1. It cannot be determined whether the above training image sample is dirty data. The above training image sample is added as a clean training image sample to the clean training image sample set.

[0071] The above-mentioned fifth sub-step and its related content, as an inventive point of an embodiment of the present disclosure, solve the technical problem four mentioned in the background technology, "directly discarding uncertain data without further processing the data in the data set that is uncertain whether it is dirty data, resulting in the loss of some data that is currently uncertain whether it is dirty data, but its true category may be determined in the subsequent cleaning process." The factors that further lead to the loss of some data that is currently uncertain whether it is dirty data, but its true category may be determined in the subsequent cleaning process are as follows: no further processing is performed on this part of the data. If the above factors are resolved, the effect of reducing data loss can be achieved. In order to achieve this effect, the present disclosure adds data that is uncertain whether it is dirty data to the cleaning training image sample set, so that it can be further determined in the next cleaning. Therefore, data loss can be reduced.

[0072] The second step is to determine the cleaned training image sample set as the training image sample set.

[0073] The training image sample set may include clean data, training image samples whose incorrect classification labels are updated to true classification labels, and uncertain training image samples.

[0074] Step 203, determining the initial image classification model as the reference image classification model, determining the value of the initial classification accuracy as the value of the reference classification accuracy, and determining the initial set of training image samples as the reference set of training image samples.

[0075] In some embodiments, the execution entity may determine the initial image classification model as a reference image classification model, determine the initial classification accuracy value as a reference classification accuracy value, and determine the initial set of training image samples as a reference set of training image samples.

[0076] Step 204, performing a training step using the training image sample set.

[0077] In some embodiments, the execution subject may use the training image sample set to perform the following training steps:

[0078] In step 2041, the execution subject may train the original image classification model using the training image sample set to obtain an image classification model.

[0079] The original image classification model may be a model of various structures, such as ResNet (Residual Neural Network) and MobileNet (Efficient Convolutional Neural Networks for Mobile Vision Applications). This application does not make any specific limitation on this.

[0080] In some embodiments, there are many methods for training the original image classification model using the training image sample set, such as SGD and Adam, which are not specifically limited in this application.

[0081] Step 2042, determining the classification accuracy of the image classification model for the validation image sample set.

[0082] In some embodiments, the execution entity may determine the classification accuracy of the image classification model for the validation image sample set.

[0083] In some embodiments, the verification image of the verification image sample in the verification image sample set can be input into the initial image classification model, and the classification accuracy of the initial image classification model for the verification image sample set can be determined according to the output result of the image classification model. For example, there are m output results of the initial image classification model, among which n (n≤m) are correct. Then the classification accuracy of the initial image classification model for the verification image sample set is

[0084] As an example, there are 50 output results of the above initial image classification model, of which 43 are correct. Therefore, the classification accuracy of the above initial image classification model for the above verification image sample set is 86%.

[0085] Step 2043, in response to determining that the classification accuracy is less than or equal to the reference classification accuracy, determining the reference image classification model as the target image classification model, and determining the training image sample reference set as the target training image sample set.

[0086] In some embodiments, in response to determining that the classification accuracy is less than or equal to the reference classification accuracy, the execution entity may determine the reference image classification model as the target image classification model, and determine the training image sample reference set as the target training image sample set.

[0087] As an example, the classification accuracy value is 0.83, and the reference classification accuracy value is 0.85. It can be determined that the classification accuracy is less than or equal to the reference classification accuracy.

[0088] Optionally, in response to determining that the classification accuracy is greater than the reference classification accuracy, the training image sample set is used as a training image sample reference set, the training image sample set is cleaned using the image classification model, the cleaned training image sample set is used as the training image sample set, the image classification model is used as the reference image classification model, and the value of the reference classification accuracy is updated to the value of the classification accuracy, and the above-mentioned training step 204 is continued.

[0089] As an example, the classification accuracy value is 0.87, and the reference classification accuracy value is 0.85, so it can be determined that the classification accuracy is greater than the initial classification accuracy.

[0090] In some embodiments, cleaning the training image sample set using the image classification model may include the following steps:

[0091] In the first step, the following cleaning steps are performed on each training image sample in the training image sample set:

[0092] In the first sub-step, the training images in the training image samples are input into the image classification model to obtain a classification information group.

[0093] The classification information in the classification information group may include classification labels and classification probabilities. The sum of the classification probabilities included in each classification information in the classification information group is 1.

[0094] As an example, the classification information in the above classification information group may be "apple: 0.81, pear: 0.09, orange: 0.07, and persimmon: 0.03".

[0095] The second sub-step is to select the classified information that meets the preset conditions from the above-mentioned classified information group as the target classified information.

[0096] The preset condition may be that the classification probability included in the classification information is the largest classification probability among the classification probabilities included in the classification information group.

[0097] As an example, the classification information "apple: 0.81" is selected as the target classification information from the classification information group "apple: 0.81, pear: 0.09, orange: 0.07, and persimmon: 0.03".

[0098] The third sub-step is, in response to determining that the classification label included in the target classification information is the same as the classification label included in the training image sample, adding the training image sample as a cleaning training image sample to the cleaning training image sample set.

[0099] As an example, if the target classification information includes a classification label of apple, and the training image sample also includes a classification label of apple, then the training image sample is clean data. The training image sample is added as a cleaned training image sample to the cleaned training image sample set.

[0100] The fourth sub-step, in response to determining that the classification label included in the above-mentioned target classification information is different from the classification label included in the above-mentioned training image sample, and the classification probability included in the above-mentioned target classification information is greater than or equal to the set first threshold, updates the classification label included in the above-mentioned training image sample to the classification label included in the above-mentioned target classification information, and adds the updated training image sample as a cleaned training image sample to the cleaned training image sample set.

[0101] The larger the first threshold is, the greater the possibility that updating the classification label included in the training image sample to the classification label included in the target classification information is correct. Generally, the first threshold can be set to be greater than or equal to 0.8.

[0102] As an example, the target classification information is "apple: 0.81". The classification label included in the target classification information is apple, and the classification probability included in the target classification information is 0.81. The classification label included in the training image sample is pear. The first threshold is set to 0.8. The above training image sample is dirty data, and the true classification label is apple. The classification label included in the above training image sample is updated to apple, and the updated training image sample is added to the cleaned training image sample set as a cleaned training image sample.

[0103] The fifth sub-step, in response to determining that the classification label included in the above-mentioned target classification information is different from the classification label included in the above-mentioned training image sample, and the classification probability included in the above-mentioned target classification information is less than the above-mentioned first threshold, and the classification probability corresponding to the classification label included in the above-mentioned training image sample is greater than or equal to the set second threshold, adds the above-mentioned training image sample as a cleaning training image sample to the cleaning training image sample set.

[0104] The smaller the second threshold is, the greater the possibility that the dirty data to be removed is real dirty data. Generally, the second threshold can be set to be less than or equal to 0.1.

[0105] As an example, the target classification information is "Tomato: 0.56". The classification label included in the target classification information is Tomato, and the classification probability included in the target classification information is 0.56. The classification label included in the training image sample is Pear, and the classification probability is 0.61. The first threshold is set to 0.8. The second threshold is set to 0.1. It cannot be determined whether the above training image sample is dirty data. The above training image sample is added as a clean training image sample to the clean training image sample set.

[0106] The second step is to determine the cleaned training image sample set as the cleaned training image sample set.

[0107] The training image sample set may include clean data, training image samples whose incorrect classification labels are updated to true classification labels, and uncertain training image samples.

[0108] Therefore, the solutions described in these embodiments can automatically clean dirty data in the original image classification data set, improve the cleaning efficiency and the quality of the data set, and thus improve the classification accuracy of the image classification model.

[0109] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the data set cleaning method of some embodiments of the present disclosure, the dirty data in the data set can be automatically cleaned, the cleaning efficiency and the quality of the data set can be improved, so that a reliable image classification model can be trained using the cleaned data set. Specifically, the reasons for the low efficiency of data set cleaning and the high content of dirty data are: there are many data in the data set, the manual review speed is slow, and there will be missed detection and false detection. Based on this, the data set cleaning method of some embodiments of the present disclosure first obtains the initial set of training image samples and the verification image sample set, wherein the training image samples in the initial set of training image samples include training images and classification labels, and the verification image samples in the verification image sample set include verification images and classification labels. Then, based on the initial set of training image samples, the initial image classification model, the initial classification accuracy and the training image sample set are determined. Thus, the image classification model trained by the original data set with a high content of dirty data and the classification accuracy of the model are determined. Then, the initial image classification model is determined as the reference image classification model, and the value of the initial classification accuracy is determined as the value of the reference classification accuracy, and the initial set of training image samples is determined as the reference set of training image samples. Thus, it is convenient to compare the accuracy of the image classification model trained by the cleaned data set with the accuracy of the image classification model trained by the original data set. Finally, the following training steps are performed using the training image sample set: the original image classification model is trained using the training image sample set to obtain an image classification model; the classification accuracy of the image classification model for the above-mentioned verification image sample set is determined; in response to determining that the classification accuracy is less than or equal to the reference classification accuracy, the reference image classification model is determined as the target image classification model, and the training image sample reference set is determined as the target training image sample set. Thus, after iterative cleaning of the data set, the image classification model with the highest classification accuracy for the image can be quickly selected. Therefore, the above-mentioned various embodiments of the present disclosure can quickly and automatically clean dirty data in the image classification data set, improve the cleaning efficiency and the quality of the data set, and thus improve the classification accuracy of the image classification model for the image.

[0110] Further references Figure 3 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a data set cleaning device. These device embodiments are similar to Figure 2 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0111] like Figure 3 As shown, a data set cleaning device 300 of some embodiments includes: an acquisition unit 301 , a determination unit 302 , a conversion unit 303 and a training unit 304 . Among them, the acquisition unit 301 is configured to acquire an initial set of training image samples and a verification image sample set, wherein the training image samples in the above-mentioned initial set of training image samples include training images and classification labels, and the verification image samples in the above-mentioned verification image sample set include verification images and classification labels; the determination unit 302 is configured to determine an initial image classification model, an initial classification accuracy and a training image sample set based on the above-mentioned initial set of training image samples; the conversion unit 303 is configured to determine the above-mentioned initial image classification model as a reference image classification model, and determine the value of the above-mentioned initial classification accuracy as a reference classification accuracy value, and determine the above-mentioned initial set of training image samples as a reference set of training image samples; the training unit 304 is configured to use the above-mentioned training image sample set to perform the following training steps: train the original image classification model using the training image sample set to obtain an image classification model; determine the classification accuracy of the image classification model for the above-mentioned verification image sample set; in response to determining that the classification accuracy is less than or equal to the reference classification accuracy, determine the reference image classification model as a target image classification model, and determine the training image sample reference set as a target training image sample set.

[0112] It is understood that the units described in the device 300 are similar to those described in the reference Figure 2 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 300 and the units contained therein, and will not be described in detail here.

[0113] Reference below Figure 4 , which shows an electronic device 400 (eg, Figure 1 Schematic diagram of the structure of the electronic equipment in FIG. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0114] like Figure 4As shown, the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0115] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The electronic device 400 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 4 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0116] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0117] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0118] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0119] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains an initial set of training image samples and a verification image sample set, wherein the training image samples in the initial set of training image samples include training images and classification labels, and the verification image samples in the verification image sample set include verification images and classification labels; determines an initial image classification model, an initial classification accuracy and a training image sample set using the initial set of training image samples; determines the initial image classification model as a reference image classification model, determines the value of the initial classification accuracy as a value of the reference classification accuracy, and determines the initial set of training image samples as a reference set of training image samples; performs the following training steps using the training image sample set: trains the original image classification model using the training image sample set to obtain an image classification model; determines the classification accuracy of the image classification model for the verification image sample set; in response to determining that the classification accuracy is less than or equal to the reference classification accuracy, determines the reference image classification model as a target image classification model, and determines the reference set of training image samples as a target training image sample set.

[0120] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0121] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0122] The units described in some embodiments of the present disclosure may be implemented by software or hardware. The described units may also be provided in a processor, for example, may be described as: a processor including an acquisition unit, a determination unit, a conversion unit, and a training unit. The names of these units do not, in some cases, constitute limitations on the units themselves, for example, the acquisition unit may also be described as a "unit for acquiring an initial set of training image samples and a verification image sample set".

[0123] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

Claims

1. A data set cleaning method, include: Acquire an initial set of training image samples and a verification image sample set, wherein the training image samples in the initial set of training image samples include training images and classification labels, and the verification image samples in the verification image sample set include verification images and classification labels; Determining an initial image classification model, an initial classification accuracy, and a training image sample set based on the initial set of training image samples; Determining the initial image classification model as a reference image classification model, determining the value of the initial classification accuracy as a reference classification accuracy value, and determining the initial set of training image samples as a reference set of training image samples; Using the training image sample set, perform the following training steps: The original image classification model is trained using the training image sample set to obtain an image classification model; Determining the classification accuracy of the image classification model for the verification image sample set; In response to determining that the classification accuracy is less than or equal to the reference classification accuracy, determining the reference image classification model as the target image classification model, and determining the reference set of training image samples as the target training image sample set; The initial image classification model is used to clean the initial set of training image samples to obtain a training image sample set, which includes the following steps: In the first step, the following cleaning steps are performed on each training image sample in the initial set of training image samples: The first sub-step is to input the training images in the training image samples into the initial image classification model to obtain a classification information group; The second sub-step is to select the classified information satisfying the preset conditions from the classified information group as the target classified information; A third sub-step, in response to determining that the classification label included in the target classification information is the same as the classification label included in the training image sample, adding the training image sample as a cleaning training image sample to a cleaning training image sample set; In a fourth sub-step, in response to determining that the classification label included in the target classification information is different from the classification label included in the training image sample, and the classification probability included in the target classification information is greater than or equal to a set first threshold, the classification label included in the training image sample is updated to the classification label included in the target classification information, and the updated training image sample is added as a cleaning training image sample to the cleaning training image sample set; A fifth sub-step, in response to determining that the classification label included in the target classification information is different from the classification label included in the training image sample, and the classification probability included in the target classification information is less than the first threshold, and the classification probability corresponding to the classification label included in the training image sample is greater than or equal to a set second threshold, adding the training image sample as a cleaning training image sample to the cleaning training image sample set; The second step is to determine the cleaned training image sample set as the training image sample set.

2. The method according to claim 1, in, The method further comprises: In response to determining that the classification accuracy is greater than the reference classification accuracy, the training image sample set is used as a training image sample reference set, the training image sample set is cleaned using the image classification model, the cleaned training image sample set is used as the training image sample set, the image classification model is used as the reference image classification model, and the value of the reference classification accuracy is updated to the value of the classification accuracy, and the training step is continued.

3. The method according to claim 2, in, The method of using the image classification model to clean the training image sample set includes: Perform the following cleaning steps on each training image sample in the training image sample set: Inputting the training images in the training image samples into the image classification model to obtain a classification information group, wherein the classification information in the classification information group includes a classification label and a classification probability, and the sum of the classification probabilities included in each classification information in the classification information group is 1; Selecting classification information that meets a preset condition from the classification information group as target classification information, wherein the preset condition is that the classification probability included in the classification information is the largest classification probability among the classification probabilities included in the classification information group; In response to determining that the classification label included in the target classification information is the same as the classification label included in the training image sample, the training image sample is added to the cleaned training image sample set as a cleaned training image sample.

4. The method according to claim 3, in, The cleaning of the training image sample set by using the image classification model also includes: The cleaned training image sample set is determined to be the cleaned training image sample set.

5. The method according to claim 1, in, The method of using the initial set of training image samples to determine an initial image classification model, an initial classification accuracy, and a training image sample set includes: Using the initial set of training image samples to train the original image classification model to obtain an initial image classification model; Determining the classification accuracy of the initial image classification model for the verification image sample set to obtain an initial classification accuracy; The initial image classification model is used to clean the initial set of training image samples to obtain a training image sample set.

6. A data set cleaning device, include: an acquisition unit configured to acquire an initial set of training image samples and a verification image sample set, wherein the training image samples in the initial set of training image samples include training images and classification labels, and the verification image samples in the verification image sample set include verification images and classification labels; a determination unit configured to determine an initial image classification model, an initial classification accuracy, and a training image sample set based on the initial set of training image samples; a conversion unit configured to determine the initial image classification model as a reference image classification model, determine the value of the initial classification accuracy as a reference classification accuracy value, and determine the initial set of training image samples as a reference set of training image samples; The training unit is configured to perform the following training steps using the training image sample set: The original image classification model is trained using the training image sample set to obtain an image classification model; Determining the classification accuracy of the image classification model for the verification image sample set; In response to determining that the classification accuracy is less than or equal to the reference classification accuracy, determining the reference image classification model as the target image classification model, and determining the reference set of training image samples as the target training image sample set; The initial image classification model is used to clean the initial set of training image samples to obtain a training image sample set, which includes the following steps: In the first step, the following cleaning steps are performed on each training image sample in the initial set of training image samples: The first sub-step is to input the training images in the training image samples into the initial image classification model to obtain a classification information group; The second sub-step is to select the classified information satisfying the preset conditions from the classified information group as the target classified information; A third sub-step, in response to determining that the classification label included in the target classification information is the same as the classification label included in the training image sample, adding the training image sample as a cleaning training image sample to a cleaning training image sample set; In a fourth sub-step, in response to determining that the classification label included in the target classification information is different from the classification label included in the training image sample, and the classification probability included in the target classification information is greater than or equal to a set first threshold, the classification label included in the training image sample is updated to the classification label included in the target classification information, and the updated training image sample is added as a cleaning training image sample to the cleaning training image sample set; A fifth sub-step, in response to determining that the classification label included in the target classification information is different from the classification label included in the training image sample, and the classification probability included in the target classification information is less than the first threshold, and the classification probability corresponding to the classification label included in the training image sample is greater than or equal to a set second threshold, adding the training image sample as a cleaning training image sample to the cleaning training image sample set; The second step is to determine the cleaned training image sample set as the training image sample set.

7. An electronic device, include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

8. A computer readable medium having a computer program stored thereon, in, When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Sample data cleaning method, device, computer device, and storage medium

    CN109241903A

  • Data cleaning method, device and equipment and storage medium

    CN111667003A