Sample Data Processing Method, Device, Storage Medium, and Electronic Device
Through cross-training and error rate threshold update methods on sample data sets, the problems of waste of human resources and incomplete improvement of data quality caused by manual circular cleaning in the prior art are solved, and efficient and automated sample data quality improvement is achieved.
Patent Information
- Application Number
- CN202110184725.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-02-10
AI Technical Summary
When improving the quality of sample data, the existing technology requires a large amount of human resources to perform manual circular cleaning, and lacks effective data cleaning indicators, resulting in the problem of waste of manpower and incomplete cleaning.
By using the first sample data set to train the target model, the error rate of each sample data and the accuracy of the data set are obtained, and the sample data with an error rate higher than the threshold are eliminated, the second sample data set is formed, and the data quality is gradually improved through cross-training and error rate threshold update.
It realizes efficient positioning and removing abnormal sample data, improves the quality of sample data, avoids the waste of human resources caused by manual circular cleaning, and provides quantifiable indicators for improving data quality.
Smart Images

Figure CN113780323B_ABST
Abstract
Description
Background Art
[0002] With the continuous development of deep learning algorithms, deep learning algorithms are widely used in scenarios such as image recognition and natural language processing. Since the learning effect of deep learning algorithms is closely related to the quality of the sample data being learned, it is very important to improve the quality of sample data.
[0003] Currently, the main method to improve data quality is to manually and repeatedly clean the sample data with incorrect output results of the deep learning model to improve the quality of sample data. However, this not only causes a large amount of waste of human resources, but also there is no metric to stop data cleaning.
[0004] It should be noted that the information disclosed in the above Background Art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The present disclosure provides a method for processing sample data, a device for processing sample data, a computer-readable storage medium, and an electronic device, thereby at least to some extent solving the problem of excessively high human resource costs required in related technologies while improving the quality of sample data.
[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.
[0007] According to a first aspect of the present disclosure, there is provided a method for processing sample data, including: training a target model using a first sample data set to obtain the error rate of each sample data in the first sample data set and the accuracy rate of the first sample data set; removing the sample data with an error rate higher than an error rate threshold from the first sample data set to obtain a second sample data set; training the target model using the second sample data set to obtain the accuracy rate of the second sample data set; when the accuracy rate of the second sample data set is greater than or equal to the accuracy rate of the first sample data set, outputting the second sample data set.
[0008] In an exemplary embodiment of the present disclosure, the training the target model using the first sample data set to obtain the error rate of each sample data in the first sample data set and the accuracy rate of the first sample data set includes: performing cross-training on the target model through the first sample data set to obtain prediction data of each sample data in the first sample data set; and obtaining the error rate of each sample data in the first sample data set and the accuracy rate of the first sample data set according to the prediction data of each sample data in the first sample data set.
[0009] In an exemplary embodiment of the present disclosure, the cross-training of the target model using the first sample data set to obtain prediction data for each sample data in the first sample data set includes: splitting the first sample data set into n sample data subsets, using each of the sample data subsets as a test set, and using the remaining sample data subsets as a training set to perform n times of cross-training on the target model.
[0010] In an exemplary embodiment of the present disclosure, the cross-training of the target model using the first sample data set to obtain prediction data for each sample data in the first sample data set further includes: iteratively performing the cross-training until the number of times of the cross-training reaches a preset number of times.
[0011] In an exemplary embodiment of the present disclosure, the method further includes: when the accuracy rate of the second sample data set is less than the accuracy rate of the first sample data set, updating the second sample data set.
[0012] In an exemplary embodiment of the present disclosure, the updating the second sample data set when the accuracy rate of the second sample data set is less than the accuracy rate of the first sample data set includes: when the accuracy rate of the second sample data set is less than the accuracy rate of the first sample data set, updating the error rate threshold and jumping to the step of removing sample data with an error rate higher than the error rate threshold from the first sample data set to update the second sample data set.
[0013] In an exemplary embodiment of the present disclosure, the updating the error rate threshold includes: adding a preset step size to the error rate threshold as the updated error rate threshold.
[0014] In an exemplary embodiment of the present disclosure, the first sample data set includes any one of the following types of data: image data, text data, audio data.
[0015] According to a second aspect of the present disclosure, there is provided a sample data processing device, including: a first accuracy rate acquisition module, configured to train a target model using a first sample data set to obtain the error rate of each sample data in the first sample data set and the accuracy rate of the first sample data set; a sample data removal module, configured to remove sample data with an error rate higher than an error rate threshold from the first sample data set to obtain a second sample data set; a second accuracy rate acquisition module, configured to train the target model using the second sample data set to obtain the accuracy rate of the second sample data set; a sample data set output module, configured to output the second sample data set when the accuracy rate of the second sample data set is greater than or equal to the accuracy rate of the first sample data set.
[0016] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor, implements the above-described sample data processing method.
[0017] According to a fourth aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the above-described sample data processing method by executing the executable instructions.
[0018] The technical solution of the present disclosure has the following beneficial effects:
[0019] During the above sample data processing, sample data with an error rate higher than an error rate threshold is removed from the first sample data set to obtain a second sample data set. When the accuracy rate of the second sample data set is greater than or equal to the accuracy rate of the first sample data set, the second sample data set is output. By removing sample data with an error rate higher than the error rate threshold from the sample data set and then comparing the accuracy rate of the sample data set after removal with the accuracy rate of the original sample data set, the abnormal sample data removed is located, and a high-quality sample data set is obtained. When the accuracy rate is improved or remains the same as the original sample data, the prediction effect of the target model is slightly improved or remains unchanged, indicating that no abnormal problems can be detected and the quality of the sample data has been improved. This process avoids the waste of a large amount of human resources caused by manual cyclic data cleaning, can efficiently locate abnormal sample data, and thus improve the quality of sample data.
[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0022] Figure 1 A flowchart showing a sample data processing method in this exemplary embodiment;
[0023] Figure 2 A flowchart showing a method for obtaining the error rate of each sample data in a first sample data set and the accuracy rate of the first sample data set in this exemplary embodiment;
[0024] Figure 3 A flowchart showing a sample data removal in this exemplary embodiment;
[0025] Figure 4 The structural block diagram of a sample data processing device in this exemplary embodiment is shown;
[0026] Figure 5 An electronic device for implementing the above method in this exemplary embodiment is shown. Detailed implementation manners
[0027] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.
[0028] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0029] In this document, "first", "second", etc. are labels for specific objects and do not limit the quantity or order of the objects.
[0030] In the related art, problem sample data has brought about incorrect learning, resulting in inaccurate cleaning ranges of the sample data and requiring multiple manual loop cleanings. However, there is no indicator for stopping data cleaning, causing unnecessary waste of manpower.
[0031] In view of the above one or more problems, the exemplary embodiments of the present disclosure provide a sample data processing method.
[0032] Figure 1 The schematic flow of the sample data processing method in this exemplary embodiment is shown, including the following steps S110 to S140:
[0033] Step S110: Train the target model using the first sample dataset to obtain the error rate of each sample data in the first sample dataset and the accuracy rate of the first sample dataset.
[0034] Step S120: Remove the sample data with an error rate higher than the error rate threshold from the first sample dataset to obtain the second sample dataset.
[0035] Step S130: Train the target model using the second sample dataset to obtain the accuracy rate of the second sample dataset.
[0036] Step S140: When the accuracy rate of the second sample dataset is greater than or equal to the accuracy rate of the first sample dataset, output the second sample dataset.
[0037] During the above sample data processing, remove the sample data with an error rate higher than the error rate threshold from the first sample dataset to obtain the second sample dataset. When the accuracy rate of the second sample dataset is greater than or equal to the accuracy rate of the first sample dataset, output the second sample dataset. By removing the sample data with an error rate higher than the error rate threshold from the sample dataset and then comparing the accuracy rate of the removed sample dataset with the accuracy rate of the original sample dataset, the abnormal sample data to be removed is located, and a high-quality sample dataset is obtained. When the accuracy rate is improved or remains the same as the original sample data, the prediction effect of the target model is slightly improved or remains unchanged, indicating that no abnormal problems can be detected and the quality of the sample data has been improved. This process avoids the waste of a large amount of human resources caused by manual cyclic data cleaning, can efficiently locate the abnormal sample data, and thus improve the quality of the sample data.
[0038] The following will separately Figure 1 describe each step in
[0039] Step S110: Train the target model using the first sample dataset to obtain the error rate of each sample data in the first sample dataset and the accuracy rate of the first sample dataset.
[0040] The first sample data set refers to a set composed of each sample data, which may contain abnormal sample data. By using the sample data to train the target model, the target model can achieve data prediction. The target model can be, for example, a supervised neural network model. Each sample data in the first sample data set is labeled with a corresponding sample label. When the sample label is mislabeled, the corresponding sample data becomes abnormal sample data. The existence of abnormal sample data will lead to a decline in the data quality of the entire sample data set. The sample data with mislabeled labels may cause the target model to learn incorrectly, resulting in deviation in the data predicted by using the target model. Therefore, it is very necessary to improve the data quality of the first sample data set. The error rate of each sample data in the first sample data set refers to the probability that the predicted data obtained by using the target model for each sample data is inconsistent with the corresponding sample label. The accuracy rate of the first sample data set refers to the proportion of the sample data in the first sample data set for which the predicted result obtained by using the target model is inconsistent with the corresponding sample label.
[0041] In an alternative implementation manner, step S110 can obtain the error rate of each sample data in the first sample data set and the accuracy rate of the first sample data set through steps S210 to S220 as shown below. The specific implementation steps are as follows: Figure 2 In an alternative implementation manner, step S110 can obtain the error rate of each sample data in the first sample data set and the accuracy rate of the first sample data set through steps S210 to S220 as shown below. The specific implementation steps are as follows:
[0042] Step S210: Cross-train the target model with the first sample data set to obtain the predicted data of each sample data in the first sample data set.
[0043] The purpose of cross-training is to obtain a reliable and stable model. The specific method is to use most of the sample data for model training, and leave a small part of the samples to be predicted by the trained model. This process is cycled until all the samples have been predicted exactly once. By cyclically performing cross-training, the predicted data of all the sample data in the first sample data set can be obtained.
[0044] In an alternative implementation manner, in step S210, cross-training the target model with the first sample data set to obtain the predicted data of each sample data in the first sample data set can be specifically implemented in the following way: Split the first sample data set into n sample data subsets, use each sample data subset as the test set respectively, and use the remaining sample data subsets as the training set, and perform n times of cross-training on the target model.
[0045] The above training dataset is used to train the target model, and the trained target model is used to determine the predicted data of each sample data in the above test dataset. For example, when n is 5, the first sample dataset is divided into 5 parts. Each time, 4 parts are taken as the training dataset and 1 part is taken as the test set, and cross-training is performed 5 times, which can generate the predicted data of all sample data. Each sample is used as both the training dataset and the test dataset, without being affected by random factors, ensuring that this process can be replicated. Through cross-training, comparative data is provided for obtaining the error rate of the sample data and the accuracy rate of the first sample dataset in the subsequent process.
[0046] Step S220: Obtain the error rate of each sample data in the first sample dataset and the accuracy rate of the first sample dataset according to the predicted data of each sample data in the first sample dataset.
[0047] In this step S220, according to the predicted data of each sample data in the first sample dataset, the predicted result of each sample data can be determined first, then the error rate of each sample data in the first sample dataset can be determined according to the predicted result of each sample data, and then the accuracy rate of the first sample dataset in the first sample dataset can be determined according to the predicted result of each sample data.
[0048] When determining the predicted result of each sample data, it is possible to determine whether the predicted result of each sample data is correct by comparing the predicted data of each sample data in the first sample dataset with the sample label. When the comparison is consistent, the prediction of each sample data is correct; when the comparison is inconsistent, the prediction of each sample data is incorrect.
[0049] When determining the error rate of each sample data in the first sample dataset, the number of incorrect predictions of each sample data can be counted according to the predicted result of each sample data, and the ratio of the number of incorrect predictions of each sample data to the total number of predictions is used as the error rate of the sample data. For example, if the number of incorrect predictions of sample data 1 is a and the sample data has been predicted b times in total, then the error rate of sample data 1 can be expressed as a / b.
[0050] When determining the accuracy rate of the first sample dataset in the first sample dataset, the ratio of the number of sample data with correct predicted results in the first sample dataset to the total number of sample data in the first sample dataset can be used as the accuracy rate of the first sample dataset. For example, if there are a total of n sample data in the first sample dataset and m sample data have correct predictions, then the accuracy rate of the first sample dataset can be p1 = m / n.
[0051] Figure 2In the steps shown, by obtaining the error rate of each sample data in the first sample data set and the accuracy of the first sample data set, a parameter basis is provided for further locating abnormal sample data in the first sample data set, and the acquisition method is simple and easy to implement.
[0052] In an optional implementation, the above cross-training may be performed iteratively until the number of cross-training times reaches a preset number.
[0053] For example, when n is 5, the first sample data set is cut into 5 parts, 4 parts are taken as training data sets each time, and 1 part is taken as test set, and cross-training is performed 5 times. The preset number of times can be set to 500 times. Through iterative cross-training, each sample data can be predicted 100 times. This process makes the error rate of each sample data in the first sample data set and the accuracy of the first sample data set more accurate through multiple cross-training.
[0054] Step S120: remove sample data with an error rate higher than an error rate threshold from the first sample data set to obtain a second sample data set.
[0055] The second sample data set is a subset of the first sample data set. The initial error rate threshold can be set to 0.5, and sample data with an error rate higher than 0.5 in each sample data in the first sample data set are found, and these sample data with an error rate higher than 0.5 are removed from the first sample data set to construct the second sample data set. For example, the first sample data set contains four sample data {a1, a2, a3, a4, a5}, and the error rates corresponding to these four sample data are {0.18, 0.66, 0.53, 0.28, 0.49}, respectively, then the second sample data set is {a1, a4, a5}.
[0056] It should be noted that, in actual application, the size of the error rate threshold can be set according to the requirements for sample data quality, and is not limited to the size of the error rate threshold set above.
[0057] Step S130: Use the second sample data set to train the target model to obtain the accuracy of the second sample data set.
[0058] The accuracy of the second sample data set can be obtained by taking the ratio of the number of sample data with correct prediction results in the second sample data set to the total number of sample data in the second sample data set. For example, if there are a total of x sample data in the second sample data set, of which y sample data are predicted correctly, then the accuracy of the second sample data set can be p2=x / y.
[0059] Step S140: when the accuracy of the second sample data set is greater than or equal to the accuracy of the first sample data set, output the second sample data set.
[0060] When p2 ≥ p1, the second sample data set can be output, where p1 represents the accuracy rate of the first sample data set, p2 represents the accuracy rate of the second sample data set, and the second sample data set is the set obtained by improving the sample data quality of the first sample data set.
[0061] In an alternative embodiment, when the accuracy rate of the second sample data set is less than the accuracy rate of the first sample data set, the second sample data set is updated.
[0062] When p2 < p1, it indicates that the quality of the sample data in the first sample data set has not been well improved, and the second sample data set needs to be updated, where p1 represents the accuracy rate of the first sample data set, and p2 represents the accuracy rate of the second sample data set. Through continuous updating of the second sample data set, the data quality of the first sample data set is gradually improved.
[0063] In an alternative embodiment, when the accuracy rate of the second sample data set is less than the accuracy rate of the first sample data set, the second sample data set can be updated in the following manner: when the accuracy rate of the second sample data set is less than the accuracy rate of the first sample data set, the error rate threshold is updated, and it jumps to the step of removing the sample data with an error rate higher than the error rate threshold from the first sample data set to update the second sample data set.
[0064] This process can re - locate the abnormal sample data in the first sample data set by increasing the error rate threshold, remove the abnormal sample data, and obtain the updated second sample data set.
[0065] In an alternative embodiment, the updated error rate threshold can be obtained by adding a preset step size to the error rate threshold..
[0066] It should be noted that the preset step size for increasing the error rate threshold can be a fixed step size. For example, the error rate threshold can be increased by 0.05 each time. The preset step size for increasing the error rate threshold can also be a non - fixed step size. For example, when the precision requirement for data quality is relatively high, a decreasing step size is adopted. The step size is increased by 0.05 for the first update, 0.025 for the second update, 0.0125 for the third update, etc. In the actual process application, it can be adaptively set according to the requirement for data quality precision.
[0067] Figure 3 A flowchart of sample data removal is shown, which shows the execution order of the above - mentioned sample data processing method, including steps S301 to S311, and the following is a detailed description:
[0068] Step S301, start;
[0069] Step S302: Data segmentation. In this process, the sample data set is segmented into a training set and a test set.
[0070] Step S303: Use the training set to train the model, which is the target model.
[0071] Step S304: Record the prediction results of the test set.
[0072] Step S305: When the number of training times is less than N, go back to Step S302. When the number of training times is greater than or equal to N, proceed to Step S306. Here, N refers to the number of training times of the sample data.
[0073] Step S306: Record the error rate of each sample data.
[0074] Step S307: Eliminate the samples with an error rate of R, where R is the error rate of each sample data.
[0075] Step S308: Determine whether the accuracy rate of the sample data set after eliminating the samples with an error rate of R is greater than or equal to the accuracy rate of the original sample data set. When the accuracy rate of the sample data set after eliminating the samples with an error rate of R is greater than or equal to the accuracy rate of the original sample data set, go to Step S310 and proceed. When the accuracy rate of the sample data set after eliminating the samples with an error rate of R is less than the accuracy rate of the original sample data set, execute Step S309.
[0076] Step S309: Increment R by 1 and jump to Step S307 to perform the elimination operation on the original sample data set.
[0077] Step S310: Correct the sample data with a dislocation rate above R, that is, the sample data that has been eliminated.
[0078] Step S311: End.
[0079] Figure 3 The steps shown can efficiently find abnormal samples and improve the quality of sample data through iterative loops.
[0080] It should be noted that the first sample data set mentioned above may include any of the following types of data: image data, text data, and audio data. Selecting corresponding target models for different types of sample data can realize the processing and analysis of different types of sample data. For example, when the data in the first sample data set is image data, target model 1 can be used to process the first sample data; when the data in the first sample data set is text data, target model 2 can be used to process the first sample data; when the data in the first sample data set is audio data, target model 3 can be used to process the first sample data. The types of sample data that can be processed are diverse and can be applied to different application scenarios such as images, texts, and audios, thereby improving the quality of sample data in different application scenarios, and having strong universality.
[0081] Taking image data as an example, when the first sample data set is the first image data set and the target model used is an image processing model (such as a convolutional neural network model), the image data set can be processed as follows: use the first image data set to train the image processing model to obtain the error rate of each image data in the first image data set and the accuracy of the first image data set; eliminate image data with an error rate higher than an error rate threshold from the first image data set to obtain a second image data set; use the second image data set to train the image processing model to obtain the accuracy of the second image data set; when the accuracy of the second image data set is greater than or equal to the accuracy of the first image data set, output the second image data set.
[0082] The exemplary embodiment of the present disclosure also provides a sample data processing device. Figure 4 As shown, the sample data processing device 400 may include:
[0083] A first accuracy acquisition module 410 is used to train a target model using the first sample data set to obtain an error rate of each sample data in the first sample data set and an accuracy rate of the first sample data set;
[0084] The sample data elimination module 420 is used to eliminate the sample data with an error rate higher than an error rate threshold from the first sample data set to obtain a second sample data set;
[0085] A second accuracy acquisition module 430 is used to train the target model using the second sample data set to obtain the accuracy of the second sample data set;
[0086] The sample data set output module 440 is configured to output the second sample data set when the accuracy of the second sample data set is greater than or equal to the accuracy of the first sample data set.
[0087] In an alternative embodiment, the first accuracy acquisition module 410 includes: a cross-training module for cross-training a target model with a first sample data set to obtain prediction data for each sample data in the first sample data set; and a first accuracy acquisition sub-module for obtaining the error rate of each sample data in the first sample data set and the accuracy of the first sample data set according to the prediction data of each sample data in the first sample data set.
[0088] In an alternative embodiment, the first accuracy acquisition sub-module is configured to: divide the first sample data set into n sample data subsets, use each sample data subset as a test set respectively, and use the remaining sample data subsets as training sets to perform n times of cross-training on the target model.
[0089] The first accuracy acquisition sub-module is further configured to: iteratively perform cross-training until the number of times of cross-training reaches a preset number of times.
[0090] In an alternative embodiment, the sample data processing device 400 further includes an update module: when the accuracy of the second sample data set is less than the accuracy of the first sample data set, update the second sample data set.
[0091] In an alternative embodiment, the update module is configured to: when the accuracy of the second sample data set is less than the accuracy of the first sample data set, update the error rate threshold and jump to the step of removing the sample data with an error rate higher than the error rate threshold from the first sample data set to update the second sample data set.
[0092] In an alternative embodiment, the update module further includes an update sub-module for updating the error rate threshold, which is configured to: add a preset step size to the error rate threshold as the updated error rate threshold.
[0093] In an alternative embodiment, in the first accuracy acquisition module 410, the first sample data set includes any one of the following types of data: image data, text data, audio data.
[0094] The specific details of each part in the above sample data processing device 400 have been described in detail in the embodiments of the method part. For the details not disclosed, please refer to the embodiments of the method part, and thus will not be elaborated here.
[0095] Exemplary embodiments of the present disclosure also provide a computer-readable storage medium, on which a program product is stored that can implement the above-described sample data processing method of this specification. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. The program product can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on an electronic device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0096] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0097] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and the readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0098] The program code contained on the readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0099] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0100] Exemplary embodiments of the present disclosure also provide an electronic device capable of implementing the above sample data processing method. The following refers to Figure 5 to describe the electronic device 500 according to such an exemplary embodiment of the present disclosure. Figure 5 The illustrated electronic device 500 is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.
[0101] As Figure 5 shown, the electronic device 500 may be presented in the form of a general-purpose computing device. The components of the electronic device 500 may include, but are not limited to: at least one processing unit 510, at least one storage unit 520, a bus 530 connecting different system components (including the storage unit 520 and the processing unit 510), and a display unit 540.
[0102] The storage unit 520 stores program code, which can be executed by the processing unit 510, so that the processing unit 510 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 510 may execute Figures 1 to 3 any one or more of the method steps in
[0103] The storage unit 520 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 521 and / or a cache storage unit 522, and may further include a read-only storage unit (ROM) 523.
[0104] The storage unit 520 may further include a program / utility 524 having a set (at least one) of program modules 525. Such program modules 525 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0105] The bus 530 can represent one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor, or a local bus using any of the various bus architectures.
[0106] The electronic device 500 can also communicate with one or more external devices 600 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 500, and / or communicate with any device that enables the electronic device 500 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 550. Moreover, the electronic device 500 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 560. As shown in the figure, the network adapter 560 communicates with other modules of the electronic device 500 through the bus 530. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0107] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software in combination with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the exemplary embodiments of the present disclosure.
[0108] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed, for example, synchronously or asynchronously in multiple modules.
[0109] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the exemplary embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0110] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here. After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0111] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only defined by the appended claims.
Claims
1. A method for processing sample data, characterized in that, Including: Training a target model using a first sample data set to obtain the error rate of each sample data in the first sample data set and the accuracy rate of the first sample data set; the error rate of the sample data is the ratio of the number of incorrect predictions of the sample data to the total number of predictions of the sample data; The accuracy rate of the first sample data set is the ratio of the number of sample data with correct prediction results in the first sample data set to the total number of sample data in the first sample data set; Removing the sample data with an error rate higher than the error rate threshold from the first sample data set to obtain a second sample data set; Training the target model using the second sample data set to obtain the accuracy rate of the second sample data set; When the accuracy rate of the second sample data set is greater than or equal to the accuracy rate of the first sample data set, outputting the second sample data set; When the accuracy rate of the second sample data set is less than the accuracy rate of the first sample data set, updating the error rate threshold and jumping to the step of removing the sample data with an error rate higher than the error rate threshold from the first sample data set to update the second sample data set.
2. The method according to claim 1, characterized in that, The training the target model using the first sample data set to obtain the error rate of each sample data in the first sample data set and the accuracy rate of the first sample data set includes: Performing cross-training on the target model through the first sample data set to obtain the prediction data of each sample data in the first sample data set; According to the prediction data of each sample data in the first sample data set, obtaining the error rate of each sample data in the first sample data set and the accuracy rate of the first sample data set.
3. The method according to claim 2, characterized in that, The performing cross-training on the target model through the first sample data set to obtain the prediction data of each sample data in the first sample data set includes: Splitting the first sample data set into n sample data subsets, using each sample data subset as a test set, and the remaining sample data subsets as a training set, and performing n times of cross-training on the target model.
4. The method according to claim 2, characterized in that, The performing cross-training on the target model through the first sample data set to obtain the prediction data of each sample data in the first sample data set further includes: Iteratively performing the cross-training until the number of times of cross-training reaches a preset number of times.
5. The method according to claim 1, characterized in that, The updating the error rate threshold includes: Adding a preset step size to the error rate threshold as the updated error rate threshold.
6. The method according to claim 1, characterized in that, The first sample data set includes any one of the following types of data: Image data, text data, audio data.
7. A sample data processing device, characterized in that, Including: A first accuracy rate obtaining module, configured to train a target model using a first sample data set to obtain the error rate of each sample data in the first sample data set and the accuracy rate of the first sample data set; the error rate of the sample data is the ratio of the number of incorrect predictions of the sample data to the total number of predictions of the sample data; The accuracy rate of the first sample data set is the ratio of the number of sample data with correct prediction results in the first sample data set to the total number of sample data in the first sample data set; A sample data elimination module, configured to eliminate sample data with an error rate higher than an error rate threshold from the first sample data set to obtain a second sample data set; A second accuracy rate acquisition module, configured to train a target model using the second sample data set to obtain the accuracy rate of the second sample data set; A sample data set output module, configured to output the second sample data set when the accuracy rate of the second sample data set is greater than or equal to the accuracy rate of the first sample data set; An update module, configured to update the error rate threshold when the accuracy rate of the second sample data set is less than the accuracy rate of the first sample data set, and jump to the step of eliminating sample data with an error rate higher than the error rate threshold from the first sample data set to update the second sample data set.
8. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
9. An electronic device, characterized in that, Comprising: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1 to 6 by executing the executable instructions.
Citation Information
Patent Citations
Method and device for processing feature traversal in sample set, equipment and readable medium
CN110472743A