A cross-validation dataset pruning and evaluation method based on data balance
Through the cross-validation method based on data balance, the data classification model data set is cropped and evaluated, which solves the problems of data set redundancy and low sample quality, and improves the efficiency and generalization ability of the image classification model.
Patent Information
- Application Number
- CN202410367109.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-03-28
AI Technical Summary
The existing image classification model datasets have problems such as redundant datasets and low sample quality, which leads to the limitation of the efficiency and generalization capabilities of deep learning image classification models.
The cross-validation data set cropping and evaluation method based on data balance is adopted. The preset first image classification model is trained through k-fold cross-validation, and the prediction accuracy and prediction probability value of each sample are recorded. After sorting, low-quality samples and redundant samples are deleted to construct a high-quality core data set, and the preset second image classification model is trained for performance evaluation using this data set.
It effectively reduces the computing resource consumption of model training, builds a high-quality data set that can better represent the distribution of the original data set, and improves the performance and generalization capabilities of the image classification model.
Smart Images

Figure CN118279696B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image classification and recognition, and in particular to a cross-validation data set clipping and evaluation method based on data balance. Background Art
[0002] Image data often contains a lot of effective information. Image recognition technology can automatically extract and analyze information from a large amount of image data, thereby promoting intelligent applications and development in different fields. For example, in industrial production, image recognition can be used to detect product defects; in medical imaging, image recognition can be used to assist doctors in disease diagnosis. In the field of image recognition, the quality of the data set plays a key role in improving image recognition performance. However, sample data sets often contain similar or repeated samples. These redundant samples not only increase the size of the data set, but also occupy valuable resources in the training process. In addition, there may be low-quality samples with incorrect or noisy annotations in the data set. These data will reduce the performance of the image classification model and have a negative impact on the generalization ability of the image classification model.
[0003] At present, the dataset pruning methods mainly include pruning difficult samples, pruning simple samples, or pruning simple and difficult samples. The method of pruning difficult samples does not take into account those samples with rich information. Pruning the remaining core dataset will cause greater redundancy in samples. After pruning simple samples, the remaining core set will further increase the disturbance of difficult samples to the image classification model training; pruning simple and difficult samples, the remaining core set retains those information-rich samples, which is more conducive to the image classification model training. However, at present, the dataset pruning methods participate in the training process when evaluating the dynamics of samples. The image classification model will memorize the samples, which will cause deviations when evaluating the characteristics of the samples.
[0004] Therefore, it is necessary to propose a cross-validation dataset pruning method based on data balance, which can solve the problems of dataset redundancy, computing resource limitations, and low-quality samples in the dataset, and improve the efficiency and generalization ability of deep learning image classification models. Summary of the invention
[0005] In view of this, the present invention provides a cross-validation dataset cropping and evaluation method based on data balance, which is used to solve the problems of dataset redundancy and low sample quality in existing image classification model datasets, thereby improving the efficiency and generalization ability of deep learning image classification models.
[0006] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a cross-validation data set pruning and evaluation method based on data balance, comprising:
[0008] Obtain a sample data set, and divide the sample data set into k subsets;
[0009] Input the divided sample data set into the image classification model, traverse the k subsets in turn, use the current subset as the validation set and the remaining subsets as the training set to train the preset first image classification model, and record the prediction accuracy and prediction probability value of the preset first image classification model for each sample in the validation set when each subset is used as the validation set;
[0010] Sort the samples based on the prediction accuracy and prediction probability value, delete the samples according to the sorting results, and obtain the core data set;
[0011] Using the core data set to train the preset second image classification model to obtain a fully trained preset second image classification model;
[0012] A test data set is obtained, and the test data set is input into a fully trained preset second image classification model for testing, and a performance evaluation is performed on the fully trained preset second image classification model according to the test results.
[0013] Furthermore, the sorting of samples based on the prediction accuracy and the prediction probability value, and deleting the samples according to the sorting result, includes:
[0014] Sort the samples with incorrect predictions from small to large according to the predicted probability value, and sort the samples with correct predictions from large to small according to the predicted probability value;
[0015] Based on the preset core data set ratio, samples with incorrect predictions are deleted in ascending order of predicted probability values;
[0016] If the proportion of the remaining sample data set after deletion is greater than the preset core data set proportion, the correctly predicted samples are deleted in descending order of predicted probability.
[0017] Further, the test data set is preprocessed, and the preprocessed test data set is input into a fully trained preset second image classification model for classification testing;
[0018] The preprocessing includes performing Gaussian blurring and adding Gaussian noise to the images in the test set.
[0019] Further, Gaussian blurring of the image includes:
[0020] Generate a two-dimensional Gaussian weight matrix according to the size and standard deviation of the preset Gaussian kernel;
[0021] For each pixel point of each sample in the test set, taking each pixel point as the center, weighted averaging is performed on the pixels in the neighborhood of the current pixel point based on the two-dimensional Gaussian weight matrix.
[0022] Furthermore, adding Gaussian noise to the image includes:
[0023] Generate random numbers that conform to Gaussian distribution to obtain Gaussian noise;
[0024] Weighting the Gaussian noise to each pixel value of the image;
[0025] The Gaussian noise intensity is adjusted by adjusting the standard deviation of the random number.
[0026] Furthermore, the preset first image classification model and the preset second image classification model are image classification models with different structures; the preset first image classification model is constructed based on a Transformer neural network; and the preset second image classification model is constructed based on a convolutional neural network.
[0027] Further, the performance of the preset second image classification model is evaluated according to the classification test result, including: the performance of the model is evaluated according to the accuracy of the preset second image classification model, and the accuracy is calculated by:
[0028] Accuracy = number of correctly predicted samples / total number of samples.
[0029] In a second aspect, the present invention further provides a cross-validation data set clipping and evaluation system based on data balance, comprising:
[0030] A data set acquisition module is used to acquire a sample data set and divide the sample data set into k subsets;
[0031] The data set analysis module is used to input the divided sample data set into the image classification model, traverse the k subsets in turn, train the preset first image classification model with the current subset as the verification set and the remaining subsets as the training set, and record the prediction accuracy and prediction probability value of the preset first image classification model for each sample in the verification set when each subset is used as the verification set;
[0032] The trimming module is used to sort the samples based on the prediction accuracy and prediction probability value, and delete the samples according to the sorting results to obtain the core data set.
[0033] A model training module, used to train the preset second image classification model using the core data set to obtain a fully trained preset second image classification model;
[0034] The performance evaluation module is used to obtain a test data set, input the test data set into a fully trained preset second image classification model for testing, and perform performance evaluation on the fully trained preset second image classification model based on the test results.
[0035] In a third aspect, the present invention further provides an electronic device, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, it can implement any of the above-mentioned cross-validation data set clipping and evaluation methods based on data balance.
[0036] In a fourth aspect, the present invention also provides a computer-readable storage medium for storing computer-readable programs or instructions, which, when executed by a processor, can implement the steps in the cross-validation data set pruning and evaluation method based on data balance in any of the above-mentioned implementation methods.
[0037] The present invention provides a cross-validation data set clipping and evaluation method based on data balance. First, the preset first classification model is trained by the k-fold cross-validation method, and each sample in the training set is evaluated to ensure that the sample evaluation will not produce errors due to the memory of the image classification model; secondly, the prediction accuracy and prediction probability value of each sample in the validation set are calculated and sorted by the trained preset first classification model, and the data set is clipped based on the sorting result, and low-quality samples and redundant samples are deleted to construct a high-quality core data set. At the same time, the high-quality core data set can better represent the distribution of the original data set, so that the trained image classification model has better performance and generalization ability. Finally, the preset second image classification model is trained using the high-quality core data set to obtain a fully trained classification model, and the performance of the trained model is evaluated using the test data set, and the core data set is objectively evaluated according to the evaluation results. The present invention not only deletes simple samples and difficult samples and retains information-rich uncertain samples, but also ensures that when measuring samples, the samples do not participate in the training of the image classification model, thereby ensuring that sample evaluation will not produce errors due to the memory of the image classification model. The present invention is applicable to any type of image classification model and any scale of image classification data set. Whether it is a small data set or a large data set, this method can be used to construct a high-quality core data set, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A schematic diagram of a process flow of a cross-validation data set pruning and evaluation method based on data balance provided by the present invention;
[0039] Figure 2 A flowchart of an embodiment of training a model by a k-fold cross-validation method provided by the present invention;
[0040] Figure 3 A schematic diagram of an embodiment of deleting samples provided by the present invention;
[0041] Figure 4A flowchart of an embodiment of the present invention for evaluating the performance of a model using a test data set;
[0042] Figure 5 A schematic diagram of the structure of an embodiment of a cross-validation data set clipping and evaluation system based on data balance provided by the present invention;
[0043] Figure 6 A schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0045] The present invention provides a cross-validation data set clipping and evaluation method, system, electronic device and computer-readable storage medium based on data balance, which are described below respectively.
[0046] Combination Figure 1 As shown, a specific embodiment of the present invention discloses a cross-validation data set pruning and evaluation method based on data balance, comprising:
[0047] Step S101, obtaining a sample data set, and dividing the sample data set into k subsets;
[0048] Step S102, input the divided sample data set into the image classification model, traverse k subsets in turn, use the current subset as the verification set and the remaining subsets as the training set to train the preset first image classification model, and record the prediction accuracy and prediction probability value of the preset first image classification model for each sample in the verification set when each subset is used as the verification set;
[0049] Step S103, sorting the samples based on the prediction accuracy and the prediction probability value, and deleting the samples according to the sorting result to obtain a core data set;
[0050] Step S104: training the preset second image classification model using the core data set to obtain a fully trained preset second image classification model;
[0051] Step S105: obtain a test data set, input the test data set into a fully trained preset second image classification model for testing, and perform a performance evaluation on the fully trained preset second image classification model based on the test results.
[0052] Compared with the prior art, the method of this embodiment first trains the preset first classification model through the k-fold cross-validation method, evaluates each sample in the training set, and ensures that the sample evaluation will not produce errors due to the memory of the image classification model; secondly, the prediction accuracy and prediction probability value of each sample in the validation set are calculated and sorted through the trained preset first classification model, and the data set is trimmed based on the sorting result, and low-quality samples and redundant samples are deleted to construct a high-quality core data set, which reduces the computing resource consumption of model training. At the same time, the high-quality core data set can better represent the distribution of the original data set, so that the trained image classification model has better performance and generalization ability. Finally, the preset second image classification model is trained using the high-quality core data set to obtain a well-trained classification model, and the performance of the trained model is evaluated using the test data set, and the core data set is objectively evaluated according to the evaluation results. The method of this embodiment not only realizes the deletion of simple samples and difficult samples, but also retains information-rich uncertainty samples; and when measuring samples, the samples do not participate in the image classification model training, ensuring that the sample evaluation will not produce errors due to the memory of the image classification model, and can be applied to any type of image classification model and any scale of image classification data set. Whether it is a small data set or a large data set, this method can be used to construct a high-quality core data set, which has broad application prospects.
[0053] The following is a specific embodiment and Figure 2 The specific process of dividing the sample data set in step S102 to train the preset first image classification model is described.
[0054] Assume that the original data set uses the CUB-200-2011 training set. The CUB-200-2011 training set is a data set for image classification and target recognition, which contains image samples of hundreds of bird species. Each species classification contains multiple corresponding images. Take the pictures under each category in CUB-200-2011 as a sample data set, divide it into k parts in equal proportion, select 1 part from each category to form a 1-fold data set, and divide the training set samples into k folds in equal proportion.
[0055] Figure 2 The data set division and model training process when k=5 is shown. The original data set (the leftmost box in the figure, the training set) is evenly divided into 5 folds to obtain D 1 -D 5 Five subsets. 1 When D is the validation set, 2 -D 5 is the training set, trains the image classification model A, and obtains the trained image classification model A_1;
[0056] D 2 When D is the validation set, 1 and D 3 -D 5 is the training set, trains the image classification model A, and obtains the trained image classification model A_2;
[0057] D 3 When D is the validation set, 1 -D 2 and D 4 -D 5 is the training set, trains the image classification model A, and obtains the trained image classification model A_3;
[0058] D 4 When D is the validation set, 1 -D 3 and D 5 is the training set, trains the image classification model A, and obtains the trained image classification model A_4;
[0059] D 5 When D is the validation set, 1 -D 4 is the training set, trains the image classification model A, and obtains the trained image classification model A_5.
[0060] Models A_1 to A_5 are evaluated based on prediction accuracy and prediction probability values. Cross-validation can better evaluate the performance of the model on unseen data. Compared with directly using the entire data set for model training, this method can reduce the dependence on specific training sets and validation sets, which helps to evaluate model performance more objectively and avoid the problem of overfitting to specific data sets that may occur when training the entire data set. In addition, partitioning and cross-validating the training set can also help us better understand the performance of the model on different data subsets, so as to better adjust model parameters and improve model performance.
[0061] As a preferred embodiment, the step of sorting samples based on prediction accuracy and prediction probability values, and deleting samples according to the sorting results, includes:
[0062] Sort the samples with incorrect predictions from small to large according to the predicted probability value, and sort the samples with correct predictions from large to small according to the predicted probability value;
[0063] Based on the preset core data set ratio, samples with incorrect predictions are deleted in ascending order of predicted probability values;
[0064] If the proportion of the remaining sample data set after deletion is greater than the preset core data set proportion, the correctly predicted samples are deleted in descending order of predicted probability.
[0065] As a specific embodiment, the preset core data set ratio can be set as needed. Suppose it is set to 90%, then the training set samples are deleted in order, and 10% of the training set samples need to be deleted. The samples with incorrect predictions are directly deleted; if the ratio still does not reach 90% after deleting the erroneous samples, the samples with correct predictions are deleted.
[0066] For sample data sets with different categories, when deleting correctly predicted samples, first determine whether the number of categories in which the sample belongs is greater than the core set divided by the number of categories. If it is satisfied, delete the sample; if not, skip the sample and delete samples in other categories.
[0067] The above process is different from the current use of two models to denoise the data set. The purpose of data set denoising is to improve the quality of the data set and enhance the model performance. The purpose of the method in this embodiment is to find a high-quality core set, retain high-quality samples in the data set, and achieve the same performance as the original data set training model. It solves the problems of redundancy, computing resource limitations, and low-quality samples in the data set, making the training data cleaner and more balanced, and more in line with actual application scenarios, thereby improving the efficiency and generalization ability of the deep learning image classification model.
[0068] Combine the following Figure 3 The sample deletion method is described in more detail in the specific embodiment of FIG. Figure 3 As shown in , the model needs to classify the sample into one of the three categories ABC. A1 and A2 represent two different samples of category A; C3 and C4 represent two different samples of category C. Assume that the probability predicted by the model and the final classification result are as follows Figure 3 As shown, Figure 3 The value in represents the prediction probability. When deleting samples, for samples with incorrect predictions, the lower the prediction probability, the more serious the error, so when deleting incorrect samples, they are sorted from small to large; for samples with correct predictions, the higher the prediction probability, the simpler the sample, so when deleting correct samples, they are sorted from large to small.
[0069] In order to verify the data quality of the high-quality core data set, a preset second image classification model is constructed. The preset second image classification model is trained through the core data set to obtain a fully trained preset second image classification model, so as to intuitively reflect the training effect of the core data set according to the performance of the model.
[0070] As a preferred embodiment, the test data set is input into a fully trained preset second image classification model for testing, and further includes:
[0071] Preprocessing the test data set, and inputting the preprocessed test data set into a fully trained preset second image classification model for classification testing;
[0072] The preprocessing includes performing Gaussian blurring and adding Gaussian noise to the images in the test set.
[0073] As a preferred embodiment, performing Gaussian blurring on an image includes:
[0074] Generate a two-dimensional Gaussian weight matrix according to the size and standard deviation of the preset Gaussian kernel;
[0075] For each pixel point of each sample in the test set, taking each pixel point as the center, weighted averaging is performed on the pixels in the neighborhood of the current pixel point based on the two-dimensional Gaussian weight matrix.
[0076] It should be noted that the core idea of Gaussian blur is to perform weighted averaging on the value of each pixel in the image, so that the value of each pixel is affected by the values of its surrounding pixels, thereby achieving the purpose of blurring the image.
[0077] When processing, we first need to determine the size and standard deviation of the Gaussian kernel. Based on the Gaussian kernel size and standard deviation, we generate a two-dimensional Gaussian weight matrix. Each element value in this matrix represents the distance from the center pixel. The farther the distance, the smaller the weight. For each sample in the test set, we perform a weighted average on each pixel: we apply the Gaussian weight matrix to each pixel of the image, and calculate the weighted average to generate the blurred image. For each pixel, we consider the values of its neighboring pixels and perform a weighted average according to the Gaussian weight matrix to complete the Gaussian blur processing of the image.
[0078] As a preferred embodiment, adding Gaussian noise to the image includes:
[0079] Generate random numbers that conform to Gaussian distribution to obtain Gaussian noise;
[0080] Weighting the Gaussian noise to each pixel value of the image;
[0081] The Gaussian noise intensity is adjusted by adjusting the standard deviation of the random number.
[0082] The intensity of the noise can be controlled by controlling the standard deviation of the generated Gaussian random numbers. The standard deviation determines the width of the distribution. The larger the standard deviation, the stronger the noise.
[0083] By introducing data enhancement techniques, such as Gaussian blur and Gaussian noise, the robustness of the model trained on the core dataset can be more comprehensively evaluated. Figure 4 As shown, Figure 4The flowchart of the performance evaluation of the trained classification model using the original test data set, the test data set after Gaussian blur processing and the test data set after Gaussian noise processing is shown. Model A in the figure is the preset second image classification model, and the Gaussian noise blur test set is the data set expansion transformation operation performed on the test set.
[0084] As a preferred embodiment, the preset first image classification model and the preset second image classification model are image classification models with different structures; the preset first image classification model and the preset second image classification model are image classification models with different structures; the preset first image classification model is constructed based on the Transformer neural network; the preset second image classification model is constructed based on a convolutional neural network.
[0085] It should be noted that the preset first image classification model and the preset second image classification model can be based on the same basic network structure or on different basic network structures. The dataset cropping method established in this embodiment can be applicable to any image classification model, any type, and any scale of image classification dataset.
[0086] As a preferred embodiment, the performance of the preset second image classification model is evaluated according to the classification test result, including: the performance of the model is evaluated according to the accuracy of the preset second image classification model, and the accuracy is calculated by:
[0087] Accuracy = number of correctly predicted samples / total number of samples.
[0088] In summary, the cross-validation dataset pruning method based on data balance proposed by us not only achieves the deletion of simple samples and difficult samples, retains information-rich uncertainty samples, but also achieves that when measuring samples, the samples do not participate in the training of the image classification model, ensuring that the sample evaluation will not produce errors due to the memory of the image classification model. In addition, the method of this application can be applied to coarse-grained and fine-grained image classification datasets, and is suitable for any image classification model and any scale dataset.
[0089] This embodiment also provides a cross-validation data set clipping and evaluation system based on data balance, such as Figure 5 As shown, the cross-validation dataset pruning and evaluation system 500 based on data balance includes:
[0090] The data set acquisition module 501 is used to acquire a sample data set and divide the sample data set into k subsets;
[0091] The data set analysis module 502 is used to input the divided sample data set into the image classification model, traverse the k subsets in turn, train the preset first image classification model with the current subset as the verification set and the remaining subsets as the training set, and record the prediction accuracy and prediction probability value of the preset first image classification model for each sample in the verification set when each subset is used as the verification set;
[0092] The trimming module 503 is used to sort the samples based on the prediction accuracy and the prediction probability value, and delete the samples according to the sorting result to obtain the core data set.
[0093] A model training module 504 is used to train the preset second image classification model using the core data set to obtain a fully trained preset second image classification model;
[0094] The performance evaluation module 505 is used to obtain a test data set, input the test data set into a fully trained preset second image classification model for testing, and perform performance evaluation on the fully trained preset second image classification model according to the test results.
[0095] like Figure 6 As shown, the cross-validation dataset clipping and evaluation method based on data balance, the present invention also provides an electronic device 600, which can be a computing device such as a mobile terminal, a desktop computer, a notebook, a palm computer and a server. The electronic device includes a processor 601, a memory 602 and a display 603.
[0096] In some embodiments, the memory 602 may be an internal storage unit of a computer device, such as a hard disk or memory of a computer device. In other embodiments, the memory 602 may also be an external storage device of a computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Further, the memory 602 may also include both an internal storage unit of a computer device and an external storage device. The memory 602 is used to store application software and various types of data installed on the computer device, such as program codes for installing the computer device. The memory 602 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a cross-validation data set clipping and evaluation method program 604 based on data balance is stored on the memory 602, and the cross-validation data set clipping and evaluation method program 604 based on data balance can be executed by the processor 601, thereby realizing a cross-validation data set clipping and evaluation method based on data balance of each embodiment of the present invention.
[0097] In some embodiments, the processor 601 can be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 602, such as executing a cross-validation data set clipping and evaluation method program based on data balance.
[0098] In some embodiments, the display 603 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 603 is used to display information on the computer device and to display a visual user interface. The components 601-603 of the computer device communicate with each other through a system bus.
[0099] This embodiment further provides a computer-readable storage medium on which a cross-validation dataset clipping and evaluation program based on data balance is stored. When the cross-validation dataset clipping and evaluation program based on data balance is executed by a processor, the steps in the above embodiment can be implemented.
[0100] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A cross-validation dataset pruning and evaluation method based on data balance, characterized in that: include: Obtain a sample data set, and divide the sample data set into k subsets; Input the divided sample data set into the image classification model, traverse the k subsets in turn, use the current subset as the validation set and the remaining subsets as the training set to train the preset first image classification model, and record the prediction accuracy and prediction probability value of the preset first image classification model for each sample in the validation set when each subset is used as the validation set; The samples are sorted based on the prediction accuracy and prediction probability values, and the samples are deleted according to the sorting results to obtain the core data set; including: Sort the samples with incorrect predictions from small to large according to the predicted probability values, and sort the samples with correct predictions from large to small according to the predicted probability values; based on the preset core data set ratio, delete the samples with incorrect predictions from small to large according to the predicted probability values; if the proportion of the remaining sample data set after deletion is greater than the preset core data set ratio, delete the samples with correct predictions from large to small according to the predicted probability values; for sample data sets with different categories, when deleting the samples with correct predictions, first determine whether the number of categories to which the sample belongs is greater than the core set divided by the number of categories. If so, delete the sample; if not, skip the sample and delete the samples in other categories; Using the core data set to train the preset second image classification model to obtain a fully trained preset second image classification model; A test data set is obtained, and the test data set is input into a fully trained preset second image classification model for testing, and a performance evaluation is performed on the fully trained preset second image classification model according to the test results.
2. The cross-validation dataset pruning and evaluation method based on data balance according to claim 1, characterized in that: The test data set is fed into the fully trained preset second image classification model for testing, which also includes: Preprocessing the test data set, and inputting the preprocessed test data set into a fully trained preset second image classification model for classification testing; The preprocessing includes performing Gaussian blurring and adding Gaussian noise to the images in the test set.
3. The cross-validation dataset pruning and evaluation method based on data balance according to claim 2, characterized in that: Gaussian blurring an image involves: Generate a two-dimensional Gaussian weight matrix according to the size and standard deviation of the preset Gaussian kernel; For each pixel point of each sample in the test set, taking each pixel point as the center, weighted averaging is performed on the pixels in the neighborhood of the current pixel point based on the two-dimensional Gaussian weight matrix.
4. The cross-validation dataset pruning and evaluation method based on data balance according to claim 2, characterized in that: Adding Gaussian noise to an image includes: Generate random numbers that conform to Gaussian distribution to obtain Gaussian noise; Weighting the Gaussian noise to each pixel value of the image; The Gaussian noise intensity is adjusted by adjusting the standard deviation of the random number.
5. The cross-validation dataset pruning and evaluation method based on data balance according to claim 1, characterized in that: The preset first image classification model and the preset second image classification model are image classification models with different structures; the preset first image classification model is constructed based on a Transformer neural network; The preset second image classification model is constructed based on a convolutional neural network.
6. The cross-validation dataset pruning and evaluation method based on data balance according to claim 1, characterized in that: The performance of the preset second image classification model is evaluated according to the classification test result, including: the performance of the model is evaluated according to the accuracy of the preset second image classification model, and the accuracy is calculated by: Accuracy = number of correctly predicted samples / total number of samples.
7. A cross-validation dataset pruning and evaluation system based on data balance, characterized in that: include: A data set acquisition module is used to acquire a sample data set and divide the sample data set into k subsets; The data set analysis module is used to input the divided sample data set into the image classification model, traverse the k subsets in turn, train the preset first image classification model with the current subset as the verification set and the remaining subsets as the training set, and record the prediction accuracy and prediction probability value of the preset first image classification model for each sample in the verification set when each subset is used as the verification set; The trimming module is used to sort the samples based on the prediction accuracy and the prediction probability value, and delete the samples according to the sorting result to obtain the core data set, including: sorting the samples with incorrect prediction from small to large according to the prediction probability value, and sorting the samples with correct prediction from large to small according to the prediction probability; deleting the samples with incorrect prediction from small to large according to the prediction probability value based on the preset core data set ratio; if the proportion of the remaining sample data set after deletion is greater than the preset core data set ratio, deleting the samples with correct prediction from large to small according to the prediction probability; for sample data sets with different categories, when deleting the samples with correct prediction, first determine whether the number of categories to which the sample belongs meets the requirement of being greater than the core set divided by the number of categories. If so, delete the sample; if not, skip the sample and delete the samples in other categories; A model training module, used to train the preset second image classification model using the core data set to obtain a fully trained preset second image classification model; The performance evaluation module is used to obtain a test data set, input the test data set into a fully trained preset second image classification model for testing, and perform performance evaluation on the fully trained preset second image classification model based on the test results.
8. An electronic device, characterized in that: It comprises a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the cross-validation data set clipping and evaluation method based on data balance as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: Used to store computer-readable programs or instructions, which, when executed by a processor, can implement the steps in the cross-validation data set pruning and evaluation method based on data balance as described in any one of claims 1 to 6 above.
Citation Information
Patent Citations
Classification model training method, quality inspection prediction method and corresponding devices
CN114462465A
Coal mine underground image quality evaluation method
CN117197092A