A method for improving the accuracy of image classification using a partition decision mechanism

By adopting partition decision mechanism and integrated learning ideas in image recognition, partitioning and model training are carried out for images, the problems of unstable improvement effects, poor portability and poor interpretability in the existing technology are solved, and the accuracy of image classification and optimization of computing efficiency are achieved.

CN114743055BActive Publication Date: 2025-06-27BEIJING INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210406278.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-18
Publication Date
2025-06-27
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

The existing convolutional neural network improvement methods have unstable improvement effects in different data sets and application scenarios, and there are problems of poor portability and poor interpretability.

Method used

The partition decision mechanism is adopted to improve the accuracy of image classification by partitioning and cropping the image, and the integrated learning idea is used to train the convolutional neural network model, and comprehensive decisions are made by combining the recognition results of multiple subgraphs.

Benefits of technology

It realizes the stable improvement of image classification accuracy in different data sets and application scenarios without adding additional computing overhead, and has good portability and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114743055B_ABST
    Figure CN114743055B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for improving the accuracy of image classification using a partition decision mechanism. Based on the idea of ensemble learning, the partition decision mechanism is used to enable the model to identify different regions of the image, summarize multiple recognition results, and then infer the category to which the entire image belongs. A model improvement method that can stably and reliably improve the accuracy of image classification is provided, and the model training process is simple, improving the accuracy of the convolutional neural network model in image classification without bringing too much additional computational overhead to the training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence computer vision image recognition, and more specifically, to a method for improving the accuracy of image classification using a partition decision mechanism. Background Art

[0002] Image recognition refers to using devices such as computers to process and analyze images, extract image features, and complete tasks such as classification, object detection, and matching. Image recognition is an important research direction in the field of computer vision. With the development of artificial intelligence technology in recent years, more and more methods and application results have emerged. Image classification is an important sub-task in the field of image recognition. Many computer vision tasks are based on image classification. For example, a core problem in object detection tasks is how to correctly identify the category of the sub-image in the detection box. Currently, the most commonly used method to solve the image classification problem is the deep learning method. By constructing a deep convolutional neural network and using the optimization method of gradient descent, the model automatically learns the method of extracting image features during training to complete image classification. However, many current mainstream convolutional neural network improvement methods have the following problems:

[0003] 1) The improvement effect is unstable. In different datasets and application scenarios, it is difficult to guarantee the improvement effect of the improvement method on the model accuracy, and there may even be a situation where the accuracy is lower than that of the original model;

[0004] 2) Poor portability. Most improvement methods are mutually exclusive and cannot be used simultaneously, resulting in the effectiveness of a study often being based on the negation of other studies;

[0005] 3) Poor interpretability. Many improvement methods essentially rely on the accumulation of computing power and data scale, and the improvement effect is difficult to be reasonably explained, and it may lead to an increase in the operation cost of the system.

[0006] Ensemble learning is a model improvement idea that can effectively overcome the above problems. It trains multiple models or trains a single model multiple times based on different data, and uses the principle of probability theory to reduce the error probability of the model in simple single recognition, achieving the effect of improving the model accuracy. Based on some simple calculations related to the principle of probability theory, the improvement effect of ensemble learning on the model accuracy is easy to prove, and it can be stably and well applied to most scenarios.

[0007] Therefore, how to improve the accuracy of the image recognition model without increasing additional computational overhead is an urgent problem for those skilled in the art. Summary of the Invention

[0008] In view of this, the present invention provides a method for improving the accuracy of image classification using a partition decision mechanism. Based on the idea of ensemble learning, the partition decision mechanism is used to enable the model to identify different regions of the image, summarize multiple recognition results, and then infer the category to which the entire image belongs. A model improvement method that can stably and reliably improve the accuracy of image classification is provided, and the model training process is simple, which improves the accuracy of the convolutional neural network model in image classification without bringing too much additional computational overhead to the training.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] A method for improving the accuracy of image classification using a partition decision mechanism, comprising the following steps:

[0011] Step 1: Collect a large number of image data for the target application scenario, manually label category labels for these images, or directly use relevant public data sets, organize them into an original image data set, and divide it into a training set and a test set;

[0012] Step 2: Perform partition cropping on the images in the original image data set obtained in Step 1 to generate a cropped sub-image data set;

[0013] Among them, for the partition cropping algorithm, for different images in the data set, either all can be cropped according to the same cropping scheme, or different cropping schemes can be respectively adopted according to the differences in the shape and size of each image to crop them into a unified size; this sub-image data set will replace the original image data set and participate in the training of the convolutional neural network model;

[0014] Step 3: Construct a data set reader, and perform data preprocessing on a number of sub-images batch-selected by the data set reader from the sub-image data set to obtain training images;

[0015] The data set reader is used to control the process of reading data from the data set during each batch of training, including how many images to select, the selection algorithm, how to preprocess the images, how to obtain the true labels of the images, etc. Among them, for the sub-image data set obtained in Step 2, the selection algorithm can be implemented in ways such as randomly selecting sub-images, sequentially selecting sub-images from the same image, or a combination of both; and the data preprocessing process includes various data augmentation and normalization methods;

[0016] The specific process of the data preprocessing is as follows:

[0017] Step 31: Scale the selected number of sub-images according to a preset size;

[0018] Step 32: Perform pixel padding on the scaled sub-images, and randomly crop the padded images to the preset size to obtain re-cropped images;

[0019] Step 33: Randomly flip the re-cropped images horizontally along the vertical central axis with a probability of 0.5, that is, randomly select half of the re-cropped images and flip them horizontally along the vertical central axis for data augmentation; form an augmented image set with the flipped images and the re-cropped images;

[0020] Step 34: Standardize all the images in the augmented image set according to the preset three-channel mean and preset three-channel variance to generate the training images;

[0021] Step 4: Construct a convolutional neural network model as the basic model for the classification task; determine the loss function and optimizer according to the structure of the classification basic model, and set the training parameters for the training of the classification basic model according to the optimizer;

[0022] To improve the classification performance of the model or reduce the training time overhead, the parameters of the model can be initialized to the parameter values that have been pre-trained on a large-scale dataset, but random initialization and other methods can also be used;

[0023] Among them, for classification problems, the cross-entropy loss function is generally used, and common and frequently used optimizers include Adam, SGD, etc.;

[0024] Setting the training parameters for model training includes the initial learning rate, decay coefficient, etc.; among them, which training parameters need to be specifically set depends on the requirements of the obtained optimizer;

[0025] Step 5: In each batch of training, use the dataset reader obtained in Step 3 to select several sub-graphs from the training set of the sub-graph dataset obtained in Step 2, and input the processed training images into the model obtained in Step 4 for training;

[0026] Among them, the training process of the model includes three steps:

[0027] Step 51: Pass all the inputs through different network layers of the model in turn to complete operations such as convolution, and finally obtain the output result of the current model for the input;

[0028] Step 52: Calculate the loss value of the output according to the output result and the true label of the image through the obtained loss function;

[0029] Step 53: According to the loss value, perform backpropagation, perform gradient descent on the model, and complete the parameter update of the model;

[0030] Step 54: Use the optimizer obtained in Step 5 to adjust the training parameters involved;

[0031] Step 55: The dataset reader checks whether the selection of images is completed. If not, it selects the next batch of training images for training and returns to Step 51. If so, the model training is completed, and the training result of the current classification basic model is tested using the original image dataset in Step 1 to obtain the image classification model. That is, moving the testing step to after the entire training process and only performing one test can save training time.

[0032] Step 6: The image to be classified is partitioned and cropped and then input into the image classification model in sequence to obtain the output results. A partition decision mechanism is used to comprehensively make decisions on all output results to obtain the final classification result.

[0033] When testing using the test set of the original image dataset or when using other natural images for testing after training, the trained image classification model is used for classification testing. The testing process is different from the data processing process during model training, which is as follows:

[0034] Step 61: For the test image, use the partition and crop algorithm obtained in Step 2 and adopt the same cropping strategy as the training set images of the original image dataset to crop it into sub-images of several different partitions.

[0035] Step 62: For each sub-image obtained in Step 61, input it into the trained model in sequence to obtain the output result of the model for each sub-image.

[0036] Step 63: According to the output results of all sub-images, use a certain partition decision mechanism to calculate the decision result after combining the output results of these sub-images as the final output result of the entire original image.

[0037] Among them, the specific implementation of the partition decision mechanism in Step 63 includes many. Most of the ensemble learning methods and group decision-making methods can be used.

[0038] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method for improving the accuracy of image classification using a partition decision mechanism, modularly designing the model training and classification recognition processes. During the training process of the image classification model, the training images are partitioned and cropped, and the convolutional neural network model is trained in batches. During the process of using the trained image classification model for image recognition and classification, a partition decision mechanism is adopted to comprehensively make decisions on the model classification results of the images to be classified, and the obtained decision results are used as the final image classification results. Based on the principle of probability theory, through simple calculations, taking the 10-classification problem as an example, in an extremely harsh situation, as long as the accuracy of the basic model is greater than 0.27, it can ensure the improvement effect of the method of the present invention on the accuracy, and it has strong reliability. Since the method of the present invention is based on modular design, it has strong portability and can be easily migrated to any basic model structure and application scenario, and has a certain degree of robustness. Compared with general ensemble learning methods, the method proposed by the present invention also has distinct interpretability, the principle of improving the model accuracy is easy to understand, and it will not bring excessive additional computational overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0040] Figure 1 The drawings are schematic flowcharts of the method for improving the accuracy of image classification using a partition decision mechanism provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0042] The embodiments of the present invention disclose a method for improving the accuracy of image classification using a partition decision mechanism. The method flow is as Figure 1 shown, and specifically includes the following steps:

[0043] S1: Collect a large amount of image data for the target application scenario, manually label category tags for these images, or directly use relevant public data sets, organize them into an original image data set, and divide it into a training set and a test set;

[0044] Download and use a total of 7 publicly available datasets, namely CIFAR-10, CIFAR-100, Cassava Disease, Imagenette, KylbergTexture, DTD, and DeepWeeds, as experimental datasets;

[0045] S2: Partition and crop the images in the original image dataset obtained in S1 to generate a cropped sub-image dataset;

[0046] For all datasets, uniformly crop them into 9 sub-images in a 3×3 grid at equal intervals according to the ratio of the original image size to the sub-image size of 15:8; for example, if the original image size is 224×224, then the sub-image size is 120×120;

[0047] S3: Construct a dataset reader and design a data preprocessing pipeline;

[0048] Set the data reading method of the dataset reader to randomly select 128 sub-images for each batch of training, and the data preprocessing pipeline is designed as follows:

[0049] S31: Resize the images to a unified size; for the CIFAR-10, CIFAR-100, and Kylberg Texture datasets, resize the images to 32×32; for the remaining four datasets, resize the images to 224×224;

[0050] S32: Pad the images with pixels of width 4 around them, and then randomly crop the images back to the same size as described in the first step;

[0051] S33: Randomly flip some images horizontally along the vertical central axis;

[0052] S34: Normalize the images using values with a three-channel mean of (0.485, 0.456, 0.406) and a variance of (0.229, 0.224, 0.225);

[0053] Among them, S32 and S33 are only used for the training set and not for the test set;

[0054] S4: Construct a convolutional neural network model as the basic model for the classification task; select appropriate loss functions and optimizers; set training parameters such as the initial learning rate and decay coefficient for model training;

[0055] Use ResNet18 as the basic model, and at the same time, for the datasets with an image preprocessing size of 32×32 in step three, reduce the size of the first-layer convolutional kernel of ResNet18 from 7×7 to 3×3 to obtain better results;

[0056] Use the cross-entropy loss function and SGD as the optimizer;

[0057] Set the initial learning rate to 0.1, the weight decay coefficient to 5×10 -4 , and the momentum to 0.9;

[0058] S5: In each batch of training, use the dataset reader obtained in S3 to select several subgraphs from the training set of the subgraph dataset obtained in S2, and input them into the model obtained in S4 for training;

[0059] The training process of the model includes three steps:

[0060] S51: Pass all inputs through different network layers of the model in sequence to complete operations such as convolution, and finally obtain the output result of the current model for the input;

[0061] S52: According to the output result and the true image label, calculate the loss value of the output through the cross-entropy loss function;

[0062] S53: According to the loss value, perform backpropagation, perform gradient descent on the model, and complete the parameter update of the model;

[0063] S54: Use the optimizer obtained in S4 to adjust the training parameters involved in S4;

[0064] The learning rate will decay to 0.01, 0.001, and 0.0001 after the 135th, 185th, and 235th iterations respectively

[0065] S55: Repeat S51 - S54, use each batch of training images to train the model until the entire training process of the model is completed; then use the test set of the original image dataset obtained in S1 to test the training result of the current model to obtain an image classification model;

[0066] S6: Input the image to be classified into the image classification model to obtain the output result, and use the partition decision mechanism to combine and judge the output result to obtain the final classification result. The specific process is as follows:

[0067] S61: For the test image, use the partition cropping algorithm obtained in S2, adopt the same cropping strategy as the training set images of the original image dataset, and crop it into several subgraphs of different partitions;

[0068] S62: For each subgraph obtained in S61, input it into the trained model in sequence to obtain the output result of the model for each subgraph;

[0069] S63: According to the output results of all subgraphs, calculate the sum of the output results of these subgraphs as the final output result of the entire original image;

[0070] S7: Conduct model evaluation; use accuracy as the model evaluation metric.

[0071] Embodiment

[0072] In this example, the hardware used is CPU: Intel(R) Xeon(R) Gold 5218 CPU @ 2.30 GHz, GPU: GeForce RTX 3090 with 24G video memory, memory: 128GB, hard disk: 8TB. The operating system is Ubuntu 18.04.5 LTS. The software is CUDA (11.2.0), cuDNN (11.2), Python (3.9.7), tensorflow - gpu (2.5.0), torch (1.9.1), torchvision (0.10.1), numpy (1.19.5), opencv - python (4.5.3.56).

[0073] The test results of the method for improving image classification accuracy using the partition decision mechanism of the present invention on each dataset in this example are shown in Table 1 below.

[0074] Table 1 Test results of the method of the present invention and the basic model on each dataset

[0075]

[0076]

[0077] As can be seen from the above table, the method of the present invention has achieved significantly higher image classification accuracy results compared to the basic model on all datasets. Therefore, the method of the present invention has the ability to stably improve image classification accuracy.

[0078] Advantages of the present invention:

[0079] 1) It has higher stability and can show a stable improvement effect in image classification accuracy in different application scenarios and datasets;

[0080] 2) Based on modular design, it has excellent portability and can be used simultaneously with other convolutional neural network model improvement methods;

[0081] 3) The principle and design idea are interpretable and the operation process is easy to understand;

[0082] 4) It has little impact on the model training efficiency and can ensure the effect without bringing too much additional computational overhead.

[0083] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method section.

[0084] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for improving the accuracy of image classification using a partition decision mechanism, characterized in that, It includes the following steps: Step 1: Collect image data for the target application scenario and perform manual annotation of category labels to form an original image dataset; Step 2: Partition and crop the images in the original image dataset to generate a sub-image dataset; Step 3: Construct a dataset reader and perform data preprocessing on several sub-images batch-selected by the dataset reader from the sub-image dataset to obtain training images; Step 31: Scale the selected several sub-images according to a preset size; Step 32: Perform pixel padding on the scaled sub-images and randomly crop the padded images to the preset size to obtain re-cropped images; Step 33: Randomly select half of the re-cropped images and flip them left and right along the vertical central axis, and combine the flipped images and the re-cropped images to form an augmented image set; Step 34: Standardize all the images in the augmented image set according to the preset three-channel mean and preset three-channel variance to generate the training images; Step 4: Construct a convolutional neural network model as a classification basic model; Step 5: Input the training images into the classification basic model for training to obtain an image classification model; In each batch of training, use the dataset reader obtained in Step 3 to select several sub-images from the training set of the sub-image dataset obtained in Step 2, and input the processed training images into the model obtained in Step 4 for training. The model training process includes: Step 51: Input the training images into the classification basic model, and pass through different network layers in turn to obtain the output result of the current model input; Step 52: Calculate the output loss value using the loss function according to the output result and the category label corresponding to the training image; Step 53: Perform backpropagation according to the output loss value, perform gradient descent on the classification basic model, and update the model parameters; Step 54: Adjust the training parameters using the optimizer; Step 55: Whether the dataset reader has finished selecting images. If not, select the next batch of training images for training and return to Step 51; If so, the model training ends, and use the original image dataset in Step 1 to test the training result of the current classification basic model to obtain the image classification model; Step 6: Partition and crop the image to be classified and input it into the image classification model in turn to obtain the output result, and use the partition decision mechanism to comprehensively decide all the output results to obtain the final classification result; The specific process of inputting the image to be classified into the image classification model for classification and recognition includes: Step 61: Partition and crop the image to be classified to obtain sub-images cropped into several different partitions; Step 62: Input each sub-image into the image classification model in turn to obtain the output result corresponding to each sub-image; Step 63: Use the partition decision mechanism for the output results of all sub-images to obtain the decision result after combining the output results of all sub-images as the final classification result of the image to be classified; The partition decision mechanism includes an ensemble learning method or a group decision-making method.

2. A method for improving the accuracy of image classification using a partition decision mechanism, characterized in that, The partition clipping method in the step 2 includes unified clipping according to a preset size, or separately clipping into a unified size according to the differences in the shape and size of the images.

3. A method for improving the accuracy of image classification using a partition decision mechanism according to claim 1, characterized in that, The dataset reader includes a preset number of images selected per batch, a selection algorithm, a preprocessing algorithm, and an algorithm for obtaining the true labels of the images; the selection algorithm includes random selection, sequential selection, or a combination of random and sequential selection; the data preprocessing includes data augmentation and normalization.

4. A method for improving the accuracy of image classification using a partition decision mechanism according to claim 1, characterized in that The loss function includes a cross-entropy loss function; the optimizer includes Adam or SGD.

Citation Information

Patent Citations

  • Mask-RCNN-based Gaofen-3 SAR image road detection method

    CN110852176A