Image recognition method based on spatial pyramid pooling and up-sampling mode optimization
By replacing the SPPF module with the SPPFCSPC module in the YOLOv5 network model and replacing the upsampling method as transposed convolution, image recognition technology is optimized, and the problem of poor performance of the existing technology in specific application scenarios is solved, and the model is reduced, parameter redundancy is reduced, and training and recognition time is shortened.
Patent Information
- Application Number
- CN202510266602.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The existing image recognition technology is not effective in certain specific application scenarios, with large models and many redundant parameters, and long training and recognition time, making it difficult to meet the needs of industrial applications.
Improve the framework of the YOLOv5 network model, replace the SPPF module with the SPPFCSPC module, and replace the nearest neighbor interpolation upsampling method with transposed convolution, and optimize the spatial pyramid pooling and upsampling method.
While maintaining the original speed and accuracy, the model size is reduced, parameter redundancy is reduced, the model training time and recognition time are shortened, and the applicability and recognition accuracy are improved.
Smart Images

Figure CN120198775A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly relates to an image recognition method with optimized spatial pyramid pooling and upsampling methods. Background Art
[0002] With the continuous improvement of people's living standards and the continuous development of machine vision technology, machine vision technology focusing on deep learning has received increasing attention and has been applied in all aspects of life. However, the models provided by the official have strong universality and lack certain pertinence. In some specific application scenarios, some of the original modules cannot play their roles and there is great room for optimization. Therefore, there is an urgent need for an image recognition technology that can improve certain modules, reduce the size of the model, reduce parameter redundancy, greatly reduce the training time and recognition time of the model while maintaining the original speed and accuracy, and meet the needs of current industrial applications. Summary of the Invention
[0003] In order to overcome the defects and deficiencies existing in the prior art, the present invention provides an image recognition method with optimized spatial pyramid pooling and upsampling methods. The present invention improves the framework of the YOLOv5 network model, replaces the SPPF module with the SPPFCSPC module, and replaces the nearest neighbor interpolation upsampling method with transposed convolution, which has the advantages of good adaptability, high accuracy, and small number of parameters, and expands the application scenarios of deep learning technology.
[0004] In order to achieve the above object, the present invention adopts the following technical solutions:
[0005] The present invention provides an image recognition method with optimized spatial pyramid pooling and upsampling methods, including the following steps:
[0006] Obtain image data and perform data augmentation on the image data;
[0007] Perform preprocessing on the image data;
[0008] Divide the preprocessed image data into a training set, a validation set, and a test set;
[0009] Construct an image recognition network model, replace the SPPF module of the YOLOv5 network model with the SPPFCSPC module, and replace the nearest neighbor interpolation upsampling method with transposed convolution;
[0010] Train the image recognition network model based on the training set to obtain a trained image recognition network model;
[0011] Test the image recognition network model based on the test set and output the accuracy of image recognition;
[0012] The predicted image recognition result is obtained based on the trained image recognition network model.
[0013] As a preferred technical solution, data augmentation is performed on the image data, specifically including:
[0014] Data augmentation operations are performed on each image, including random scaling, inversion, cropping, rotation, and optical transformation.
[0015] As a preferred technical solution, preprocessing is performed on the image data, specifically including:
[0016] The image data is denoised by means of mean filtering. A template is given to the target pixel on the image, and the template includes the adjacent pixels around it. The average value of all the pixels in the template is used to replace the original pixel value.
[0017] As a preferred technical solution, the classification labels of the image data include two labels: target images and non-target images. All targets in the training set are labeled with target boxes.
[0018] As a preferred technical solution, the image recognition network model is trained based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.
[0019] The present invention also provides an image recognition system optimized by spatial pyramid pooling and upsampling method, including: an image data acquisition module, a data augmentation module, a data preprocessing module, a data partitioning module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module;
[0020] The image data acquisition module is used to acquire image data;
[0021] The data augmentation module is used to perform data augmentation on the image data;
[0022] The data preprocessing module is used to perform preprocessing on the image data;
[0023] The data partitioning module is used to partition the preprocessed image data into a training set, a validation set, and a test set;
[0024] The image recognition network model construction module is used to construct an image recognition network model, replace the SPPF module of the YOLOv5 network model with the SPPFCSPC module, and replace the nearest neighbor interpolation upsampling method with transposed convolution;
[0025] The network model training module is used to train the image recognition network model based on the training set to obtain the trained image recognition network model;
[0026] The network model testing module is used to test the image recognition network model based on a test set and output the accuracy of image recognition.
[0027] The image recognition result output module is used to obtain the predicted image recognition result based on the trained image recognition network model.
[0028] As a preferred technical solution, the data augmentation module is used to perform data augmentation on the image data, specifically including:
[0029] Performing data augmentation operations on each image, including random scaling, inversion, cropping, rotation, and optical transformation.
[0030] As a preferred technical solution, the data preprocessing module is used to preprocess the image data, specifically including:
[0031] Performing noise reduction processing on the image data by means of mean filtering. Given a template for the target pixel on the image, the template includes its surrounding neighboring pixels, and the average value of all pixels in the template is used to replace the original pixel value.
[0032] As a preferred technical solution, the classification labels of the image data include two types of labels: target images and non-target images, and all targets in the training set are labeled with target boxes.
[0033] As a preferred technical solution, the network model training module is used to train the image recognition network model based on a training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.
[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0035] (1) By replacing the SPPF module with the SPPFCSPC module in the YOLOv5 network model in the present invention, a speed improvement is obtained while maintaining the receptive field unchanged.
[0036] (2) By replacing the nearest neighbor interpolation upsampling method with transposed convolution in the YOLOv5 network model in the present invention, the distortion and artifacts that may be generated by the nearest neighbor interpolation method are avoided, and the recognition accuracy is improved to a certain extent compared with the nearest neighbor interpolation method.
[0037] (3) The recognition accuracy of the image recognition network model of the present invention is relatively high, reaching more than 90%, which can meet the requirements in actual production, has strong applicability, and can be applied to scenarios with limited computing resources. Description of the Drawings
[0038] Figure 1 It is a schematic flow chart of an image recognition method with optimized spatial pyramid pooling and upsampling methods according to the present invention;
[0039] Figure 2 It is a schematic network structure diagram of the SPPFCSPC module of the present invention;
[0040] Figure 3 It is a schematic network structure diagram of an image recognition network model improved based on the YOLOv5 network model of the present invention. Specific Embodiments
[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0042] Embodiment 1
[0043] As Figure 1 shown, this embodiment provides an image recognition method with optimized spatial pyramid pooling and upsampling methods, including the following steps:
[0044] S1: Obtain image data and perform data augmentation on the image data;
[0045] In this embodiment, image data can be obtained online, and relevant image samples can be collected through various search engines and databases. Attention should be paid to the balance of the image samples to avoid a situation where the number of samples in different categories varies greatly, or by taking images of various scenes by oneself to ensure that the samples are diverse enough;
[0046] In order to make the distribution of the sample data set more balanced, perform data augmentation operations such as random scaling, inversion, cropping, rotation, and optical transformation on each image to increase the number of the sample data set, so that the network model can be fully trained, and make the distribution of the sample data set more balanced, avoiding the problem that the model has a tendency during the training process due to the excessive number of samples in certain categories.
[0047] S2: Preprocess the image by means of mean filtering to exclude the interference of noise;
[0048] In this embodiment, use the mean filtering method to perform noise reduction processing on the collected image to avoid the influence of noise on image recognition. Specifically, give a template to the target pixel on the image. The template includes its surrounding adjacent pixels (8 pixels surrounding the target pixel form a filtering template, that is, including the target pixel itself), and then use the average value of all pixels in the template to replace the original pixel value.
[0049] S3: Divide the preprocessed image dataset into a training set, a validation set, and a test set;
[0050] In this embodiment, the preprocessed image dataset is divided into a training set, a validation set, and a test set according to a ratio of 3:1:1. The classification labels of the image dataset include two types: target images and non-target images. In the training set, all targets are labeled with bounding boxes;
[0051] S4: Construct an image recognition network model, adjust and optimize the structure of the model. Based on the YOLOv5 network model, the backbone network uses the original framework, only improving the upsampling method and the spatial pyramid pooling method, replacing the SPPF module with the SPPFCSPC module, and replacing the nearest neighbor interpolation upsampling method with transposed convolution;
[0052] As Figure 3 shown, the image recognition network model is improved based on the YOLOv5 network model. In the figure, Input represents the input feature map, Output1, Output2, and Output3 represent the output feature maps, Conv represents the convolutional layer, Concat represents the concatenation operation, and the CSP (Cross-Stage Partial) layer is an optimization technique for improving network performance. It achieves efficient feature transfer and gradient flow by dividing different parts of the network, optimizing the computational efficiency and reducing the number of parameters;
[0053] In this embodiment, in the network of the image recognition network model, the SPPF module is replaced with the SPPFCSPC module, and a speed improvement is obtained while keeping the receptive field unchanged. As Figure 2 shown, Maxpool2d is the max pooling layer, Concat represents the concatenation operation, Conv represents the convolutional layer, and the SPPFCSPC module draws on the idea of optimizing the SPP module based on the SPPF module to optimize the SPPCSPC module and obtain a speed improvement while keeping the receptive field unchanged;
[0054] In this embodiment, the image recognition network model replaces the nearest neighbor interpolation (nearest) upsampling method with transposed convolution (ConvTranspose2d) in the network. Since the images generated by the nearest neighbor interpolation method lack smoothness and are prone to jaggedness and artifacts, while transposed convolution can generate smoother and more natural details, avoiding the distortion and artifacts that may be generated by the nearest neighbor interpolation method. Compared with the nearest neighbor interpolation method, its recognition accuracy has a certain degree of improvement;
[0055] S5: Set the training parameters, use the training set and the validation set to train and tune the convolutional neural network model, and obtain the network model with the best target recognition effect;
[0056] In this embodiment, training parameters such as the number of epochs, batch size, and learning rate are set. The Adam optimizer is used. After a large amount of training and debugging, the network model with the best recognition effect is obtained;
[0057] Specifically, set the number of epochs to 50, the batch size to 16, and the learning rate to 0.01, and use the cosine annealing algorithm to dynamically adjust the learning rate to avoid the oscillation phenomenon caused by too fast gradient descent during training, thereby improving the training stability and generalization ability of the model. Since the Adam optimizer incorporates the concept of momentum, it accumulates the exponentially decaying average of the previous gradients to help accelerate learning. At the same time, it also uses the exponentially decaying average of the squared gradients to adaptively adjust the learning rate of each parameter, and has strong robustness and is widely used in deep learning tasks. Therefore, the Adam optimizer is used for training. After a large amount of training and debugging, the network model with the best target recognition effect is obtained.
[0058] S6: Call the network model to perform recognition tests on the test set. Use the recognition accuracy rate as the model evaluation criterion to verify the model performance. By comparing the recognition results with the marked true positions, it can be detected whether the recognition method has the ability of target detection, and the recognition accuracy rate of the target is output, thereby completing the image recognition optimized by the spatial pyramid pooling and upsampling method. While ensuring that the receptive field remains unchanged, the speed is increased, and at the same time, the recognition accuracy is improved. It has the advantages of low cost, small implementation difficulty, strong applicability, fast speed, high recognition accuracy, etc., can effectively expand the actual application scenarios, and can recognize multiple targets.
[0059] Embodiment 2
[0060] This embodiment provides an image recognition system optimized by the spatial pyramid pooling and upsampling method for implementing the image recognition method optimized by the spatial pyramid pooling and upsampling method in the above Embodiment 1. The system includes: an image data acquisition module, a data augmentation module, a data preprocessing module, a data partitioning module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module;
[0061] In this embodiment, the image data acquisition module is used to acquire image data;
[0062] In this embodiment, the data augmentation module is used to perform data augmentation on the image data;
[0063] In this embodiment, the data preprocessing module is used to preprocess the image data;
[0064] In this embodiment, the data partitioning module is used to partition the preprocessed image data into a training set, a validation set, and a test set;
[0065] In this embodiment, the image recognition network model construction module is used to construct an image recognition network model, replace the SPPF module of the YOLOv5 network model with the SPPFCSPC module, and replace the nearest neighbor interpolation upsampling method with transposed convolution;
[0066] In this embodiment, the network model training module is used to train the image recognition network model based on the training set to obtain the trained image recognition network model;
[0067] In this embodiment, the network model testing module is used to test the image recognition network model based on the test set and output the accuracy of image recognition;
[0068] In this embodiment, the image recognition result output module is used to obtain the predicted image recognition result based on the trained image recognition network model.
[0069] In this embodiment, the data augmentation module is used to perform data augmentation on the image data, specifically including:
[0070] Performing data augmentation operations on each image, including random scaling, inversion, cropping, rotation, and optical transformation.
[0071] In this embodiment, the data preprocessing module is used to preprocess the image data, specifically including:
[0072] Performing noise reduction processing on the image data by means of mean filtering. Given a template for the target pixel on the image, the template includes its surrounding neighboring pixels, and the average value of all pixels in the template is used to replace the original pixel value.
[0073] In this embodiment, the classification labels of the image data include two labels: target image and non-target image, and all targets in the training set are labeled with target boxes.
[0074] In this embodiment, the network model training module is used to train the image recognition network model based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.
[0075] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An image recognition method with spatial pyramid pooling and upsampling optimization, characterized in that: The steps include: Acquire image data and perform data enhancement on the image data; Preprocess the image data; Divide the preprocessed image data into training set, validation set and test set; Build an image recognition network model, replace the SPPF module of the YOLOv5 network model with the SPPFCSPC module, and replace the nearest neighbor interpolation upsampling method with transposed convolution; The image recognition network model is trained based on the training set to obtain a trained image recognition network model; Test the image recognition network model based on the test set and output the image recognition accuracy; The predicted image recognition results are obtained based on the trained image recognition network model.
2. The image recognition method according to claim 1, characterized in that: Perform data enhancement on image data, including: Data augmentation operations are performed on each image, including random scaling, inversion, cropping, rotation, and optical transformation.
3. The image recognition method according to claim 1, characterized in that: Preprocess the image data, including: The image data is denoised by the mean filtering method. A template is given to the target pixel on the image, which includes the adjacent pixels around it, and the original pixel value is replaced by the average value of all pixels in the template.
4. The image recognition method according to claim 1, characterized in that: The classification labels of image data include two labels: target image and non-target image. All targets in the training set are marked with target boxes.
5. The image recognition method according to claim 1, characterized in that: The image recognition network model is trained based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate. The learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient. The training is performed based on the Adam optimizer.
6. An image recognition system with spatial pyramid pooling and upsampling optimization, characterized in that: include: Image data acquisition module, data enhancement module, data preprocessing module, data partitioning module, image recognition network model building module, network model training module, network model testing module, image recognition result output module; The image data acquisition module is used to acquire image data; The data enhancement module is used to perform data enhancement on the image data; The data preprocessing module is used to preprocess the image data; The data division module is used to divide the preprocessed image data into a training set, a verification set and a test set; The image recognition network model construction module is used to construct an image recognition network model, replace the SPPF module of the YOLOv5 network model with the SPPFCSPC module, and replace the nearest neighbor interpolation upsampling method with transposed convolution; The network model training module is used to train the image recognition network model based on the training set to obtain the trained image recognition network model; The network model testing module is used to test the image recognition network model based on the test set and output the accuracy of image recognition; The image recognition result output module is used to obtain predicted image recognition results based on the trained image recognition network model.
7. The image recognition system according to claim 6, characterized in that: The data enhancement module is used to perform data enhancement on the image data, specifically including: Data augmentation operations are performed on each image, including random scaling, inversion, cropping, rotation, and optical transformation.
8. The image recognition system according to claim 6, characterized in that: The data preprocessing module is used to preprocess the image data, specifically including: The image data is denoised by the mean filtering method. A template is given to the target pixel on the image, which includes the adjacent pixels around it, and the original pixel value is replaced by the average value of all pixels in the template.
9. The image recognition system according to claim 6, characterized in that: The classification labels of image data include two labels: target image and non-target image. All targets in the training set are marked with target boxes.
10. The image recognition system with spatial pyramid pooling and upsampling optimization according to claim 6, characterized in that: The network model training module is used to train the image recognition network model based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the square gradient, and the training is performed based on the Adam optimizer.