Image recognition method for realizing low-complexity and high-precision dual-objective optimization of model
By replacing the SPP module in the YOLOv5 model with the GSConv module and adding the BiFormer module, combining data augmentation and optimization training algorithms, the model complexity and accuracy problems of image recognition technology in complex environments is solved, and efficient image recognition effect is achieved.
Patent Information
- Application Number
- CN202510350199.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
In complex environments, existing image recognition technology often leads to slow training speed and limited recognition speed, making it difficult to optimize both model complexity and accuracy.
The GSConv convolution module is used to replace the SPP module in the YOLOv5 model, and the BiFormer module is added. Combined with data augmentation and preprocessing technology, the cosine annealing algorithm and Adam optimizer are used for training to build the Wogan image recognition network model.
The dual-objective optimization of the model with low complexity and high precision is achieved, and the accuracy of image recognition is improved to more than 90%, which is suitable for scenarios with limited computing resources.
Smart Images

Figure CN120299028A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly relates to an image recognition method for realizing the dual-objective optimization of low model complexity and high precision. Background Art
[0002] With the continuous improvement of people's living standards and the continuous development of machine vision technology, image recognition technology has received increasing attention. In many complex environments, traditional vision methods can no longer meet people's requirements, and deep learning has obtained its development opportunities. Various models with good recognition speed and accuracy have emerged. However, as the recognition scenarios become more and more complex, the number of model parameters increases, resulting in slow training speed and limited recognition speed. Therefore, an image recognition technology that can balance reducing model complexity and improving accuracy is crucial in the current situation where the model scene adaptation degree is higher. Summary of the Invention
[0003] In order to overcome the defects and deficiencies existing in the prior art, the present invention provides an image recognition method for realizing the dual-objective optimization of low model complexity and high precision, which is applied to the target detection of industrial ponkan oranges. The present invention improves the shortcomings of insufficient accuracy and high model complexity of the YOLOv5 model, conducts the dual-objective optimization of low model complexity and high precision, and expands the application scenarios of deep learning technology.
[0004] In order to achieve the above object, the present invention adopts the following technical solutions:
[0005] The present invention provides an image recognition method for realizing the dual-objective optimization of low model complexity and high precision, including the following steps:
[0006] Obtain ponkan orange image data and perform data augmentation on the ponkan orange image data;
[0007] Preprocess the ponkan orange image data after data augmentation;
[0008] Divide the preprocessed ponkan orange image data into a training set, a validation set, and a test set;
[0009] Construct a ponkan orange image recognition network model, replace all Conv convolution modules after the SPP module in the YOLOv5 network model with GSConv convolution modules, and add a BiFormer module after the first GSConv convolution module after the SPP module in the replaced YOLOv5 model network;
[0010] Train the ponkan orange image recognition network model based on the training set to obtain a trained ponkan orange image recognition network model;
[0011] Test the ponkan orange image recognition network model based on the test set and output the recognition accuracy of the image;
[0012] The predicted image recognition result is obtained based on the trained image recognition network model for ponkan oranges.
[0013] As a preferred technical solution, data augmentation is performed on the ponkan orange image data, specifically including:
[0014] Data augmentation operations are performed on each ponkan orange image, including random scaling, inversion, cropping, rotation, and optical transformation.
[0015] As a preferred technical solution, preprocessing is performed on the ponkan orange image data after data augmentation, specifically including:
[0016] The ponkan orange image data is denoised by means of mean filtering. A template is given to the target pixel on the image, and the template includes its surrounding neighboring pixels. The average value of all pixels in the template is used to replace the original pixel value.
[0017] As a preferred technical solution, the GSConv convolution module is provided with channel dense convolution and depthwise separable convolution. The training set is input into the channel dense convolution, and the feature information generated by the channel dense convolution is mixed into the output of the depthwise separable convolution through uniform mixing.
[0018] As a preferred technical solution, the ponkan orange image recognition network model is trained based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.
[0019] The present invention also provides an image recognition system for realizing the dual-objective optimization of low complexity and high accuracy of the model, including: an image data acquisition module, a data augmentation module, a data preprocessing module, a data partitioning module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module;
[0020] The image data acquisition module is used to acquire ponkan orange image data;
[0021] The data augmentation module is used to perform data augmentation on the ponkan orange image data;
[0022] The data preprocessing module is used to perform preprocessing on the ponkan orange image data after data augmentation;
[0023] The data partitioning module is used to partition the preprocessed ponkan orange image data into a training set, a validation set, and a test set;
[0024] The image recognition network model construction module is used to construct a ponkan image recognition network model, replace all Conv convolution modules after the SPP module in the YOLOv5 network model with GSConv convolution modules, and add a BiFormer module after the first GSConv convolution module after the SPP module in the replaced YOLOv5 model network;
[0025] The network model training module is used to train the ponkan image recognition network model based on the training set to obtain the trained ponkan image recognition network model;
[0026] The network model testing module is used to test the ponkan image recognition network model based on the test set and output the accuracy of image recognition;
[0027] The image recognition result output module is used to obtain the predicted image recognition result based on the trained ponkan image recognition network model.
[0028] As a preferred technical solution, the data augmentation module is used to perform data augmentation on the ponkan image data, specifically including:
[0029] Perform data augmentation operations on each ponkan image, including random scaling, inversion, cropping, rotation, and optical transformation.
[0030] As a preferred technical solution, the data preprocessing module is used to preprocess the ponkan image data after data augmentation, specifically including:
[0031] Perform noise reduction processing on the ponkan image data by means of mean filtering. Give a template to the target pixels on the image. This template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.
[0032] As a preferred technical solution, the GSConv convolution module is provided with channel dense convolution and depthwise separable convolution. The training set is input into the channel dense convolution, and the feature information generated by the channel dense convolution is mixed into the output of the depthwise separable convolution through uniform mixing.
[0033] As a preferred technical solution, the network model training module is used to train the ponkan image recognition network model based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.
[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0035] (1) In the present invention, all Conv convolutional modules after the SPP module in the original YOLOv5 network model are replaced with GSConv convolutional modules, saving a large amount of computational cost while maintaining the learning ability of the model.
[0036] (2) In the present invention, a BiFormer module is added after the first Conv convolutional module (the current first GSConv convolutional module) after the SPP module in the original YOLOv5 network model. The double-layer routing attention mechanism is adopted to enhance the model's ability to capture global information and improve the accuracy of image target detection.
[0037] (3) The proposed ponkan image recognition network model in the present invention has a relatively high classification accuracy, reaching more than 90%, while reducing the complexity of the model, having strong applicability, meeting the requirements in actual production, and can be applied in scenarios with limited computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic flowchart of the image recognition method for realizing the dual-objective optimization of low complexity and high accuracy of the model in the present invention;
[0039] Figure 2 It is a schematic diagram of the network structure of the GSConv convolutional module in the present invention;
[0040] Figure 3 It is a schematic diagram of the network structure of the BiFormer module in the present invention;
[0041] Figure 4 It is a schematic diagram of the network structure of the ponkan image recognition network model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0043] Embodiment 1
[0044] As Figure 1 shown, this embodiment provides an image recognition method for realizing the dual-objective optimization of low complexity and high accuracy of the model, including the following steps:
[0045] S1: Obtain ponkan image data and perform data augmentation on the ponkan image data;
[0046] In this embodiment, the image data of ponkan oranges can be obtained online. By using various search engines and databases, relevant ponkan orange image samples are collected. Attention should be paid to the balance of the ponkan orange image samples to avoid a situation where the number of samples in different categories varies greatly. Or, by taking images in various scenarios on one's own, the samples are ensured to be diverse enough;
[0047] In this embodiment, data augmentation operations such as random scaling, inversion, cropping, rotation, and optical transformation are performed on each ponkan orange image to increase the number of the sample dataset, enabling the network model to be fully trained and making the distribution of the sample dataset more balanced.
[0048] S2: Preprocess the ponkan orange image data by means of mean filtering to eliminate the interference of noise;
[0049] In this embodiment, the method of mean filtering is used to reduce the noise of the collected ponkan orange images. Specifically, a template is given to the target pixel on the image. The template includes its surrounding neighboring pixels (8 pixels surrounding the target pixel form a filtering template, that is, including the target pixel itself), and then the average value of all the pixels in the template is used to replace the original pixel value;
[0050] S3: Divide the ponkan orange image data into a training set, a validation set, and a test set;
[0051] In this embodiment, the ponkan orange image dataset is divided into a training set, a validation set, and a test set according to the ratio of 3:1:1. The classification labels of the ponkan orange images include two types of labels: target images and non-target images. In the training set, all targets are marked with target boxes;
[0052] S4: Construct a ponkan orange image recognition network model, adjust and optimize the structure of the YOLOv5 network model, and reduce the model complexity while improving the model accuracy;
[0053] Such as Figure 4As shown in the figure, in this embodiment, all Conv convolutional modules after the SPP module in the YOLOv5 network model are replaced with GSConv convolutional modules, which can save a large amount of computational cost while maintaining the learning ability of the model. After the first Conv convolutional module (now the first GSConv convolutional module) after the SPP module of the YOLOv5 model network, a BiFormer module is added, which adopts a double-layer routing attention mechanism to enhance the model's ability to capture global information and improve the accuracy of small target detection. In the figure, Input and Output represent the input feature map and the output feature map respectively, Conv represents the convolutional layer, Concat represents the concatenation operation, the C3 module is an improved BottleneckCSP module, which consists of 3 standard convolutional layers and multiple Bottleneck modules, and Upsample is the upsampling layer, which is used to increase the spatial resolution and restore the spatial information of high-level features;
[0054] As Figure 2 shown, in the GSConv convolutional module, Input and Output in the figure represent the input feature map and the output feature map. The features generated by the channel-dense convolution SC in the GsConv convolutional module penetrate into each part of the features generated by the depthwise separable convolution DSC, and through uniform mixing, the information of the channel-dense convolution SC is fully mixed into the output of the depthwise separable convolution DSC, so as to retain the hidden connections between features as much as possible. It is mainly used in the Neck layer in the network, while the standard Backbone is retained in the Backbone layer. After using the Backbone layer to perform feature transformation on the image, the image channel dimension reaches the maximum and the width dimension reaches the minimum. Using the GSConv convolutional module to process the concatenated feature map is just right;
[0055] As Figure 3 shown, in the BiFormer module, Input and Output in the figure represent the input feature map and the output feature map. The depthwise separable convolution DWConv is a lightweight convolutional layer, which is used to capture local features and maintain computational efficiency. The double-layer routing attention Bi-level Routing Attention is the core feature of the BiFormer module, which dynamically selects the key-value pairs most relevant to each query, rather than simply interacting with all features. The layer normalization LN is used to stabilize the training process of the model, and the multi-layer perceptron MLP is a fully connected network part, which is used for non-linear feature transformation. This design enables the BiFormer module to effectively reduce the computational burden when processing large-scale image data while maintaining sensitivity to long-range dependencies. The BiFormer mechanism can enhance the model's ability to capture global information and improve the accuracy of small target detection.
[0056] S5: Set the training parameters, use the training set and the validation set to train and debug the ponkan image recognition network model, and obtain the network model with the best target recognition effect;
[0057] In this embodiment, training parameters such as epoch, batch size, and learning rate are set, the Adam optimizer is used, and after a large amount of training and debugging, the network model with the best recognition effect is obtained;
[0058] Specifically, in this embodiment, epoch is set to 50, batch size is set to 16, the learning rate is set to 0.01, and the cosine annealing algorithm is used to dynamically adjust the learning rate to avoid the oscillation phenomenon caused by too fast gradient descent during training, thereby improving the training stability and generalization ability of the model. The Adam optimizer incorporates the concept of momentum, accumulates the exponentially decaying average of the previous gradients to help accelerate learning. At the same time, it also uses the exponentially decaying average of the squared gradients to adaptively adjust the learning rate of each parameter, and has strong robustness;
[0059] S6: Invoke the network model to perform recognition tests on the test set, use the recognition accuracy rate as the model evaluation criterion to verify the model performance. By comparing the recognized ponkan results with the marked true positions, the recognition ability of the model method for ponkan in the image can be detected, and the recognition accuracy rate of the image recognition is output, thus completing the low-complexity and high-precision dual-objective optimization of the ponkan image recognition network model based on deep learning.
[0060] In this embodiment, after obtaining the ponkan image recognition network model with the recognition accuracy rate meeting the set requirements, its trained parameters can be deployed to the computer vision system. After using the camera to collect on-site images in real time, the ponkan image is used as the input of the network model, and according to the output of the model, the target in the image can be recognized in real time. The present invention can reduce the model complexity while improving the model accuracy, solve the disadvantages of many model modules, many parameters, slow training and detection speeds under complex conditions, and balance the model accuracy and model complexity.
[0061] Embodiment 2
[0062] This embodiment provides an image recognition system for realizing the dual-objective optimization of low model complexity and high precision, which is used to implement the image recognition method for realizing the dual-objective optimization of low model complexity and high precision in the above-mentioned Embodiment 1. The system includes: an image data acquisition module, a data augmentation module, a data preprocessing module, a data partitioning module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module;
[0063] In this embodiment, the image data acquisition module is used to acquire ponkan image data;
[0064] In this embodiment, the data augmentation module is used to perform data augmentation on the citrus reticulata blanco image data;
[0065] In this embodiment, the data preprocessing module is used to preprocess the citrus reticulata blanco image data after data augmentation;
[0066] In this embodiment, the data partitioning module is used to partition the preprocessed citrus reticulata blanco image data into a training set, a validation set, and a test set;
[0067] In this embodiment, the image recognition network model construction module is used to construct a citrus reticulata blanco image recognition network model, replace all Conv convolutional modules after the SPP module in the YOLOv5 network model with GSConv convolutional modules, and add a BiFormer module after the first GSConv convolutional module after the SPP module in the replaced YOLOv5 model network;
[0068] In this embodiment, the network model training module is used to train the citrus reticulata blanco image recognition network model based on the training set to obtain a trained citrus reticulata blanco image recognition network model;
[0069] In this embodiment, the network model testing module is used to test the citrus reticulata blanco image recognition network model based on the test set and output the accuracy of image recognition;
[0070] In this embodiment, the image recognition result output module is used to obtain the predicted image recognition result based on the trained citrus reticulata blanco image recognition network model.
[0071] In this embodiment, the data augmentation module is used to perform data augmentation on the citrus reticulata blanco image data, specifically including:
[0072] Perform data augmentation operations on each citrus reticulata blanco image, including random scaling, inversion, cropping, rotation, and optical transformation.
[0073] In this embodiment, the data preprocessing module is used to preprocess the citrus reticulata blanco image data after data augmentation, specifically including:
[0074] Perform noise reduction processing on the citrus reticulata blanco image data by means of mean filtering. Give a template to the target pixel on the image. The template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.
[0075] In this embodiment, the GSConv convolutional module is provided with channel dense convolution and depthwise separable convolution. Input the training set into the channel dense convolution, and mix the feature information generated by the channel dense convolution into the output of the depthwise separable convolution through uniform mixing.
[0076] In this embodiment, the network model training module is used to train the ponkan image recognition network model based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is carried out based on the Adam optimizer.
[0077] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An image recognition method for realizing the dual-objective optimization of low complexity and high precision of a model, characterized in that It includes the following steps: Obtain the image data of Wogan oranges and perform data augmentation on the image data of Wogan oranges; Preprocess the image data of Wogan oranges after data augmentation; Divide the preprocessed image data of Wogan oranges into a training set, a validation set, and a test set; Construct a Wogan orange image recognition network model, replace all Conv convolution modules after the SPP module in the YOLOv5 network model with GSConv convolution modules, and add a BiFormer module after the first GSConv convolution module after the SPP module in the replaced YOLOv5 model network; Train the Wogan orange image recognition network model based on the training set to obtain the trained Wogan orange image recognition network model; Test the Wogan orange image recognition network model based on the test set and output the accuracy of image recognition; Obtain the predicted image recognition result based on the trained Wogan orange image recognition network model.
2. The image recognition method for achieving dual-objective optimization of low complexity and high precision of the model according to claim 1, wherein Perform data augmentation on the image data of Wogan oranges, specifically including: Perform data augmentation operations on each Wogan orange image, including random scaling, inversion, cropping, rotation, and optical transformation.
3. The image recognition method for realizing the dual-objective optimization of low complexity and high precision of the model according to claim 1, characterized in that, Preprocess the image data of Wogan oranges after data augmentation, specifically including: Perform noise reduction processing on the image data of Wogan oranges by means of mean filtering. Given a template for target pixels on the image, the template includes its surrounding neighboring pixels, and the average value of all pixels in the template is used to replace the original pixel value.
4. The image recognition method for realizing the dual-objective optimization of low complexity and high precision of the model according to claim 1, wherein, The GSConv convolution module is provided with channel dense convolution and depthwise separable convolution. The training set is input into the channel dense convolution, and the feature information generated by the channel dense convolution is mixed into the output of the depthwise separable convolution through uniform mixing.
5. The image recognition method for realizing the dual-objective optimization of low complexity and high precision of the model according to claim 1, wherein, Train the Wogan orange image recognition network model based on the training set. During the training process, use the cosine annealing algorithm to dynamically adjust the learning rate, adaptively adjust the learning rate of each parameter based on the exponential decay average of the squared gradient, and perform training based on the Adam optimizer.
6. An image recognition system that realizes the dual-objective optimization of low complexity and high precision of the model, characterized in that, It includes: An image data acquisition module, a data augmentation module, a data preprocessing module, a data division module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module; The image data acquisition module is used to obtain the image data of Wogan oranges; The data augmentation module is used to perform data augmentation on the image data of Wogan oranges; The data preprocessing module is used to preprocess the image data of Wogan oranges after data augmentation; The data division module is used to divide the preprocessed image data of Wogan oranges into a training set, a validation set, and a test set; The image recognition network model construction module is used to construct a Wogan orange image recognition network model, replace all Conv convolution modules after the SPP module in the YOLOv5 network model with GSConv convolution modules, and add a BiFormer module after the first GSConv convolution module after the SPP module in the replaced YOLOv5 model network; The network model training module is used to train the Wogan orange image recognition network model based on the training set to obtain the trained Wogan orange image recognition network model; The network model testing module is used to test the ponkan image recognition network model based on a test set and output the accuracy rate of image recognition. The image recognition result output module is used to obtain the predicted image recognition result based on the trained ponkan image recognition network model.
7. The image recognition system for realizing the dual-objective optimization of low complexity and high precision of the model according to claim 1, characterized in that, The data augmentation module is used to perform data augmentation on the ponkan image data, specifically including: Performing data augmentation operations on each ponkan image, including random scaling, inversion, cropping, rotation, and optical transformation.
8. The image recognition system for realizing the dual-objective optimization of low complexity and high precision of the model according to claim 1, wherein The data preprocessing module is used to preprocess the ponkan image data after data augmentation, specifically including: Performing noise reduction processing on the ponkan image data by means of mean filtering. A template is given to the target pixel on the image, and the template includes its surrounding neighboring pixels. The average value of all pixels in the template is used to replace the original pixel value.
9. The image recognition system for realizing the dual-objective optimization of low complexity and high precision of the model according to claim 1, characterized in that, The GSConv convolution module is provided with channel dense convolution and depthwise separable convolution. The training set is input into the channel dense convolution, and the feature information generated by the channel dense convolution is mixed into the output of the depthwise separable convolution through uniform mixing.
10. The image recognition system for realizing the dual-objective optimization of low complexity and high precision of the model according to claim 1, characterized in that, The network model training module is used to train the ponkan image recognition network model based on a training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.