Image recognition method for realizing balance between model recognition speed and precision
By adding the CBRM module and Shuffle_Block module to the YOLOv5 network model, and combining data enhancement, preprocessing and optimization algorithms, the problem of imbalance in the recognition speed and accuracy of deep learning model is solved, and efficient Wogan image recognition is achieved.
Patent Information
- Application Number
- CN202510349730.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
The recognition accuracy and speed of deep learning models are difficult to balance, resulting in the recognition effect not meeting industrial needs.
The CBRM module is added to the YOLOv5 network model, and the first to ninth layers are replaced with the Shuffle_Block module. At the same time, data augmentation, preprocessing, training optimization technologies are used, including cosine annealing algorithm and Adam optimizer, to build the Wogan image recognition network model.
The optimal trade-off between recognition speed and accuracy is achieved. The accuracy of Wogan image recognition reaches more than 90%, which is suitable for scenarios with limited computing resources.
Smart Images

Figure CN120299026A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly relates to an image recognition method for achieving a balance between model recognition speed and accuracy. Background Art
[0002] With the continuous improvement of people's living standards and the continuous development of machine vision technology, machine vision technology with deep learning as the focus has received increasing attention and has been applied to all aspects of life. However, the recognition accuracy of the deep learning model is generally positively correlated with the model parameters. Too many parameters will lead to a slow recognition speed, which in turn affects the recognition effect and does not meet the current industrial requirements. Therefore, there is an urgent need for an image recognition technology that can achieve a balance between model recognition speed and accuracy to meet the current development needs. Summary of the Invention
[0003] In order to overcome the deficiencies and drawbacks of the prior art, the present invention provides an image recognition method for achieving a balance between model recognition speed and accuracy, which is applied to the recognition scenario of ponkan oranges, improves the framework of the YOLOv5 network model, reduces the number of parameters in its framework, and at the same time retains its response speed. It has the advantages of good adaptability, high accuracy, and fast speed, and expands the application scenarios of deep learning technology.
[0004] In order to achieve the above object, the present invention adopts the following technical solutions:
[0005] The present invention provides an image recognition method for achieving a balance between model recognition speed and accuracy, including the following steps:
[0006] Obtain ponkan orange image data and perform data augmentation on the ponkan orange image data;
[0007] Preprocess the ponkan orange image data after data augmentation;
[0008] Divide the preprocessed ponkan orange image data into a training set, a validation set, and a test set;
[0009] Construct a ponkan orange image recognition network model, add a CBRM module after the input layer of the YOLOv5 network model, and replace the first to ninth layers in the YOLOv5 network model with Shuffle_Block modules;
[0010] Train the ponkan orange image recognition network model based on the training set to obtain a trained ponkan orange image recognition network model;
[0011] Test the ponkan orange image recognition network model based on the test set and output the recognition accuracy of the image;
[0012] Obtain the predicted image recognition result based on the trained ponkan orange image recognition network model.
[0013] As a preferred technical solution, data augmentation is performed on the Wogan image data, specifically including:
[0014] Perform data augmentation operations on each Wogan image, including random scaling, inversion, cropping, rotation, and optical transformation.
[0015] As a preferred technical solution, preprocessing is performed on the Wogan image data after data augmentation, specifically including:
[0016] Perform noise reduction processing on the Wogan image data by means of mean filtering. Give a template to the target pixel on the image. This template includes the adjacent pixels around it, and replace the original pixel value with the average value of all pixels in the template.
[0017] As a preferred technical solution, the classification labels of the image data include two labels: target image and non-target image, and all targets in the training set are labeled with target boxes.
[0018] As a preferred technical solution, train the Wogan image recognition network model based on the training set. During the training process, use the cosine annealing algorithm to dynamically adjust the learning rate, adaptively adjust the learning rate of each parameter based on the exponential decay average of the squared gradient, and perform training based on the Adam optimizer.
[0019] The present invention also provides an image recognition system for achieving a balance between model recognition speed and accuracy, including: an image data acquisition module, a data augmentation module, a data preprocessing module, a data partitioning module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module;
[0020] The image data acquisition module is used to acquire Wogan image data;
[0021] The data augmentation module is used to perform data augmentation on the Wogan image data;
[0022] The data preprocessing module is used to preprocess the Wogan image data after data augmentation;
[0023] The data partitioning module is used to partition the preprocessed Wogan image data into a training set, a validation set, and a test set;
[0024] The image recognition network model construction module is used to construct a Wogan image recognition network model, add a CBRM module after the input layer of the YOLOv5 network model, and replace the first to ninth layers in the YOLOv5 network model with Shuffle_Block modules;
[0025] The network model training module is used to train the ponkan image recognition network model based on the training set to obtain the trained ponkan image recognition network model;
[0026] The network model testing module is used to test the ponkan image recognition network model based on the test set and output the accuracy rate of image recognition;
[0027] The image recognition result output module is used to obtain the predicted image recognition result based on the trained ponkan image recognition network model.
[0028] As a preferred technical solution, the data augmentation module is used to perform data augmentation on the ponkan image data, specifically including:
[0029] Perform data augmentation operations on each ponkan image, including random scaling, inversion, cropping, rotation, and optical transformation.
[0030] As a preferred technical solution, the data preprocessing module is used to preprocess the ponkan image data after data augmentation, specifically including:
[0031] Perform noise reduction processing on the ponkan image data by means of mean filtering. Given a template for the target pixel on the image, the template includes its surrounding neighboring pixels, and the average value of all pixels in the template is used to replace the original pixel value.
[0032] As a preferred technical solution, the classification labels of the image data include two labels: target image and non-target image, and all targets in the training set are labeled with target boxes.
[0033] As a preferred technical solution, the network model training module is used to train the ponkan image recognition network model based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.
[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0035] (1) The present invention adds a CBRM module after the input layer of the original YOLOv5 network model, which can extract more abstract and high-level features, increase the number of channels, and replace the first to ninth layers in the original YOLOv5 network model with Shuffle_Block modules to achieve the best trade-off between speed and accuracy.
[0036] (2) The recognition accuracy of the ponkan image recognition network model of the present invention is relatively high, the accuracy rate of recognizing ponkan reaches more than 90%, and it has strong applicability, can meet the requirements in actual production, and can be applied to scenarios with limited computing resources. Description of the Drawings
[0037] Figure 1 It is a schematic flow chart of the image recognition method for achieving the balance between model recognition speed and accuracy in the present invention;
[0038] Figure 2 It is a schematic network structure diagram of the Shuffle_Block module of the present invention;
[0039] Figure 3 It is a schematic network structure diagram of the CBRM module of the present invention;
[0040] Figure 4 It is a schematic network structure diagram of the ponkan image recognition network model of the present invention. Detailed implementation manners
[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0042] Embodiment 1
[0043] As Figure 1 shown, this embodiment provides an image recognition method for achieving the balance between model recognition speed and accuracy, including the following steps:
[0044] S1: Obtain ponkan image data and perform data augmentation on the ponkan image data;
[0045] In this embodiment, relevant ponkan image samples are collected through various search engines and databases, and attention needs to be paid to the balance of the ponkan image samples to avoid the situation where the number of samples in different categories varies greatly, or through various scenarios of self-shot images to ensure that the samples are diverse enough;
[0046] In this embodiment, data augmentation operations such as random scaling, inversion, cropping, rotation, and optical transformation are performed on each ponkan image to increase the number of the sample data set, so that the network model can be fully trained and avoid the problem that the model has a tendency during the training process due to the excessive number of samples in certain categories;
[0047] S2: Preprocess the image by means of mean filtering to exclude the interference of noise;
[0048] In this embodiment, the method of mean filtering is used to denoise the collected citrus reticulata blanco images, avoiding the influence of noise on the recognition of citrus reticulata blanco images. A template is given to the target pixel on the image, and the template includes its surrounding adjacent pixels (8 pixels surrounding the target pixel form a filtering template, that is, including the target pixel itself), and then the average value of all pixels in the template is used to replace the original pixel value.
[0049] S3: Divide the citrus reticulata blanco image dataset into a training set, a validation set, and a test set;
[0050] In this embodiment, the citrus reticulata blanco image dataset is divided into a training set, a validation set, and a test set according to the ratio of 3:1:1. The classification labels of citrus reticulata blanco images include two labels: target images and non-target images. All targets in the training set are marked with target boxes;
[0051] S4: Construct a citrus reticulata blanco image recognition network model, and adjust and optimize the structure of the YOLOv5 network model;
[0052] As Figure 4 shown, a CBRM module is added after the input layer of the original YOLOv5 network model to extract more abstract and high-level features and increase the number of channels; the first to ninth layers in the YOLOv5 network model are replaced with Shuffle_Block modules to achieve the best trade-off between speed and accuracy. In the figure, Input and Output represent the input feature map and the output feature map respectively, Conv represents the convolutional layer, Concat represents the splicing operation, and the C3 module is an improved BottleneckCSP module, which consists of 3 standard convolutional layers and multiple Bottleneck modules. Upsample is the upsampling layer, which is used to increase the spatial resolution and restore the spatial information of high-level features;
[0053] As Figure 2 shown, in the Shuffle_Block module, the DWConv layer represents the depth convolutional layer, which is a variant of the convolutional operation, mainly used to reduce the computational complexity. The biggest difference from the standard convolution is that it has an independent convolutional kernel for each channel, reducing the amount of computation and the number of parameters. The Shuffle layer represents shuffling the channels. The Shuffle_Block module is designed based on the design principle of ShuffleNet, making the input and output channels, the number of grouped convolutional groups, the degree of network fragmentation, and the speed of the element-wise operation and the memory access volume MAC on different hardware reach the best, achieving the best trade-off between speed and accuracy;
[0054] As Figure 3As shown in the figure, in the CBRM module, Input and Output in the figure represent the input feature map and the output feature map, MaxPool2d represents the max pooling layer, and Conv represents the convolutional layer, mainly to increase the number of channels and extract more abstract and high-level features.
[0055] S5: Set the training parameters, use the training set and the validation set to train and tune the convolutional neural network model, and obtain the network model with the best target recognition effect;
[0056] In this embodiment, training parameters such as epoch, batch size, and learning rate are set, the Adam optimizer is used, and after a large number of trainings and debuggings, the network model with the best recognition effect is obtained.
[0057] Specifically, set epoch to 50, batch size to 16, and learning rate to 0.01, and use the cosine annealing algorithm to dynamically adjust the learning rate to avoid the oscillation phenomenon caused by too fast gradient descent during training, thereby improving the training stability and generalization ability of the model. Since the Adam optimizer contains the concept of momentum, it accumulates the exponentially decaying average of the previous gradients to help accelerate learning. At the same time, it also uses the exponentially decaying average of the squared gradients to adaptively adjust the learning rate of each parameter, which has strong robustness and is widely used in deep learning tasks. Therefore, the Adam optimizer is used for training.
[0058] S6: Call the network model, perform recognition tests on the test set, use the recognition accuracy rate as the model evaluation criterion to verify the model performance. By comparing the recognition result of the ponkan and the marked true position, it can be detected whether the recognition method has the ability to detect ponkan, and the accuracy rate of target recognition is output, thereby completing the ponkan image recognition that balances the implementation speed and accuracy based on deep learning.
[0059] The present invention can achieve the balance between the detection and training speed and the recognition accuracy. In the case where the training model is too large due to complex usage scenarios, this method can effectively reduce the number of parameters, while maintaining the recognition speed, and can also effectively expand the actual application scenarios and can recognize multiple targets.
[0060] Embodiment 2
[0061] This embodiment provides an image recognition system that realizes the balance between the model recognition speed and the accuracy, including: an image data acquisition module, a data augmentation module, a data preprocessing module, a data partitioning module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module;
[0062] In this embodiment, the image data acquisition module is used to acquire ponkan image data;
[0063] In this embodiment, the data augmentation module is used to perform data augmentation on the citrus reticulata blanco cv. ponkan image data;
[0064] In this embodiment, the data preprocessing module is used to preprocess the citrus reticulata blanco cv. ponkan image data after data augmentation;
[0065] In this embodiment, the data division module is used to divide the preprocessed citrus reticulata blanco cv. ponkan image data into a training set, a validation set, and a test set;
[0066] In this embodiment, the image recognition network model construction module is used to construct a citrus reticulata blanco cv. ponkan image recognition network model. After the input layer of the YOLOv5 network model, a CBRM module is added, and the first to ninth layers in the YOLOv5 network model are replaced with Shuffle_Block modules;
[0067] In this embodiment, the network model training module is used to train the citrus reticulata blanco cv. ponkan image recognition network model based on the training set to obtain a trained citrus reticulata blanco cv. ponkan image recognition network model;
[0068] In this embodiment, the network model testing module is used to test the citrus reticulata blanco cv. ponkan image recognition network model based on the test set and output the accuracy rate of image recognition;
[0069] In this embodiment, the image recognition result output module is used to obtain the predicted image recognition result based on the trained citrus reticulata blanco cv. ponkan image recognition network model.
[0070] In this embodiment, the data augmentation module is used to perform data augmentation on the citrus reticulata blanco cv. ponkan image data, specifically including:
[0071] Perform data augmentation operations on each citrus reticulata blanco cv. ponkan image, including random scaling, inversion, cropping, rotation, and optical transformation.
[0072] In this embodiment, the data preprocessing module is used to preprocess the citrus reticulata blanco cv. ponkan image data after data augmentation, specifically including:
[0073] Perform noise reduction processing on the citrus reticulata blanco cv. ponkan image data by means of mean filtering. Given a template for the target pixel on the image, the template includes its surrounding adjacent pixels, and the average value of all pixels in the template is used to replace the original pixel value.
[0074] In this embodiment, the classification labels of the image data include two types of labels: target images and non-target images. All targets in the training set are labeled with target boxes.
[0075] In this embodiment, the network model training module is used to train the ponkan image recognition network model based on a training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.
[0076] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An image recognition method for achieving a balance between model recognition speed and accuracy, characterized in that Including the following steps: Obtain the image data of Wogan oranges and perform data augmentation on the image data of Wogan oranges; Preprocess the image data of Wogan oranges after data augmentation; Divide the preprocessed image data of Wogan oranges into a training set, a validation set, and a test set; Construct a Wogan orange image recognition network model, add a CBRM module after the input layer of the YOLOv5 network model, and replace the first layer to the ninth layer in the YOLOv5 network model with Shuffle_Block modules; Train the Wogan orange image recognition network model based on the training set to obtain the trained Wogan orange image recognition network model; Test the Wogan orange image recognition network model based on the test set and output the accuracy rate of image recognition; Obtain the predicted image recognition result based on the trained Wogan orange image recognition network model.
2. The image recognition method for achieving the balance between model recognition speed and accuracy according to claim 1, wherein Perform data augmentation on the image data of Wogan oranges, specifically including: Perform data augmentation operations on each Wogan orange image, including random scaling, inversion, cropping, rotation, and optical transformation.
3. The image recognition method for achieving a balance between model recognition speed and accuracy according to claim 1, characterized in that, Preprocess the image data of Wogan oranges after data augmentation, specifically including: Perform noise reduction processing on the image data of Wogan oranges by means of mean filtering. Give a template to the target pixels on the image. The template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.
4. The image recognition method for achieving the balance between model recognition speed and accuracy according to claim 1, characterized in that The classification labels of the image data include two labels: target images and non-target images. All targets in the training set are labeled with target boxes.
5. The image recognition method for achieving a balance between model recognition speed and accuracy according to claim 1, wherein Train the Wogan orange image recognition network model based on the training set. During the training process, use the cosine annealing algorithm to dynamically adjust the learning rate, adaptively adjust the learning rate of each parameter based on the exponential decay average of the squared gradient, and perform training based on the Adam optimizer.
6. An image recognition system that achieves a balance between model recognition speed and accuracy, characterized in that Including: An image data acquisition module, a data augmentation module, a data preprocessing module, a data division module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module; The image data acquisition module is used to obtain the image data of Wogan oranges; The data augmentation module is used to perform data augmentation on the image data of Wogan oranges; The data preprocessing module is used to preprocess the image data of Wogan oranges after data augmentation; The data division module is used to divide the preprocessed image data of Wogan oranges into a training set, a validation set, and a test set; The image recognition network model construction module is used to construct a Wogan orange image recognition network model, add a CBRM module after the input layer of the YOLOv5 network model, and replace the first layer to the ninth layer in the YOLOv5 network model with Shuffle_Block modules; The network model training module is used to train the Wogan orange image recognition network model based on the training set to obtain the trained Wogan orange image recognition network model; The network model testing module is used to test the Wogan orange image recognition network model based on the test set and output the accuracy rate of image recognition; The image recognition result output module is used to obtain the predicted image recognition result based on the trained Wogan orange image recognition network model.
7. The image recognition system for achieving the balance between model recognition speed and accuracy according to claim 6, wherein The data augmentation module is used to perform data augmentation on the image data of Wogan oranges, specifically including: Perform data augmentation operations on each Wogan orange image, including random scaling, inversion, cropping, rotation, and optical transformation.
8. The image recognition system for achieving the balance between model recognition speed and accuracy according to claim 6, characterized in that The data preprocessing module is used to preprocess the Wogan orange image data after data augmentation, specifically including: Perform noise reduction processing on the Wogan orange image data through the method of mean filtering. Given a template for the target pixel on the image, the template includes its surrounding neighboring pixels, and the average value of all pixels in the template is used to replace the original pixel value.
9. The image recognition system for achieving the balance between model recognition speed and accuracy according to claim 6, characterized in that The classification labels of the image data include two types of labels: target images and non-target images. All targets in the training set are labeled with target boxes.
10. The image recognition system for achieving the balance between the recognition speed and accuracy according to claim 6, wherein The network model training module is used to train the Wogan orange image recognition network model based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.