Lightweight image recognition method considering both positioning precision and classification precision

By improving the YOLOv5 network model, using the C3_faster module and aLRPLoss loss function, the correlation problem of positioning and classification tasks in the machine vision model is solved, and high-precision Apple recognition is achieved, suitable for industrial applications.

CN120299029APending Publication Date: 2025-07-11SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510350207.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the positioning task and classification task of machine vision models lack correlation, the positive and negative samples are unbalanced, and the numerous parameters make it difficult to adjust parameters, making it difficult to take into account both positioning accuracy and classification accuracy in industrial applications.

Method used

Improve the YOLOv5 network model, replace the C3 module with the C3_faster module, and introduce aLRPLoss loss function, combining data enhancement and optimization algorithms to improve the lightweight and parameter adjustment efficiency of the model.

Benefits of technology

It realizes high-precision Apple recognition, taking into account positioning accuracy and classification accuracy, reducing the amount of parameter adjustment parameters, and is suitable for industrial scenarios with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299029A_ABST
    Figure CN120299029A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight image recognition method considering both positioning precision and classification precision. The method comprises the following steps: acquiring apple image data and performing data enhancement; performing data preprocessing on the apple image data after data enhancement, and dividing the apple image data into a training set, a verification set and a test set; constructing an apple image recognition network model, and replacing all C3 modules in the YOLOv5 network model with C3faster modules; obtaining a training set, and training the apple image recognition network model through an aLRPLoss loss function to obtain a trained apple image recognition network model; testing the apple image recognition network model based on the test set, and outputting the accuracy of image recognition; and obtaining a predicted image recognition result based on the trained apple image recognition network model. According to the method, the classification precision is improved, high-quality positioning is implemented, provable balance is provided between positive and negative samples, and the parameter adjustment quantity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image recognition technology, and particularly to a lightweight image recognition method that takes into account both positioning accuracy and classification accuracy. Background Art

[0002] With the continuous improvement of people's living standards and the continuous development of machine vision technology, machine vision technology focusing on deep learning has received increasing attention and has been applied in all aspects of life. However, the loss function provided by the official has its drawbacks. There is a lack of association between the positioning task and the classification task in the model, the positive and negative samples are unbalanced, and at the same time, there are many parameters, which makes it difficult to adjust the parameters. There is a large room for optimization. By improving some modules, the positioning accuracy and classification accuracy can be balanced, the number of parameters can be reduced, and the requirements of current industrial applications can be met. Summary of the Invention

[0003] In order to overcome the defects and deficiencies of the existing technology, the present invention provides a lightweight image recognition method that takes into account both positioning accuracy and classification accuracy, and is applied to apple recognition in industrial apple picking. The present invention improves the framework of the YOLOv5 network model, replaces some C3 modules with C3_faster modules, and introduces the aLRP Loss function to improve the classification accuracy, implement high-quality positioning, and provide a provable balance between positive and negative samples, taking into account both positioning accuracy and classification accuracy, reducing the number of parameters to be adjusted, and expanding the application scenario of deep learning technology in apple recognition.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions:

[0005] The present invention provides a lightweight image recognition method that takes into account both positioning accuracy and classification accuracy, including the following steps:

[0006] Obtain apple image data and perform data augmentation on the apple image data;

[0007] Perform data preprocessing on the apple image data after data augmentation;

[0008] Divide the apple image data after data preprocessing into a training set, a validation set, and a test set;

[0009] Construct an apple image recognition network model, and replace all C3 modules in the YOLOv5 network model with C3_faster modules;

[0010] Obtain the training set, and train the apple image recognition network model through the aLRP Loss function to obtain a trained apple image recognition network model;

[0011] Based on the test set, test the apple image recognition network model and output the accuracy rate of image recognition;

[0012] Based on the trained apple image recognition network model, the predicted image recognition result is obtained.

[0013] As a preferred technical solution, data augmentation is performed on the apple image data, specifically including:

[0014] Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0015] As a preferred technical solution, data preprocessing is performed on the apple image data after data augmentation, specifically including:

[0016] Perform noise reduction processing on the apple image data by means of mean filtering. Give a template to the target pixel on the image. This template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.

[0017] As a preferred technical solution, the classification labels of the apple image data include two labels: target image and non-target image. All targets in the training set are labeled with target boxes.

[0018] As a preferred technical solution, the apple image recognition network model is trained by the aLRPLoss loss function. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is based on the Adam optimizer.

[0019] The present invention also provides a lightweight image recognition system that takes into account both positioning accuracy and classification accuracy, including: an image data acquisition module, a data augmentation module, a data preprocessing module, a data partitioning module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module;

[0020] The image data acquisition module is used to acquire apple image data;

[0021] The data augmentation module is used to perform data augmentation on the apple image data;

[0022] The data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation;

[0023] The data partitioning module is used to partition the apple image data after data preprocessing into a training set, a validation set, and a test set;

[0024] The image recognition network model construction module is used to construct an apple image recognition network model, and replace all C3 modules in the YOLOv5 network model with C3_faster modules;

[0025] The network model training module is used to obtain a training set, train the apple image recognition network model through the aLRPLoss loss function, and obtain the trained apple image recognition network model.

[0026] The network model testing module is used to test the apple image recognition network model based on a test set and output the accuracy rate of image recognition.

[0027] The image recognition result output module is used to obtain the predicted image recognition result based on the trained apple image recognition network model.

[0028] As a preferred technical solution, the data augmentation module is used to perform data augmentation on the apple image data, specifically including:

[0029] Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0030] As a preferred technical solution, the data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation, specifically including:

[0031] Perform noise reduction processing on the apple image data by means of mean filtering. Give a template to the target pixel on the image. The template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.

[0032] As a preferred technical solution, the classification labels of the apple image data include two types of labels: target images and non-target images, and all targets in the training set are labeled with target boxes.

[0033] As a preferred technical solution, the network model training module is used to obtain a training set, train the apple image recognition network model through the aLRPLoss loss function, dynamically adjust the learning rate using the cosine annealing algorithm during the training process, adaptively adjust the learning rate of each parameter based on the exponential decay average of the squared gradient, and perform training based on the Adam optimizer.

[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0035] (1) The present invention replaces all C3 modules in the YOLOv5 network model with C3_faster modules, realizes the lightweight of the model, and improves the detection speed.

[0036] (2) The present invention trains the apple image recognition network model through the aLRPLoss loss function, takes into account both the positioning accuracy and the classification accuracy, and reduces the number of parameters to be adjusted.

[0037] (3) The recognition accuracy of the apple image recognition network model of the present invention is relatively high, reaching more than 90%, with strong applicability, which can meet the requirements in actual production and can be applied to scenarios with limited computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic flow chart of the lightweight image recognition method that takes into account both positioning accuracy and classification accuracy of the present invention;

[0039] Figure 2 It is a schematic network structure diagram of the C3_faster module of the present invention;

[0040] Figure 3 It is a schematic network structure diagram of the apple image recognition network model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] Embodiment 1

[0043] As Figure 1 shown, this embodiment provides a lightweight image recognition method that takes into account both positioning accuracy and classification accuracy, including the following steps:

[0044] S1: Obtain apple image data and perform data augmentation on the apple image data;

[0045] In this embodiment, apple image data is collected through various search engines and databases to collect relevant apple image samples, and attention needs to be paid to the balance of the apple image samples to avoid a situation where the number of samples in different categories varies too much, or through various scenarios of self-shot images to ensure that the samples are diverse enough;

[0046] In this embodiment, data augmentation operations such as random scaling, inversion, cropping, rotation, and optical transformation are performed on each apple image to increase the number of the sample data set, so that the network model can be fully trained and avoid the problem that the model has a tendency during the training process due to too many samples in some categories;

[0047] S2: Preprocess the image by means of mean filtering to exclude the interference of noise;

[0048] In this embodiment, the method of mean filtering is used to denoise the collected apple images. A template is given to the target pixel on the image. The template includes its surrounding adjacent pixels (8 pixels surrounding the target pixel, forming a filtering template, that is, including the target pixel itself), and then the average value of all pixels in the template is used to replace the original pixel value.

[0049] S3: Divide the apple image dataset into a training set, a validation set, and a test set;

[0050] In this embodiment, the apple image dataset is divided into a training set, a validation set, and a test set according to the ratio of 3:1:1. The classification labels of the apple images include two labels: target images and non-target images;

[0051] S4: Build an apple image recognition network model, and adjust and optimize the structure of the YOLOv5 network model;

[0052] As Figure 3 shown, replace all C3 modules in the YOLOv5 network model with C3_faster modules to achieve model lightweight and improve the detection speed;

[0053] As Figure 2 shown, the C3_faster module is proposed based on the FasterNet lightweight model. By combining partial convolution and MLP layers, it reduces the computational amount and memory access while maintaining efficient feature extraction. Residual connections ensure the effective transmission of information, and optional hierarchical scaling provides additional model flexibility, which helps to improve the stability and efficiency of training and inference. The introduction of the C3_faster module makes the model less complex and the detection speed faster.

[0054] S5: Set the training parameters, use the training set and the validation set to train and tune the convolutional neural network model, and obtain the network model with the best target recognition effect;

[0055] In this embodiment, set training parameters such as epoch, batch size, and learning rate, use the Adam optimizer, and after a large amount of training and debugging, obtain the network model with the best recognition effect;

[0056] Specifically, set the number of epochs to 50, the batch size to 16, and the learning rate to 0.01. The cosine annealing algorithm is used to dynamically adjust the learning rate to avoid the oscillation phenomenon caused by too fast gradient descent during training, thereby improving the training stability and generalization ability of the model. Since the Adam optimizer incorporates the concept of momentum, it accumulates the exponentially decaying average of the previous gradients to help accelerate learning. At the same time, it also uses the exponentially decaying average of the squared gradients to adaptively adjust the learning rate of each parameter, which has strong robustness and is widely used in deep learning tasks. Therefore, the Adam optimizer is used for training.

[0057] In this embodiment, the apple image recognition network model is trained by the aLRP Loss function, taking into account both the localization accuracy and the classification accuracy, and reducing the number of parameters to be adjusted.

[0058] S6: Call the network model to perform recognition tests on the test set. Use the recognition accuracy as the model evaluation criterion to verify the model performance. By comparing the recognized apple results with the marked true positions, the recognition ability of the recognition method for apples in the image can be detected, and the accuracy of the target recognition is output, thus taking into account both the localization accuracy and the classification accuracy.

[0059] The present invention can take into account both the localization accuracy and the classification accuracy, improve the high-quality localization for high-precision classification, reduce the number of parameters to be adjusted and the model complexity. Based on deep learning detection, the actual application scenarios can be effectively expanded, enabling the method to recognize multiple targets.

[0060] Embodiment 2

[0061] This embodiment provides a lightweight image recognition system that takes into account both the localization accuracy and the classification accuracy, and is used to implement the lightweight image recognition method that takes into account both the localization accuracy and the classification accuracy in the above Embodiment 1. The system includes: an image data acquisition module, a data augmentation module, a data preprocessing module, a data partitioning module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module.

[0062] In this embodiment, the image data acquisition module is used to acquire apple image data.

[0063] In this embodiment, the data augmentation module is used to perform data augmentation on the apple image data.

[0064] In this embodiment, the data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation.

[0065] In this embodiment, the data partitioning module is used to partition the apple image data after data preprocessing into a training set, a validation set, and a test set.

[0066] In this embodiment, the image recognition network model construction module is used to construct an apple image recognition network model, and replace all C3 modules in the YOLOv5 network model with C3_faster modules;

[0067] In this embodiment, the network model training module is used to obtain a training set, and train the apple image recognition network model through the aLRPLoss loss function to obtain a trained apple image recognition network model;

[0068] In this embodiment, the network model testing module is used to test the apple image recognition network model based on a test set and output the accuracy of image recognition;

[0069] In this embodiment, the image recognition result output module is used to obtain a predicted image recognition result based on the trained apple image recognition network model.

[0070] In this embodiment, the data augmentation module is used to perform data augmentation on apple image data, specifically including:

[0071] Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0072] In this embodiment, the data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation, specifically including:

[0073] Perform noise reduction processing on the apple image data through mean filtering. Give a template to the target pixel on the image. This template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.

[0074] In this embodiment, the classification labels of apple image data include two types of labels: target images and non-target images, and all targets in the training set are labeled with target boxes.

[0075] In this embodiment, the network model training module is used to obtain a training set, train the apple image recognition network model through the aLRPLoss loss function, use the cosine annealing algorithm to dynamically adjust the learning rate during the training process, adaptively adjust the learning rate of each parameter based on the exponential decay average of the squared gradient, and perform training based on the Adam optimizer.

[0076] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A lightweight image recognition method that takes into account both positioning accuracy and classification accuracy, characterized in that, It includes the following steps: Obtain apple image data and perform data augmentation on the apple image data; Perform data preprocessing on the apple image data after data augmentation; Divide the apple image data after data preprocessing into a training set, a validation set, and a test set; Construct an apple image recognition network model and replace all C3 modules in the YOLOv5 network model with C3_faster modules; Obtain the training set, train the apple image recognition network model through the aLRPLoss loss function to obtain the trained apple image recognition network model; Test the apple image recognition network model based on the test set and output the accuracy of image recognition; Obtain the predicted image recognition result based on the trained apple image recognition network model.

2. The lightweight image recognition method that takes into account both positioning accuracy and classification accuracy according to claim 1, characterized in that, Perform data augmentation on the apple image data, specifically including: Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

3. The lightweight image recognition method that takes into account both positioning accuracy and classification accuracy according to claim 1, wherein Perform data preprocessing on the apple image data after data augmentation, specifically including: Perform noise reduction processing on the apple image data by means of mean filtering. Give a template to the target pixels on the image. This template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.

4. The lightweight image recognition method that takes into account both positioning accuracy and classification accuracy according to claim 1, characterized in that, The classification labels of the apple image data include two labels: target image and non-target image. All targets in the training set are labeled with target boxes.

5. The lightweight image recognition method that takes into account both positioning accuracy and classification accuracy according to claim 1, characterized in that, Train the apple image recognition network model through the aLRPLoss loss function. During the training process, use the cosine annealing algorithm to dynamically adjust the learning rate, adaptively adjust the learning rate of each parameter based on the exponential decay average of the squared gradient, and perform training based on the Adam optimizer.

6. A lightweight image recognition system that takes into account both positioning accuracy and classification accuracy, characterized in that It includes: An image data acquisition module, a data augmentation module, a data preprocessing module, a data division module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module; The image data acquisition module is used to obtain apple image data; The data augmentation module is used to perform data augmentation on the apple image data; The data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation; The data division module is used to divide the apple image data after data preprocessing into a training set, a validation set, and a test set; The image recognition network model construction module is used to construct an apple image recognition network model and replace all C3 modules in the YOLOv5 network model with C3_faster modules; The network model training module is used to obtain the training set, train the apple image recognition network model through the aLRPLoss loss function to obtain the trained apple image recognition network model; The network model testing module is used to test the apple image recognition network model based on the test set and output the accuracy of image recognition; The image recognition result output module is used to obtain the predicted image recognition result based on the trained apple image recognition network model.

7. The lightweight image recognition system that takes into account both positioning accuracy and classification accuracy according to claim 6, wherein The data augmentation module is used to perform data augmentation on the apple image data, specifically including: Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

8. The lightweight image recognition system that takes into account both positioning accuracy and classification accuracy according to claim 6, wherein The data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation, specifically including: Perform noise reduction processing on the apple image data by means of mean filtering. Give a template to the target pixel on the image. The template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.

9. The lightweight image recognition system that takes into account both positioning accuracy and classification accuracy according to claim 6, characterized in that, The classification labels of the apple image data include two types of labels: target images and non-target images. All targets in the training set are labeled with bounding boxes.

10. The lightweight image recognition system that takes into account both positioning accuracy and classification accuracy according to claim 6, wherein The network model training module is used to obtain the training set, train the apple image recognition network model through the aLRPLoss loss function, dynamically adjust the learning rate using the cosine annealing algorithm during the training process, adaptively adjust the learning rate of each parameter based on the exponential decay average of the squared gradient, and perform training based on the Adam optimizer.