Image recognition method for enhancing small target detection and feature extraction capability

By improving the structure and training method of the YOLOv5 network model, the problem of insufficient ability of deep learning models in small object recognition and feature extraction is solved, and high-precision Apple image recognition is achieved.

CN120299024APending Publication Date: 2025-07-11SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510349673.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Deep learning models lack accuracy and insufficient ability to extract detailed features when identifying small targets, which cannot meet industrial needs.

Method used

Improve the framework of the YOLOv5 network model, replace the C3 module with the C2f module, and combine data enhancement, mean filtering and noise reduction, NWD loss function and Adam optimizer for training, and use the cosine annealing algorithm to adjust the learning rate to improve feature extraction capabilities and detection accuracy.

Benefits of technology

It improves the accuracy of small object detection, enhances feature extraction capabilities, and has an identification accuracy of more than 90%, which is suitable for scenarios with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299024A_ABST
    Figure CN120299024A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition method for enhancing small target detection and feature extraction capabilities, and the method comprises the following steps: obtaining apple image data, and carrying out the data enhancement of the apple image data; performing data preprocessing on the apple image data after data enhancement; dividing the apple image data after data preprocessing into a training set, a verification set and a test set; an apple image recognition network model is constructed, and a C3 module in front of an SPP module of the YOLOv5 network model is used as a C2f module; obtaining a training set, and training the Apple image recognition network model through an NWD loss function to obtain a trained Apple image recognition network model; testing the apple image recognition network model based on the test set, and outputting the accuracy of image recognition; and obtaining a predicted image recognition result based on the trained apple image recognition network model. According to the invention, the detection precision of the small target is improved, and the feature extraction capability is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to an image recognition method for enhancing the detection and feature extraction capabilities of small targets. Background Art

[0002] With the continuous improvement of people's living standards and the continuous development of machine vision technology, machine vision technology focusing on deep learning has received increasing attention and has been applied in all aspects of life. However, the accuracy of deep learning models in identifying small targets does not meet the ideal requirements. At the same time, when facing complex objects, their ability to extract detailed features is insufficient, which does not meet the current industrial needs. Therefore, there is an urgent need for a technology to improve the feature extraction ability and the detection accuracy of small targets to meet the current development needs. Summary of the Invention

[0003] In order to overcome the defects and deficiencies existing in the prior art, the present invention provides an image recognition method for enhancing the detection and feature extraction capabilities of small targets, which is applied to the target recognition of apple images. The present invention improves the framework of the YOLOv5 network model, improves the detection accuracy of small targets, enhances the feature extraction ability at the same time, and expands the application scenarios of deep learning technology in apple image recognition.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions:

[0005] The present invention provides an image recognition method for enhancing the detection and feature extraction capabilities of small targets, including the following steps:

[0006] Obtain apple image data and perform data augmentation on the apple image data;

[0007] Perform data preprocessing on the apple image data after data augmentation;

[0008] Divide the apple image data after data preprocessing into a training set, a validation set, and a test set;

[0009] Construct an apple image recognition network model, and change the C3 module before the SPP module of the YOLOv5 network model to a C2f module;

[0010] Obtain the training set, and train the apple image recognition network model through the NWD loss function to obtain a trained apple image recognition network model;

[0011] Based on the test set, test the apple image recognition network model and output the accuracy rate of image recognition;

[0012] Based on the trained apple image recognition network model, obtain the predicted image recognition result.

[0013] As a preferred technical solution, data augmentation is performed on the apple image data, specifically including:

[0014] Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0015] As a preferred technical solution, data preprocessing is performed on the apple image data after data augmentation, specifically including:

[0016] Perform noise reduction processing on the apple image data by means of mean filtering. Given a template for the target pixel on the image, the template includes its surrounding neighboring pixels, and the average value of all pixels in the template is used to replace the original pixel value.

[0017] As a preferred technical solution, the classification labels of the apple image data include two types of labels: target images and non-target images, and all targets in the training set are labeled with target boxes.

[0018] As a preferred technical solution, the apple image recognition network model is trained by the NWD loss function. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is based on the Adam optimizer.

[0019] The present invention also provides a lightweight image recognition system that takes into account both positioning accuracy and classification accuracy, including: an image data acquisition module, a data augmentation module, a data preprocessing module, a data division module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module;

[0020] The image data acquisition module is used to acquire apple image data;

[0021] The data augmentation module is used to perform data augmentation on the apple image data;

[0022] The data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation;

[0023] The data division module is used to divide the apple image data after data preprocessing into a training set, a validation set, and a test set;

[0024] The image recognition network model construction module is used to construct an apple image recognition network model, and change the C3 module before the SPP module of the YOLOv5 network model to a C2f module;

[0025] The network model training module is used to obtain the training set, and train the apple image recognition network model through the NWD loss function to obtain the trained apple image recognition network model;

[0026] The network model testing module is used to test the apple image recognition network model based on a test set and output the accuracy rate of image recognition.

[0027] The image recognition result output module is used to obtain the predicted image recognition result based on the trained apple image recognition network model.

[0028] As a preferred technical solution, the data augmentation module is used to perform data augmentation on the apple image data, specifically including:

[0029] Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0030] As a preferred technical solution, the data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation, specifically including:

[0031] Perform noise reduction processing on the apple image data by means of mean filtering. Give a template to the target pixel on the image. This template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.

[0032] As a preferred technical solution, the classification labels of the apple image data include two types of labels: target images and non-target images. All targets in the training set are labeled with target boxes.

[0033] As a preferred technical solution, the apple image recognition network model is trained by the NWD loss function. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is based on the Adam optimizer.

[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0035] (1) In the present invention, the C3 module before the SPP module of the YOLOv5 network model is changed to a C2f module, enhancing the feature extraction ability.

[0036] (2) In the present invention, the apple image recognition network model is trained by the NWD loss function, improving the detection ability for small targets in images.

[0037] (3) The recognition accuracy of the apple image recognition network model of the present invention is relatively high, reaching more than 90%, with strong applicability, which can meet the requirements in actual production and can be applied in scenarios with limited computing resources. Description of the Drawings

[0038] Figure 1Schematic flowchart of the image recognition method for enhancing small target detection and feature extraction ability of the present invention;

[0039] Figure 2 Schematic network structure diagram of the C2f module of the present invention;

[0040] Figure 3 Schematic network structure diagram of the Bottleneck module of the present invention;

[0041] Figure 4 Schematic network structure diagram of the apple image recognition network model of the present invention. Detailed implementation manners

[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0043] Embodiment 1

[0044] As Figure 1 shown, this embodiment provides an image recognition method for enhancing small target detection and feature extraction ability, including the following steps:

[0045] S1: Obtain apple image data and perform data augmentation on the apple image data;

[0046] In this embodiment, apple image data is collected through various search engines and databases to collect relevant apple image samples, and attention needs to be paid to the balance of apple image samples to avoid the situation where the number of samples of different categories varies greatly, or through various scenarios of self-shot images to ensure that the samples are diverse enough;

[0047] In this embodiment, data augmentation operations such as random scaling, inversion, cropping, rotation, and optical transformation are performed on each apple image to increase the number of the sample data set, so that the network model can be fully trained;

[0048] S2: Preprocess the image by the method of mean filtering to exclude the interference of noise;

[0049] In this embodiment, the method of mean filtering is used to perform noise reduction processing on the collected apple image. A template is given to the target pixel on the image. The template includes its surrounding adjacent pixels (8 pixels surrounding the target pixel form a filtering template, that is, including the target pixel itself), and then the average value of all pixels in the template is used to replace the original pixel value.

[0050] S3: Divide the apple image data set into a training set, a validation set, and a test set;

[0051] In this embodiment, the apple image dataset is divided into a training set, a validation set, and a test set according to a ratio of 3:1:1. The classification labels of the apple images include two types of labels: target images and non-target images;

[0052] S4: Construct an apple image recognition network model, and adjust and optimize the structure of the YOLOv5 network model;

[0053] As Figure 4 shown, the C3 module before the SPP module of the YOLOv5 network model is changed to a C2f module, which enhances the feature extraction ability. In the figure, Input and Output represent the input feature map and the output feature map respectively, Conv represents the convolutional layer, Concat represents the concatenation operation, the C2f module is a module, and Upsample is the upsampling layer, which is used to increase the spatial resolution and restore the spatial information of high-level features;

[0054] As Figure 2 shown, in the C2f module, first the input data passes through the first convolutional layer, and then the output is divided into two parts. One part is directly passed to the output, and the other part is processed through multiple Bottleneck modules. Finally, the results of the two parts are concatenated in the channel dimension and passed through the second convolutional layer to obtain the final output. As Figure 3 shown, the Bottleneck module is an important part of the C2f module. The Bottleneck module can reduce the computational complexity and improve the feature extraction ability;

[0055] S5: Set the training parameters, use the training set and the validation set to train and tune the convolutional neural network model, and obtain the network model with the best target recognition effect;

[0056] In this embodiment, the apple image recognition network model is trained through the NWD loss function to improve the detection ability of small targets in images. The NWD loss function calculates the similarity between boxes, models the boxes as Gaussian distributions, and then uses the Wasserstein distance to measure the similarity between the two distributions. The NWD loss function has scale invariance and is relatively gentle for position differences, which can enhance the model's detection ability for small targets;

[0057] In this embodiment, training parameters such as epoch, batch size, and learning rate are set, and the Adam optimizer is used. After a large amount of training and debugging, the network model with the best recognition effect is obtained;

[0058] Specifically, set the number of epochs to 50, the batch size to 16, and the learning rate to 0.01, and use the cosine annealing algorithm to dynamically adjust the learning rate to avoid the oscillation phenomenon caused by too fast gradient descent during training, thereby improving the training stability and generalization ability of the model. Since the Adam optimizer incorporates the concept of momentum, it accumulates the exponentially decaying average of the previous gradients to help accelerate learning. At the same time, it also uses the exponentially decaying average of the squared gradients to adaptively adjust the learning rate of each parameter, which has strong robustness and is widely used in deep learning tasks. Therefore, the Adam optimizer is used for training.

[0059] S6: Invoke the network model to perform recognition tests on the test set. Use the recognition accuracy as the model evaluation criterion to verify the model performance. By comparing the recognized apple results with the marked true positions, the recognition ability of this recognition method for apples can be detected, and the accuracy of the target recognition is output, thereby completing the image recognition that enhances the small target detection and feature extraction capabilities.

[0060] The present invention can achieve a balance between the detection and training speed and the recognition accuracy. In the case where the training model is too large due to complex usage scenarios, this method can effectively reduce the number of parameters while maintaining the recognition speed, effectively expand the actual application scenarios, and can recognize multiple targets.

[0061] Embodiment 2

[0062] This embodiment provides a lightweight image recognition system that takes into account both the positioning accuracy and the classification accuracy, and is used to implement the lightweight image recognition method that takes into account both the positioning accuracy and the classification accuracy in the above Embodiment 1. The system includes: an image data acquisition module, a data augmentation module, a data preprocessing module, a data partitioning module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module;

[0063] In this embodiment, the image data acquisition module is used to acquire apple image data;

[0064] In this embodiment, the data augmentation module is used to perform data augmentation on the apple image data;

[0065] In this embodiment, the data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation;

[0066] In this embodiment, the data partitioning module is used to partition the apple image data after data preprocessing into a training set, a validation set, and a test set;

[0067] In this embodiment, the image recognition network model construction module is used to construct an apple image recognition network model, and the C3 module before the SPP module of the YOLOv5 network model is changed to a C2f module;

[0068] In this embodiment, the network model training module is used to obtain a training set, and train the apple image recognition network model through the NWD loss function to obtain a trained apple image recognition network model;

[0069] In this embodiment, the network model testing module is used to test the apple image recognition network model based on a test set and output the accuracy rate of image recognition;

[0070] In this embodiment, the image recognition result output module is used to obtain a predicted image recognition result based on the trained apple image recognition network model.

[0071] In this embodiment, the data augmentation module is used to perform data augmentation on apple image data, specifically including:

[0072] Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0073] In this embodiment, the data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation, specifically including:

[0074] Perform noise reduction processing on the apple image data by means of mean filtering. Give a template to the target pixel on the image. This template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.

[0075] In this embodiment, the classification labels of apple image data include two types of labels: target images and non-target images, and all targets in the training set are labeled with target boxes.

[0076] In this embodiment, the apple image recognition network model is trained through the NWD loss function. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, and the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.

[0077] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and shall be included within the protection scope of the present invention.

Claims

1. An image recognition method for enhancing the small target detection and feature extraction capabilities, characterized in that, It includes the following steps: Obtain apple image data and perform data augmentation on the apple image data; Perform data preprocessing on the apple image data after data augmentation; Divide the apple image data after data preprocessing into a training set, a validation set, and a test set; Construct an apple image recognition network model, and change the C3 module before the SPP module of the YOLOv5 network model to a C2f module; Obtain the training set, train the apple image recognition network model through the NWD loss function, and obtain the trained apple image recognition network model; Test the apple image recognition network model based on the test set and output the accuracy of image recognition; Obtain the predicted image recognition result based on the trained apple image recognition network model.

2. The image recognition method for enhancing small target detection and feature extraction capabilities according to claim 1, wherein Perform data augmentation on the apple image data, specifically including: Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

3. The image recognition method for enhancing small target detection and feature extraction capabilities according to claim 1, wherein Perform data preprocessing on the apple image data after data augmentation, specifically including: Perform noise reduction processing on the apple image data by means of mean filtering. Give a template to the target pixel on the image. The template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.

4. The image recognition method for enhancing small target detection and feature extraction capabilities according to claim 1, wherein The classification labels of the apple image data include two labels: target image and non-target image. All targets in the training set are labeled with target boxes.

5. The lightweight image recognition method that takes into account both positioning accuracy and classification accuracy according to claim 1, characterized in that, Train the apple image recognition network model through the NWD loss function. During the training process, use the cosine annealing algorithm to dynamically adjust the learning rate, adaptively adjust the learning rate of each parameter based on the exponential decay average of the squared gradient, and perform training based on the Adam optimizer.

6. A lightweight image recognition system that takes into account both positioning accuracy and classification accuracy, characterized in that, It includes: An image data acquisition module, a data augmentation module, a data preprocessing module, a data division module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module; The image data acquisition module is used to obtain apple image data; The data augmentation module is used to perform data augmentation on the apple image data; The data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation; The data division module is used to divide the apple image data after data preprocessing into a training set, a validation set, and a test set; The image recognition network model construction module is used to construct an apple image recognition network model, and change the C3 module before the SPP module of the YOLOv5 network model to a C2f module; The network model training module is used to obtain the training set, train the apple image recognition network model through the NWD loss function, and obtain the trained apple image recognition network model; The network model testing module is used to test the apple image recognition network model based on the test set and output the accuracy of image recognition; The image recognition result output module is used to obtain the predicted image recognition result based on the trained apple image recognition network model.

7. The image recognition system for enhancing small target detection and feature extraction capabilities according to claim 6, wherein The data augmentation module is used to perform data augmentation on the apple image data, specifically including: Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

8. The image recognition system for enhancing small target detection and feature extraction capabilities according to claim 6, wherein, The data preprocessing module is used to preprocess the apple image data after data augmentation, specifically including: Noise reduction processing is performed on the apple image data by means of mean filtering. A template is given to the target pixel on the image. The template includes the neighboring pixels around it, and the average value of all pixels in the template is used to replace the original pixel value.

9. The image recognition system for enhancing the small target detection and feature extraction capabilities according to claim 6, characterized in that, The classification labels of the apple image data include two types of labels: target images and non-target images. All targets in the training set are labeled with target boxes.

10. The image recognition system for enhancing the small target detection and feature extraction capabilities according to claim 6, wherein The apple image recognition network model is trained through the NWD loss function. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, and the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient. Training is carried out based on the Adam optimizer.