Small target image recognition method based on deep learning

By introducing SimAm and SPP modules in the YOLOv5 model, the small-object image recognition capability is improved, the problem of low accuracy of small-object recognition in the existing technology is solved, and the high accuracy and applicability are improved.

CN120220137APending Publication Date: 2025-06-27SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510260276.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has poor image recognition effect on fruit picking small and small targets, especially when the detection target is small, the model trains few feature points, resulting in low recognition accuracy.

Method used

Improve the YOLOv5 model and add SimAm module and SPP module. The SimAm module adjusts attention weight by calculating the similarity between different positions or channels in the feature map. The SPP module uses pooling windows of different sizes to pool the input images at different scales, and splicing all pooled feature vectors into fixed-sized feature vectors.

Benefits of technology

The model's ability to identify small targets has been improved, the classification accuracy rate reaches more than 90%, and it is highly applicable, and can be effectively applied in scenarios with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220137A_ABST
    Figure CN120220137A_ABST
Patent Text Reader

Abstract

The invention discloses a small target image recognition method based on deep learning, and the method comprises the following steps: obtaining small target image data, and carrying out the data enhancement of the small target image data; preprocessing the small target image data; dividing the small target image data into a training set, a verification set and a test set; constructing a small target image recognition network model, and adding a SimAm module and an SPP module into the YOLOv5 network model; training the small target image recognition network model based on the training set to obtain a trained small target image recognition network model; testing the small target image recognition network model based on the test set, and outputting small target image recognition accuracy; and obtaining a predicted small target image recognition result based on the trained small target image recognition network model. According to the method, the defect of poor small target detection capability of the YOLOv5 model is improved, and the method has the advantages of good adaptability, high accuracy, high speed and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to a small target image recognition method based on deep learning. Background Art

[0002] With the continuous development of machine vision technology, small target image recognition technology has attracted more and more attention. In the fruit industry, fruit picking is an important part, which is very time-consuming and laborious, seriously affecting the supply efficiency of fruits. And an important link in intelligent fruit picking is fruit recognition. Especially small targets such as grapes, apples, and jujubes are the key breakthrough points for fruit target recognition;

[0003] In the field of fruit picking, the proportion of using intelligent picking is not high because traditional image processing has high requirements for the environment, and occlusion, uneven light, etc. will all affect the recognition effect. The emergence of machine learning has improved this situation. Machine learning enables the recognition accuracy not to overly rely on the original image, and through a large number of pre-training, the model has a certain predictability, thus can solve some complex situations. However, in some special cases, machine learning still cannot achieve good results. For example, when the detected target is small, the number of feature points extracted during training is small, and the training effect of the model is not good. Therefore, by improving the applicability of the original model for small target detection, the small target recognition ability of the model can be enhanced to achieve better recognition effects for specific small targets. Therefore, there is an urgent need for a small target recognition technology with good applicability, high speed, and high recognition accuracy. Summary of the Invention

[0004] In order to overcome the defects and deficiencies existing in the prior art, the present invention provides a small target image recognition method based on deep learning. The present invention improves the shortcoming of the poor small target detection ability of the YOLOv5 model, and has the advantages of good adaptability, high accuracy, high speed, etc., expanding the application scenarios of deep learning technology.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] The present invention provides a small target image recognition method based on deep learning, including the following steps:

[0007] Obtain small target image data and perform data augmentation on the small target image data;

[0008] Preprocess the small target image data;

[0009] Divide the small target image data into a training set, a validation set, and a test set;

[0010] Build a small target image recognition network model. Add the SimAm module and the SPP module to the YOLOv5 network model. The SimAm module adjusts the attention weights by calculating the similarity between different positions or channels in the feature map. The SPP module performs pooling operations on the input image at different scales using pooling windows of different sizes, and splices all the pooled feature vectors to form a feature vector of a fixed size;

[0011] Train the small target image recognition network model based on the training set to obtain the trained small target image recognition network model;

[0012] Test the small target image recognition network model based on the test set and output the recognition accuracy of the small target image;

[0013] Obtain the predicted recognition result of the small target image based on the trained small target image recognition network model.

[0014] As a preferred technical solution, perform data augmentation on the small target image data, specifically including:

[0015] Perform data augmentation operations on each small target image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0016] As a preferred technical solution, perform preprocessing on the small target image data, specifically including:

[0017] Preprocess the small target image by means of mean filtering. Give a template to the target pixels on the image. The template includes the adjacent pixels around it, and then replace the original pixel value with the average value of all the pixels in the template.

[0018] As a preferred technical solution, the classification labels of the small target images include two labels: target images and non-target images. All targets in the training set are labeled with target boxes.

[0019] As a preferred technical solution, the SPP module includes an input layer, multiple max pooling layers, a splicing layer, and an output layer;

[0020] The input layer inputs the feature map. Each max pooling layer performs pooling operations on the input feature map using pooling windows of different sizes, evenly divides the feature maps of different sizes into several grids of different sizes, performs pooling operations on the feature maps within each grid. The splicing layer splices all the pooled feature vectors, and the output layer outputs the spliced feature vector of a fixed size.

[0021] The present invention provides a small target image recognition system based on deep learning, including: a small target image data acquisition module, a data augmentation module, a data preprocessing module, a data partitioning module, a network model construction module, a network model training module, a network model testing module, and an image recognition result output module;

[0022] The small target image data acquisition module is used to acquire small target image data;

[0023] The data augmentation module is used to perform data augmentation on the small target image data;

[0024] The data preprocessing module is used to preprocess the small target image data;

[0025] The data partitioning module is used to partition the small target image data into a training set, a validation set, and a test set;

[0026] The network model construction module is used to construct a small target image recognition network model. The SimAm module and the SPP module are added to the YOLOv5 network model. The SimAm module adjusts the attention weights by calculating the similarity between different positions or channels in the feature map. The SPP module performs pooling operations on the input image at different scales using pooling windows of different sizes, and splices all the pooled feature vectors to form a feature vector of a fixed size;

[0027] The network model training module is used to train the small target image recognition network model based on the training set to obtain a trained small target image recognition network model;

[0028] The network model testing module is used to test the small target image recognition network model based on the test set and output the recognition accuracy of the small target image;

[0029] Based on the trained small target image recognition network model, the predicted small target image recognition result is obtained.

[0030] As a preferred technical solution, the data augmentation module is used to perform data augmentation on the small target image data, specifically including:

[0031] Performing data augmentation operations on each small target image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0032] As a preferred technical solution, the data preprocessing module is used to preprocess the small target image data, specifically including:

[0033] Preprocessing the small target image by the method of mean filtering. A template is given to the target pixel on the image, and the template includes its surrounding adjacent pixels. Then, the average value of all the pixels in the template is used to replace the original pixel value.

[0034] As a preferred technical solution, the classification labels of the small target images include two types of labels: target images and non-target images, and all targets in the training set are labeled with target boxes.

[0035] As a preferred technical solution, the SPP module includes an input layer, multiple max pooling layers, a concatenation layer, and an output layer;

[0036] The input layer inputs a feature map. Each max pooling layer performs a pooling operation on the input feature map using pooling windows of different sizes, evenly divides the feature maps of different sizes into several grids of different sizes, and performs a pooling operation on the feature map within each grid. The concatenation layer concatenates all the pooled feature vectors, and the output layer outputs the concatenated feature vectors of a fixed size.

[0037] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0038] (1) By introducing the SPP module into the YOLOv5 network model, the present invention effectively avoids problems such as image distortion caused by cropping and scaling operations on image regions.

[0039] (2) By introducing the SimAm module into the YOLOv5 network model, the present invention enhances the model's recognition ability for small targets.

[0040] (3) The classification accuracy of the present invention is relatively high, reaching more than 90%, which can meet the requirements in actual production, has strong applicability, and can be applied in scenarios with limited computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a schematic flow chart of the method for identifying small target images based on deep learning of the present invention;

[0042] Figure 2 is a schematic network structure diagram of the SPP module of the present invention;

[0043] Figure 3 is a schematic network structure diagram of the small target image recognition network model improved based on the YOLOv5 network model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0045] Embodiment 1

[0046] As Figure 1As shown in the figure, this embodiment provides a small target image recognition method based on deep learning, including the following steps:

[0047] S1: Obtain small target image data and perform data augmentation on the small target image data;

[0048] In this embodiment, small target images can be obtained online. Through various search engines and databases, relevant small target image samples are collected, and attention needs to be paid to the balance of the small target image samples to avoid the situation where the number of samples of different categories varies greatly, or by taking pictures of various scenes of small target images by oneself to ensure that the samples are diverse enough;

[0049] In this embodiment, in order to make the distribution of the sample dataset more balanced, data augmentation operations such as random scaling, inversion, cropping, rotation, and optical transformation are performed on each small target image to increase the number of the sample dataset, so that the network model can be fully trained and avoid the problem that the model has a tendency during the training process due to the excessive number of samples of certain categories;

[0050] S2: Preprocess the small target image by means of mean filtering to exclude the interference of noise;

[0051] In this embodiment, the mean filtering method is used to perform noise reduction processing on the collected small target image. A template is given to the target pixel on the image. The template includes its surrounding adjacent pixels (8 pixels surrounding the target pixel form a filtering template, that is, including the target pixel itself), and then the average value of all pixels in the template is used to replace the original pixel value;

[0052] S3: Divide the small target image data into a training set, a validation set, and a test set;

[0053] In this embodiment, the small target image dataset is divided into a training set, a validation set, and a test set according to a ratio of 3:1:1. The classification labels of the small target images include two labels: target images and non-target images. All targets in the training set are marked with target boxes;

[0054] S4: Build a small target image recognition network model, adjust and optimize the structure of the YOLOv5 network model, and add a SimAm module and an SPP module to the model, which has little impact on the calculation amount and training time while improving the applicability and classification accuracy of the model;

[0055] Such as Figure 3As shown in the figure, Input in the figure represents the input feature map, Output1 and Output2 represent the output feature maps, Conv represents the convolutional layer, Concat represents the concatenation operation, the CSP layer is used to optimize the computational efficiency and enhance the expressive power of the model. By introducing the CSP layer, the YOLOv5 network model solves the problems of gradient disappearance and explosion, enabling the network to better maintain the gradient flow and improve the stability of training; through the segmentation and partial calculation of the feature map, the redundant computational amount is reduced, accelerating the training and inference processes, and it is widely used in the field of deep learning. Upsample is the upsampling layer, which is used to increase the spatial resolution and restore the spatial information of high-level features.

[0056] In this embodiment, the small target image recognition network model is improved based on the YOLOv5 network model, and the SimAm module is added to the network to enhance the model's recognition ability for small target objects. The SimAm module adjusts the attention weights by calculating the similarity between different positions or channels in the feature map, without relying on an explicit weighting mechanism (such as the common weighted sum operation). Its goal is to control the information flow in the feature map through similarity, thereby enhancing the expressive power of the model. Different from the traditional attention mechanism that calculates through weighted sum, the SimAm module does not rely on complex weighted operations, thus reducing the computational amount; SimAm pays more attention to the similarity relationship between features, which helps to improve the representation ability of the model. Especially in tasks such as contrastive learning and unsupervised learning, it can better capture the potential associations between features.

[0057] In this embodiment, the small target image recognition network model adds a Spatial Pyramid Pooling (SPP) module to the network to avoid image distortion problems caused by operations such as cropping and scaling of image regions, such as Figure 2As shown, Input and Output represent the input feature map and the output feature map respectively, MaxPool2d represents the max pooling layer, n×n represents the size of the pooling kernel, Concat represents the concatenation operation. The SPP module effectively avoids problems such as image distortion caused by cropping and scaling operations on image regions, solves the problem of the convolutional neural network extracting duplicate features related to the graph, saves computational costs, and the multi-scale features extracted by spatial pyramid pooling significantly enhance the feature extraction ability of the network, which helps to improve the performance of the network model. It performs pooling operations on the input image at different scales using pooling windows of different sizes, and finally concatenates these pooling results to form a feature vector of a fixed size. Specifically, it evenly divides feature maps of different sizes into several grids of different sizes, performs pooling operations on the feature maps within each grid, and then connects all the pooled feature vectors as the input of the network. In this way, even if the input image sizes are different, a feature vector of a fixed length can be obtained for network training and inference.

[0058] In this embodiment, the NWD loss function is added to better calculate the loss in cooperation with the IOU loss function;

[0059] S5: Set the training parameters, use the training set and the validation set to train and tune the convolutional neural network model, and obtain the network model with the best small target recognition effect;

[0060] In this embodiment, training parameters such as epoch, batch size, and learning rate are set, the Adam optimizer is used, and after a large number of trainings and debuggings, the network model with the best small target recognition effect is obtained;

[0061] Specifically, in this embodiment, epoch is set to 50, batch size is set to 16, the learning rate is set to 0.01, and the cosine annealing algorithm is used to dynamically adjust the learning rate to avoid the oscillation phenomenon caused by too fast gradient descent during training, thereby improving the training stability and generalization ability of the model. Since the Adam optimizer contains the concept of momentum, it accumulates the exponentially decaying average of the previous gradients to help accelerate learning. At the same time, it also uses the exponentially decaying average of the squared gradients to adaptively adjust the learning rate of each parameter, and has strong robustness and is widely used in deep learning tasks. Therefore, the Adam optimizer is used for training. After a large number of trainings and debuggings, the network model with the best small target recognition effect is obtained.

[0062] S6: Invoke the network model to conduct classification tests on the test set. Use the recognition accuracy rate as the model evaluation criterion to verify the model performance. By comparing the recognition results with the marked true positions, it can be detected whether the recognition method has the ability to recognize small targets, and the accuracy rate of small target recognition is output, thus completing the small target image recognition based on deep learning.

[0063] In this embodiment, after obtaining a network model with a recognition accuracy rate meeting the requirements, deploy its trained parameters to the computer vision system. After using the camera to collect on-site images in real time, use the images as the input of the network model. According to the output of the model, the targets in the images can be recognized in real time. Cooperating with motion control systems such as robotic arms, the picking work can be realized, which can effectively improve the efficiency and reduce the consumption of unnecessary human and material resources, effectively alleviate the difficulty of manual picking in the case of a large number of fruits, effectively expand the actual application scenarios, and has the advantages of low cost, small implementation difficulty, strong applicability, high speed, good detection effect, etc.

[0064] Embodiment 2

[0065] This embodiment provides a small target image recognition system based on deep learning for implementing the small target image recognition method based on deep learning in the above Embodiment 1. The system includes: a small target image data acquisition module, a data augmentation module, a data preprocessing module, a data partitioning module, a network model construction module, a network model training module, a network model testing module, and an image recognition result output module;

[0066] In this embodiment, the small target image data acquisition module is used to acquire small target image data;

[0067] In this embodiment, the data augmentation module is used to perform data augmentation on the small target image data;

[0068] In this embodiment, the data preprocessing module is used to preprocess the small target image data;

[0069] In this embodiment, the data partitioning module is used to partition the small target image data into a training set, a validation set, and a test set;

[0070] In this embodiment, the network model construction module is used to construct a small target image recognition network model. Add the SimAm module and the SPP module to the YOLOv5 network model. The SimAm module adjusts the attention weights by calculating the similarity between different positions or channels in the feature map. The SPP module performs pooling operations on the input image at different scales using pooling windows of different sizes, and splices all the pooled feature vectors to form a feature vector of a fixed size;

[0071] In this embodiment, the network model training module is used to train the small target image recognition network model based on the training set to obtain the trained small target image recognition network model;

[0072] In this embodiment, the network model testing module is used to test the small target image recognition network model based on the test set and output the recognition accuracy of the small target image;

[0073] Based on the trained small target image recognition network model, the predicted small target image recognition result is obtained.

[0074] In this embodiment, the data augmentation module is used to perform data augmentation on the small target image data, specifically including:

[0075] Perform data augmentation operations on each small target image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0076] In this embodiment, the data preprocessing module is used to preprocess the small target image data, specifically including:

[0077] Preprocess the small target image by means of mean filtering. Give a template to the target pixel on the image. The template includes its surrounding adjacent pixels, and then replace the original pixel value with the average value of all pixels in the template.

[0078] In this embodiment, the classification labels of the small target images include two types of labels: target images and non-target images. All targets in the training set are labeled with target boxes.

[0079] In this embodiment, the SPP module includes an input layer, multiple max-pooling layers, a concatenation layer, and an output layer;

[0080] The input layer inputs the feature map. Each max-pooling layer performs a pooling operation on the input feature map using pooling windows of different sizes, evenly divides the feature maps of different sizes into several grids of different sizes, and performs a pooling operation on the feature map within each grid. The concatenation layer concatenates all the pooled feature vectors, and the output layer outputs the concatenated feature vectors of a fixed size.

[0081] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and shall be included in the protection scope of the present invention.

Claims

1. A small target image recognition method based on deep learning, characterized in that: The steps include: Acquire small target image data, and perform data enhancement on the small target image data; Preprocess small target image data; Divide the small object image data into training set, validation set and test set; Construct a small object image recognition network model, add SimAm module and SPP module to the YOLOv5 network model, the SimAm module adjusts the attention weight by calculating the similarity between different positions or channels in the feature map, and the SPP module uses pooling windows of different sizes to perform pooling operations on the input image at different scales, and concatenates all pooled feature vectors to form a feature vector of fixed size; The small object image recognition network model is trained based on the training set to obtain a trained small object image recognition network model; Test the small object image recognition network model based on the test set and output the small object image recognition accuracy; The predicted small object image recognition results are obtained based on the trained small object image recognition network model.

2. The small target image recognition method based on deep learning according to claim 1 is characterized in that: Data enhancement is performed on small target image data, including: Data augmentation operations are performed on each small target image, including random scaling, inversion, cropping, rotation, and optical transformation.

3. The small target image recognition method based on deep learning according to claim 1, characterized in that: Preprocess the small target image data, including: The small target image is preprocessed by the mean filtering method. A template is given to the target pixel on the image, which includes the neighboring pixels around it, and then the original pixel value is replaced by the average value of all pixels in the template.

4. The small target image recognition method based on deep learning according to claim 1, characterized in that: The classification labels of the small target images include two labels: target image and non-target image. All targets in the training set are annotated with target boxes.

5. The small target image recognition method based on deep learning according to claim 1, characterized in that: The SPP module includes an input layer, multiple maximum pooling layers, a splicing layer and an output layer; The input layer inputs a feature map, each maximum pooling layer uses pooling windows of different sizes to perform pooling operations on the input feature map, and evenly divides the feature maps of different sizes into a number of grids of different sizes. The feature map in each grid is pooled, the splicing layer splices all the pooled feature vectors, and the output layer outputs the spliced ​​feature vector of a fixed size.

6. A small object image recognition system based on deep learning, characterized in that: include: Small target image data acquisition module, data enhancement module, data preprocessing module, data partitioning module, network model construction module, network model training module, network model testing module, image recognition result output module; The small target image data acquisition module is used to acquire small target image data; The data enhancement module is used to perform data enhancement on small target image data; The data preprocessing module is used to preprocess the small target image data; The data division module is used to divide the small object image data into a training set, a verification set and a test set; The network model building module is used to build a small target image recognition network model. A SimAm module and an SPP module are added to the YOLOv5 network model. The SimAm module adjusts the attention weight by calculating the similarity between different positions or channels in the feature map. The SPP module uses pooling windows of different sizes to perform pooling operations on input images at different scales, and concatenates all pooled feature vectors to form a feature vector of a fixed size. The network model training module is used to train the small object image recognition network model based on the training set to obtain the trained small object image recognition network model; The network model testing module is used to test the small object image recognition network model based on the test set and output the small object image recognition accuracy; The predicted small object image recognition results are obtained based on the trained small object image recognition network model.

7. The small object image recognition system based on deep learning according to claim 6, characterized in that: The data enhancement module is used to perform data enhancement on small target image data, specifically including: Data augmentation operations are performed on each small target image, including random scaling, inversion, cropping, rotation, and optical transformation.

8. The small object image recognition system based on deep learning according to claim 6, characterized in that: The data preprocessing module is used to preprocess the small target image data, specifically including: The small target image is preprocessed by the mean filtering method. A template is given to the target pixel on the image, which includes the neighboring pixels around it, and then the original pixel value is replaced by the average value of all pixels in the template.

9. The small object image recognition system based on deep learning according to claim 6, characterized in that: The classification labels of the small target images include two labels: target image and non-target image. All targets in the training set are annotated with target boxes.

10. The small object image recognition system based on deep learning according to claim 6, characterized in that: The SPP module includes an input layer, multiple maximum pooling layers, a splicing layer and an output layer; The input layer inputs a feature map, each maximum pooling layer uses pooling windows of different sizes to perform pooling operations on the input feature map, and evenly divides the feature maps of different sizes into a number of grids of different sizes. The feature map in each grid is pooled, the splicing layer splices all the pooled feature vectors, and the output layer outputs the spliced ​​feature vector of a fixed size.