Method for improving small target detection precision and coordinate change speed

By adding the Coord-Conv convolution module to the YOLOv5 network model and combining the Shape-IOU method, the coordinate transformation of the convolution neural network is optimized, and the problem of insufficient speed and accuracy of small object detection and coordinate transformation is solved, and the efficient recognition effect of Apple recognition is achieved.

CN120299027APending Publication Date: 2025-07-11SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510349731.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing convolutional neural networks have problems with insufficient speed and accuracy in small object detection and coordinate transformation, especially in Apple identification applications, which are difficult to meet industrial needs.

Method used

改进YOLOv5网络模型,通过添加Coord-Conv卷积模块并结合Shape-IOU方法,优化卷积神经网络的坐标变换,提高小目标检测精度和速度。

Benefits of technology

It realizes high-precision and fast coordinate transformation of Apple's recognition network model, with an identification accuracy of more than 90%, and is suitable for actual production environments with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299027A_ABST
    Figure CN120299027A_ABST
Patent Text Reader

Abstract

The invention discloses a method for improving small target detection precision and coordinate change speed. The method comprises the following steps: acquiring apple image data and performing data enhancement; performing data preprocessing on the apple image data after data enhancement, and dividing a data set; constructing an apple image recognition network model, replacing a first Conv convolution module behind an SPP module in a YOLOv5 network model with a Coord-Conv convolution module, adding the Coord-Conv convolution module in front of three detection heads, and optimizing the network model in combination with Shape-IOU and IOU methods; training the apple image recognition network model based on the training set; testing the apple image recognition network model based on the test set, and outputting the accuracy of image recognition; and obtaining a predicted image recognition result based on the trained apple image recognition network model. According to the invention, the speed of coordinate transformation and the precision of small target detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a method for improving the detection accuracy of small targets and the coordinate change speed. Background Art

[0002] With the continuous improvement of people's living standards and the continuous development of machine vision technology, machine vision technology focusing on deep learning has attracted more and more attention and has been applied in all aspects of life. However, the convolution operation in the existing provided models has natural drawbacks. In standard convolution, when the convolution kernel slides on the input, it only focuses on the pixel intensity in the local area and ignores its absolute position, leaving a large room for optimization. By improving certain modules, the drawbacks of standard convolution can be overcome, the speed and accuracy of coordinate transformation can be improved, and the requirements of current industrial applications can be met. Summary of the Invention

[0003] In order to overcome the defects and deficiencies of the existing technology, the present invention provides a method for improving the detection accuracy of small targets and the coordinate change speed, which is applied to the pre-link in the apple picking industry for apple recognition. The present invention improves the framework of the YOLOv5 network model, adds Coord-Conv convolution and incorporates the Shape-IOU method, solves the defect of the convolutional neural network in coordinate transformation, improves the speed of coordinate transformation and the detection accuracy of small targets, and expands the application scenario of deep learning technology in apple recognition.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions:

[0005] The present invention provides a method for improving the detection accuracy of small targets and the coordinate change speed, including the following steps:

[0006] Obtain apple image data and perform data augmentation on the apple image data;

[0007] Perform data preprocessing on the apple image data after data augmentation;

[0008] Divide the preprocessed apple image data into a training set, a validation set, and a test set;

[0009] Construct an apple image recognition network model, replace the first Conv convolution module after the SPP module in the YOLOv5 network model with a Coord-Conv convolution module, and add a Coord-Conv convolution module before the three detection heads, and optimize the apple image recognition network model by combining the Shape-IOU method and the IOU method;

[0010] Train the apple image recognition network model based on the training set to obtain a trained apple image recognition network model;

[0011] Test the apple image recognition network model based on the test set and output the accuracy of image recognition.

[0012] Obtain the predicted image recognition result based on the trained apple image recognition network model.

[0013] As a preferred technical solution, perform data augmentation on the apple image data, specifically including:

[0014] Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0015] As a preferred technical solution, perform data preprocessing on the data-augmented apple image data, specifically including:

[0016] Perform noise reduction on the apple image data by means of mean filtering. Given a template for the target pixel on the image, the template includes its surrounding neighboring pixels, and replace the original pixel value with the average value of all pixels in the template.

[0017] As a preferred technical solution, the classification labels of the apple image data include two labels: target image and non-target image, and all targets in the training set are labeled with target boxes.

[0018] As a preferred technical solution, train the apple image recognition network model based on the training set. During the training process, use the cosine annealing algorithm to dynamically adjust the learning rate, adaptively adjust the learning rate of each parameter based on the exponential decay average of the squared gradient, and perform training based on the Adam optimizer.

[0019] The present invention also provides a system for improving the detection accuracy of small targets and the coordinate change speed, including: an image data acquisition module, a data augmentation module, a data preprocessing module, a data division module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module;

[0020] The image data acquisition module is used to acquire apple image data;

[0021] The data augmentation module is used to perform data augmentation on the apple image data;

[0022] The data preprocessing module is used to perform data preprocessing on the data-augmented apple image data;

[0023] The data division module is used to divide the preprocessed apple image data into a training set, a validation set, and a test set;

[0024] The image recognition network model construction module is used to construct an apple image recognition network model. The first Conv convolution module after the SPP module in the YOLOv5 network model is replaced with a Coord-Conv convolution module, and a Coord-Conv convolution module is added before the three detection heads. The apple image recognition network model is optimized by combining the Shape-IOU method and the IOU method;

[0025] The network model training module is used to train the apple image recognition network model based on the training set to obtain the trained apple image recognition network model;

[0026] The network model testing module is used to test the apple image recognition network model based on the test set and output the accuracy of image recognition;

[0027] The image recognition result output module is used to obtain the predicted image recognition result based on the trained apple image recognition network model.

[0028] As a preferred technical solution, the data augmentation module is used to perform data augmentation on the apple image data, specifically including:

[0029] Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0030] As a preferred technical solution, the data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation, specifically including:

[0031] Perform noise reduction processing on the apple image data by means of mean filtering. Given a template for the target pixel on the image, the template includes its surrounding adjacent pixels, and the average value of all pixels in the template is used to replace the original pixel value.

[0032] As a preferred technical solution, the classification labels of the apple image data include two types of labels: target images and non-target images. All targets in the training set are labeled with target boxes.

[0033] As a preferred technical solution, the network model training module is used to train the apple image recognition network model based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.

[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0035] (1) In the present invention, the first Conv convolution module after the SPP module in the YOLOv5 network model is replaced with a Coord-Conv convolution module, and a Coord-Conv convolution module is added before the three detection heads, which solves the defect of the convolutional neural network in coordinate transformation and improves the speed of coordinate transformation;

[0036] (2) The present invention combines the IOU method and the Shape-IOU method to optimize the apple image recognition network model, improving the speed and accuracy of small target detection;

[0037] (3) The apple image recognition network model of the present invention has a relatively high recognition accuracy, with an accuracy reaching more than 90%, strong applicability, can meet the requirements in actual production, and can be applied in scenarios with limited computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a schematic flowchart of the method for improving the accuracy of small target detection and the speed of coordinate change of the present invention;

[0039] Figure 2 is a schematic network structure diagram of the Coord-Conv convolution module of the present invention;

[0040] Figure 3 is a schematic network structure diagram of the apple image recognition network model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] Embodiment 1

[0043] As Figure 1 shown, this embodiment provides a method for improving the accuracy of small target detection and the speed of coordinate change, including the following steps:

[0044] S1: Obtain apple image data and perform data augmentation on the apple image data;

[0045] In this embodiment, for the apple image data, relevant apple image samples are collected through various search engines and databases, and attention should be paid to the balance of the apple image samples to avoid the situation where the number of samples of different categories varies greatly, or through various scenarios of self-shot images to ensure that the samples are diverse enough;

[0046] In this embodiment, data augmentation operations such as random scaling, inversion, cropping, rotation, and optical transformation are performed on each apple image to increase the number of the sample dataset, enabling the network model to be fully trained and avoiding the problem of the model having a tendency during the training process due to an excessive number of samples of certain categories;

[0047] S2: Preprocess the image by means of mean filtering to exclude the interference of noise;

[0048] In this embodiment, the collected apple images are denoised by using the mean filtering method. A template is given to the target pixel on the image. The template includes its surrounding adjacent pixels (8 pixels surrounding the target pixel form a filtering template, that is, including the target pixel itself), and then the average value of all pixels in the template is used to replace the original pixel value.

[0049] S3: Divide the apple image dataset into a training set, a validation set, and a test set;

[0050] In this embodiment, the apple image dataset is divided into a training set, a validation set, and a test set according to the ratio of 3:1:1. The classification labels of the apple images include two labels, namely target images and non-target images. All targets in the training set are labeled with target boxes;

[0051] S4: Construct an apple image recognition network model and adjust and optimize the structure of the YOLOv5 network model;

[0052] As Figure 3 shown, replace the first Conv convolution module after the SPP module in the YOLOv5 network model with a Coord-Conv convolution module, and add a Coord-Conv convolution module before the three detection heads to solve the defect of the convolutional neural network in coordinate transformation and improve the speed of coordinate transformation. In the figure, Input and Output respectively represent the input feature map and the output feature map, Conv represents the convolutional layer, Concat represents the splicing operation, and the C3 layer is a convolutional operation that combines multiple convolutional kernels, which can increase the receptive field of the model and reduce the number of parameters;

[0053] As Figure 2As shown, on the basis of the standard convolution, the Conv convolution module adds two-channel implementations to the input, one at the i coordinate and the other at the j coordinate. There are mainly two problems with the coordinate transformation task in the convolutional network: the transformation from the Cartesian space to the one-hot pixel space and other methods. When there is only one pixel, even if all training cases are easily available, the convolution still cannot be smoothly transformed; in addition, the best-performing convolutional model is huge in volume and takes a long time to train. The CoordConv convolution module solves the defect of the convolutional neural network in coordinate transformation and improves the speed of coordinate transformation by adding coordinate information to the input feature map, enabling the convolutional kernel to determine the exact position of each pixel.

[0054] In this embodiment, the Shape-IOU method is added and used together with the originally used IOU method. The apple image recognition network model is optimized by combining the Shape-IOU method and the IOU method, improving the speed and accuracy of small target detection.

[0055] S5: Set the training parameters, use the training set and the validation set to train and tune the convolutional neural network model, and obtain the network model with the best target recognition effect.

[0056] In this embodiment, training parameters such as epoch, batch size, and learning rate are set. The Adam optimizer is used. After a large number of trainings and debuggings, the network model with the best recognition effect is obtained.

[0057] Specifically, set epoch to 50, batch size to 16, and learning rate to 0.01, and use the cosine annealing algorithm to dynamically adjust the learning rate to avoid the oscillation phenomenon caused by too fast gradient descent during training, thereby improving the training stability and generalization ability of the model. Since the Adam optimizer contains the concept of momentum, it accumulates the exponentially decaying average of the previous gradients to help accelerate learning. At the same time, it also uses the exponentially decaying average of the squared gradients to adaptively adjust the learning rate of each parameter, and has strong robustness and is widely used in deep learning tasks. Therefore, the Adam optimizer is used for training.

[0058] S6: Call the network model to perform recognition tests on the test set, use the recognition accuracy as the model evaluation criterion to verify the model performance. By comparing the recognized apple results with the marked true positions, the recognition ability of the recognition method for apples in the image can be detected, and the accuracy of target recognition is output, thereby improving the small target detection accuracy and the coordinate change speed.

[0059] The present invention can solve the defects of convolutional neural networks in coordinate transformation, improve the speed of coordinate transformation, improve the speed and accuracy of small target detection, effectively expand the actual application scenarios, and enable this method to identify various targets.

[0060] Embodiment 2

[0061] This embodiment provides a system for improving the accuracy of small target detection and the speed of coordinate change, which is used to implement the method for improving the accuracy of small target detection and the speed of coordinate change in the above Embodiment 1. The system includes: an image data acquisition module, a data enhancement module, a data preprocessing module, a data division module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module;

[0062] In this embodiment, the image data acquisition module is used to acquire apple image data;

[0063] In this embodiment, the data enhancement module is used to perform data enhancement on the apple image data;

[0064] In this embodiment, the data preprocessing module is used to perform data preprocessing on the apple image data after data enhancement;

[0065] In this embodiment, the data division module is used to divide the preprocessed apple image data into a training set, a validation set, and a test set;

[0066] In this embodiment, the image recognition network model construction module is used to construct an apple image recognition network model, replace the first Conv convolution module after the SPP module in the YOLOv5 network model with a Coord-Conv convolution module, and add a Coord-Conv convolution module before the three detection heads, and optimize the apple image recognition network model by combining the Shape-IOU method and the IOU method;

[0067] In this embodiment, the network model training module is used to train the apple image recognition network model based on the training set to obtain a trained apple image recognition network model;

[0068] In this embodiment, the network model testing module is used to test the apple image recognition network model based on the test set and output the accuracy of image recognition;

[0069] In this embodiment, the image recognition result output module is used to obtain the predicted image recognition result based on the trained apple image recognition network model.

[0070] In this embodiment, the data enhancement module is used to perform data enhancement on the apple image data, specifically including:

[0071] Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

[0072] In this embodiment, the data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation, specifically including:

[0073] Perform noise reduction processing on the apple image data by means of mean filtering. Give a template to the target pixel on the image. This template includes its surrounding adjacent pixels, and replace the original pixel value with the average value of all pixels in the template.

[0074] In this embodiment, the classification labels of the apple image data include two types of labels: target images and non-target images. All targets in the training set are labeled with target boxes.

[0075] In this embodiment, the network model training module is used to train the apple image recognition network model based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.

[0076] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for improving the detection accuracy and coordinate change speed of small targets, characterized in that, Including the following steps: Obtain apple image data and perform data augmentation on the apple image data; Perform data preprocessing on the apple image data after data augmentation; Divide the preprocessed apple image data into a training set, a validation set, and a test set; Construct an apple image recognition network model, replace the first Conv convolutional module after the SPP module in the YOLOv5 network model with a Coord-Conv convolutional module, and add a Coord-Conv convolutional module before the three detection heads, and optimize the apple image recognition network model by combining the Shape-IOU method and the IOU method; Train the apple image recognition network model based on the training set to obtain the trained apple image recognition network model; Test the apple image recognition network model based on the test set and output the accuracy of image recognition; Obtain the predicted image recognition result based on the trained apple image recognition network model.

2. The method for improving the detection accuracy of small targets and the coordinate change speed according to claim 1, characterized in that, Perform data augmentation on the apple image data, specifically including: Perform data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

3. The method for improving the detection accuracy of small targets and the coordinate change speed according to claim 1, characterized in that Perform data preprocessing on the apple image data after data augmentation, specifically including: Perform noise reduction processing on the apple image data by means of mean filtering. Give a template to the target pixels on the image. This template includes its surrounding neighboring pixels, and replace the original pixel value with the average value of all pixels in the template.

4. The method for improving the detection accuracy of small targets and the coordinate change speed according to claim 1, characterized in that, The classification labels of the apple image data include two labels: target image and non-target image. All targets in the training set are labeled with target boxes.

5. The method for improving the detection accuracy of small targets and the coordinate change speed according to claim 1, characterized in that, Train the apple image recognition network model based on the training set. During the training process, use the cosine annealing algorithm to dynamically adjust the learning rate, adaptively adjust the learning rate of each parameter based on the exponential decay average of the squared gradient, and perform training based on the Adam optimizer.

6. A system for improving the detection accuracy and coordinate change speed of small targets, characterized in that, Including: An image data acquisition module, a data augmentation module, a data preprocessing module, a data division module, an image recognition network model construction module, a network model training module, a network model testing module, and an image recognition result output module; The image data acquisition module is used to obtain apple image data; The data augmentation module is used to perform data augmentation on the apple image data; The data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation; The data division module is used to divide the preprocessed apple image data into a training set, a validation set, and a test set; The image recognition network model construction module is used to construct an apple image recognition network model, replace the first Conv convolutional module after the SPP module in the YOLOv5 network model with a Coord-Conv convolutional module, and add a Coord-Conv convolutional module before the three detection heads, and optimize the apple image recognition network model by combining the Shape-IOU method and the IOU method; The network model training module is used to train the apple image recognition network model based on the training set to obtain the trained apple image recognition network model; The network model testing module is used to test the apple image recognition network model based on a test set and output the accuracy rate of image recognition. The image recognition result output module is used to obtain the predicted image recognition result based on the trained apple image recognition network model.

7. The system for improving the detection accuracy of small targets and the coordinate change speed according to claim 6, characterized in that, The data augmentation module is used to perform data augmentation on the apple image data, specifically including: Performing data augmentation operations on each apple image, including random scaling, inversion, cropping, rotation, and optical transformation.

8. The system for improving the detection accuracy of small targets and the coordinate change speed according to claim 6, characterized in that, The data preprocessing module is used to perform data preprocessing on the apple image data after data augmentation, specifically including: Performing noise reduction processing on the apple image data through mean filtering. A template is given to the target pixel on the image, and the template includes its surrounding adjacent pixels. The average value of all pixels in the template is used to replace the original pixel value.

9. The system for improving the detection accuracy of small targets and the coordinate change speed according to claim 6, characterized in that, The classification labels of the apple image data include two types of labels: target images and non-target images. All targets in the training set are labeled with target boxes.

10. The system for improving the detection accuracy of small targets and the coordinate change speed according to claim 6, characterized in that, The network model training module is used to train the apple image recognition network model based on the training set. During the training process, the cosine annealing algorithm is used to dynamically adjust the learning rate, the learning rate of each parameter is adaptively adjusted based on the exponential decay average of the squared gradient, and the training is performed based on the Adam optimizer.