Strawberry maturity detection method based on improved YOLOv8
By improving the YOLOv8 model, including adding small object detection head branches, replacing the Bottleneck module and adding Shuffle Attention attention mechanism, the problem of insufficient detection accuracy and high computational complexity in strawberry ripening detection is solved, and more efficient and accurate strawberry ripening detection is achieved.
Patent Information
- Application Number
- CN202510216631.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-17
AI Technical Summary
The existing YOLOv8 model has problems such as insufficient detection accuracy and high computational complexity in the strawberry maturity detection task, making it difficult to realize real-time application in resource-constrained agricultural scenarios.
Improvements to the YOLOv8 model include adding small object detection head branches to the Neck network, replacing the Bottleneck module with PCGGC residual module, adding Shuffle Attention attention mechanism, and performing image enhancement and preprocessing operations on the dataset.
It improves the efficiency and accuracy of strawberry ripening, enhances the model's detection ability of small targets, reduces the computational complexity, and enables the model to be applied in real-time in agricultural scenarios.
Smart Images

Figure CN120164210A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to machine vision and image processing, and specifically to a strawberry maturity detection method based on an improved YOLOv8 model. Background Art
[0002] In the fields of agricultural production and agricultural product quality inspection, the accurate detection of strawberry maturity is of great significance for strawberry planting management, picking timing, and subsequent sorting and sales. Traditional strawberry maturity detection methods mainly rely on manual visual judgment, which is not only inefficient but also the detection results are easily affected by the subjective factors of the inspectors, making it difficult to ensure accuracy and consistency.
[0003] With the development of machine vision technology, object detection methods based on deep learning have provided a new way for strawberry maturity detection. Among them, the YOLO (You Only Look Once) series of models stand out with their unique advantages such as fast detection speed and relatively high accuracy, and have been widely used in the field of object detection. As a relatively new member of this series, the YOLOv8 model can quickly locate and identify objects in images in a short time, providing the possibility for real-time detection.
[0004] However, when the existing YOLOv8 model is used to process the strawberry maturity detection task, there are still some deficiencies. For strawberries, which have significant changes in color, shape, size, etc. at different maturity stages and have many small targets, the YOLOv8 model may not be accurate enough in feature extraction and small target detection, resulting in the need to improve the detection accuracy; at the same time, the computational complexity of the model is relatively high, which is not conducive to real-time application in some resource-constrained agricultural scenarios.
[0005] Therefore, in order to make the YOLOv8 model better adapt to the requirements of the strawberry maturity detection task, it is necessary to improve it to achieve accurate detection of strawberry maturity and promote the intelligent development of the strawberry industry. Summary of the Invention
[0006] In view of the above-listed technical problems, the present invention provides a strawberry maturity detection method based on an improved YOLOv8. This method can improve the efficiency and accuracy of strawberry maturity detection, and achieve rapid and accurate evaluation of strawberry maturity to solve the current technical deficiencies.
[0007] To achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:
[0008] A strawberry maturity detection method based on an improved YOLOv8, the steps of which include:
[0009] Step 1: Collect fruit images of strawberry orchards as dataset samples for strawberry maturity, which cover strawberry images in different maturity states, including strawberry images in the immature state, semi-ripe state, and fully ripe state;
[0010] Step 2: Perform image enhancement and preprocessing operations on the obtained strawberry images to expand the samples and form a dataset, and divide the above strawberry maturity dataset into a training set and a validation set;
[0011] Step 3: According to the task requirements of detecting strawberry maturity, improve the network structure of YOLOv8, and the improved YOLOv8 model is used as the strawberry maturity detection model;
[0012] Step 4: Input the above training set into the strawberry maturity detection model for training to obtain a trained strawberry maturity detection model;
[0013] Step 5: Use the trained strawberry maturity detection model to detect strawberries to obtain strawberry maturity detection results.
[0014] As a further optimization of the technical solution of the present invention, step 2 specifically includes:
[0015] 2.1. Expand the strawberry dataset images, including horizontal flipping and vertical flipping, and then rotate the flipped images. The rotation directions and angles are 90 degrees, 180 degrees, and 270 degrees counterclockwise.
[0016] 2.2. Add Gaussian noise to the strawberry dataset images after the expansion process,
[0017] The expression for adding Gaussian noise to the image is:
[0018] I new (x,y) = I(x,y) + n(x,y)
[0019] where I new (x,y) is the pixel value after adding Gaussian noise, I(x,y) is the pixel value of the original image, n(x,y) is the noise value obeying the Gaussian distribution, x and y represent the coordinates of the image pixels, x represents the coordinate position in the horizontal direction, and y represents the coordinate position in the vertical direction; its generation method is randomly sampled from the Gaussian distribution probability density function;
[0020] The probability density function is:
[0021]
[0022] where z is a random variable, which is the noise value of the pixel point in the image data, μ is the mean value, and σ is the standard deviation.
[0023] 2.3. The expanded strawberry maturity dataset is divided into a training set and a validation set according to an appropriate ratio;
[0024] 2.4. Use Labelimg to label different categories (immature state, semi-ripe state, fully ripe state) in the images, frame the corresponding strawberry areas on the images according to the maturity state of the strawberries, and correspond to the corresponding category labels. After completing the annotation of each image, generate a txt file in YOLO format.
[0025] As a further optimization of the technical solution of the present invention, step 3 includes:
[0026] 3.1. In the Neck network of the YOLOv8 model, add a small object detection head branch;
[0027] 3.2. Construct a residual module PCGGC, and replace the Bottleneck module of its C2f structure with a PCGGC module in the Backbone network and the Neck network of the YOLOv8 model;
[0028] 3.3. In the Backbone network and the Neck network of the YOLOv8 model, add a Shuffle Attention mechanism behind the Conv2 module in their C2f structures.
[0029] As a further optimization of the technical solution of the present invention, the structural processing of the small object detection head branch in step 3 includes:
[0030] Add an Upsample-Concat-C2f structure, that is, an upsampling layer - concatenation layer - feature extraction and fusion module structure, to the Neck network of the original YOLOv8 model for outputting a feature map with a size of 160×160 pixels for the small object branch;
[0031] Among them, the upsampling layer is connected to the feature extraction and fusion module of the fifteenth layer of the Neck network, and the concatenation layer is connected to the feature extraction and fusion module of the second layer of the Backbone network;
[0032] Add a CBS-Concat-C2f structure, that is, a convolutional module - concatenation layer - feature extraction and fusion module, behind the feature extraction and fusion module of the above Upsample-Concat-C2f structure, and convert a feature map with a width and height of 160 pixels and a size of 160×160 pixels into a feature map with a width and height of 80 pixels and a size of 80×80 for output;
[0033] Among them, the convolution module is connected to the feature extraction and fusion module of the Upsample-Concat-C2f structure.
[0034] As a further optimization of the technical solution of the present invention, the PCGGC residual module structure and principle in step 3 include:
[0035] 3.2.1. The PCGGC residual module sequentially includes a partial convolution layer, a point convolution layer, a GN group normalization layer, a GELU activation function layer, and a point convolution layer. Among them, there is a residual connection after the second point convolution layer, and the original input and the input after passing through the partial convolution layer will be directly connected here;
[0036] 3.2.2. The input feature A enters the 3×3 partial convolution layer for preliminary transformation and extraction to obtain the output feature B;
[0037] 3.2.3. The output feature B enters the 1×1 point convolution layer to adjust the number of channels of feature B and integrate cross-channel information, and perform a linear combination of different channel information without changing the spatial size of the feature map. After processing, the output feature C is obtained;
[0038] 3.2.4. The output feature C enters the GN group normalization layer and the GELU activation function layer. The GN layer normalizes the feature C, and then the GELU layer introduces non-linear factors to obtain the output feature D;
[0039] 3.2.5. The output feature D enters the next 1×1 point convolution layer to further transform and extract the feature, adjust the number of channels again and integrate the feature information. After processing, the output feature E1 is obtained;
[0040] 3.2.6. Finally, the output feature E1 is added to the original input feature A of the module and the output feature B after passing through the 3×3 partial convolution layer through the residual connection E1 + A + B = E2 to obtain the final output feature E2.
[0041] As a further optimization of the technical solution of the present invention, the working principle of the Shuffle Attention in step 3 includes:
[0042] 3.3.1. The input feature map X with a size of h×w×c, where h represents the height, w represents the width, and c represents the number of channels;
[0043] 3.3.2. Perform a grouping operation on the input feature map, divide it into g groups on average, and the number of channels of each group is to obtain g grouped feature maps X1, X2,..., X g ;
[0044] 3.3.3. For each group, further split it into two parts, namely the upper and lower branches, and the number of channels for each branch is The upper branch feature map of the output is Y i (i = 1, 2, …, g) is used to implement channel attention, and the lower branch feature map is Z i (i = 1, 2, …, g) is used to implement spatial attention;
[0045] 3.3.4. Perform global average pooling operation on the upper branch to aggregate spatial information, map it to a new space through a fully connected layer, and obtain a feature vector O of, apply the Sigmoid activation function to vector O to obtain the channel attention weight A c , and reassign A c to Y i (i = 1, 2, …, g) for element-wise multiplication operation to obtain the weighted grouped feature map Y ci (i = 1, 2, …, g);
[0046] 3.3.5. Perform group normalization operation on the lower branch to calculate the spatial attention of each group of features. Similarly, through a fully connected layer, obtain the spatial attention weight A after passing through the Sigmoid activation function s , and reassign A s to Z i (i = 1, 2, …, g) for element-wise multiplication operation to obtain the weighted grouped feature map Z si (i = 1, 2, …, g);
[0047] The processing formula for the fully connected layer of both parts is:
[0048] F c (x) = Wx + b
[0049] where, F c (x) represents the output result after being processed by the fully connected layer, W is the weight matrix of the fully connected layer, x is the feature vector input to the fully connected layer, and b is the bias vector;
[0050] 3.3.6. Concatenate Y ci (i = 1, 2, …, g) and Z si (i = 1, 2, …, g) to form g groups of new feature maps X1 ′ , X2 ′ , …, X g ′ , and then aggregate to obtain the output feature map X ′ , whose dimension is the same as that of the input feature map X, still h × w × c;
[0051] 3.3.7. Before output, perform channel rearrangement on the feature map X ′ to obtain the final output feature map K.
[0052] As a further optimization of the technical solution of the present invention, step 4 specifically includes:
[0053] 4.1. Use the above training set as the input of the strawberry maturity detection model, and set training parameters such as the number of iterations, the number of input samples per iteration, and the input size;
[0054] 4.2. The input data passes through the Backbone network, and after a series of CBS modules and C2f-PS modules (improved C2f) and the final SPPF fast spatial pyramid pooling module, the initial feature extraction of the image is completed, and feature maps of different scales are output as the input of the Neck network;
[0055] 4.3. The feature map output by the Backbone network enters the Neck network, and through upsampling and splicing operations, the feature information of different scales is fused as the input of the Head network;
[0056] 4.4. The fused feature map output by the Neck network enters the Head network, and after passing through the CBS module and the Conv2d two-dimensional convolution operation, the Bbox Loss bounding box loss and the Cls Loss classification loss are output, and finally the detection box position information of the strawberry target and the corresponding maturity classification result, that is, the immature state, the semi-ripe state, and the fully ripe state, are obtained.
[0057] The present invention has the following beneficial effects:
[0058] (1) By performing horizontal and vertical flipping and counterclockwise rotation operations on the strawberry dataset images at different angles, the number of samples in the dataset is expanded, so that the model is no longer limited to learning the characteristics of strawberries from a specific perspective, but can more comprehensively and deeply understand the various morphological characteristics of strawberries, adapt to the appearance changes of strawberries from different perspectives, and help improve the generalization ability of the model.
[0059] (2) Gaussian noise is added to the strawberry dataset images after the expansion process to simulate the possible noise interference situations in the image acquisition process, such as uneven illumination and noise generated by the limitations of the sensor itself, which may easily mask or distort some details. Reducing the degradation of detection performance caused by noise can improve the adaptability and robustness of the model under different environmental conditions.
[0060] (3) Small targets occupy fewer pixels in the image and have less obvious features. By adding a small target detection head branch, the model further enhances its ability to detect small targets on the basis of its original ability to detect medium and large targets. This helps the model detect smaller strawberries in the image more accurately, capture their feature information better, improve the multi-scale target detection ability, and enhance the detection accuracy of the maturity of small strawberries.
[0061] (4) Replace the Bottleneck module of the C2f structure in the Backbone backbone network and Neck neck network of the YOLOv8 model with the PCGGC residual module. This module can improve the running speed and inference efficiency, reduce redundant calculations to a certain extent. The GELU activation function is smoother than the traditional ReLU activation function, and the GN group normalization is insensitive to the batch size and can better normalize features during small-batch training, while maintaining the accuracy on this basis.
[0062] (5) Add the Shuffle Attention attention mechanism after the Conv2 module in the C2f structure of the YOLOv8 model, which can promote the interaction and fusion of information in different channels and spatial positions. From the channel dimension, through specific calculation methods, the potential connections between channels can be mined, enabling the model to focus on the important features carried by different channels. From the spatial position dimension, this mechanism guides the model to focus on information in different regions, and whether it is the edge, top, or bottom of the strawberry, it can be fully considered in the model. This improvement can increase the model's attention to key features, enable the model to use information more effectively in the feature fusion stage, and improve the accuracy of judging the maturity of strawberries. Description of the Drawings
[0063] Figure 1 It is a flowchart of the improved YOLOv8 strawberry maturity detection method of the present invention.
[0064] Figure 2 It is a schematic diagram of the network structure of the strawberry maturity detection model of the present invention.
[0065] Figure 3 It is a schematic diagram of the small target detection head branch structure of the present invention.
[0066] Figure 4 It is a schematic diagram of the PCGGC residual module structure of the present invention.
[0067] Figure 5 It is a flowchart of the Shuffle Attention attention mechanism of the present invention.
[0068] Figure 6 It is a schematic diagram of the improved C2f (C2f-PS) structure of the present invention.
[0069] Figure 7 Schematic diagram of the detection effect of the strawberry maturity detection model of the present invention. Specific implementation manner
[0070] A strawberry maturity detection method based on improved YOLOv8, the steps of which include:
[0071] Step 1: Collect fruit images of strawberry orchards as dataset samples of strawberry maturity, which cover strawberry images in different maturity states, including strawberry images in the immature state, strawberry images in the semi-ripe state, and strawberry images in the fully ripe state.
[0072] Step 2: Perform image enhancement and preprocessing operations on the obtained strawberry images to expand the samples and form a dataset, and divide the above strawberry maturity dataset into a training set and a validation set;
[0073] Step 2 specifically includes:
[0074] 2.1. Expand the strawberry dataset images, including horizontal flipping and vertical flipping, and then rotate the flipped images. The rotation directions and angles are 90 degrees, 180 degrees, and 270 degrees counterclockwise;
[0075] 2.2. Add Gaussian noise to the strawberry dataset images after the expansion process,
[0076] The expression for adding Gaussian noise to the image is:
[0077] I new (x,y) = I(x,y) + n(x,y)
[0078] where, I new (x,y) is the pixel value after adding Gaussian noise, I(x,y) is the pixel value of the original image, n(x,y) is the noise value subject to Gaussian distribution, x and y represent the coordinates of the image pixels, x represents the coordinate position in the horizontal direction, and y represents the coordinate position in the vertical direction. Its generation method is obtained by randomly sampling from the Gaussian distribution probability density function;
[0079] The probability density function is:
[0080]
[0081] where, z is a random variable, which is the noise value of the pixel point in the image data, μ is the mean value, and σ is the standard deviation.
[0082] 2.3. Divide the strawberry maturity dataset after adding Gaussian noise into a training set and a validation set according to an appropriate ratio;
[0083] 2.4. Use Labelimg to label different categories (immature state, semi-ripe state, fully ripe state) in the images, frame the corresponding strawberry areas on the images according to the maturity state of the strawberries, and correspond to the corresponding category labels; after completing the annotation of each image, generate a txt file in YOLO format.
[0084] Step 3: For the task requirement of detecting strawberry maturity, improve the YOLOv8 model, and the improved YOLOv8 model is used as the strawberry maturity detection model;
[0085] Step 3 includes:
[0086] 3.1. In the Neck network of the YOLOv8 model, add a small target detection head branch; the specific operation is as follows:
[0087] Add an Upsample-Concat-C2f structure, that is, an upsampling layer - concatenation layer - feature extraction and fusion module structure, in the Neck network of the YOLOv8 model for the output of the feature map with a size of 160×160 for the small target branch;
[0088] Among them, the upsampling layer is connected to the feature extraction and fusion module of the fifteenth layer of the Neck network, and the concatenation layer is connected to the feature extraction and fusion module of the second layer of the Backbone network;
[0089] Add a CBS-Concat-C2f structure, that is, a convolutional module - concatenation layer - feature extraction and fusion module, after the feature extraction and fusion module of the above Upsample-Concat-C2f structure, and convert the feature map with a width and height of 160 pixels and a size of 160×160 pixels into a feature map with a width and height of 80 pixels and a size of 80×80 for output;
[0090] Among them, the convolutional module is connected to the feature extraction and fusion module of the Upsample-Concat-C2f structure.
[0091] 3.2. Construct a residual module PCGGC, and replace the Bottleneck module of its C2f structure with the PCGGC module in the Backbone network and Neck network of the YOLOv8 model;
[0092] The structure and working process of the PCGGC residual module include:
[0093] 3.2.1. The PCGGC residual module sequentially includes a partial convolution layer, a point convolution layer, a GN group normalization layer, a GELU activation function layer, a point convolution layer, and a residual connection block in order; the original input, i.e., input feature A, and the output after the partial convolution layer, i.e., output feature B, are directly connected to the residual connection block;
[0094] 3.2.2. Input feature A enters the 3×3 partial convolution layer for preliminary transformation and extraction to obtain output feature B;
[0095] 3.2.3. Output feature B enters the 1×1 point convolution layer to adjust the number of channels of output feature B and integrate cross-channel information, and perform a linear combination of different channel information without changing the spatial size of the feature map. After processing, output feature C is obtained;
[0096] 3.2.4. Output feature C enters the GN group normalization layer and the GELU activation function layer. The GN group normalization layer normalizes output feature C, and then the GELU activation function layer introduces non-linear factors to obtain output feature D;
[0097] 3.2.5. Output feature D enters the next 1×1 point convolution layer to further transform and extract features, adjust the number of channels again, and integrate feature information. After processing, output feature E1 is obtained;
[0098] 3.2.6. Finally, output feature E1 is added to the original input feature A and output feature B after the 3×3 partial convolution layer through the residual connection block for the operation E1 + A + B = E2 to obtain the final output feature E2.
[0099] 3.3. In the Backbone backbone network and Neck neck network of the YOLOv8 model, a Shuffle Attention attention mechanism is added behind the last Conv2 module in the C2f structure;
[0100] The working process of the Shuffle Attention attention mechanism includes:
[0101] 3.3.1. The input feature map X with a size of h×w×c, where h represents the height, w represents the width, and c represents the number of channels;
[0102] 3.3.2. The input feature map is grouped, and it is evenly divided into g groups, and the number of channels in each group is to obtain g grouped feature maps X1, X2,..., X g ;
[0103] 3.3.3. For each group, it is further split into upper and lower branches, and the number of channels in each branch is The upper-branch output feature map is Y i (i = 1, 2, …, g) is used to implement channel attention, and the lower-branch feature map is Z i (i = 1, 2, …, g) is used to implement spatial attention;
[0104] 3.3.4. Perform global average pooling operation on the upper branch to aggregate spatial information, map it to a new space through a fully connected layer, and obtain a feature vector O of, apply the Sigmoid activation function to the feature vector O to obtain the channel attention weight A c , and reallocate the channel attention A c to Y i (i = 1, 2, …, g) for element-wise multiplication operation to obtain the weighted grouped feature map Y ci (i = 1, 2, …, g);
[0105] 3.3.5. Perform group normalization operation on the lower branch, calculate the spatial attention of each group of features, pass through a fully connected layer, and obtain the spatial attention weight A after passing through the Sigmoid activation function s , and reallocate A s to Z i (i = 1, 2, …, g) for element-wise multiplication operation to obtain the weighted grouped feature map Z si (i = 1, 2, …, g);
[0106] The processing formula for the fully connected layers of the upper and lower branches is:
[0107] F c (x) = Wx + b
[0108] where F c (x) represents the output result after being processed by the fully connected layer, W is the weight matrix of the fully connected layer, x is the feature vector input to the fully connected layer, and b is the bias vector;
[0109] 3.3.6. Concatenate Y ci (i = 1, 2, …, g) and Z si (i = 1, 2, …, g) to form g groups of new feature maps X1 ′ , X2 ′ , …, X g ′ , and then aggregate to obtain the output feature map X ′ , whose dimension is the same as that of the input feature map X, still h × w × c;
[0110] 3.3.7. Perform channel rearrangement on the feature map X ′ to obtain the final output feature map K.
[0111] Step 4: Input the above training set into the strawberry maturity detection model for training to obtain a trained strawberry maturity detection model;
[0112] Step 4 specifically includes:
[0113] 4.1. Use the training set as the input of the strawberry maturity detection model, and set training parameters such as the number of iterations, the number of input samples in one iteration, and the input size;
[0114] 4.2. The input data passes through the Backbone network, and after a series of CBS modules, the improved C2f structure, and the final SPPF (Spatial Pyramid Pooling Fast) module, the initial feature extraction of the image is completed, and feature maps of different scales are output as the input of the Neck network;
[0115] 4.3. The feature maps output by the Backbone network enter the Neck network, and through upsampling and splicing operations, the feature information of different scales is fused as the input of the Head network;
[0116] 4.4. The fused feature maps output by the Neck network enter the Head network. After passing through the CBS module and Conv2d (2D convolution) operation, the Bbox Loss (bounding box loss) and Cls Loss (classification loss) are output, and the position information of the detection box of the strawberry target and the corresponding maturity classification results, namely the immature state, semi-ripe state, and fully ripe state, are obtained.
[0117] Step 5: Use the trained strawberry maturity detection model to detect strawberries to obtain the strawberry maturity detection results.
Claims
1. A strawberry maturity detection method based on improved YOLOv8, characterized in that: The steps include: Step 1: Collect fruit images from a strawberry orchard as a dataset sample of strawberry maturity, which includes strawberry images of different maturity states, including unripe strawberry images, half-ripe strawberry images, and fully ripe strawberry images; Step 2: Perform image enhancement and preprocessing operations on the acquired strawberry image to expand the sample and form a data set, and divide the above strawberry maturity data set into a training set and a verification set; Step 3: According to the task requirements of detecting strawberry maturity, the YOLOv8 model is improved, and the improved YOLOv8 model is used as the strawberry maturity detection model; Step 4: Input the above training set into the strawberry maturity detection model for training to obtain a trained strawberry maturity detection model; Step 5: Use the trained strawberry maturity detection model to detect strawberries and obtain the strawberry maturity detection results.
2. The strawberry maturity detection method according to claim 1, wherein: The step 2 specifically includes: 2.
1. Expand the strawberry dataset images, including horizontal flipping and vertical flipping. Rotate the flipped images again. The rotation direction and angle are 90, 180, and 270 degrees counterclockwise. 2.
2. Add Gaussian noise to the expanded strawberry dataset images. The expression for adding Gaussian noise to an image is: I new (x,y)=I(x,y)+n(x,y) (1) Among them, I new (x, y) is the pixel value after adding Gaussian noise, I(x, y) is the pixel value of the original image, n(x, y) is the noise value that obeys Gaussian distribution, x and y represent the coordinates of the image pixel, x represents the horizontal coordinate position, and y represents the vertical coordinate position; it is generated by random sampling from the Gaussian distribution probability density function; The probability density function is: Among them, z is a random variable, which is the noise value of the pixel in the image data, μ is the mean, and σ is the standard deviation; 2.
3. The strawberry maturity dataset after adding Gaussian noise is divided into a training set and a validation set according to an appropriate ratio; 2.
4. Use Labelimg to label the different categories in the image, namely unripe, half-ripe, and fully ripe. Select the corresponding strawberry area on the image according to the maturity status of the strawberry and correspond to the corresponding category label. After completing the labeling of each image, generate a txt file in YOLO format.
3. The strawberry maturity detection method according to claim 1, wherein: The step 3 comprises: 3.
1. Add a small target detection head branch to the Neck network of the YOLOv8 model; 3.
2. Construct the residual module PCGGC. In the Backbone network and Neck network of the YOLOv8 model, replace the Bottleneck module of the C2f structure with the PCGGC module. 3.
3. In the Backbone network and Neck network of the YOLOv8 model, add the Shuffle Attention mechanism after the last Conv2 module in the C2f structure.
4. The strawberry maturity detection method according to claim 3, wherein: Adding a small target detection head branch in step 3.1 includes: Add an Upsample-Concat-C2f structure in the Neck network of the YOLOv8 model, i.e., an upsampling layer-concatenation layer-feature extraction fusion module structure, for outputting feature maps of small target branches with a size of 160×160 pixels; Among them, the upsampling layer is connected to the feature extraction and fusion module of the fifteenth layer of the Neck network, and the splicing layer is connected to the feature extraction and fusion module of the second layer of the Backbone network; A CBS-Concat-C2f structure is added after the feature extraction and fusion module of the Upsample-Concat-C2f structure, i.e., convolution module-concatenation layer-feature extraction and fusion module, to convert the feature map with a width and height of 160 pixels and a size of 160×160 pixels into a feature map with a width and height of 80 pixels and a size of 80×80 for output; Among them, the convolution module is connected to the feature extraction fusion module of the Upsample-Concat-C2f structure.
5. The strawberry maturity detection method according to claim 3, wherein: The PCGGC residual module structure and working process in step 3.2 include: 3.2.1, PCGGC residual module includes a partial convolution layer, a point convolution layer, a GN group normalization layer, a GELU activation function layer, a point convolution layer, and a residual connection block in sequence; the original input, i.e., input feature A, and the output after the partial convolution layer, i.e., output feature B, are directly connected to the residual connection block; 3.2.2, Input feature A enters the 3×3 partial convolution layer for preliminary transformation and extraction to obtain output feature B; 3.2.3, the output feature B enters the 1×1 point convolution layer, adjusts the number of channels of the output feature B and integrates the cross-channel information, and linearly combines the information of different channels without changing the spatial size of the feature map, and obtains the output feature C after processing; 3.2.4, the output feature C enters the GN group normalization layer and the GELU activation function layer. The GN group normalization layer normalizes the output feature C, and then the GELU activation function layer introduces nonlinear factors to obtain the output feature D; 3.2.
5. The output feature D enters the next 1×1 point convolution layer to further transform and extract the features, adjust the number of channels again and integrate the feature information, and obtain the output feature E1 after processing; 3.2.
6. Finally, the output feature E1 is added to the original input feature A and the output feature B after the 3×3 partial convolution layer through the residual connection block, and the final output feature E2 is obtained.
6. The strawberry maturity detection method according to claim 3, characterized in that: The working process of the ShuffleAttention mechanism in step 3.3 includes: 3.3.
1. Input feature map X of size h×w×c, where h represents height, w represents width, and c represents the number of channels. 3.3.
2. Group the input feature map and divide it into g groups on average. The number of channels in each group is Get g group feature maps X1, X2, ..., X g ; 3.3.
3. For each group, further split it into two branches, the number of channels in each branch is The output upper branch feature map is Y i (i=1,2,…,g) is used to realize channel attention, and the feature map of the lower branch is Z i (i=1,2,…,g) is used to implement spatial attention; 3.3.
4. Perform a global average pooling operation on the upper branch to aggregate the spatial information and map it to the new space through a fully connected layer to obtain a The feature vector O is obtained by applying the Sigmoid activation function to the feature vector O to obtain the channel attention weight A c , the channel attention A c Reassign to Y i (i=1,2,…,g) to obtain the weighted group feature map Y ci (i=1,2,…,g); 3.3.
5. Perform group normalization on the lower branch, calculate the spatial attention of each group of features, pass through a fully connected layer, and obtain the spatial attention weight A through the Sigmoid activation function s , A s Reassign to Z i (i=1,2,…,g) to obtain the weighted group feature map Z si (i=1,2,…,g); The fully connected layer processing formulas of the upper and lower branches are: F c (x)=Wx+b (3) Among them, F c (x) represents the output result after being processed by the fully connected layer, W is the weight matrix of the fully connected layer, x is the feature vector input to the fully connected layer, and b is the bias vector; 3.3.
6. Y ci (i=1,2,…,g) and Z si (i=1,2,…,g) are concatenated to form g groups of new feature maps X1 ′ , X2 ′ , …, X g ′ , and then aggregate to get the output feature map X ′ , its dimension is the same as the input feature map X, still h×w×c; 3.3.
7. Feature map X ′ Channels are rearranged to obtain the final output feature map K.
7. The strawberry maturity detection method according to claim 3, characterized in that: The step 4 specifically includes: 4.
1. Use the training set as the input of the strawberry maturity detection model, set the number of iterations, the number of input samples per iteration, and the input size training parameters; 4.
2. The input data passes through the Backbone network, a series of CBS modules, an improved C2f structure and the final SPPF fast spatial pyramid pooling module to complete the preliminary feature extraction of the image and output feature maps of different scales as the input of the Neck network; 4.
3. The feature map output by the Backbone network enters the Neck network, and through upsampling and splicing operations, it fuses feature information of different scales as the input of the Head network; 4.
4. The fused feature map output by the Neck network enters the Head network. After the CBS module and Conv2d two-dimensional convolution operation, the Bbox Loss bounding box loss and the Cls Loss classification loss are output to obtain the detection box position information of the strawberry target and the corresponding maturity classification results, namely, unripe state, half-ripe state and fully ripe state.
Citation Information
Cited By
Light-weight strawberry maturity detection method and device based on FDET-YOLOv11n
CN121600311A
A lightweight strawberry maturity detection method and device based on FDET-YOLOv11n
CN121600311B