Strawberry fruit detection system and method based on improved YOLOv8n network
By improving the YOLOv8n network structure, combining the receptive field attention convolution, extrusion excitation network and multi-scale parallel convolution feature fusion module, the strawberry fruit detection model is optimized, which improves the detection accuracy and efficiency, and solves the problems of high-parameter model calculation cost and insufficient detection accuracy of lightweight models.
Patent Information
- Application Number
- CN202510462519.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing computer vision fruit detection model has high calculation cost under high parameters and insufficient detection accuracy under lightweight models, resulting in low detection efficiency and high error detection rate.
The improved YOLOv8n network structure is adopted, including backbone network, neck network and prediction head network, and the feature extraction and detection performance is optimized using the receptive field attention convolution module, cross-stage dual convolution feature fusion module, extrusion excitation network and multi-scale parallel convolution feature fusion module.
It improves the detection performance and accuracy of strawberry fruit detection, reduces calculation costs, enhances the generalization ability of the model, and solves the bottlenecks in speed and performance of traditional models.
Smart Images

Figure CN120299034A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a strawberry fruit detection system and method based on an improved YOLOv8n network. Background Art
[0002] With the rapid development of agricultural technology, the modern agricultural production management mode centered on intelligence and automation is increasingly becoming the mainstream direction of the industry development. As a representative of high-value cash crops, strawberries have attracted much attention due to their short growth cycle and strong market demand. However, its production process has long faced practical challenges such as insufficient efficiency of manual picking, high labor costs, and great difficulty in quality control. In particular, it is worth noting that the traditional manual detection method not only has disadvantages such as low efficiency and high labor intensity, but also it is difficult to ensure the objectivity and accuracy of the detection results due to being easily affected by subjective judgment.
[0003] Therefore, in the prior art, computer vision detection is used to replace the manual detection method, effectively improving the efficiency of fruit detection and reducing the harvesting cost. As a fusion and improved version of the YOLO series, YOLOv8 adopts an anchor-free detection mechanism, directly predicting the object center position instead of relying on predefined anchor boxes, simplifying the training process and accelerating the post-processing step of non-maximum suppression (NMS), improving the detection speed and accuracy. However, in the practice of applying the YOLOv8 network to fruit detection, increasing the model parameter quantity can usually improve the feature extraction ability and enhance the detection accuracy for small target fruits or complex occlusion scenarios. However, an overly large model may lead to an increase in computational cost, affect real-time performance, and increase the risk of overfitting. Especially when the training data is limited, although the lightweight version model has fewer parameter quantities and faster inference speed, it may reduce the ability to distinguish dense fruits or morphological similar interfering objects (such as leaves) due to insufficient feature expression ability, resulting in missed detections or false detections, thus there is a contradictory relationship between the traditional model parameter quantity and detection accuracy in the existing computer vision fruit detection models. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the defects that the computer vision fruit detection method in the prior art has a relatively high computational cost under a high-parameter quantity model and insufficient detection accuracy in a lightweight model, so as to provide a strawberry fruit detection system and method based on an improved YOLOv8n network.
[0005] An improved YOLOv8n network structure for strawberry fruit detection, comprising a backbone network (Backbone), a neck network (Neck), and a prediction head network (Head) connected in sequence; The backbone network includes connected receptive field attention convolution modules (RFAConv), cross-stage double convolution feature fusion modules (C2f), squeeze-and-excitation networks (SENet), and fast spatial pyramid pooling modules (SPPF); the receptive field attention convolution module is used to obtain feature maps in the height and width directions of the input feature map, form an attention weight map based on the feature maps in the height and width directions, and is used to weight the input feature map rows to form an output feature map; the cross-stage double convolution feature fusion module is used to extract a basic feature map from the input feature map, generate a deep feature map by passing the basic feature map through a bottleneck layer, and fuse the basic feature map and the deep feature map to form an output feature map; the squeeze-and-excitation network is used to adjust the weights of each channel of the input feature map through a channel attention mechanism to obtain an output feature map; the fast spatial pyramid pooling module is used to perform pooling on the input feature map multiple times, and splice the input feature map and the pooling result along the channel dimension to form an output feature map; The neck network includes connected multi-scale parallel convolution cross-stage double convolution feature fusion modules (C2f_InceptionNext), feature splicing modules (Concat), upsampling modules (Upsample), and standard convolution modules (Conv); the multi-scale parallel convolution cross-stage double convolution feature fusion module is used to extract a basic feature map from the input feature map, generate multi-scale feature maps by passing the basic feature map through multi-scale parallel convolution layers, and fuse the basic feature map and the multi-scale feature maps to form an output feature map; the feature splicing module is used to splice the feature maps output by different branches to form a spliced feature map; the upsampling module is used to increase the resolution of the input feature map to form an output feature map; the standard convolution module is used to perform convolution on the input feature map to form an output feature map; The prediction head network includes multiple detection head modules (Detect); the detection head module is used to receive the feature map output by the neck network to form raw prediction values, including bounding boxes, confidence levels, and class probabilities.
[0006] Further, the backbone network includes a first receptive field attention convolution module, a second receptive field attention convolution module, a first cross-stage double convolution feature fusion module, a third receptive field attention convolution module, a second cross-stage double convolution feature fusion module, a fourth receptive field attention convolution module, a third cross-stage double convolution feature fusion module, a fifth receptive field attention convolution module, a fourth cross-stage double convolution feature fusion module, and a fast spatial pyramid pooling module connected in sequence.
[0007] Further, the neck network includes a first upsampling module, a first feature splicing module, a first multi-scale parallel convolutional cross-stage double convolutional feature fusion module, a second upsampling module, a second feature splicing module, a second multi-scale parallel convolutional cross-stage double convolutional feature fusion module, a first standard convolutional module, a third feature splicing module, a third multi-scale parallel convolutional cross-stage double convolutional feature fusion module, a second standard convolutional module, a fourth feature splicing module, and a fourth multi-scale parallel convolutional cross-stage double convolutional feature fusion module connected in sequence; the first upsampling module is connected to the output end of the fast spatial pyramid pooling module, the first feature splicing module is connected to the output end of the third cross-stage double convolutional feature fusion module, the second feature splicing module is connected to the output end of the second cross-stage double convolutional feature fusion module, the third feature splicing module is connected to the output end of the first multi-scale parallel convolutional cross-stage double convolutional feature fusion module, and the fourth feature splicing module is connected to the output end of the fast spatial pyramid pooling module.
[0008] Further, the prediction head network includes a first detection head module, a second detection head module, and a third detection head module; the first detection head module is connected to the output end of the second multi-scale parallel convolutional cross-stage double convolutional feature fusion module, the second detection head module is connected to the output end of the third multi-scale parallel convolutional cross-stage double convolutional feature fusion module, and the third detection head module is connected to the output end of the fourth multi-scale parallel convolutional cross-stage double convolutional feature fusion module.
[0009] Further, the receptive field attention convolution module is used for: expanding the number of channels through group convolution operations; adjusting the size of the feature map through normalization and the nonlinear activation function ReLU; respectively performing average pooling operations on the feature map in the height and width directions to generate intermediate feature maps; splicing the two intermediate feature maps and fusing them through a convolutional layer to generate a fused feature map; normalizing and nonlinearly transforming the fused feature map, splitting it into two parts of feature maps, respectively generating attention weight maps through convolutional operations and the Sigmoid function, weighting the input feature map based on the attention weight map, and generating an output feature map through a convolutional layer.
[0010] A strawberry fruit detection system based on an improved YOLOv8n network includes a preprocessing module, an improved YOLOv8n network, and a postprocessing module; the preprocessing module is used to process the input image to generate an original feature map; the structure of the improved YOLOv8n network is as described above, and it is used to process the original feature map to obtain original prediction values; the postprocessing module is used to process the original prediction values to obtain prediction results.
[0011] Further, the preprocessing module includes a dimension normalization unit and a numerical standardization unit; the dimension normalization unit is used to adapt the image size of the input image to the preset input dimension of the improved YOLOv8n network, and the numerical standardization unit is used to constrain the pixel data distribution of the input image within the range of [0, 1] through linear transformation of pixel values.
[0012] Further, the postprocessing module includes a distance intersection over union non-maximum suppression module, and the distance intersection over union non-maximum suppression includes a sorting unit, an iterative selection unit, a distance intersection over union calculation unit, and a filtering unit; the sorting unit is used to sort the original preset values in descending order based on the confidence scores of the candidate boxes, the iterative selection unit is used to sequentially select the candidate box with the highest confidence score that has not been suppressed from the highest score and add it to the prediction result, the distance intersection over union calculation unit is used to calculate the distance intersection over union between the remaining candidate boxes and the candidate box with the highest confidence score, and the filtering unit is used to suppress the candidate boxes whose distance intersection over union with the candidate box with the highest confidence score is greater than the preset threshold.
[0013] A strawberry fruit detection method based on an improved YOLOv8n network includes the following method steps: training the improved YOLOv8n network; inputting an input image into a strawberry fruit detection system based on the improved YOLOv8n network to obtain a prediction result, and the strawberry fruit detection system based on the improved YOLOv8n network is as described above; the loss function for training the improved YOLOv8n network is a dynamic weighted intersection over union loss function.
[0014] Further, the loss function for training the improved YOLOv8n network is a dynamic weighted intersection over union loss function, which is expressed as: ; Wherein, represents the dynamic weighted intersection over union, represents the power form of the intersection over union loss, γ represents the focusing degree parameter, represents the weight factor; In a closed box with width W g and height H g , X and Y represent the center coordinates of the predicted box, X gt and Y gt represent the center coordinates of the ground truth box, and the weight factor is expressed as: .
[0015] Beneficial effects: The present invention discloses a strawberry fruit detection system and method based on an improved YOLOv8n network. The improved YOLOv8n network structure includes a backbone network, a neck network, and a prediction head network connected in sequence. The backbone network includes a receptive field attention convolution module, a cross-stage double convolution feature fusion module, a squeeze-and-excitation network, and a fast spatial pyramid pooling module connected together. The neck network includes a multi-scale parallel convolution cross-stage double convolution feature fusion module, a feature splicing module, an upsampling module, and a standard convolution module. The receptive field attention convolution module replaces the standard convolution, improving the network performance and significantly enhancing the detection performance in the detection of strawberry fruits. At the same time, the introduction of the squeeze-and-excitation network strengthens the learning ability of the backbone network for strawberry color and shape features. Meanwhile, the optimized decoupled head design of YOLOv8 better adapts to the requirements of classification and regression tasks, avoiding the interference that may be brought by shared parameters and improving the accuracy and generalization ability of the model. And the multi-scale parallel convolution cross-stage double convolution feature fusion module reduces the memory access cost while maintaining a large receptive field, reducing the computational cost and improving the efficiency, solving the bottleneck of traditional CNN in terms of speed and performance. Brief Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 It is a schematic block diagram of the structure of the strawberry fruit detection system of the present invention; Figure 2 It is a schematic block diagram of the structure of the receptive field attention convolution module of the improved YOLOv8n network of the present invention; Figure 3 It is a schematic block diagram of the structure of the squeeze-and-excitation network of the improved YOLOv8n network of the present invention; Figure 4 It is a schematic block diagram of the structure of the multi-scale parallel convolution layer of the improved YOLOv8n network of the present invention.
[0018] Figure 5 It is a schematic block diagram of the structure of the cross-stage double convolution feature fusion module of the improved YOLOv8n network of the present invention. Detailed Embodiments
[0019] To make the above objects, features, and advantages of the present application more apparent and understandable, the following detailed description of the specific embodiments of the present application will be provided in conjunction with the accompanying drawings. A lot of specific details are set forth in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0020] In the description of the present application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0021] Embodiment 1: Referring to Figures 1 to 4 As shown, this embodiment provides an improved YOLOv8n network structure for strawberry fruit detection, including a backbone network (Backbone), a neck network (Neck), and a prediction head network (Head) connected in sequence; The backbone network includes a connected receptive field attention convolution module (RFAConv), a cross-stage double convolution feature fusion module (C2f), a squeeze-and-excitation network (SENet), and a fast spatial pyramid pooling module (SPPF); the receptive field attention convolution module is used to obtain feature maps in the height and width directions of the input feature map, form an attention weight map based on the feature maps in the height and width directions, and be used to weight the input feature map rows to form an output feature map; the cross-stage double convolution feature fusion module is used to extract a basic feature map from the input feature map, generate a deep feature map by passing the basic feature map through a bottleneck layer, and fuse the basic feature map and the deep feature map to form an output feature map; the squeeze-and-excitation network is used to adjust the weights of each channel of the input feature map through a channel attention mechanism to obtain an output feature map; the fast spatial pyramid pooling module is used to perform multiple poolings on the input feature map and splice the input feature map and the pooling results along the channel dimension to form an output feature map; The neck network includes a connected multi-scale parallel convolutional cross-stage double convolutional feature fusion module (C2f_InceptionNext), a feature concatenation module (Concat), an upsampling module (Upsample), and a standard convolutional module (Conv); the multi-scale parallel convolutional cross-stage double convolutional feature fusion module is used to extract a basic feature map from the input feature map, generate a multi-scale feature map through a multi-scale parallel convolutional layer for the basic feature map, and fuse the basic feature map and the multi-scale feature map to form an output feature map; the feature concatenation module is used to concatenate the feature maps output by different branches to form a concatenated feature map; the upsampling module is used to increase the resolution of the input feature map to form an output feature map; the standard convolutional module is used to perform convolution on the input feature map to form an output feature map; The prediction head network includes multiple detection head modules (Detect); the detection head module is used to receive the feature map output by the neck network to form original prediction values, including bounding boxes, confidence levels, and class probabilities.
[0022] Specifically, the backbone network includes a first receptive field attention convolutional module, a second receptive field attention convolutional module, a first cross-stage double convolutional feature fusion module, a third receptive field attention convolutional module, a second cross-stage double convolutional feature fusion module, a fourth receptive field attention convolutional module, a third cross-stage double convolutional feature fusion module, a fifth receptive field attention convolutional module, a fourth cross-stage double convolutional feature fusion module, and a fast spatial pyramid pooling module connected in sequence.
[0023] The neck network includes a first upsampling module, a first feature concatenation module, a first multi-scale parallel convolutional cross-stage double convolutional feature fusion module, a second upsampling module, a second feature concatenation module, a second multi-scale parallel convolutional cross-stage double convolutional feature fusion module, a first standard convolutional module, a third feature concatenation module, a third multi-scale parallel convolutional cross-stage double convolutional feature fusion module, a second standard convolutional module, a fourth feature concatenation module, and a fourth multi-scale parallel convolutional cross-stage double convolutional feature fusion module connected in sequence; the first upsampling module is connected to the output end of the fast spatial pyramid pooling module, the first feature concatenation module is connected to the output end of the third cross-stage double convolutional feature fusion module, the second feature concatenation module is connected to the output end of the second cross-stage double convolutional feature fusion module, the third feature concatenation module is connected to the output end of the first multi-scale parallel convolutional cross-stage double convolutional feature fusion module, and the fourth feature concatenation module is connected to the output end of the fast spatial pyramid pooling module.
[0024] The prediction head network includes a first detection head module, a second detection head module, and a third detection head module; the first detection head module is connected to the output end of the second multi-scale parallel convolutional cross-stage double convolutional feature fusion module, the second detection head module is connected to the output end of the third multi-scale parallel convolutional cross-stage double convolutional feature fusion module, and the third detection head module is connected to the output end of the fourth multi-scale parallel convolutional cross-stage double convolutional feature fusion module.
[0025] In this embodiment, the receptive field attention convolution module is used for: when the input feature map has a size of C×H×W, the number of channels is expanded to CK through a grouped convolution operation 2 , so that the new feature map has a size of CK 2 ×H×W; through normalization and the non-linear activation function ReLU processing, the feature map size is adjusted to C×KH×W; average pooling operations are respectively performed on the feature map in the height and width directions to generate intermediate feature maps; two intermediate feature maps with sizes of C×1×KW and C×KH×1 are concatenated, and through convolution layer fusion, the generated fused feature map has a size of C×(KW + KH)×1; the fused feature map is normalized and non-linearly transformed, split into two parts of feature maps, and attention weight maps are generated respectively through convolution operations and the Sigmoid function. Based on the attention weight maps, the input feature map is weighted, and the output feature map is generated through a convolution layer. The receptive field attention convolution module enables the convolution operation to adaptively adjust the receptive field size by introducing a dynamic receptive field mechanism and an attention mechanism, so as to better capture target features of different scales. In this embodiment, the size of the spatial features of the input feature map is C×H×W, which consists of non-overlapping sliding windows. When a 3×3 convolution kernel is used to extract features, each 3×3-sized window in the receptive field spatial features represents a receptive field slider. After conversion, the height and width of the receptive field spatial features are tripled, and the size is expanded to 3C×3H×3W, so that the receptive field area is expanded, and thus multi-scale information can be better captured.
[0026] Refer to Figure 5As shown, the cross-stage double convolutional feature fusion module is used to optimize feature extraction, enhancement, and multi-scale feature fusion. It extracts the base feature map from the input feature map, generates the deep feature map by passing the base feature map through the bottleneck layer (Bottleneck), and fuses the base feature map and the deep feature map to form the output feature map. Specifically, it uses the convolutional layer to extract the base features from the input feature map, and then divides the generated feature map into two paths. One path is directly sent to the feature concatenation layer (Concat), and the other path undergoes deep processing through multiple bottleneck layers. The bottleneck layer contains convolution, normalization, and activation functions, and is good at capturing complex features, especially suitable for targets with multi-scale and complex backgrounds. Finally, the two paths of features are merged in the feature concatenation layer to achieve the effective integration of multi-level information and improve the detection ability of multi-scale objects. The depth of the model is controlled by the depth_multiple parameter, which allows flexible adjustment of the number of bottleneck layers, thus achieving the best balance between accuracy and computational complexity. The cross-stage double convolutional feature fusion module can dynamically optimize the network depth according to specific task requirements by adjusting the number of bottleneck layers, and is suitable for scenarios with limited computing resources. In addition, the concatenated feature map is processed by an additional convolutional layer to enhance the feature expression ability and provide high-quality input for subsequent tasks.
[0027] In this embodiment, the Squeeze-and-Excitation Networks (SENet) is an innovative convolutional neural network architecture that introduces an attention mechanism in the channel dimension. By modeling the dynamic dependencies between feature channels, this network significantly improves the model's representational ability and classification performance without significantly increasing the computational burden. Its core idea is to adaptively learn the weight coefficients of each feature channel, thereby enhancing the key features and suppressing redundant information. The Squeeze-and-Excitation Networks achieve feature recalibration through two key operations: Squeeze and Excitation. In the Squeeze stage, global average pooling aggregates the spatial information of each channel into a scalar representation to capture the global context. Subsequently, in the Excitation stage, the fully connected layer and the non-linear activation function are used to learn the complex dependencies between channels and generate the adaptive weights. Finally, these weights dynamically adjust the original features, enabling the network to focus on the most discriminative features, thus significantly improving the model's expressive ability.
[0028] Specifically, for any feature transformation operation F_tr that maps the input X to the feature map U ∈ ℝ^{H×W×C}, this work realizes feature recalibration by constructing the Squeeze-and-Excitation Networks. This module adaptively adjusts the original feature map by dynamically learning the channel attention weights, and its core process includes two serial operations: Global information embedding is performed during the Squeeze operation. Global Average Pooling (GAP) is used to compress the spatial dimension (H×W) of the feature map U, generating a channel statistical descriptor z ∈ ℝ C . This operation aggregates the two-dimensional feature responses U ∈ ℝ H×W of each channel into a scalar representation, effectively capturing the global spatial distribution of channel features: ; The Squeeze operation eliminates redundant position-sensitive information through spatial dimension compression and constructs channel-level statistics with a global receptive field, providing a basis for subsequent channel correlation modeling.
[0029] Channel weight generation is performed during the Excitation operation. An adaptive gating mechanism based on a bottleneck layer is designed to learn the non-linear dependence relationship between channels through a two-layer fully connected network. Specifically: the first layer implements dimensionality reduction and applies ReLU activation, and the second layer restores the dimension and generates a normalized weight s ∈ [0, 1] through the Sigmoid function. C .
[0030] Finally, the learned channel weight s is applied to the original feature map U to obtain a recalibrated feature map, which is then input into the subsequent network layer. This gating mechanism enables the model to explicitly model the importance of channels, strengthen discriminative features, and suppress redundant information.
[0031] The attention mechanism of the Squeeze-and-Excitation network realizes the adaptive recalibration of feature channels by establishing an explicit inter-channel dependence relationship modeling. Compared with traditional convolution operations, it automatically evaluates the importance of each channel's information using a lightweight structure, selectively enhances key features, and suppresses redundant or interfering information. The integration of this mechanism in the YOLOv8n backbone network only requires adding a lightweight structure composed of two fully connected layers after the convolutional layer to achieve feature recalibration through global information aggregation and adaptive weight allocation. This improvement can significantly enhance the model's sensitivity to the subtle features of strawberry fruits (such as fruit surface texture and maturity color difference). Especially in complex orchard environments, it can effectively distinguish target features under similar color backgrounds, reducing the false detection rate caused by foliage occlusion and lighting changes. Since the proportion of newly added structural parameters in the original network is small, without increasing the computational overhead almost, the feature discrimination ability and localization accuracy of the detection model are systematically optimized in dense fruit scenarios.
[0032] Refer to Figure 1As shown, in this embodiment, the backbone network structure includes a complex structure with ten layers, enabling efficient feature extraction. However, introducing the squeeze-and-excitation network in each layer will significantly increase the computational burden and model complexity, which is not ideal for application scenarios with limited resources. Therefore, in practical applications, it is necessary to balance the performance improvement brought by the attention mechanism and the computational cost.
[0033] To balance the accuracy and efficiency of the model, three schemes are designed in this paper. After adding the squeeze-and-excitation network to the first and third layers, the third and fifth layers, and the fifth and seventh layers respectively, the optimal introduction method is selected as the basis for subsequent model improvement through comparative experiments. As shown in Table 1, in this embodiment, the introduction of the squeeze-and-excitation network is not counted in the number of layers of the backbone network.
[0034] Table 1: Results of comparative experiments on SENet improvement
[0035] It can be seen that through the comparative experiments on different hierarchical optimization schemes of the YOLOv8n network architecture, the optimal improvement strategy suitable for this dataset is determined in this embodiment. The experimental results show that although the adjustment of the shallow network (the first and third layers) slightly increases the mean average precision (mAP50) by 0.12 percentage points, it causes the detection precision (Precision) and recall rate (Recall) to decrease by 0.53% and 0.21% respectively, indicating that the optimization of shallow parameters is likely to trigger negative fluctuations in the basic detection performance. In contrast, the optimization scheme of the deep network (the fifth and seventh layers) shows significant advantages. While the Precision increases by 0.55% and the Recall increases by 0.73%, the mAP50 index achieves a breakthrough increase of 0.83%. It should be noted that although the improvement of the middle network (the third and fifth layers) increases the Precision and Recall by 0.15% and 0.18% respectively, it is accompanied by an abnormal decrease of 0.35% in the mAP50 index, revealing that the adjustment of this layer will weaken the multi-scale target adaptation ability of the model. Based on the above deep-dependence characteristics (the contribution rate of deep optimization to the improvement of comprehensive performance reaches 98.7%), as the preference of this embodiment, the fifth and seventh layers are finally selected as the integration positions of the squeeze-and-excitation network to optimize the detection effect by enhancing the deep feature extraction ability.
[0036] In this embodiment, the multi-scale parallel convolution layer (InceptionNeXt) integrates the multi-scale feature extraction mechanism of the classical Inception module and the large-kernel depth convolution design concept of ConvNeXt, achieving a balance between computational efficiency and model performance. The kernel decomposition strategy of the multi-scale parallel convolution layer decouples the traditional large-size convolution kernel into a multi-path small-kernel combination while retaining the identity mapping characteristics of the key channels. Specifically, its core module uses a three-way parallel structure to process with 3×3, 1×k, and k×1 convolution kernels respectively, and the rest is directly passed through the identity mapping, thereby reducing the memory access cost while maintaining a large receptive field. In the neck network of this improved YOLOv8n network, the multi-scale parallel convolution layer is introduced into the cross-stage double convolution feature fusion module to form a multi-scale parallel convolution cross-stage double convolution feature fusion module, solving the bottleneck problem of the traditional model in terms of speed and performance and achieving efficient image classification.
[0037] Embodiment 2: This embodiment provides a strawberry fruit detection system based on an improved YOLOv8n network, including a preprocessing module, an improved YOLOv8n network, and a postprocessing module; the preprocessing module is used to process the input image to generate an original feature map; the structure of the improved YOLOv8n network is as described in Embodiment 1, and is used to process the original feature map to obtain an original prediction value; the postprocessing module is used to process the original prediction value to obtain a prediction result.
[0038] Specifically, the preprocessing module includes a dimension normalization unit and a numerical standardization unit; the dimension normalization unit is used to adapt the image size of the input image to the preset input dimension of the improved YOLOv8n network, and the numerical standardization unit is used to constrain the pixel data distribution of the input image within the range of [0,1] through pixel value linear transformation.
[0039] In this embodiment, the postprocessing module includes a distance intersection over union non-maximum suppression module (DIoU-NMS, Distance Intersection over Union - Non Maximum Suppression), and the distance intersection over union non-maximum suppression includes a sorting unit, an iterative selection unit, a distance intersection over union calculation unit, and a filtering unit; the sorting unit is used to sort the original preset values in descending order based on the confidence scores of the candidate boxes, the iterative selection unit is used to sequentially select the candidate box with the highest confidence score that has not been suppressed from the highest score and add it to the prediction result, the distance intersection over union calculation unit is used to calculate the distance intersection over union of the remaining candidate boxes and the candidate box with the highest confidence score, and the filtering unit is used to suppress the candidate boxes whose distance intersection over union with the candidate box with the highest confidence score is greater than the preset threshold.
[0040] Specifically, DIoU-NMS uses DIoU as the criterion for NMS, and at the same time considers the distance between the centers of the two bounding boxes as the evaluation criterion. Let the Euclidean distance between the centers of the two bounding boxes be d, and the length of the diagonal of the smallest closed rectangle containing the two bounding boxes be c. The calculation formula of DIoU is: .
[0041] Example 3: This embodiment provides a strawberry fruit detection method based on an improved YOLOv8n network, including the following method steps: training the improved YOLOv8n network; inputting the input image into the strawberry fruit detection system based on the improved YOLOv8n network to obtain the prediction result, and the strawberry fruit detection system based on the improved YOLOv8n network is as described in Example 2; the loss function for training the improved YOLOv8n network is the dynamic weighted intersection over union loss function (Wise-IoU).
[0042] Specifically, the loss function for training the improved YOLOv8n network is the dynamic weighted intersection over union loss function, which is expressed as: ; Among them, represents the dynamic weighted intersection over union, represents the power form of the intersection over union loss, γ represents the focusing degree parameter, represents the weight factor; In the closed box with width W g and height H g , X and Y represent the center coordinates of the predicted box, X gt and Y gt represent the center coordinates of the ground truth box, and the weight factor is expressed as: .
[0043] In this embodiment, the dynamic weighted intersection over union loss function redefines the quality evaluation standard of the anchor box by introducing the outlier degree, effectively reduces the competition of high-quality anchor boxes during training through the gradient gain distribution strategy, and at the same time significantly reduces the harmful gradient impact that may be brought by low-quality samples.
[0044] In this embodiment, the improved YOLOv8n network has the following optimizations in terms of structure: (1) replacing the standard convolution with a receptive field attention convolution module in the backbone network; (2) introducing a squeeze-and-excitation network in the backbone network to enhance the backbone network's learning ability for strawberry color and shape features; (3) replacing the convolutional layer in the bottleneck layer of the cross-stage double convolutional feature fusion module in the neck network with a multi-scale parallel convolution. An ablation experiment is conducted on the strawberry dataset. The experimental environment and parameters for each experiment are the same. The results of the ablation experiment are shown in Table 2.
[0045] Table 2: Results of the ablation experiment
[0046] It can be seen that the results of the ablation experiment show that different improvement combinations have improvements compared to the baseline model. Replacing the standard convolution with a receptive field attention convolution module can significantly improve the performance of the model, but the model size increases. The squeeze-and-excitation network enhances the backbone network's learning ability for strawberry color and shape features, effectively improving the precision without increasing the model size. The multi-scale parallel convolutional layer can effectively reduce the computational amount and significantly reduce the model size while the performance increases. When any two improvement measures are introduced simultaneously, the overall performance of the model is still significantly better than the baseline model, thus verifying the effectiveness of each improvement measure. This detection method effectively solves the contradiction between the number of traditional model parameters and the detection accuracy while maintaining the high-precision feature extraction ability, and achieves high-precision detection under the premise of low computational resource consumption in the working condition of strawberry fruit detection.
[0047] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0048] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. An improved YOLOv8n network structure for strawberry fruit detection, characterized in that, It includes a backbone network, a neck network, and a prediction head network connected in sequence; The backbone network includes a receptive field attention convolution module, a cross-stage double convolution feature fusion module, a squeeze-and-excitation network, and a fast spatial pyramid pooling module connected in sequence. The receptive field attention convolution module is used to obtain feature maps in the height and width directions of the input feature map, form an attention weight map based on the feature maps in the height and width directions, and use it to weight the rows of the input feature map to form an output feature map. The cross-stage double convolution feature fusion module is used to extract a basic feature map from the input feature map, generate a deep feature map by passing the basic feature map through a bottleneck layer, and fuse the basic feature map and the deep feature map to form an output feature map; The squeeze-and-excitation network is used to adjust the channel weights of the input feature map through a channel attention mechanism to obtain an output feature map. The fast spatial pyramid pooling module is used to perform multiple poolings on the input feature map, and splice the input feature map and the pooling result along the channel dimension to form an output feature map; The neck network includes a multi-scale parallel convolution cross-stage double convolution feature fusion module, a feature splicing module, an upsampling module, and a standard convolution module connected in sequence. The multi-scale parallel convolution cross-stage double convolution feature fusion module is used to extract a basic feature map from the input feature map, generate a multi-scale feature map by passing the basic feature map through a multi-scale parallel convolution layer, and fuse the basic feature map and the multi-scale feature map to form an output feature map; The feature splicing module is used to splice the feature maps output by different branches to form a spliced feature map. The upsampling module is used to increase the resolution of the input feature map to form an output feature map. The standard convolution module is used to perform convolution on the input feature map to form an output feature map; The prediction head network includes a plurality of detection head modules. The detection head module is used to receive the feature map output by the neck network and form original prediction values, including bounding boxes, confidence levels, and class probabilities.
2. An improved YOLOv8n network structure for strawberry fruit detection according to claim 1, characterized in that, The backbone network includes a first receptive field attention convolution module, a second receptive field attention convolution module, a first cross-stage double convolution feature fusion module, a third receptive field attention convolution module, a second cross-stage double convolution feature fusion module, a fourth receptive field attention convolution module, a third cross-stage double convolution feature fusion module, a fifth receptive field attention convolution module, a fourth cross-stage double convolution feature fusion module, and a fast spatial pyramid pooling module connected in sequence.
3. An improved YOLOv8n network structure for strawberry fruit detection according to claim 2, characterized in that, The neck network includes a first upsampling module, a first feature splicing module, a first multi-scale parallel convolutional cross-stage double convolutional feature fusion module, a second upsampling module, a second feature splicing module, a second multi-scale parallel convolutional cross-stage double convolutional feature fusion module, a first standard convolutional module, a third feature splicing module, a third multi-scale parallel convolutional cross-stage double convolutional feature fusion module, a second standard convolutional module, a fourth feature splicing module, and a fourth multi-scale parallel convolutional cross-stage double convolutional feature fusion module connected in sequence; the first upsampling module is connected to the output end of the fast spatial pyramid pooling module, the first feature splicing module is connected to the output end of the third cross-stage double convolutional feature fusion module, the second feature splicing module is connected to the output end of the second cross-stage double convolutional feature fusion module, the third feature splicing module is connected to the output end of the first multi-scale parallel convolutional cross-stage double convolutional feature fusion module, and the fourth feature splicing module is connected to the output end of the fast spatial pyramid pooling module.
4. An improved YOLOv8n network structure for strawberry fruit detection according to claim 3, characterized in that, The prediction head network includes a first detection head module, a second detection head module, and a third detection head module; the first detection head module is connected to the output end of the second multi-scale parallel convolutional cross-stage double convolutional feature fusion module, the second detection head module is connected to the output end of the third multi-scale parallel convolutional cross-stage double convolutional feature fusion module, and the third detection head module is connected to the output end of the fourth multi-scale parallel convolutional cross-stage double convolutional feature fusion module.
5. An improved YOLOv8n network structure for strawberry fruit detection according to claim 1, characterized in that, The receptive field attention convolution module is used to: expand the number of channels through group convolution operations; adjust the feature map size through normalization and the nonlinear activation function ReLU; perform average pooling operations on the feature map in the height and width directions respectively to generate intermediate feature maps; splice the two intermediate feature maps and fuse them through a convolutional layer to generate a fused feature map; normalize and nonlinearly transform the fused feature map, divide it into two parts of feature maps, generate attention weight maps through convolutional operations and the Sigmoid function respectively, weight the input feature map based on the attention weight maps, and generate an output feature map through a convolutional layer.
6. A strawberry fruit detection system based on an improved YOLOv8n network, characterized in that, It includes a preprocessing module, an improved YOLOv8n network, and a postprocessing module; the preprocessing module is used to process the input image to generate an original feature map; the structure of the improved YOLOv8n network is as described in any one of claims 1 to 4, and is used to process the original feature map to obtain an original prediction value; the postprocessing module is used to process the original prediction value to obtain a prediction result.
7. A strawberry fruit detection system based on an improved YOLOv8n network according to claim 6, characterized in that, The preprocessing module includes a dimension normalization unit and a numerical standardization unit; the dimension normalization unit is used to adapt the image size of the input image to the preset input dimension of the improved YOLOv8n network, and the numerical standardization unit is used to constrain the pixel data distribution of the input image within the range of [0, 1] through linear transformation of pixel values.
8. The strawberry fruit detection system based on the improved YOLOv8n network according to claim 6, characterized in that, The post-processing module includes a distance intersection over union non-maximum suppression module, and the distance intersection over union non-maximum suppression includes a sorting unit, an iterative selection unit, a distance intersection over union calculation unit, and a filtering unit; the sorting unit is used to sort the original preset values in descending order based on the confidence scores of the candidate boxes, the iterative selection unit is used to sequentially select the candidate box with the highest confidence score that has not been suppressed from the highest score and add it to the prediction result, the distance intersection over union calculation unit is used to calculate the distance intersection over union between the remaining candidate boxes and the candidate box with the highest confidence score, and the filtering unit is used to suppress the candidate boxes whose distance intersection over union with the candidate box with the highest confidence score is greater than the preset threshold.
9. A strawberry fruit detection method based on an improved YOLOv8n network, characterized in that, It includes the following method steps: training an improved YOLOv8n network; inputting the input image into the strawberry fruit detection system based on the improved YOLOv8n network to obtain the prediction result, and the strawberry fruit detection system based on the improved YOLOv8n network is as described in claims 6 to 8; The loss function for training the improved YOLOv8n network is a dynamic weighted intersection over union loss function.
10. A strawberry fruit detection method based on an improved YOLOv8n network according to claim 9, characterized in that, The loss function for training the improved YOLOv8n network, which is a dynamic weighted intersection over union loss function, is expressed as: ; Among them, represents the dynamic weighted intersection over union (IoU), represents the power form of the IoU loss, and γ represents the focusing degree parameter, represents the weight factor; In a closed box with width W g and height H g where X and Y represent the center coordinates of the predicted box, X gt and Y gt represent the center coordinates of the ground truth box, the weight factor is expressed as: 。
Citation Information
Patent Citations
Lightweight target detection network and method based on YOLO and electronic equipment
CN115546620A
Fruit detection method and system in complex environment based on improved YOLOv8n and application
CN118537718A
Citrus maturity detection method based on improved YOLOv8 and related device
CN118968502A
MULTI-TASK PANOPTIC DRIVING PERCEPTION METHOD AND SYSTEM BASED ON IMPROVED YOU ONLY LOOK ONCE VERSION 5 (YOLOv5)
US20250005914A1
Cited By
Lightweight model-based chemical instrument identification method, apparatus and device, and medium
CN121617078A
Voltage sag source classification method based on multi-modal feature extraction and double attention mechanism
CN122196507A