Blueberry fruit focusing detection method and system based on lightweight YOLO model
By constructing a multi-dimensional dataset, introducing FasterNet and DAttention mechanisms, and optimizing the loss function of the YOLO model, the problems of low efficiency and insufficient accuracy in blueberry fruit detection were solved, and high-precision target detection in complex environments was achieved.
Patent Information
- Application Number
- CN202511040356.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-11
Smart Images

Figure CN120932096A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection technology, and in particular to a method and system for focusing and detecting blueberry fruits based on a lightweight YOLO model. Background Technology
[0002] In the process of modern agricultural development, the blueberry industry, as a sector with high economic value, is constantly pursuing improvements in production efficiency and quality. However, the blueberry fruit testing and harvesting process faces many severe challenges.
[0003] In the early days, blueberry inspection relied primarily on manual sorting, with workers identifying the fruit based on sight and experience. However, this method was extremely inefficient, unable to meet the needs of large-scale cultivation, and highly susceptible to subjective factors, leading to judgment biases and failing to guarantee accuracy and consistency. With technological advancements, traditional machine vision threshold segmentation technology was introduced, distinguishing fruit from the background by setting thresholds for color, shape, and other parameters. However, the complex blueberry growing environment, variable lighting conditions, and the fact that fruit is often obscured by leaves to varying degrees can easily cause color thresholds to fail, resulting in a significant reduction in detection accuracy and failing to meet actual production needs.
[0004] The rise of deep learning technology has revolutionized object detection. While the YOLO series models excel in balancing real-time performance and accuracy, their direct application to blueberry detection reveals a series of technical bottlenecks. Regarding model lightweighting, existing methods often sacrifice some detection accuracy to reduce the number of parameters, failing to meet the stringent high-precision requirements of agricultural detection. In terms of attention mechanisms, traditional modules cannot dynamically adapt to localized occlusion in blueberry detection. When parts like the stem are partially obscured by leaves, it's difficult to focus on the occlusion edge features, resulting in poor detection performance. Furthermore, in terms of loss function optimization, existing functions cannot specifically constrain the aspect ratio when dealing with small targets like blueberry stems, leading to easily distorted bounding boxes and large localization errors. Therefore, a blueberry fruit focusing detection method and system based on a lightweight YOLO model is currently needed. Summary of the Invention
[0005] To address the problems of low detection efficiency and lack of accuracy in traditional blueberry fruit detection, this invention provides a blueberry fruit focusing detection method and system based on a lightweight YOLO model.
[0006] In a first aspect, the present invention provides a blueberry fruit focusing detection method based on a lightweight YOLO model, which adopts the following technical solution:
[0007] A blueberry fruit focus detection method based on a lightweight YOLO model includes:
[0008] Blueberry images were collected to form a multidimensional dataset, and the collected multidimensional dataset was labeled.
[0009] Data augmentation is performed based on the acquired multidimensional dataset, including geometric transformation of the acquired original dataset and optical and occlusion simulations by adding Gaussian noise and gradient operators.
[0010] The preprocessed cube was used as input to build the detection model, including the introduction of the FasterNet ultra-lightweight network into the Backbone layer of the YOLOv10n model;
[0011] A dynamic attention mechanism called DAttention is introduced on the basis of the detection model, including a dynamic attention mechanism of DAttention embedded after the PSA layer based on the improved Backbone layer.
[0012] The completed detection model is used for model training and optimization, including optimizing the detection box localization using the CIoU loss function;
[0013] Using the optimal weights obtained during training, the test set is input for detection, generating the final detection results.
[0014] Furthermore, the optical and occlusion simulations performed by adding Gaussian noise and gradient operators include adding zero-mean Gaussian noise using a probability density function, adjusting brightness through linear transformation, covering part of the fruit area with an irregular polygon mask, calculating the image gradient using the Sobel operator, generating mask edges along high-gradient regions, simulating leaf occlusion contours, and thus completing data augmentation. The image gradient calculation formula is as follows:
[0015]
[0016] Among them, G x Represented as a horizontal kernel, G y It is represented as a vertical core.
[0017] Furthermore, the step of using the preprocessed multidimensional dataset as input to construct the detection model includes replacing the C2f module of the YOLOv10n model with a C2f-Faster module using the FasterNet ultra-lightweight network. The FasterNet ultra-lightweight network consists of multiple stages of FasterNet modules, with an embedding layer set before each stage for spatial downsampling and channel expansion. The FasterNet module has a built-in PConv layer, an intermediate Conv layer, and an output Conv layer, and a BatchNorm regularization layer and a ReLU activation function are connected after the intermediate Conv layer.
[0018] Furthermore, the C2f-Faster module uses a main branch and a side branch for feature fusion. The main branch extracts fine-grained features of the fruit step by step through three cascaded FasterNet modules. Downsampling is achieved through an embedding layer at each stage. The side branch directly shorts the original feature map to retain the spatial location information of the fruit. Finally, the features of the main branch and the side branch are merged through a Concat operation.
[0019] Furthermore, the dynamic attention mechanism DAttention embedded after the PSA layer based on the improved Backbone layer includes linearly transforming the query features using weight prediction parameters and offset prediction parameters to generate the weight distribution and spatial offset of the sampling points. Then, the weights are normalized using Softmax, and the offsets are added to the reference point coordinates to obtain irregular sampling positions. Finally, features are extracted and weighted fused using bilinear interpolation. The calculation formula for the DAttention dynamic attention mechanism is as follows:
[0020]
[0021] Where, x q Let K represent the query feature, and W represent the total number of sampling points. A Represented as the weighted prediction parameter, W Δ The offset prediction parameters are represented by Bilinear, which represents the bilinear interpolation function, and Softmax represents the normalization operation.
[0022] Furthermore, the optimization of detection box localization using the CIoU loss function includes calculating the Euclidean distance between the center point of the predicted box and the center point of the ground truth box, normalizing the Euclidean distance to the diagonal length of the smallest closed region containing both boxes, and obtaining the center point distance term. Then, by calculating the angle difference between the aspect ratios of the predicted and ground truth boxes, the shape of the predicted box is constrained to match the ground truth box. Based on the center point distance term and the aspect ratio, the CIoU loss function is calculated to obtain the optimized detection box. Adaptive NMS processing is then applied to detection boxes with confidence scores greater than a confidence threshold. The CIoU loss function calculation formula is as follows:
[0023]
[0024] Where IoU represents the intersection-union ratio between the predicted bounding box and the ground truth bounding box. The value is represented as the center point distance, where b represents the center point coordinates of the predicted bounding box. gt Represented as the center point coordinates of the true bounding box, ρ 2 It is represented by the Euclidean distance function, c represents the diagonal length of the smallest closure region containing the predicted box and the ground truth box, v represents the aspect ratio consistency term, and α represents the dynamic weight coefficient.
[0025] Furthermore, the optimal weights obtained through training include employing an early stopping mechanism. When the average precision improvement on the validation set is less than a set threshold, the current detection model weights are saved as the optimal weights. Then, a cosine decay learning rate strategy is used to dynamically adjust the learning rate. This is combined with SGD optimizer and AMP hybrid precision training, using a momentum term to skip shallow local optima and find a better weight combination. The formula for the cosine decay learning rate strategy is:
[0026]
[0027] Where lr0 represents the initial learning rate, epoch represents the current training epoch, and max_epoch represents the maximum training epoch.
[0028] Secondly, a blueberry fruit focusing detection system based on a lightweight YOLO model includes:
[0029] The data acquisition module is configured to: acquire blueberry images to form a multi-dimensional dataset, and label the acquired multi-dimensional dataset;
[0030] The preprocessing module is configured to perform data augmentation based on the acquired cube, including geometric transformation of the acquired original dataset and optical and occlusion simulations by adding Gaussian noise and gradient operators.
[0031] The model module is configured to: build a detection model by taking the preprocessed cube as input, including introducing the FasterNet ultralight network into the backbone layer of the YOLOv10n model;
[0032] The attention module is configured to introduce a DAttention dynamic attention mechanism on the basis of the detection model, including a DAttention dynamic attention mechanism embedded after the PSA layer based on the improved Backbone layer.
[0033] The optimization module is configured to: train and optimize the model based on the completed detection model, including optimizing the detection box localization using the CIoU loss function;
[0034] The output module is configured to use the optimal weights obtained during training to perform detection on the test set input and generate the final detection results.
[0035] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the blueberry fruit focusing detection method based on a lightweight YOLO model.
[0036] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a blueberry fruit focusing detection method based on a lightweight YOLO model.
[0037] In summary, the present invention has the following beneficial technical effects:
[0038] 1. This invention constructs training data that closely resembles the actual blueberry planting scenario through geometric transformation, Gaussian noise addition, and gradient operator-based occlusion simulation. This effectively solves the problem of poor adaptability of traditional data augmentation methods to greenhouse environments, and significantly improves the detection stability of the model under different lighting, fruit posture, and occlusion conditions.
[0039] 2. This invention utilizes the Sobel operator to generate irregular masks that conform to the natural leaf occlusion characteristics, and combines the fruit contour gradient characteristics to simulate the real occlusion edge, breaking through the limitations of traditional rectangular occlusion simulation. This significantly improves the detection success rate of the model in fruit overlap or leaf occlusion scenarios, effectively reducing missed detections and false detections.
[0040] 3. This invention introduces the FasterNet network to replace the traditional module. By using partial channel convolution (PConv) and double convolution, it significantly reduces the amount of computation while retaining the ability to extract key features, solving the accuracy loss problem that is common in lightweight models and providing a feasible solution for real-time detection of edge devices.
[0041] 4. The C2f-Faster module of the present invention adopts a collaborative design of main branch and side branch. The main branch extracts the semantic features of the fruit step by step, while the side branch retains the spatial location information. Through feature fusion, the dual optimization of "semantic understanding and precise positioning" is achieved, which effectively improves the detection accuracy of overlapping fruits and small targets (such as fruit stems).
[0042] 5. This invention utilizes the DAttention mechanism to achieve feature focusing on occluded and low-light regions through deformable convolution and dynamic weight adjustment, solving the problem of insufficient response of traditional attention mechanisms in complex environments, and enabling the model to accurately identify targets even when the fruit is partially occluded or the lighting is uneven.
[0043] 6. This invention overcomes the limitations of traditional IoU loss by simultaneously optimizing the overlap of the detection box, the center point position, and the aspect ratio through the CIoU loss function. This significantly improves the matching degree between the detection box and the actual shape of the fruit, especially for irregularly shaped targets such as fruit stems, effectively reducing positioning errors and shape distortion. Attached Figure Description
[0044] Figure 1This is a schematic diagram of the overall structure of a blueberry fruit focusing detection method based on a lightweight YOLO model according to Embodiment 1 of the present invention.
[0045] Figure 2 This is a schematic diagram of the detection model structure in a blueberry fruit focusing detection method based on a lightweight YOLO model according to Embodiment 1 of the present invention.
[0046] Figure 3 This is a FasterNet network structure diagram in a blueberry fruit focusing detection method based on a lightweight YOLO model according to Embodiment 1 of the present invention.
[0047] Figure 4 This is a diagram of the FasterBlock structure in a blueberry fruit focusing detection method based on a lightweight YOLO model, according to Embodiment 1 of the present invention.
[0048] Figure 5 This is a C2f-Faster structure diagram in a blueberry fruit focusing detection method based on a lightweight YOLO model according to Embodiment 1 of the present invention. Detailed Implementation
[0049] The present invention will be further described in detail below with reference to the accompanying drawings.
[0050] Example 1
[0051] Reference Figure 1 This embodiment of a blueberry fruit focusing detection method based on a lightweight YOLO model includes:
[0052] Blueberry images were collected to form a multidimensional dataset, and the collected multidimensional dataset was labeled.
[0053] Data augmentation is performed based on the acquired multidimensional dataset, including geometric transformation of the acquired original dataset and optical and occlusion simulations by adding Gaussian noise and gradient operators.
[0054] The preprocessed cube was used as input to build the detection model, including the introduction of the FasterNet ultra-lightweight network into the Backbone layer of the YOLOv10n model;
[0055] A dynamic attention mechanism called DAttention is introduced on the basis of the detection model, including a dynamic attention mechanism of DAttention embedded after the PSA layer based on the improved Backbone layer.
[0056] The completed detection model is used for model training and optimization, including optimizing the detection box localization using the CIoU loss function;
[0057] Using the optimal weights obtained during training, the test set is input for detection, generating the final detection results.
[0058] Specifically, a blueberry fruit focusing detection method based on a lightweight YOLO model includes the following steps:
[0059] like Figure 1 As shown, S1, obtain blueberry images to form a multi-dimensional dataset, and label the obtained multi-dimensional dataset;
[0060] Images were collected in a greenhouse during three time periods: morning, noon, and evening. The variations in natural light at different times (e.g., oblique sunlight in the morning, direct sunlight at noon, and diffused light in the evening) simulated the light fluctuations in a real harvesting environment. By controlling the shooting angle of the imaging equipment (tilt angle ±15° to simulate the vertical grasping perspective of a robotic arm, horizontal rotation angle 0-360° to cover the circumference of the fruit), the dataset includes morphological features of the fruit from different angles, solving the problem of insufficient model generalization ability caused by the single perspective of traditional datasets. For example, when the fruit is obscured by leaves from above, adjusting the tilt angle can capture the side profile of the fruit, avoiding missed detections caused by a fixed perspective.
[0061] Next, the LabelImg tool was used to label the boundaries of the stem (including residual calyx structures) and the fruit body (the main spherical flesh), with the error controlled within 2 pixels to ensure that the labeled boxes closely matched the actual physical boundaries. This operation enabled the model to learn subtle morphological differences between the stem and the fruit body, such as the difference between the columnar structure of the stem and the spherical curvature of the fruit body. The blueberry dataset was then divided into training, validation, and test sets in a 7:2:1 ratio for model training and testing.
[0062] S2. Data augmentation is performed based on the acquired multidimensional dataset, including geometric transformation of the acquired original dataset, and optical and occlusion simulations are performed by adding Gaussian noise and gradient operators.
[0063] Geometric transformations were performed on a subset of blueberry image data. Random cropping was used to retain at least 60% of the fruit area, preventing the loss of key features such as the stem and body due to over-cropping. Specifically, the fruit's bounding box was first located, and then, using the center of the bounding box as a reference, a cropping area was randomly generated proportionally. The calculation formula is as follows:
[0064]
[0065] Among them, S 果实 S represents the area of the fruit in pixels within the labeled box. 裁剪后图像This represents the area of the cropped image. To ensure the fruit remains intact after cropping, for example, for a fruit with a diameter of 200 pixels, the fruit's pixel area in the cropped image must account for ≥60% to preserve the stem structure and texture features of the fruit. Then, the image is rotated ±15° to simulate the natural tilt of the fruit during growth, allowing the model to learn the target features under different postures. During rotation, bilinear interpolation is used to fill the edges, avoiding pixel distortion and ensuring a clear fruit outline after rotation.
[0066] Then Gaussian noise was added for optical simulation, using the formula:
[0067]
[0068] Where x represents the noise value, μ represents the noise mean, and σ represents the noise standard deviation. The noise distribution simulates the noise distribution of a camera sensor under different lighting conditions, and the noise distribution satisfies:
[0069] I 噪声 =I 原始 +N(0,σ 2 ),
[0070] Where N(0,σ) 2 ) represents a sequence with mean μ and variance σ. 2 Gaussian noise is used; for example, in low-light environments, the noise intensity increases, and the model needs to learn to extract fruit features from the noise. In strong-light environments, the noise intensity decreases, and the model focuses on fruit color differences, then uses a linear transformation to adjust the brightness. 调整 =I 原始 ×(1±0.2), covering the illumination span from cloudy days (brightness reduced by 20%) to sunny days (brightness increased by 20%).
[0071] Finally, an irregular polygonal mask is used to cover 10%-40% of the fruit area to simulate natural shading by leaves. The mask shape is generated by a gradient operator; in this embodiment, the Sobel operator is used, with a horizontal kernel G. x and vertical core G y The image edge gradient is calculated to generate a jagged contour similar to the edge of a leaf. The formula for calculating the image gradient is as follows:
[0072]
[0073] Among them, G x Represented as a horizontal kernel, G yRepresented as a vertical kernel, a mask boundary is generated along the high gradient region, making the occlusion edge closer to the real leaf occlusion effect. When the fruit stem is partially occluded, the model infers the position of the fruit stem through the boundary relationship between the fruit body and the occluder. The gradient abrupt change point at the edge of the fruit body can indicate the occlusion boundary. Combined with the spatial prior of the fruit stem and the fruit body, the occluded fruit stem region can be inferred.
[0074] S3. Use the preprocessed multidimensional dataset as input to build a detection model, including introducing the FasterNet ultra-lightweight network into the Backbone layer of the YOLOv10n model;
[0075] like Figure 2 , Figure 3 and Figure 4 As shown, in the Backbone layer of the YOLOv10n model, the FasterNet ultra-lightweight network is introduced to replace the original C2f module with the C2f-Faster module in this implementation. The core goal is to significantly reduce the computational load while maintaining detection accuracy. FasterNet adopts a "staged feature extraction" architecture, consisting of four stages of FasterNetBlock cascaded. Different types of downsampling layers are set before each stage, including embedding layers or merging layers. The embedding layer is located before the first stage and uses a 4×4 convolutional layer with a stride of 4 to quickly downsample the input feature map (e.g., 320×320 pixels) to 80×80 pixels, while expanding the number of channels from 64 to 128, achieving preliminary feature extraction of "spatial dimensionality reduction and channel dimensionality enhancement". The merging layer is located before the second, third and fourth stages and uses a 2×2 convolutional layer with a stride of 2. Each time, the feature map resolution is halved and the number of channels is doubled, gradually extracting more abstract semantic features.
[0076] Each FasterNetBlock contains three core layers: a 3×3 PConv layer (partial channel convolution), where PConv uses a subset of channels from the input feature map for feature extraction while keeping the number of other channels unchanged. This partial channel representation is C. P That is, 1 / 4 of channel C. The FLOPS of PConv are expressed as: Where h is the height of the feature map, w is the width of the feature map, k is the kernel size, and C P PConv performs convolution operations on only 1 / 4 of the input channels, retaining the remaining 3 / 4. For example, with 64 input channels, a traditional 3×3 convolution would require calculation on all 64 channels, while PConv only processes 16 channels, reducing computation by 75%. The PConv layer calculation formula is as follows:
[0077] X pconv =[Conv3×3(X part ),Xremain ],
[0078] Among them, X part ∈R h×w×C / 4 For partial channel input, X remain ∈R h×w×3C / 4 To preserve the channels, the number of output channels remains C, but the computation is reduced by 3 / 4. This design takes advantage of the redundancy of the characteristic channels of blueberries—the color features of the fruit body are mainly concentrated in the RGB channels, and the texture features of the stem are mainly concentrated in the gradient channels, so there is no need to calculate all channels.
[0079] 1×1 Conv layer (intermediate layer): Reduces the number of channels to half of the input channels (e.g., 64→32). Dimensionality reduction forces the network to focus on key features (e.g., the texture of the fur on the twig). It is then connected to a BatchNorm regularization layer and a ReLU activation function to enhance the diversity of features and the ability to express non-linearity.
[0080] 1×1 Conv layer (output layer): restores the number of channels to the original dimension (e.g., 32→64), completing the optimization and reorganization of features.
[0081] like Figure 5 As shown, the C2f-Faster module adopts a dual-branch parallel architecture, achieving feature complementarity through deep feature extraction in the main branch and preservation of positional information in the side branches. The main branch consists of three cascaded FasterNetBlocks, extracting fine-grained features of the fruit step by step in stages.
[0082] The first stage involves downsampling to 80×80 pixels through the embedding layer to extract the basic color features of the fruit body. The second stage involves downsampling to 40×40 pixels through the merging layer to extract the preliminary texture features of the fruit stem. The third stage involves downsampling to 20×20 pixels through the merging layer to extract the semantic features (such as category distinction) of the fruit stem and the fruit body. During each downsampling stage, the number of channels gradually increases from 64 to 512, which conforms to the principle of "shallow networks extract positional information and deep networks extract semantic information".
[0083] The bypass branch directly shorts the original feature map without performing any downsampling or feature transformation, thus fully preserving the spatial location information of the fruit (such as the coordinate offset of the bounding box) and avoiding the loss of localization accuracy caused by downsampling of the main branch. The features of the main branch and the bypass branch are merged through the Concat operation, and the deep features output by the main branch contain rich semantic information.
[0084] The shallow features output by the bypass branch contain precise location information. These two features are fused into a new feature map using channel concatenation (Concat), as shown in the formula:
[0085] F fusion =Concat(Fmain ,F bypass ),
[0086] Among them, F main F represents the main branch output feature. bypass Represented as a side branch feature, it possesses both "semantic understanding" and "precise localization" capabilities. For example, when detecting overlapping fruits, the main branch is responsible for identifying the fruit category, while the side branch provides the precise boundary coordinates of each fruit, solving the problem that traditional single-branch networks cannot simultaneously handle classification and localization.
[0087] S4. Introduce the DAttention dynamic attention mechanism on the basis of the detection model, including the DAttention dynamic attention mechanism embedded after the PSA layer based on the improved Backbone layer.
[0088] The DAttention dynamic attention mechanism is embedded after the PSA (partial self-attention) layer in the improved Backbone layer, forming a cascaded structure of feature extraction and attention enhancement. The Backbone layer first extracts multi-scale features through the C2f-Faster module, and the PSA layer performs partial self-attention processing on the features to capture local region dependencies. The DAttention mechanism further dynamically enhances key region features for complex scenes such as occlusion and low light. The mechanism achieves dynamic enhancement of blueberry fruit features in the following way, with the core calculation formula as follows:
[0089]
[0090] Where, x q Let K represent the query feature, and W represent the total number of sampling points. A Represented as the weighted prediction parameter, W Δ The offset prediction parameters are represented by Bilinear, which is the bilinear interpolation function, and Softmax is the normalization operation. First, W... A and W Δ For query feature x q A linear transformation is performed to generate the weight distribution and spatial offset of the sampling points. Then, the weights are normalized using Softmax, and the offsets are added to the reference point coordinates P. q The irregular sampling locations are obtained, and finally, features are extracted and weighted fused using bilinear interpolation for the irregular sampling points p. k = (u,v), and the coordinates of its four nearest surrounding pixels are (i,j), (i+1,j), (i,j+1), (i+1,j+1). Then the bilinear interpolation formula is:
[0091]
[0092] Where x(i+m,j+n) represents the coordinates (i+m,j+n) in the feature map x, |uim| and |vjn| represent the absolute distances between the sampling point's horizontal and vertical coordinates and its adjacent pixels, respectively. The weighting factor is weighted by distance, with closer pixels receiving higher weights, ensuring smooth and continuous feature extraction at irregular locations. This achieves feature enhancement for occluded or unevenly lit areas. The degree of occlusion is determined by calculating the gradient variance of the feature map, using the following formula:
[0093]
[0094] Among them, g i Represented as the gradient magnitude of a pixel, through the horizontal gradient G x and vertical gradient G y The square root of the sum of squares is used to determine the edge strength of a pixel. The mean gradient magnitude within a region is used to measure the overall gradient level. Var represents the gradient variance, describing the degree of gradient fluctuation within the region. When the gradient variance of a region exceeds a preset threshold (Var > 150, which is 150 in this embodiment), it is determined that the region is occluded (gradient abrupt changes caused by leaf edges or fruit overlap). The weights of sampling points in the occluded region are automatically increased by 1.5-2 times. For example, when the fruit is occluded by the upper left leaf, the offset generation module automatically calculates an offset of 5-10 pixels to the right and 3-5 pixels downward, so that the sampling points focus on the outline of the unoccluded fruit, guiding the model to focus on the features of the occluded edge. The sampling points of traditional convolution are fixed on a 3×3 grid, while DAttention uses W... Δ The generated offset allows sampling points to adaptively shift to key areas such as occlusion edges. Finally, the feature map is divided into 8×8 sub-regions, and the mean and standard deviation of brightness in each region are calculated.
[0095] Weight′ k =Weight k ×(1+α·I(L≤μ / 2)),
[0096] Where L is the brightness value within the sub-region, calculated using the RGB-to-grayscale formula, and I (L≤μ / 2) represents the indicator function. When the brightness L is less than half the mean, the function value is 1; otherwise, it is 0, meaning that only the weight of the dark region (low-light environment) is enhanced. α is an adjustment factor. The features of the dark region are extracted using bilinear interpolation before local contrast enhancement.
[0097]
[0098] in, It is expressed as the ratio of the current brightness to the regional mean, and β is the contrast adjustment factor, which enhances the characteristic difference between the fruit stem (green) and the fruit body (blue) under low light.
[0099] S5. Based on the completed detection model, perform model training and optimization, including optimizing the detection box localization using the CIoU loss function;
[0100] The CIoU loss function is used to optimize the detection box localization of the detection model after introducing the attention mechanism. In object detection, the traditional IoU loss only focuses on the overlap area between the predicted box and the ground truth box, which has problems such as inability to optimize in non-overlapping scenes and ignoring differences in box position and shape. The CIoU loss function achieves more accurate regression optimization of the detection box by integrating constraints of multiple dimensions. Its working principle is as follows: First, the IoU term, that is, the intersection-union ratio of the predicted box and the ground truth box, is calculated, and the formula is:
[0101] Where B represents the predicted bounding box output by the detection model, B gt The predicted bounding box represents the ground truth bounding box, and the area represents the area of the computational region, i.e., the number of overlapping pixels. IoU reflects the degree of overlap between the two boxes; a larger value indicates more overlap. However, when the predicted bounding box and the ground truth bounding box do not overlap, the IoU is 0. In this case, traditional IoU loss cannot provide optimization direction and therefore cannot guide the model to adjust the position of the predicted bounding box. CIoU loss, on the other hand, introduces a center point distance term to calculate this distance. Calculate the Euclidean distance between the center point of the predicted bounding box and the center point of the ground truth bounding box:
[0102]
[0103] Where (x,y) and (x gt ,y gt The coordinates of the center points of the predicted and ground truth bounding boxes are shown below, respectively. This distance reflects the positional deviation of the center points of the two boxes. To eliminate the influence of image scale on the distance, the square of the Euclidean distance is divided by the square of the diagonal length *c* of the smallest closure region containing both boxes, resulting in a normalized center point distance term. This way, regardless of the size of the detected target, the deviation of the center points can be measured at a uniform scale, thereby guiding the center point of the predicted box towards the center point of the ground truth bounding box and solving the problem of IoU loss being insensitive to box position. Even when the predicted and ground truth bounding boxes do not overlap and the IoU is 0, the loss value can still be calculated using the center point distance term, providing optimization direction for the model.
[0104] The diagonal length of the minimum closure region containing both the predicted and ground truth bounding boxes is specifically expressed as:
[0105] width = max(x, x gt )-min(x,x gt ),
[0106] height = max(y, ygt )-min(y,y gt ),
[0107]
[0108] Where x and x gt These represent the coordinates of the predicted bounding box and the ground truth bounding box on the horizontal axis (x-axis), respectively, and y and y'. gt ρ(b,b) represents the coordinates of the predicted bounding box and the ground truth bounding box on the vertical axis (y-axis), respectively. `width` and `height` represent the width and height of the minimum closure region, respectively. Distance normalization is applied to make the loss function unaffected by the target scale. For example, when the target is far away in the image (small scale), `c` is larger, and the impact of distance deviation is reduced; when the target is close (large scale), `c` is smaller, and the impact of distance deviation is amplified, thus ensuring optimization consistency across different scales. When the center point of the predicted bounding box deviates from the center of the ground truth bounding box, ρ(b,b) = 1. gt ) increases, leading to As the value of the loss function increases, the model will adjust the weights to move the predicted box closer to the true box. If multiple blueberries overlap, the closure region c will change dynamically according to the overlap range to ensure that the distance deviation under different overlap scenarios is reasonably normalized, avoiding over-optimization of the deviation of large-scale fruits or the neglect of the deviation of small-scale fruits.
[0109] Next, calculate the aspect ratio consistency term v, using the following formula:
[0110]
[0111] Among them, w gt h gt 'v' represents the width and height of the ground truth bounding box, while 'w' and 'h' represent the width and height of the predicted bounding box. This is the aspect ratio consistency term, which constrains the shape of the predicted bounding box to match the ground truth bounding box by calculating the angular difference between their aspect ratios. The aspect ratio is mapped to angular space using the arctangent function, and then the square of the angular difference is calculated, avoiding the potential numerical instability issues that can occur when directly using the aspect ratio. This consistency term measures the shape difference between the predicted and ground truth bounding boxes. The greater the difference in their aspect ratios, the larger the v value, and the greater the penalty, thus forcing the shape of the predicted bounding box to match the ground truth bounding box. This is particularly suitable for detecting non-rectangular targets such as blueberry stems.
[0112] Then, based on the above calculation results, using the formula:
[0113]
[0114] Where IoU represents the intersection-union ratio between the predicted bounding box and the ground truth bounding box. The value is represented as the center point distance, where b represents the center point coordinates of the predicted bounding box. gtRepresented as the center point coordinates of the true bounding box, ρ 2 Represented as the Euclidean distance function, c represents the diagonal length of the smallest closure region containing both the predicted and ground truth boxes, v represents the aspect ratio consistency term, and α represents the dynamic weight coefficient, which is given by the formula... It is determined that α will automatically adjust the strength of the aspect ratio constraint based on the overlap (IoU value) between the predicted bounding box and the ground truth bounding box. In low-overlap scenarios, which typically correspond to small targets or occlusion scenarios, ((1-IoU)) is relatively large, and α will automatically increase, increasing the weight of the aspect ratio consistency term v in the loss function, and the model prioritizes optimizing the shape of the predicted bounding box. In high-overlap scenarios, α approaches 1, and the center point distance term plays a dominant role in the loss function, realizing a two-stage optimization logic of first optimizing the shape and then accurately locating.
[0115] Finally, after obtaining the detection boxes, adaptive NMS processing is performed. First, the gradient variance Var of the detection box region is calculated to reflect the degree of occlusion. The NMS threshold is dynamically adjusted based on the gradient variance. When Var ≤ 150, it is considered a normal scene, and a lower NMS threshold of 0.4 is used to retain high-confidence, non-overlapping detection boxes. When Var > 150, it is considered an occluded scene, and the NMS threshold is increased to 0.6 to allow more overlapping boxes to be retained, preventing the accidental deletion of fruits obscured by leaves. Simultaneously, the confidence of the detection boxes is recalibrated based on the aspect ratio consistency term v, using the following formula:
[0116] Confidence level′ = confidence level × (1 + β × (1 - v)),
[0117] β is an adjustment factor. When the aspect ratio of the detection box is close to the true value, v approaches 0, and the confidence level increases. When the aspect ratio deviates significantly from the true value, v approaches 1, and the confidence level decreases. This reduces false detections caused by shape distortion and ultimately yields more accurate detection results.
[0118] Example 2
[0119] S6. Using the optimal weights obtained from training, input the test set for detection to generate the final detection results.
[0120] In training a blueberry fruit detection model, obtaining and applying the optimal weights is crucial for ensuring detection accuracy and efficiency. The core process revolves around an early stopping mechanism, a cosine decay learning rate strategy, optimizer configuration, and mixed-precision training. The specific working principle is as follows:
[0121] An early stopping mechanism is employed to select the optimal weights. During model training, the mean accuracy (mAP) of the validation set is continuously monitored. When the mAP of the validation set increases by less than a preset threshold for five consecutive rounds (0.1% in this embodiment), it indicates that the model performance has stabilized and continued training may lead to overfitting. At this point, the system automatically saves the weight parameters of the current detection model and identifies them as the optimal weights. This mechanism avoids overtraining by dynamically evaluating the model's performance on the validation set and effectively improves the model's generalization ability on the test set.
[0122] The learning rate is dynamically adjusted using a cosine decay learning rate strategy. This strategy is based on the current training epoch and the preset maximum training epoch, using the formula:
[0123]
[0124] Here, lr0 represents the initial learning rate, epoch represents the current training epoch, and max_epoch represents the maximum training epoch. The initial learning rate determines the step size for parameter updates in the early stages of training. In the early stages (the first 100 epochs), the learning rate is at a relatively high level, and the model quickly learns the basic features of blueberries, such as color and overall shape, with larger step sizes. As training progresses, the learning rate gradually decreases according to a cosine function curve. In the last 200 epochs, the learning rate decreases, and the model finely adjusts the parameters with smaller step sizes, focusing on the extraction of detailed features of small targets such as the stem. This dynamic adjustment method enables the model to converge quickly in the early stages of training and avoid getting stuck in local optima in the later stages, thus improving detection accuracy.
[0125] Then, the weights were further optimized by combining the SGD optimizer with AMP (Automatic Mixed Precision) training. For the optimizer configuration, the Stochastic Gradient Descent (SGD) algorithm was used, with a momentum of 0.937 and a weight decay of 5e-4. The introduction of the momentum term gives gradient updates inertia, allowing them to skip shallow local optima and guide the model to find better weight combinations when optimizing complex tasks such as fruit stem and body boundary features. Weight decay, through parameter regularization, prevents overfitting. Simultaneously, AMP training technology was applied, converting floating-point operations from FP32 to FP16, reducing memory usage by 50% while maintaining computational accuracy. This allows the model to use a larger batch size (e.g., increasing from 8 to 16), accelerating the training process, while a dynamic loss scaling mechanism ensures numerical stability, maintaining detection accuracy comparable to FP32 training. Finally, the optimal weights obtained from training were applied to the test set detection. The blueberry images in the test set are input into the detection model with optimal weights. The model identifies and locates the fruits in the images based on the learned parameters, and outputs the final detection results containing information such as the location of the detection box and the category confidence. Through the above complete weight optimization and application process, the model achieves high-precision and high-efficiency target detection performance in the blueberry fruit detection task, providing reliable support for practical applications.
[0126] Example 2
[0127] The difference between this embodiment and Embodiment 1 is that this embodiment provides a blueberry fruit focusing detection system based on a lightweight YOLO model, including:
[0128] The data acquisition module is configured to: acquire blueberry images to form a multi-dimensional dataset, and label the acquired multi-dimensional dataset;
[0129] The preprocessing module is configured to perform data augmentation based on the acquired cube, including geometric transformation of the acquired original dataset and optical and occlusion simulations by adding Gaussian noise and gradient operators.
[0130] The model module is configured to: build a detection model by taking the preprocessed cube as input, including introducing the FasterNet ultralight network into the backbone layer of the YOLOv10n model;
[0131] The attention module is configured to introduce a DAttention dynamic attention mechanism on the basis of the detection model, including a DAttention dynamic attention mechanism embedded after the PSA layer based on the improved Backbone layer.
[0132] The optimization module is configured to: train and optimize the model based on the completed detection model, including optimizing the detection box localization using the CIoU loss function;
[0133] The output module is configured to use the optimal weights obtained during training to perform detection on the test set input and generate the final detection results.
[0134] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the blueberry fruit focusing detection method based on a lightweight YOLO model.
[0135] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a blueberry fruit focusing detection method based on a lightweight YOLO model.
[0136] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A blueberry fruit focusing detection method based on a lightweight YOLO model, characterized in that, include: Blueberry images were collected to form a multidimensional dataset, and the collected multidimensional dataset was labeled. Data augmentation is performed based on the acquired multidimensional dataset, including geometric transformation of the acquired original dataset and optical and occlusion simulations by adding Gaussian noise and gradient operators. The preprocessed cube was used as input to build the detection model, including the introduction of the FasterNet ultra-lightweight network into the Backbone layer of the YOLOv10n model; A dynamic attention mechanism called DAttention is introduced on the basis of the detection model, including a dynamic attention mechanism of DAttention embedded after the PSA layer based on the improved Backbone layer. The completed detection model is used for model training and optimization, including optimizing the detection box localization using the CIoU loss function; Using the optimal weights obtained during training, the test set is input for detection, generating the final detection results.
2. The blueberry fruit focusing detection method based on a lightweight YOLO model according to claim 1, characterized in that, The optical and occlusion simulations performed by adding Gaussian noise and gradient operators include adding zero-mean Gaussian noise using a probability density function, adjusting brightness through linear transformation, covering part of the fruit area with an irregular polygon mask, calculating the image gradient using the Sobel operator, generating mask edges along high-gradient regions, simulating leaf occlusion contours, and thus completing data augmentation. The image gradient calculation formula is as follows: Among them, G x Represented as a horizontal kernel, G y It is represented as a vertical core.
3. The blueberry fruit focusing detection method based on a lightweight YOLO model according to claim 1, characterized in that, The step of using the preprocessed multidimensional dataset as input to construct a detection model includes replacing the C2f module of the YOLOv10n model with a C2f-Faster module using the FasterNet ultra-lightweight network. The FasterNet ultra-lightweight network consists of multiple stages of FasterNet modules, with an embedding layer set before each stage for spatial downsampling and channel expansion. The FasterNet module has a built-in PConv layer, an intermediate Conv layer, and an output Conv layer, and a BatchNorm regularization layer and a ReLU activation function are connected after the intermediate Conv layer.
4. The blueberry fruit focusing detection method based on a lightweight YOLO model according to claim 3, characterized in that, The C2f-Faster module uses a main branch and a side branch for feature fusion. The main branch extracts fine-grained features of the fruit step by step through three cascaded FasterNet modules. At each stage, downsampling is achieved through an embedding layer. The side branch directly shorts the original feature map to retain the spatial location information of the fruit. Finally, the features of the main branch and the side branch are merged through a Concat operation.
5. The blueberry fruit focusing detection method based on a lightweight YOLO model according to claim 4, characterized in that, The proposed dynamic attention mechanism, DAttention, is embedded after the PSA layer of the improved Backbone layer. It involves linearly transforming the query features using weight prediction parameters and offset prediction parameters to generate the weight distribution and spatial offset of the sampling points. Then, it normalizes the weights using Softmax and adds the offsets to the reference point coordinates to obtain irregular sampling positions. Finally, it extracts features through bilinear interpolation and performs weighted fusion. The calculation formula for the DAttention dynamic attention mechanism is as follows: Where, x q Let K represent the query feature, and W represent the total number of sampling points. A Represented as the weighted prediction parameter, W Δ The offset prediction parameters are represented by Bilinear, which represents the bilinear interpolation function, and Softmax represents the normalization operation.
6. The blueberry fruit focusing detection method based on a lightweight YOLO model according to claim 1, characterized in that, The optimization of detection box localization using the CIoU loss function includes calculating the Euclidean distance between the center point of the predicted box and the center point of the ground truth box, normalizing the Euclidean distance to the diagonal length of the smallest closed region containing both boxes, and obtaining the center point distance term. Then, by calculating the angle difference between the aspect ratios of the predicted and ground truth boxes, the shape of the predicted box is constrained to match the ground truth box. Based on the center point distance term and aspect ratio, the CIoU loss function is calculated to obtain the optimized detection box. Adaptive NMS processing is then applied to detection boxes with confidence scores greater than a confidence threshold. The CIoU loss function calculation formula is as follows: Where IoU represents the intersection-union ratio between the predicted bounding box and the ground truth bounding box. The value is represented as the center point distance, where b represents the coordinates of the center point of the predicted bounding box. gt Represented as the center point coordinates of the true bounding box, ρ 2 It is represented by the Euclidean distance function, c represents the diagonal length of the smallest closure region containing the predicted box and the ground truth box, v represents the aspect ratio consistency term, and α represents the dynamic weight coefficient.
7. The blueberry fruit focusing detection method based on a lightweight YOLO model according to claim 1, characterized in that, The optimal weights obtained through training include employing an early stopping mechanism. When the average precision improvement on the validation set is less than a set threshold, the current detection model weights are saved as the optimal weights. Then, a cosine decay learning rate strategy is used to dynamically adjust the learning rate. This is combined with SGD optimizer and AMP hybrid precision training, using a momentum term to skip shallow local optima and find a better weight combination. The formula for the cosine decay learning rate strategy is: Where lr0 represents the initial learning rate, epoch represents the current training epoch, and max_epoch represents the maximum training epoch.
8. A blueberry fruit focusing detection system based on a lightweight YOLO model, comprising the method described in claim 1, characterized in that, include: The data acquisition module is configured to: acquire blueberry images to form a multi-dimensional dataset, and label the acquired multi-dimensional dataset; The preprocessing module is configured to perform data augmentation based on the acquired cube, including geometric transformation of the acquired original dataset and optical and occlusion simulations by adding Gaussian noise and gradient operators. The model module is configured to: build a detection model by taking the preprocessed cube as input, including introducing the FasterNet ultralight network into the backbone layer of the YOLOv10n model; The attention module is configured to introduce a DAttention dynamic attention mechanism on the basis of the detection model, including a DAttention dynamic attention mechanism embedded after the PSA layer based on the improved Backbone layer. The optimization module is configured to: train and optimize the model based on the completed detection model, including optimizing the detection box localization using the CIoU loss function; The output module is configured to use the optimal weights obtained during training to perform detection on the test set input and generate the final detection results.
9. A computer-readable storage medium storing a plurality of instructions, characterized in that, The instructions are adapted to be loaded and executed by the processor of the terminal device as described in claim 1, which is a blueberry fruit focusing detection method based on a lightweight YOLO model.
10. A terminal device, comprising a processor and a computer-readable storage medium, wherein the processor is configured to implement instructions; and the computer-readable storage medium is configured to store multiple instructions, characterized in that, The instructions are adapted to be loaded by a processor and executed as described in claim 1, a blueberry fruit focusing detection method based on a lightweight YOLO model.
Citation Information
Cited By
Image processing-based hose coupler assembly deformation detection method
CN121685400A