Improved YOLOv8 network model-based flat peach fruit identification method

通过改进YOLOv8模型的多模态融合和时空同步注意力机制,解决了复杂农业环境下的果实检测精度和实时性问题,实现了高效果实识别。

CN120299030APending Publication Date: 2025-07-11SHANDONG AGRICULTURAL UNIVERSITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510355469.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing YOLOv8 model has problems such as single feature extraction, high model redundancy, and rigid multi-task optimization in complex agricultural scenarios, which are difficult to adapt to the multi-scale feature interaction and lighting changes in the orchard environment, resulting in insufficient fruit detection accuracy and real-time performance.

Method used

Build an improved multimodal fusion YOLOv8 model, combine multimodal dynamic fusion network, space-time synchronization attention mechanism and dynamic path selection, and optimize the feature extraction and recognition process through light compensation and three-dimensional spatial deformation enhancement technology.

Benefits of technology

It improves the fruit detection accuracy and recognition efficiency in complex environments, enhances the robustness and real-timeness of the model, reduces the number of model parameters, and improves work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299030A_ABST
    Figure CN120299030A_ABST
Patent Text Reader

Abstract

The invention provides a flat peach fruit identification method based on an improved YOLOv8 network model. The method comprises the following steps: constructing a multi-modal fusion model; performing light elimination and shadow restoration on the flat peach fruit image through an illumination compensation algorithm to obtain an extinction image; performing three-dimensional space deformation enhancement on the extinction image to obtain a space enhanced image; performing dynamic path selection on the spatial enhanced image through the hybrid perception backbone network, and extracting local detail features to obtain a feature enhanced image; performing space-time collaborative optimization on the feature enhanced image to obtain a space-time joint attention map; performing dynamic information interaction on the space-time joint attention map to obtain a preliminary result; and performing loss optimization on the preliminary result through a multi-task joint loss function to obtain a fruit identification result. According to the method, through the multi-modal dynamic fusion network, the space-time synchronization attention mechanism and the dynamic path selection, the detection precision and the recognition efficiency are improved, and the robustness of the model is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and particularly to a method for identifying flat peach fruits based on an improved YOLOv8 network model. Background Art

[0002] With the increasing shortage of agricultural labor and the growing demand for large-scale cultivation, automated harvesting technology has become the core direction of modern agricultural development. As a fruit with high economic value, the harvesting operation of flat peaches has long relied on manual labor, suffering from problems such as low efficiency, high cost, and labor intensity. In recent years, object detection algorithms based on deep learning (such as the YOLO series) have been gradually applied in agricultural scenarios. However, the particularity of the orchard environment poses severe challenges to the algorithms: fruits often have blurred features due to foliage occlusion, bag reflection, and dense distribution; under natural light conditions, direct light, backlight, and side light alternate, affecting image quality and feature stability; devices such as mobile robots and drones are limited by computing power and power consumption, and lightweight models are required to balance detection accuracy and real-time performance. Although the existing YOLOv8 model has improved multi-scale detection capabilities through the C2f module and SPPF structure, it still has limitations such as single feature extraction, high model redundancy, and rigid multi-task optimization in complex agricultural scenarios.

[0003] The main deficiencies are reflected in the following aspects: The traditional bidirectional feature pyramid (Bi-FPN) mostly uses fixed-weight fusion and cannot adaptively interact with multi-scale features in complex environments, resulting in a decline in the detection performance of small or occluded targets; existing light compensation techniques rely on empirical parameter adjustment (such as histogram equalization) and are difficult to achieve pixel-level light reconstruction, with poor fruit texture restoration effects in backlight or shadow areas; existing models cannot dynamically adjust the calculation path according to device computing power during deployment and are difficult to balance accuracy and real-time performance. Therefore, it is very necessary to design a method for identifying flat peach fruits based on an improved YOLOv8 network model. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for identifying flat peach fruits based on an improved YOLOv8 network model, which realizes fruit recognition in complex environments through a multi-modal dynamic fusion network, a spatio-temporal synchronous attention mechanism, and dynamic path selection operations, so as to improve detection accuracy and recognition efficiency, and enhance the robustness and real-time performance of the model.

[0005] To achieve the above object, the present invention provides the following solution:

[0006] A method for identifying flat peach fruits based on an improved YOLOv8 network model, comprising the following steps:

[0007] Construct an improved multi-modal fusion YOLOv8 model; the multi-modal fusion YOLOv8 model includes: a hybrid perception backbone network, a multi-scale interaction fusion module, and an adaptive dynamic detection head;

[0008] Perform light elimination and shadow repair operations on the flat peach fruit image through a light compensation algorithm to obtain a light-eliminated image;

[0009] Perform three-dimensional spatial deformation enhancement operations on the light-eliminated image to obtain a spatially enhanced image;

[0010] Perform dynamic path selection on the spatially enhanced image through the hybrid perception backbone network and extract local detail features to obtain a feature-enhanced image;

[0011] Perform spatio-temporal collaborative optimization on the feature-enhanced image through the bidirectional feature pyramid and spatio-temporal synchronous attention mechanism of the multi-scale interaction fusion module to obtain a spatio-temporal joint attention map;

[0012] Perform dynamic information interaction on the spatio-temporal joint attention map through the adaptive dynamic detection head to obtain a preliminary result;

[0013] Perform loss optimization on the preliminary result through a multi-task joint loss function and dynamically adjust the weights of the multi-task joint loss function in real time to obtain the fruit recognition result.

[0014] Optionally, the hybrid perception backbone network includes: a lightweight convolution branch, a global attention branch, and a deformable feature enhancement branch; the adaptive dynamic detection head includes: an object localization unit, a confidence prediction unit, a category discrimination unit, and a gated feature recombination unit.

[0015] Optionally, performing light elimination and shadow repair operations on the flat peach fruit image through a light compensation algorithm to obtain a light-eliminated image includes:

[0016] Obtain a global illumination distribution map based on the low-frequency illumination characteristics of the flat peach fruit image and the physical properties of direct light and diffuse light;

[0017] Perform dynamic brightness mapping on the illumination distribution map through an S-shaped curve compression function and adaptive gamma correction to obtain a light compensation result;

[0018] Perform shadow area segmentation and texture reconstruction operations on the global illumination distribution map to obtain a shadow repair result;

[0019] Perform weight fusion and edge sharpening on the light compensation result and the shadow repair result to obtain a light-eliminated image.

[0020] Optionally, performing three-dimensional spatial deformation enhancement operations on the light-eliminated image to obtain a spatially enhanced image includes:

[0021] Generate a pseudo-depth map of the extinction image through a monocular depth estimation network, and construct a three-dimensional point cloud space model based on the pseudo-depth map;

[0022] Simulate multi-angle shooting in the three-dimensional point cloud space model through a random perspective projection algorithm to generate two-dimensional projection images under virtual perspectives;

[0023] Perform non-rigid deformation compensation on the two-dimensional projection images to obtain spatially enhanced images.

[0024] Optionally, perform dynamic path selection on the spatially enhanced images through a hybrid perception backbone network and extract local detail features to obtain feature-enhanced images, including:

[0025] Extract local detail features of the spatially enhanced images through lightweight convolutional layers built into the hybrid perception backbone network; The local detail features include: texture complexity, illumination uniformity, and occlusion density;

[0026] Select branch activation paths according to the local detail features to obtain branch output results;

[0027] Generate dynamic branch weights through the Sigmoid function, and perform weighted aggregation on the branch output results according to the dynamic branch weights to obtain feature-enhanced images.

[0028] Optionally, the calculation formula for texture complexity is: TC = 0.6·C n + 0.4·E n ; where TC is the quantization value of texture complexity, C n is the contrast of the local region of the image, and E n is the entropy of the local region of the image;

[0029] The calculation formula for illumination uniformity is: IU = 1 - (0.5·MSE n + 0.5·G n ); where IU is the quantization value of illumination uniformity, MSE n is the mean square error of the luminance channel, and G n is the local luminance gradient.

[0030] Optionally, perform spatio-temporal collaborative optimization on the feature-enhanced images through the bidirectional feature pyramid and spatio-temporal synchronous attention mechanism of the multi-scale interaction fusion module to obtain a spatio-temporal joint attention map, including:

[0031] Extract multi-level feature maps of the feature-enhanced images through the hybrid perception backbone network; The multi-level feature maps include: low-level feature maps and high-level feature maps;

[0032] After bilinearly upsampling the high-level feature maps, add them element-wise to the low-level feature maps to obtain forward feature maps;

[0033] The low-level feature map is downsampled by a 3×3 depthwise separable convolution and then concatenated with the high-level feature map to obtain an inverse feature map;

[0034] The forward feature map and the inverse feature map are weighted and fused through the Softmax function to obtain a bidirectional feature map;

[0035] The bidirectional feature map is subjected to deformable convolution through the offset field generated by a 3×3 convolutional layer, and combined with an average pooling operation to obtain a spatial attention map;

[0036] The consecutive frames in the bidirectional feature map are temporally aligned and then input into an LSTM network to obtain a temporal correlation degree;

[0037] The spatial attention map and the temporal correlation degree are subjected to a Hadamard product operation to obtain a spatio-temporal joint attention map.

[0038] Optionally, the spatio-temporal joint attention map is subjected to dynamic information interaction through an adaptive dynamic detection head to obtain preliminary results, including:

[0039] The spatio-temporal joint attention map is convolutionally compressed to obtain a normalized feature map;

[0040] The localization information, confidence score, and class probability of the normalized feature map are obtained through a target localization unit, a confidence prediction unit, and a class discrimination unit respectively;

[0041] Based on the confidence score, a gating weight is obtained through the Sigmoid function;

[0042] A feature recombination scheme is determined based on the comparison result between the confidence score and the confidence threshold to obtain a recombined feature;

[0043] The recombined feature and the class probability are concatenated according to the spatial position to obtain preliminary results.

[0044] Optionally, a feature recombination scheme is determined based on the comparison result between the confidence score and the confidence threshold to obtain a recombined feature, including:

[0045] When the confidence score < 0.3, based on the gating weight, the localization information is error-compensated by the class probability to obtain a calibrated feature, and the normalized feature map is replaced with the calibrated feature to recalculate the confidence score;

[0046] When the confidence score ≥ 0.3, the localization information and the confidence score are weighted and fused through the gating weight to obtain a recombined feature.

[0047] Optionally, the multi-task joint loss function includes: a localization loss function, a confidence loss function, and a classification loss function; the weights of the multi-task joint loss function are obtained through gradient magnitude calculation.

[0048] According to the specific embodiments provided by the present invention, the following technical effects are disclosed: The present invention provides a flat peach fruit recognition method based on an improved YOLOv8 network model, and the method includes: constructing an improved multi-modal fusion YOLOv8 model; performing light elimination and shadow repair operations on flat peach fruit images through a light compensation algorithm to obtain a light-eliminated image; performing three-dimensional spatial deformation enhancement operations on the light-eliminated image to obtain a spatially enhanced image; performing dynamic path selection on the spatially enhanced image through a hybrid perception backbone network and extracting local detail features to obtain a feature-enhanced image; performing spatio-temporal collaborative optimization on the feature-enhanced image through the bidirectional feature pyramid and spatio-temporal synchronous attention mechanism of a multi-scale interaction fusion module to obtain a spatio-temporal joint attention map; performing dynamic information interaction on the spatio-temporal joint attention map through an adaptive dynamic detection head to obtain a preliminary result; performing loss optimization on the preliminary result through a multi-task joint loss function and dynamically adjusting the weights of the multi-task joint loss function in real time to obtain a fruit recognition result. This method improves the detection accuracy in complex environments by constructing a multi-modal dynamic fusion network and a spatio-temporal synchronous attention mechanism; reduces the model parameter quantity and improves the working efficiency through dynamic path selection; and significantly improves the accuracy, robustness, and real-time performance of fruit recognition through the light compensation algorithm and three-dimensional spatial deformation enhancement. Description of the Drawings

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a flow chart of the flat peach fruit recognition method of the present invention;

[0051] Figure 2 It is a flow chart of the light elimination and shadow repair operations of the present invention;

[0052] Figure 3 It is a flow chart of the three-dimensional spatial deformation enhancement operations of the present invention;

[0053] Figure 4 It is a flow chart of the dynamic path selection of the present invention;

[0054] Figure 5 It is a flow chart of the spatio-temporal collaborative optimization of the present invention;

[0055] Figure 6 It is a flow chart of the dynamic information interaction of the present invention. Detailed Embodiments

[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0057] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0058] As Figure 1 shown, the present invention provides a flat peach fruit recognition method based on an improved YOLOv8 network model, including the following steps:

[0059] Step 100: Construct an improved multi-modal fusion YOLOv8 model; the multi-modal fusion YOLOv8 model includes: a hybrid perception backbone network, a multi-scale interaction fusion module, and an adaptive dynamic detection head;

[0060] Specifically, the hybrid perception backbone network includes: a lightweight convolution branch, a global attention branch, and a deformable feature enhancement branch; the adaptive dynamic detection head includes: a target localization unit, a confidence prediction unit, a category discrimination unit, and a gated feature recombination unit.

[0061] Step 200: Perform light elimination and shadow repair operations on the flat peach fruit image through a light compensation algorithm to obtain a light-eliminated image; the specific steps are as Figure 2 shown, including:

[0062] Step 201: Obtain a global illumination distribution map according to the low-frequency illumination characteristics of the flat peach fruit image and the physical properties of direct light and diffuse light;

[0063] Specifically, the low-frequency illumination characteristics (such as large-area shadows and highlight regions) are extracted through a ResNet-50 network. The physical properties include: reflectivity and incident angle. A two-dimensional illumination intensity matrix is generated according to the low-frequency illumination characteristics and physical properties, and the illumination weight of each pixel in the matrix is marked, where 0 represents a full shadow and 1 represents an overexposed area.

[0064] Step 202: Perform dynamic brightness mapping on the illumination distribution map through an S-shaped curve compression function and adaptive gamma correction to obtain a light compensation result;

[0065] Specifically, the S-shaped curve compression function is used to adjust the brightness of the overexposed area, and the expression of the function is: where I out is the adjusted brightness, and I inLet \(I\) be the image brightness, \(\theta\) be the brightness threshold, which is 0.7 in this embodiment, and \(k\) be the curve steepness coefficient, which is 10 in this embodiment. Adaptive gamma correction is performed on the shadow area with illumination weight \(< 0.3\), and the expression is: where \(q\) is the illumination weight. Adaptive histogram equalization is used in the illumination transition area with illumination weight from 0.3 to 0.8 to improve the detail visibility of this area.

[0066] Step 203: Perform shadow area segmentation and texture reconstruction operations on the global illumination distribution map to obtain the shadow repair result;

[0067] Specifically, the threshold segmentation method is used to obtain the shadow mask of the global illumination distribution map, and morphological closing operation is performed with a \(5\times5\) elliptical kernel to fill the holes in the shadow mask, so as to ensure continuity.

[0068] Step 204: Perform weight fusion and edge sharpening on the illumination compensation result and the shadow repair result to obtain the extinction image.

[0069] Specifically, through the calculation formula \(I\) f \(=(1 - M)\cdot I\) c \(+M\cdot I\) r weight fusion is performed, where \(I\) f is the fusion result, \(M\) is the binary mask of the shadow area, \(I\) c is the illumination compensation result, and \(I\) r is the shadow repair result; then the unsharp masking algorithm is used to enhance the contour clarity of the fusion result, so as to obtain the extinction image. The expression of the unsharp masking algorithm is: \(I\) s \(=I\) f \(+\lambda_1\cdot I\) f \(-G\) σ \(I\) f ; where \(I\) s is the extinction image, \(\lambda_1\) is the sharpening intensity coefficient, which is 0.8 in the embodiment, and \(G\) σ is the Gaussian blur kernel.

[0070] Step 300: Perform three-dimensional space deformation enhancement operation on the extinction image to obtain the space enhancement image; the specific steps are as Figure 3 shown, including:

[0071] Step 301: Generate the pseudo-depth map of the extinction image through the monocular depth estimation network, and construct a three-dimensional point cloud space model according to the pseudo-depth map;

[0072] Specifically, by using the lightweight depth estimation network DepthFormer-S, the extinction image is first normalized, and then the relative depth map is output through forward propagation. Next, the relative depth map is converted into absolute depth values to obtain a pseudo-depth map. Finally, based on the camera internal parameters of the flat peach fruit image, the pixel coordinates are converted into three-dimensional points, thereby obtaining a three-dimensional point cloud space model.

[0073] Step 302: Simulate multi-angle shooting in the three-dimensional point cloud space model through the random perspective projection algorithm to generate two-dimensional projection images under virtual perspectives;

[0074] Specifically, according to the horizontal rotation angle, pitch angle, and translation vector during the simulation of multi-angle shooting, a rotation matrix and a translation matrix are respectively constructed, and the rotation matrix and the translation matrix are concatenated into an external parameter matrix; perform rigid body transformation on each three-dimensional point in the external parameter matrix, and project the transformed point cloud onto a virtual plane. Then, use bilinear interpolation to fill the pixel vacancies after projection to obtain a two-dimensional projection image.

[0075] Step 303: Perform non-rigid deformation compensation on the two-dimensional projection image to obtain a spatially enhanced image.

[0076] Specifically, in this embodiment, the two-dimensional projection image is subjected to non-rigid deformation compensation through a U-Net structure with an output channel number of 2 and a smoothness loss function.

[0077] Step 400: Perform dynamic path selection on the spatially enhanced image through the hybrid perception backbone network and extract local detail features to obtain a feature-enhanced image; the specific steps are as Figure 4 shown, including:

[0078] Step 401: Extract local detail features of the spatially enhanced image through the lightweight convolutional layer built in the hybrid perception backbone network; the local detail features include: texture complexity, illumination uniformity, and occlusion density;

[0079] Specifically, the calculation formula for texture complexity is:

[0080] TC = 0.6·C n + 0.4·E n ;

[0081] where TC is the quantization value of texture complexity, C n is the contrast of the local area of the image, and E n is the entropy of the local area of the image; the calculation formulas for C n and E n are respectively:

[0082] C n = ∑ i,j |i - j| 2·P(i,j);

[0083] E n =-∑ i,j P(i,j)·logP(i,j);

[0084] where P(i,j) is the joint probability distribution of pixel gray values for pixel coordinates (i,j).

[0085] The calculation formula for illumination uniformity is:

[0086] IU = 1 - 0.5·MSE n + 0.5·G n );

[0087]

[0088] where IU is the quantization value of illumination uniformity, MSE n is the mean square error of the luminance channel, G n is the local luminance gradient, μ L is the luminance mean, L i is the luminance of the luminance channel, is the gradient magnitude calculated by the Sobel operator.

[0089] The occlusion density is the ratio of the number of occluded pixels to the total number of pixels in the image. When the occlusion density is greater than 0.5, it means the fruit is severely occluded.

[0090] Step 402: Select the branch activation path according to the local detail features to obtain the branch output result;

[0091] Specifically, map the three indicators of texture complexity TC, illumination uniformity IU, and occlusion density OD to a three-dimensional space, with each dimension corresponding to an indicator, forming a decision cube. And divide the cube into 8 sub-regions, each sub-region corresponding to a branch activation strategy, which are: not activating the lightweight convolutional branch LCB, the global attention branch GAB, and the deformable feature enhancement branch DFEB (A0), only activating LCB (A1), only activating GAB (A2), only activating DFEB (A3), activating LCB and GAB (A4), activating LCB and DFEB (A5), activating GAB and DFEB (A6), and activating LCB, GAB, and DFEB (A7).

[0092] More specifically, the selection of the activation path is shown in Table 1:

[0093] Table 1 Activation Path Selection Table

[0094] TC IU OD Activation Strategy High Low High A5 Low Low Low A4 High High High A7 High High Low A3 High Low Low A4 Low High High A6 Low High Low A1 Low Low High A6

[0095] When TC > 0.7, it is determined as high; otherwise, it is low. When IU > 0.3, it is determined as high; otherwise, it is low. When OD > 0.5, it is determined as high; otherwise, it is low.

[0096] Step 403: Generate dynamic branch weights through the Sigmoid function, and perform weighted aggregation on the branch output results according to the dynamic branch weights to obtain a feature-enhanced image.

[0097] Step 500: Perform spatio-temporal collaborative optimization on the feature-enhanced image through the bidirectional feature pyramid and spatio-temporal synchronization attention mechanism of the multi-scale interaction fusion module to obtain a spatio-temporal joint attention map; the specific steps are as Figure 5 shown, including:

[0098] Step 501: Extract multi-level feature maps of the feature-enhanced image through the hybrid perception backbone network; the multi-level feature maps include: low-level feature maps and high-level feature maps;

[0099] Step 502: After bilinearly upsampling the high-level feature maps, add them element-wise to the low-level feature maps to obtain a forward feature map;

[0100] Step 503: After downsampling the low-level feature maps through 3×3 depthwise separable convolutions, concatenate them with the high-level feature maps to obtain a reverse feature map;

[0101] Step 504: Perform weighted fusion on the forward feature map and the reverse feature map through the Softmax function to obtain a bidirectional feature map;

[0102] Step 505: Perform deformable convolution on the bidirectional feature map through the offset field generated by the 3×3 convolutional layer, and combine it with the average pooling operation to obtain a spatial attention map;

[0103] Step 506: Align consecutive frames in the bidirectional feature map and input them into the LSTM network to obtain a temporal correlation;

[0104] Step 507: Perform Hadamard product operation on the spatial attention map and the temporal correlation to obtain a spatio-temporal joint attention map.

[0105] It should be noted that through multi-layer deep convolution operations, the problems of occlusion and motion blur in the process of capturing fruit deformation features are solved. The weighted fusion strategy of the bidirectional feature pyramid balances the contributions of multi-scale features and avoids information redundancy. The spatio-temporal joint attention mechanism of deformable convolution significantly improves the recall rate of occluded fruits and also reduces the false detection rate.

[0106] Step 600: Perform dynamic information interaction on the spatio-temporal joint attention map through the adaptive dynamic detection head to obtain a preliminary result; the specific steps are as Figure 6 shown, including:

[0107] Step 601: Convolve and compress the spatio-temporal joint attention map to obtain a normalized feature map;

[0108] Specifically, in this embodiment, the spatio-temporal joint attention map is convolved and compressed to 256 dimensions.

[0109] Step 602: Obtain the localization information, confidence score, and class probability of the normalized feature map through the target localization unit, confidence prediction unit, and class discrimination unit respectively;

[0110] Specifically, the target localization unit consists of 3 layers of deformable convolutions, which are used to output the bounding box offset of each pixel (i.e., the localization information). The confidence prediction unit adopts a double-branch structure. The main branch extracts the global confidence feature through global average pooling, and then extracts the local details through the 3×3 convolution of the auxiliary branch, and finally outputs the confidence score. The class discrimination unit interacts the normalized feature map with the pre-trained class prototype vector based on the cross-attention mechanism and outputs the distribution of class probabilities.

[0111] Step 603: Based on the confidence score, obtain the gating weight through the Sigmoid function;

[0112] Specifically, the confidence score is processed through a network containing two fully connected layers and used as the variable of the Sigmoid function to obtain the gating weight.

[0113] Step 604: Determine the feature recombination scheme based on the comparison result of the confidence score and the confidence threshold to obtain the recombined feature;

[0114] Specifically, when the confidence score < 0.3, based on the gating weight, the localization information is error-compensated through the class probability to obtain the calibrated feature, and the normalized feature map is replaced with the calibrated feature to recalculate the confidence score; until the confidence score ≥ 0.3, the localization information and the confidence score are weighted and fused through the gating weight to obtain the recombined feature.

[0115] More specifically, the expression for error compensation is:

[0116] F r = F o + g·ConvF c ;

[0117] where, F r is the calibrated feature, F o is the feature of the localization information, g is the gating weight, and F c is the feature of the class probability.

[0118] The expression for the recombined feature is: F f = (1 - g)·Fo +g·F c 。

[0119] Step 605: Concatenate the reconstructed features and class probabilities according to the spatial positions to obtain a preliminary result.

[0120] Step 700: Optimize the loss of the preliminary result through a multi-task joint loss function and dynamically adjust the weights of the multi-task joint loss function in real time to obtain the fruit recognition result.

[0121] Specifically, the multi-task joint loss function includes: a localization loss function, a confidence loss function, and a classification loss function. The localization loss function uses a deformation-sensitive GIoU loss function, the confidence loss function uses a cross-entropy loss function, and the classification loss function uses a distillation loss function. The weights of the multi-task joint loss function are calculated through gradient magnitudes.

[0122] The beneficial effects of the present invention are as follows:

[0123] 1) Through the multi-modal dynamic fusion network and the spatio-temporal synchronous attention mechanism, the fruit detection accuracy in complex orchard environments is significantly improved;

[0124] 2) By combining the three-dimensional space deformation enhancement technology and the light compensation algorithm, problems such as backlighting, shadows, and fruit deformation are effectively solved;

[0125] 3) The number of model parameters is reduced through the dynamic path strategy, resulting in a significant improvement in the inference speed;

[0126] 4) Through the multi-task joint loss function, the fine-grained classification ability is improved on the premise of ensuring the bounding box regression accuracy and recall rate of the model, and the robustness and real-time performance of the model are enhanced.

[0127] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0128] Specific examples are used in the present invention to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for identifying flat peach fruits based on an improved YOLOv8 network model, characterized in that, The method includes the following steps: Construct an improved multi-modal fusion YOLOv8 model; the multi-modal fusion YOLOv8 model includes: a hybrid perception backbone network, a multi-scale interaction fusion module, and an adaptive dynamic detection head; Perform light elimination and shadow repair operations on the flat peach fruit image through a light compensation algorithm to obtain a light-eliminated image; Perform three-dimensional spatial deformation enhancement operations on the light-eliminated image to obtain a spatially enhanced image; Perform dynamic path selection on the spatially enhanced image through the hybrid perception backbone network and extract local detail features to obtain a feature-enhanced image; Perform spatio-temporal collaborative optimization on the feature-enhanced image through the bidirectional feature pyramid and spatio-temporal synchronous attention mechanism of the multi-scale interaction fusion module to obtain a spatio-temporal joint attention map; Perform dynamic information interaction on the spatio-temporal joint attention map through the adaptive dynamic detection head to obtain a preliminary result; Optimize the loss of the preliminary result through a multi-task joint loss function and dynamically adjust the weights of the multi-task joint loss function in real time to obtain a fruit recognition result.

2. The flat peach fruit recognition method based on the improved YOLOv8 network model according to claim 1, characterized in that The hybrid perception backbone network includes: a lightweight convolution branch, a global attention branch, and a deformable feature enhancement branch; the adaptive dynamic detection head includes: an object localization unit, a confidence prediction unit, a class discrimination unit, and a gated feature recombination unit.

3. The flat peach fruit recognition method based on the improved YOLOv8 network model according to claim 1, characterized in that, Performing light elimination and shadow repair operations on the flat peach fruit image through a light compensation algorithm to obtain a light-eliminated image, including: Obtain a global light distribution map according to the low-frequency light characteristics of the flat peach fruit image and the physical properties of direct light and diffuse light; Perform dynamic brightness mapping on the light distribution map through an S-shaped curve compression function and adaptive gamma correction to obtain a light compensation result; Perform shadow area segmentation and texture reconstruction operations on the global light distribution map to obtain a shadow repair result; Perform weight fusion and edge sharpening on the light compensation result and the shadow repair result to obtain the light-eliminated image.

4. The flat peach fruit recognition method based on the improved YOLOv8 network model according to claim 1, characterized in that, Performing three-dimensional spatial deformation enhancement operations on the light-eliminated image to obtain a spatially enhanced image, including: Generate a pseudo-depth map of the light-eliminated image through a monocular depth estimation network and construct a three-dimensional point cloud space model according to the pseudo-depth map; Simulate multi-angle shooting in the three-dimensional point cloud space model through a random perspective projection algorithm to generate a two-dimensional projection image under a virtual perspective; Perform non-rigid deformation compensation on the two-dimensional projection image to obtain the spatially enhanced image.

5. The flat peach fruit recognition method based on the improved YOLOv8 network model according to claim 1, characterized in that, Perform dynamic path selection on the spatially enhanced image through the hybrid perception backbone network and extract local detail features to obtain a feature-enhanced image, including: Extract the local detail features of the spatially enhanced image through the lightweight convolution layer built in the hybrid perception backbone network; the local detail features include: texture complexity, light uniformity, and occlusion density; Select a branch activation path according to the local detail features to obtain a branch output result; Generate dynamic branch weights through a Sigmoid function and weighted aggregate the branch output results according to the dynamic branch weights to obtain the feature-enhanced image.

6. The flat peach fruit recognition method based on the improved YOLOv8 network model according to claim 5, characterized in that The calculation formula for the texture complexity is: TC = 0.6·C n + 0.4@E n ; where TC is the quantization value of the texture complexity, C n is the contrast of the local area of the image, and E n is the entropy of the local area of the image; The calculation formula for the illumination uniformity is: IU = 1 - (0.5·MSE n + 0.5·G n ); where, IU is the quantization value of the illumination uniformity, MSE n is the mean square error of the luminance channel, and G n is the local luminance gradient.

7. The flat peach fruit recognition method based on the improved YOLOv8 network model according to claim 1, characterized in that, Performing spatio-temporal collaborative optimization on the feature-enhanced image through the bidirectional feature pyramid and spatio-temporal synchronization attention mechanism of the multi-scale interaction fusion module to obtain a spatio-temporal joint attention map, including: Extracting multi-level feature maps of the feature-enhanced image through the hybrid perception backbone network; the multi-level feature maps include: low-level feature maps and high-level feature maps; Performing element-wise addition of the high-level feature map after bilinear upsampling and the low-level feature map to obtain a forward feature map; Concatenating the low-level feature map after downsampling by a 3×3 depthwise separable convolution and the high-level feature map to obtain a reverse feature map; Performing weighted fusion of the forward feature map and the reverse feature map through a Softmax function to obtain a bidirectional feature map; Performing deformable convolution on the bidirectional feature map through an offset field generated by a 3×3 convolutional layer and combining with an average pooling operation to obtain a spatial attention map; Aligning consecutive frames in the bidirectional feature map and inputting them into an LSTM network to obtain a temporal correlation degree; Performing a Hadamard product operation on the spatial attention map and the temporal correlation degree to obtain the spatio-temporal joint attention map.

8. The flat peach fruit recognition method based on the improved YOLOv8 network model according to claim 2, wherein, Performing dynamic information interaction on the spatio-temporal joint attention map through the adaptive dynamic detection head to obtain a preliminary result, including: Performing convolutional compression on the spatio-temporal joint attention map to obtain a normalized feature map; Respectively obtaining the localization information, confidence score, and class probability of the normalized feature map through the target localization unit, the confidence prediction unit, and the class discrimination unit; Obtaining a gating weight through a Sigmoid function based on the confidence score; Determining a feature recombination scheme based on the comparison result of the confidence score and a confidence threshold to obtain a recombined feature; Concatenating the recombined feature and the class probability according to the spatial position to obtain the preliminary result.

9. The flat peach fruit recognition method based on the improved YOLOv8 network model according to claim 8, wherein, Determining a feature recombination scheme based on the comparison result of the confidence score and a confidence threshold to obtain a recombined feature, including: When the confidence score < 0.3, based on the gating weight, compensating the error of the localization information through the class probability to obtain a calibrated feature, and replacing the normalized feature map with the calibrated feature to recalculate the confidence score; When the confidence score ≥ 0.3, performing weighted fusion of the localization information and the confidence score through the gating weight to obtain the recombined feature.

10. The flat peach fruit recognition method based on the improved YOLOv8 network model according to claim 1, characterized in that, The multi-task joint loss function includes: a localization loss function, a confidence loss function, and a classification loss function; the weights of the multi-task joint loss function are calculated through gradient magnitude.

Citation Information

Cited By

  • Unmanned aerial vehicle cluster target long-time robust tracking method in low-altitude airspace complex environment

    CN120949800A

  • Unmanned aerial vehicle cluster target long-time robust tracking method under low-altitude airspace complex environment

    CN120949800B

  • Navel orange defect detection method and system

    CN122435604A

  • Method and system for navel orange defect detection

    CN122435604B