Apple identification method suitable for dwarf close planting and high-shielding orchard environment
By using RGB cameras, data augmentation, and feature extraction optimization techniques in apple picking robots, the accuracy and robustness issues of apple identification in dwarf dense-planting orchards have been solved, achieving efficient apple detection and positioning.
Patent Information
- Application Number
- CN202511216186.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-18
AI Technical Summary
In dwarf and densely planted orchards, intelligent apple picking robots face problems such as severe shading caused by high-density planting, drastic changes in light, limited feature extraction capabilities of target detection models, and incompatibility of bounding box loss functions with shading conditions, resulting in low accuracy and high false negative rate in apple recognition.
Images are acquired using an RGB camera. Combined with data augmentation techniques and annotation rules, the WIoU loss function is used to optimize the bounding box regression gradient. The C2f module and SPPELAN structure are embedded, and the feature extraction and bounding box regression are enhanced through the attention mechanism. The ODConv module and SE attention module are introduced to improve the robustness of the model to changes in illumination and occlusion.
It significantly improves the robustness and accuracy of target detection in complex scenarios for apple picking robots, especially in cases of small targets and occlusion, enabling precise location of apple areas and reducing the false negative rate.
Smart Images

Figure CN120976767A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of agricultural automation and computer vision, and particularly relates to an apple recognition method suitable for dwarf dense planting and high-shading orchard environments. BACKGROUND
[0002] Due to the serious shortage of labor in the apple picking link, intelligent apple picking robots are widely used in apple picking labor to adapt to the picking work of large-yield apple orchards.
[0003] However, in dwarf dense planting orchards, high-density planting leads to serious branch and leaf shading, combined with severe natural light changes, traditional recognition algorithms face technical bottlenecks such as high small target missed detection rate and large shading scene positioning deviation, which seriously restricts the practical process of the picking robot.
[0004] Under the environment of dwarf dense planting orchard, the intelligent apple picking robot faces severe challenges in apple recognition. Due to the high complexity of the environment, it is characterized by light spots, shadows caused by severe natural light changes, and significant day-night differences. These factors can easily cause color distortion or contrast reduction, interfere with the accurate extraction of fruit outlines, especially the strong light area high light reflection and the weak light area deep shadow have a significant impact. At the same time, under the dense planting mode, the branch and leaf layering leads to a large number of fruits being partially or completely shaded, the target boundary is seriously blurred, the small size apple missed detection phenomenon is prominent, and the target separation is extremely difficult due to the mutual overlap of fruits.
[0005] The existing apple recognition algorithm of the intelligent apple picking robot has limited feature extraction capability of the target detection model, the standard convolution cannot effectively adapt to the severe change of light, the feature fusion structure is inefficient when processing small target multi-scale information, which can easily lead to the loss of key semantic information, and the distinction between the interference such as branches, leaves, soil and support in the orchard is insufficient, especially the channel modeling method is not sufficient enough, which can easily misidentify the shadow area as fruit; and the bounding box loss function is sensitive to the target center distance but ignores the normalized weight distribution under the shading condition, which can easily cause positioning deviation, so that the existing apple recognition algorithm of the intelligent apple picking robot cannot meet the requirements of accurate apple recognition in dwarf dense planting orchards. SUMMARY
[0006] The purpose of the present application is to provide an apple recognition method suitable for dwarf dense planting and high-shading orchard environments, which can completely identify partially or completely shaded apples, accurately locate the visible area of the apples under the shading condition, and complete high-precision bounding box regression on this basis, thereby significantly improving the target detection robustness and recognition accuracy of the apple picking robot in complex scenes.
[0007] The technical solution of the present application is: A method for identifying apples in a dwarf and high-density planting environment with high shading, comprising the following steps: An RGB camera is used to collect apple images under different light conditions and shading levels; light conditions include strong light, shadow, and day-night changes. The resolution of the RGB camera image is 416x416 pixels.
[0008] After preprocessing the collected apple images, the apples in the preprocessed images are labeled using LabelImg software, and the minimum bounding rectangle is drawn; the preprocessing of the apple image includes uniform size adjustment, spatial voxelization grid division, and sampling. Data augmentation operations are applied, including random translation, scaling, adding Gaussian noise, random point rejection, and color deactivation (such as brightness and contrast adjustment); the labeling rule is that apples with more than 60% of the area shaded are not labeled. Apples with less than 60% of the edge exposed area are not labeled.
[0009] The labeled images are input into the target detection model for processing: The WIoU loss function optimizes the bounding box regression gradient through a dynamic distance attention weight mechanism; In the C2f module embedded with the ODConv module, the convolution kernel weight is dynamically adjusted in the Backbone feature extraction stage to obtain enhanced Backbone features; to improve the feature extraction capability under complex lighting and shading conditions; Backbone feature extraction refers to: after inputting an image, the Backbone network extracts discriminative multi-scale feature maps such as edges, textures, contours, and target shapes from the original image through multiple layers of convolution and nonlinear transformation, which are used for classification and positioning by the subsequent detection head.
[0010] The enhanced Backbone features are processed through the SPPELAN structure with attention mechanism, which recalibrates the feature channels of the target features in the unevenly lit areas of the image through the attention mechanism, and gives higher weight to the local features of the shaded target, strengthening the robustness of the target detection model to lighting changes and local shading, and suppressing the interference of the orchard background; the SPPELAN structure further multi-scale fuses the enhanced Backbone features to provide multi-scale features for the WIoU loss function, improving the small target detection capability; through multi-scale pooling and attention mechanism fusion, the detection capability of multi-scale small targets is enhanced; The target detection model outputs an image with predicted bounding boxes, completing the identification of apples in a dwarf and high-density planting environment with high shading.
[0011] Further, the method for optimizing the bounding box regression gradient of the WIoU loss function through a dynamic distance attention weight mechanism includes the following steps: The IoU value of the prediction box and the real box is calculated as a basic overlap index; The Euclidean distance between the center points of the prediction box and the real box is calculated to quantify the center point deviation of the prediction box and the real box. The Euclidean distance is normalized by target size to obtain a normalized distance factor; the scale effect is eliminated; The distance attention weight is constructed according to the IoU value of the prediction box and the real box and the normalized distance factor, which is self-adaptive to different target sizes, and enhances the boundary box regression optimization ability of small targets and partially occluded targets.
[0012] Through the above weighting mechanism, in the small target apple or occluded target detection scene, the stability and effectiveness of the boundary box regression loss function can still be maintained, thereby improving the robustness and accuracy of the detection model in the complex orchard scene.
[0013] Further, the ODConv module is embedded in the C2f module, and the ODConv module includes multiple parallel convolution branches for extracting features of different scales and directions; the dynamic weight value generated by the attention mechanism is controlled by the attention mechanism, and the learnable attention weight is introduced in the spatial dimension and the channel dimension to dynamically weight and fuse each dimension, realizing dynamic adjustment of feature selection and fusion; In each forward propagation process, the ODConv module adaptively adjusts the weighted response of the convolution kernel according to the spatial distribution and semantic content of the input features, thereby enhancing the network's expression ability for complex light changes, fruit occlusion, and fruit tree branch and leaf interference.
[0014] Further, the SPPELAN structure is introduced into the hierarchical aggregation mechanism based on the original SPPF structure, and the multi-branch convolution and cross-layer connection structure are used to aggregate features of different scales to enhance the model's adaptability to multi-scale targets; the SPPELAN structure uses multiple maximum pooling operations of different sizes, such as 1x1, 3x3, and 5x5, to extract multi-scale features from the input features. The extracted features are channel-by-channel compressed and reconstructed by lightweight convolution to enhance the expression ability of small targets in the feature map.
[0015] Further, the SE attention module is introduced in the SPPELAN structure during the pooling result fusion process to weight and strengthen important features in the channel dimension. The obtained features are attention-enhanced multi-scale features. Thus, the network's response ability to key semantic information is improved, and the interference of redundant background information is suppressed.
[0016] Further, in order to maintain the compactness of feature expression and information flow, the pooling branch multi-scale features are used as shallow features, the attention-enhanced multi-scale features are used as deep features, both of which are efficiently fused through cross-layer residual connection to improve the detection sensitivity of apple targets in dense fruit tree sheltered environment.
[0017] The SPPELAN structure as a whole adopts lightweight design, and takes into account detection accuracy and inference efficiency, effectively improves the multi-scale detection capability of apple small targets without significantly increasing the network parameter amount, and is particularly suitable for the target detection scene of dense fruit in the dwarf and dense planting orchard.
[0018] Further, the way of fusing the attention mechanism includes the following steps: The attention mechanism is embedded in the key convolutional layer in the Backbone main network through the SA module, and the importance of different channel information and spatial regions is adaptively weighted through the joint modeling of channel attention and spatial attention. Further, the attention mechanism adopts a channel grouping strategy, divides the input feature map into multiple subgroups, and models the spatial attention and channel attention of each subgroup respectively, and finally fuses the multi-dimensional attention information through channel scrambling operation.
[0019] Further, in the spatial attention modeling process, a local context extraction operation is introduced to enhance the model's perception ability of local fruit outlines and occlusion edges; in the channel attention modeling process, a global pooling method is used to enhance the ability to capture significant features.
[0020] The SA module has a significant inhibitory effect on the background noise region and complex light region (such as sunspot) in the orchard image, and enhances the features of the occluded fruit region, thereby improving the accuracy and stability of target detection; Through the fusion with the original convolution structure in the Backbone, the SA effectively improves the feature extraction quality on the basis of keeping the network lightweight, especially in the scene of dense interlacing of fruits and branches in the dwarf and dense planting orchard, the target boundary is blurred, etc. It shows stronger adaptability and robustness.
[0021] Compared with the prior art, the beneficial effects of the present application are: The WIoU loss function in the target detection model of the application optimizes the boundary box regression gradient through a dynamic distance attention weight mechanism to improve the detection robustness of small target apples and occluded targets; the C2f module can adaptively adjust the convolution strategy according to the content of the input image, so as to more accurately capture the key region features, especially in the case of uneven light, strong background interference, and fruit part occlusion, and better recognition robustness is shown, the dynamic convolution enhances the modeling ability of the model to the dependence between space and channel, and the accuracy and stability of apple detection are significantly improved; the SPPELAN structure fuses shallow and deep information through a multi-scale parallel pooling path, and simultaneously combines an attention mechanism to significantly enhance small target regions, not only avoiding the feature blurring problem caused by traditional pooling, but also enhancing the model's ability to distinguish detailed regions (such as small apples); the WIoUv3 loss function in the target detection model, the ODConv module embedded in the C2f module, the SPPELAN structure and the attention mechanism synergistically improve the apple detection performance in the dwarfing and dense planting scene. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 Flow chart for identifying and detecting apples in the method of the application.
[0023] Figure 2 Schematic diagram for data set annotation of the collected orchard images.
[0024] Figure 3 Network structure diagram of the target detection model of the application.
[0025] Figure 4 Structure diagram of the full-dimensional dynamic convolution of the application.
[0026] Figure 5 Structure diagram of the existing C2f module and the C2f module embedded with the ODConv module of the application, wherein, Figure 5 The upper part of the frame is the structure diagram of the existing C2f module, and the lower part of the frame is the structure diagram of the C2f module embedded with the ODConv module of the application.
[0027] Figure 6 Structure diagram of the existing SPPE structure and structure diagram of the SPPELAN structure of the application, wherein, Figure 6 The upper part of the frame is the structure diagram of the existing SPPE structure, and the lower part of the frame is the structure diagram of the SPPELAN structure of the application.
[0028] Figure 7 Structure diagram of the attention mechanism module of the application.
[0029] Figure 8 Comparison diagram of convergence of different loss functions.
[0030] Figure 9 The target detection model is used to detect the recognition effect of the orchard apple image. DETAILED DESCRIPTION
[0031] The specific embodiments of the present application will be described in detail below. Figures 1 to 9 The specific embodiments of the present application will be described in detail below.
[0032] It should be noted that the circuit connections involved in the present application all adopt conventional circuit connection methods and do not involve any innovation.
[0033] EMBODIMENT As shown in Figure 1 An apple recognition method suitable for a dwarf dense planting and high shelter orchard environment, comprising the following steps: An RGB camera is used to collect apple images covering different light conditions (strong light, shadow, day and night changes) and shelter degrees; the image resolution is 416x416 pixels.
[0034] As shown in Figure 2 After the collected apple images are preprocessed, the apples in the preprocessed images are labeled by LabelImg software, and the minimum circumscribed rectangle is drawn; the preprocessing of the apple image includes uniform size adjustment, spatial voxelization grid division and sampling of the image; data enhancement operations are applied, including random translation, scaling, adding Gaussian noise, random point rejection, color inactivation (such as brightness and contrast adjustment) and the like;
[0035] The labeling rule is that apples with a shelter area of more than 60% are not labeled. Apples with an image edge exposed area of less than 60% are not labeled.
[0036] The labeled images are input into the target detection model for processing. The target detection model of this embodiment is a YOLOv8 model, wherein the target detection model is as shown in Figure 3 The WIoU loss function optimizes the boundary box regression gradient through a dynamic distance attention weight mechanism; As shown in Figure 4 Through the ODConv module embedded in the C2f module, the convolution kernel weight is dynamically adjusted in the Backbone feature extraction stage to obtain enhanced Backbone features; to improve the feature extraction capability under complex light and shelter conditions; Backbone feature extraction refers to: after inputting an image, the Backbone network extracts multi-scale feature maps with discriminative characteristics such as edges, textures, contours and target shapes from the original image through multiple layers of convolution and nonlinear transformation, which are used for classification and positioning by the subsequent detection head.
[0037] As shown in Figure 7 As shown, the enhanced Backbone feature is processed by the SPPELAN structure with attention mechanism, the target feature in the unevenly illuminated area of the image is recalibrated by the attention mechanism, and the local feature of the occluded target is given a higher weight, which enhances the robustness of the target detection model to light changes and local occlusions, and suppresses the interference of the orchard background; the SPPELAN structure further multi-scale fuses the enhanced Backbone feature to provide multi-scale features for the WIoU loss function, and improves the small target detection capability; through multi-scale pooling and attention mechanism fusion, the detection capability of multi-scale small targets is enhanced, and an image with a prediction box is generated, completing apple recognition in the environment of dwarf dense planting and high occlusion orchard.
[0038] The method for optimizing the boundary box regression gradient of the WIoU loss function through a dynamic distance attention weight mechanism comprises the following steps: Calculate the IoU value of the prediction box and the real box as the basic overlap index; Calculate the Euclidean distance between the center points of the prediction box and the real box to quantify the center point deviation of the prediction box and the real box, normalize the Euclidean distance according to the target size, and obtain the normalized distance factor; eliminate the scale effect; According to the IoU value of the prediction box and the real box and the normalized distance factor, a distance attention weight is constructed to adaptively adjust different target sizes and enhance the boundary box regression optimization capability for small targets and partially occluded targets.
[0039] Through the above weighting mechanism, the stability and effectiveness of the boundary box regression loss function can still be maintained in the small target apple or occluded target detection scene, thereby improving the robustness and accuracy of the detection model in the complex orchard scene.
[0040] As shown in Figure 5 The ODConv module is embedded in the C2f module, the ODConv module includes multiple parallel convolution branches for extracting features of different scales and directions; the dynamic weight value generated by the attention mechanism is controlled by the attention mechanism, and the learnable attention weight is introduced in the spatial dimension and the channel dimension to dynamically weight and fuse each dimension, realizing dynamic adjustment of feature selection and fusion; In each forward propagation process, the ODConv module adaptively adjusts the weighted response of the convolution kernel according to the spatial distribution and semantic content of the input feature, thereby enhancing the expression ability of the network to complex light changes, fruit occlusions, and fruit tree branch interference.
[0041] As shown in Figure 6As shown, in the YOLOv8 network, the SPPELAN structure is introduced into the original SPPF structure based on a hierarchical aggregation mechanism. Different scale feature maps are aggregated through multi-branch convolution and cross-layer connection structure to enhance the adaptability of the model to multi-scale targets. The SPPELAN structure uses multiple maximum pooling operations of different sizes, such as 1x1, 3x3, and 5x5, to perform multi-scale extraction on the input features, and then performs channel-by-channel compression and reconstruction through lightweight convolution to enhance the expression ability of small targets in the feature map.
[0042] The SPPELAN structure introduces an SE attention module in the pooling result fusion process to weight and strengthen important features in the channel dimension. This improves the network's response to key semantic information and suppresses redundant background information interference.
[0043] To maintain the compactness and information flow of feature expression, shallow and deep multi-scale features are efficiently fused using cross-layer residual connection. This improves the detection sensitivity of apple targets in dense orchard environments.
[0044] The SPPELAN structure is designed with lightweight overall, balancing detection accuracy and inference efficiency. Without significantly increasing the number of network parameters, it effectively improves the multi-scale detection capability of small apple targets, especially in the target detection scenarios of dense and multi-scale mixed fruits in dwarf and dense planting orchards. CT: Concat (Concatenation); S: Sigmoid activation function; GN: Group Normalization; GAP: Global Average Pooling; CS: Channel Shuffle; EWP: Element-wise Multiplication; Fc(X) = Wx + b; FG: Feature Grouping; AG: Aggregation;
[0045] The way to introduce attention mechanism in the Backbone part of YOLOv8 network includes the following steps: The attention mechanism embeds the key convolutional layers in the Backbone backbone network through the SA module. Through the joint modeling of channel attention and spatial attention, the importance of different channel information and spatial regions is adaptively weighted.
[0046] The attention mechanism uses a channel grouping strategy to divide the input feature map into multiple subgroups, and models the spatial attention and channel attention for each subgroup. Finally, the multi-dimensional attention information is fused through channel shuffling.
[0047] In the spatial attention modeling process, a local context extraction operation is introduced to enhance the model's perception of local fruit outlines and occluded edges. In the channel attention modeling process, global pooling is used to enhance the ability to capture significant features.
[0048] The SA module has significant inhibitory effect on background noise regions and complex illumination regions (such as sunflecks) in orchard images, and performs feature enhancement on the fruit regions of the occluded parts, thereby improving the accuracy and stability of target detection. By fusing with the original convolution structure in Backbone, the SA effectively improves the feature extraction quality on the basis of keeping the network lightweight, and has stronger adaptability and robustness in scenes such as dense interlacing of fruits and branches in dwarf and dense planting orchards and blurred target boundaries.
[0049] As shown in Figure 8 Compared with a plurality of mainstream regression loss functions, the WIoU loss function has significant advantages in comprehensive training performance, convergence efficiency and model robustness. After using the WIoU loss function, the target detection model not only shows faster convergence speed in the training stage, but also realizes stable improvement in a plurality of detection performance indicators, verifying the effectiveness and applicability of the loss function in the present application.
[0050] As shown in Figure 9 The improved target detection model shows good detection effect in the detection of sparse and dense distribution apple targets in the orchard environment, and no obvious missed detection phenomenon occurs. It can be seen that the improved target detection model has strong adaptability in different target density scenes, further verifying the effectiveness and practicality of the proposed improvement strategy in the complex orchard environment.
[0051] The above disclosure is only a few preferred specific embodiments of the present application, but the embodiments of the present application are not limited thereto, and any changes that can be thought of by those skilled in the art shall fall within the protection scope of the present application.
Claims
1. A method for apple identification suitable for dwarf, densely planted, and highly shaded orchard environments, characterized in that, Includes the following steps: Capture apple images; The acquired images are input into the target detection model for processing, including the following steps: The ODConv module embedded in the C2f module dynamically adjusts the convolution kernel weights during the Backbone feature extraction stage to obtain enhanced Backbone features; The enhanced Backbone features are processed through the SPPELAN structure that integrates an attention mechanism, and the feature channels of the target features in the unevenly lit areas of the image are recalibrated through the attention mechanism, and higher weights are given to the local features of the occluded targets. The SPPELAN structure performs multi-scale fusion of enhanced backbone features to generate images with prediction boxes, enabling apple recognition in dwarf, densely planted, and highly shaded orchard environments.
2. The apple identification method according to claim 1, applicable to dwarf dense planting and high-shading orchard environments, is characterized in that, The ODConv module contains multiple parallel convolutional branches, which are used to extract feature responses at different scales and directions. Attention weights are introduced in both the spatial and channel dimensions, and dynamic weighted fusion is performed on each dimension to achieve dynamic adjustment of feature selection and fusion.
3. The apple identification method according to claim 2, applicable to dwarf dense planting and high-shading orchard environments, is characterized in that... The SPPELAN structure employs multiple max pooling operations of different sizes to extract input features at multiple scales. The extracted features are multi-scale features of the pooling branches, and are compressed and reconstructed channel by channel through lightweight convolution to enhance the expressive power of small objects in the feature map.
4. The apple identification method according to claim 3, applicable to dwarf dense planting and high-shading orchard environments, is characterized in that... The SPPELAN structure introduces an SE attention module during the pooling result fusion process to weight and enhance important features in the channel dimension, resulting in attention-enhanced multi-scale features.
5. The apple identification method according to claim 4, applicable to dwarf dense planting and high-shading orchard environments, is characterized in that... Pooling branch multi-scale features are used as shallow features, while attention-enhanced multi-scale features are used as deep features. The two are efficiently integrated using a cross-layer residual connection method.
6. The apple identification method according to claim 1, applicable to dwarf dense planting and high-shading orchard environments, is characterized in that... The method of integrating the attention mechanism includes the following steps: The attention mechanism is embedded in the key convolutional layers of the backbone network through the SA module. By jointly modeling channel attention and spatial attention, the importance of different channel information and spatial regions is adaptively weighted.
7. The apple identification method according to claim 6, applicable to dwarf dense planting and high-shading orchard environments, is characterized in that... The attention mechanism employs a channel grouping strategy, dividing the input feature map into multiple subgroups and performing spatial attention modeling and channel attention modeling on each subgroup. Finally, the multi-dimensional attention information is fused through a channel shuffling operation.
8. The apple identification method according to claim 7, applicable to dwarf dense planting and high-shading orchard environments, is characterized in that... In the process of spatial attention modeling, a local context extraction operation is introduced; in the process of channel attention modeling, global pooling is used to enhance the ability to capture salient features.
Citation Information
Cited By
Improved YOLOv8n-based cloudy orchard multi-class sheltered pear detection method
CN121661511A
Fruit maturity detection and grading picking method and system
CN121999481A