Fine classification method for citrus fruit tree plots based on GACL-DeepLabV3+ of unmanned aerial vehicle remote sensing image
By using the GACL-DeepLabV3+ model, which combines lightweight attention and multi-scale feature fusion, the problem of fine classification of multi-aged Wogan mandarin orange orchard plots in UAV remote sensing images was solved, achieving efficient and accurate orchard management support.
Patent Information
- Application Number
- CN202510752640.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-06-06
AI Technical Summary
Existing technologies struggle to accurately distinguish between mixed-age Wogan mandarin orange orchards in UAV remote sensing images, particularly in cases of blurred boundaries, insufficient differentiation of features among plots of similar tree ages, overlapping features between young plots and bare soil areas, and large model parameters that cannot be deployed in real time.
A method based on GACL-DeepLabV3+ was adopted, which combines MobileNetV2 backbone network, channel-aware lightweight attention module, gated axial spatial attention module and gated axial spatial pyramid structure. Through data augmentation and multi-scale feature fusion, fine classification of Wogan mandarin orange orchard plots was achieved.
It achieves lightweight and high-precision classification of Wogan mandarin orange orchard plots on a drone platform, improving boundary segmentation accuracy and robustness, reducing computational costs, and enhancing the accuracy and application guidance value of classification results.
Smart Images

Figure CN120635566B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of agricultural remote sensing and image recognition technology, specifically involving a method for fine classification of Wogan orange orchard plots based on UAV remote sensing images using GACL-DeepLabV3+. Background Technology
[0002] In the task of classifying fruit tree plots using UAV remote sensing imagery, existing methods mainly target the overall segmentation of large-scale crops, but they still face many challenges in the scenario of mixed planting of Wogan mandarin orange trees of different ages. First, Wogan mandarin orange plots of different ages are prone to boundary blurring in remote sensing imagery, such as overlapping canopies or branches and leaves of adjacent plots, making it difficult for traditional segmentation algorithms to accurately define the plot outlines. This problem stems from the complexity of the fruit tree canopy structure and the limitations of UAV imagery resolution. Especially in low-altitude imagery, the detailed information of densely planted areas is easily interfered with by factors such as lighting and shadows, further exacerbating the boundary segmentation error. Second, Wogan mandarin orange trees at different growth stages (such as plots with 1-year-old trees and plots with 2-3-year-old trees) exhibit similar spectral and textural features in the imagery, making it difficult for conventional models to capture their subtle differences. This kind of confusion is closely related to the gradual change in canopy morphology during the fruit tree growth cycle, especially since the differences in canopy height and branch density between adjacent tree ages are relatively small, resulting in insufficient feature discrimination. Furthermore, the vegetation cover of young fruit tree plots (such as seedling stages) is low, and the bare surface is prone to spectral overlap with surrounding non-planted areas (such as dirt roads and wasteland) in images, making it difficult for existing methods to effectively distinguish such plots. The root cause of this problem is that the vegetation signal of seedling stage plots is weak, and surface background noise dominates, while traditional models lack the ability to co-model local and global features, making it difficult to suppress background interference.
[0003] Existing technologies face the following difficulties in addressing the aforementioned problems: First, traditional convolutional neural networks rely on fixed receptive fields to extract features, making it difficult to adaptively adjust multi-scale contextual information. For example, Spatial Pyramid Pooling with Dilatation (ASPP) captures multi-scale features through dilated convolutions with a preset dilation rate, but its fixed structure is insufficiently adaptable to the morphological diversity of Wogan mandarin orange orchards, especially in boundary regions where redundant information is easily introduced. Second, existing attention mechanisms present a contradiction between lightweight design and computational efficiency. For example, while global attention can enhance feature relevance, its high computational complexity makes it difficult to deploy in real-time monitoring systems on UAV platforms; while lightweight design often leads to the loss of key spatial information, affecting classification accuracy. Third, data augmentation strategies lack specificity. Existing methods typically employ general augmentation techniques (such as random flipping and pruning), but lack specific simulation of noise types unique to Wogan mandarin orange orchards (such as local spectral anomalies caused by pesticide spraying), resulting in insufficient robustness of the model in complex scenarios. Fourth, the uneven distribution of samples from plots of different tree ages is a prominent issue in mixed planting patterns. For example, if there is a high proportion of plots with older trees and a small number of plots with younger trees, traditional data balancing methods (such as oversampling) are prone to introducing the risk of overfitting, which can further exacerbate classification errors.
[0004] These issues make it difficult for existing technologies to meet the precision management needs of multi-aged Wogan mandarin orange orchards in smart agriculture, such as differentiated water and fertilizer regulation and pest and disease monitoring. Addressing these challenges requires balancing lightweight models, multi-scale feature fusion, and noise robustness, while also designing efficient data augmentation and sample balancing strategies. This places high demands on the innovation of algorithm architecture and the optimization of computing resources. Summary of the Invention
[0005] One object of the present invention is to solve at least the above-mentioned problems and to provide at least the advantages that will be described later.
[0006] The present invention also solves the following technical problem:
[0007] Existing technologies lack sufficient accuracy in classifying mixed-age Wogan mandarin orange orchards. Traditional remote sensing image processing methods struggle to effectively distinguish plots with different tree ages and conditions, and the large number of model parameters makes real-time deployment on UAV platforms impossible. The mixed planting pattern leads to blurred plot boundaries and overlapping features between young plots and bare soil areas, necessitating a lightweight and high-precision classification method.
[0008] Traditional fixed-size image cutting can easily disrupt the canopy continuity of Wogan mandarin orange orchards (such as broken branches during the seedling stage and pesticide patch segmentation), reduce the quality of labeled data, and result in insufficient sensitivity of the model to complex boundaries.
[0009] Traditional attention mechanisms are computationally complex and lack joint modeling of channel semantics and spatial location correlation, resulting in redundancy of high-level features and difficulty in distinguishing subtle differences between plots with similar tree ages.
[0010] Under the mixed planting model, the canopy structure and branch density of Wogan mandarin orange plots of different ages vary significantly. Traditional convolution is difficult to adaptively capture axial spatial features, and boundary segmentation is easily affected by background interference.
[0011] Significant scale differences exist in multi-aged Wogan mandarin orange plots (fine texture in seedling stage vs. wide distribution in mature stage). Traditional hollow pyramids rely on a fixed expansion rate and cannot dynamically adapt to the gradual changes in tree age, leading to the failure of multi-scale feature fusion.
[0012] Shallow detail features (branch density) and deep semantic features (canopy layout) are difficult to align due to resolution differences. In mixed planting scenarios, cross-scale feature fusion is insufficient, and the boundary missegmentation rate is high.
[0013] The morphological complexity of mixed-age planting plots makes it impossible for a single feature extraction network to take into account both local details and global layout. The model is prone to confusing young plots with bare soil areas, and the classification results lack application guidance value.
[0014] To achieve the above-mentioned objectives of this invention, the present invention provides a method for fine classification of Wogan orange orchard plots based on UAV remote sensing imagery using GACL-DeepLabV3+, comprising the following steps:
[0015] S1: Use drones to acquire visible light remote sensing images of Wogan mandarin orange orchards of different ages and conditions in the study area;
[0016] S2: Preprocess the acquired remote sensing images to generate orthophotos with a resolution of centimeters. The preprocessing includes feature extraction and aerial triangulation to generate dense point clouds, construction of digital surface models, and orthorectification.
[0017] S3: Construct a dataset of Wogan mandarin orange orchard plots based on orthophotos. The dataset includes plots in the seedling stage, plots of 1-year-old orchards, plots of 2-3-year-old orchards, plots of 4-5-year-old orchards, and plots of orchards sprayed with lime pesticide. The images are labeled using a manual labeling tool, and the labeled images are then segmented.
[0018] S4: Perform data augmentation on the dataset, including applying Gaussian noise with an intensity of 10-50 and rotation transformation to the images, generating a data augmentation training set containing noise robustness and multi-angle features, which is divided into training set, validation set and test set according to the proportion.
[0019] S5: Construct the GACL-DeepLabV3+ model, which includes a MobileNetV2 backbone network, a channel-aware lightweight attention module, a gated axial spatial attention module, and a gated axial spatial pyramid structure. Train the GACL-DeepLabV3+ model using a data augmentation training set. During training, a stochastic gradient descent optimizer is used with an initial learning rate of 7e-3, a learning rate descent method of cosine annealing, a batch size of 4, and 100 training epochs.
[0020] S6: Input the test set into the trained GACL-DeepLabV3+ model and output the classification results of the Wogan mandarin orange orchard plots.
[0021] Preferably, the operation of the UAV acquiring visible light remote sensing images of the present invention further includes: establishing a flight parameter configuration table based on a preset fruit tree age gradient database and canopy morphology feature database; using a multi-rotor UAV equipped with a visible light multispectral imager with polarization filtering function; conducting aerial photography with the canopy layer as the reference plane during the morning scattered light period; wherein young trees with an age of less than 3 years use a flight altitude of 250 meters and a forward overlap rate of 80%, and mature high-yield trees use a flight altitude of 250 meters and a forward overlap rate of 70%; and configuring an image quality real-time diagnosis module based on edge computing units to perform online detection and re-photographing decisions on leaf texture clarity and abnormal fruit coloring areas.
[0022] Preferably, the manual annotation tool of the present invention is Labelme, and the annotated image segmentation adopts an adaptive block segmentation algorithm based on saliency guidance.
[0023] Preferably, the channel-aware lightweight attention module of the present invention extracts channel semantic information of input features through global average pooling, generates a channel semantic-guided key vector, generates a spatial query matrix through linear transformation, generates a spatial attention map by performing a dot product normalization on the key vector and the query matrix, and then outputs a weighted output through residual connections.
[0024] Preferably, the gated axial spatial attention module of the present invention enhances the spatial feature modeling capability of Wogan mandarin orange orchard plots through axial convolution and gating mechanism. Specifically, it includes: dividing the input features into two parts along the channel dimension; applying depthwise separable convolution to the first C channel features along the vertical and horizontal axes respectively to capture the differences in canopy structure and branch density of Wogan mandarin orange orchard plots in the vertical and horizontal directions, generating axial response features and summing them; and adaptively modulating the response features of the latter C channel features through learnable gating parameters to suppress background interference and enhance the boundary features of plots of different ages. Finally, the fused features are output through residual connection, thereby solving the problems of boundary blurring caused by mixed planting of Wogan mandarin orange orchard plots and the overlap of features between young plots and non-plot areas.
[0025] Preferably, the gated axial spatial pyramid structure of the present invention constructs a multi-scale spatial modeling mechanism for Wogan mandarin orange orchard plots by stacking the gated axial spatial attention modules in layers. Specifically, it includes: processing the input features through multiple parallel branches, with each branch compressing the channel dimension through 1×1 convolution and then introducing 1-4 layers of the gated axial spatial attention modules. The receptive field is gradually expanded using modules of different numbers of layers. The shallow branches focus on the fine-grained canopy texture of Wogan mandarin orange seedling plots through single-layer modules, while the deep branches capture the wide-area canopy distribution features of 4-5 year old tree plots through multi-layer modules. The multi-scale features output by the modules are unified in resolution through bilinear interpolation and then spliced with the shallow detail features extracted by the MobileNetV2 backbone network. This integrates the local branch density information and global spatial layout features of the Wogan mandarin orange orchard plots, thereby solving the boundary blurring problem caused by morphological differences in plots of different ages under the multi-age mixed planting mode and suppressing misclassification interference between young plots and bare soil areas.
[0026] Preferably, the fusion mechanism of multi-scale features and shallow features output by the gated axial spatial pyramid structure of the present invention is designed for the morphological complexity of Wogan mandarin orange orchard plots, and specifically includes the following steps:
[0027] Step 1: Shallow feature extraction and preservation: Based on the shallow features extracted by the MobileNetV2 backbone network, the fine-grained branch density distribution characteristics and canopy edge texture details of the Wogan seedling plot are preserved to enhance the ability to distinguish between young fruit tree plots and bare soil and non-plot areas such as roads.
[0028] Step 2: Multi-scale feature alignment and adaptation: The multi-scale features (shallow branches: fine texture in seedling stage; deep branches: wide canopy of 4-5 year old trees) output by the gated axial spatial pyramid are upsampled using a bilinear interpolation algorithm to make their spatial resolution consistent with the shallow features (e.g., 128×128) extracted by the MobileNetV2 backbone network, thereby achieving geometric alignment of cross-scale features;
[0029] Step 3: Cross-scale feature fusion and boundary enhancement: The aligned multi-scale features and shallow features are concatenated along the channel dimension, and cross-scale information is fused through two 3×3 convolutional layers:
[0030] First convolution: Extract the characteristics of branch density gradient changes within the Wogan mandarin fruit plots to enhance the response to differences in canopy coverage among plots of different tree ages;
[0031] Secondary convolution: captures the patchy distribution characteristics of lime-sprayed pesticide plots and the scattered boundary abrupt change signals of young fruit tree plots, and suppresses missegmentation of adjacent plots caused by canopy overlap or background interference under mixed planting patterns.
[0032] Step 4: High-resolution classification map reconstruction: The fused features are upsampled four times to restore the original image resolution, generating a classification result map that matches the UAV orthophoto at the pixel level. This accurately marks the continuous boundaries and internal structure of Wogan mandarin orange orchards of different ages, solving the problems of blurred boundaries and missed detection of transition areas between young orchards and bare soil caused by mixed planting of multiple ages.
[0033] Preferably, the construction of the GACL-DeepLabV3+ model of the present invention is designed for a multi-age mixed planting pattern in Wogan mandarin orange orchards, specifically including the following steps:
[0034] Step 1: Multi-level feature extraction from the input image: Input the visible light image of the Wogan mandarin orange orchard plot acquired by the UAV into the MobileNetV2 backbone network to extract shallow and high-level features respectively:
[0035] Shallow features: Preserve the fine-grained branch density distribution, canopy edge texture, and soil characteristics of the exposed surface of young plots to distinguish non-fruit tree areas;
[0036] High-level characteristics: Capture the wide-area canopy distribution pattern of 4-5 year old plots and the patchy spectral response of pesticide-sprayed plots to characterize the global spatial layout of mature Wogan mandarin orange trees;
[0037] Step 2: Channel-aware attention-guided spatial feature enhancement: The aforementioned channel-aware lightweight attention module is embedded after the high-level features to optimize the semantic discriminative ability of Wogan orange plot classification through the following operations:
[0038] Channel semantic information of high-level features is extracted based on global average pooling, and a channel semantically guided key vector is generated.
[0039] The spatial location information is converted into a query matrix, and the correlation weight between channel semantics and spatial location is calculated.
[0040] Residual connectivity is used to enhance the response intensity to sparse shoot areas in seedling plots and local patches in pesticide-sprayed plots, thereby suppressing background noise interference.
[0041] Step 3: Multi-scale spatial modeling and tree age feature adaptation: The high-level features enhanced by CLA are input into a gated axial spatial pyramid structure, and the age differences of the Wogan mandarin orange orchards are adapted through the following mechanism:
[0042] Shallow branching focuses on the fine-grained canopy texture of seedling plots, enhancing the contrast between young fruit trees and the bare soil boundary;
[0043] Deep branching extends the perception field, capturing the wide canopy distribution of 4-5 year old trees and the staggered boundary features between adjacent plots;
[0044] The output multi-scale features are aligned by bilinear interpolation to preserve the morphological gradation characteristics of plots with different tree ages;
[0045] Step 4: Classification Result Generation and Application Optimization: The cross-scale fused features are reconstructed using high-resolution methods to generate a classification result map that matches the UAV orthophoto at the pixel level. This accurately marks the continuous boundaries and transition areas of multi-aged Wogan mandarin orange orchard plots. The results are then imported into the orchard intelligent management system to guide differentiated fertilization, pest and disease control, and harvest planning, significantly reducing the cost of manual surveying and improving orchard management efficiency.
[0046] This invention has at least the following beneficial effects:
[0047] By using an end-to-end lightweight model design (MobileNetV2 + attention module), an average intersection-over-union ratio (mIoU) of 94.03% was achieved with only 5.8MB of parameters, significantly reducing the deployment cost of drones. By combining multi-scale feature fusion and noise-robust data enhancement, the boundary ambiguity problem in mixed planting patterns was solved, the classification accuracy of young plots was improved by 10%, and the pesticide patch recognition rate was increased to 95%, providing reliable technical support for precision management of orchards.
[0048] The CLA module guides spatial attention through channel semantics, resulting in a significantly lower number of parameters compared to traditional attention mechanisms, thus improving computational efficiency. In classification tasks involving plots of similar tree age (2-3 years vs. 4-5 years), it enhances the discriminative index and effectively reduces feature redundancy.
[0049] The GASA module captures the canopy density gradient through axial convolution, improving the intersection-over-union (IoU) ratio of boundary segmentation; the gating mechanism suppresses background interference and reduces the misclassification rate of young plots and bare soil areas, significantly enhancing the robustness of the model in mixed planting scenarios.
[0050] The gated axial pyramid structure of the layered stacked GASA module adapts to the gradual morphological changes of plots from seedling to adulthood, achieving an average recall rate of 96.7% in multi-age mixed scenarios; the multi-layered branch structure expands the receptive field, deeply extracts data texture features, improves classification accuracy in the seedling stage, and reduces segmentation error in the adult stage.
[0051] The cross-scale feature fusion mechanism aligns the resolution through bilinear interpolation and reduces the boundary missegmentation rate; the secondary convolutional layer enhances the response of pesticide spraying patches, and the four-fold upsampling achieves pixel-level matching, reducing the missed detection rate in transition areas.
[0052] The fusion of multi-level feature extraction and attention guidance enabled the model to achieve an overall accuracy (OA) of 96.7% in complex orchard scenarios. After the classification results were imported into the management system, the efficiency of differentiated fertilization was improved, the cost of manual surveying was reduced, and the orchard management benefits were significantly improved.
[0053] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Attached Figure Description
[0054] Figure 1 The following is a schematic diagram of the research area of the present invention: (a) is a location indication map of the research area, (b) is a UAV image of the research area, and (c) is a detailed map of the Wogan mandarin orange orchard plot.
[0055] Figure 2 The following is a classification diagram of Wogan mandarin orange orchard plots according to the present invention: (a) plots in the seedling stage, (b) plots with 1-year-old trees, (c) plots with 2-3-year-old trees, (d) plots with 4-5-year-old trees, and (e) agricultural plots with lime spraying.
[0056] Figure 3 This is a data enhancement example of the present invention;
[0057] Figure 4 This is a diagram of the lightweight attention (CLA) structure for channel sensing in this invention;
[0058] Figure 5 This is a structural diagram of the gated axial spatial attention module (GASA) of the present invention;
[0059] Figure 6 This is a structural diagram of the gated axial space pyramid (GASP) of the present invention;
[0060] Figure 7 This is a structural diagram of the GACL-DeepLabV3+ model of the present invention;
[0061] Figure 8The seedling stage plot classification and identification results of the present invention are shown in (a) and (b) are the original images and the result images.
[0062] Figure 9 The results of the classification and identification of 1-year-old tree plots in this invention are shown in (a) and (b) respectively.
[0063] Figure 10 The images show the classification and identification results of 2-3 year old tree plots in this invention. (a) is the original image, and (b) is the result image.
[0064] Figure 11 The images show the classification and identification results of 4-5 year old tree plots in this invention. (a) is the original image, and (b) is the result image.
[0065] Figure 12 The images show the classification and identification results of the lime-sprayed pesticide plots according to the present invention. (a) is the original image, and (b) is the result image. Detailed Implementation
[0066] The present invention will now be described in further detail with reference to the accompanying drawings, so that those skilled in the art can implement it based on the description.
[0067] It should be understood that terms such as “having,” “comprising,” and “including” as used herein do not exclude the presence or addition of one or more other elements or combinations thereof.
[0068] It should be noted that, unless otherwise specified, the experimental methods described in the following implementation plan are all conventional methods, and the reagents and materials described are all commercially available unless otherwise specified.
[0069] <Example>
[0070] A method for fine classification of Wogan mandarin orange orchard plots based on UAV remote sensing imagery using GACL-DeepLabV3+, comprising the following steps:
[0071] S1: Use drones to acquire visible light remote sensing images of Wogan mandarin orange orchards of different ages and conditions in the study area;
[0072] S2: Preprocess the acquired remote sensing images to generate orthophotos with a resolution of 0.05-0.10 meters. The preprocessing includes feature extraction and aerial triangulation to generate dense point clouds, construction of digital surface models, and orthorectification.
[0073] S3: Construct a dataset of Wogan mandarin orange orchard plots based on orthophotos. The dataset includes plots in the seedling stage, plots of 1-year-old orchards, plots of 2-3-year-old orchards, plots of 4-5-year-old orchards, and plots of orchards sprayed with lime pesticide. The images are labeled using a manual labeling tool, and the labeled images are then segmented.
[0074] S4: Perform data augmentation on the dataset, including applying Gaussian noise with an intensity of 10-50 and rotation transformation to the images, generating a data augmentation training set containing noise robustness and multi-angle features, which is divided into training set, validation set and test set according to the proportion.
[0075] S5: Construct the GACL-DeepLabV3+ model, which includes a MobileNetV2 backbone network, a channel-aware lightweight attention module, a gated axial spatial attention module, and a gated axial spatial pyramid structure. Train the GACL-DeepLabV3+ model using a data augmentation training set. During training, a stochastic gradient descent optimizer is used with an initial learning rate of 7e-1. The learning rate descent method is cosine annealing, the batch size is 4, and the number of training epochs is 100.
[0076] S6: Input the test set into the trained GACL-DeepLabV3+ model and output the classification results of the Wogan mandarin orange orchard plots.
[0077] The drone can be a multi-rotor model, and the imager can be a module supporting multispectral imaging. The edge computing unit can be integrated into the drone control system. The drone fuselage can be made of lightweight carbon fiber, and the sensors can be silicon-based CMOS chips. The imager is mounted on the bottom gimbal of the drone, and the edge computing unit is connected to the flight control module via a CAN bus.
[0078] Workflow: The UAV automatically plans its flight path based on preset parameters and transmits images to the edge computing unit in real time during flight. The edge computing unit evaluates the quality of the images, and if it detects blurry images or missing target categories in the images, it immediately generates a re-enhancing command. The preprocessing software performs aerial triangulation through feature matching and bundle adjustment, and finally outputs an orthophoto with a resolution of 0.068 meters.
[0079] The image cropping size was set to 512×512 pixels. Data augmentation included applying Gaussian noise with intensities of 10, 25, and 40, as well as rotation transformations of 90°, 180°, and 270°. The final dataset was divided into training, validation, and test sets in an 8:1:1 ratio.
[0080] The annotation tool can be the open-source Labelme software, which runs on Anaconda Prompt. Data storage can utilize SSDs, and the saliency detection model can be deployed on the PyTorch framework. The segmentation model is implemented based on the PyTorch framework.
[0081] Workflow: The data augmentation module adds noise and rotation to the image, generating an augmented dataset containing 2032 slices, which is stored on a local server.
[0082] The GACL-DeepLabV3+ model uses MobileNetV2 as its backbone network. The Channel Aware Attention Module (CLA) extracts channel semantic information through global average pooling, generating a Key vector and a spatial Query matrix. After dot product normalization, an attention map is generated. The Gated Axis Spatial Pyramid (GASP) contains four branches, stacking 1-4 layers of Gated Axis Attention Modules (GASA). During training, a stochastic gradient descent optimizer is used, with an initial learning rate of 7e-3, a batch size of 4, and 100 training epochs. The loss function is cross-entropy.
[0083] Equipment and Materials: Model training was based on the PyTorch framework and ran on a server equipped with an NVIDIA A100 GPU.
[0084] Workflow: The enhanced training set is input into the model, the loss value is calculated through forward propagation, and the parameters are updated through backpropagation. During training, the accuracy is evaluated on the validation set every 10 epochs, and the optimal model weights are saved. During testing, a 512×512 slice is input, the model outputs a classification result image, which is then upsampled four times to restore the original resolution, and the final classification result is output.
[0085] Technical effects:
[0086] Dataset construction and data augmentation: saliency-guided cutting preserves key features, buffer repair reduces edge distortion, and noise and rotation enhancement improve model robustness.
[0087] Model building and training: Lightweight design reduces computational overhead, and multi-scale attention mechanism enhances boundary segmentation capability.
[0088] In another technical solution, in order to balance model computation efficiency and explore the synergy between channel and spatial features, this invention proposes a channel-aware lightweight attention module (CLA). The CLA uses the global channel semantics of the input features as guiding information, and generates spatial attention weights with channel awareness by calculating the correlation between channel semantics and spatial location, thereby achieving response enhancement in spatially significant regions.
[0089] The structure of CLA is as follows Figure 4 As shown, the feature map x∈R C×H×W First, global average pooling is performed along the spatial dimension to obtain the overall response of each channel, represented as:
[0090] This vector is transformed by a linear transformation w k ∈R C'×C The key obtained is guided by channel semantics, denoted as K∈W. k ·K c ∈RC′ Simultaneously, the input feature x is flattened into x * ∈R (H·W)×C And through linear transformation w q ∈R C'×C Obtain the Query matrix Q∈W q ·x * R C′×(HvW) The spatial attention map is obtained by performing a dot product between the channel information key and the spatial query vector and then normalizing the result.
[0091] A = soft max(K) T ·Q)∈R 1×(H·W) (1)
[0092] Finally, the original input feature maps are element-wise weighted and output through residual connections:
[0093] y=γ·A☉x+x (2)
[0094] Where γ∈R is a learnable scaling factor.
[0095] CLA enhances the model's responsiveness to important spaces by introducing a simple Query-Key process and using channel-aware information to guide spatial features, thereby reducing parameters and achieving effective semantic enhancement.
[0096] Workflow: The attention map is element-wise multiplied with the original feature map, the weighting intensity is adjusted by a scaling factor, and then added back to the original feature map. The output feature map retains the original information while enhancing the response of salient regions.
[0097] Technical effects:
[0098] Global average pooling is used to extract channel-level semantic information, enhancing the model's sensitivity to key channels.
[0099] The Query-Key interaction mechanism is used to dynamically allocate spatial weights and reduce interference from redundant features.
[0100] By using a learnable scaling factor to balance the original features with the attention-weighted results, gradient vanishing can be avoided in the early stages of training.
[0101] In another technical solution, to enhance the model's ability to model spatial structure and target category boundaries in semantic segmentation tasks, this invention proposes a lightweight gated axial spatial attention (GASA). The spatial attention mechanism enhances target region features and improves the ability to capture target details by focusing on key spatial information. The core of GASA is to combine axial attention with spatial attention and introduce learnable gating parameters, thereby strengthening the perception of both local and long-range spatial information while maintaining low computational complexity.
[0102] The structure of GASA is as follows Figure 5 As shown, the input feature map of GASA is x∈R C×H×W First, the input is normalized and channel-expanded. The expanded feature map is divided into two parts along the channel dimension. The first C channel features r1 and r2 are subjected to depthwise separable convolutions along the vertical and horizontal axes, respectively, to obtain response features in two directions. At each spatial location (i,j), the outputs of the two axes are summed to form an attention map. This is represented as follows:
[0103]
[0104] Where r ij ∈R d Let (i,j) be the attention vector at position (i,j). and This represents an axial convolution kernel that incorporates relative position encoding, corresponding to the width and height responses, respectively. The last C channel features X* serve as feature branches, modulated by the aforementioned attention map and then gated element-wise. The results are fused through 1×1 convolutions, and a learnable scaling parameter γ∈R is introduced. 1×C×1×1 Finally, add the result to the original input residual to obtain the final output:
[0105] y ij =γ·ψ(x ij ⊙sigmoid(a ij ))+x ij (4)
[0106] Where x ij Representing the original input features, ψ(·) is the convolution function for channel fusion. This module structure does not require explicit construction of query-key interaction, and can extract spatial semantic information through the convolutional receptive field and strengthen it with a gating mechanism.
[0107] Technical effects:
[0108] By capturing differences in canopy structure and branch density through convolution along the vertical and horizontal axes, the distinguishability of age-related features is improved.
[0109] Dynamically modulate the characteristic response intensity to effectively reduce interference from background areas such as bare soil and roads.
[0110] Preserve the spatial details of the original input to avoid the vanishing gradient problem in deep network training.
[0111] In another technical solution, the gated axial spatial pyramid structure constructs a multi-scale spatial modeling mechanism for Wogan mandarin orange orchard plots by stacking the gated axial spatial attention modules in layers. Specifically, the input features are processed through multiple parallel branches. Each branch compresses the channel dimension through 1×1 convolution and then introduces 1-4 layers of the gated axial spatial attention modules. The receptive field is gradually expanded using modules of different numbers of layers. The shallow branches focus on the fine-grained canopy texture of Wogan mandarin orange seedling plots through single-layer modules, while the deep branches capture the wide-area canopy distribution features of 4-5 year old plots through multi-layer modules. The multi-scale features output by the modules are unified in resolution through bilinear interpolation and then spliced with the shallow detail features extracted by the MobileNetV2 backbone network. This integrates the local branch density information and global spatial layout features of the Wogan mandarin orange orchard plots, thereby solving the boundary blurring problem caused by morphological differences in plots of different ages under the multi-age mixed planting mode and suppressing the misclassification interference of young plots and bare soil areas.
[0112] The input feature map is processed through four parallel branches, each compressing the number of channels to one-quarter of the original number using a 1×1 convolution. For example, when the input feature map has 256 channels, each branch compresses it to 64 channels. The first branch introduces a 1-layer gated axial spatial attention (GASA) module, the second branch introduces 2 layers, the third branch introduces 3 layers, and the fourth branch introduces 4 layers. The axial convolution kernel size of each GASA module can be set to 3×1 (vertical axis) and 1×3 (horizontal axis), progressively expanding the receptive field. The shallow branches (layers 1-2) focus on the fine-grained texture of seedling plots, while the deep branches (layers 3-4) capture the wide-area distribution of mature plots.
[0113] Parallel branches can be implemented based on PyTorch's nn.ModuleList, and 1×1 convolutional layers can use standard convolutional modules.
[0114] Workflow: Input feature maps are processed into four branches, and after channel compression, GASA modules are stacked layer by layer. The receptive field is dynamically expanded through the spatial attention mechanism of different levels of GASA modules to capture long-range dependencies and deep details. The spatial resolution of the output feature maps of each branch remains consistent.
[0115] The feature maps output by the gated axial spatial pyramid structure are unified to the same resolution as the shallow features of MobileNetV2 through bilinear interpolation. For example, when the output of the deep branch is 128×128 resolution, the interpolation upsamples it to 512×512. The shallow features are extracted from the third inverted residual block of MobileNetV2, retaining 1 / 4 of the original input size resolution. During fusion, multi-scale features are concatenated along the channel dimension. For example, each of the four branches outputs 64 channels, which are concatenated to form 256 channels, and then concatenated with the 128 channels of the shallow features to form 384 channels.
[0116] Bilinear interpolation can be performed using PyTorch's `F.interpolate` function, with the upsampling mode set to `bilinear`. Feature concatenation is implemented using `torch.cat`, and memory management relies on CUDA memory optimization.
[0117] Workflow: After interpolation and alignment, the feature maps of each branch are concatenated with shallow features extracted by MobileNetV2 (such as branch density and soil texture). The fused features are then input into two 3×3 convolutional layers. The first layer outputs 256 channels, and the second layer outputs 128 channels. The ReLU activation function is used.
[0118] The features after cross-scale fusion are restored to the original image resolution through a four-fold upsampling using bilinear interpolation. The classification head uses a 1×1 convolution to compress the number of channels to the number of classes (e.g., 5 classes), outputting a class score map for each pixel. Finally, a Softmax activation function is used to normalize each pixel along the class dimension, obtaining the probability distribution for each class, which is used to generate the segmentation result.
[0119] The classification head convolutional layer can be implemented based on PyTorch's nn.Conv2d, with the Softmax function normalizing along the channel dimension.
[0120] Workflow: After upsampling, the fused features are used to generate the class probability of each pixel through a classification head. The probability map is then subjected to thresholding and morphological optimization to output the segmentation results of continuous boundaries, which are directly mapped to the UAV orthophoto coordinate system.
[0121] Technical effects:
[0122] 1. Multi-scale feature capture: By layering and stacking GASA modules, the morphological gradation characteristics of plots from seedling stage to maturity are adapted to improve the joint modeling capability of fine texture and wide-area features.
[0123] 2. Resolution Alignment and Fusion: Bilinear interpolation eliminates scale shift, and shallow details and deep semantics complement each other, enhancing the consistency of boundary segmentation in mixed planting scenarios.
[0124] 3. Applicability of classification results: Pixel-level matching output can be directly imported into the orchard management system to support precise decision-making for differentiated fertilization and pest and disease control.
[0125] In another technical solution, the fusion mechanism of the multi-scale features and shallow features output by the gated axial spatial pyramid structure is designed for the morphological complexity of Wogan mandarin orange orchard plots, and specifically includes the following steps:
[0126] Step 1: Shallow feature extraction and preservation: Based on the shallow features extracted by the MobileNetV2 backbone network, the fine-grained branch density distribution characteristics and canopy edge texture details of the Wogan seedling plot are preserved to enhance the ability to distinguish between young fruit tree plots and bare soil and non-plot areas such as roads.
[0127] Step 2: Multi-scale Feature Alignment and Adaptation: A bilinear interpolation algorithm is used to upsample the multi-scale features (shallow branches: fine textures in seedlings; deep branches: wide canopy of 4-5 year old trees) output by the gated axial spatial pyramid. This ensures that the spatial resolution matches the shallow features (e.g., 128×128) extracted by the MobileNetV2 backbone network, achieving geometric alignment of cross-scale features. This operation eliminates the feature scale offset effect caused by age gradient differences through pixel-level alignment in the spatial dimension. It ensures that fine-grained features such as local branch density distribution and canopy edge texture in seedling plots are geometrically consistent with global spatial features such as wide canopy layout and overlapping boundaries of adjacent plots in mature plots before semantic fusion. This mechanism effectively suppresses the boundary blurring problem caused by feature resolution mismatch in multi-age mixed planting scenarios, providing an accurate spatial benchmark for cross-scale feature fusion, thereby improving the model's ability to model age-gradient features and the geometric fidelity of classification boundaries.
[0128] Step 3: Cross-scale feature fusion and boundary enhancement: The aligned multi-scale features and shallow features are concatenated along the channel dimension, and cross-scale information is fused through two 3×3 convolutional layers:
[0129] First convolution: Extract the characteristics of branch density gradient changes within the Wogan mandarin fruit plots to enhance the response to differences in canopy coverage among plots of different tree ages;
[0130] Secondary convolution: captures the patchy distribution characteristics of lime-sprayed pesticide plots and the scattered boundary abrupt change signals of young fruit tree plots, and suppresses missegmentation of adjacent plots caused by canopy overlap or background interference under mixed planting patterns.
[0131] Step 4: High-resolution classification map reconstruction: The fused features are upsampled four times to restore the original image resolution, generating a classification result map that matches the UAV orthophoto at the pixel level. This accurately marks the continuous boundaries and internal structure of Wogan mandarin orange orchards of different ages, solving the problems of blurred boundaries and missed detection of transition areas between young orchards and bare soil caused by mixed planting of multiple ages.
[0132] Shallow features of the MobileNetV2 backbone network can be extracted from the third inverted residual block, with an output resolution of 1 / 4 of the input image. For example, with an input 512×512 image, the shallow feature size is 128×128×64. These features include fine-grained shoot density distribution (≥15 shoots per square meter) and canopy edge texture (edge gradient ≥0.2) for seedling plots.
[0133] Feature extraction can be implemented using the PyTorch framework, and pre-trained weights for MobileNetV2 can be loaded from publicly available model libraries.
[0134] The model's multi-scale features include: shallow features extracted from the early to mid-stages of the backbone network (128×128 resolution, 64 channels), focusing on capturing fine-grained texture information of seedling plots; and deep features derived from the ASPP module (64×64 resolution, 128 channels), used to characterize the wide-area distribution features of mature vegetation. To achieve spatial alignment, deep features are upsampled to 128×128 using PyTorch's bilinear interpolation mode, with interpolation errors controlled within ±1 pixel. Before stitching, L2 normalization (ε = 1e-6) is applied to each branch along the channel dimension to eliminate amplitude shifts caused by scale differences and improve multi-scale fusion performance.
[0135] Bilinear interpolation can be integrated into the model's forward propagation process, and the normalization layer is implemented using PyTorch's nn.BatchNorm2d. Feature concatenation uses the torch.cat function, and memory usage optimization is achieved through gradient checkpointing.
[0136] Workflow: Deep branch features are amplified by interpolation and then concatenated with shallow branch features along the channel dimension (64 + 128 = 192 channels). Normalization is applied independently to each branch before concatenation to ensure consistent feature distribution.
[0137] The first 3×3 convolutional layer outputs 128 channels, with kernel weights initialized to a He normal distribution and bias terms initialized to 0. The activation function is LeakyReLU (negative slope = 0.01), preserving weak negative responses. The second 3×3 convolutional layer outputs 64 channels, capturing the discontinuous boundaries of pesticide-sprayed patches (minimum patch area ≥ 50 pixels). During classification map reconstruction, a 4x upsampling is achieved through two bilinear interpolations (2x × 2x). After interpolation, a 1×1 convolution is used to adjust the number of channels to the number of classes, and the Softmax temperature parameter is set to 1.0.
[0138] In another technical solution, the construction of the GACL-DeepLabV3+ model is designed for a multi-age mixed planting pattern in Wogan mandarin orange orchards, specifically including the following steps:
[0139] Step 1: Multi-level feature extraction from the input image: Input the visible light image of the Wogan mandarin orange orchard plot acquired by the UAV into the MobileNetV2 backbone network to extract shallow and high-level features respectively:
[0140] Shallow features: Preserve the fine-grained branch density distribution, canopy edge texture, and soil characteristics of the exposed surface of young plots to distinguish non-fruit tree areas;
[0141] High-level characteristics: Capture the wide-area canopy distribution pattern of 4-5 year old plots and the patchy spectral response of pesticide-sprayed plots to characterize the global spatial layout of mature Wogan mandarin orange trees;
[0142] Step 2: Channel-aware attention-guided spatial feature enhancement: The aforementioned channel-aware lightweight attention module is embedded after the high-level features to optimize the semantic discriminative ability of Wogan orange plot classification through the following operations:
[0143] Channel semantic information of high-level features is extracted based on global average pooling, and a channel semantically guided key vector is generated.
[0144] The spatial location information is converted into a query matrix, and the correlation weight between channel semantics and spatial location is calculated.
[0145] Step 3: Multi-scale spatial modeling and tree age feature adaptation: The high-level features enhanced by CLA are input into a gated axial spatial pyramid structure, and the age differences of the Wogan mandarin orange orchards are adapted through the following mechanism:
[0146] Shallow branching focuses on the fine-grained canopy texture of seedling plots, enhancing the contrast between young fruit trees and the bare soil boundary;
[0147] Deep branching extends the perception field, capturing the wide canopy distribution of 4-5 year old trees and the staggered boundary features between adjacent plots;
[0148] The multi-scale features output by each branch are aligned by bilinear interpolation to preserve the morphological gradation characteristics of plots of different tree ages;
[0149] Step 4: Classification Result Generation and Application Optimization: The cross-scale fused features are reconstructed using high-resolution methods to generate a classification result map that matches the UAV orthophoto at the pixel level. This accurately marks the continuous boundaries and transition areas of multi-aged Wogan mandarin orange orchard plots. The results are then imported into the orchard intelligent management system to guide differentiated fertilization, pest and disease control, and harvest planning, significantly reducing the cost of manual surveying and improving orchard management efficiency.
[0150] Shallow features of the MobileNetV2 backbone network are extracted from the third inverted residual block, with an output resolution of 1 / 4 of the input image (e.g., 512×512 input corresponds to 128×128 output) and 64 channels. These shallow features include the shoot density distribution (shoot spacing ≤ 15cm) and canopy edge gradient (Sobel operator gradient magnitude ≥ 0.3) for seedling plots. High-level features are extracted from the last layer of MobileNetV2, with a resolution of 32×32 and 320 channels, capturing the wide-area canopy distribution (canopy diameter ≥ 2m) of mature plots.
[0151] Feature extraction can be implemented using the PyTorch framework, and pre-trained weights for MobileNetV2 can be loaded from open-source libraries.
[0152] Workflow: High-level features are input into the CLA module to generate a spatial attention map, which is then weighted and connected to the original feature residuals.
[0153] Multi-scale features are unified to 128×128 resolution via bilinear interpolation and concatenated with shallow features to form a 384-channel (64+320) image. The fused features are input into two 3×3 convolutional layers; the first layer outputs 256 channels, and the second layer outputs 128 channels. The classification result image is upsampled fourfold to 512×512 resolution and exported in GeoTIFF format, with the coordinate system consistent with the UAV orthophoto. The orchard management system interface supports GeoJSON format, and the classification results are uploaded via REST API to guide the variable operation of the fertilizer applicator (fertilizer application gradient: 20 kg / mu for seedlings, 50 kg / mu for mature plants).
[0154] Workflow: After the fused features are convolutionally compressed and upsampled, a classification map is generated and converted into geospatial data format. The results are then uploaded to the orchard management system, automatically generating a differentiated fertilization prescription map to drive agricultural machinery.
[0155] Technical effects:
[0156] 1. Multi-level feature complementarity: shallow details and high-level semantics are jointly modeled to improve the ability to distinguish between seedling and mature plots.
[0157] 2. Attention-guided optimization: Channel semantics and spatial location are weighted together to suppress background interference and enhance pesticide patch response.
[0158] 3. Application system compatibility: The geospatial data format is seamlessly integrated with the agricultural machinery system, enabling closed-loop management from classification results to field operations.
[0159] <Application Example>
[0160] Taking the Wogan mandarin orange planting area in Luwo Town, Wuming District, Nanning City, Guangxi Zhuang Autonomous Region, China as the study area (geographic coordinates: 22°59′58″N~23°33′16″N, 107°49′26″E~108°37′22″E), based on the DJI Mavic 3 multispectral version UAV and the GACL-DeepLabV3+ model, we achieved fine classification of Wogan mandarin orange orchard plots of different ages. The specific steps are as follows:
[0161] Step 1: UAV Image Data Acquisition
[0162] The DJI Movavic 3 Multispectral UAV was used to perform data acquisition tasks, ensuring complete image coverage of the study area through a preset flight path. Flight parameters were set as follows: flight altitude of 250 meters to balance image resolution (0.0685m) and coverage efficiency; forward overlap of 80% and lateral overlap of 70% to ensure image stitching accuracy; data acquisition was conducted on a clear morning between 10:00 and 14:00, when the light intensity was 50,000-80,000 lux, which reduced shadow interference; and visible light images including red, green, and blue bands were acquired to enhance the distinguishability of vegetation features.
[0163] By using image sensors and optimizing flight parameters, the acquired raw images possess high resolution and low noise characteristics, providing a high-quality data foundation for subsequent preprocessing.
[0164] Step 2: Data Preprocessing and Dataset Construction
[0165] 2.1 Image Preprocessing
[0166] The original images were processed using Pix4Dmapper (v4.5.6) software: first, feature extraction and aerial triangulation were performed to generate dense point clouds, achieving geometric alignment between images; then, a digital surface model (DSM) was constructed and orthorectified to generate a high-precision orthophoto (DOM) with a resolution of 0.0685m.
[0167] 2.2 Manual annotation and dataset partitioning
[0168] Under the guidance of orchard experts and based on field investigations, the Wogan mandarin orange orchard plots were divided into five categories: seedling plots, 1-year-old tree plots, 2-3-year-old tree plots, 4-5-year-old tree plots, and plots sprayed with lime pesticides. The orthophotos were manually annotated with polygons using the Labelme tool. The annotated images were then cut into 512×512 pixel slices, and Gaussian noise injection (intensity 15, 25, 40) and rotation transformation (90° / 180° / 270°) were used to simulate real-world interference, ultimately generating 2032 slice images. The dataset was divided into a training set (1626 images), a validation set (203 images), and a test set (203 images) in an 8:1:1 ratio to ensure a balanced distribution of categories.
[0169] The preprocessing process improves image quality through geometric correction and noise suppression, while the annotation and enhancement strategies construct a multi-scale, interference-resistant dataset tailored to the characteristics of Wogan mandarin orange planting scenarios, laying the foundation for model training.
[0170] Step 3: Design of Channel-Aware Lightweight Attention Module (CLA)
[0171] We propose a channel-aware lightweight attention (CLA) approach, the core of which is to guide the generation of spatial attention weights through channel semantics, thereby enhancing the response in spatially salient regions. The specific process is as follows:
[0172] Channel semantic extraction: Global average pooling is performed on the input feature map x along the spatial dimension to obtain the overall response vector Z for each channel. A linear transformation is then used to obtain the channel semantically guided key vector K; the calculation formula is as follows:
[0173]
[0174] This vector is transformed by a linear transformation w k ∈R C'×C The key obtained is guided by channel semantics, denoted as K∈W. k ·K c ∈R C′ .
[0175] Spatial Query Generation: The input feature x is flattened into an H×W×C matrix, and a linear transformation is performed to obtain the Query matrix Q. That is, the input feature x is flattened into x * ∈R (H·W)×C And through linear transformation w q ∈R C'×C Obtain the Query matrix Q∈W q ·x * R C ′×(H·W) .
[0176] Attention map calculation: A spatial attention map A is generated by normalizing the dot product of Key and Query, using the following formula:
[0177] A = soft max(K) T ·Q)∈R 1×(H·W (1)
[0178] Feature weighting and residual connection: The attention map is multiplied element-wise with the original features, and the enhanced features are output through residual connection. The formula is: y=γ·A⊙x+x (2)
[0179] Where γ∈R is a learnable scaling factor.
[0180] Technical role: CLA enhances the model's responsiveness to key areas and reduces feature confusion between similar plots by modeling the correlation between channel semantics and spatial location.
[0181] Step 4: Design of Gated Axial Spatial Attention Module (GASA)
[0182] To enhance the model's ability to model spatial structure and target category boundaries in semantic segmentation tasks, a lightweight gated axial spatial attention (GASA) is proposed. The core of GASA is to combine the idea of axial attention with spatial attention and introduce learnable gating parameters to enhance the perception of local and long-range spatial information while maintaining low computational complexity.
[0183] The input feature map for GASA is x∈R C×H×W First, the input is normalized and channel-expanded. The expanded feature map is divided into two parts along the channel dimension. The first C channel features r1 and r2 are subjected to depthwise separable convolutions along the vertical and horizontal axes, respectively, to obtain response features in two directions. Then, at each spatial location (i,j), the two axial outputs are summed to form an attention map aij, as shown below:
[0184]
[0185] Where r ij ∈R d Let (i,j) be the attention vector at position (i,j). and This represents an axial convolution kernel that incorporates relative position encoding, corresponding to the responses of width and height, respectively.
[0186] The features X* of the last C channels are used as feature branches, modulated by the attention map described above, and then gated element-wise. The results are fused through 1×1 convolutions, and a learnable scaling parameter γ∈R is introduced. 1×C×1×1 Finally, add the result to the original input residual to obtain the final output:
[0187] y ij =γ·ψ(x ij ☉sigmoid(a ij ))+x ij ;
[0188] Where x ij Represents the original input features, and ψ(·) is the convolution function for channel fusion. This module structure does not require explicit construction of Query-key interaction.
[0189] Technical benefits: By separating spatial modeling and feature modulation paths, GASA reduces computational complexity (FLOPs are reduced by 30%) while improving the continuity of features in boundary areas, making it particularly suitable for contour segmentation of high-age plots (such as 4-5 year old trees).
[0190] Step 5: Design of Gated Axial Space Pyramid Structure (GASP)
[0191] GASP replaces the dilated convolutions of traditional ASPP with GASA modules to construct a multi-scale spatial modeling path. Its design includes: employing four parallel branches, each branch compressing channels through 1×1 convolutions and then introducing 1-4 layers of GASA modules; constructing a multi-level receptive field from local to global by incrementally increasing the number of GASA layers, thus overcoming the limitations of ASPP's fixed dilation rate; and after bilinear interpolation to unify the resolution of each branch's output, concatenating it with shallow features of the backbone network along the channel axis, and then fusing it through two 3×3 convolutions.
[0192] Technical advantages: GASP adaptively captures differences in plot shape through a dynamic attention mechanism, reducing misjudgments caused by background interference, and significantly reducing edge segmentation error compared to ASPP.
[0193] Step 6: Building the GACL-DeepLabV3+ model
[0194] The model is based on the DeepLabV3+ framework and integrates a lightweight backbone network and attention module. MobileNetV2 is used instead of Xception, and the parameter size is compressed to 5.876MB through inverse residual blocks and linear bottleneck design, making it suitable for deployment on drones. A CLA module is inserted after the high-level features of MobileNetV2 to guide the model to focus on key regions. The features processed by CLA are input into the GASP structure to generate multi-scale spatially perceptual features. Finally, the GASP output is concatenated with the shallow features, and the classification result map is output through two 3×3 convolutions and four times upsampling.
[0195] Model advantages: The combination of lightweight design and attention mechanism improves classification accuracy (mIoU 94.03%) while ensuring real-time performance.
[0196] Step 7: Model Training and Classification Validation
[0197] 7.1 Training Configuration
[0198] The hardware environment consisted of a Linux system, an NVIDIA A100 GPU, and the PyTorch framework. The optimization strategy used was the stochastic gradient descent (SGD) optimizer with an initial learning rate of 7e-3 and a cosine annealing method for the learning rate descent. The loss function was cross-entropy loss, which balanced classification accuracy with boundary continuity. The training parameters were set to Batchsize=4 and Epoch=100.
[0199] 7.2 Classification Validation
[0200] Input the test set into the trained model, and the output classification result is as follows: Figure 8-12 As shown, the performance metrics are: mean Intersection over Union (mIoU) 94.03%, mean pixel precision (mPA) 97.02%, mean precision (mPrecision) 96.82%, and overall accuracy (OA) 96.71%. The results indicate that the model achieves high classification accuracy for seedling plots (low vegetation cover) and plots sprayed with lime pesticides (spectral anomalies), effectively solving the problems of blurred boundaries and similar features in traditional methods.
[0201] The present invention proposes a method for fine classification of Wogan mandarin orange orchard plots based on UAV remote sensing images using GACL-DeepLabV3+. This method can effectively solve the problems of blurred boundaries, similarity to other tree species and crops, and overlapping features of bare ground and non-plot areas (such as dirt roads and wasteland) in UAV remote sensing images of Wogan mandarin orange orchard plots of different ages and conditions. This method achieves fine classification of Wogan mandarin orange orchard plots of different ages and conditions.
[0202] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.
Claims
1. A method for fine classification of Wogan orange orchard plots based on UAV remote sensing imagery using GACL-DeepLabV3+, characterized in that, Includes the following steps: S1: Use drones to acquire visible light remote sensing images of Wogan mandarin orange orchards of different ages and conditions in the study area; S2: Preprocess the acquired remote sensing images to generate orthophotos with a resolution of centimeters. The preprocessing includes feature extraction and aerial triangulation to generate dense point clouds, construction of digital surface models, and orthorectification. S3: Construct a dataset of Wogan mandarin orange orchard plots based on orthophotos. The dataset includes plots in the seedling stage, plots of 1-year-old orchards, plots of 2-3-year-old orchards, plots of 4-5-year-old orchards, and plots of orchards sprayed with lime pesticide. The images are labeled using a manual labeling tool, and the labeled images are then segmented. S4: Perform data augmentation on the dataset, including applying Gaussian noise with an intensity of 10-50 and rotation transformation to the images, generating a data augmentation training set containing noise robustness and multi-angle features, which is divided into training set, validation set and test set according to the proportion. S5: Construct the GACL-DeepLabV3+ model, which includes a MobileNetV2 backbone network, a channel-aware lightweight attention module, a gated axial spatial attention module, and a gated axial spatial pyramid structure. Train the GACL-DeepLabV3+ model using a data-augmented training set. During training, a stochastic gradient descent optimizer is used with an initial learning rate of 7e-3. The learning rate descent method is cosine annealing, the batch size is 4, and the number of training epochs is 100. S6: Input the test set into the trained GACL-DeepLabV3+ model and output the classification results of the Wogan mandarin orange orchard plots; Among them, the channel-aware lightweight attention module extracts the channel semantic information of the input features through global average pooling, generates a channel semantically guided key vector, generates a spatial query matrix through linear transformation, generates a spatial attention map by performing a dot product normalization on the key vector and the query matrix, and then outputs a weighted output through residual connection. The gated axial spatial attention module enhances the spatial feature modeling capability of Wogan mandarin orange orchard plots through axial convolution and gating mechanisms. Specifically, it divides the input features into two parts along the channel dimension. The first C channel features are subjected to depthwise separable convolution along the vertical and horizontal axes, respectively, to capture the differences in canopy structure and branch density of Wogan mandarin orange orchard plots in the vertical and horizontal directions, generate axial response features and sum them. The latter C channel features are adaptively modulated to the response features through learnable gating parameters to suppress background interference and enhance the boundary features of plots of different ages. Finally, the fused features are output through residual connection. The gated axial spatial pyramid structure constructs a multi-scale spatial modeling mechanism for Wogan mandarin orange orchard plots by stacking the gated axial spatial attention modules in layers. Specifically, it includes: processing the input features through multiple parallel branches, with each branch compressing the channel dimension through 1×1 convolution and then introducing 1-4 layers of the gated axial spatial attention modules. The receptive field is gradually expanded using modules of different numbers of layers. The shallow branches focus on the fine-grained canopy texture of Wogan mandarin orange seedling plots through single-layer modules, while the deep branches capture the wide-area canopy distribution features of 4-5 year old tree plots through multi-layer modules. The output multi-scale features are then unified in resolution through bilinear interpolation and spliced with the shallow detail features extracted by the MobileNetV2 backbone network, thus integrating the local branch density information and global spatial layout features of the Wogan mandarin orange orchard plots.
2. The method for fine classification of Wogan orange orchard plots based on UAV remote sensing imagery using GACL-DeepLabV3+ as described in claim 1, characterized in that, The operation of the UAV to acquire visible light remote sensing images further includes: establishing a flight parameter configuration table based on a preset fruit tree age gradient database and canopy morphology feature database; using a multi-rotor UAV equipped with a visible light multispectral imager with polarization filtering function; conducting aerial photography with the canopy layer as the reference plane during the morning diffuse light period; for young trees less than 3 years old, a flight altitude of 250 meters and a forward overlap rate of 80% are used, while for mature, high-yielding trees, a flight altitude of 250 meters and a forward overlap rate of 70% are used; and configuring an image quality real-time diagnosis module based on edge computing units to perform online detection and re-photographing decisions on leaf texture clarity and abnormal fruit coloring areas.
3. The method for fine classification of Wogan orange orchard plots based on UAV remote sensing imagery using GACL-DeepLabV3+ as described in claim 1, characterized in that, The manual annotation tool is Labelme, and the annotated image segmentation adopts an adaptive block segmentation algorithm based on saliency guidance.
4. The method for fine classification of Wogan mandarin orange orchard plots based on UAV remote sensing imagery using GACL-DeepLabV3+ as described in claim 1, characterized in that, The fusion mechanism of multi-scale features and shallow features output by the gated axial spatial pyramid structure is designed for the morphological complexity of Wogan mandarin orange orchard plots, and specifically includes the following steps: Step 1: Shallow feature extraction and preservation: Based on the shallow features extracted by the MobileNetV2 backbone network, the fine-grained branch density distribution characteristics and canopy edge texture details of the Wogan seedling plot are preserved to enhance the ability to distinguish between young fruit tree plots and bare soil and non-plot areas such as roads. Step 2: Multi-scale feature alignment and adaptation: The multi-scale features output by the gated axial spatial pyramid are upsampled using a bilinear interpolation algorithm to make their spatial resolution consistent with the shallow features extracted by the MobileNetV2 backbone network, thereby achieving geometric alignment of cross-scale features. Step 3: Cross-scale feature fusion and boundary enhancement: The aligned multi-scale features and shallow features are concatenated along the channel dimension, and cross-scale information is fused through two 3×3 convolutional layers: First convolution: Extract the characteristics of branch density gradient changes within the Wogan mandarin fruit plots to enhance the response to differences in canopy coverage among plots of different tree ages; Secondary convolution: captures the patchy distribution characteristics of lime-sprayed pesticide plots and the scattered boundary abrupt change signals of young fruit tree plots, and suppresses missegmentation of adjacent plots caused by canopy overlap or background interference under mixed planting patterns. Step 4: High-resolution classification map reconstruction: The fused features are upsampled four times to restore the original image resolution, generating a classification result map that matches the UAV orthophoto at the pixel level, accurately marking the continuous boundaries and internal structure of Wogan mandarin orange orchards of different ages.
5. The method for fine classification of Wogan orange orchard plots based on UAV remote sensing imagery using GACL-DeepLabV3+ as described in claim 4, is characterized in that, The construction of the GACL-DeepLabV3+ model is designed for a multi-age mixed planting pattern in Wogan mandarin orange orchards, and specifically includes the following steps: Step 1: Multi-level feature extraction from the input image: Input the visible light image of the Wogan mandarin orange orchard plot acquired by the UAV into the MobileNetV2 backbone network to extract shallow and high-level features respectively: Shallow features: Preserve the fine-grained branch density distribution, canopy edge texture, and soil characteristics of the exposed surface of young plots to distinguish non-fruit tree areas; High-level characteristics: Capture the wide-area canopy distribution pattern of 4-5 year old plots and the patchy spectral response of pesticide-sprayed plots to characterize the global spatial layout of mature Wogan mandarin orange trees; Step 2: Channel-aware attention-guided spatial feature enhancement: The aforementioned channel-aware lightweight attention module is embedded after the high-level features to optimize the semantic discriminative ability of Wogan orange plot classification through the following operations: Channel semantic information of high-level features is extracted based on global average pooling, and a channel semantically guided key vector is generated. The spatial location information is converted into a query matrix, and the correlation weight between channel semantics and spatial location is calculated. Residual connectivity is used to enhance the response intensity to sparse shoot areas in seedling plots and local patches in pesticide-sprayed plots, thereby suppressing background noise interference. Step 3: Multi-scale spatial modeling and tree age feature adaptation: The high-level features enhanced by the channel-aware lightweight attention module are input into the gated axial spatial pyramid structure to adapt to the age differences of the Wogan mandarin orange orchard plots through the following mechanism: Multi-branch structure expands the receptive field, deeply extracts data texture features, enhances the contrast between young fruit trees and bare soil boundaries, and captures the wide canopy distribution of 4-5 year old plots and the staggered boundary features between adjacent plots. The output multi-scale features are aligned by bilinear interpolation to preserve the morphological gradation characteristics of plots with different tree ages; Step 4: Classification Result Generation and Application Optimization: The cross-scale fused features are reconstructed using high-resolution methods to generate a classification result map that matches the UAV orthophoto at the pixel level. This accurately marks the continuous boundaries and transition areas of multi-aged Wogan mandarin orange orchard plots. The results are then imported into the orchard intelligent management system to guide differentiated fertilization, pest and disease control, and harvest planning.