Citrus tree management method based on image semantic segmentation

By adopting image semantic segmentation technology in complex orchard environments, combining CSFAN and multiple feature interaction modules, the problems of intra-class differences, ground-like confusion and natural noise robustness in citrus tree detection are solved, and the precise segmentation and efficient detection of citrus tree are achieved.

CN120182976AInactive Publication Date: 2025-06-20CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510256311.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, citrus tree detection has problems such as significant in-class differences, confusion of land, insufficient adaptability of multi-scale targets, and poor robustness to natural noise in complex orchard environments.

Method used

Using a method based on image semantic segmentation, an encoder-decoder architecture for cascading feature aggregation of CSFAN is constructed by using a drone to capture images in the air over time, and combining BVPE, MDFPB and RFM modules to perform cross-spatial domain feature interaction and frequency information refinement to achieve accurate segmentation of citrus trees.

Benefits of technology

It significantly improves the recognition accuracy and robustness of citrus trees in complex backgrounds, improves the generalization ability and real-time detection efficiency of the model, and supports the high-frequency detection needs of precision agriculture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182976A_ABST
    Figure CN120182976A_ABST
Patent Text Reader

Abstract

The invention discloses a citrus tree management method based on image semantic segmentation. The method comprises the following steps: step 1, shooting an image of a citrus orchard; 2, the shot original images are screened, and a data set used for training is obtained; 3, accurately marking the reserved image, carrying out post-processing, and reasonably dividing a data set at the same time; step 4, constructing an encoder-decoder architecture of CSFAN cascade feature aggregation, and obtaining heterogeneous features through a deep residual network in an encoding stage; 5, applying spatial constraints to the heterogeneous features through an RFM module, and effectively screening and optimizing information; 6, a decoder processes the features obtained in the step 5 through a BVPE module and an MDFPB module, decoding is conducted through a multi-level composite loss supervision system, and accurate segmentation of the citrus tree is achieved; the accurate segmentation result of the citrus tree is applied to intelligent management of an orchard, and disease spots and newborn shoot conditions are monitored in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, in particular to remote sensing image processing technology, and specifically relates to a citrus tree management method based on image semantic segmentation. Background Art

[0002] As the world's largest producer and consumer of citrus fruits, the citrus industry holds an important economic position in the rural revitalization strategy. With the rapid development of high-resolution remote sensing technology, agricultural management is facing an urgent need to transform towards precision and intelligence. However, there are a series of key problems to be solved in the current field of citrus tree detection, especially in the support of large-scale agricultural monitoring and precision agriculture applications. First, the traditional manual census method is inefficient and unreliable, making it difficult to meet the high-precision and high-efficiency management requirements in large-scale agricultural production. Especially in citrus orchards with dense trees, the mode relying on manual visual recognition is difficult to ensure systematicness and repeatability, significantly restricting the advancement of modern agricultural management. Second, the existing semi-automated image processing technology has poor adaptability in complex environments. Especially under adverse conditions such as intercropping of fruit trees, foliage occlusion, or extreme weather, the accuracy of existing recognition algorithms drops significantly, failing to effectively capture the detailed features of fruit trees and difficult to meet the refined management requirements of precision agriculture. Third, when facing large-scale agricultural monitoring, the existing monitoring systems have significant bottlenecks in data processing capabilities and response speeds. Data processing delays and feedback lags may affect the timeliness of pest control, thereby increasing potential agricultural risks. Finally, the construction of a link in the smart orchard has not reached an ideal level of integration and optimization, lacking an efficient real-time monitoring, data collection, processing, and feedback mechanism, resulting in time lags in data flow and decision-making responses between various links, and it is difficult to achieve data-driven precision agricultural decision-making and efficient collaboration. Overall, the current technical system has not been able to fully support the real-time and precision requirements of large-scale agricultural monitoring, nor effectively promote the full implementation and practice of smart agriculture applications.

[0003] Deep learning technology based on convolutional neural networks provides a new path to solve the above problems. Through a multi-scale feature fusion network, citrus trees with different crown widths can be effectively identified, showing strong detection capabilities in remote sensing images. Its spatial pyramid structure can adapt to diverse planting environments and achieve progressive modeling from local details to global semantics through hierarchical feature extraction. However, the industrial application of this technology still faces key bottlenecks: the regional limitation of training data leads to insufficient model generalization ability, and it is vulnerable to geographical environment differences during cross-regional deployment; the problem of feature confusion under complex background interference has not been fundamentally solved, especially in areas where vegetation with similar spectral characteristics to citrus trees is mixed; the contradiction between model operation efficiency and real-time detection requirements is becoming increasingly prominent. Existing algorithms face serious computational redundancy problems when processing high-resolution remote sensing images, resulting in the analysis time of a single-frame image far exceeding the threshold allowed by actual operations, making it difficult to support the high-frequency detection requirements in continuous dynamic monitoring scenarios. Especially when dealing with sudden pests and diseases, the response delay of traditional computing architectures may lead to the missed critical prevention and control window period, which has become an important bottleneck restricting the practical application of intelligent detection systems.

[0004] More severely, the lack of a dedicated dataset for citrus trees in existing orchards seriously restricts the implementation of the technology. Public datasets generally have problems such as insufficient sample size, single coverage scenario, inconsistent annotation standards, etc., and lack time-series data resources covering different growth stages, pest and disease characteristics, and climatic conditions. The weakness of this data foundation directly leads to the difficulty for the model to capture the dynamic growth laws of citrus trees and is even less able to support the precise detection requirements across production areas and varieties, severely limiting the practical application value of intelligent detection systems in the citrus industry. These technical bottlenecks and data shortfalls form a double constraint, urgently requiring breakthroughs through algorithm innovation, data system construction, and hardware collaborative optimization to provide reliable technical support for smart agriculture.

[0005] Recently, with the rise of the research direction of remote sensing semantic segmentation, researchers in different fields have begun to explore how to guide the model to learn target features that are difficult to capture in complex backgrounds. For example, the YongChang Xie team proposed a new multi-spatial constraint learning network in the complex scene in the paper "MMLN: Multi-directional and Multi-constraint Learning Network for Remote Sensing Imagery Semantic Segmentation", which enhances the interaction between spatial domains to cope with the confusing category features in complex scenes. However, such methods face many challenges in complex remote sensing scenes, including significant intra-class differences of targets, ground object confusion, insufficient adaptability to multi-scale targets, and lack of robustness to natural noise, which seriously restrict their effectiveness in practical applications. Therefore, it is still a certain challenge to design a method that uses cross-spatial domain features to guide the network to learn citrus tree features in the complex orchard scene. Summary of the Invention

[0006] Object of the Invention: The present invention aims to solve many challenges existing in the complex orchard environment, such as significant intra-class differences, ground object confusion, insufficient adaptability to multi-scale targets, and poor robustness to natural noise.

[0007] To solve the above technical problems, the present invention discloses a citrus tree management method based on image semantic segmentation, including the following steps:

[0008] Step 1: Use a drone to take aerial images of the citrus orchard at multiple times.

[0009] Step 2: Screen the taken original images, retain the images with rich categories and diverse ground object features, and eliminate those images that only contain a single category or have insufficient information to obtain a dataset for training.

[0010] The single category: refers to an image that only contains one type of ground object category (such as only vegetation) and lacks other relevant elements (such as soil, weeds, roads, equipment, or fruit trees at different growth stages); insufficient information: refers to an image that cannot provide effective information due to quality or content defects, such as blurring, overexposure, low resolution, repeated scenes, or single-angle shooting.

[0011] Step 3: Accurately annotate the retained images, label each ground object category with an accurate label, and perform post-processing including: removing non-standard labels, correcting boundary details, ensuring the accuracy and consistency of the labels, and reasonably dividing the dataset.

[0012] The dataset division: The screened data science is divided into a training set (for model learning), a validation set (for hyperparameter tuning), and a test set (for final evaluation). Here, the ratio is approximately 6:2:2;

[0013] Step 4: Construct an encoder-decoder architecture with CSFAN-level cascaded feature aggregation. In the encoding stage, multi-level feature extraction is implemented through the deep residual network ConvNext-tiny, and a cross-layer feature interaction strategy is adopted, that is, the feature maps of each level are concatenated with the features of their adjacent scales in the channel dimension to obtain heterogeneous features;

[0014] Step 5: Apply spatial constraints to the heterogeneous features obtained in Step 4 through the RFM module. After effectively screening and optimizing the information, it is passed to the decoder for further processing;

[0015] Step 6: The decoder performs cross-spatial domain feature interaction on the features passed in Step 5 through the BVPE module, and refines the frequency information through the MDFPB module. A multi-level composite loss supervision system is used for decoding to achieve the precise segmentation of citrus trees.

[0016] In Step 4, the ConvNext-tiny model is used as the encoder to extract features in four stages, which are represented as:

[0017]

[0018] Among them, Input represents the input feature map, F1, F2, F3, and F4 respectively represent the feature maps extracted by the ConvNext-tiny model in four stages. Each stage's ConvNexBlock k represents a block operation. k corresponds to the stage and takes values of 1, 2, 3, or 4. Starting from the second stage, the input of each stage is the output feature map of the previous stage. In this way, multi-level feature extraction is performed to obtain features F1, F2, F3, and F4. The shallow and deep feature maps of features F1, F2, F3, and F4 are concatenated to obtain heterogeneous features.

[0019] In Step 6, the BVPE module performs cross-spatial domain feature interaction as follows:

[0020] Step 6-1-1: According to the spatial dimension, the input feature map Input is divided into two parts through the split function, namely Input_1 and Input_2;

[0021] Step 6-1-2: Linear transformations are respectively performed on Input_1 and Input_2 to generate query vector Q1, key vector K1, and two value vectors V1 and V2;

[0022] Step 6-1-3: Calculate the dot product of the query vector Q1 and the key vector K1, scale the dot product result to the dimension d of the query vector, then perform a Hadamard product calculation with the scalar Λ1, and normalize the calculation result through the Sigmoid function to obtain the self-attention weight map S1;

[0023]

[0024] Λ1 = exp(λ q ·λ k )

[0025] λ is a learnable scaling factor used to adjust the similarity between the query and the key. The value range of λ is usually from 0 to 1. Increasing it will enhance long-range dependencies and is suitable for long-range tasks, but it will lose local details. Decreasing it will enhance short-range dependencies and is suitable for local structure modeling, but it will weaken long-range logic. By dynamically adjusting λ, a balance is achieved between long-range and local modeling. Its subscripts q and k correspond to the position encoding parameters of the query vector and the key vector respectively. Λ1 is a dynamic position weight matrix, representing the intensity of the attention weight, used to adjust the attention score. Λ1 is obtained by applying the exponential function to the dot product result of λ q and λ k . The position weight matrix Λ1 is multiplied and superimposed on the attention score through the Hadamard product to dynamically adjust the association strength of different positions;

[0026] Step 6-1-4: Perform weighted summation by multiplying the weight map S1 and the feature map V1 through dot product to obtain the feature map F1 with attention weights; the weight map S1 acts on the feature map V2, and the second feature map F2 with attention weights is obtained through weighted summation; the calculation formulas of F1 and F2 are as follows:

[0027] F1 = S1·V1

[0028] F2 = S1·V2

[0029] Step 6-1-5: Subtract F1 and F2 to obtain the differential feature map D1, and the formula is as follows;

[0030] D1 = F1 - F2

[0031] Step 6-1-6: Perform channel screening on the differential feature map D1 and F1 and F2 respectively, that is, screen the feature information in the channels to generate two channel attention weight matrices W1 and W2, specifically as follows:

[0032] E1 = ChannelAttn(reshape(F1))

[0033] E2 = ChannelAttn(reshape(F2))

[0034] E3 = ChannelAttn(reshape(D1))

[0035] Strengthen the F1, F2 calculated in step 6-1-4 and their differential information D1 in the channel dimension, that is, change the input dimension through reshape to match the input dimension requirements of the channel screening ChannelAttn, filter some redundant information through channel screening and enhance the channel information with larger weights, so as to obtain the channel-enhanced feature matrices E1, E2 and E3;

[0036] W1 = Softmax(E3(E1) T )

[0037] W2 = Softmax(E3(E2) T )

[0038] Perform transposed inner product calculation on the above channel-enhanced feature matrices E1, E2 and E3, and apply the Softmax function to the result of the inner product calculation to generate the attention weight matrices W1 and W2;

[0039] Step 6-1-7: Multiply W1 by F1 and perform residual connection (the addition operation in the corresponding formula) to obtain the attention feature map A1, and multiply W2 by F2 and perform residual connection to obtain the attention feature map A2;

[0040] A1 = W1·F1 + F1

[0041] A2 = W2·F2 + F2

[0042] By calculating the attention weight matrix and updating the weighted attention result, and then enhancing the feature representation through residual connection, the feature outputs A1 and A2 in different spatial domains are finally obtained;

[0043] Step 6-1-8: Concatenate the attention feature maps A1 and A2 in the spatial dimension to obtain the output OutPut that fuses the information from different spatial domains, that is, the final BVPE output result:

[0044]

[0045] In the formula, represents that the feature dimension belongs to B×N×C, B: batch size, indicating the number of samples processed at one time, C: channel dimension, indicating the embedding vector dimension or hidden layer dimension of each element. N: sequence length, indicating the number of elements in a single sample, and OutPut represents the output.

[0046] In step 6, the MDFPB module performs frequency information refinement as follows:

[0047] Step 6-2-1: Pass the input feature map to the discrete wavelet transform DWT for Haar wavelet transform to obtain four feature maps with different frequencies, namely LL, LH, HL, and HH; where LL represents the low-frequency part, containing the main information and general structure of the image; LH, HL, and HH respectively represent the high-frequency detail information in different directions, including the edge features in the horizontal, vertical, and diagonal directions;

[0048] Step 6-2-2: The four frequency feature maps LL, LH, HL, and HH are concatenated by channel dimension to obtain a new feature map Frequency_All;

[0049] Step 6-2-3: Perform a frequency information direction-aware enhancement operation on the feature map Frequency_All. This operation performs feature extraction and enhancement through three different branch structures respectively;

[0050] In the upper branch, perform convolution operations of central difference, horizontal difference, and vertical difference on the input image in sequence. The central difference convolution captures the detail changes in the image, the horizontal difference convolution extracts the horizontal edge information of the image, and the vertical difference convolution extracts the vertical edge information of the image. Through these difference convolution operations, the different direction details of the image are fused with each other. Finally, the results of all convolution operations are summed up to obtain a feature map M_F1 containing the detail features in each direction;

[0051] The lower branch uses strip-shaped convolution operations of 1×3 and 3×1. This operation can capture frequency information at different scales in the horizontal and vertical directions respectively. Through the 1×3 convolution, the horizontal texture features of the image are captured; while the 3×1 convolution focuses on the feature extraction in the vertical direction. Then, after the 1×1 convolution, the results of the two convolutions are integrated, and finally the feature map M_F3 is generated;

[0052] The middle branch performs a 3×3 convolution operation on the input. This operation can effectively extract the global features of the image, capture the global context information in the image through a larger convolution kernel. After convolution, the feature map is screened by channels to select the most representative and important channel features, and the feature map M_F2 is obtained;

[0053] Step 6-2-4: Apply the Sigmoid function normalization operation to the feature map F2 to map its values to between 0 and 1, generating two weight matrices M_W1 and M_W2;

[0054] Step 6-2-5: Multiply M_W1 by the feature map M_F1 of the upper branch to obtain the weighted feature map M_F4, and multiply M_W2 by the feature map M_F3 of the lower branch to obtain the weighted feature map M_F5;

[0055] Step 6-2-6: M_F4 and M_F5 are fused through 1×1 convolution to obtain the final output result.

[0056] 5. A citrus tree management method based on image semantic segmentation according to claim 4, wherein: in step 6-2-1, the discrete wavelet transform DWT is specifically as follows:

[0057] Step 6-2-1-1: Perform preliminary filtering through a 2×1 strip convolution kernel to extract low-frequency (smooth) and high-frequency (detail) information, obtaining two feature maps: High_frequency and Low_frequency;

[0058] Step 6-2-1-2: Apply a horizontal convolution operation with a convolution kernel of 2×1 and a vertical convolution operation with a convolution kernel of 1×2 to the High_frequency and Low_frequency feature maps respectively to further extract different-direction details of the signal in the two-dimensional space; obtain four frequency feature maps representing low-frequency - low-frequency, low-frequency - high-frequency, high-frequency - low-frequency, and high-frequency - high-frequency: LL, LH, HL, and HH.

[0059] In step 5, the RFM module filters the high-frequency and low-frequency spatial channel information, specifically as follows:

[0060] Step 5-1: Receive the heterogeneous features passed by the encoder and fuse them through 1×1 convolution to obtain the feature map X1;

[0061] Step 5-2: Input the feature map X1 into a 3×3 convolution layer to obtain the local feature map X1_3×3, and generate the weight matrix X1_W3 through AVGPOOL and Softmax operations; perform spatial information screening on the feature map X1, and obtain the compressed spatial feature maps X_AVG and Y_AVG by performing average-pooling operations on the X-axis and Y-axis respectively;

[0062] Step 5-3: Use the Sigmoid function to normalize X_AVG and Y_AVG to generate the weight maps R_W1 and R_W2;

[0063] Step 5-4: Multiply the compressed spatial feature maps X_AVG and Y_AVG element-wise with the weight maps R_W1 and R_W2 to obtain the spatial screening feature maps R_A1 and R_A2;

[0064] Step 5-5: Concatenate the two feature maps R_A1 and R_A2 along the channel dimension to obtain the final feature map X1_A1;

[0065] Step 5-6: Multiply the feature map X1_A and the weight matrix X1_W3 to obtain the spatially filtered local feature information X1_A2; Multiply the local feature map X1_3×3 obtained in Step 7-2 and the weight map X1_A1_W of the feature map X1_A1 to obtain the weighted local feature map X1_A3;

[0066] Step 5-7: Add the spatially filtered local feature information X1_A2 and the local feature map X1_A3 to obtain the fused spatial perception feature and weighted local feature map X1_X2;

[0067] Step 5-8: Normalize the feature map X1_X2 to generate the weight map X1_X2_W, and perform weighted calculation with the feature map X1 to obtain the refined output.

[0068] In Step 6, the multi-level composite loss supervision system is as follows: At each feature scale during the decoding process, pixel-level smoothed factor cross-entropy loss and contour-aware Dice loss are synchronously applied, combining the main loss and the auxiliary loss, and through interpolation and weighting of features of different layers, it is expressed as follows:

[0069]

[0070] Case 1: The main output uses the main loss, and the auxiliary output uses the joint loss, for layers 2 to 4; Case 2: The number of output values of the model is 2, the main output uses the main loss, and the auxiliary output uses the auxiliary loss, when the condition of Case 1 is not satisfied, it is used for the first layer; Condition 3: The model only returns the main output and directly uses the main loss;

[0071] Among them, x represents the feature map obtained from each layer; y represents the true value corresponding to x; g ф represents the auxiliary output of the corresponding layer, k represents the corresponding layer, L total represents the multi-level loss function, f θ (x) represents the main output, H k (x) represents the auxiliary output of the corresponding layer, L main represents the main loss, L aux represents the auxiliary loss, L joint represents the joint loss, the hyperparameter λ represents the weight of the corresponding loss, and the value range is between 0 and 1,

[0072] It means that during the training process, the loss is calculated according to different conditions. If the network output has two parts, the main loss L main and the auxiliary loss L aux are calculated respectively, and weighted according to the hyperparameter λ to achieve more precise optimization; When there are multiple auxiliary outputs (features H from the intermediate layer of the network K(x), where k ∈ {2, 3, 4}), the joint loss of each auxiliary output will be calculated to enhance the learning of features at different scales. The design of this formula aims to promote the co-optimization of the network at different feature scales through a multi-level loss supervision mechanism, thereby improving the overall performance and generalization ability of the model.

[0073] The joint loss L joint is calculated as follows:

[0074] L joint = αL SCE + βL Dice L joint combines two loss functions: weighted cross-entropy loss L sce and Dice loss L Dice , where α and β are hyperparameters that control the weights of these two loss functions, and their value ranges are both between 0 and 1.

[0075] The weighted cross-entropy loss L sce calculates the pixel-level classification loss, which measures the difference between the predicted class and the true label.

[0076] The weighted cross-entropy loss L sce measures the classification performance of the model by calculating the logarithmic loss between the true label and the predicted probability. The formula is as follows:

[0077]

[0078] In the formula, H and W are the height and width of the image respectively, H×W represents the number of pixels in each image, the number of classes C represents the total number of possible classes in the prediction task, c represents the class, j represents the j-th class, i represents the pixel at the i-th spatial position of the network, represents the probability that pixel i belongs to class c, y i,c represents the original target label of the i-th pixel, represents the smoothed target label of the i-th pixel, z i,c represents the original predicted value of the network for the pixel at the i-th spatial position belonging to the c-th class, z i,j represents the original predicted value of the network for the pixel at the i-th spatial position corresponding to other classes j, ∈ represents a label smoothing factor, and its value range is between 0 and 1. Its role is to adjust the distribution of the true label to prevent the model from overfitting to the training data.

[0079] The Dice loss L Dice measures the accuracy of the model's image segmentation by calculating the Dice coefficient. The smaller the value of this loss function, the closer the model's prediction is to the true label. The formula is as follows:

[0080]

[0081] In the formula, i represents the pixel at the i-th spatial position corresponding to the network, and p i represents the predicted value of the model at the pixel i, which is a probability value, while y i represents the true label of the pixel; H and W represent the height and width of the image, and H×W represents the number of pixels in each image; γ is a smoothing factor to avoid a zero denominator, and its value ranges from 0.1 to 10.

[0082] Beneficial effects:

[0083] 1. The newly collected orchard citrus tree semantic segmentation dataset of the present invention covers high-resolution remote sensing images of different seasons, weather conditions, and different growth stages, with a pixel resolution of 6000×4000. This dataset not only includes orchard images under normal weather conditions, but also specifically includes image data under foggy and defogged weather conditions, fully considering the impact of climate change on the growth of citrus trees and the monitoring accuracy. By performing precise semantic segmentation processing on these images of multiple time periods and multiple scenarios, this dataset can effectively support the automatic recognition of citrus trees and the monitoring of pests and diseases, improve the efficiency of precision agriculture management, and provide reliable data support for the construction of intelligent orchards and the development of real-time monitoring systems.

[0084] 2. The Boundary Variation Perception Enhancer (BVPE), Multi-directional Frequency Processing Block (MDFPB), and Redundancy-Free Fusion Module (RFM) proposed in the present invention together constitute an innovative citrus tree segmentation solution. BVPE adopts a cross-spatial domain feature interaction mechanism. Through self-attention weight calculation and spatial information differentiation enhancement, it accurately captures context semantic features with significant distinctiveness. Especially in orchard scenarios, it can effectively identify the subtle differences between citrus trees and surrounding vegetation, ensuring pixel-level accurate segmentation of citrus trees. MDFPB, on the other hand, refines and extracts target gray-scale features by introducing frequency information of wavelet transform, resists the interference of natural noise, and addresses the challenges of target segmentation under complex backgrounds such as crown overlap and branch and leaf occlusion, significantly improving the robustness and generalization ability of the network. RFM, through a dual-path feature screening mechanism, fuses high-frequency texture features and low-frequency semantic features, and uses cross-frequency domain information interaction guided by channel importance to alleviate the problem of uneven class distribution between citrus trees and background vegetation. Through multi-level feature complementarity and redundancy suppression, RFM significantly improves the segmentation accuracy of citrus groves and the generalization ability of the model, ensuring the clear boundary of citrus trees and the connectivity of dense crown areas in orchard remote sensing monitoring. Combining these three innovative technologies can provide reliable technical support for precision agriculture and orchard management, greatly improving the accuracy and stability of citrus tree segmentation. At the same time, a multi-level composite loss supervision system is innovatively designed, which synchronously applies pixel-level smoothed factor cross-entropy loss and contour-aware Dice loss at each feature scale during the decoding process, combines the main loss and the auxiliary loss, and through interpolation and weighting of features at different layers, the model achieves a citrus crown segmentation accuracy of 85.10% on unmanned aerial vehicle (UAV) aerial images, which is 4.56 percentage points higher than the baseline model. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] Figure 1 is the algorithm flow chart of the present invention;

[0086] Figure 2 is the structural schematic diagram of the CSFAN (Cross-Spatial Frequency Attention Network) algorithm model in the present invention;

[0087] Figure 3 is the structural schematic diagram of the BVPE module;

[0088] Figure 4 is the structural schematic diagram of the MDFPB module;

[0089] Figure 5 It is a structural diagram of the RFM module;

[0090] Figure 6 Schematic diagram of the Haar wavelet transform structure. DETAILED DESCRIPTION

[0091] Example 1

[0092] A citrus tree management method based on image semantic segmentation comprises the following steps:

[0093] Step 1: Use drones to take aerial photos of citrus orchards in Yiling District, Yichang City at multiple time periods. The obtained 6000×4000 resolution high-quality images cover the growth cycles of spring, summer, autumn and winter, helping to monitor the conditions of citrus trees at different growth stages. At the same time, it also includes images under foggy and defogging weather conditions, further improving the adaptability to changes in the orchard environment and the diversity of image data;

[0094] Step 2: Screen the original images, retain images with rich categories and diverse features, and remove those containing only a single category or insufficient information;

[0095] Step 3: Use image annotation tools to accurately annotate the retained images, label each feature category accurately, and ensure the quality of the labels through post-processing. At the same time, the data set is reasonably divided for training, verification, and testing.

[0096] Step 4: Construct the encoder-decoder architecture of CSFAN cascade feature aggregation. In the encoding stage, multi-level feature extraction is implemented through a deep residual network (ConvNext-tiny), and a cross-layer feature interaction strategy is used to splice the feature maps of each level with the features of the adjacent layers in the channel dimension.

[0097] Step 5: The information screening module applies spatial constraints to the spliced ​​heterogeneous features, effectively screens and optimizes the information, and then passes it to the decoder for further processing;

[0098] Step 6: The decoder performs cross-spatial feature interaction and frequency information refinement on the features obtained in step 5, and finally achieves accurate segmentation of the citrus tree. Based on the above steps, a complete citrus tree segmentation model is constructed;

[0099] The precise segmentation results of citrus trees are applied to intelligent orchard management to monitor disease spots and the condition of new shoots in real time.

[0100] In Step 1, a drone is used to conduct multi - period aerial photography of citrus orchards in Yiling District, Yichang City, to obtain high - quality images with a resolution of 6000×4000, covering different growth cycles in spring, summer, autumn and winter. It can accurately monitor the growth of citrus trees at various stages such as budding, flowering, fruiting and ripening. The image acquisition not only covers normal weather conditions, but also specifically includes images in foggy and defogged weather, enhancing the adaptability of the orchard under different climate and environmental conditions, helping to improve the diversity of image data, and providing more comprehensive support for the health monitoring of citrus trees, pest and disease detection, etc.;

[0101] In Step 2, when screening the original images, first, they are classified according to the land cover types contained in the images, and those images containing multiple different land cover features are preferentially retained. For example, an image containing multiple elements such as citrus trees, surrounding vegetation, soil, water bodies, etc. will be retained because it provides rich scene information, which helps to improve the model's recognition ability for complex environments. At the same time, those images containing only a single category, such as images containing only citrus trees or only a background grassland, will be excluded because their features are too single and lack diversity, and they cannot effectively support the training of the model. Images with insufficient information, such as blurred, overly dark or severely occluded images, will also be excluded because they cannot provide clear and effective feature information. Through this strict screening, the finally retained images can provide more diverse, complete and high - quality land cover information, laying a more solid foundation for subsequent analysis and modeling;

[0102] In Step 3, when using an image annotation tool to accurately annotate the retained images, first, each land cover type (such as citrus trees, grasslands, shrubs, water bodies, etc.) is manually annotated to ensure that the areas of each category are accurately framed and labeled. During the annotation process, special attention is paid to the fine adjustment of the boundaries to ensure the accuracy of the labels. After the annotation is completed, post - processing steps are carried out, mainly including removing overlapping labels, correcting incorrect annotation areas, and unifying the label format to ensure the quality of the dataset. Subsequently, the dataset is reasonably divided according to a certain ratio (such as 70% for training, 15% for validation, and 15% for testing) to ensure the diversity and balance of the training, validation and test data, providing high - quality basic data for the training and performance evaluation of the model.

[0103] As Figure 2 shown, the constructed CSFAN model is specifically as follows:

[0104] In constructing the encoder - decoder structure of cascaded feature aggregation, a deep residual network (ConvNext - tiny) model is used as the encoder to extract four - stage features, which can be expressed as:

[0105]

[0106] Among them, F1, F2, F3, and F4 respectively represent the feature maps extracted by the four stages of the ConvNext-tiny model. Each stage's ConvNexBlock k represents a block operation. k corresponds to the stage and takes values of 1, 2, 3, or 4. Starting from the second stage, the input of each stage is the output feature map of the previous stage. The finally obtained feature map F4 will contain rich hierarchical features of the input image. For example, by performing multi-level feature extraction to obtain features F1, F2, F3, and F4, the shallow feature maps and deep feature maps of features F1, F, F, and F4 are concatenated to obtain heterogeneous features.

[0107] Such as Figure 2 As shown, the specific training process of the CSFAN model is as follows:

[0108] The image dataset is parsed through the first Conv convolution and pooling operations to obtain the feature map Feature0;

[0109] The feature map Feature0 is input into the first ConvNextBlock block for parsing and downsampling to obtain the feature map Feature1, whose number of channels is 64;

[0110] The feature map Feature1 is input into the second ConvNextBlock block for parsing and downsampling to obtain the feature map Feature2, whose number of channels is 128;

[0111] The feature map Feature1 and the feature map Feature2 are concatenated in the channel dimension to obtain Feature1_2;

[0112] Feature1_2 is input into the first RFM module, and the feature map Feature1_skip is obtained by filtering the spatial and channel information of high-frequency and low-frequency;

[0113] The feature map Feature1_skip and the feature map Feature2 are concatenated in the channel dimension to obtain Feature1skip_2;

[0114] Feature1skip_2 is input into the second RFM module, and the feature map Feature2_skip is obtained by filtering the spatial and channel information of high-frequency and low-frequency;

[0115] The feature map Feature2 is input into the third ConvNextBlock block for parsing and downsampling to obtain the feature map Feature3, whose number of channels is 256;

[0116] Concatenate the feature map Feature2_skip and the feature map Feature3 along the channel dimension to obtain Feature2skip_3;

[0117] Input Feature2skip_3 into the third RFM module, and filter the spatial and channel information of high-frequency and low-frequency to obtain the feature map Feature3_skip;

[0118] Input the feature map Feature3 into the fourth ConvNextBlock block for parsing to obtain the feature map Feature4;

[0119] Use the feature map Feature4 as the feature of the skip connection to be passed from the encoder to the decoder. Input the feature map Feature4 into the fourth BVPE module for cross-spatial self-attention calculation to obtain the feature map Feature4_BV;

[0120] Input the feature map Feature4_BV into the fourth MDFPB module for wavelet transform and enhancing the frequency-domain directional features to obtain the frequency-refined feature map Feature4_MD;

[0121] Input the feature map Feature4_MD into the fourth feature mapping module for the dimension expansion-compression strategy to obtain the non-linearly enhanced feature map Feature4_decorder;

[0122] Concatenate the feature map Feature4_decorder and the feature map Feature3_skip passed from the skip connection along the channel dimension to form a composite feature that fuses multi-level information. Then, compress the number of channels from 512 to 256 through a linear transformation layer (1×1 convolution) and input its feature map into the third BVPE module for cross-spatial self-attention calculation to obtain the feature map Feature3_BV;

[0123] Input the feature map Feature3_BV into the third MDFPB module for wavelet transform and enhancing the frequency-domain directional features to obtain the frequency-refined feature map Feature3_MD;

[0124] Input the feature map Feature3_MD into the third feature mapping module for the dimension expansion-compression strategy to obtain the non-linearly enhanced feature map Feature3_decorder;

[0125] The feature map Feature3_decorder and the feature map Feature2_skip passed from the skip connection are concatenated in the channel dimension to form a composite feature that fuses multi-level information. Subsequently, the number of channels is compressed from 256 to 128 through a linear transformation layer (1×1 convolution), and the feature map is input into the second BVPE module for cross-spatial self-attention calculation to obtain the feature map Feature2_BV;

[0126] The feature map Feature2_BV is input into the second MDFPB module for wavelet transform and enhancing the frequency-domain directional features to obtain the frequency-refined feature map Feature2_MD;

[0127] The feature map Feature2_MD is input into the second feature mapping module for the dimension expansion-compression strategy to obtain the non-linearly enhanced feature map Feature2_decorder;

[0128] The feature map Feature2_decorder and the feature map Feature1_skip passed from the skip connection are concatenated in the channel dimension to form a composite feature that fuses multi-level information. Subsequently, the number of channels is compressed from 128 to 64 through a linear transformation layer (1×1 convolution), and the feature map is input into the first BVPE module for cross-spatial self-attention calculation to obtain the feature map Feature1_BV;

[0129] The feature map Feature1_BV is input into the first MDFPB module for wavelet transform and enhancing the frequency-domain directional features to obtain the frequency-refined feature map Feature1_MD;

[0130] The feature map Feature1_MD is input into the first feature mapping for the dimension expansion-compression strategy to obtain the non-linearly enhanced feature map Feature1_decorder;

[0131] Finally, the feature map Feature1_decoder is input into the Segmentation Head module for processing to obtain the final classification output.

[0132] As Figure 3 shown, the structure of the BVPE module is specifically as follows:

[0133] First, after the input data is processed by the BVPE module, the input feature map Input is divided into two parts, namely Input_1 and Input_2, according to the spatial dimension through the split function. By splitting the input feature map into different parts, the model can independently focus on different regions during processing, enhancing the understanding of local features, which is particularly important for the subtle differences between citrus trees and the surrounding vegetation and fruits.

[0134] Linear transformations are performed on Input_1 and Input_2 respectively to generate the query vector Q1, the key vector K1, and two value vectors V1 and V2. The linear transformation maps the input feature map into different spaces to generate key components suitable for the self-attention mechanism, laying the foundation for subsequent attention calculations.

[0135] The dot product of Q1 and K1 is calculated, and the calculation result is normalized through the Sigmoid function to obtain the self-attention weight map S1. The dot product calculation helps to quantify the relationship between different spatial positions, and the Sigmoid function normalizes the result to ensure that the weights are within the range of [0,1], thus ensuring the stability and controllability of the attention mechanism.

[0136] Next, the weighted sum of the feature map V1 is calculated using the weight map S1 to obtain the feature map F1 with attention weights; similarly, the weight map S1 acts on the feature map V2, and the weighted sum is performed to obtain the second feature map F2 with attention weights. By weighting V1 and V2 with the calculated attention weights, the differences between the citrus tree area and the background vegetation are highlighted, unimportant parts are suppressed, and the representation ability of fruits and tree canopies is improved.

[0137] Next, F1 and F2 are subtracted to obtain the differential feature map D1. This subtraction operation helps the model to emphasize the subtle differences between citrus trees and the surrounding vegetation. Especially at the junctions of fruit trees and grasslands and shrubs, the model can improve the sensitivity to these small changes through the differential feature map, thus achieving more accurate segmentation.

[0138] Then, the differential feature map D1, as well as F1 and F2, are respectively subjected to channel screening (ChannelAttention) to generate two channel attention weight matrices W1 and W2. By performing channel screening on the feature map, the model can further extract important channel information related to citrus trees, dynamically adjust the feature representation of different channels, and enhance the focusing ability on citrus trees, fruits, and tree canopies.

[0139] Finally, multiply W1 by F1 to obtain the attention feature map A1, and multiply W2 by F2 to obtain the attention feature map A2. By applying the attention weights to the feature maps F1 and F2, the responses of each channel are further adjusted, enabling the model to better focus on the key regions of citrus trees and fruits.

[0140] Finally, concatenate (connect) the attention feature maps A1 and A2 in the spatial dimension to obtain the final BVPE output result.

[0141] Through the processing of the BVPE module, the model can more precisely focus on the subtle differences between citrus trees and their fruits and the background vegetation, thereby achieving pixel-level precise segmentation of citrus trees and the background, significantly improving the segmentation performance and accuracy. This process is particularly helpful for the precise identification of citrus trees in complex backgrounds in orchard scenes, providing reliable technical support for precision agriculture and orchard management.

[0142] In the present invention, the MDFPB module receives the feature map transmitted from the BVPE module and includes operations of wavelet transform processing and enhancement of frequency information direction perception.

[0143] As Figure 4 shown, the structure of the Haar wavelet transform of the MDFPB module is as follows according to the specific process:

[0144] First, the MDFPB module transmits the input feature map to the discrete wavelet transform (DWT) for Haar wavelet transform to obtain four feature maps with different frequencies, namely LL, LH, HL, and HH. Among them, LL represents the low-frequency part, containing the main information and general structure of the image; LH, HL, and HH respectively represent the high-frequency detail information in different directions, including the edge features in the horizontal, vertical, and diagonal directions.

[0145] The four frequency feature maps of LL (low frequency - low frequency), LH (low frequency - high frequency), HL (high frequency - low frequency), and HH (high frequency - high frequency) are concatenated (concat) along the channel dimension to obtain a new feature map Frequency_All.

[0146] Perform frequency information direction perception enhancement operations on the feature map Frequency_All, and this operation performs feature extraction and enhancement through three different branch structures respectively.

[0147] In the upper branch, convolution operations of central difference, horizontal difference, and vertical difference are sequentially performed on the input image. The central difference convolution captures the detailed changes in the image, the horizontal difference convolution extracts the horizontal edge information of the image, and the vertical difference convolution extracts the vertical edge information of the image. Through these difference convolution operations, the details in different directions of the image are fused with each other, and finally the results of all convolution operations are summed up to obtain a feature map M_F1 containing detailed features in each direction.

[0148] In the lower branch, strip-shaped convolution operations of 1×3 and 3×1 are used. This kind of operation can capture frequency information at different scales in the horizontal and vertical directions respectively. Through the 1×3 convolution, the horizontal texture features of the image are captured; while the 3×1 convolution focuses on feature extraction in the vertical direction. Then, the results of the two convolutions are integrated through a 1×1 convolution, and finally the feature map M_F3 is generated to further optimize the detailed performance of the image.

[0149] In the middle branch, a 3×3 convolution operation is performed on the input. This operation can effectively extract the global features of the image and capture the global context information in the image through a larger convolution kernel. After convolution, the feature map is screened through channels. Specifically, by using AvgPool and MaxPool to reduce the spatial dimension of the feature map while retaining key information, MaxPool highlights significant features, AvgPool retains global information, and the ReLu operation non-linearly activates the input, retaining positive values and setting negative values to zero. Finally, the most representative and important channel features are selected to obtain the feature map M_F2. Next, the Sigmoid function normalization operation is applied to the feature map M_F2 to map its values to between 0 and 1, generating two weight matrices M_W1 and M_W2. The Sigmoid normalization helps balance the importance of different feature maps, and the generated weight matrices can guide subsequent weighted operations.

[0150] Finally, M_W1 is multiplied by the feature map M_F1 of the upper branch to obtain the weighted feature map M_F4, and M_W2 is multiplied by the feature map M_F3 of the lower branch to obtain the weighted feature map M_F5. Through these weighted operations, the model can focus more on regions or features that make important contributions.

[0151] Finally, M_F4 and M_F5 are fused through a 1×1 convolution to obtain the final output result. This process effectively integrates the feature information extracted by each branch and provides a more discriminative feature representation for subsequent tasks.

[0152] The wavelet transform of the DWT (Discrete Wavelet Transform) is specifically as follows:

[0153] First, initial filtering is performed using a 2×1 strip-shaped convolutional kernel to extract low-frequency (smooth) and high-frequency (detail) information, obtaining two feature maps: High_frequency and Low_frequency. Specifically, the 2×1 convolutional kernel performs a convolution operation on the input signal in the horizontal direction, thereby decomposing the signal into low-frequency and high-frequency parts. Next, convolution operations with convolutional kernels of 2×1 (horizontal) and 1×2 (vertical) are respectively applied to the High_frequency and Low_frequency feature maps to further extract different-direction details of the signal in the two-dimensional space. After this series of processes, four frequency feature maps are finally obtained: LL (low-frequency - low-frequency), LH (low-frequency - high-frequency), HL (high-frequency - low-frequency), and HH (high-frequency - high-frequency).

[0154] In the present invention, the designs of the upper branch, middle branch, and lower branch particularly emphasize enhancing the directionality of frequency information, thereby improving the sensitivity of the model to features in different directions. The upper branch focuses on capturing the detail changes in the horizontal, vertical, and diagonal directions in the image through differential convolution, which helps to enhance the directional features of the image edges and local structures. The middle branch uses conventional convolution and combines channel screening to obtain more extensive context information, while further strengthening the features in the key directions and removing redundant information; the lower branch processes the frequency information of different scales through strip-shaped convolution to enhance the model's perception ability of image details in multiple scales and directions. The combination of the three can comprehensively improve the performance of the model in complex image structures and details, especially in capturing and strengthening the directional frequency information.

[0155] In the present invention, an RFM module is added to the skip connection to optimize the information from the encoder.

[0156] As Figure 5 shown, the structure of the RFM module is specifically as follows:

[0157] Receives the heterogeneous features transmitted by the encoder and fuses them through 1×1 convolution to obtain the feature map X1, aiming to integrate different-frequency information and reduce the number of channels, thereby improving the compactness of feature expression and the information fusion efficiency;

[0158] Next, the feature map X1 is input into a 3×3 convolutional layer to obtain the local feature map X1_3×3, and a weight matrix X1_W3 is generated through AVGPOOL and Softmax operations. This operation helps the model focus on the extraction of local information and improve the sensitivity to details. At the same time, spatial information screening is performed on the feature map X1. By performing AVGPOOL calculations on the X-axis and Y-axis respectively, the compressed spatial feature maps X_AVG and Y_AVG are obtained. In this way, the model can compress the spatial dimension and capture the macroscopic structural information in the image;

[0159] Then, the Sigmoid function is used to normalize X_AVG and Y_AVG to generate the weight maps R_W1 and R_W2, ensuring that the weights are between 0 and 1, enabling the model to adjust the influence of different spatial regions according to importance; Next, the compressed spatial feature maps X_AVG and Y_AVG are multiplied element-wise with the weight maps R_W1 and R_W2 to obtain the spatially screened feature maps R_A1 and R_A2;

[0160] The two feature maps R_A1 and R_A2 are concatenated along the channel dimension to obtain the final feature map X1_A1. Through this process, the model can more accurately screen and fuse information from different spatial dimensions, improving the expression ability of the feature map and the capture accuracy of key information;

[0161] The feature map X1_A just obtained is multiplied by the weight matrix X1_W3 to obtain the locally featured information X1_A2 after spatial screening. At the same time, the locally obtained local feature map X1_3×3 is multiplied by the weight map X1_A1_W of the feature map X1_A1 to obtain the weighted local feature map X1_A3;

[0162] Then, the locally featured information X1_A2 after spatial screening and the local feature map X1_A3 are added together to obtain the fused spatially aware feature and weighted local feature map X1_X2;

[0163] Finally, by normalizing the feature map X1_X2, a weight map X1_X2_W is generated and weighted calculation is performed with the feature map X1 to obtain a refined output.

[0164] Through the above process, the model can capture tiny details and local information in the image more accurately. First, by screening spatial and local features, the model can remove redundant information and focus on key regions in the image. Then, through weighted and normalization operations, the refinement process can effectively enhance important features while suppressing noise and irrelevant details, thus improving the expressiveness and accuracy of the feature map. Finally, the refined output can not only better retain the details and structure of the image, but also improve the sensitivity and recognition ability to complex image information, enhance the accuracy and robustness of the model in detail processing, and thus help the decoder to generate a clearer and more accurate output.

[0165] In addition, at each stage of the decoder, the present invention innovatively designs a multi-level composite loss supervision system, synchronously applying pixel-level smooth factor cross-entropy loss and contour-aware Dice loss at each feature scale during the decoding process, combining the main loss and the auxiliary loss, and interpolating and weighting the features of different layers. It can be expressed as:

[0166]

[0167] Case 1: The main output uses the main loss, and the auxiliary output uses the combined loss, for layers 2 to 4; Case 2: The number of output values of the model is 2, the main output uses the main loss, and the auxiliary output uses the auxiliary loss, which is used for the first layer when the conditions of Case 1 are not met; Condition 3: The model only returns the main output and directly uses the main loss;

[0168] This formula defines a multi-level loss function for training a deep neural network containing a main output f θ (x) and multiple auxiliary outputs H k (x). During the training process, the network calculates the loss according to different conditions. If the network output has two parts, the main loss L main and the auxiliary loss L aux are calculated respectively and weighted according to the hyperparameter λ to achieve more precise optimization. When there are multiple auxiliary outputs (features H k (x) from the intermediate layers of the network, where k ∈ {2, 3, 4}), the combined loss L joint of each auxiliary output will be calculated to enhance the learning of features at different scales. The design of this formula aims to promote the co-optimization of the network at different feature scales through a multi-level loss supervision mechanism, thereby improving the overall performance and generalization ability of the model.

[0169] L joint = αL SCE + βL Dice

[0170] This formula defines a combined loss function L joint , which combines two loss functions: the weighted cross-entropy loss L SCE and the Dice loss L Dice , for the image segmentation task in multi-task learning. Here, α and β are hyperparameters used to control the weights of these two loss functions. In this formula, L SCE is responsible for calculating the pixel-level classification loss, measuring the difference between the predicted class and the true label, while L SCE measures the accuracy of the model in image segmentation by calculating the Dice coefficient, especially in the case of dealing with imbalanced classes. The design purpose of the combined loss function is to combine these two losses, balance the classification accuracy and the segmentation accuracy, so as to improve the performance of the model in the segmentation task.

[0171]

[0172] This formula represents the weighted cross-entropy loss (Soft Cross-Entropy Loss, SCE), which is used to calculate the difference between the model prediction and the true label. In this loss function, H and W are the height and width of the image respectively, H×W represents the number of pixels in each image, the number of classes C represents the total number of possible classes in the prediction task, c represents the class, j represents the j-th class, i represents the i-th pixel, represents the probability that pixel i belongs to class c, y i,c represents the original target label of the i-th pixel, represents the smoothed target label of the i-th pixel, Z i,c represents the prediction value of the network for pixel i and class c, ∈ represents a label smoothing factor, with a value range between 0 and 1, and its role is to adjust the distribution of the true label to prevent the model from overfitting to the training data. This formula measures the classification performance of the model by calculating the logarithmic loss between the true label and the predicted probability.

[0173]

[0174] This formula represents the Dice loss (Dice Loss), which is commonly used to measure the similarity between the model prediction result and the true label in the image segmentation task. In the formula, H and W are the height and width of the image, representing the number of pixels in each image. p i is the prediction value of the model on pixel i, usually a probability value, while y i is the true label of this pixel. γ is a smoothing factor used to avoid the denominator being zero and ensure numerical stability. The Dice coefficient essentially evaluates the model performance by calculating the overlap degree between the prediction and the true label. The calculation method in the formula ensures that the model is particularly sensitive to the segmentation effect of small regions. The smaller the value of this loss function, the closer the model prediction is to the true label.

[0175] The precise segmentation results of citrus trees are applied to the intelligent management of orchards to monitor disease spots and the status of new buds in real time: for the detection of new young leaves with a proportion of less than 5%, through BVPE and MDFPN, the recall rate is increased to 92.5% (only 72.1% by traditional methods), and the false detection rate is controlled at 3.1%; for disease spots of 0.1 - 2 cm 2 Combined with RFM, an F1-score of 0.87 (precision 89.8%, recall 84.3%) is achieved, among which the detection accuracy of 0.1 cm 2 tiny spots is increased by 34% compared with traditional methods; in the orchards where this system is deployed in the experimental base, the response speed of disease recognition is increased by 40%, the accurate spraying coverage rate of pesticides is increased by 62% (the amount of pesticides used per unit area is reduced by 28%), and the monitoring error of the growth of new buds is controlled within ±1.2 mm / day.

[0176] The present invention constructs a citrus tree management method based on image semantic segmentation, and the specific applications include: feature processing, health status assessment, and growth cycle prediction;

[0177] Among them, BVPE, MDFPB, and RFM are specifically designed for the citrus tree recognition scenario. This part processes the three-channel image to generate a multi-level feature map with spatial-semantic relationships, so as to achieve the precise separation of citrus trees from the complex agricultural background. The algorithm is used to accurately capture the canopy features of different scales. Among them, the low-level features can finely identify the microscopic details of the leaves, the middle-level features can accurately capture the trend of the branches, and the high-level features can deeply analyze the overall tree shape structure. Whether the citrus tree is in the seedling stage or the adult stage, this module can perfectly adapt to its entire growth cycle. In terms of light robustness processing, the module will enhance the weight of the texture features in the shadow area to highlight the effective information in this area; at the same time, it will suppress the noise interference in the overexposed area to ensure stable and accurate feature extraction under different light conditions. In addition, the training data includes pixel-level semantic annotations, covering the RGB three channels and the Alpha channel, near-infrared band auxiliary annotations for distinguishing withered branches and leaves, and 3D point cloud spatial annotations to help verify the canopy volume features, and cover the different growth stage features in spring, autumn, and winter, enhancing the adaptability to seasonal changes. The data augmentation strategy includes 30% branch and leaf occlusion, 15% fruit occlusion, and 5% extreme weather samples to ensure the robustness of the algorithm in complex occlusion scenarios. The 128-dimensional channel features output by this module, including 32-dimensional citrus tree-specific feature channels, can be directly docked with the orchard robot navigation system to achieve a spatial positioning error of less than 5 cm, or combined with the yield prediction model to reach an accuracy of 85.10% for maturity recognition, thus providing comprehensive support for the intelligent agriculture solution.

[0178] The present invention innovatively designs a multi-level composite loss supervision system, synchronously applying pixel-level smoothed factor cross-entropy loss and contour-aware Dice loss at each feature scale during the decoding process, combining the main loss and the auxiliary loss, and interpolating and weighting the features of different layers to ensure accurate citrus tree recognition and health assessment in various agricultural scenarios. A multi-dimensional biometric fusion analysis method is adopted to accurately locate key regions such as the citrus tree canopy, branches, and fruits based on the image segmentation results, and a dynamic health evaluation system is constructed by combining multi-source data such as the normalized difference vegetation index, leaf area index, and three-dimensional canopy morphological features derived from near-infrared spectral reflectance. By analyzing 16 quantitative indicators such as abnormal offsets in the leaf chromaticity space, such as the attenuation of the normalized difference vegetation index value caused by a decrease in chlorophyll content, the damaged rate of the branch epidermis, and the canopy density distribution, and combining machine learning models to identify typical pest and disease texture features, such as Huanglong disease patches and citrus canker depressions. The system generates a structured health assessment report including growth vitality scores, disease risk levels, and nutrient deficiency warnings in real time through an adaptive weight allocation algorithm, providing data support for precision plant protection.

[0179] In addition, the growth cycle prediction module predicts key factors such as the growth trend, yield, and climate adaptability in the future for a period of time by analyzing the growth cycle characteristics of citrus trees and environmental factors. Finally, combining the comprehensive information output by these modules, the system provides scientific decision-making support for precision agricultural management, promoting the development of orchard management towards intelligence and high efficiency. Through the application of this series of innovative algorithms, the present invention provides comprehensive support for smart agriculture solutions, significantly improving the accuracy and efficiency of citrus tree recognition and management.

[0180] The present invention provides a citrus tree management method based on image semantic segmentation. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.

Claims

1. A citrus tree management method based on image semantic segmentation, characterized in that: The following steps are involved: Step 1: Take images of the citrus orchard; Step 2: Screen the original images taken, retain the features of the objects, and remove images that only contain a single category or have insufficient information to obtain a data set for training; Step 3: Accurately annotate the images in the training set, set labels for each feature category, and divide the data set; Step 4: Construct the encoder-decoder architecture of CSFAN cascade feature aggregation. In the encoding stage, the deep residual network ConvNext-tiny is used to implement multi-level feature extraction. The cross-layer feature interaction strategy is used to splice the feature maps of each level with the features of the adjacent layers of the scale in the channel dimension to obtain heterogeneous features. Step 5: The RFM module applies spatial constraints to the heterogeneous features obtained in step 4, filters and optimizes the information, and then passes it to the decoder for further processing; Step 6: The decoder uses the BVPE module to perform cross-spatial feature interaction on the features passed in step 5, and refines the frequency information through the MDFPB module. It uses a multi-level composite loss supervision system for decoding to achieve accurate segmentation of citrus tree images.

2. The citrus tree management method based on image semantic segmentation according to claim 1, characterized in that: In step 4, the ConvNext-tiny model acts as an encoder to extract four-stage features, expressed as: Among them, Input represents the input feature map, F1, F2, F3 and F4 represent the feature maps extracted in the four stages of the ConvNext-tiny model, and ConvNexBlock of each stage k Represents a block operation, k corresponds to the stage value of 1, 2, 3 or 4, starting from the second stage, the input of each stage is the output feature map of the previous stage, such as multi-level feature extraction to obtain features F1, F2, F3 and F4, and the shallow feature maps and deep feature maps of features F1, F2, F3 and F4 are spliced ​​to obtain heterogeneous features.

3. The citrus tree management method based on image semantic segmentation according to claim 2, characterized in that: In step 6, the BVPE module performs cross-spatial domain feature interaction as follows: Step 6-1-1: Divide the input feature map Input into two parts according to the spatial dimension through the split function, namely feature map Input_1 and feature map Input_2; Step 6-1-2: Perform linear transformation on the feature graph Input_1 and the feature graph Input_2 respectively, so as to generate the query vector Q1, the key vector K1, and the value vector V1 and the value vector V2; Step 6-1-3: Calculate the dot product of the query vector Q1 and the key vector K1, scale the dot product result to the query vector dimension d, and then calculate the Hadamard product with the scalar Λ1, and normalize the calculation result through the Sigmoid function to obtain the self-attention weight map S1; Λ1=exp(λ q ·l k ) λ is a learnable scaling factor used to adjust the similarity between the query and the key. The value range of λ is usually 0 to 1. Its subscripts q and k correspond to the position encoding parameters of the query vector and the key vector, respectively. Λ1 is a dynamic position weight matrix, which indicates the strength of the attention weight. Λ1 is calculated by λ q and λ k The dot product result of is obtained by applying the exponential function, and the position weight matrix Λ1 is multiplied by the Hadamard product and superimposed on the attention score to dynamically adjust the association strength of different positions; Step 6-1-4: Multiply the weight map S1 and the feature map V1 by dot product to obtain the feature map F1 with attention weight; the weight map S1 acts on the feature map V2, and the second feature map F2 with attention weight is obtained after weighted summation; the calculation formulas of feature maps F1 and F2 are as follows: F1=S1·V1 F2=S1·V2 Step 6-1-5: Subtract the feature map F1 from the feature map F2 to obtain the differential feature map D1. The formula is as follows; D1=F1-F2 Step 6-1-6: Perform channel screening on the differential feature map D1, feature map F1 and feature map F2, that is, screen the feature information in the channel to generate two channel attention weight matrices W1 and W2, as follows: E1=ChannelAttn(reshape(F1)) E2=ChannelAttn(reshape(F2)) E3=ChannelAttn(reshape(D1)) The feature graphs F1, F2 and their differential information D1 calculated in step 6-1-4 are enhanced in channel dimension, that is, the input dimension is changed by reshaping to match the input dimension requirement of the channel screening ChannelAttn, and the channel enhanced feature matrices E1, E2 and E3 are obtained by channel screening; W1=Softmax(E3(E1) T ) <h2 style=";text-align:left;direction:ltr">W2=Softmax(E3(E2)<h2 style=";text-align:left;direction:ltr"> T <h2 style=";text-align:left;direction:ltr"> ) The above channel enhancement feature matrices E1, E2 and E3 are transposed and inner product calculated, and the results of the inner product calculation are applied to the Softmax function to generate the attention weight matrix W1 and matrix W2; Step 6-1-7: Multiply the matrix W1 and the feature map F1 and perform residual connection to obtain the attention feature map A1, and multiply the matrix W2 and the feature map F2 and perform residual connection to obtain the attention feature map A2; A1=W1·F1+F1 A2=W2·F2+F2 By calculating the attention weight matrix and updating the weighted attention result, and then enhancing the feature representation through residual connection, we finally get the feature output feature map A1 and feature map A2 in different spatial domains; Step 6-1-8: Concatenate the attention feature maps A1 and A2 in the spatial dimension to obtain the output feature map OutPut that integrates information from different spatial domains, which is the final BVPE output result: In the formula, Indicates that the feature dimension is B×N×C, B: batch size, indicating the number of samples processed at one time, C: channel dimension, indicating the embedding vector dimension or hidden layer dimension of each element, N: sequence length, indicating the number of elements in a single sample.

4. The citrus tree management method based on image semantic segmentation according to claim 3, characterized in that: In step 6, the MDFPB module refines the frequency information as follows: Step 6-2-1: Pass the input feature map to discrete wavelet transform DWT for Haar wavelet transform, and obtain four feature maps of different frequencies, namely feature map LL, feature map LH, feature map HL and feature map HH; wherein feature map LL represents the overall structure of the low-frequency part; feature map LH, feature map HL and feature map HH represent the edge features in the horizontal, vertical and diagonal directions respectively; Step 6-2-2: The four frequency feature maps, feature map LL, feature map LH, feature map HL and feature map HH, are concatenated according to the channel dimension to obtain a new feature map Frequency_All; Step 6-2-3: Perform frequency information direction perception enhancement operation on the feature map Frequency_All. This operation extracts and enhances features through three different branch structures respectively. In the upper branch, the input image is subjected to convolution operations of center difference, horizontal difference and vertical difference in sequence. The center difference convolution captures the detail changes in the image, the horizontal difference convolution extracts the horizontal edge information of the image, and the vertical difference convolution extracts the vertical edge information of the image. Through these difference convolution operations, the details of the image in different directions are fused with each other. Finally, the results of all convolution operations are summed up to obtain a feature map M_F1 containing the detail features of each direction. The lower branch uses 1×3 and 3×1 strip convolution operations, which can capture frequency information of different scales in the horizontal and vertical directions respectively. Through 1×3 convolution, the horizontal texture features of the image are captured. The 3×1 convolution focuses on feature extraction in the vertical direction. Then, the two convolution results are integrated through 1×1 convolution to finally generate the feature map M_F3; The middle branch performs a 3×3 convolution operation on the input to capture the global context information in the image through a larger range of convolution kernels. The feature map after convolution is filtered through channels to select the most representative and important channel features to obtain the feature map M_F2; Step 6-2-4: Apply the Sigmoid function normalization operation to the feature map M_F2, map its value to between 0 and 1, and generate two weight matrices M_W1 and M_W2; Step 6-2-5: Multiply M_W1 by the feature map M_F1 of the upper branch to obtain the weighted feature map M_F4, and multiply M_W2 by the feature map M_F3 of the lower branch to obtain the weighted feature map M_F5; Step 6-2-6: M_F4 and M_F5 are fused through 1×1 convolution to obtain the final output result.

5. The citrus tree management method based on image semantic segmentation according to claim 4, characterized in that: In step 6-2-1, the discrete wavelet transform DWT is specifically as follows: Step 6-2-1-1: Perform preliminary filtering through a 2×1 strip convolution kernel to extract low-frequency and high-frequency information and obtain two feature maps: feature map High_frequency and feature map Low_frequency; Step 6-2-1-2: Apply horizontal convolution operation with convolution kernel of 2×1 and vertical convolution operation with convolution kernel of 1×2 to feature map High_frequency and feature map Low_frequency respectively, and further extract details of the signal in different directions in two-dimensional space; obtain four frequency feature maps representing low frequency-low frequency, low frequency-high frequency, high frequency-low frequency and high frequency-high frequency: feature map LL, feature map LH, feature map HL and feature map HH.

6. The citrus tree management method based on image semantic segmentation according to claim 2, characterized in that: In step 5, the RFM module filters the high-frequency and low-frequency spatial channel information as follows: Step 5-1: Receive the heterogeneous features transmitted by the encoder and fuse them through 1×1 convolution to obtain the feature map X1; Step 5-2: Input the feature map X1 into the 3×3 convolution layer to obtain the local feature map X1_3×3, and generate the weight matrix X1_W3 through AVGPOOL and Softmax operations; filter the spatial information of the feature map X1, and obtain the compressed spatial feature maps X_AVG and Y_AVG by performing average-pooling operations on the X-axis and Y-axis respectively; Step 5-3: Use the Sigmoid function to normalize X_AVG and Y_AVG to generate weight graphs R_W1 and R_W2; Step 5-4: Multiply the compressed spatial feature maps X_AVG and Y_AVG by element with the weight maps R_W1 and R_W2 to obtain the spatial screening feature maps R_A1 and R_A2; Step 5-5: Concatenate the feature map R_A1 and the feature map R_A2 along the channel dimension to obtain the final feature map X1_A1; Step 5-6: Multiply the feature map X1_A and the weight matrix X1_W3 to obtain the local feature information X1_A2 after spatial screening; multiply the local feature map X1_3×3 obtained in step 5-2 by the weight map X1_A1_W of the feature map X1_A1 to obtain the weighted local feature map X1_A3; Step 5-7: Add the spatially filtered local feature information X1_A2 and the local feature map X1_A3 to obtain the fused spatial perception feature and the weighted local feature map X1_X2; Step 5-8: Normalize the feature map X1_X2 to generate the weight map X1_X2_W, and perform weighted calculation with the feature map X1 to obtain a refined output.

7. The citrus tree management method based on image semantic segmentation according to claim 1, characterized in that: In step 6, the multi-level composite loss supervision system is specifically as follows: pixel-level smoothing factor cross entropy loss and contour-aware Dice loss are simultaneously applied at each feature scale in the decoding process, combining the main loss and auxiliary loss, and interpolating and weighting the features of different layers, which is expressed as follows: Case 1: The main output uses the main loss, the auxiliary output uses the joint loss, and is used for layers 2 to 4; Case 2: The number of output values ​​of the model is 2, the main output uses the main loss, and the auxiliary output uses the auxiliary loss. When the condition in case 1 is not met, it is used for the first layer; Condition 3: The model only returns the main output, and the main loss is used directly; Among them, x represents the feature map obtained from each layer; y represents the true value corresponding to x; g ф represents the auxiliary output of the corresponding level, k represents the corresponding level, L total Represents the multi-level loss function, f θ (x) indicates the main output, H k (x) represents the auxiliary output of the corresponding level, L main represents the main loss, L aux represents auxiliary loss, L joint represents the joint loss, and the hyperparameter λ represents the weight of the corresponding loss, ranging from 0 to 1; The combined loss L joint The calculation is as follows: L joint =αL SCE +βL Dice L joint Combines two loss functions: weighted cross entropy loss L sce and Dice loss L Dice , α and β are hyperparameters that control the weights of the two loss functions, and their values ​​range from 0 to 1.

8. The citrus tree management method based on image semantic segmentation according to claim 7, characterized in that: The weighted cross entropy loss L sce Calculate pixel-level classification loss to measure the difference between the predicted category and the true label. Weighted cross entropy loss L sce The classification performance of the model is measured by calculating the logarithmic loss between the true label and the predicted probability. The formula is as follows: Where H and W are the height and width of the image, respectively. H×W represents the number of pixels in each image. The number of categories C represents the total number of possible categories in the prediction task. c represents the category. j represents the jth category. i represents the pixel at the i-th spatial position of the network. y represents the probability that pixel i belongs to category c. i,c represents the original target label of the i-th pixel, represents the target label of the i-th pixel after smoothing, z i,c represents the original prediction value of the network that the pixel at the ith spatial position belongs to the cth category, z i,j It represents the original prediction value of the network for the pixel at the ith spatial position corresponding to other categories j. ∈ represents a label smoothing factor, which ranges from 0 to 1. Its function is to adjust the distribution of the true label to prevent the model from overfitting the training data.

9. The citrus tree management method based on image semantic segmentation according to claim 7, characterized in that: The Dice loss L Dice The accuracy of the model in image segmentation is measured by calculating the Dice coefficient. The formula is as follows: In the formula, i represents the pixel corresponding to the i-th spatial position of the network, p i represents the predicted value of the model at the pixel i, which is a probability value, and y i represents the true label of the pixel; H and W represent the height and width of the image, and H×W represents the number of pixels in each image; γ is a smoothing factor to avoid the denominator being zero, and its value range is between 0.1 and 10.

10. The citrus tree management method based on image semantic segmentation according to claim 1, characterized in that: In step 3, the data set is divided into a training set, a validation set and a test set.