A mango tree extraction method coupling multi-scale convolution and dual-branch network
By using a method of coupling multi-scale convolution and dual-branch network in the extraction of single-wood canopy of mango trees, the problems of low extraction efficiency, low accuracy and poor edge recognition in the prior art are solved, and efficient and accurate recognition effect in complex scenarios is achieved.
Patent Information
- Application Number
- CN202411234280.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-09-04
AI Technical Summary
The prior art has problems such as the object detection method in the extraction of single-wood canopy of mango trees, the slow inference speed of example segmentation method, and the rough canopy edge extraction of ordinary semantic segmentation method, resulting in low extraction efficiency, low accuracy, poor edge recognition, and difficult to apply in environments with high precision, high closure and high scenario complexity.
Using a method of coupling multi-scale convolution and dual-branch network, surface drawing is performed through drone images, feature dimension improvement and dimensionality reduction is performed using exponential calculation and principal component analysis, and a shared module and semantic segmentation and edge detection branches are constructed. Combined with the dual-branch feature fusion module and regularized loss function, the model is trained to achieve accurate extraction of the canopy of mango single plant.
It realizes efficient and accurate identification of mango tree single-wood canopy in complex scenarios, improves extraction efficiency and accuracy, and optimizes edge details recognition, which is suitable for environments with high precision, high closure and high scenario complexity.
Smart Images

Figure CN119091304B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of single tree image segmentation, and particularly relates to a mango single tree extraction method coupling multi-scale convolution and double-branch network. Background Art
[0002] Mango is an important tropical cash crop in the world. my country is the original production place and one of the main producing countries of mango. In 2023, in view of mango, the pillar and characteristic industry of the local rural economy, the Hainan government issued the "Sanya Mango White Paper and Sanya Mango Standard" to promote the modernization of the mango industry. Foreign countries have integrated "sky-ground integration" to carry out personalized and precise management of fruit trees, while the management model of domestic mango orchards relies heavily on manual labor. How to efficiently, accurately and automatically extract mango crowns is an important basis for monitoring fruit tree growth changes, flowering rate, loss rate, and yield. It is also an important prerequisite for agricultural insurance production estimation and claims. At present, in the field of remote sensing crown extraction, there are technical bottlenecks such as field background confusion and high crown overlap. New methods with accurate classification and efficient identification are urgently needed to promote the development of application industries.
[0003] In recent years, the performance of semantic segmentation, target detection, instance segmentation, etc. in tree crown extraction tasks has surpassed traditional computer vision algorithms. Luo et al. improved Faster R-CNN to detect single trees in coal mine afforestation areas, but the crown width was not accurate. Instance segmentation was improved from this model: Huang Xinxi et al. used it to extract the single tree area of ginkgo trees in the city, which was relatively accurate, but the scattered distribution of the crown could not verify its applicability in densely planted orchards; the instance segmentation reasoning speed of the two-stage architecture was also inferior to the semantic segmentation of the single-stage architecture. The rise of precision agriculture has promoted the development of semantic segmentation in single tree canopy extraction. For CNN architecture, Freudenberg et al. used a combination of deep and shallow U-Net to achieve higher accuracy. Song Haoxin used PSPNet, Deeplabv3+, U-Net, U 2 Net realizes the comparison of single tree extraction of citrus trees. SegNeXt is based on CNN and is inspired by the ViT architecture and combines multiscale convolution attention (MSCA), which has a high accuracy improvement. At present, only related research has been carried out in farmland classification in China. Semantic segmentation and edge detection have a mutually reinforcing effect: Xuandong et al., Fenglei Chen et al. all applied edge detection to improve the accuracy of farmland and building segmentation, but there is no relevant method to discuss the application of dual-branch network and multi-scale convolution ideas to single tree canopy segmentation.
[0004] Therefore, most of the current single tree crown extraction models have the following problems: the target detection method does not have accurate canopy edges, the instance segmentation method has slow inference speed, the ordinary semantic segmentation method has rough canopy edge extraction, and most of the existing semantic segmentation studies have sparse canopy distribution and low canopy density. The accuracy and applicability of the developed single tree segmentation algorithm for the extraction of single tree canopies of mango trees with mixed periods and complex scenarios needs to be explored, resulting in low efficiency, low accuracy and poor edge recognition of single tree extraction, which makes it difficult to be widely used in the extraction of single tree canopies of mango economic forests with high accuracy, high canopy density and high scenario complexity. Summary of the invention
[0005] The problem to be solved by the present invention is to overcome complex scenes, improve inversion accuracy, optimize edge details, take into account training efficiency and other problems in single mango tree canopy recognition, and propose a single mango tree extraction method that couples multi-scale convolution and dual-branch network.
[0006] To achieve the above object, the present invention is implemented through the following technical solutions:
[0007] A method for extracting a single mango tree by coupling multi-scale convolution with a dual-branch network comprises the following steps:
[0008] S1. Select UAV images of mango trees with complex scenes, draw the surface of the mango tree canopy, and make labels for the mango tree canopy in complex scenes;
[0009] S2. For the drone image selected in step S1, apply index calculation to perform feature dimension upgrading to obtain a multi-feature image;
[0010] S3. Apply principal component analysis to the multi-feature image obtained in step S2 to perform data dimensionality reduction, and then combine it with the label drawn in step S1 to form an image-label pair, cut it into a size of 512×512 to obtain input data, and construct the input data set of the model;
[0011] S4. Construct a shared module Share_Stem, and apply the lightweight trunk first module and the lightweight trunk second module to extract the shallow features and sub-shallow features of the input data of step S3;
[0012] S5. Construct a semantic segmentation branch, apply the dilation enhanced multi-scale backbone network feature extraction model D_E_MCAN, extract multi-scale semantic features from the sub-shallow features obtained in step S4, and combine the loss function to obtain the predicted semantic segmentation result;
[0013] S6. construct an edge detection branch, alienate and splice the shallow features obtained in step S4 and the multi-scale semantic features obtained in step S5 into edge features, and combine the binary cross entropy loss function to obtain a predicted edge mask;
[0014] S7. Construct an edge detection-semantic segmentation dual-branch model D_E_BSNet, apply the dual-branch feature fusion module AFD to fuse the multi-scale semantic features obtained in step S5 and the edge features obtained in step S6, and combine the constructed overall loss function to obtain the final fused feature prediction mask;
[0015] S8. The input data set obtained in step S3 is input into the edge detection-semantic segmentation dual-branch model constructed in step S7 for training to obtain the optimal mango single plant canopy recognition model;
[0016] S9. Using the optimal mango single plant canopy recognition model obtained in step S8, extract the mango canopy in the target area.
[0017] Furthermore, the specific implementation method of step S1 includes the following steps:
[0018] S1.1. Set up drone imagery of mango forests in complex scenarios, including background weeds, overlapping canopies, undulating terrain, pruning period, and mixed crops;
[0019] The periods of complex scenarios include pre-flowering period, bud period, flowering period, small fruit period, and swelling period;
[0020] S1.2. For the drone images collected in step S1.1, the canopy of a single mango plant is mapped and the precision agriculture module of ENVI5.6.3 is used for preliminary identification to obtain a canopy vector in *.SHP format as the canopy label of a single mango plant in a complex scenario.
[0021] Furthermore, the specific implementation method of step S2 includes the following steps:
[0022] S2.1. Set index features for the drone images covered by the target dataset. The calculation formula is as follows, where R, G, and B are the grayscale values of the red, green, and blue bands of the corresponding images, respectively; r, g, and b are the normalized grayscale values of the red, green, and blue bands of the corresponding images, respectively; r = R / (R+G+B), g = G / (R+G+G), and b = G / (R+G+B):
[0023] S2.1.1.Extra green index ExG, the formula is as follows:
[0024] ExG = 2g-rb;
[0025] S2.1.2. The super red index ExR is as follows:
[0026] ExR = 1.4rg;
[0027] S2.1.3.Extra Blue Index ExB, the formula is as follows:
[0028] ExB = 1.4bg;
[0029] S2.1.4. Plant color extraction index CIVE, the formula is as follows:
[0030] CIVE=0.441r-0.881g+0.385b+18.78745;
[0031] S2.1.5. Green-blue vegetation index GBVI, the formula is as follows:
[0032]
[0033] S2.1.6. Vegetation factor index VEG, the formula is as follows:
[0034]
[0035] S2.1.7. The super green and super red differential index ExGR is as follows:
[0036] ExGR=EXG-ExR;
[0037] S2.1.8. Green leaf index GLI, the formula is as follows:
[0038]
[0039] S2.1.9. Normalized green-red difference index NGRDI, the formula is as follows:
[0040]
[0041] S2.1.10. Red-Green Ratio Index (RGRI), the formula is as follows:
[0042]
[0043] S2.2. Use ENVI5.6.3 to calculate the index data obtained in step S2.1, and integrate the calculated index data with the channel data of the drone image into a new *.TIF format raster to obtain a multi-feature image.
[0044] Furthermore, the specific implementation method of step S3 includes the following steps:
[0045] S3.1. First, apply ENVI5.6.3 to perform principal component analysis and dimensionality reduction on the image of the sample data set obtained in step S2 to generate a 4-dimensional principal component;
[0046] S3.2. Convert the label samples in the *.SHP format obtained in step S1 into the *.TIF format of the same length and width as the image after dimensionality reduction in step S3.1, and form an image-label pair with the 4-dimensional principal components of step S3.1. Then cut them into 512×512 sizes in sequence with a row and column overlap rate of 10% as input data to construct the input data set I of the model.
[0047] Furthermore, in step S4, the method for constructing the shared module Share_Stem is to apply two improved lightweight backbone MobileNetV3 initial modules to be connected successively, the shallow feature F1 extracted by the lightweight backbone first module M1 will be input into the edge detection branch, and the sub-shallow feature F2 extracted by the lightweight backbone second module M2 will be input into the semantic segmentation branch;
[0048] The shared module Share_Stem obtained is two 3×3 convolutional layers with a step size of 2 and a padding of 1 + a batch normalization layer + a HSigmoid activation layer, and the expression is as follows:
[0049] F1=σ(BatchNorm(Conv 3×3 (I)))
[0050] F2 = σ(BatchNorm(Conv 3×3 (F1)))
[0051] Among them, BatchNorm is batch normalization and σ is the HSigmoid activation layer.
[0052] Furthermore, the specific implementation method of step S5 includes the following steps:
[0053] S5.1. Construct an enhanced multi-scale convolutional attention mechanism E_MCA attention module. The mathematical expression of the E_MCA attention module is:
[0054]
[0055] Among them, x is the input feature; conv 1×1 is 1×1 convolution; DWConv 5×5 is the initial 5×5 depth convolution of x; DWConv is the depth convolution; i∈{0,1,2,3,4,5}, Scale0 is the identity connection, and the others are the i-th branches. The branches are the depth convolutions with kernel sizes of 3, 7, 11, 17, and 21, respectively. The depth strip convolution is used instead of the large kernel convolution, and E_MAC att For E_MCA attention, is an element-by-element matrix multiplication operation;
[0056] S5.2. Apply the E_MCA attention module of step S5.1 to construct the spatial attention mechanism E-MCA Spatial attention module, which is expressed as:
[0057]
[0058] Among them, GELU is the activation function;
[0059] S5.3. Apply the spatial attention mechanism E-MCA Spatial attention module constructed in step S5.2 to construct the basic building block E_MCABlock, and apply two layers of residual connections, the expression is:
[0060]
[0061] Among them, BatchNorm is batch normalization, E_MCA spa_att The spatial attention mechanism E-MCA Spatial attention module constructed in step S5.2, layer_scale is the layer scaling factor, and the initial value is 1e -2 , drop_path is layer regularization;
[0062] The second residual connection connects E_MCA spa_att Replaced with MLP multi-layer perception module, the expression is:
[0063]
[0064] Among them, y' out The input feature of this step is the output feature of the previous step; MLP is a multi-layer perception module, layer_scale is the layer scaling factor, and the initial value is 1e -2 , drop_path is layer regularization;
[0065] S5.4. Construct the backbone network D_E_MCAN, including a hierarchical structure, which includes four stages, each level contains a downsampling module and a stack of building modules and batch normalization. In the first stage, the initial downsampling module is the initial Stem part, and the Share_Stem constructed in step S4 is used to extract the secondary shallow feature F2; in the second to fourth stages, the downsampling module refers to the stem part of the ViT architecture and adopts the overlapping patch OverlapPatchEmbed module. In the second to fourth stages, 2, 4, and 2 E_MCABlock building blocks are connected to the downsampling module respectively;
[0066] S5.5. Construct the semantic segmentation branch and apply the D_E_MCAN constructed in step S5.4. as the model backbone network. The feature F extracted by the semantic segmentation branch sThe formula is as follows;
[0067] F s =S(F2)
[0068] Among them, S is the semantic segmentation branch;
[0069] And apply a 1×1 2d convolution to make a simple classification head to predict the semantic segmentation result S m ;
[0070] S5.6. Construct the Dice loss function L used by the semantic segmentation branch Dice , the expression is:
[0071]
[0072] L Dice =1-Dice
[0073] Among them, Dice is a set similarity measurement function.
[0074] The specific implementation method of the further step S6 comprises the following steps:
[0075] S6.1. Construct an edge detection branch, and send the shallow feature F1 described in step S4.1 and the features extracted from the four stages of the hierarchical structure in the semantic segmentation branch backbone network described in step S5 to a 3×3 convolutional layer, a group normalization layer, and a GELU layer respectively to transform the semantic features into edge features;
[0076] Bilinear interpolation is used to upsample the multi-scale boundary features and connect them together, and a 1×1 2D convolution is applied to make a simple classification head to predict the edge detection results; the feature F extracted by the edge detection branch b The expression is:
[0077] F b =Concat(Upsample([GELU(Conv 3×3 ([F1,F S_multi-scale (F2)]))]))
[0078] Among them, F S_multi-scale Multi-scale features extracted for semantic segmentation network, Conv 3×3 It is 3×3 convolution, Upsample is bilinear interpolation upsampling, and Concat is concatenation;
[0079] S6.2. Set the binary cross entropy loss function L described in step S6 bce for:
[0080]
[0081] Where: N is N samples, y i is the true label of the i-th sample, which takes the value 0 or 1. is the probability of the i-th sample predicted by the model, with a value of [0,1];
[0082] A simple 1×1 convolutional classification head is applied to predict the predicted edge mask Bm as described in step S6.
[0083] Furthermore, the specific implementation method of step S7 includes the following steps:
[0084] S7.1. Constructing a dual-branch feature fusion module The active fusion decoder connects the semantic segmentation branch constructed in step S5 and the edge detection branch constructed in step S6 to fuse the semantic features F extracted in step S5 s The edge feature F extracted in step S6 b , generate fusion features to achieve the final prediction mask F m ; For the dual-branch feature fusion module, the dual-branch features perform the following steps in this module:
[0085] The dual-branch features are first globally averaged pooled to fuse the semantic features with the edge features. After passing through a multi-layer perception module, semantic attention and edge attention are obtained. The two are then spliced at the channel level to obtain fused attention. The expression is:
[0086]
[0087] Based on the multi-head attention mechanism, we divide the data into H groups and perform linear projection calculation to obtain the query vector q, key vector k and value vector v. We calculate the association matrix A. The value A is calculated in the i-th row and j-th column of the A matrix. i,j As follows, apply v and A i,j Multiply them together to get the fusion weight matrix w i , the expression is:
[0088]
[0089] w i =Av i
[0090] Then, through the inverse operation of channel-level splicing, it is divided into the semantic branch weight vector w s_att and the edge branch weight vector w b_att , apply residual connection and add them together to get the fusion feature F f After fusing the features, a 1×1 convolution is connected as the classification head to obtain the final semantic segmentation prediction map F m , the expression is:
[0091] F f =(1+w s_att )F s +(1+w b_att )F b
[0092] S7.2. Set the regularization loss function of step S7 to apply a bidirectional consistency loss function, including a boundary-to-semantic consistency loss function and a semantic-to-boundary consistency loss function;
[0093] S7.3. Set the calculation steps of the semantic-to-boundary consistency loss function as:
[0094] Set the sobel edge detection operator C with a kernel size of 3×3 and apply it to the fusion semantic prediction graph F. m Sliding convolution operations are performed on the X-axis and Y-axis, so that pixels on one side of the edge are given higher positive weights and pixels on the other side are given lower negative weights, while the weight of the center pixel is 0, in order to detect the boundary where the image intensity changes sharply. After calculating the gradients in the X and Y directions, the sum of the absolute values of the two gradient components is calculated as the pseudo-semantic boundary prediction. The average absolute loss is used to supervise the pseudo-semantic boundary, and the expression is:
[0095] b ps =max(||C⊙F m ||)
[0096]
[0097] Among them, max is the maximum value, ⊙ is the sliding operation, |||| is the absolute value, b ps The pseudo edge labels generated by the operator sliding, is the true semantic boundary label;
[0098] S7.4. In order to maintain the semantic consistency between the subject and the boundary, a boundary-to-semantic consistency loss function is formulated, and the calculation steps are as follows;
[0099] Calculate the boundary to semantic consistency loss function as follows:
[0100]
[0101] is the semantic true label, where c and p are the traversal categories and pixels, Mark the ground truth pixels and high confidence pixels on the boundary prediction map b, ∈ is the confidence threshold, selected as 0.8;
[0102] S7.5. Regularization function is the boundary-to-semantic consistency loss function With semantic to boundary loss function The sum of is expressed as:
[0103]
[0104] S7.6. The enhanced edge detection-semantic segmentation dual-branch dilated convolutional network D_E_BSNet is now assembled, and the overall loss function formula is:
[0105]
[0106]
[0107] Shape Symbols with superscripts are the corresponding true value labels, λ1 and λ2 are hyperparameters that control the classification loss and dual-task regularization weights, both set to 1;
[0108] S7.7. Apply a simple 1×1 classification head to make the overall final prediction mask F m .
[0109] The method for extracting a single mango tree by coupling a multi-scale convolution with a dual-branch network according to claim 8 is characterized in that step S8 is to input the input data set I obtained in step S3 into the combined model constructed in step S7 for model training, and the average loss curve of the training set, the average intersection-over-union curve of the validation set, the overall recognition accuracy of the test set, the average intersection-over-union ratio, the precision, the recall rate and the F1 score are used as criteria for evaluating the quality of the model to evaluate the quality of the model.
[0110] Beneficial effects of the present invention:
[0111] The invention discloses a method for extracting a mango tree by coupling multi-scale convolution with a dual-branch network. In order to enable the model to have the generalization ability of identifying a single tree in a complex orchard scenario, a data set of a mango tree in a complex scenario is prepared. In order to enrich the feature set of a mango tree canopy and optimize the features, a data dimension reduction method based on PCA is designed, the dimension reduction of RGB-index composite data of a drone is realized, and important features are selectively retained. An E_MCA attention mechanism is designed, the residual connection is an E_MCA spatial attention mechanism, and the D_E_MCAN backbone network is strategically stacked to extract attention vectors of multi-scale features respectively, so that context information of a mango tree identification object can be effectively extracted. An enhanced semantic-edge dual-branch dilated convolutional network structure D_E_BSNet is constructed to realize the organic fusion of semantic and edge features. The regularization loss function is set as a bidirectional consistency loss function to alleviate the conflict of the dual branches during back propagation. Finally, the optimal model is trained, saved, and applied to the prediction of the mango tree canopy, so that the accurate extraction of the mango tree canopy can be realized with less training cost.
[0112] The method for extracting individual mango trees by coupling multi-scale convolution with a dual-branch network described in the present invention can efficiently and accurately identify the canopy of individual mango trees from UAV remote sensing images, broadens the technical means for the application field of individual tree extraction, lays a technical foundation and scientific reference for the engineering application of remote sensing identification product services in the economic forest field in the future, and can help the intensive counting and extraction of more economic crops. BRIEF DESCRIPTION OF THE DRAWINGS
[0113] Figure 1 This is a flow chart of a method for extracting a single mango tree by coupling multi-scale convolution with a dual-branch network according to the present invention;
[0114] Figure 2 This is an example of a complex scene mango sample set described in the present invention, where (a) is the pre-flowering stage, (b) is the bud stage, (c) is the flowering stage, (d) is the small fruit stage, (e) is the swelling stage, (f) is the background with weeds, (g) is the canopy overlap, (h) is the undulating terrain, (i) is the pruning stage, and (j) is the mixed cropping stage;
[0115] Figure 3 This is an architecture diagram of the sharing module Share-Stem described in the present invention;
[0116] Figure 4 This is a diagram of the enhanced multi-scale convolutional attention architecture described in the present invention;
[0117] Figure 5 This is a diagram of the enhanced multi-scale spatial convolutional attention architecture described in the present invention;
[0118] Figure 6 This is a diagram of the enhanced multi-scale spatial convolution building block architecture of the present invention;
[0119] Figure 7 This is an architecture diagram of the expansion-enhanced multi-scale backbone network described in the present invention;
[0120] Figure 8 This is a schematic diagram of the double-branch semantic-edge model of a single mango tree canopy according to the present invention;
[0121] Fig. 9 is a structural diagram of the active fusion decoder of the present invention;
[0122] Fig.10 The figures are the accuracy evaluation diagrams of the optimal mango single plant canopy extraction model described in the present invention, wherein (a) is the average loss curve of the training set, (b) is the average intersection over union (IoU) curve of the validation set, and (c) is the qualitative analysis comparison diagram of the model test;
[0123] Fig.11 This is the optimal mango tree canopy identification result diagram described in the present invention. DETAILED DESCRIPTION
[0124] In order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.
[0125] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected specific embodiments of the present invention. Based on the specific embodiments of the present invention, all other specific embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0126] In order to further understand the content, features and effects of the present invention, the following specific implementation methods are given as examples, and the attached Figure 1 -Attached Fig.11 The detailed instructions are as follows:
[0127] Embodiment 1:
[0128] A method for extracting a single mango tree by coupling multi-scale convolution with a dual-branch network comprises the following steps:
[0129] S1. Select UAV images of mango trees with complex scenes, draw the surface of the mango tree canopy, and make labels for the mango tree canopy in complex scenes;
[0130] Furthermore, the specific implementation method of step S1 includes the following steps:
[0131] S1.1. Set up drone imagery of mango forests in complex scenarios, including background weeds, overlapping canopies, undulating terrain, pruning period, and mixed crops;
[0132] The periods of complex scenarios include pre-flowering period, bud period, flowering period, small fruit period, and swelling period;
[0133] S1.2. Surface mapping of the mango canopy is performed on the drone image collected in step S1.1, and preliminary identification is performed using the precision agriculture module of ENVI5.6.3 to obtain a canopy vector in *.SHP format as a mango canopy label in a complex scenario;
[0134] S2. For the drone image selected in step S1, apply index calculation to perform feature dimension upgrading to obtain a multi-feature image;
[0135] Furthermore, the specific implementation method of step S2 includes the following steps:
[0136] S2.1. Set index features for the drone images covered by the target dataset. The calculation formula is as follows, where R, G, and B are the grayscale values of the red, green, and blue bands of the corresponding images, respectively; r, g, and b are the normalized grayscale values of the red, green, and blue bands of the corresponding images, respectively; r = R / (R+G+B), g = G / (R+G+B), and b = B / (R+G+B):
[0137] S2.1.1.Extra green index ExG, the formula is as follows:
[0138] ExG = 2g-rb;
[0139] S2.1.2. The super red index ExR is as follows:
[0140] ExR = 1.4rg;
[0141] S2.1.3.Extra Blue Index ExB, the formula is as follows:
[0142] ExB = 1.4bg;
[0143] S2.1.4. Plant color extraction index CIVE, the formula is as follows:
[0144] CIVE=0.441r-0.881g+0.385b+18.78745;
[0145] S2.1.5. Green-blue vegetation index GBVI, the formula is as follows:
[0146]
[0147] S2.1.6. Vegetation factor index VEG, the formula is as follows:
[0148]
[0149] S2.1.7. The super green and super red differential index ExGR is as follows:
[0150] ExGR=EXG-ExR;
[0151] S2.1.8. Green leaf index GLI, the formula is as follows:
[0152]
[0153] S2.1.9. Normalized green-red difference index NGRDI, the formula is as follows:
[0154]
[0155] S2.1.10. Red-Green Ratio Index (RGRI), the formula is as follows:
[0156]
[0157] S2.2. Use ENVI5.6.3 to calculate the index data obtained in step S2.1, and integrate the calculated index data with the channel data of the drone image into a new *.TIF format raster to obtain a multi-feature image;
[0158] S3. Apply principal component analysis to the multi-feature image obtained in step S2 to perform data dimensionality reduction, and then combine it with the label drawn in step S1 to form an image-label pair, cut it into a size of 512×512 to obtain input data, and construct the input data set of the model;
[0159] Furthermore, the specific implementation method of step S3 includes the following steps:
[0160] S3.1. First, apply ENVI5.6.3 to perform principal component analysis and dimensionality reduction on the image of the sample data set obtained in step S2 to generate a 4-dimensional principal component;
[0161] S3.2. Convert the label samples in the *.SHP format obtained in step S1 into the *.TIF format of the same length and width as the image after dimensionality reduction in step S3.1, and form an image-label pair with the 4-dimensional principal components of step S3.1. Then cut them into 512×512 sizes in sequence with a row and column overlap rate of 10% as input data to construct the input data set I of the model.
[0162] S4. Construct a shared module Share_Stem, and apply the lightweight trunk first module and the lightweight trunk second module to extract the shallow features and sub-shallow features of the input data of step S3;
[0163] Furthermore, in step S4, the method for constructing the shared module Share_Stem is to apply two improved lightweight backbone MobileNetV3 initial modules to be connected successively, the shallow feature F1 extracted by the lightweight backbone first module M1 will be input into the edge detection branch, and the sub-shallow feature F2 extracted by the lightweight backbone second module M2 will be input into the semantic segmentation branch;
[0164] The shared module Share_Stem obtained is two 3×3 convolutional layers with a step size of 2 and a padding of 1 + a batch normalization layer + a HSigmoid activation layer, and the expression is as follows:
[0165] F1=σ(BatchNorm(Conv 3×3 (I)))
[0166] F2 = σ(BatchNorm(Conv 3×3 (F1)))
[0167] Among them, BatchNorm is batch normalization, σ is the HSigmoid activation layer;
[0168] S5. Construct a semantic segmentation branch, apply the dilation enhanced multi-scale backbone network feature extraction model D_E_MCAN, extract multi-scale semantic features from the sub-shallow features obtained in step S4, and combine the loss function to obtain the predicted semantic segmentation result;
[0169] Furthermore, the specific implementation method of step S5 includes the following steps:
[0170] S5.1. Construct an enhanced multi-scale convolutional attention mechanism E_MCA attention module. The mathematical expression of the E_MCA attention module is:
[0171]
[0172] Among them, x is the input feature; Conv 1×1 is 1×1 convolution; DWConv 5×5 is the initial 5×5 depth convolution of x; DWConv is the depth convolution; i∈{0,1,2,3,4,5}, Scale0 is the identity connection, and the others are the i-th branches. The branches are the depth convolutions with kernel sizes of 3, 7, 11, 17, and 21, respectively. The depth strip convolution is used instead of the large kernel convolution. E_MCA att For E_MCA attention, is an element-by-element matrix multiplication operation;
[0173] Furthermore, the main idea is to decompose the large k×k convolution kernel into k×1 and 1×k convolutions, and use the multi-scale features extracted by the decomposed convolution kernel for segmentation. A total of 5 pairs of strip convolutions are used to extract contextual features of more scales to meet the extraction of mango trees with different canopy sizes.
[0174] Furthermore, Scale1 represents replacing the 3×3 convolution with a 1×3 and 3×1 convolution kernel.
[0175] S5.2. Apply the E_MCA attention module of step S5.1 to construct the spatial attention mechanism E-MCA Spatial attention module, which is expressed as:
[0176]
[0177] Among them, GELU is the activation function;
[0178] S5.3. Apply the spatial attention mechanism E-MCA Spatial attention module constructed in step S5.2 to construct the basic building block E_MCABlock, and apply two layers of residual connections, the expression is:
[0179]
[0180] Among them, BatchNorm is batch normalization, E_MCA spa_att The spatial attention mechanism E-MCA Spatial attention module constructed in step S5.2, layer_scale is the layer scaling factor, and the initial value is 1e -2 , drop_path is layer regularization;
[0181] The second residual connection connects E_MCA spa_att Replaced with MLP multi-layer perception module, the expression is:
[0182]
[0183] Among them, y' out The input feature of this step is the output feature of the previous step; MLP is a multi-layer perception module, layer_scale is the layer scaling factor, and the initial value is 1e -2 , drop_path is layer regularization;
[0184] S5.4. Construct the backbone network D_E_MCAN, including a hierarchical structure, which includes four stages, each level contains a downsampling module and a stack of building modules and batch normalization. In the first stage, the initial downsampling module is the initial Stem part, and the Share_Stem constructed in step S4 is used to extract the secondary shallow feature F2; in the second to fourth stages, the downsampling module refers to the stem part of the ViT architecture and adopts the overlapping patch OverlapPatchEmbed module. In the second to fourth stages, 2, 4, and 2 E_MCABlock building blocks are connected to the downsampling module respectively;
[0185] Referring to the backbone network MSCAN of the SegNeXt network, the stacking of E_MCABlock building blocks is applied to obtain the dilated enhanced multiscale convolution attention backbone (D_E_MCAN);
[0186] Furthermore, the downsampling module is different from ViT in that it uses an overlapping patch OverlapPatchEmbed module to obtain richer local and global context features, and applies a 3×3 convolution with a dilation rate of 3 and a normalized stack. This is used to improve small object detection, enhance robustness to deformation and occlusion, and help the model network understand complex scenes and object boundaries at the pixel level. The last three stages are followed by 2, 4, and 2 E_MCABlock blocks after the downsampling module, respectively.
[0187] S5.5. Construct the semantic segmentation branch and apply the D_E_MCAN constructed in step S5.4 as the model backbone network. The feature F extracted by the semantic segmentation branch s The formula is as follows;
[0188] F s =S(F2)
[0189] Among them, S is the semantic segmentation branch;
[0190] And apply a 1×1 2d convolution to make a simple classification head to predict the semantic segmentation result S m ;
[0191] S5.6. Construct the Dice loss function L used by the semantic segmentation branch Dice , the expression is:
[0192]
[0193] L Dice =1-Dice
[0194] Among them, Dice is a set similarity measurement function.
[0195] S6. construct an edge detection branch, alienate and splice the shallow features obtained in step S4 and the multi-scale semantic features obtained in step S5 into edge features, and combine the binary cross entropy loss function to obtain a predicted edge mask;
[0196] Furthermore, the specific implementation method of step S6 includes the following steps:
[0197] S6.1. Construct an edge detection branch, and send the shallow feature F1 described in step S4.1 and the features extracted from the four stages of the hierarchical structure in the semantic segmentation branch backbone network described in step S5 to a 3×3 convolutional layer, a group normalization layer, and a GELU layer respectively to transform the semantic features into edge features;
[0198] Bilinear interpolation is used to upsample the multi-scale boundary features and connect them together, and a 1×1 2D convolution is applied to make a simple classification head to predict the edge detection results; the feature F extracted by the edge detection branchb The expression is:
[0199] F b =Concat(Upsample([GELU(Conv 3×3 ([F1,F S_multi-scale (F2)]))]))
[0200] Among them, F S_multi-scale Multi-scale features extracted for semantic segmentation network, Conv 3×3 It is 3×3 convolution, Upsample is bilinear interpolation upsampling, and Concat is concatenation;
[0201] S6.2. Set the binary cross entropy loss function L described in step S6 bce for:
[0202]
[0203] Where: N is N samples, y i is the true label of the i-th sample, which takes the value 0 or 1. is the probability of the i-th sample predicted by the model, with a value of [0,1];
[0204] A simple 1×1 convolutional classification head is applied to predict the predicted edge mask Bm as described in step S6.
[0205] S7. Construct an edge detection-semantic segmentation dual-branch model D_E_BSNet, apply the dual-branch feature fusion module AFD to fuse the multi-scale semantic features obtained in step S5 and the edge features obtained in step S6, and combine the constructed overall loss function to obtain the final fused feature prediction mask;
[0206] Furthermore, the specific implementation method of step S7 includes the following steps:
[0207] S7.1. Constructing a dual-branch feature fusion module The active fusion decoder connects the semantic segmentation branch constructed in step S5 and the edge detection branch constructed in step S6 to fuse the semantic features F extracted in step S5 s The edge feature F extracted in step S6 b , generate fusion features to achieve the final prediction mask F m ; For the dual-branch feature fusion module, the dual-branch features perform the following steps in this module:
[0208] The dual-branch features are first globally averaged pooled to fuse the semantic features with the edge features. After passing through a multi-layer perception module, semantic attention and edge attention are obtained. The two are then spliced at the channel level to obtain fused attention. The expression is:
[0209]
[0210] Based on the multi-head attention mechanism, we divide the data into H groups and perform linear projection calculation to obtain the query vector q, key vector k and value vector v. We calculate the association matrix A. The value A is calculated in the i-th row and j-th column of the A matrix. i,j As follows, apply v and A i,j Multiply them together to get the fusion weight matrix w i , the expression is:
[0211]
[0212] w i =Av i
[0213] Then, through the inverse operation of channel-level splicing, it is divided into the semantic branch weight vector w s_att and the edge branch weight vector w b_att , apply residual connection and add them together to get the fusion feature F f After fusing the features, a 1×1 convolution is connected as the classification head to obtain the final semantic segmentation prediction map F m , the expression is:
[0214] F f =(1+w s_att )F s +(1+w b_att )F b
[0215] S7.2. Set the regularization loss function of step S7 to apply a bidirectional consistency loss function, including a boundary-to-semantic consistency loss function and a semantic-to-boundary consistency loss function;
[0216] S7.3. Set the calculation steps of the semantic-to-boundary consistency loss function as:
[0217] Set the sobel edge detection operator C with a kernel size of 3×3 and apply it to the fusion semantic prediction graph F. m Sliding convolution operations are performed on the X-axis and Y-axis, so that pixels on one side of the edge are given higher positive weights and pixels on the other side are given lower negative weights, while the weight of the center pixel is 0, in order to detect the boundary where the image intensity changes sharply. After calculating the gradients in the X and Y directions, the sum of the absolute values of the two gradient components is calculated as the pseudo-semantic boundary prediction. The average absolute loss is used to supervise the pseudo-semantic boundary, and the expression is:
[0218] b ps =max(||C⊙F m ||)
[0219]
[0220] Among them, max is the maximum value, ⊙ is the sliding operation, |||| is the absolute value, b ps The pseudo edge labels generated by the operator sliding, is the true semantic boundary label;
[0221] S7.4. In order to maintain the semantic consistency between the subject and the boundary, a boundary-to-semantic consistency loss function is formulated, and the calculation steps are as follows;
[0222] Calculate the boundary to semantic consistency loss function as follows:
[0223]
[0224] is the semantic true label, where c and p are the traversal categories and pixels, Mark the ground truth pixels and high confidence pixels on the boundary prediction map b, ∈ is the confidence threshold, selected as 0.8;
[0225] S7.5. Regularization function is the boundary-to-semantic consistency loss function With semantic to boundary loss function The sum of is expressed as:
[0226]
[0227] S7.6. The enhanced edge detection-semantic segmentation dual-branch dilated convolutional network D_E_BSNet is now assembled, and the overall loss function formula is:
[0228]
[0229]
[0230] Shape Symbols with superscripts are the corresponding true value labels, λ1 and λ2 are hyperparameters that control the classification loss and dual-task regularization weights, both set to 1;
[0231] S7.7. Apply a simple 1×1 classification head to make the overall final prediction mask F m .
[0232] S8. The input data set obtained in step S3 is input into the edge detection-semantic segmentation dual-branch model constructed in step S7 for training to obtain the optimal mango single plant canopy recognition model;
[0233] Furthermore, step S8 is to input the input data set I obtained in step S3 into the combined model constructed in step S7 for model training, and use the average loss curve of the training set, the average intersection-over-union curve of the validation set, the overall recognition accuracy of the test set, the average intersection-over-union ratio, precision, recall rate and F1 score as the criteria for evaluating the quality of the model to evaluate the quality of the model.
[0234] After training, the best performing model is saved after comprehensive evaluation. The evaluation indicators and formulas are as follows:
[0235] S8.1. The formula for average intersection-over-union ratio is:
[0236]
[0237] S8.2. Formula for overall recognition accuracy:
[0238]
[0239] S8.3. Formula for precision:
[0240]
[0241] S8.4. Recall formula:
[0242]
[0243] S8.5. The formula for the F1 score is:
[0244]
[0245] Among them, k is the total number of pixel categories; true positive (TP) means the number of pixels predicted to be true and actually true; false positive (FP) means the number of pixels predicted to be true and actually false; false negative (FN) means the number of pixels predicted to be false and actually true; true negative (TN) means the number of pixels predicted to be false and actually false.
[0246] S9. Using the optimal mango single plant canopy recognition model obtained in step S8, extract the mango canopy in the target area.
[0247] The software and hardware parameters of the model training are shown in Table 1. The loss change curve of the training set and the mIoU curve of the validation set are shown in the attached Fig.10 (a), 10(b):
[0248] Table 1 Model software and hardware parameters
[0249]
[0250] In order to facilitate understanding of the technical effect of this embodiment, the comparative effect of the method of this embodiment and the prior art is shown in Table 2. Table 2 is a quantitative accuracy evaluation of the research model of the patent of this embodiment.
[0251] Table 2 Technical effect accuracy evaluation table
[0252]
[0253] As can be seen from Table 2, the improved model, SegNeXt and U-Net proposed in this embodiment have very high overall accuracy for pixels. In particular, the improved model proposed in this embodiment leads in all indicators and shows the best overall performance. In addition, the model volume is quantitatively measured and compared, and the results are shown in Table 3:
[0254] Table 3 Comparison of different model parameters
[0255]
[0256] From Table 3, we can get the parameter comparison results, and we can intuitively draw the conclusion that Deeplabv3+ is an improved model based on Deeplabv3. The accuracy improvement is accompanied by a doubling of the training cost. However, the constructed model in this paper, which is improved on the basis of SegNeXt, achieves a balance between accuracy improvement and model efficiency. While maintaining lightweight, it still achieves the best recognition accuracy of 99.2%.
[0257] Secondly, the optimal mango canopy extraction model of each model algorithm was used for qualitative analysis and comparison test to observe whether the mango canopy extraction results of each model had over-segmentation, under-segmentation, missed segmentation, and wrong segmentation. The comparison diagram is attached. Fig.10 (c). From the attached Fig.10 (c) It can be intuitively found that the improved model has significantly better performance. In terms of recognition details, the classification is more detailed, the edges are rounder and smoother, closer to the actual edges, and the above-mentioned poor segmentation phenomenon is less likely to occur.
[0258] Finally, the optimal model developed in this embodiment with good balanced performance in all aspects was used to identify some examples of randomly selected Hainan mango orchards. The final recognition results are shown in the attached Fig.11 .
[0259] Compared with other mango tree crown segmentation methods, this embodiment has the following advantages:
[0260] (1) Compared with traditional field survey and manual measurement, it saves time and effort, improves efficiency and saves costs;
[0261] (2) The optimal feature set with dimension increase and then dimension reduction is used for training to improve the extraction accuracy of mango canopy. Combined with the index feature, accurate segmentation results can be obtained;
[0262] (3) Through the dual-branch network, combined with the active fusion decoder, the edge features and semantic features are strategically fused to enhance the accuracy of edge recognition;
[0263] (4) Through multi-scale convolutional attention, the fusion of deep and shallow multi-scale features is improved, overcoming the problem of canopy recognition of different canopy sizes under complex backgrounds;
[0264] (5) The bidirectional consistency loss function is used to mitigate the performance degradation caused by the different back-propagation directions of semantic and edge features.
[0265] In summary, in the method for extracting a mango tree canopy described in this embodiment, a mango tree data set under complex scenarios is prepared to enable the model to have the generalization ability to identify individual trees in complex orchard scenarios; in order to enrich the feature set of the mango tree canopy and optimize the features, a data dimension reduction method based on PCA is designed to achieve the dimension reduction of the drone RGB-index composite data, and selectively retain important features; an E_MCA attention mechanism is designed, the residual connection is the E_MCA spatial attention mechanism, and the D_E_MCAN backbone network is strategically stacked to extract the attention vectors of multi-scale features respectively, which can effectively extract the context information of the mango tree identification object; an enhanced semantic-edge dual-branch dilated convolutional network structure D_E_BSNet is constructed, and an active fusion decoder is applied to realize the organic fusion of semantic and edge features; the regularization loss function is set to a bidirectional consistency loss function to alleviate the conflict of the two branches during back propagation; finally, the model is trained, and the optimal model is saved and applied to the mango tree canopy prediction, so that the accurate extraction of the mango tree canopy can be achieved with less training cost. The method of the invention can efficiently and accurately identify the canopy of individual mango trees from UAV remote sensing images, broadens the technical means for the application of single tree extraction, lays a technical foundation and scientific reference for the engineering application of remote sensing identification product services in the economic forest field in the future, and can help the intensive counting and extraction of more economic crops.
[0266] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0267] Although the present application has been described above with reference to specific embodiments, various modifications may be made thereto and parts thereof may be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application may be used in combination with each other in any manner, and the fact that these combinations are not exhaustively described in this specification is only for the sake of omitting space and saving resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A method for extracting single mango trees by coupling multi-scale convolution with a dual-branch network, characterized in that: The steps include: S1. Select UAV images of mango trees with complex scenes, draw the surface of the mango tree canopy, and make labels for the mango tree canopy in complex scenes; S2. For the drone image selected in step S1, apply index calculation to perform feature dimension upgrading to obtain a multi-feature image; S3. Apply principal component analysis to the multi-feature image obtained in step S2 to perform data dimensionality reduction, and then combine it with the label drawn in step S1 to form an image-label pair, cut it into a size of 512×512 to obtain input data, and construct the input data set of the model; S4. Construct a shared module Share_Stem, and apply the lightweight trunk first module and the lightweight trunk second module to extract the shallow features and sub-shallow features of the input data of step S3; Step S4 constructs a shared module Share_Stem by applying two improved lightweight backbone MobileNetV3 initial modules to connect in sequence, the shallow feature F1 extracted by the first lightweight backbone module M1 is input into the edge detection branch, and the sub-shallow feature F2 extracted by the second lightweight backbone module M2 is input into the semantic segmentation branch; The shared module Share_Stem obtained is two 3×3 convolutional layers with a step size of 2 and a padding of 1 + a batch normalization layer + a HSigmoid activation layer, and the expression is as follows: F1=σ(BatchNorm(Conv 3×3 (I))) F2=σ(BatchNorm(Conv 3×3 (F1))) Among them, BatchNorm is batch normalization, σ is the HSigmoid activation layer; S5. Construct a semantic segmentation branch, apply the dilation enhanced multi-scale backbone network feature extraction model D_E_MCAN, extract multi-scale semantic features from the sub-shallow features obtained in step S4, and combine the loss function to obtain the predicted semantic segmentation result; The specific implementation method of step S5 includes the following steps: S5.
1. Construct an enhanced multi-scale convolutional attention mechanism E_MCA attention module. The mathematical expression of the E_MCA attention module is: Among them, x is the input feature; Conv 1×1 is 1×1 convolution; DWConv 5×5 is the initial 5×5 depth convolution of x; DWConv is the depth convolution; i∈{0,1,2,3,4,5}, Scale0 is the identity connection, and the others are the i-th branches. The branches are the depth convolutions with kernel sizes of 3, 7, 11, 17, and 21, respectively. The depth strip convolution is used instead of the large kernel convolution. E_MCA att For E_MCA attention, is an element-by-element matrix multiplication operation; S5.
2. Apply the E_MCA attention module of step S5.1 to construct the spatial attention mechanism E-MCA Spatial attention module, which is expressed as: Among them, GELU is the activation function; S5.
3. Apply the spatial attention mechanism E-MCA Spatial attention module constructed in step S5.2 to construct the basic building block E_MCABlock, and apply two layers of residual connections, the expression is: Among them, BatchNorm is batch normalization, E_MCA spa_att The spatial attention mechanism E-MCASpatial attention module constructed in step S5.2, layer_scale is the layer scaling factor, and the initial value is 1e -2 , drop_path is layer regularization; The second residual connection connects E_MCA spa_att Replaced with MLP multi-layer perception module, the expression is: Among them, y' out is the output feature of the previous step and the input feature of this step; MLP is a multi-layer perception module, layer_scale is the layer scaling factor, and the initial value is 1e -2 , drop_path is layer regularization; S5.
4. Construct the backbone network D_E_MCAN, including a hierarchical structure, which includes four stages, each level contains a downsampling module and a stack of building modules and batch normalization. In the first stage, the initial downsampling module is the initial Stem part, and the Share_Stem constructed in step S4 is used to extract the secondary shallow feature F2; in the second to fourth stages, the downsampling module refers to the stem part of the ViT architecture and adopts the overlapping patch OverlapPatchEmbed module. In the second to fourth stages, 2, 4, and 2 E_MCABlock building blocks are connected to the downsampling module respectively; S5.
5. Construct the semantic segmentation branch and apply the D_E_MCAN constructed in step S5.4 as the model backbone network. The feature F extracted by the semantic segmentation branch s The formula is as follows; F s =S(F2) Among them, S is the semantic segmentation branch; And apply a 1×1 2d convolution to make a simple classification head to predict the semantic segmentation result S m ; S5.
6. Construct the Dice loss function L used by the semantic segmentation branch Dice , the expression is: THE Dice =1-Dice Among them, Dice is the set similarity measurement function; S6. construct an edge detection branch, alienate and splice the shallow features obtained in step S4 and the multi-scale semantic features obtained in step S5 into edge features, and combine the binary cross entropy loss function to obtain a predicted edge mask; The specific implementation method of step S6 includes the following steps: S6.
1. Construct an edge detection branch, and send the shallow feature F1 described in step S4.1 and the features extracted from the four stages of the hierarchical structure in the semantic segmentation branch backbone network described in step S5 to a 3×3 convolutional layer, a group normalization layer, and a GELU layer respectively to transform the semantic features into edge features; Bilinear interpolation is used to upsample the multi-scale boundary features and connect them together, and a 1×1 2D convolution is applied to make a simple classification head to predict the edge detection results; the feature F extracted by the edge detection branch b The expression is: Among them, F S_multi-scale Multi-scale features extracted for semantic segmentation network, Conv 3×3 It is 3×3 convolution, Upsample is bilinear interpolation upsampling, and Concat is concatenation; S6.
2. Set the binary cross entropy loss function L described in step S6 bce for: Where: N is N samples, y i is the true label of the i-th sample, which takes the value 0 or 1. is the probability of the i-th sample predicted by the model, with a value of [0,1]; Apply a simple 1×1 convolutional classification head to predict the predicted edge mask Bm as described in step S6; S7. Construct an edge detection-semantic segmentation dual-branch model D_E_BSNet, apply the dual-branch feature fusion module AFD to fuse the multi-scale semantic features obtained in step S5 and the edge features obtained in step S6, and combine the constructed overall loss function to obtain the final fused feature prediction mask; The specific implementation method of step S7 includes the following steps: S7.
1. Constructing a dual-branch feature fusion module The active fusion decoder connects the semantic segmentation branch constructed in step S5 and the edge detection branch constructed in step S6 to fuse the semantic features F extracted in step S5 s The edge feature F extracted in step S6 b , generate fusion features to achieve the final prediction mask F m ; For the dual-branch feature fusion module, the dual-branch features perform the following steps in this module: The dual-branch features are first globally averaged pooled to fuse the semantic features with the edge features. After passing through a multi-layer perception module, semantic attention and edge attention are obtained. The two are then spliced at the channel level to obtain fused attention. The expression is: Based on the multi-head attention mechanism, we divide the data into H groups and perform linear projection calculation to obtain the query vector q, key vector k and value vector v. We calculate the association matrix A. The value A is calculated in the i-th row and j-th column of the A matrix. i,j As follows, apply v and A i,j Multiply them together to get the fusion weight matrix w i , the expression is: w i =Off i Then, through the inverse operation of channel-level splicing, it is divided into the semantic branch weight vector w s_att and the edge branch weight vector w b_att , apply residual connection and add them together to get the fusion feature F f After fusing the features, a 1×1 convolution is connected as the classification head to obtain the final semantic segmentation prediction map F m , the expression is: F f =(1+w s_att )F s +(1+w b_att )F b S7.
2. Set the regularization loss function of step S7 to apply a bidirectional consistency loss function, including a boundary-to-semantic consistency loss function and a semantic-to-boundary consistency loss function; S7.
3. Set the calculation steps of the semantic-to-boundary consistency loss function as: Set the sobel edge detection operator C with a kernel size of 3×3 and apply it to the fusion semantic prediction graph F. m Sliding convolution operations are performed on the X-axis and Y-axis, so that pixels on one side of the edge are given higher positive weights and pixels on the other side are given lower negative weights, while the weight of the center pixel is 0, in order to detect the boundary where the image intensity changes sharply. After calculating the gradients in the X and Y directions, the sum of the absolute values of the two gradient components is calculated as the pseudo-semantic boundary prediction. The average absolute loss is used to supervise the pseudo-semantic boundary, and the expression is: b ps =max(||C⊙F m ||) Among them, max is the maximum value, ⊙ is the sliding operation, || || is the absolute value, b ps The pseudo edge labels generated by the operator sliding, is the true semantic boundary label; S7.
4. In order to maintain the semantic consistency between the subject and the boundary, a boundary-to-semantic consistency loss function is formulated, and the calculation steps are as follows; Calculate the boundary to semantic consistency loss function as follows: is the semantic true label, where c and p are the traversal categories and pixels, Mark the ground truth pixels and high confidence pixels on the boundary prediction map b, ∈ is the confidence threshold, selected as 0.8; S7.
5. Regularization function is the boundary-to-semantic consistency loss function With semantic to boundary loss function The sum of is expressed as: S7.
6. The enhanced edge detection-semantic segmentation dual-branch dilated convolutional network D_E_BSNet is now assembled, and the overall loss function formula is: Shape Symbols with superscripts are the corresponding true value labels, λ1 and λ2 are hyperparameters that control the classification loss and dual-task regularization weights, both set to 1; S7.
7. Apply a simple 1×1 classification head to make the overall final prediction mask F m ; S8. The input data set obtained in step S3 is input into the edge detection-semantic segmentation dual-branch model constructed in step S7 for training to obtain the optimal mango single plant canopy recognition model; S9. Using the optimal mango single plant canopy recognition model obtained in step S8, extract the mango canopy in the target area.
2. The method for extracting a single mango tree by coupling multi-scale convolution with a dual-branch network according to claim 1, characterized in that: The specific implementation method of step S1 includes the following steps: S1.
1. Set up drone imagery of mango forests in complex scenarios, including background weeds, overlapping canopies, undulating terrain, pruning period, and mixed crops; The periods of complex scenarios include pre-flowering period, bud period, flowering period, small fruit period, and swelling period; S1.
2. For the drone images collected in step S1.1, the canopy of a single mango plant is mapped and the precision agriculture module of ENVI5.6.3 is used for preliminary identification to obtain a canopy vector in *.SHP format as the canopy label of a single mango plant in a complex scenario.
3. The method for extracting a single mango tree by coupling multi-scale convolution with a dual-branch network according to claim 2, characterized in that: The specific implementation method of step S2 includes the following steps: S2.
1. Set index features for the drone images covered by the target dataset. The calculation formula is as follows, where R, G, and B are the grayscale values of the red, green, and blue bands of the corresponding images, respectively; r, g, and b are the normalized grayscale values of the red, green, and blue bands of the corresponding images, respectively; r = R / (R+G+B), g = G / (R+G+B), and b = B / (R+G+B): S2.1.1.Extra green index ExG, the formula is as follows: ExG = 2g-rb; S2.1.
2. The super red index ExR is as follows: ExR = 1.4rg; S2.1.3.Extra Blue Index ExB, the formula is as follows: ExB = 1.4bg; S2.1.
4. Plant color extraction index CIVE, the formula is as follows: CIVE=0.441r-0.881g+0.385b+18.78745; S2.1.
5. Green-blue vegetation index GBVI, the formula is as follows: S2.1.
6. Vegetation factor index VEG, the formula is as follows: S2.1.
7. The super green and super red differential index ExGR is as follows: ExGR=EXG-ExR; S2.1.
8. Green leaf index GLI, the formula is as follows: S2.1.
9. Normalized green-red difference index NGRDI, the formula is as follows: S2.1.
10. Red-Green Ratio Index (RGRI), the formula is as follows: S2.
2. Use ENVI5.6.3 to calculate the index data obtained in step S2.1, and integrate the calculated index data with the channel data of the drone image into a new *.TIF format raster to obtain a multi-feature image.
4. The method for extracting a single mango tree by coupling multi-scale convolution with a dual-branch network according to claim 3, characterized in that: The specific implementation method of step S3 includes the following steps: S3.
1. First, apply ENVI5.6.3 to perform principal component analysis and dimensionality reduction on the image of the sample data set obtained in step S2 to generate a 4-dimensional principal component; S3.
2. Convert the label samples in the *.SHP format obtained in step S1 into the *.TIF format of the same length and width as the image after dimensionality reduction in step S3.1, and form an image-label pair with the 4-dimensional principal components of step S3.
1. Then cut them into 512×512 sizes in sequence with a row and column overlap rate of 10% as input data to construct the input data set I of the model.
5. The method for extracting a single mango tree by coupling multi-scale convolution with a dual-branch network according to claim 4, characterized in that: Step S8 is to input the input data set I obtained in step S3 into the combined model constructed in step S7 for model training, and use the average loss curve of the training set, the average intersection-over-union curve of the validation set, the overall recognition accuracy of the test set, the average intersection-over-union ratio, precision, recall rate and F1 score as the criteria for evaluating the quality of the model to evaluate the quality of the model.
Citation Information
Patent Citations
Edge perception image semantic segmentation method based on adaptive feature fusion
CN113658200A
Remote sensing image semantic segmentation method based on shared convolution kernel and boundary loss function
CN115035295A
Image semantic segmentation model and segmentation method
CN116468740A
Cited By
Sparse planting type fruit tree single plant growth period extraction method, electronic equipment and storage medium
CN120913111A
A method for extracting a single plant growth period of a sparse-planting fruit tree, an electronic device, and a storage medium
CN120913111B