Tree height and biomass collaborative inversion method, system and device based on codec double focusing and medium
Through the multi-task learning model based on codec, the problem of synergistic accuracy of tree height and biomass inversion is solved, efficient synergistic inversion of tree height and biomass is achieved, and the inversion accuracy and model generalization ability are improved.
Patent Information
- Application Number
- CN202511015383.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-07-23
AI Technical Summary
The existing technology is difficult to effectively coordinate the inversion of tree height and biomass. There are technical bottlenecks in the fusion of optical remote sensing and synthetic aperture radar data. Traditional models fail to effectively capture the physical correlation between data, resulting in limited improvement in inversion accuracy.
A multi-task learning model based on codec dual focus is adopted to achieve collaborative inversion of tree height and biomass through residual fusion encoding network and dual-task feature interaction network. The model includes an encoder focus module and a decoder focus module. It uses a cross-branch interaction module and a pyramid pooling module to improve feature sharing and task interaction and enhance inversion accuracy.
The synergistic accuracy of tree height and biomass inversion is improved, the efficiency of multi-source data fusion is enhanced, the cross-modal feature synergy effect on optical and synthetic aperture radar data is realized, and the generalization ability and accuracy of the inversion model are improved.
Smart Images

Figure CN120522692A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing data application, and in particular relates to a tree height and biomass collaborative inversion method, system, equipment and medium based on codec dual focusing. Background Art
[0002] Determining forest structural parameters is fundamental to ecological monitoring and climate research. Tree height, a key parameter of forest vertical structure, is directly linked to core applications such as estimating aboveground biomass, assessing timber stock, and monitoring forest degradation. Furthermore, as a direct proxy for carbon storage, accurately deriving biomass is crucial for quantifying forest carbon sequestration capacity.
[0003] However, the collaborative inversion of tree height and biomass parameters currently faces the following technical bottlenecks: At the data level, a single data source cannot meet the objectives. Optical remote sensing offers high spectral resolution, but is limited by electromagnetic wave penetration and can only obtain two-dimensional spectral information from the canopy surface. It cannot directly invert forest vertical structure, such as tree height, and is susceptible to interference from weather factors such as clouds and aerosols. Synthetic aperture radar (SAR) allows for all-day, all-weather observations, but its backscattered signal exhibits a complex nonlinear relationship with forest parameters and is significantly affected by factors such as topographic relief and multiple scattering from vegetation layers. Therefore, when used alone to invert biomass, accuracy fluctuates significantly. Although spaceborne lidar (LIDAR) can penetrate the canopy to obtain true tree height values through laser pulses, the data exhibits a discrete spot distribution, making it impossible to directly generate continuous, planar grids of tree height and biomass. Therefore, it urgently needs to be combined with other remote sensing data to achieve scale expansion. However, data fusion still faces technical bottlenecks. Optical, SAR, and LIDAR data differ significantly in spatial resolution, observation mechanism, and data morphology. Traditional fusion methods struggle to effectively capture the physical correlations between data, resulting in limited improvements in inversion accuracy.
[0004] At the model level, single-task models lack synergy. Traditional tree height and biomass inversion often use independent modeling methods: tree height inversion relies on lidar point clouds and optical vegetation indices, while biomass inversion focuses on synthetic aperture radar backscatter coefficients and optical spectral characteristics. However, these two methods sever the inherent relationship between tree height and biomass, leading to potential contradictions in the inversion results based on physical logic. The architecture of multi-task models is limited. Although existing multi-task learning models can achieve feature reuse through shared encoders, the decoders lack explicit interaction between tasks and have difficulty handling the strong correlation between tree height and biomass. Existing models do not design targeted cross-task interaction mechanisms, making it impossible to achieve information complementarity between tree height and biomass. For example, biomass inversion cannot utilize tree height's ability to represent the vertical structure of vegetation, resulting in reduced accuracy in complex forest environments. Summary of the Invention
[0005] Purpose of the invention: The first purpose of the present invention is to provide a tree height and biomass collaborative inversion method based on codec dual focusing, which improves the collaborative inversion accuracy of strongly correlated targets.
[0006] The second object of the present invention is to provide a tree height and biomass collaborative inversion system based on codec dual focusing.
[0007] A third object of the present invention is to provide an electronic device.
[0008] A fourth object of the present invention is to provide a computer storage medium.
[0009] Technical solution: To achieve the above objectives, the present invention provides a tree height and biomass collaborative inversion method based on codec dual focusing, comprising the following steps: (1) Use spaceborne lidar to obtain tree height and biomass data in the target area, and obtain optical and synthetic aperture radar data at the corresponding locations; (2) Preprocess the acquired data and construct a training data set; (3) Establish a multi-task learning model based on encoder and decoder dual focus, train the training data set through multi-task learning, and generate a collaborative inversion model of tree height and biomass; the multi-task learning model includes an encoder focusing module and a decoder focusing module, wherein the encoder focusing module inputs the input data of the training data set into the residual fusion encoding network to generate tree height branch features and biomass branch features respectively, and the decoder focusing module inputs the tree height branch features and biomass branch features into the dual-task feature interaction network to output the tree height inversion results and biomass inversion results; The residual fusion coding network includes two parallel branch structures: a tree height encoder and a biomass encoder. Each branch passes through five residual modules in sequence. The convolution kernel size of the residual module decreases from large to small, and the number of channels doubles from module to module. After processing each residual module, it is connected to a cross-branch interaction module. In the cross-branch interaction module, the feature maps output by the upper and lower branch residual modules are subjected to channel addition operation, and then the added feature maps are randomly shuffled. After being processed by the residual module and the cross-branch interaction module five times, they are respectively input into the pyramid pooling module for multi-scale pooling operation, and the tree height branch features and the biomass branch features are output. The dual-task feature interaction network includes two parallel branch structures: a tree height decoder and a biomass decoder. Each branch sequentially obtains intermediate tree height features and intermediate biomass features through three groups of residual modules and upsampling operations. The intermediate tree height features and intermediate biomass features are then input into the feature interaction module for cross-task feature fusion, and the final tree height results and biomass results are output. (4) Obtain optical and synthetic aperture radar data within the study area and input them into the collaborative inversion model of tree height and biomass to perform collaborative inversion of tree height and biomass. Compare the predicted values with the actual values to evaluate the inversion accuracy.
[0010] Optionally, step (1) specifically includes the following steps: (1.1) Download the GEDI data of the target area. The tree height data comes from the Level 2 product of the spaceborne lidar, and the biomass data comes from the Level 4 product of the spaceborne lidar. Simultaneously read the L2A and L4A files of the spaceborne lidar, extract the tree height, biomass data, latitude and longitude of the spaceborne lidar light spots, raw waveform data, and elevation information, and filter the light spots using the following filtering conditions: 1) When sensitivity is less than 0.95, the corresponding light spot is removed; 2) When the elevation difference between the GEDI elevation and the SRTM elevation is greater than 30 m, the corresponding light spot is removed; 3) When the degrade flag degrade>0, the stale flag stale_return_flag=0, the degrade flag degrade_flag=1, and the quality flag quality_flag=0, the corresponding light spots are removed; 4) When the error condition flag rx_assess_flag = 1, the corresponding light spot is removed; (1.2) Add a buffer to the latitude and longitude information to generate a vector information file of the corresponding location; (1.3) Download the optical and synthetic aperture radar data of the corresponding location according to the vector information file.
[0011] Optionally, step (2) specifically includes the following steps: (2.1) Optical data preprocessing: Based on the time of tree height and biomass data, select images with a collection time deviation of ≤15 days. If there is no data for the same day, select the image with the closest time. Mask the pixels of the remote sensing image according to the QA_PIXEL band, and then complete the image de-clouding. (2.2) Synthetic Aperture Radar Data Preprocessing: Filter images corresponding to the time of tree height and biomass data. If no corresponding time exists for the current day, select the closest time. Convert radar observation data into backscatter coefficients through radiometric calibration, convert the original data dB values into intensity information, and remove speckle noise from the image. (2.3) Training dataset preparation: The tree height, biomass, optical, and synthetic aperture radar data were cut into multiple data sets at a size of 300 × 300 pixels. The data sets that did not contain valid true values of tree height and biomass were removed. The filtered data were randomly sampled and divided into training, test, and validation sets at a ratio of 40%:30%:30%.
[0012] Optionally, the input data in step (3) is taken from a training data set, and the input data includes optical data and synthetic aperture radar data, wherein the optical data includes 11 band information extracted from the Sentinel-2 image, and the synthetic aperture radar data includes radar backscatter coefficients of two polarization modes extracted from the Sentinel-1 image, and the input data size is uniformly regularized to 300×300×13.
[0013] Optionally, the cross-branch interaction module first adds the two branch features according to the channel dimension to obtain the fusion feature ,in For tree height branch i The output of the residual module, The biomass branch i The output of the residual module, where Representative i The number of channels output by the residual module; Randomly shuffling the channels can be expressed as a random rearrangement of the channel indices: , in, A random permutation function for channel indices.
[0014] Optionally, the pyramid pooling module includes a 3×3 pooling core, a 15×15 pooling core, a 30×30 pooling core and a global pooling core. Perform multi-scale downsampling and average pooling to obtain multi-scale pooling features, which can be expressed as: , , , , in, are input features, , , , ; Then perform channel dimension fusion on the multi-scale pooling features, that is, first 、 and Perform bilinear upsampling respectively to make the image size 100×100, and get 、 and , then add the channel dimensions to get the fusion feature , expressed as: , Final output tree height branch features and biomass branch characteristics .
[0015] Optionally, the feature interaction module uses the tree height intermediate feature and biomass intermediate characteristics Splicing by channel dimension to construct joint features and ,in Represents the channel dimension splicing operation; Then through the learnable weight matrix , 100 is the number of joint feature channels, 50 is the number of single task feature channels, and the dependency weight of the tree height task on the biomass task is calculated. And the dependency weight of the biomass task on the tree height task , the calculation formula is: , , Among them, softmax is the activation function, and then based on the dependency weight, the intermediate features of tree height and biomass are fused across tasks: , , in is the feature after tree height fusion, The features after biomass fusion are finally output through the fully connected layer to complete the output of tree height and biomass inversion results.
[0016] Based on the same inventive concept, the present invention discloses a tree height and biomass collaborative inversion system based on codec dual focusing, comprising: The data acquisition module is used to obtain tree height and biomass data of the target area using space-borne lidar, and to obtain optical and synthetic aperture radar data of the corresponding locations; Data preprocessing module, used to preprocess the acquired data and construct a training data set; A collaborative inversion model construction module is used to establish a multi-task learning model based on encoder and decoder dual focusing. The training dataset is trained through multi-task learning to generate a collaborative inversion model of tree height and biomass. The multi-task learning model includes an encoder focusing module and a decoder focusing module. The encoder focusing module inputs the input data of the training dataset into the residual fusion encoding network to generate tree height branch features and biomass branch features respectively. The decoder focusing module inputs the tree height branch features and biomass branch features into the dual-task feature interaction network to output the tree height inversion results and the biomass inversion results. The residual fusion coding network includes two parallel branch structures: a tree height encoder and a biomass encoder. Each branch passes through five residual modules in sequence. The convolution kernel size of the residual module decreases from large to small, and the number of channels doubles from module to module. After processing each residual module, it is connected to the cross-branch interaction module. In the cross-branch interaction module, the feature maps output by the upper and lower branch residual modules are subjected to channel addition operation, and then the added feature maps are randomly shuffled. After being processed by the residual module and the cross-branch interaction module five times, they are respectively input into the pyramid pooling module for multi-scale pooling operation, and finally the tree height branch feature and the biomass branch feature are output. The dual-task feature interaction network includes two parallel branch structures: a tree height decoder and a biomass decoder. Each branch sequentially obtains intermediate tree height features and intermediate biomass features through three groups of residual modules and upsampling operations. The intermediate tree height features and intermediate biomass features are then input into the feature interaction module for cross-task feature fusion, and the final tree height results and biomass results are output. The collaborative inversion prediction module is used to obtain optical and synthetic aperture radar data in the study area, and input them into the collaborative inversion model of tree height and biomass to perform collaborative inversion of tree height and biomass, and compare the predicted values with the actual values to evaluate the inversion accuracy.
[0017] Based on the same inventive concept, the electronic device described in the present invention includes a processor and a storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method described above.
[0018] Based on the same inventive concept, the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the program implements the steps of the above-mentioned method when executed by a processor.
[0019] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: (1) The present invention designs an encoder focusing module, which improves the efficiency of multi-source data fusion, enhances the cross-modal feature synergy of optical and synthetic aperture radar data, and explores the complementarity between data; (2) The present invention designs a decoder focusing module to enhance the complementary mechanism between strongly correlated target inversion tasks; (3) The present invention uses a dual-focus strategy of encoder and decoder to form a closed-loop mechanism from feature sharing to task interaction, which improves the collaborative inversion accuracy of strongly correlated targets; (4) The present invention designs a residual fusion coding network, which realizes feature complementarity among multiple tasks and enhances generalization ability through cross-branch interaction modules. Through the pyramid pooling module, it takes into account both local details and global information of features, thereby enhancing the model's utilization of spatial information. (5) The present invention designs a dual-task feature interaction network, which realizes adaptive fusion by dynamically calculating task dependency weights, making full use of complementary information while retaining the core features of each task and avoiding interference from invalid information. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 Schematic diagram of the multi-task learning network structure in the present invention; Figure 3 Schematic diagram of the residual fusion coding network structure in the present invention; Figure 4 Schematic diagram of the residual module structure in the present invention; Figure 5 Schematic diagram of the cross-branch interaction mechanism in the present invention; Figure 6 Schematic diagram of the pyramid pooling structure in the present invention; Figure 7 Schematic diagram of the dual-task feature interaction network structure in the present invention; Figure 8 This is a structural diagram of the feature interaction module in the present invention. DETAILED DESCRIPTION
[0021] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0022] Example 1: Figure 1 As shown, the present invention discloses a tree height and biomass collaborative inversion method based on codec dual focusing, comprising the following steps: (1) Use spaceborne lidar to obtain tree height and biomass data in the target area, and obtain optical and synthetic aperture radar data at the corresponding locations; Step (1) specifically includes the following steps: (1.1) Download the GEDI data of the target area from the spaceborne lidar. The tree height data comes from the Level 2 product of the spaceborne lidar, and the biomass data comes from the Level 4 product of the spaceborne lidar. Simultaneously read the L2A and L4A files of the spaceborne lidar, extract the tree height and biomass data, the latitude and longitude of the spaceborne lidar light spots, the raw waveform data, and the elevation information. Use the filtering conditions to filter the light spots and retain the ones with good quality. To ensure the quality of GEDI spaceborne lidar data, the screening criteria are as follows: 1) When sensitivity is less than 0.95, the corresponding light spot is removed; 2) When the elevation difference between the GEDI elevation and the SRTM elevation is greater than 30 m, the corresponding light spot is removed; 3) When the degrade flag degrade>0, the stale flag stale_return_flag=0, the degrade flag degrade_flag=1, and the quality flag quality_flag=0, the corresponding light spots are removed; 4) When the error condition flag rx_assess_flag = 1, the corresponding light spot is removed; The screening information is shown in Table 1 and the screening conditions are shown in Table 2; Table 1 Screening parameters of spaceborne lidar
[0023] Table 2 Screening conditions for spaceborne lidar GEDI
[0024] (1.2) Add a buffer to the latitude and longitude information to generate a vector information file of the corresponding location; (1.3) Download the optical and synthetic aperture radar data for the corresponding location based on the vector information file; the optical and synthetic aperture radar data are from Sentinel-1 and Sentinel-2, as shown in Tables 3 and 4.
[0025] Table 3 Sentinel-1 image extraction parameters
[0026] Table 4 Sentinel-2 image extraction parameters
[0027] (2) Preprocess the acquired data, construct a training data set, and divide it into training set, test set, and validation set; Step (2) specifically includes the following steps: (2.1) Optical data preprocessing: Based on the time of tree height and biomass data, select images with a collection time deviation of ≤15 days. If there is no data for the same day, select the image with the closest time. Mask the pixels of the remote sensing image according to the QA_PIXEL band and use the Fmask algorithm to complete the image cloud removal. (2.2) Synthetic Aperture Radar Data Preprocessing: Filter images corresponding to the time of tree height and biomass data. If no corresponding time exists for the current day, select the closest time. Convert radar observation data into backscatter coefficients through radiometric calibration. The original data information is dB values, which need to be converted into intensity information using the following formula: , Lee filtering is used to remove speckle noise from the image, reducing the impact of noise while retaining image details; (2.3) Training data set preparation: The tree height, biomass, optical and synthetic aperture radar data were cut into multiple data groups with a size of 300×300 pixels, and the data groups without valid true values of tree height and biomass were eliminated to ensure the quality of the data set for model training. Through random sampling, the screened data were divided into training set, test set and validation set in a ratio of 40%:30%:30%. Among them, the tree height and biomass data are the labels of the multi-task learning model, and the optical and synthetic aperture radar data are the input features of the multi-task learning model.
[0028] (3) Establish a multi-task learning model based on dual focus of encoder and decoder, train the training dataset through multi-task learning, and generate a collaborative inversion model of tree height and biomass; like Figure 2 As shown in the figure, the multi-task learning model includes an encoder focusing module and a decoder focusing module, wherein the encoder focusing module inputs the input data of the training dataset into the residual fusion encoding network to generate tree height branch features and biomass branch features respectively, and the decoder focusing module inputs the tree height branch features and biomass branch features into the dual-task feature interaction network to output the tree height inversion results and biomass inversion results.
[0029] The input data are taken from the training dataset and include optical data and synthetic aperture radar data. The optical data are the 11 band information extracted from the Sentinel-2 image in Table 4, and the synthetic aperture radar data are the radar backscatter coefficients of the two polarization modes of the Sentinel-1 image in Table 3. The input data size is uniformly regularized to 300 × 300 × 13.
[0030] like Figure 3As shown in the figure, the residual fusion coding network includes two parallel branch structures: tree height encoder and biomass encoder. Each branch passes through 5 residual modules in sequence. The convolution kernel size of the residual module decreases from large to small, and the number of channels doubles from module to module. After each residual module is processed, it is connected to the cross-branch interaction module. In the cross-branch interaction module, the feature maps output by the upper and lower branch residual modules are added together, and then the added feature maps are randomly shuffled. After being processed by the residual module and the cross-branch interaction module 5 times, they are respectively input into the multi-scale pooling operation in the pyramid pooling module, and finally the tree height branch features and biomass branch features are output.
[0031] The first residual module uses a 7×7 convolution kernel with the following structure: Figure 4 As shown in the figure, the second residual module uses a 5×5 convolution kernel, the third residual module uses a 3×3 convolution kernel, and the fourth and fifth residual modules both use a 1×1 convolution kernel; the feature image size in the five residual modules remains unchanged, the feature image size is 300×300, and the number of feature channels is doubled successively, namely 26, 52, 104, 208 and 416.
[0032] like Figure 4 As shown in the figure, the input features in the first residual module pass through the 7×7 convolution kernel Conv, the normalization layer BN, the ReLU activation function, the 7×7 convolution kernel Conv and the normalization layer BN in sequence, and then are summed with the input features that only pass through the 1×1 convolution kernel Conv, and then pass through the ReLU activation function to output the result.
[0033] In the second residual module, the input features pass through the 5×5 convolution kernel Conv, the normalization layer BN, the ReLU activation function, the 5×5 convolution kernel Conv and the normalization layer BN in sequence, and then are summed with the input features that only pass through the 1×1 convolution kernel Conv, and then pass through the ReLU activation function to output the result.
[0034] In the third residual module, the input features pass through the 3×3 convolution kernel Conv, the normalization layer BN, the ReLU activation function, the 3×3 convolution kernel Conv and the normalization layer BN in sequence, and then are summed with the input features that only pass through the 1×1 convolution kernel, and then pass through the ReLU activation function to output the result.
[0035] The input features in the fourth and fifth residual modules pass through the 1×1 convolution kernel Conv, the normalization layer BN, the ReLU activation function, the 1×1 convolution kernel Conv and the normalization layer BN in turn, and then are summed with the input features that only pass through the 1×1 convolution kernel Conv, and then pass through the ReLU activation function to output the results.
[0036] like Figure 5As shown in the figure, in order to achieve the feature fusion between tree height and biomass branches, the cross-branch interaction module first adds the two branch features according to the channel dimension to obtain the fusion feature , this operation enables the two branch features to be directly fused in the channel dimension and share complementary information; For tree height branch i The output of the residual module, The biomass branch i The output of the residual module, where Representative i The number of channels output by the residual module; Randomly shuffle the channels, destroy the fixed channel order, and enhance feature diversity. Mathematically, it can be expressed as a random rearrangement of channel indices: , in, It is a random permutation function of channel indexes, ensuring that the channel order changes dynamically during each training and improving model generalization.
[0037] After processing by the quintic residual module and the cross-branch interaction module, considering that the spot range of the spaceborne lidar data is 25 meters and the spatial resolution of Sentinel-1 and Sentinel-2 is 10 meters, the resolution of the spaceborne lidar is about three times that of optical and synthetic aperture radar data. In order to achieve spatial matching between the input features and the label information, the pyramid pooling module is input and the pyramid pooling operation is introduced.
[0038] like Figure 6 As shown, the pyramid pooling module includes 3×3 pooling core, 15×15 pooling core, 30×30 pooling core and global pooling core. The input features in the pyramid pooling module are Perform multi-scale downsampling and average pooling to obtain multi-scale pooling features: , , , , in, are input features, , , , ; In order to keep the number of channels unchanged, the channel dimension fusion is performed on the multi-scale pooling features. First, bilinear upsampling is performed on , and respectively, so that the image size is 100×100. That is, first 、 and Perform bilinear upsampling respectively to make the image size 100×100, and get 、 and , then add the channel dimensions to get the fusion feature , expressed as: , After fusion, the number of feature channels is still 416, the size is 100×100, and the final output is the tree height branch feature and biomass branch characteristics , as the input of the decoder focusing module, completing the feature extraction process of the residual fusion coding network.
[0039] In order to achieve matching between input features (optical data and synthetic aperture radar data) and label information (tree height and biomass) in the spatial dimension, a pyramid pooling operation is used to reduce the feature image size to 1 / 3 of the original size while keeping the number of feature channels unchanged, and finally output tree height branch features and biomass branch features.
[0040] like Figure 7 As shown in the figure, the dual-task feature interaction network includes two parallel branch structures: tree height decoder and biomass decoder. Each branch structure passes through three sets of residual modules and upsampling operations in sequence, and then enters the feature interaction module for cross-task fusion to output the final tree height and biomass results.
[0041] The input data of the dual-task feature interaction network are the tree height branch features and biomass branch features generated by the encoder focus module. The two sets of feature data are processed in parallel. The dual-task feature interaction network structure is as follows: Figure 7 As shown, the input data is the tree height branch feature generated by the encoder focus module and biomass branch characteristics ,The two sets of data are input into the two parallel branches of tree height decoder and biomass decoder respectively, and the processing flow is started in parallel.
[0042] First, the data passes through three sets of residual modules and upsampling operations. After processing by the residual modules, the feature image size remains unchanged, while the number of feature channels is reduced. Because upsampling gradually increases the feature image size, effectively extracting features at different scales, the convolution kernel sizes used by the residual modules are set from small to large. After three upsampling operations, the feature image size is restored to the original input size.
[0043] The tree height decoder and biomass decoder branches each pass through three sets of residual modules followed by upsampling operations. Because upsampling increases the size of the feature image, to ensure approximately the same receptive field, the convolution kernels used in the residual modules increase in size, to 3×3, 5×5, and 7×7, respectively. After the residual modules, the feature image size remains unchanged, while the number of features decreases, to 200, 100, and 50, respectively. Each residual module is followed by an upsampling operation, with the feature image size increasing in size, to 150×150, 200×200, and 300×300, before returning to the original input size. The number of feature channels remains unchanged.
[0044] After three sets of residual modules and upsampling operations, the intermediate features of the tree height decoder and the biomass decoder are respectively the tree height intermediate features Intermediate characteristics with biomass , enter the feature interaction module to learn the cross-task dependency between tree height and biomass tasks.
[0045] like Figure 8 As shown, the feature interaction module will be the middle feature of the tree height and biomass intermediate characteristics Splicing by channel dimension to construct joint features and ,in Represents the channel dimension splicing operation; Then through the learnable weight matrix , 100 is the number of joint feature channels, 50 is the number of single task feature channels, and the dependency weight of the tree height task on the biomass task is calculated. And the dependency weight of the biomass task on the tree height task , the specific calculation formula is as follows: , , Among them, the softmax function ensures weight normalization so that Reflects the dependence ratio of tree height task on biomass task characteristics, Reflects the proportion of biomass task dependence on tree height task characteristics; Then, based on the dependency weights, cross-task fusion is performed on the tree height and biomass task features: , , in and It is the intermediate feature of the tree height task decoder and the biomass task decoder, which is spliced through channels in the feature interaction module. Forming joint features, and is the learnable weight matrix, and It is the feature vector of the tree height task and biomass task after cross-task interactive processing; finally, the output of the tree height and biomass inversion results is completed through the fully connected layer.
[0046] The encoder and decoder dual-focus strategy in this paper is optimized by a multi-task joint loss function, where the loss function consists of: using mean square error (MSE) as the basic loss, and dynamically adjusting the tree height loss in combination with the task uncertainty weighting mechanism. and biomass loss The weight of the total loss L The calculation formula is: ,
[0047] in, is the uncertainty parameter of the tree height task, Uncertain parameters for biomass tasks are automatically learned through the network; is the L2 regularization term of the encoder parameters, is the regularization coefficient.
[0048] Optimization strategy: The Adam optimizer is used, the initial learning rate is set to 1e-4, and it is decayed by 0.8 after every 50 rounds of training; the training batch size is set to 32, and 200 rounds are iterated until the loss function converges. The convergence threshold is the loss fluctuation amplitude.
[0049] (4) Obtain optical and synthetic aperture radar data within the study area and input them into the tree height and biomass collaborative inversion model to perform tree height and biomass collaborative inversion. Compare the predicted values with the actual values to evaluate the inversion accuracy. The root mean square error (RMSE) is used as the accuracy evaluation indicator, and the formula is as follows: , Where n is the total number of samples, is the model prediction value, is the true value, and the range of the root mean square error is , when the predicted value is completely consistent with the true value, it is equal to 0, that is, a perfect prediction; the worse the prediction effect, the larger the root mean square error value.
[0050] Example 2: The present invention discloses a tree height and biomass collaborative inversion system based on codec dual focusing, comprising: The data acquisition module is used to obtain tree height and biomass data in the target area using spaceborne lidar, and obtain optical and synthetic aperture radar data at the corresponding location. It downloads the GEDI data of the spaceborne lidar in the target area. The tree height data comes from the level 2 product of the spaceborne lidar, and the biomass data comes from the level 4 product of the spaceborne lidar. It also reads the L2A and L4A files of the spaceborne lidar, extracts the tree height and biomass data, the latitude and longitude of the spaceborne lidar light spots, the original waveform data, and the elevation information. It then uses the filtering conditions to filter the light spots and retain the ones with good quality. To ensure the quality of GEDI spaceborne lidar data, the screening criteria are as follows: 1) When sensitivity is less than 0.95, the corresponding light spot is removed; 2) When the elevation difference between the GEDI elevation and the SRTM elevation is greater than 30 m, the corresponding light spot is removed; 3) When the degrade flag degrade>0, the stale flag stale_return_flag=0, the degrade flag degrade_flag=1, and the quality flag quality_flag=0, the corresponding light spots are removed; 4) When the error condition flag rx_assess_flag = 1, the corresponding light spot is removed; the screening information is shown in Table 1, and the screening conditions are shown in Table 2; Add a buffer to the latitude and longitude information to generate a vector information file of the corresponding location; Download the optical and synthetic aperture radar data of the corresponding location according to the vector information file; the optical and synthetic aperture radar data come from Sentinel-1 and Sentinel-2, as shown in Tables 3 and 4.
[0051] The data preprocessing module is used to preprocess the acquired data and construct a training dataset. Optical data preprocessing: Based on the time of tree height and biomass data, images with a collection time deviation of ≤15 days are selected. If there is no data for the same day, the image with the closest time is selected. The pixels of the remote sensing image are masked according to the QA_PIXEL band, and the Fmask algorithm is used to complete the image cloud removal. Synthetic aperture radar data preprocessing: Based on the time of tree height and biomass data, select images of the corresponding time. If there is no corresponding time on the same day, select the closest time. Convert the radar observation data into backscatter coefficients through radiometric calibration. The original data information is dB value, which needs to be converted into intensity information. The formula is as follows: , Lee filtering is used to remove speckle noise from the image, reducing the impact of noise while retaining image details; Preparation of training dataset: The tree height, biomass, optical, and synthetic aperture radar data were cut into multiple data groups at a size of 300×300 pixels. The data groups without valid true values of tree height and biomass were removed to ensure the quality of the dataset for model training. Through random sampling, the screened data were divided into training set, test set, and validation set in a ratio of 40%:30%:30%. The tree height and biomass data were labels for the multi-task learning model, and the optical and synthetic aperture radar data were input features for the multi-task learning model.
[0052] The collaborative inversion model construction module is used to establish a multi-task learning model based on encoder and decoder dual focusing, train the training data set through multi-task learning, and generate a collaborative inversion model of tree height and biomass; the multi-task learning model includes an encoder focusing module and a decoder focusing module, wherein the encoder focusing module inputs the input data of the training data set into the residual fusion encoding network to generate tree height branch features and biomass branch features respectively, and the decoder focusing module inputs the tree height branch features and biomass branch features into the dual-task feature interaction network to output the tree height inversion results and biomass inversion results.
[0053] The residual fusion coding network consists of two parallel branch structures: tree height encoder and biomass encoder. Each branch passes through five residual modules in sequence. The convolution kernel size of the residual module decreases from large to small, and the number of channels doubles from module to module. After each residual module is processed, it is connected to the cross-branch interaction module. In the cross-branch interaction module, the feature maps output by the upper and lower branch residual modules are added together, and then the added feature maps are randomly shuffled. After being processed by the residual module and the cross-branch interaction module five times, they are respectively input into the pyramid pooling module for multi-scale pooling operations, and finally the tree height branch features and biomass branch features are output.
[0054] The dual-task feature interaction network consists of two parallel branch structures: tree height decoder and biomass decoder. Each branch obtains intermediate tree height features and intermediate biomass features through three groups of residual modules and upsampling operations in turn. The intermediate tree height features and intermediate biomass features are then input into the feature interaction module for cross-task feature fusion, and the final tree height results and biomass results are output.
[0055] The input data are taken from the training dataset and include optical data and synthetic aperture radar data. The optical data are the 11 band information extracted from the Sentinel-2 image in Table 4, and the synthetic aperture radar data are the radar backscatter coefficients of the two polarization modes of the Sentinel-1 image in Table 3. The input data size is uniformly regularized to 300 × 300 × 13.
[0056] The first residual module uses a 7×7 convolution kernel with the following structure: Figure 4As shown in the figure, the second residual module uses a 5×5 convolution kernel, the third residual module uses a 3×3 convolution kernel, and the fourth and fifth residual modules both use a 1×1 convolution kernel; the feature image size in the five residual modules remains unchanged, the feature image size is 300×300, and the number of feature channels is doubled successively, namely 26, 52, 104, 208 and 416.
[0057] like Figure 4 As shown in the figure, the input features in the first residual module pass through the 7×7 convolution kernel Conv, the normalization layer BN, the ReLU activation function, the 7×7 convolution kernel Conv and the normalization layer BN in sequence, and then are summed with the input features that only pass through the 1×1 convolution kernel Conv, and then pass through the ReLU activation function to output the result.
[0058] In the second residual module, the input features pass through the 5×5 convolution kernel Conv, the normalization layer BN, the ReLU activation function, the 5×5 convolution kernel Conv and the normalization layer BN in sequence, and then are summed with the input features that only pass through the 1×1 convolution kernel Conv, and then pass through the ReLU activation function to output the result.
[0059] In the third residual module, the input features pass through the 3×3 convolution kernel Conv, the normalization layer BN, the ReLU activation function, the 3×3 convolution kernel Conv and the normalization layer BN in sequence, and then are summed with the input features that only pass through the 1×1 convolution kernel, and then pass through the ReLU activation function to output the result.
[0060] The input features in the fourth and fifth residual modules pass through the 1×1 convolution kernel Conv, the normalization layer BN, the ReLU activation function, the 1×1 convolution kernel Conv and the normalization layer BN in turn, and then are summed with the input features that only pass through the 1×1 convolution kernel Conv, and then pass through the ReLU activation function to output the results.
[0061] like Figure 5 As shown in the figure, in order to achieve the feature fusion between tree height and biomass branches, the cross-branch interaction module first adds the two branch features according to the channel dimension to obtain the fusion feature , this operation enables the two branch features to be directly fused in the channel dimension and share complementary information; For tree height branch i The output of the residual module, The biomass branch i The output of the residual module, where Representative i The number of channels output by the residual module; Randomly shuffle the channels, destroy the fixed channel order, and enhance feature diversity. Mathematically, it can be expressed as a random rearrangement of channel indices: , in, It is a random permutation function of channel indexes, ensuring that the channel order changes dynamically during each training and improving model generalization.
[0062] After processing by the quintic residual module and the cross-branch interaction module, considering that the spot range of the spaceborne lidar data is 25 meters and the spatial resolution of Sentinel-1 and Sentinel-2 is 10 meters, the resolution of the spaceborne lidar is about three times that of optical and synthetic aperture radar data. In order to achieve spatial matching between the input features and the label information, the pyramid pooling module is input and the pyramid pooling operation is introduced.
[0063] like Figure 6 As shown, the pyramid pooling module includes 3×3 pooling core, 15×15 pooling core, 30×30 pooling core and global pooling core. The input features in the pyramid pooling module are Perform multi-scale downsampling and average pooling to obtain multi-scale pooling features: , , , , in, are input features, , , , ; In order to keep the number of channels unchanged, channel dimension fusion is performed on the multi-scale pooling features. 、 and Perform bilinear upsampling respectively to make the image size 100×100, and get 、 and , then add the channel dimensions to get the fusion feature , expressed as: , After fusion, the number of feature channels is still 416, the size is 100×100, and the final output is the tree height branch feature and biomass branch characteristics , as the input of the decoder focusing module, completing the feature extraction process of the residual fusion coding network.
[0064] In order to achieve matching between input features (optical data and synthetic aperture radar data) and label information (tree height and biomass) in the spatial dimension, a pyramid pooling operation is used to reduce the feature image size to 1 / 3 of the original size while keeping the number of feature channels unchanged, and finally output tree height branch features and biomass branch features.
[0065] The input data of the dual-task feature interaction network is the tree-high branch features generated by the encoder focus module and biomass branch characteristics , two sets of feature data are processed in parallel; the dual-task feature interaction network structure is as follows Figure 7 As shown, the input data are the tree height branch features and biomass branch features generated by the encoder focus module. The two sets of data are respectively input into the two parallel branches of the tree height decoder and the biomass decoder, and the processing flow is started in parallel.
[0066] First, the data passes through three sets of residual modules and upsampling operations. After processing by the residual modules, the feature image size remains unchanged, while the number of feature channels is reduced. Because upsampling gradually increases the feature image size, effectively extracting features at different scales, the convolution kernel sizes used by the residual modules are set from small to large. After three upsampling operations, the feature image size is restored to the original input size.
[0067] The tree height decoder and biomass decoder branches each pass through three sets of residual modules followed by upsampling operations. Because upsampling increases the size of the feature image, to ensure approximately the same receptive field, the convolution kernels used in the residual modules increase in size, to 3×3, 5×5, and 7×7, respectively. After the residual modules, the feature image size remains unchanged, while the number of features decreases, to 200, 100, and 50, respectively. Each residual module is followed by an upsampling operation, with the feature image size increasing in size, to 150×150, 200×200, and 300×300, before returning to the original input size. The number of feature channels remains unchanged.
[0068] After three sets of residual modules and upsampling operations, the intermediate features of the tree height decoder and the biomass decoder are respectively the tree height intermediate features Intermediate characteristics with biomass , enter the feature interaction module to learn the cross-task dependency between tree height and biomass tasks.
[0069] like Figure 8 As shown, the feature interaction module will be the middle feature of the tree height and biomass intermediate characteristics Splicing by channel dimension to construct joint features and ,in Represents the channel dimension splicing operation; Then through the learnable weight matrix , 100 is the number of joint feature channels, 50 is the number of single task feature channels, and the dependency weight of the tree height task on the biomass task is calculated. And the dependency weight of the biomass task on the tree height task , the specific calculation formula is as follows: , , Among them, the softmax function ensures weight normalization so that Reflects the dependence ratio of tree height task on biomass task characteristics, Reflects the proportion of biomass task dependence on tree height task characteristics; Then, based on the dependency weights, cross-task fusion is performed on the tree height and biomass task features: , , in and It is the intermediate feature of the tree height task decoder and the biomass task decoder, which is spliced through channels in the feature interaction module. Forming joint features, and is the learnable weight matrix, and It is the feature vector of the tree height task and biomass task after cross-task interactive processing; finally, the output of the tree height and biomass inversion results is completed through the fully connected layer.
[0070] The encoder and decoder dual-focus strategy in this paper is optimized by a multi-task joint loss function, where the loss function consists of: using mean square error (MSE) as the basic loss, and dynamically adjusting the tree height loss in combination with the task uncertainty weighting mechanism. and biomass loss The weight of the total loss L The calculation formula is: , in, is the uncertainty parameter of the tree height task, Uncertain parameters for biomass tasks are automatically learned through the network; is the L2 regularization term of the encoder parameters, is the regularization coefficient.
[0071] Optimization strategy: The Adam optimizer is used, the initial learning rate is set to 1e-4, and it is decayed by 0.8 after every 50 rounds of training; the training batch size is set to 32, and 200 rounds are iterated until the loss function converges. The convergence threshold is the loss fluctuation amplitude.
[0072] The collaborative inversion prediction module is used to obtain optical and synthetic aperture radar data in the study area, and input them into the collaborative inversion model of tree height and biomass to perform collaborative inversion of tree height and biomass, and compare the predicted values with the actual values to evaluate the inversion accuracy.
[0073] The root mean square error (RMSE) is used as the accuracy evaluation indicator, and the formula is as follows: , Where n is the total number of samples, is the model prediction value, is the true value, and the range of the root mean square error is , when the predicted value is completely consistent with the true value, it is equal to 0, that is, a perfect prediction; the worse the prediction effect, the larger the root mean square error value.
[0074] Embodiment 3: In this embodiment, an electronic device includes a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method described above.
[0075] Embodiment 4: In this embodiment, a computer-readable storage medium stores a computer program, which implements the steps of the above method when executed by a processor.
Claims
1. A tree height and biomass collaborative inversion method based on codec dual focusing, characterized by: The following steps are involved: (1) Use spaceborne lidar to obtain tree height and biomass data in the target area, and obtain optical and synthetic aperture radar data at the corresponding locations; (2) Preprocess the acquired data and construct a training data set; (3) Establish a multi-task learning model based on dual focus of encoder and decoder, train the training dataset through multi-task learning, and generate a collaborative inversion model of tree height and biomass; The multi-task learning model includes an encoder focusing module and a decoder focusing module. The encoder focusing module feeds the input data of the training dataset into the residual fusion encoding network to generate tree height branch features and biomass branch features respectively. The decoder focusing module feeds the tree height branch features and biomass branch features into the dual-task feature interaction network to output tree height inversion results and biomass inversion results. The residual fusion coding network includes two parallel branch structures: a tree height encoder and a biomass encoder. Each branch passes through five residual modules in sequence. The convolution kernel size of the residual module decreases from large to small, and the number of channels doubles from module to module. After processing each residual module, it is connected to a cross-branch interaction module. In the cross-branch interaction module, the feature maps output by the upper and lower branch residual modules are subjected to channel addition operation, and then the added feature maps are randomly shuffled. After being processed by the residual module and the cross-branch interaction module five times, they are respectively input into the pyramid pooling module for multi-scale pooling operation, and the tree height branch features and the biomass branch features are output. The dual-task feature interaction network includes two parallel branch structures: a tree height decoder and a biomass decoder. Each branch sequentially obtains intermediate tree height features and intermediate biomass features through three groups of residual modules and upsampling operations. The intermediate tree height features and intermediate biomass features are then input into the feature interaction module for cross-task feature fusion, and the final tree height results and biomass results are output. (4) Obtain optical and synthetic aperture radar data within the study area and input them into the collaborative inversion model of tree height and biomass to perform collaborative inversion of tree height and biomass. Compare the predicted values with the actual values to evaluate the inversion accuracy.
2. The tree height and biomass collaborative inversion method based on codec dual focusing according to claim 1 is characterized by: The step (1) specifically includes the following steps: (1.1) Download the GEDI data of the target area. The tree height data comes from the Level 2 product of the spaceborne lidar, and the biomass data comes from the Level 4 product of the spaceborne lidar. Simultaneously read the L2A and L4A files of the spaceborne lidar, extract the tree height, biomass data, latitude and longitude of the spaceborne lidar light spots, raw waveform data, and elevation information, and filter the light spots using the following filtering conditions: 1) When sensitivity is less than 0.95, the corresponding light spot is removed; 2) When the elevation difference between the GEDI elevation and the SRTM elevation is greater than 30 m, the corresponding light spot is removed; 3) When the degrade flag degrade>0, the stale flag stale_return_flag=0, the degrade flag degrade_flag=1, and the quality flag quality_flag=0, the corresponding light spots are removed; 4) When the error condition flag rx_assess_flag = 1, the corresponding light spot is removed; (1.2) Add a buffer to the latitude and longitude information to generate a vector information file of the corresponding location; (1.3) Download the optical and synthetic aperture radar data of the corresponding location according to the vector information file.
3. The tree height and biomass collaborative inversion method based on codec dual focusing according to claim 1 is characterized by: The step (2) specifically includes the following steps: (2.1) Optical data preprocessing: Based on the time of tree height and biomass data, select images with a collection time deviation of ≤15 days. If there is no data for the same day, select the image with the closest time. Mask the pixels of the remote sensing image according to the QA_PIXEL band, and then complete the image de-clouding. (2.2) Synthetic Aperture Radar Data Preprocessing: Filter images corresponding to the time of tree height and biomass data; If there is no corresponding time on the day, the closest time will be selected; The radar observation data is converted into backscatter coefficients through radiometric calibration, the original data information dB value is converted into intensity information, and then the speckle noise in the image is removed; (2.3) Training dataset preparation: The tree height, biomass, optical, and synthetic aperture radar data were cut into multiple data sets at a size of 300 × 300 pixels. The data sets that did not contain valid true values of tree height and biomass were removed. The filtered data were randomly sampled and divided into training, test, and validation sets at a ratio of 40%:30%:30%.
4. The tree height and biomass collaborative inversion method based on codec dual focusing according to claim 1, characterized in that: The input data in step (3) is taken from the training data set, and the input data includes optical data and synthetic aperture radar data, wherein the optical data includes 11 band information extracted from the Sentinel-2 image, and the synthetic aperture radar data includes radar backscatter coefficients of two polarization modes extracted from the Sentinel-1 image. The input data size is uniformly regularized to 300×300×13.
5. The tree height and biomass collaborative inversion method based on codec dual focusing according to claim 1, characterized in that: In the cross-branch interaction module, the two branch features are first added according to the channel dimension to obtain the fusion feature ,in For tree height branch i The output of the residual module, The biomass branch i The output of the residual module, where Representative i The number of channels output by the residual module; Again Randomly shuffling the channels can be expressed as a random rearrangement of the channel indices: , in, A random permutation function for channel indices.
6. The tree height and biomass collaborative inversion method based on codec dual focusing according to claim 1, characterized in that: The pyramid pooling module includes 3×3 pooling core, 15×15 pooling core, 30×30 pooling core and global pooling core. Perform multi-scale downsampling and average pooling to obtain multi-scale pooling features, which can be expressed as: , , , , in, are input features, , , , ; Then perform channel dimension fusion on the multi-scale pooling features, that is, first 、 and Perform bilinear upsampling respectively to make the image size 100×100, and get 、 and , then add the channel dimensions to get the fusion feature , expressed as: , Final output tree height branch features and biomass branch characteristics .
7. The tree height and biomass collaborative inversion method based on codec dual focusing according to claim 1, characterized in that: In the feature interaction module, the tree height intermediate feature and biomass intermediate characteristics Splice by channel dimension to construct joint features and ,in Represents the channel dimension splicing operation; Then through the learnable weight matrix , 100 is the number of joint feature channels, 50 is the number of single task feature channels, and the dependency weight of the tree height task on the biomass task is calculated. And the dependency weight of the biomass task on the tree height task , the calculation formula is: , , Among them, softmax is the activation function, and then based on the dependency weight, the intermediate features of tree height and biomass are fused across tasks: , , in is the feature after tree height fusion, The features after biomass fusion are finally output through the fully connected layer to complete the output of tree height and biomass inversion results.
8. A tree height and biomass collaborative inversion system based on codec dual focusing, characterized by: include: The data acquisition module is used to obtain tree height and biomass data of the target area using space-borne lidar, and to obtain optical and synthetic aperture radar data of the corresponding locations; Data preprocessing module, used to preprocess the acquired data and construct a training data set; The collaborative inversion model construction module is used to establish a multi-task learning model based on the dual focus of the encoder and decoder. The training dataset is trained through multi-task learning to generate a collaborative inversion model of tree height and biomass. The multi-task learning model includes an encoder focusing module and a decoder focusing module. The encoder focusing module feeds the input data of the training dataset into the residual fusion encoding network to generate tree height branch features and biomass branch features respectively. The decoder focusing module feeds the tree height branch features and biomass branch features into the dual-task feature interaction network to output tree height inversion results and biomass inversion results. The residual fusion coding network includes two parallel branch structures: a tree height encoder and a biomass encoder. Each branch passes through five residual modules in sequence. The convolution kernel size of the residual module decreases from large to small, and the number of channels doubles from module to module. After processing each residual module, it is connected to the cross-branch interaction module. In the cross-branch interaction module, the feature maps output by the upper and lower branch residual modules are subjected to channel addition operation, and then the added feature maps are randomly shuffled. After being processed by the residual module and the cross-branch interaction module five times, they are respectively input into the pyramid pooling module for multi-scale pooling operation, and finally the tree height branch feature and the biomass branch feature are output. The dual-task feature interaction network includes two parallel branch structures: a tree height decoder and a biomass decoder. Each branch sequentially obtains intermediate tree height features and intermediate biomass features through three groups of residual modules and upsampling operations. The intermediate tree height features and intermediate biomass features are then input into the feature interaction module for cross-task feature fusion, and the final tree height results and biomass results are output. The collaborative inversion prediction module is used to obtain optical and synthetic aperture radar data in the study area, and input them into the collaborative inversion model of tree height and biomass to perform collaborative inversion of tree height and biomass, and compare the predicted values with the actual values to evaluate the inversion accuracy.
9. An electronic device, characterized in that: including processor and storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Power transmission line channel tree height inversion method based on laser radar and optical remote sensing
CN111414891A
Ground penetrating radar inversion method and device, electronic equipment and storage medium
CN116893412A
Cloud detection method and system based on coding and decoding attention interaction, medium and equipment
CN117292276A
Ground penetrating radar data inversion method based on multi-scale supervised generative adversarial network
CN118709524A
High-precision soil three-dimensional surveying and mapping method based on radar detection
CN119355719A