A high-resolution and large-scale forest biomass remote sensing prediction method and system

By acquiring multi-source data for geo-registration and feature extraction, and using the backbone network of forest biomass prediction model for prediction, the limitations of forest biomass prediction in traditional methods are solved, and high-resolution, large-scale accurate prediction and model generalization capabilities are achieved.

CN119323736BActive Publication Date: 2025-05-06UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411866559.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Traditional forest biomass prediction methods have limitations in both time scale and spatial range, making it difficult to achieve accurate predictions at high resolution and large scale.

Method used

High-resolution large-scale forest biomass remote sensing prediction method is adopted to obtain multi-source data (microwave radar remote sensing image data, optical remote sensing image data, meteorological data and topographic data), and joint geographic registration of geometric features and regional features is performed, feature variables are extracted, and input them into the forest biomass prediction model backbone network for prediction. The model includes an image serialization module, a two-branch feature extraction network that integrates the channel prior and a regression network. Feature extraction is achieved through the channel prior module and the state space feature extraction module to achieve more accurate forest biomass estimation.

Benefits of technology

Accurate prediction of high-resolution large-scale forest biomass is achieved, which alleviates the saturation problem of biomass estimation in traditional methods, and improves the generalization ability and maintenance efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119323736B_ABST
    Figure CN119323736B_ABST
Patent Text Reader

Abstract

The invention discloses a high-resolution large-scale forest biomass remote sensing prediction method and system, comprising: obtaining multi-source data within a geographic range to be predicted; combining geometric features and regional features to perform geo-registration on the multi-source data and extracting feature variables; inputting a forest biomass prediction model backbone network, including an image serialization module, a dual-branch feature extraction network fused with channel priors, and a regression network, wherein the image serialization module slices the image and converts it into an image sequence, and then inputs the dual-branch feature extraction network, wherein the dual-branch feature extraction network first extracts a feature map after channel and space weighting with a channel prior module, then performs layer-by-layer downsampling on the weighted feature map in four stages, and uses a state space feature extraction module to perform dual-branch feature extraction, and finally performs upsampling regression through a regression network to output a forest biomass prediction result within the predicted geographic range. The invention can predict forest biomass.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of forest biomass prediction, and in particular to a high-resolution large-scale forest biomass remote sensing prediction method and system. Background Art

[0002] As the largest organic carbon pool, forests play an indispensable role in promoting the global carbon cycle and buffering global warming. Forest biomass is the energy basis and material source for the operation of the entire forest ecosystem. It is the basis for studying forest productivity, carbon cycle and global change, and is also one of the most important indicators for assessing the carbon sequestration capacity of forests. Forest biomass refers to the total amount of organic matter of all living substances in a forest community per unit area over a period of time, including the biomass of trees in the forest (total mass of roots, stems, leaves, branches and litter) and the biomass of the understory vegetation layer. Therefore, accurate and large-scale prediction of forest biomass and acquisition of spatial distribution information are of great significance for guiding reasonable forest management activities, ensuring ecosystem security and maintaining global carbon balance. However, traditional prediction methods have limitations in both time scale and spatial scope. Summary of the invention

[0003] The present invention provides a high-resolution large-scale forest biomass remote sensing prediction method and system to solve the problems existing in the above-mentioned prior art. The technical solution is as follows:

[0004] On the one hand, a high-resolution and large-scale forest biomass remote sensing prediction method is provided, including:

[0005] S1. Acquire multi-source data within the geographic range to be predicted, wherein the multi-source data includes microwave radar remote sensing image data, optical remote sensing image data, meteorological data and terrain data;

[0006] S2. Geo-referencing the multi-source data by combining geometric features and regional features;

[0007] S3, extracting characteristic variables from the multi-source data after georeferencing;

[0008] S4. Input the image of the extracted feature variables into the backbone network of the forest biomass prediction model, the backbone network includes an image serialization module, a dual-branch feature extraction network that integrates channel priors, and a regression network. The image serialization module slices the image and converts it into an image sequence, which is then input into the dual-branch feature extraction network. The dual-branch feature extraction network first extracts the feature map after channel and space weighting with the channel prior module, and then downsamples the weighted feature map layer by layer in four stages, and uses the state space feature extraction module to perform dual-branch feature extraction, and finally performs upsampling regression through the regression network to output the forest biomass prediction result within the predicted geographical range.

[0009] Optionally, the S2 specifically includes:

[0010] S21, using the optical remote sensing image data as a reference image and other multi-source data as images to be registered, performing geometric feature pre-matching to obtain an initial matching point set, including:

[0011] Each image is divided into 20×20 non-overlapping blocks, and the FAST value of each pixel in each non-overlapping block is calculated. The FAST value is defined as the sum of the absolute values ​​of the pixel differences between a total of 16 pixels on a circle with a radius of 3 and each pixel as the center and the center of the circle. The pixel with the largest FAST value is selected as the feature point in each non-overlapping block, and the feature points in 20×20×1 images are obtained;

[0012] Assume that the coordinates of a feature point in the reference image are , calculate the Euclidean distance between the feature point and all feature points in the image to be registered , select the feature point in the image to be registered that minimizes the Euclidean distance as the matching feature point, assuming the coordinates are ,but ;

[0013] Complete the matching of all feature points in the reference image to form the initial matching point set;

[0014] S22, performing regional grayscale constrained registration optimization on the initial matching point set, eliminating mismatched point pairs, and obtaining an accurate matching point set, including:

[0015] Based on the initial matching point set, the window grayscale similarity coefficient of the matching point pairs in the two images is calculated using the following formula: :

[0016]

[0017] In the above formula , Respectively represent the grayscale values ​​of the pixels in the 5×5 window centered on the matching point pair in the reference image and the image to be registered, , are the average values ​​of the pixels in the corresponding windows respectively;

[0018] When the grayscale similarity coefficient of the window If the maximum value of is greater than the preset threshold T, the matching point pair is retained, otherwise, the mismatched point pairs are removed to obtain an accurate matching point set;

[0019] S23. Correct the image using an affine transformation model to reduce significant geometric deformation and complete the registration.

[0020] Optionally, a total of 21 characteristic variables are extracted in S3, including:

[0021] Four characteristic variables are extracted from the microwave radar remote sensing image data, namely, VV polarization backscattering coefficient, VH polarization backscattering coefficient, backscattering coefficient polarization ratio VV / VH, and standard depolarization rate (VV-VH) / (VH+VV);

[0022] Extracting spectral information and vegetation index from the optical remote sensing image data, specifically including eight characteristic variables, namely, four single-band reflectances, normalized difference vegetation index NDVI, soil adjusted vegetation index SAVI, simple ratio vegetation index SR and difference vegetation index DVI;

[0023] Based on the gray-level co-occurrence matrix, the texture features of the four single bands and the backscattering coefficient polarization ratio VV / VH are calculated, and energy values ​​are constructed to describe the texture features, and 5 energy values ​​are obtained as 5 feature variables. Each element in the gray-level co-occurrence matrix refers to the joint distribution of the gray levels of two pixels with a certain spatial position relationship. and direction angle When As the starting point, grayscale appears The probability of the whole image is arranged into a matrix to obtain the gray-level co-occurrence matrix , specifically including: moving any point in the image and a point away from it Constitute a point pair, and set the gray value of the point pair to be , assuming that the maximum gray level of the image is L, then and The combination of For the entire image, count each The number of times the value appears, and then arrange it into a square matrix, and then The total number of occurrences is normalized to get the probability , the resulting matrix is ​​the gray-level co-occurrence matrix; the energy value It is a measure of the uniformity of the grayscale distribution and the coarseness of the texture of the image. If the element values ​​of the grayscale co-occurrence matrix are similar, the energy value is small, indicating a fine texture; when the image texture is coarse, the energy value is large, which is expressed by the following formula:

[0024] ;

[0025] In addition, two characteristic variables, elevation and slope, were extracted from the digital elevation model data of the Space Shuttle Radar Topography Mission, and two characteristic variables, annual mean temperature and annual precipitation, were selected, for a total of 21 characteristic variables.

[0026] Optionally, the channel prior module first performs average pooling and maximum pooling on the input image sequence, then generates a channel attention map using an activation function through a shared multilayer perceptron, and obtains a channel prior image sequence with channel attention by multiplying the input image sequence and elements of the channel attention map;

[0027] Then, the channel prior image sequence is input into a multi-scale separable convolution module to generate a spatial attention map. The multi-scale separable convolution module sets three separable convolution branches for parallel processing to capture multi-scale spatial features. The three separable convolution branches are respectively composed of 1×5 convolution and 5×1 convolution, 1×7 convolution and 7×1 convolution, and 1×9 convolution and 9×1 convolution. The obtained features are mixed with the features previously obtained by 3×3 convolution through channels to obtain the spatial attention map. The spatial attention map is input into a 1×1 convolution for channel adjustment and then combined with the channel prior image sequence. Through element-by-element multiplication operations, the spatial attention weights and channel attention weights are mapped to the feature map to obtain the weighted feature map, thereby realizing the recognition of key features in the image.

[0028] Optionally, the state space feature extraction module is composed of layer normalization, a dual-branch feature extraction module, a multi-layer perceptron, layer normalization, a downsampling module and a residual connection, and the dual-branch feature extraction module is used to further accurately extract features, and the residual connection and layer normalization are used to alleviate the gradient vanishing problem;

[0029] The dual-branch feature extraction module includes an omnidirectional selective state space branch and a hybrid convolution branch. The input is first divided into two equal-sized input branches 1 and 2, and then input into the two branches respectively. The omnidirectional selective state space branch is used to capture multi-scanning direction features, and the hybrid convolution branch is used to extract contextual information of different scales. The outputs of the two branches are merged along the channel dimension of the feature map, and channel shuffling is used to promote information interaction between the two sub-input channels.

[0030] Optionally, after the input branch 1 inputs the omnidirectional selective state space branch, it is input into the omnidirectional selective scanning module after layer normalization, linear layer, mixed convolution, and activation function, and the input is flattened into 8 groups of sequences along the horizontal, vertical, oblique, reverse oblique and opposite directions, and the input sequence is modeled using a selective state space model to achieve information compression. The scanning results in all directions are accumulated together to form an output. The output integrates the features in 8 directions, so that the model can capture and model the scale space features of the image in all directions, and obtain image features with rich spatial position information. The output is then multiplied with the result obtained by the other branch after layer normalization, the linear layer and the activation function, and finally the output of the omnidirectional selective state space branch is obtained after layer normalization. The selective state space model introduces a selective mechanism, and through the selection of input sequence information, the relationship between input features and biomass is learned during continuous training, and effective information is identified and captured.

[0031] Optionally, after the input branch 2 is input into the hybrid convolution branch, it undergoes four cascaded hybrid convolutions, wherein each of the hybrid convolutions includes an expanded convolution and a multi-scale convolution, and the expansion rates of the expanded convolution are 1, 2, 3, and 1, respectively, to avoid the grid effect caused by discontinuous data, and to use expanded convolutions with different expansion rates to improve the receptive field and capture a wider range of contextual information; the multi-scale convolution uses parallel 1×1 convolution, 3×3 convolution, and 5×5 convolution to obtain multi-scale information of feature variables, and through aggregation to fully learn the geographic spatial effects of forests, learn their spatial correlations at different spatial sizes, and improve the understanding of the complexity of forest ecosystems.

[0032] On the other hand, a high-resolution large-scale forest biomass remote sensing prediction system is provided, the system comprising:

[0033] An acquisition module, used to acquire multi-source data within the geographic range to be predicted, wherein the multi-source data includes microwave radar remote sensing image data, optical remote sensing image data, meteorological data and terrain data;

[0034] A geo-referencing module, used for combining geometric features and regional features to geo-referencing the multi-source data;

[0035] The extraction module is used to extract characteristic variables from multi-source data after georeferencing;

[0036] The prediction module is used to input the image of the extracted feature variables into the backbone network of the forest biomass prediction model, and the backbone network includes an image serialization module, a dual-branch feature extraction network fused with channel priors, and a regression network. The image serialization module slices the image and converts it into an image sequence, which is then input into the dual-branch feature extraction network. The dual-branch feature extraction network first extracts the feature map after channel and space weighting with the channel prior module, and then downsamples the weighted feature map layer by layer in four stages, and uses the state space feature extraction module to perform dual-branch feature extraction, and finally performs upsampling regression through the regression network to output the forest biomass prediction result within the predicted geographical range.

[0037] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned high-resolution large-scale forest biomass remote sensing prediction method.

[0038] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned high-resolution large-scale forest biomass remote sensing prediction method.

[0039] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0040] The present invention integrates multi-source remote sensing data under the premise of geo-referencing, enriches the characteristic variables of forest biomass, and at the same time, the backbone network of the designed forest biomass prediction model extracts the channel and spatial weighted feature maps through the channel prior module, and uses the state space feature extraction module to perform double-branch feature extraction, so as to achieve more accurate forest biomass estimation and improve the generalization ability of the model. The specific effects are as follows:

[0041] 1. The fusion of multi-source remote sensing image data can fully alleviate the data missing existing in single data, and the added meteorological data and terrain data improve the model's understanding and reflection of the complexity of forest ecosystems, effectively alleviating the biomass estimation saturation problem that may exist in the previous single remote sensing image data when processing forest biomass estimation.

[0042] 2. A multi-source data geo-referencing method based on joint geometric features and regional features uses the block extraction strategy of the FAST operator features to obtain the initial matching point set through coarse matching, and then uses regional grayscale constraints to obtain accurate matching point pairs. Finally, an affine transformation model is used for registration. This geo-referencing method makes full use of the subtle connections between different data, realizes the deep fusion of multi-source data, and improves the accuracy of geo-referencing between data.

[0043] 3. The forest biomass prediction algorithm is based on the fusion of state-space model and multi-scale features. It identifies key features through the channel prior module, uses the basic structure of the state-space model, introduces hybrid convolution to capture multi-scale spatial features, and uses selective mechanism to compress information, thus achieving high-precision prediction of forest biomass. When new forest types or geographical environmental factors appear, the model can make predictions adaptively, which improves the generalization ability and maintenance efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0045] Figure 1 This is a flow chart of a high-resolution large-scale forest biomass remote sensing prediction method provided by an embodiment of the present invention;

[0046] Figure 2 is a multi-source data geo-referenced flow chart provided by an embodiment of the present invention;

[0047] Figure 3 This is a main network structure diagram of the forest biomass prediction model provided by an embodiment of the present invention;

[0048] Figure 4 is a structural diagram of a channel priori module provided by an embodiment of the present invention;

[0049] Figure 5 is a structural diagram of a dual-branch feature extraction module provided by an embodiment of the present invention;

[0050] Figure 6 is a scanning schematic diagram of an omnidirectional selective scanning module provided by an embodiment of the present invention;

[0051] Figure 7 is a regression network structure diagram provided by an embodiment of the present invention;

[0052] Figure 8 It is an overall block diagram of a high-resolution large-scale forest biomass remote sensing prediction method provided by an embodiment of the present invention;

[0053] Fig. 9 This is a block diagram of a high-resolution, large-scale forest biomass remote sensing prediction system provided by an embodiment of the present invention;

[0054] Fig.10 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0056] The embodiment of the present invention provides a high-resolution large-scale forest biomass remote sensing prediction method, which can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of a high-resolution large-scale forest biomass remote sensing prediction method is shown in FIG. The processing flow of the method may include the following steps:

[0057] S1. Acquire multi-source data within the geographic range to be predicted, wherein the multi-source data includes microwave radar remote sensing image data, optical remote sensing image data, meteorological data and terrain data;

[0058] The embodiment of the present invention downloads microwave radar remote sensing image data with a spatial resolution of 10m collected by the Sentinel-1 microwave radar of the European Space Agency and optical remote sensing image data with a spatial resolution of 10m in four bands collected by the Sentinel-2 high-resolution multispectral imager, and simultaneously downloads terrain data with a spatial resolution of 30m from the seamless data distribution system and meteorological data such as the annual average temperature and annual rainfall with a spatial resolution of 30m from the data geographic remote sensing ecological network.

[0059] S2. Geo-referencing the multi-source data by combining geometric features and regional features;

[0060] Alternatively, if Figure 2 As shown, the S2 specifically includes:

[0061] S21, using the optical remote sensing image data as a reference image and other multi-source data as images to be registered, performing geometric feature pre-matching to obtain an initial matching point set, including:

[0062] Each image is divided into 20×20 non-overlapping blocks, and the FAST value of each pixel in each non-overlapping block is calculated. The FAST value is defined as the sum of the absolute values ​​of the pixel differences between a total of 16 pixels on a circle with a radius of 3 and each pixel as the center and the center of the circle. The pixel with the largest FAST value is selected as the feature point in each non-overlapping block, and the feature points in 20×20×1 images are obtained;

[0063] Assume that the coordinates of a feature point in the reference image are , calculate the Euclidean distance between the feature point and all feature points in the image to be registered , select the feature point in the image to be registered that minimizes the Euclidean distance as the matching feature point, assuming the coordinates are ,but ;

[0064] Complete the matching of all feature points in the reference image to form the initial matching point set;

[0065] S22, performing regional grayscale constrained registration optimization on the initial matching point set, eliminating mismatched point pairs, and obtaining an accurate matching point set, including:

[0066] Based on the initial matching point set, the window grayscale similarity coefficient of the matching point pairs in the two images is calculated using the following formula: :

[0067]

[0068] In the above formula , Respectively represent the grayscale values ​​of the pixels in the 5×5 window centered on the matching point pair in the reference image and the image to be registered, , are the average values ​​of the pixels in the corresponding windows respectively;

[0069] When the grayscale similarity coefficient of the window If the maximum value of is greater than the preset threshold T, the matching point pair is retained, otherwise, the mismatched point pairs are removed to obtain an accurate matching point set;

[0070] S23. Correct the image using an affine transformation model to reduce significant geometric deformation and complete the registration.

[0071] S3, extracting characteristic variables from the multi-source data after georeferencing;

[0072] Optionally, a total of 21 characteristic variables are extracted in S3, including:

[0073] Four characteristic variables are extracted from the microwave radar remote sensing image data, namely, VV polarization backscattering coefficient, VH polarization backscattering coefficient, backscattering coefficient polarization ratio VV / VH, and standard depolarization rate (VV-VH) / (VH+VV);

[0074] Extracting spectral information and vegetation index from the optical remote sensing image data, specifically including eight characteristic variables, namely, four single-band reflectances, normalized difference vegetation index NDVI, soil adjusted vegetation index SAVI, simple ratio vegetation index SR and difference vegetation index DVI;

[0075] Based on the gray-level co-occurrence matrix, the texture features of the four single bands and the backscattering coefficient polarization ratio VV / VH are calculated, and energy values ​​are constructed to describe the texture features, and 5 energy values ​​are obtained as 5 feature variables. Each element in the gray-level co-occurrence matrix refers to the joint distribution of the gray levels of two pixels with a certain spatial position relationship. and direction angle When As the starting point, grayscale appears The probability of the whole image is arranged into a matrix to obtain the gray-level co-occurrence matrix , specifically including: moving any point in the image and a point away from it Constitute a point pair, and set the gray value of the point pair to be , assuming that the maximum gray level of the image is L, then and The combination has a total of For the entire image, count each The number of times the value appears, and then arrange it into a square matrix, and then The total number of occurrences is normalized to get the probability , the resulting matrix is ​​the gray-level co-occurrence matrix; the energy value It is a measure of the uniformity of the grayscale distribution and the coarseness of the texture of the image. If the element values ​​of the grayscale co-occurrence matrix are similar, the energy value is small, indicating a fine texture; when the image texture is coarse, the energy value is large, which is expressed by the following formula:

[0076] ;

[0077] In addition, two characteristic variables, elevation and slope, were extracted from the digital elevation model data of the Space Shuttle Radar Topography Mission, and two characteristic variables, annual mean temperature and annual precipitation, were selected, for a total of 21 characteristic variables.

[0078] In order to enable the model to capture characteristic information at the same resolution and achieve high-precision forest biomass prediction, the embodiment of the present invention uses ArcGIS software to unify the image spatial resolution of all characteristic variables to 30m.

[0079] S4. Input the extracted feature variable image into the backbone network of the forest biomass prediction model, such as Figure 3 As shown, the backbone network includes an image serialization module, a dual-branch feature extraction network integrating channel priors, and a regression network. The image serialization module slices the image (slices each of the 21 images of feature variables into 9 pieces) and converts it into an image sequence, which is then input into the dual-branch feature extraction network. The dual-branch feature extraction network first extracts the channel- and space-weighted feature map with the channel prior module, then downsamples the weighted feature map layer by layer in four stages, and uses the state-space feature extraction module to perform dual-branch feature extraction. Finally, upsampling regression is performed through the regression network to output the forest biomass prediction result within the predicted geographical range.

[0080] Alternatively, if Figure 4 As shown, the channel prior module first performs average pooling and maximum pooling on the input image sequence, and then generates a channel attention map using an activation function through a shared multilayer perceptron, and obtains a channel prior image sequence with channel attention by multiplying the input image sequence and the elements of the channel attention map;

[0081] Then, the channel prior image sequence is input into a multi-scale separable convolution module to generate a spatial attention map. The multi-scale separable convolution module sets three separable convolution branches for parallel processing to capture multi-scale spatial features. The three separable convolution branches are respectively composed of 1×5 convolution and 5×1 convolution, 1×7 convolution and 7×1 convolution, and 1×9 convolution and 9×1 convolution. The obtained features are mixed with the features previously obtained by 3×3 convolution through channels to obtain the spatial attention map. The spatial attention map is input into a 1×1 convolution for channel adjustment and then combined with the channel prior image sequence. Through element-by-element multiplication operations, the spatial attention weights and channel attention weights are mapped to the feature map to obtain the weighted feature map, thereby realizing the recognition of key features in the image.

[0082] Alternatively, if Figure 3 As shown, the state space feature extraction module is composed of layer normalization, a dual-branch feature extraction module, a multi-layer perceptron, layer normalization, a downsampling module and a residual connection. The dual-branch feature extraction module is used to further accurately extract features, and the residual connection and layer normalization are used to alleviate the gradient vanishing problem.

[0083] The dual-branch feature extraction module, such as Figure 5 As shown, it includes an omnidirectional selective state space branch and a hybrid convolution branch. The input is first divided into two equal-sized input branches 1 and 2, and then input into the two branches respectively. The omnidirectional selective state space branch is used to capture multi-scanning direction features, and the hybrid convolution branch is used to extract contextual information of different scales. The outputs of the two branches are merged along the channel dimension of the feature map, and the information interaction between the two sub-input channels is promoted through channel shuffling.

[0084] Alternatively, if Figure 5 As shown on the right side of the figure, after the input branch 1 is input into the omnidirectional selective state space branch, it is input into the omnidirectional selective scanning module after layer normalization, linear layer, mixed convolution, and activation function, as shown in FIG. Figure 6As shown, the input is flattened into 8 groups of sequences along the horizontal, vertical, oblique, anti-oblique and opposite directions, and the input sequence is modeled using a selective state space model to achieve information compression. The scanning results in all directions are accumulated together to form an output, which integrates the features in 8 directions (the existing four-way selective scanning does not fully utilize the two-dimensional structure of the image, thereby limiting its ability to analyze from multiple angles), so that the model can capture and model the scale space features of the image in all directions and obtain image features with rich spatial position information. The output is then multiplied with the result obtained by the linear layer and the activation function of another branch after layer normalization, and finally the output of the omnidirectional selective state space branch is obtained after layer normalization, as shown in FIG. Figure 5 As shown on the right side of the figure, the selective state space model introduces a selective mechanism, which learns the relationship between input features and biomass during continuous training by selecting input sequence information, and identifies and captures effective information.

[0085] Alternatively, if Figure 5 As shown on the left side of the figure, after the input branch 2 is input into the hybrid convolution branch, it passes through four cascaded hybrid convolutions, each of which includes dilated convolution and multi-scale convolution. The dilation rates of the dilated convolution are 1, 2, 3, and 1, respectively, to avoid the grid effect caused by discontinuous data. The dilated convolutions with different dilation rates are used to improve the receptive field and capture a wider range of contextual information. The multi-scale convolution uses parallel 1×1 convolution, 3×3 convolution, and 5×5 convolution to obtain multi-scale information of feature variables, and through aggregation, it fully learns the geographic spatial effects of forests, learns their spatial correlations at different spatial sizes, and improves the understanding of the complexity of forest ecosystems.

[0086] The regression network, by upsampling the extracted features and refining them using convolutional layers, consists of four convolution-based decoding blocks, and finally reduces the number of channels to a single channel through a 1×1 convolution, and finally maps the features to biomass data through an activation function, such as Figure 7 shown.

[0087] The training process of the backbone network of the forest biomass prediction model of the embodiment of the present invention is as follows:

[0088] 1) Collection of forest biomass data in the sample area;

[0089] The latitude and longitude of the sample area, dominant tree species, diameter at breast height, height, number of trees per hectare and other data were obtained through national forest inventory data and field measurement data. For different tree species, the forest biomass data (Aboveground biomass, AGB) was obtained using the following allometric growth equation:

[0090]

[0091] in is the breast diameter, is the tree height, a and b are tree species parameters;

[0092] The forest biomass data of the embodiment of the present invention will be converted into image data according to the longitude and latitude and the corresponding forest biomass values ​​and input into the subsequent model.

[0093] 2) Multi-source data acquisition in the sample area - multi-source data georeferencing - feature variable extraction;

[0094] The specific process is similar to the process of acquiring and processing multi-source data within the aforementioned geographical scope to be predicted, and will not be repeated here.

[0095] 3) Divide the image pairs of biomass and characteristic variables obtained in the previous two steps into training set and validation set in a ratio of 4:1, such as Figure 8 As shown, a complete data set is formed. The algorithm will be trained on the training set and the model effect will be verified on the validation set. The model with the best performance on the validation set will be selected as the final model. In the validation process, the determination coefficient is used and Root Mean Square Error (RMSE) as indicators to evaluate model performance, and the coefficient of determination The model's ability to explain biomass trends was measured, and the root mean square error reflected the prediction accuracy of the regression model.

[0096] like Fig. 9 As shown, the embodiment of the present invention also provides a high-resolution large-scale forest biomass remote sensing prediction system, the system comprising:

[0097] An acquisition module 910 is used to acquire multi-source data within the geographic range to be predicted, wherein the multi-source data includes microwave radar remote sensing image data, optical remote sensing image data, meteorological data and terrain data;

[0098] A geo-referencing module 920, for geo-referencing the multi-source data by combining geometric features and regional features;

[0099] An extraction module 930 is used to extract characteristic variables from the multi-source data after geo-referenced;

[0100] Prediction module 940 is used to input the image of the extracted feature variables into the backbone network of the forest biomass prediction model, the backbone network includes an image serialization module, a dual-branch feature extraction network fused with channel priors, and a regression network. The image serialization module slices the image and converts it into an image sequence, which is then input into the dual-branch feature extraction network. The dual-branch feature extraction network first extracts the feature map after channel and space weighting with the channel prior module, and then downsamples the weighted feature map layer by layer in four stages, and uses the state space feature extraction module to perform dual-branch feature extraction, and finally performs upsampling regression through the regression network to output the forest biomass prediction result within the predicted geographical range.

[0101] A high-resolution, large-scale forest biomass remote sensing prediction system provided in an embodiment of the present invention has a functional structure corresponding to a high-resolution, large-scale forest biomass remote sensing prediction method provided in an embodiment of the present invention, which will not be repeated here.

[0102] Fig.10 It is a structural schematic diagram of an electronic device 1000 provided in an embodiment of the present invention. The electronic device 1000 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 1001 and one or more memories 1002, wherein the memory 1002 stores at least one instruction, and the at least one instruction is loaded and executed by the processor 1001 to implement the steps of the above-mentioned high-resolution large-scale forest biomass remote sensing prediction method.

[0103] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, which can be executed by a processor in a terminal to complete the above-mentioned high-resolution large-scale forest biomass remote sensing prediction method. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0104] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0105] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A high-resolution large-scale forest biomass remote sensing prediction method, characterized in that: The method comprises: S1. Acquire multi-source data within the geographic range to be predicted, wherein the multi-source data includes microwave radar remote sensing image data, optical remote sensing image data, meteorological data and terrain data; S2. Geo-referencing the multi-source data by combining geometric features and regional features; S3, extracting characteristic variables from the multi-source data after georeferencing; S4, inputting the image of the extracted characteristic variable into the backbone network of the forest biomass prediction model, the backbone network includes an image serialization module, a dual-branch feature extraction network fused with channel priors, and a regression network, the image serialization module slices the image and converts it into an image sequence, which is then input into the dual-branch feature extraction network, the dual-branch feature extraction network first extracts a feature map after channel and space weighting with a channel prior module, then performs dual-branch feature extraction and downsampling on the weighted feature map in four serially connected state space feature extraction modules, and finally performs upsampling regression through the regression network, and outputs the forest biomass prediction result within the predicted geographical range; The state space feature extraction module is composed of layer normalization, a dual-branch feature extraction module, a multi-layer perceptron, layer normalization, a downsampling module and a residual connection. The dual-branch feature extraction module is used to further accurately extract features, and the residual connection and layer normalization are used to alleviate the gradient vanishing problem. The dual-branch feature extraction module includes an omnidirectional selective state space branch and a hybrid convolution branch. The input is first divided into two equal-sized input branches 1 and 2, and then input into the two branches respectively. The omnidirectional selective state space branch is used to capture multi-scanning direction features, and the hybrid convolution branch is used to extract contextual information of different scales. The outputs of the two branches are merged along the channel dimension of the feature map, and channel shuffling is used to promote information interaction between the two sub-input channels.

2. The method according to claim 1, characterized in that The S2 specifically includes: S21, using the optical remote sensing image data as a reference image and other multi-source data as images to be registered, performing geometric feature pre-matching to obtain an initial matching point set, including: Each image is divided into 20×20 non-overlapping blocks, and the FAST value of each pixel in each non-overlapping block is calculated. The FAST value is defined as the sum of the absolute values ​​of the pixel differences between a total of 16 pixels on a circle with a radius of 3 and each pixel as the center and the center of the circle. The pixel with the largest FAST value is selected as the feature point in each non-overlapping block, and the feature points in 20×20×1 images are obtained; Assume that the coordinates of a feature point in the reference image are , calculate the Euclidean distance between the feature point and all feature points in the image to be registered , select the feature point in the image to be registered that minimizes the Euclidean distance as the matching feature point, assuming the coordinates are ,but ; Complete the matching of all feature points in the reference image to form the initial matching point set; S22, performing regional grayscale constrained registration optimization on the initial matching point set, eliminating mismatched point pairs, and obtaining an accurate matching point set, including: Based on the initial matching point set, the window grayscale similarity coefficient of the matching point pairs in the two images is calculated using the following formula: : ; In the above formula , Respectively represent the grayscale values ​​of the pixels in the 5×5 window centered on the matching point pair in the reference image and the image to be registered, , are the average grayscale values ​​of the corresponding windows respectively; When the grayscale similarity coefficient of the window If the maximum value of is greater than the preset threshold T, the matching point pair is retained, otherwise, the mismatched point pairs are removed to obtain an accurate matching point set; S23. Correct the image using an affine transformation model to reduce significant geometric deformation and complete the registration.

3. The method according to claim 1, characterized in that: A total of 21 characteristic variables are extracted from S3, including: Four characteristic variables are extracted from the microwave radar remote sensing image data, namely, VV polarization backscattering coefficient, VH polarization backscattering coefficient, backscattering coefficient polarization ratio VV / VH, and standard depolarization rate (VV-VH) / (VH+VV); Extracting spectral information and vegetation index from the optical remote sensing image data, specifically including eight characteristic variables, namely, four single-band reflectances, normalized difference vegetation index NDVI, soil adjusted vegetation index SAVI, simple ratio vegetation index SR and difference vegetation index DVI; Based on the gray-level co-occurrence matrix, the texture features of the four single bands and the backscattering coefficient polarization ratio VV / VH are calculated, and energy values ​​are constructed to describe the texture features, and 5 energy values ​​are obtained as 5 feature variables. Each element in the gray-level co-occurrence matrix refers to the joint distribution of the gray levels of two pixels with a certain spatial position relationship. and direction angle When As the starting point, grayscale appears The probability of the whole image is arranged into a matrix to obtain the gray-level co-occurrence matrix , specifically including: moving any point in the image and a point away from it Constitute a point pair, and set the gray value of the point pair to be , assuming that the maximum gray level of the image is L, then and The combination of For the entire image, count each The number of times the value appears, and then arrange it into a square matrix, and then The total number of occurrences is normalized to get the probability , the resulting matrix is ​​the gray-level co-occurrence matrix; the energy value It is a measure of the uniformity of the grayscale distribution and the coarseness of the texture of the image. If the element values ​​of the grayscale co-occurrence matrix are similar, the energy value is small, indicating a fine texture; when the image texture is coarse, the energy value is large, which is expressed by the following formula: ; In addition, two characteristic variables, elevation and slope, were extracted from the digital elevation model data of the Space Shuttle Radar Topography Mission, and two characteristic variables, annual mean temperature and annual precipitation, were selected, for a total of 21 characteristic variables.

4. The method according to claim 1, characterized in that: The channel prior module first performs average pooling and maximum pooling on the input image sequence, then generates a channel attention map using an activation function through a shared multilayer perceptron, and obtains a channel prior image sequence with channel attention by multiplying the input image sequence and the elements of the channel attention map; Then, the channel prior image sequence is input into a multi-scale separable convolution module to generate a spatial attention map. The multi-scale separable convolution module sets three separable convolution branches for parallel processing to capture multi-scale spatial features. The three separable convolution branches are respectively composed of 1×5 convolution and 5×1 convolution, 1×7 convolution and 7×1 convolution, and 1×9 convolution and 9×1 convolution. The obtained features are mixed with the features previously obtained by 3×3 convolution through channels to obtain the spatial attention map. The spatial attention map is input into a 1×1 convolution for channel adjustment and then combined with the channel prior image sequence. Through element-by-element multiplication operations, the spatial attention weights and channel attention weights are mapped to the feature map to obtain the weighted feature map, thereby realizing the recognition of key features in the image.

5. The method according to claim 1, characterized in that After the input branch 1 inputs the omnidirectional selective state space branch, it is input into the omnidirectional selective scanning module after layer normalization, linear layer, mixed convolution, and activation function. The input is flattened into 8 groups of sequences along the horizontal, vertical, oblique, reverse oblique and opposite directions, and the input sequence is modeled using a selective state space model to achieve information compression. The scanning results in all directions are accumulated together to form an output. The output integrates the features in 8 directions, so that the model can capture and model the scale space features of the image in an all-round way, and obtain image features with rich spatial position information. The output is then multiplied with the result obtained by the other branch after layer normalization, the linear layer and the activation function, and finally the output of the omnidirectional selective state space branch is obtained after layer normalization. The selective state space model introduces a selective mechanism. By selecting the input sequence information, the relationship between the input features and the biomass is learned in continuous training, and effective information is identified and captured.

6. The method according to claim 1, characterized in that After the input branch 2 is input into the hybrid convolution branch, it passes through four cascaded hybrid convolutions, each of which includes an expanded convolution and a multi-scale convolution. The expansion rates of the expanded convolution are 1, 2, 3, and 1, respectively, to avoid the grid effect caused by discontinuous data. The expanded convolutions with different expansion rates are used to improve the receptive field and capture a wider range of contextual information. The multi-scale convolution uses parallel 1×1 convolution, 3×3 convolution, and 5×5 convolution to obtain multi-scale information of feature variables, and through aggregation, fully learn the geographic spatial effects of forests, learn their spatial correlations at different spatial sizes, and improve the understanding of the complexity of forest ecosystems.

7. A high-resolution, large-scale forest biomass remote sensing prediction system, characterized in that: The system comprises: An acquisition module, used to acquire multi-source data within the geographic range to be predicted, wherein the multi-source data includes microwave radar remote sensing image data, optical remote sensing image data, meteorological data and terrain data; A geo-referencing module, used for combining geometric features and regional features to geo-referencing the multi-source data; The extraction module is used to extract characteristic variables from multi-source data after georeferencing; A prediction module is used to input the image of the extracted characteristic variables into the backbone network of the forest biomass prediction model, wherein the backbone network includes an image serialization module, a dual-branch feature extraction network fused with channel priors, and a regression network. The image serialization module slices the image and converts it into an image sequence, which is then input into the dual-branch feature extraction network. The dual-branch feature extraction network first extracts a feature map after channel and space weighting with a channel prior module, then performs dual-branch feature extraction and downsampling on the weighted feature map in four serially connected state space feature extraction modules, and finally performs upsampling regression through the regression network to output the forest biomass prediction result within the predicted geographical range; The state space feature extraction module is composed of layer normalization, a dual-branch feature extraction module, a multi-layer perceptron, layer normalization, a downsampling module and a residual connection. The dual-branch feature extraction module is used to further accurately extract features, and the residual connection and layer normalization are used to alleviate the gradient vanishing problem. The dual-branch feature extraction module includes an omnidirectional selective state space branch and a hybrid convolution branch. The input is first divided into two equal-sized input branches 1 and 2, and then input into the two branches respectively. The omnidirectional selective state space branch is used to capture multi-scanning direction features, and the hybrid convolution branch is used to extract contextual information of different scales. The outputs of the two branches are merged along the channel dimension of the feature map, and channel shuffling is used to promote information interaction between the two sub-input channels.

8. An electronic device, comprising a processor and a memory, wherein at least one instruction is stored in the memory, wherein: The at least one instruction is loaded and executed by the processor to implement the high-resolution, large-scale forest biomass remote sensing prediction method as described in any one of claims 1-6.

9. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, characterized in that: The at least one instruction is loaded and executed by the processor to implement the high-resolution large-scale forest biomass remote sensing prediction method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Transform and dense feature fusion-based remote sensing image change detection method and system

    CN115690002A

  • Forest aboveground biomass estimation method, device and equipment based on multi-source remote sensing data

    CN118918477A