Method and system for fusing satellite-borne laser radar data based on particle swarm optimization random forest model

By fusing ICESat-2 and GEDI data using a particle swarm optimization random forest model, the problems of insufficient accuracy and coverage of spaceborne lidar data in forest canopy height inversion were solved, achieving higher accuracy and more stable forest structure monitoring.

CN121388480APending Publication Date: 2026-01-23LANZHOU JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511636842.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing technologies for forest canopy height inversion using ICESat-2 and GEDI spaceborne lidar data have failed to fully leverage the complementary advantages of both, resulting in limitations in inversion accuracy. ICESat-2 is susceptible to interference from forest cover and topographic slope, while GEDI has insufficient spatial coverage and tends to underestimate data. Furthermore, most fusion research methods have failed to effectively coordinate data differences, affecting the accuracy of forest structure monitoring.

Method used

A method based on particle swarm optimization random forest model is adopted. By spatially matching ICESat-2 and GEDI data, effective overlapping light spots are screened out. Key photon features are screened by combining random forest-recursive feature elimination method. Finally, the hyperparameters of random forest model are optimized by particle swarm optimization algorithm to establish the optimal fusion model.

Benefits of technology

It improves the accuracy and stability of forest canopy height inversion, enhances the model's adaptability to complex terrain and land cover, reduces model complexity, improves overall operating efficiency, and expands the sample coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121388480A_ABST
    Figure CN121388480A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of remote sensing data fusion, and discloses a method and system for fusing satellite-borne laser radar data based on a particle swarm optimization random forest model, and the method comprises the steps: obtaining multi-source satellite-borne laser radar data; performing space matching; carrying out feature screening on all photon features by adopting a random forest-recursive feature elimination method; optimizing the random forest model by adopting a particle swarm optimization algorithm to obtain an optimal random forest model; and performing forest canopy height inversion on the target area by adopting the optimal random forest model and taking the first satellite-borne laser radar data of the target area as input. According to the method, the random forest regression model based on particle swarm optimization is constructed, a new path is provided for efficient processing of satellite-borne laser radar data and ground feature parameter prediction, and the method has good popularization and application values and can be widely applied to remote sensing analysis scenes such as forestry investigation and ecological monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing technology and forest resource monitoring technology, specifically relating to a method and system for fusing spaceborne lidar data based on a particle swarm optimization random forest model. Background Technology

[0002] With the continuous development of remote sensing technology, lidar, as an active remote sensing method, has been widely used in forest canopy height inversion tasks due to its ability to penetrate the forest canopy and acquire information on the vertical structure of the forest understory. Especially in large-scale forest height mapping, spaceborne lidar systems have demonstrated unique advantages. Among them, the ICESat-2 (Ice, Cloud and Land Elevation Satellite-2), equipped with the advanced altimetry system ATLAS, can acquire higher density and more detailed photon point cloud data; while the full-waveform lidar sensor equipped in the Global Ecosystem Dynamics Investigation (GEDI) not only has strong resistance to solar noise, but its laser point density and spot size have also been significantly improved compared to the previous generation GLAS system. Therefore, spaceborne lidar is increasingly becoming an important data source for global forest structure monitoring.

[0003] In recent years, some studies have applied deep learning methods to the task of retrieving forest canopy height from spaceborne lidar data, and in conjunction with lidar and optical imagery, have achieved high-resolution canopy height mapping. However, current research still has the following shortcomings: (1) Although existing studies have used convolutional neural networks and ensemble probabilistic deep learning models for modeling, the complementary advantages of the two spaceborne lidar data, ICESat-2 and GEDI, have not been fully utilized when retrieving forest canopy height, resulting in limited accuracy in canopy height estimation. (2) Multiple experiments have shown that the canopy height obtained by ICESat-2 is easily affected by forest cover and topographic slope, and its inversion accuracy is generally lower than that of GEDI, with larger root mean square error and mean difference. (3) Although GEDI has high inversion accuracy, its observation range only covers the area between 51.6° north and south latitude, resulting in insufficient spatial coverage and an underestimation of canopy height. This may lead to the accumulation of biases during mapping, affecting the accuracy of forest structure monitoring. (4) Most current fusion research methods fail to fully consider the differences between ICESat-2 and GEDI in terms of observation mechanism, spatial coverage, and photon sampling method, and lack effective data coordination mechanism and unified modeling framework, resulting in limited fusion effect. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a method and system for fusing spaceborne lidar data based on a particle swarm optimization random forest model, which can effectively solve the aforementioned problems.

[0005] The technical solution adopted in this invention is as follows:

[0006] This invention provides a method for fusing spaceborne lidar data based on a particle swarm optimization random forest model, comprising the following steps:

[0007] Step S1: Acquire multi-source spaceborne lidar data, including first spaceborne lidar data and second spaceborne lidar data; wherein, the first spaceborne lidar data is spaceborne lidar data in the form of high-density photons; and the second spaceborne lidar data is full-waveform lidar data.

[0008] Step S2: Spatial matching is performed on the first spaceborne lidar data and the second spaceborne lidar data to extract several sets of effective overlapping light spots of spaceborne lidar in the spatially overlapping region; the effective overlapping light spots of spaceborne lidar simultaneously contain the first spaceborne lidar data and the second spaceborne lidar data in the spatially overlapping region.

[0009] Step S3: Take each of the effective overlapping light spots of the spaceborne lidar as a raw sample to construct a raw sample set; each raw sample in the raw sample set has multiple photon features and forest canopy height features that affect the height of the forest canopy.

[0010] Step S4: Based on the original sample set, the random forest-recursive feature elimination method is used to filter all photon features and optimize them to obtain several target photon features that have a significant impact on forest canopy height, forming a target photon feature set;

[0011] Step S5: Based on the target photon feature set and the forest canopy height feature, process the original sample set, retaining only the target photon feature and the forest canopy height feature, to construct a sample set;

[0012] Step S6: Based on the sample set, the random forest model is optimized using the particle swarm optimization algorithm to obtain the optimal random forest model;

[0013] Step S7: Using the optimal random forest model, and taking the first satellite-borne lidar data of the target area as input, the forest canopy height of the target area is inverted.

[0014] Furthermore, the first spaceborne lidar data is ICESat-2 spaceborne lidar data; the second spaceborne lidar data is GEDI spaceborne lidar data.

[0015] Furthermore, step S2 specifically involves:

[0016] Step S21: Using the footprint point S1 of the first satellite-borne lidar data and the footprint point S2 of the second satellite-borne lidar data as units, for each footprint point, low-quality footprint points are removed according to the echo quality indicator of the footprint point to achieve the screening of footprint points.

[0017] Step S22: For each of the filtered footprint points, for each footprint point S1 in the first satellite-borne lidar data, a buffer zone with a radius of R is set as the search radius, centered on footprint point S1, to generate a circular coverage area D1 for each footprint point S1.

[0018] For each footprint point S2 of the second satellite-borne lidar data, a buffer zone with radius R is set as the search radius, centered on footprint point S2, to generate a circular coverage area D2 for each footprint point S2.

[0019] Spatial matching is performed on circular coverage areas D1 and D2, and the overlapping area between the two is selected, which is the effective overlapping light spot of the spaceborne lidar.

[0020] Furthermore, step S3 specifically includes:

[0021] For each of the said spaceborne lidar effective overlapping light spots, there are multiple first spaceborne lidar data and multiple second spaceborne lidar data;

[0022] The average of the original values ​​of each photon feature of all the first spaceborne lidar data within the effective overlapping spot of the spaceborne lidar is taken as the feature value of the corresponding photon feature of the original sample.

[0023] The average GEDI waveform of all second-spaceborne lidar data within the effective overlapping spot of the spaceborne lidar is taken, and then 98 percent of the height of the average GEDI waveform is used as the forest canopy height feature of the original sample.

[0024] Furthermore, step S4 specifically involves:

[0025] Step S41, for each original sample in the original sample set, it is represented as: ;in, Indicates the first One original sample; Indicates the number of photon features; These represent the eigenvalues ​​of the first photon feature, the second photon feature, ..., the eigenvalues ​​of the third photon feature, respectively. Eigenvalues ​​of each photon feature; The feature values ​​representing the characteristics of forest canopy height serve as sample labels;

[0026] Step S42: Construct a first random forest model. The first random forest model has basic parameters, including: the number of decision trees, the minimum number of samples for node splitting, the minimum number of samples for leaf nodes, the mean squared error as the splitting criterion, and all photon features are considered in each split. The training strategy adopts sampling with replacement.

[0027] Step S43: Input the original sample set into the first random forest model. The first random forest model evaluates the importance of each photon feature and recursively eliminates them to obtain the target photon feature set. , , The number of target photon features is determined in the following way:

[0028] Step S431, each original sample has The original sample set of each photon feature is input into the first random forest model. The first random forest model obtains the importance of each photon feature in each decision tree based on the importance of each photon feature in the prediction of forest canopy height features. Then, the importance of each photon feature in each decision tree is averaged to obtain the importance score of each photon feature. Then, the importance score is calculated... The importance of each photon feature is ranked to obtain a photon feature sequence; the first random forest model simultaneously outputs the RMSE based on the original sample set;

[0029] Step S432: In the photon feature sequence, delete the M1 photon features with the lowest importance to obtain the optimized photon feature sequence; based on the optimized photon feature sequence, update the original sample set, retaining only the optimized photon features in each original sample, thereby obtaining the optimized sample set;

[0030] The optimized sample set is input into the first random forest model, and the first random forest model outputs the importance sequence of each optimized photon feature and the RMSE.

[0031] Step S433: Determine whether the output RMSE is stable at this time; if it is stable, proceed to step S424; if it is unstable, repeat step S422.

[0032] Step S434: The optimized photon features are the target photon features, and the importance scores of each target photon feature are output.

[0033] Furthermore, step S6 specifically involves:

[0034] Step S61: Divide the sample set into a feature training set and a feature test set using stratified sampling techniques;

[0035] Step S62: Select the second random forest model and train the second random forest model using the feature training set to obtain the trained second random forest model;

[0036] Step S63: Optimize the trained second random forest model using the particle swarm optimization algorithm to obtain the optimal random forest model;

[0037] Step S64: Train the optimal random forest model using the feature training set, and evaluate the performance of the trained optimal random forest model using the feature test set.

[0038] Furthermore, step S63 specifically includes:

[0039] Step S631: Set the parameters of the particle swarm optimization algorithm, initialize the position and velocity of the particles, and calculate the fitness value of each particle using the fitness function, where each particle represents the parameter configuration of the random forest model.

[0040] Step S632: Based on the fitness value of each particle and the historical best fitness value, obtain the individual best position and the global best position;

[0041] Step S633: Based on the individual best position and the global best position, update the position and velocity of each particle, and calculate the fitness value of each particle according to the updated particle position.

[0042] Step S634: Repeat steps S632 to S633 until the preset number of iterations is reached to obtain the optimal parameter configuration;

[0043] Step S635: Obtain the optimal random forest model based on the optimal parameter configuration.

[0044] Furthermore, step S63 is more specifically as follows:

[0045] Step 1: Initialize the particle swarm. Each particle represents a set of hyperparameters for the random forest model, including: maximum depth, maximum number of leaf nodes, minimum number of samples, and number of decision trees.

[0046] Step 2: Set the particle swarm optimization algorithm parameters, including the number of particles N and the maximum number of iterations. Maximum value of inertia weight Minimum value of inertia weight Individual and group learning factors , And randomly initialize the velocity of each particle. and location ;

[0047] Step 3: Construct the fitness function. With the goal of minimizing prediction error and model complexity, the fitness function is defined as follows:

[0048] )

[0049] in: Represents the fitness function. The root mean square error of the model on the feature test set; The number of decision trees in the random forest model. and These are weight parameters, used to balance the impact of model prediction accuracy and complexity;

[0050] Step 4: Calculate the fitness value of the particle. If the current fitness is better than its historical best value, update the individual's best position; at the same time, update the global best position, which is the solution with the best fitness among all particles.

[0051] Specifically, the velocity, position, and inertial weight of each particle are iteratively adjusted according to the following update formula:

[0052] 1) Speed ​​of updates:

[0053]

[0054] 2) Location update:

[0055]

[0056] 3) Inertia weight update:

[0057]

[0058] in: and Particles In the In the nth iteration 3D velocity and position vector; and Particles In the In the nth iteration 3D velocity and position vector; Corresponding to the random forest model, the first One hyperparameter to be determined; Inertial weights; and For learning factors; and It is a random number within the interval [0,1]. For particles In the In the nth iteration Dimension's historical best position; For the group in the first In the nth iteration Dimension's historical best position; This represents the maximum number of iterations. This represents the current iteration number;

[0059] Step 5: Under the new hyperparameter combination, use the random forest model for training and validation, and re-evaluate the fitness value of each particle;

[0060] Step 6: Determine whether the maximum number of iterations or fitness convergence condition has been reached. If the condition is met, proceed to Step 7; otherwise, return to Step 4 for the next round of iteration.

[0061] Step 7: Output the optimal combination of hyperparameters, train the final random forest model, and save the model inertia weights and hyperparameter configurations for subsequent predictions.

[0062] Furthermore, when evaluating the performance of the optimal random forest model after training using a feature test set, the following correlation coefficient and error metric are used for evaluation: coefficient of determination. The metrics are: Root Mean Square Error (RMSE) and Mean Absolute Error (MAE); where the coefficient of determination is closer to 1, indicating a better fit of the random forest model; and smaller RMSE and MAE, indicating higher prediction accuracy. The calculation formulas are:

[0063]

[0064]

[0065]

[0066] in: These are the characteristic values ​​of forest canopy height, derived from GEDI observation data; These are the feature values ​​for the predicted forest canopy height characteristics; For n The mean; This represents the number of samples.

[0067] This invention also provides a system for fusing spaceborne lidar data based on a particle swarm optimization random forest model, comprising:

[0068] The acquisition module is used to acquire multi-source spaceborne lidar data, including first spaceborne lidar data and second spaceborne lidar data; wherein, the first spaceborne lidar data is spaceborne lidar data in the form of high-density photons; and the second spaceborne lidar data is full-waveform lidar data.

[0069] The spatial matching module is used to perform spatial matching on the first spaceborne lidar data and the second spaceborne lidar data to extract several sets of effective overlapping light spots of the spaceborne lidar in the spatially overlapping region; the effective overlapping light spots of the spaceborne lidar simultaneously contain the first spaceborne lidar data and the second spaceborne lidar data in the spatially overlapping region.

[0070] The original sample set construction module is used to construct an original sample set by taking each effective overlapping spot of the spaceborne lidar as an original sample; each original sample in the original sample set has multiple photon features and forest canopy height features that affect the height of the forest canopy.

[0071] The Random Forest-Recursive Feature Elimination Module is used to perform feature filtering on all photon features based on the original sample set using the Random Forest-Recursive Feature Elimination Method, and optimize to obtain several target photon features that have a significant impact on forest canopy height, forming a target photon feature set;

[0072] The sample set construction module is used to process the original sample set according to the target photon feature set and the forest canopy height feature, retaining only the target photon feature and the forest canopy height feature, and constructing a sample set;

[0073] The optimization module is used to optimize the random forest model based on the sample set using the particle swarm optimization algorithm to obtain the optimal random forest model.

[0074] The regression prediction module is used to perform forest canopy height inversion in the target area by using the optimal random forest model and taking the first satellite-borne lidar data of the target area as input.

[0075] The present invention provides a method and system for fusing spaceborne lidar data based on a particle swarm optimization random forest model, which has the following advantages:

[0076] This invention improves the quality and representativeness of spaceborne lidar data by refining feature extraction and selection methods, enabling the random forest model to extract key information more effectively and enhancing its adaptability to complex terrain and features. By introducing particle swarm optimization (PSO) to optimize the hyperparameters of the random forest model, the prediction accuracy and stability of the model are significantly improved, while model complexity is reduced and overall operating efficiency is increased. This invention constructs a random forest regression model based on PSO optimization, providing a new path for efficient processing of spaceborne lidar data and prediction of ground feature parameters. It has good application value and can be widely used in remote sensing analysis scenarios such as forestry surveys and ecological monitoring. Attached Figure Description

[0077] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0078] Figure 1 This is a flowchart of a method for fusing spaceborne lidar data based on a particle swarm optimization random forest model according to an embodiment of the present invention;

[0079] Figure 2 This is a process diagram of a random forest-recursive feature elimination method based on a particle swarm optimization random forest model to fuse spaceborne lidar data according to an embodiment of the present invention.

[0080] Figure 3 This is an algorithm flowchart of a method for fusing spaceborne lidar data based on a particle swarm optimization random forest model according to an embodiment of the present invention;

[0081] Figure 4 This is the verification result of a fusion model based on a particle swarm optimization random forest model for fusing spaceborne lidar data in Qilian Mountain National Park, according to an embodiment of the present invention. Detailed Implementation

[0082] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0083] This invention proposes a method for fusing spaceborne lidar data based on a particle swarm optimization (PSO) random forest model. By spatially matching ICESat-2 (high-density photon) and GEDI (high-precision full-waveform lidar) data, the sample coverage is expanded. A random forest-recursive feature elimination method is used to dynamically filter photon features, reducing redundant feature interference. Furthermore, a particle swarm optimization algorithm is employed to dynamically adjust the hyperparameters of the random forest (maximum depth, maximum number of leaf nodes, minimum number of samples, and number of decision trees) to establish a PSO-optimized random forest spaceborne lidar data fusion model, which is then applied to the target area. This method effectively achieves complementary fusion of multi-source lidar data, providing a technical path to improve the accuracy and reliability of forest structure parameter inversion.

[0084] Please see Figures 1 to 4As shown, this invention provides a method for fusing spaceborne lidar data based on a particle swarm optimization random forest model, comprising the following steps:

[0085] Step S1: Acquire multi-source spaceborne lidar data, including first spaceborne lidar data and second spaceborne lidar data; wherein, the first spaceborne lidar data is spaceborne lidar data in the form of high-density photons; and the second spaceborne lidar data is full-waveform lidar data.

[0086] As one embodiment, the first spaceborne lidar data is ICESat-2 spaceborne lidar data; the second spaceborne lidar data is GEDI spaceborne lidar data.

[0087] Step S2: Spatial matching is performed on the first spaceborne lidar data and the second spaceborne lidar data to extract several sets of effective overlapping light spots of spaceborne lidar in the spatially overlapping region; the effective overlapping light spots of spaceborne lidar simultaneously contain the first spaceborne lidar data and the second spaceborne lidar data in the spatially overlapping region.

[0088] Step S2 is as follows:

[0089] Step S21: Using the footprint points S1 of the first satellite-borne lidar data and S2 of the second satellite-borne lidar data as units, for each footprint point, low-quality footprint points are removed according to the echo quality indicator of the footprint point to achieve the screening of footprint points; of course, in practical applications, the terrain height or the terrain slope of each segment can also be used as the screening feature to screen each footprint point.

[0090] Step S22: For each of the filtered footprint points, for each footprint point S1 in the first satellite-borne lidar data, a buffer zone with a radius of R is set as the search radius, centered on footprint point S1, to generate a circular coverage area D1 for each footprint point S1.

[0091] For each footprint point S2 of the second satellite-borne lidar data, a buffer zone with radius R is set as the search radius, centered on footprint point S2, to generate a circular coverage area D2 for each footprint point S2.

[0092] Spatial matching is performed on circular coverage areas D1 and D2, and the overlapping area between the two is selected, which is the effective overlapping light spot of the spaceborne lidar.

[0093] Step S3: Take each of the effective overlapping light spots of the spaceborne lidar as a raw sample to construct a raw sample set; each raw sample in the raw sample set has multiple photon features and forest canopy height features that affect the height of the forest canopy; wherein, the photon features include, but are not limited to, the percentage of canopy coverage, the slope of each segment of terrain, and the percentile height.

[0094] Step S3 is as follows:

[0095] For each of the said spaceborne lidar effective overlapping light spots, there are multiple first spaceborne lidar data and multiple second spaceborne lidar data;

[0096] The average of the original values ​​of each photon feature of all the first spaceborne lidar data within the effective overlapping spot of the spaceborne lidar is taken as the feature value of the corresponding photon feature of the original sample.

[0097] The average GEDI waveform of all second-spaceborne lidar data within the effective overlapping spot of the spaceborne lidar is taken, and then 98 percent of the height of the average GEDI waveform is used as the forest canopy height feature of the original sample.

[0098] Step S4: Based on the original sample set, the random forest-recursive feature elimination method is used to filter all photon features and optimize them to obtain several target photon features that have a significant impact on forest canopy height, forming a target photon feature set;

[0099] This step specifically involves using the random forest-recursive feature elimination method to evaluate the importance of all photon features and recursively eliminate them, thereby obtaining the target photon feature set.

[0100] Step S4 is as follows:

[0101] Step S41, for each original sample in the original sample set, it is represented as: ;in, Indicates the first One original sample; Indicates the number of photon features; These represent the eigenvalues ​​of the first photon feature, the second photon feature, ..., the eigenvalues ​​of the third photon feature, respectively. Eigenvalues ​​of each photon feature; The feature values ​​representing the characteristics of forest canopy height serve as sample labels;

[0102] Step S42: Construct a first random forest model. The first random forest model has basic parameters, including: the number of decision trees, the minimum number of samples for node splitting, the minimum number of samples for leaf nodes, the mean squared error as the splitting criterion, and all photon features are considered in each split. The training strategy adopts sampling with replacement.

[0103] Step S43: Input the original sample set into the first random forest model. The first random forest model evaluates the importance of each photon feature and recursively eliminates them to obtain the target photon feature set. , , The number of target photon features is determined in the following way:

[0104] Step S431, each original sample has The original sample set of each photon feature is input into the first random forest model. The first random forest model obtains the importance of each photon feature in each decision tree based on the importance of each photon feature in the prediction of forest canopy height features. Then, the importance of each photon feature in each decision tree is averaged to obtain the importance score of each photon feature. Then, the importance score is calculated... The importance of each photon feature is ranked to obtain a photon feature sequence; the first random forest model simultaneously outputs the RMSE based on the original sample set;

[0105] Step S432: In the photon feature sequence, delete the M1 photon features with the lowest importance to obtain the optimized photon feature sequence; based on the optimized photon feature sequence, update the original sample set, retaining only the optimized photon features in each original sample, thereby obtaining the optimized sample set;

[0106] The optimized sample set is input into the first random forest model, and the first random forest model outputs the importance sequence of each optimized photon feature and the RMSE.

[0107] Step S433: Determine whether the output RMSE is stable at this time; if it is stable, proceed to step S424; if it is unstable, repeat step S422.

[0108] Step S434: The optimized photon features are the target photon features, and the importance scores of each target photon feature are output.

[0109] Step S5: Based on the target photon feature set and the forest canopy height feature, process the original sample set, retaining only the target photon feature and the forest canopy height feature, to construct a sample set;

[0110] Step S6: Based on the sample set, the random forest model is optimized using the particle swarm optimization algorithm to obtain the optimal random forest model;

[0111] Step S6 is as follows:

[0112] Step S61: Divide the sample set into a feature training set and a feature test set using stratified sampling techniques;

[0113] Step S62: Select the second random forest model and train the second random forest model using the feature training set to obtain the trained second random forest model;

[0114] Step S63: Optimize the trained second random forest model using the particle swarm optimization algorithm to obtain the optimal random forest model;

[0115] Step S63 is as follows:

[0116] Step S631: Set the parameters of the particle swarm optimization algorithm, initialize the position and velocity of the particles, and calculate the fitness value of each particle using the fitness function, where each particle represents the parameter configuration of the random forest model.

[0117] Step S632: Based on the fitness value of each particle and the historical best fitness value, obtain the individual best position and the global best position;

[0118] Step S633: Based on the individual best position and the global best position, update the position and velocity of each particle, and calculate the fitness value of each particle according to the updated particle position.

[0119] Step S634: Repeat steps S632 to S633 until the preset number of iterations is reached to obtain the optimal parameter configuration;

[0120] Step S635: Obtain the optimal random forest model based on the optimal parameter configuration.

[0121] Step S63 is more specifically as follows:

[0122] Step 1: Initialize the particle swarm. Each particle represents a set of hyperparameter configurations for the random forest model, including: maximum depth (max_depth), maximum number of leaf nodes (max_leaf_nodes), minimum number of samples (min_samples_leaf), and number of decision trees (n_estimators).

[0123] Step 2: Set the particle swarm optimization algorithm parameters, including the number of particles N and the maximum number of iterations. Maximum value of inertia weight Minimum value of inertia weight Individual and group learning factors , And randomly initialize the velocity of each particle. and location ;

[0124] Step 3: Construct the fitness function. With the goal of minimizing prediction error and model complexity, the fitness function is defined as follows:

[0125] )

[0126] in: Represents the fitness function. The root mean square error of the model on the feature test set; The number of decision trees in the random forest model. and These are weight parameters, used to balance the impact of model prediction accuracy and complexity;

[0127] Step 4: Calculate the fitness value of the particle. If the current fitness is better than its historical best value, update the individual's best position; at the same time, update the global best position, which is the solution with the best fitness among all particles.

[0128] Specifically, the velocity, position, and inertial weight of each particle are iteratively adjusted according to the following update formula:

[0129] 1) Speed ​​of updates:

[0130]

[0131] 2) Location update:

[0132]

[0133] 3) Inertia weight update:

[0134]

[0135] in: and Particles In the In the nth iteration 3D velocity and position vector; and Particles In the In the nth iteration 3D velocity and position vector; Corresponding to the random forest model, the first One hyperparameter to be determined; Inertial weights; and For learning factors; and It is a random number within the interval [0,1]. For particles In the In the nth iteration Dimension's historical best position; For the group in the first In the nth iteration Dimension's historical best position; This represents the maximum number of iterations. This represents the current iteration number;

[0136] Step 5: Under the new hyperparameter combination, use the random forest model for training and validation, and re-evaluate the fitness value of each particle;

[0137] Step 6: Determine whether the maximum number of iterations or fitness convergence condition has been reached. If the condition is met, proceed to Step 7; otherwise, return to Step 4 for the next round of iteration.

[0138] Step 7: Output the optimal combination of hyperparameters, train the final random forest model, and save the model inertia weights and hyperparameter configurations for subsequent predictions.

[0139] Step S64: Train the optimal random forest model using the feature training set, and evaluate the performance of the trained optimal random forest model using the feature test set.

[0140] When evaluating the performance of the optimal random forest model after training using a feature test set, the following correlation coefficient and error metric are used for evaluation: Coefficient of Determination The metrics are: Root Mean Square Error (RMSE) and Mean Absolute Error (MAE); where the coefficient of determination is closer to 1, indicating a better fit of the random forest model; and smaller RMSE and MAE, indicating higher prediction accuracy. The calculation formulas are:

[0141]

[0142]

[0143]

[0144] in: These are the characteristic values ​​of forest canopy height, derived from GEDI observation data; These are the feature values ​​for the predicted forest canopy height characteristics; For n The mean; This represents the number of samples.

[0145] Step S7: Using the optimal random forest model, and taking the first satellite-borne lidar data of the target area as input, the forest canopy height of the target area is inverted.

[0146] The main ideas of steps S6 to S7 can be summarized as follows:

[0147] Step A1: Optimize the random forest model using the particle swarm optimization algorithm to obtain the optimal random forest model;

[0148] Step A2: Train the optimal random forest model using the feature training set, and evaluate the performance of the trained optimal random forest model using the feature test set.

[0149] Step A3: Based on the evaluation metrics, compare the performance of the random forest model with that of the optimal random forest model, and obtain the optimization results based on the comparative analysis.

[0150] Step A4: If the optimization result meets the preset requirements, output the optimization result; if it does not meet the preset requirements, perform iterative optimization until the requirements are met, save the model weights, and apply the optimal random forest model to the spaceborne lidar data of the target area.

[0151] Step A41: Analyze the optimization results, and evaluate the performance of the optimized random forest model by combining the feature dataset and evaluation metrics.

[0152] Step A411: Based on the optimization results, use the optimized random forest regression model to perform regression prediction on the feature dataset and obtain the regression prediction results;

[0153] Step A412: Compare the regression prediction results with the actual observed values ​​in the feature dataset and calculate the prediction error;

[0154] Step A413: Combine prediction error and regression evaluation index to evaluate the regression performance of the optimized random forest model.

[0155] Step A42: Compare and analyze the performance evaluation results of the random forest model with the preset requirements;

[0156] Step A43: Based on the comparative analysis results, if the optimization results meet the preset requirements, output the optimization results; if they do not meet the preset requirements, return to step A1 and perform iterative optimization until the requirements are met.

[0157] Step A44: Save the model weights and apply the optimal model to the spaceborne lidar data of the target area.

[0158] Step A441: Serialize the parameters of the random forest model optimized by the particle swarm optimization algorithm (maximum depth, maximum number of leaf nodes, minimum number of samples, number of decision trees) into a configuration file (JSON), and record the hyperparameter combination and feature importance ranking results;

[0159] Step A442: Use the machine learning framework joblib to save the trained particle swarm optimization algorithm optimized random forest model object (including decision tree set, node splitting rules, and feature weights) as a binary file.

[0160] Step A443: Input the preprocessed photon feature parameters of the target area into the model to generate the predicted canopy height.

[0161] This invention also provides a system for fusing spaceborne lidar data based on a particle swarm optimization random forest model, comprising:

[0162] The acquisition module is used to acquire multi-source spaceborne lidar data, including first spaceborne lidar data and second spaceborne lidar data; wherein, the first spaceborne lidar data is spaceborne lidar data in the form of high-density photons; and the second spaceborne lidar data is full-waveform lidar data.

[0163] The spatial matching module is used to perform spatial matching on the first spaceborne lidar data and the second spaceborne lidar data to extract several sets of effective overlapping light spots of the spaceborne lidar in the spatially overlapping region; the effective overlapping light spots of the spaceborne lidar simultaneously contain the first spaceborne lidar data and the second spaceborne lidar data in the spatially overlapping region.

[0164] The original sample set construction module is used to construct an original sample set by taking each effective overlapping spot of the spaceborne lidar as an original sample; each original sample in the original sample set has multiple photon features and forest canopy height features that affect the height of the forest canopy.

[0165] The Random Forest-Recursive Feature Elimination Module is used to perform feature filtering on all photon features based on the original sample set using the Random Forest-Recursive Feature Elimination Method, and optimize to obtain several target photon features that have a significant impact on forest canopy height, forming a target photon feature set;

[0166] The sample set construction module is used to process the original sample set according to the target photon feature set and the forest canopy height feature, retaining only the target photon feature and the forest canopy height feature, and constructing a sample set;

[0167] The optimization module is used to optimize the random forest model based on the sample set using the particle swarm optimization algorithm to obtain the optimal random forest model.

[0168] The regression prediction module is used to perform forest canopy height inversion in the target area by using the optimal random forest model and taking the first satellite-borne lidar data of the target area as input.

[0169] Two examples are described below:

[0170] Example 1: Please refer to Figures 1 to 4 As shown in the figure, the method for fusing spaceborne lidar data based on a particle swarm optimization random forest model described in this embodiment includes the following steps:

[0171] S1: Preprocess spaceborne lidar data and extract photon point cloud features and canopy height features from the data.

[0172] Specifically, preprocessing spaceborne lidar data and extracting photon point cloud features and canopy height features from the data includes the following steps:

[0173] S11. Obtain footprint points from the spaceborne lidar data and acquire filtering features, such as echo quality indicators of footprint points, terrain height, and terrain slope of each segment.

[0174] S12. Eliminate low-quality footprints based on quality indicators;

[0175] It should be noted that the conditions for removing quality flags are: 1) Exclude values ​​where the quality flag is 0 (quality_flag_a <n>= 0); 2) Exclude light spots with declining pointing and / or positioning information (rade_flag=1); 3) Exclude data collected during the day (night_flag = 0);

[0176] S13. Based on the selected footprint points, extract photon point cloud features and canopy height features, such as percentile height, canopy coverage percentage, and terrain slope per segment.

[0177] It should be further explained that percentile height is the height value located at a specific percentile point after sorting the elevation values ​​of all photon points within a single laser footpoint from lowest to highest. It typically represents the distribution characteristics of relative height above the Earth's surface. The formula for calculating percentile height is:

[0178]

[0179] In the formula, This represents the relative height of the p-th percentile. The p-th percentile of the photon elevation value (p is usually taken as 25 / 50 / 75 / 95). Reference baseline elevation (usually the elevation of the lowest photon or the elevation of the terrain surface).

[0180] S2: Extract overlapping light spots of spaceborne lidar based on spatial location matching of footprint points, obtain lidar dataset, and optimize feature dataset by filtering features based on random forest-recursive feature elimination method.

[0181] Specifically, the overlapping light spots of the spaceborne lidar are extracted based on the spatial location matching of footprint points to obtain the lidar dataset. The optimization of the feature dataset involves the following steps:

[0182] S21. Using the footprint points of the spaceborne lidar as units, perform spatial matching on different spaceborne lidar data within a preset spatial neighborhood radius, extract lidar waveforms or point cloud data in the spatially overlapping areas, and construct a spaceborne lidar overlapping spot dataset.

[0183] It should be noted that the data footprints of the ICESat-2 and GEDI spaceborne lidars in Qilian Mountain National Park were used as units, and a buffer zone with a radius of 100 meters was set as the search radius to extract spaceborne lidar data in the spatially overlapping areas and construct a spaceborne lidar overlapping spot dataset.

[0184] S22. For the constructed LiDAR dataset, the Random Forest-Recursive Feature Elimination method is used to evaluate the importance of all extracted features and recursively eliminate them to obtain the optimal feature dataset.

[0185] It should be further explained that during the recursive feature removal process, the random forest model uses squared error as the splitting criterion, constructs a forest using 100 decision trees, does not limit the maximum tree depth, sets the minimum number of samples for node splits to 2, and the minimum number of samples for leaf nodes to 1, and considers all features in each split. It also employs a bootstrap sampling method for training. Figure 2 As shown, when the number of photon feature parameters reaches more than 8, the RMSE tends to be stable; when the number of parameters is 11, the RMSE is the smallest and the standard deviation is small. Finally, the last three feature variables are removed to form the optimized feature dataset.

[0186] S23. Based on the optimized feature dataset, the feature dataset is divided into a feature training set and a feature test set using stratified sampling techniques.

[0187] It should be noted that the constructed star-borne lidar overlapping spot dataset was selected, in which the photon point cloud features of ICESat-2 were used as input features, and 98% of the waveform height of GEDI was used as input labels. The number of samples was determined and stratified sampling was adopted, and the dataset was divided into a feature training set and a feature test set in a 7:3 ratio.

[0188] S24. Based on the training requirements, select a random forest model and train the random forest model using the feature training set;

[0189] S3: Optimize the random forest model using the particle swarm optimization algorithm to obtain the optimal random forest model. Compare and analyze the random forest model with the optimal random forest model, and combine the evaluation index to obtain the optimization result.

[0190] Specifically, the random forest model is optimized using the particle swarm optimization algorithm to obtain the optimal random forest model. A comparative analysis of the random forest model and the optimal random forest model is then conducted, and the optimization results are obtained by combining evaluation metrics. The process includes the following steps:

[0191] S31. Optimize the random forest model using the particle swarm optimization algorithm to obtain the optimal random forest model;

[0192] It should be further explained that the Particle Swarm Optimization (PSO) algorithm is used to optimize the random forest model. This is achieved by leveraging PSO's efficient search capability in the parameter search space to optimize the random forest model's parameter configuration. Each particle's position update depends on its historical best position and the best position within the entire swarm. Through continuous iteration, the entire swarm tends towards the optimal solution. In this combined algorithm, each particle represents a set of parameter settings for the random forest model. Through the iterative process of PSO, the optimal parameter combination for the random forest model's performance can be explored.

[0193] Specifically, optimizing the random forest model using the particle swarm optimization algorithm to obtain the optimal random forest model includes the following steps:

[0194] S311. Set the parameters of the particle swarm optimization algorithm, initialize the position and velocity of the particles, and calculate the fitness value of each particle using the fitness function, where each particle represents the parameter configuration of the random forest model.

[0195] S312. Based on the fitness value of each particle and its historical best fitness value, obtain the individual best position and the global best position;

[0196] S313. Based on the individual best position and the global best position, update the position and velocity of each particle, and calculate the fitness value of each particle according to the updated particle position.

[0197] S314. Repeat steps S312 to S313 until the preset number of iterations is reached to obtain the optimal parameter configuration;

[0198] S315. Based on the optimal parameter configuration, the optimal random forest model is obtained.

[0199] It should be added that, such as Figure 2 As shown, the optimal values ​​for particles and the population are calculated to find the optimal parameters that minimize the difference in fitness. The specific modeling steps are as follows:

[0200] Step 1: Initialize the particle swarm. Each particle represents a set of hyperparameters for the random forest regression model, including: maximum depth (max_depth), maximum number of leaf nodes (max_leaf_nodes), minimum number of samples (min_samples_leaf), and number of decision trees (n_estimators).

[0201] Step 2: Set the parameters of the particle swarm optimization algorithm, such as the number of particles N and the maximum number of iterations. Inertia weight , Individual and group learning factors , And randomly initialize the velocity of each particle. and location .

[0202] Step 3: Construct the fitness function. With the goal of minimizing prediction error and model complexity, the fitness function is defined as follows:

[0203] )

[0204] In the formula, Represents the fitness function. The root mean square error of the model on the validation set; The number of decision trees in the random forest model. and These are weight parameters, used to balance the impact of model prediction accuracy and complexity.

[0205] Step 4: Calculate the fitness value of the particle. If the current fitness is better than its historical best value, update the individual's best position; at the same time, update the global best position (i.e. the solution with the best fitness among all particles).

[0206] Step 5: Iteratively adjust the velocity, position, and inertial weight of each particle according to the following update formula:

[0207] 1) Speed ​​of updates:

[0208]

[0209] 2) Location update:

[0210]

[0211] 3) Inertia weight update:

[0212]

[0213] In the formula and For particles In the In the nth iteration 3D velocity and position vector; Corresponding to the random forest regression model, the first One hyperparameter to be determined; Inertial weights; and For learning factors; and It is a random number within the interval [0,1]. It is a particle In the In the nth iteration Dimension's historical best position; It is the group in the first In the nth iteration Dimension's historical best position; This represents the maximum number of iterations.

[0214] Step 6: Under the new hyperparameter combination, train and validate the random forest regression model to re-evaluate the fitness value of each particle.

[0215] Step 7: Determine whether the maximum number of iterations or fitness convergence condition has been reached. If it is met, proceed to Step 8; otherwise, return to Step 4 for the next round of iterations.

[0216] Step 8: Output the optimal hyperparameter combination, train the final random forest regression model, and save the model weights and parameter configurations for subsequent predictions. Furthermore, based on the constructed feature dataset, where the training and test sets are divided in an 8:2 ratio, the parameter ranges adjusted by particle swarm optimization are as follows:

[0217] 1) max_depth: range is [4, 20];

[0218] 2) min_samples_leaf: range [1, 10];

[0219] 3) max_leaf_nodes: range [5, 50];

[0220] 4) n_estimators: range [10, 150].

[0221] Where `max_depth` represents the maximum depth of the decision tree; `min_samples_leaf` represents the minimum number of samples per leaf node; `max_leaf_nodes` represents the maximum number of leaf nodes; and `n_estimators` represents the number of decision trees. The maximum number of PSO iterations is set to 200, and the inertia weights are... Individual learning factor Group learning factor The number of particles is set to 50. The final fitness function used is:

[0222] )

[0223] In the formula, The root mean square error of the model on the validation set; The number of decision trees in the random forest model. and It can be set to empirical values ​​(e.g., 0.99 and 0.01) to ensure that accuracy is prioritized while constraining model complexity.

[0224] S33. Train the optimal random forest model using the feature training set, and evaluate the performance of the optimal random forest model after training using the feature test set.

[0225] It should be noted that, regarding the accuracy of the regression model and its applicability to high-resolution datasets, the following correlation coefficients and error metrics were used for evaluation: Coefficient of Determination (R²), Root Mean Square Error (RMSE), and Mean Absolute Error (MAE). A R² closer to 1 indicates a better fit of the regression model, while smaller RMSE and MAE indicate higher predictive accuracy. The calculation formulas are as follows:

[0226]

[0227]

[0228]

[0229] In the formula, This refers to the GEDI height value; The predicted GEDI height value; The mean of the GEDI height values; This represents the number of overlapping light spots.

[0230] S34. Based on the evaluation metrics, compare the performance of the random forest model with that of the optimal random forest model, and obtain the optimization results based on the comparative analysis.

[0231] Specifically, the hyperparameter configurations include: maximum depth (max_depth), maximum number of leaf nodes (max_leaf_nodes), minimum number of samples (min_samples_leaf), and number of decision trees (n_estimators).

[0232] S4: Analyze the optimization results. If the optimization results meet the preset requirements, output the optimization results. If they do not meet the preset requirements, return to step S3 and perform iterative optimization until the requirements are met. Save the model weights and apply the model to the spaceborne lidar data in the target area.

[0233] Specifically, the optimization results are analyzed. If the optimization results meet the preset requirements, the optimization results are output. If they do not meet the preset requirements, the process returns to step S3 for iterative optimization until the requirements are met. The model weights are then saved. Applying the model to the spaceborne lidar data in the target area includes the following steps:

[0234] S41. Analyze the optimization results, combine the feature dataset and evaluation metrics, and evaluate the performance of the optimized random forest model.

[0235] Specifically, evaluating the performance of the optimized random forest model by combining the feature dataset and evaluation metrics includes the following steps:

[0236] S411. Based on the optimization results, use the optimized random forest regression model to perform regression prediction on the feature dataset and obtain the regression prediction results.

[0237] S412. Compare the regression prediction results with the actual observed values ​​in the feature dataset and calculate the prediction error;

[0238] It should be added that the prediction error assessment includes: calculating the root mean square error (RMSE): which is the square root of MSE and more intuitively expresses the degree of deviation between the predicted value and the true value; calculating the mean absolute error (MAE): which is the average of the absolute values ​​of the prediction errors of all samples, reflecting the overall deviation; and calculating the coefficient of determination (R²): which is used to measure the goodness of fit of the model to the observations. The closer the R² value is to 1, the better the model's prediction performance.

[0239] In addition, to evaluate the stability and generalization ability of the random forest model under different parameter settings, the above error indicators can be calculated on the training set and validation set respectively, and the model can be compared and analyzed to see if there is overfitting or underfitting.

[0240] S413. Combine prediction error and regression evaluation index to evaluate the regression performance of the optimized random forest model.

[0241] S42. Compare and analyze the performance evaluation results of the random forest model with the preset requirements;

[0242] S43. Based on the comparative analysis results, if the optimization results meet the preset requirements, output the optimization results; if they do not meet the preset requirements, return to step S3 and perform iterative optimization until the requirements are met.

[0243] S44. Save the model weights and apply the optimal model to the spaceborne lidar data of the target area.

[0244] Specifically, saving the model weights and applying the optimal model to the spaceborne lidar data of the target area includes the following steps:

[0245] S441. Serialize the parameters of the random forest model optimized by the particle swarm optimization algorithm (maximum depth, maximum number of leaf nodes, minimum number of samples, number of decision trees) into a configuration file (JSON) to record the hyperparameter combination and feature importance ranking results;

[0246] S442. Use the machine learning framework joblib to save the trained particle swarm optimization algorithm optimized random forest model object (including decision tree set, node splitting rules, feature weights) as a binary file.

[0247] S443. Input the preprocessed photon feature parameters of the target area into the model to generate the predicted value of the canopy height.

[0248] Example 2:

[0249] S1: Extract the effective overlapping light spots of the spaceborne lidar through spatial matching, and remove noise and non-vegetation areas;

[0250] Load the spaceborne lidar data; execute the "buffer" tool (search radius set to 100 meters) on the ICESat-2 spot data to generate a circular coverage area for each ICESat-2 spot; use the spatial analysis module to spatially match the GEDI spot data with the ICESat-2 buffer, and filter out the overlapping areas; filter the overlapping data according to the following conditions:

[0251] 1)ICESat-2 condition: "night_flag" = 0 AND "cloud_flag_atm" < 2

[0252] 2)GEDI conditions: "quality_flag" = 1 AND "sensitivity" > 0.9 AND "modis_nonvegetated" < 70

[0253] S2: Select key features and construct an optimized feature set based on the Random Forest-Recursive Feature Elimination (RF-RFE) method;

[0254] Configure the Python environment in PyCharm, call the RFE module of the scikit-learn library, load the dataset, and extract initial features, including:

[0255] 1) Initialize the Random Forest (RF) model, set the basic parameters (number of decision trees = 100, maximum depth = 10), and score the importance of all features in the feature pool; sort the features based on the Gini index, filter the features with importance higher than the threshold (>0.01), and remove redundant features.

[0256] 2) Recursive Feature Elimination (RFE) iterative optimization: using 5-fold cross-validation, train a random forest model, calculate the out-of-bag error of the current feature subset, and remove the feature with the lowest importance.

[0257] S3: The Particle Swarm Optimization (PSO) algorithm is used to dynamically adjust the hyperparameters of the random forest (number of decision trees, maximum depth) to establish a PSO-RF spaceborne lidar satellite fusion model;

[0258] 1) Define the hyperparameters to be optimized in the random forest algorithm as dimensional variables in the particle swarm optimization. This mainly includes:

[0259] Number of decision trees (n_estimators): Set an integer value between 50 and 500; Maximum tree depth (max_depth): Set an integer value between 5 and 50; Minimum number of leaf samples (min_samples_leaf): Set an integer value between 1 and 10; Minimum number of samples for node split (min_samples_split): Set an integer value between 2 and 20.

[0260] 2) Particle swarm parameter settings: Initialize the number of particles to 50; learning factors c1=c2=1.5; inertia weight range ω_max=0.9, ω_min=0.4; maximum number of iterations iter_max=200 times.

[0261] 3) The out-of-bag (OOB) error of the random forest model on the training set is used as the fitness value, where OOB_score is obtained by calculating the coefficient of determination R² between the out-of-bag predicted value and the true value of the GEDI canopy height.

[0262] 4) Termination condition judgment: Terminate early when the change of the global optimal fitness value is less than 1e-4 in 10 consecutive iterations, or stop optimization when the maximum number of iterations of 200 is reached.

[0263] S4: Save the model weights applied to the spaceborne lidar data of the target area.

[0264] The optimized PSO-RF model parameters (number of decision trees, maximum depth, minimum number of samples per leaf node, etc.) are serialized into a configuration file (JSON), recording the hyperparameter combinations and feature importance ranking results. The trained PSO-RF model object (including the decision tree set, node splitting rules, and feature weights) is saved as a binary file using the machine learning framework joblib. The preprocessed photon feature parameters of the target region are input into the model to generate canopy height predictions.

[0265] Specifically, Python scripts can configure parameters, train and update the PSO-RF fusion model, and deploy its applications, enabling the forest canopy height inversion system to integrate the latest fusion algorithms in a timely manner, achieving efficient processing and automatic estimation of multi-source spaceborne lidar data, thereby ensuring mapping accuracy and processing efficiency.

[0266] This embodiment solves the problem of missing GEDI coverage in areas above 51.6° north and south latitude by fusing ICESat-2 (spot diameter 17m) and GEDI (spot diameter 25m) data, and addresses the issues of insufficient samples and data discrepancies from a single data source, thus achieving data complementarity and expanded coverage.

[0267] The beneficial effects of this invention are as follows:

[0268] 1. This invention improves feature selection and datasets, enhancing data quality and depth, thereby deepening feature engineering. This enables the random forest model to capture key information in the data more accurately, thus better coping with complex and ever-changing environments and providing stronger support for decision-making.

[0269] 2. This invention uses the particle swarm optimization algorithm to globally optimize the key hyperparameters of the random forest model (including maximum depth, maximum number of leaf nodes, minimum number of samples, and number of decision trees), which effectively improves the prediction accuracy and generalization ability of the model.

[0270] 3. This invention, through a random forest regression model based on particle swarm optimization, can effectively mine structural information among lidar data from different platforms, providing an accurate and stable modeling scheme for high-resolution remote sensing applications such as forest structure parameter estimation, and helping to promote the development of fields such as remote sensing ecological monitoring and carbon storage assessment.

[0271] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.< / n>

Claims

1. A method for fusing spaceborne lidar data based on a particle swarm optimization random forest model, characterized in that, Includes the following steps: Step S1: Acquire multi-source spaceborne lidar data, including first spaceborne lidar data and second spaceborne lidar data; wherein, the first spaceborne lidar data is spaceborne lidar data in the form of high-density photons; and the second spaceborne lidar data is full-waveform lidar data. Step S2: Spatial matching is performed on the first spaceborne lidar data and the second spaceborne lidar data to extract several sets of effective overlapping light spots of spaceborne lidar in the spatially overlapping region; the effective overlapping light spots of spaceborne lidar simultaneously contain the first spaceborne lidar data and the second spaceborne lidar data in the spatially overlapping region. Step S3: Take each of the effective overlapping light spots of the spaceborne lidar as a raw sample to construct a raw sample set; each raw sample in the raw sample set has multiple photon features and forest canopy height features that affect the height of the forest canopy. Step S4: Based on the original sample set, the random forest-recursive feature elimination method is used to filter all photon features and optimize them to obtain several target photon features that have a significant impact on forest canopy height, forming a target photon feature set; Step S5: Based on the target photon feature set and the forest canopy height feature, process the original sample set, retaining only the target photon feature and the forest canopy height feature, to construct a sample set; Step S6: Based on the sample set, the random forest model is optimized using the particle swarm optimization algorithm to obtain the optimal random forest model; Step S7: Using the optimal random forest model, and taking the first satellite-borne lidar data of the target area as input, the forest canopy height of the target area is inverted.

2. The method for fusing spaceborne lidar data based on a particle swarm optimization random forest model according to claim 1, characterized in that, The first satellite-borne lidar data is ICESat-2 satellite-borne lidar data; the second satellite-borne lidar data is GEDI satellite-borne lidar data.

3. The method for fusing spaceborne lidar data based on a particle swarm optimization random forest model according to claim 1, characterized in that, Step S2 is as follows: Step S21: Using the footprint point S1 of the first satellite-borne lidar data and the footprint point S2 of the second satellite-borne lidar data as units, for each footprint point, low-quality footprint points are removed according to the echo quality indicator of the footprint point to achieve the screening of footprint points. Step S22: For each of the filtered footprint points, for each footprint point S1 in the first satellite-borne lidar data, a buffer zone with a radius of R is set as the search radius, centered on footprint point S1, to generate a circular coverage area D1 for each footprint point S1. For each footprint point S2 of the second satellite-borne lidar data, a buffer zone with radius R is set as the search radius, centered on footprint point S2, to generate a circular coverage area D2 for each footprint point S2. Spatial matching is performed on circular coverage areas D1 and D2, and the overlapping area between the two is selected, which is the effective overlapping light spot of the spaceborne lidar.

4. The method for fusing spaceborne lidar data based on a particle swarm optimization random forest model according to claim 1, characterized in that, Step S3 is as follows: For each of the said spaceborne lidar effective overlapping light spots, there are multiple first spaceborne lidar data and multiple second spaceborne lidar data; The average of the original values ​​of each photon feature of all the first spaceborne lidar data within the effective overlapping spot of the spaceborne lidar is taken as the feature value of the corresponding photon feature of the original sample. The average GEDI waveform of all second-spaceborne lidar data within the effective overlapping spot of the spaceborne lidar is taken, and then 98 percent of the height of the average GEDI waveform is used as the forest canopy height feature of the original sample.

5. The method for fusing spaceborne lidar data based on a particle swarm optimization random forest model according to claim 1, characterized in that, Step S4 is as follows: Step S41, for each original sample in the original sample set, it is represented as: ;in, Indicates the first One original sample; Indicates the number of photon features; These represent the eigenvalues ​​of the first photon feature, the second photon feature, ..., the eigenvalues ​​of the third photon feature, respectively. Eigenvalues ​​of each photon feature; The feature values ​​representing the characteristics of forest canopy height serve as sample labels; Step S42: Construct a first random forest model. The first random forest model has basic parameters, including: the number of decision trees, the minimum number of samples for node splitting, the minimum number of samples for leaf nodes, the mean squared error as the splitting criterion, and all photon features are considered in each split. The training strategy adopts sampling with replacement. Step S43: Input the original sample set into the first random forest model. The first random forest model evaluates the importance of each photon feature and recursively eliminates them to obtain the target photon feature set. , , The number of target photon features is determined in the following way: Step S431, each original sample has The original sample set of each photon feature is input into the first random forest model. The first random forest model obtains the importance of each photon feature in each decision tree based on the importance of each photon feature in the prediction of forest canopy height features. Then, the importance of each photon feature in each decision tree is averaged to obtain the importance score of each photon feature. Then, the importance score is calculated... The importance of each photon feature is ranked to obtain a photon feature sequence; the first random forest model simultaneously outputs the RMSE based on the original sample set; Step S432: In the photon feature sequence, delete the M1 photon features with the lowest importance to obtain the optimized photon feature sequence; based on the optimized photon feature sequence, update the original sample set, retaining only the optimized photon features in each original sample, thereby obtaining the optimized sample set; The optimized sample set is input into the first random forest model, and the first random forest model outputs the importance sequence of each optimized photon feature and the RMSE. Step S433: Determine whether the output RMSE is stable at this time; if it is stable, proceed to step S424; if it is unstable, repeat step S422. Step S434: The optimized photon features are the target photon features, and the importance scores of each target photon feature are output.

6. The method for fusing spaceborne lidar data based on a particle swarm optimization random forest model according to claim 1, characterized in that, Step S6 is as follows: Step S61: Divide the sample set into a feature training set and a feature test set using stratified sampling techniques; Step S62: Select the second random forest model and train the second random forest model using the feature training set to obtain the trained second random forest model; Step S63: Optimize the trained second random forest model using the particle swarm optimization algorithm to obtain the optimal random forest model; Step S64: Train the optimal random forest model using the feature training set, and evaluate the performance of the trained optimal random forest model using the feature test set.

7. The method for fusing spaceborne lidar data based on a particle swarm optimization random forest model according to claim 6, characterized in that, Step S63 is as follows: Step S631: Set the parameters of the particle swarm optimization algorithm, initialize the position and velocity of the particles, and calculate the fitness value of each particle using the fitness function, where each particle represents the parameter configuration of the random forest model. Step S632: Based on the fitness value of each particle and the historical best fitness value, obtain the individual best position and the global best position; Step S633: Based on the individual best position and the global best position, update the position and velocity of each particle, and calculate the fitness value of each particle according to the updated particle position. Step S634: Repeat steps S632 to S633 until the preset number of iterations is reached to obtain the optimal parameter configuration; Step S635: Obtain the optimal random forest model based on the optimal parameter configuration.

8. The method for fusing spaceborne lidar data based on a particle swarm optimization random forest model according to claim 7, characterized in that, Step S63 is more specifically as follows: Step 1: Initialize the particle swarm. Each particle represents a set of hyperparameters for the random forest model, including: maximum depth, maximum number of leaf nodes, minimum number of samples, and number of decision trees. Step 2: Set the particle swarm optimization algorithm parameters, including the number of particles N and the maximum number of iterations. Maximum value of inertia weight Minimum value of inertia weight Individual and group learning factors , And randomly initialize the velocity of each particle. and location ; Step 3: Construct the fitness function. With the goal of minimizing prediction error and model complexity, the fitness function is defined as follows: ) in: Represents the fitness function. The root mean square error of the model on the feature test set; This represents the number of decision trees in the random forest model. and These are weight parameters, used to balance the impact of model prediction accuracy and complexity; Step 4: Calculate the fitness value of the particle. If the current fitness is better than its historical best value, update the individual's best position; at the same time, update the global best position, which is the solution with the best fitness among all particles. Specifically, the velocity, position, and inertial weight of each particle are iteratively adjusted according to the following update formula: 1) Speed ​​of updates: , 2) Location update: , 3) Inertia weight update: , in: and Particles In the In the nth iteration 3D velocity and position vector; and Particles In the In the nth iteration 3D velocity and position vector; Corresponding to the random forest model, the first One hyperparameter to be determined; Inertial weights; and For learning factors; and It is a random number within the interval [0,1]. For particles In the In the nth iteration Dimension's historical best position; For the group in the first In the nth iteration Dimension's historical best position; This represents the maximum number of iterations. This represents the current iteration number; Step 5: Under the new hyperparameter combination, use the random forest model for training and validation, and re-evaluate the fitness value of each particle; Step 6: Determine whether the maximum number of iterations or fitness convergence condition has been reached. If the condition is met, proceed to Step 7; otherwise, return to Step 4 for the next round of iteration. Step 7: Output the optimal combination of hyperparameters, train the final random forest model, and save the model inertia weights and hyperparameter configurations for subsequent predictions.

9. The method for fusing spaceborne lidar data based on a particle swarm optimization random forest model according to claim 6, characterized in that, When evaluating the performance of the optimal random forest model after training using a feature test set, the following correlation coefficient and error metric are used for evaluation: Coefficient of Determination The metrics are: Root Mean Square Error (RMSE) and Mean Absolute Error (MAE); where the coefficient of determination is closer to 1, indicating a better fit of the random forest model; and smaller RMSE and MAE, indicating higher prediction accuracy. The calculation formulas are: , , , in: These are the characteristic values ​​of forest canopy height, derived from GEDI observation data; These are the feature values ​​for the predicted forest canopy height characteristics; For n The mean; This represents the number of samples.

10. A system for fusing spaceborne lidar data based on a particle swarm optimization random forest model, characterized in that, include: The acquisition module is used to acquire multi-source spaceborne lidar data, including first spaceborne lidar data and second spaceborne lidar data; wherein, the first spaceborne lidar data is spaceborne lidar data in the form of high-density photons; and the second spaceborne lidar data is full-waveform lidar data. The spatial matching module is used to perform spatial matching on the first spaceborne lidar data and the second spaceborne lidar data to extract several sets of effective overlapping light spots of the spaceborne lidar in the spatially overlapping region; the effective overlapping light spots of the spaceborne lidar simultaneously contain the first spaceborne lidar data and the second spaceborne lidar data in the spatially overlapping region. The original sample set construction module is used to construct an original sample set by taking each effective overlapping spot of the spaceborne lidar as an original sample; each original sample in the original sample set has multiple photon features and forest canopy height features that affect the height of the forest canopy. The Random Forest-Recursive Feature Elimination Module is used to perform feature filtering on all photon features based on the original sample set using the Random Forest-Recursive Feature Elimination Method, and optimize to obtain several target photon features that have a significant impact on forest canopy height, forming a target photon feature set; The sample set construction module is used to process the original sample set according to the target photon feature set and the forest canopy height feature, retaining only the target photon feature and the forest canopy height feature, and constructing a sample set; The optimization module is used to optimize the random forest model based on the sample set using the particle swarm optimization algorithm to obtain the optimal random forest model. The regression prediction module is used to perform forest canopy height inversion in the target area by using the optimal random forest model and taking the first satellite-borne lidar data of the target area as input.