An inversion method to determine the spatial distribution of biological crusts

By collecting multidimensional environmental factor data and constructing a machine learning model, the problem that the distribution characteristics of biological crusts in soil erosion models in arid and semi-arid regions are difficult to reflect was solved, and high-precision biological crust distribution inversion and model prediction accuracy were achieved.

CN122287316APending Publication Date: 2026-06-26黄河流域水土保持生态环境监测中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610356739.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-23
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing soil erosion models have insufficient prediction accuracy in arid and semi-arid regions, mainly because they fail to fully consider the soil and water conservation function of biocrusts and lack multi-source data fusion and accurate quantification methods, resulting in the distribution characteristics of biocrusts not being accurately reflected, thus affecting the accuracy of erosion simulation by the models.

Method used

By collecting multi-dimensional environmental factor data through field surveys combined with UAVs and high-resolution remote sensing images, a biological crust distribution prediction model was constructed using machine learning methods. Key driving factors were identified, their spatiotemporal evolution patterns were inverted, and combined with spatiotemporal dynamic analysis, a distribution prediction model applicable to different ecological environments was constructed.

Benefits of technology

It has achieved high-precision inversion of the distribution of biological crusts, improved the prediction accuracy of soil erosion models, provided high-precision distribution data support, and provided a reliable foundation for regional soil and water conservation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122287316A_ABST
    Figure CN122287316A_ABST
Patent Text Reader

Abstract

This invention relates to the field of remote sensing monitoring technology for soil erosion and the ecological environment, and particularly to an inversion method for clearly defining the spatial distribution of biocrusts. The method involves collecting multidimensional environmental factor data influencing biocrust distribution through field surveys combined with UAV and high-resolution remote sensing imagery, and constructing a comprehensive environmental database. Key environmental driving factors are screened through statistical testing, and their effects on biocrust distribution and spatial differences are analyzed. Machine learning methods such as random forest, support vector machine, and gradient boosting decision tree are used to construct a static distribution model of biocrusts. This is further combined with time series decomposition technology to construct a spatiotemporally coupled probabilistic model of biocrust distribution. The model prediction results are visualized using GIS technology, inverting the spatiotemporal distribution characteristics of biocrusts and conducting uncertainty analysis. This achieves a refined and quantitative inversion of the spatial distribution of biocrusts in typical small watersheds in arid and semi-arid regions, clarifying the spatiotemporal evolution law and environmental driving mechanism of biocrust distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing monitoring technology for soil erosion and ecological environment, and in particular to an inversion method for determining the spatial distribution of biological crusts. Background Technology

[0002] my country's arid and semi-arid regions are widely distributed, with fragile ecosystems and prominent soil erosion problems. In addition to vascular plants, these regions are also rich in biocrusts (biocrusts), formed by the cementation of microorganisms, mosses, lichens, and other organisms with topsoil particles, which have become an important component of land cover. Biocrusts significantly enhance soil's resistance to erosion and scour, improve surface hydrological processes, and have a significant impact on the spatial and temporal patterns of soil erosion, making them a key factor in maintaining the ecological stability of these regions.

[0003] Soil erosion models are key tools for quantitatively estimating soil loss and guiding soil and water conservation planning. Currently, the model system represented by the Chinese Soil Loss Equation (CSLE) still has insufficient prediction accuracy when applied to arid and semi-arid regions. One of the core problems is that existing models have not fully and reasonably considered the soil and water conservation function of biological crusts.

[0004] The developmental types, coverage, and thickness of biocrusts are influenced by a combination of environmental factors, including precipitation, soil texture, topography, and vegetation, exhibiting significant spatial heterogeneity at the small watershed and even slope scales. Existing models typically base land cover management factors on vascular vegetation or simplify biocrusts to uniform parameters, failing to accurately reflect their patchy and dynamic distribution characteristics. This leads to discrepancies between model input data and field conditions, consequently affecting the accuracy of subsequent erosion simulations.

[0005] Furthermore, existing studies on the distribution of biocrusts are mostly qualitative descriptions or local field surveys, lacking multi-source data fusion and precise quantification methods. This makes it difficult to systematically reveal the spatiotemporal evolution of biocrusts at the small watershed scale, and also fails to provide high-precision distribution data support for soil erosion models.

[0006] In summary, overcoming the accuracy bottleneck of existing soil erosion models in arid and semi-arid regions hinges on solving the problem of fine spatial inversion of biocrusts. Developing a technical method capable of accurately characterizing the spatial distribution and spatiotemporal evolution of biocrusts has significant theoretical and applied value for improving the accuracy of model input data and providing reliable basic data for precise control of regional soil erosion and ecological restoration. Summary of the Invention

[0007] This invention takes biocrusts in typical small watersheds in arid and semi-arid regions as the research object. It constructs a comprehensive environmental database through multi-source data collection, screens key driving factors of biocrust distribution, analyzes their effects and spatial differences, constructs a biocrust distribution prediction model, inverts the spatiotemporal distribution characteristics and evolution laws of biocrusts, and conducts uncertainty analysis to provide high-precision biocrust distribution data for soil erosion model construction and regional ecological governance.

[0008] The purpose of this invention is to overcome the above-mentioned problems and provide an inversion method for clearly determining the spatial distribution of biological crusts. To achieve the above objective, this invention adopts the following technical solution: A method for clearly defining the spatial distribution of biocrusts is proposed. Taking biocrusts in typical small watersheds in arid and semi-arid regions as the research object, this method collects multidimensional environmental factor data affecting biocrust distribution through field surveys combined with UAV and high-resolution remote sensing imagery, and constructs a comprehensive environmental database. Machine learning methods are used to analyze the nonlinear relationships, interactions, and spatial differences among factors, identify key driving factors, and, combined with spatiotemporal dynamic analysis, construct a biocrust distribution prediction model applicable to different ecological environments to invert the spatiotemporal evolution patterns and main factor characteristics of biocrusts. The method includes the following steps: Step S1: Field survey: Set up repeating plots in woodland, shrubland and grassland, and record the slope aspect, slope, altitude, landform location and biological crust data and vegetation data of the plots. Obtain spectral data of algal crust, moss crust and bare soil through spectral measurement. Smooth the spectral data and calculate the first derivative spectrum to identify characteristic bands. Collect soil samples using the five-point sampling method and determine the soil physicochemical properties. Simultaneously monitor meteorological data and soil moisture content. Step S2, Factor Screening and Model Building: Key environmental factors related to biological crusts are screened through statistical tests, and their effects and spatial differences are analyzed. After data preprocessing, machine learning methods are used to build a distribution model. Combined with relevant data standardization and splitting of training and validation sets, the model is validated through evaluation indicators.

[0009] Step S3, Distribution of biocrust and its prediction model: Using the land use type map as the base map, and combining key environmental factor data with the biocrust distribution model, a biocrust distribution map and an uncertainty map are generated. Based on the seasonal / interannual variation characteristics of environmental variables, the key environmental variables are transformed into time dynamic parameters through time series decomposition technology, and a spatiotemporally coupled biocrust distribution probability model is constructed.

[0010] Furthermore, parameters related to the distribution of biological crusts include coverage, type, and the ratio of algae to moss; multidimensional environmental factors include rainfall, soil, and topography.

[0011] Furthermore, in step S2, factor selection and model construction include the following steps: S21. Factor Screening: Collinearity and independence tests were used to screen environmental factors that were highly correlated with biological crusts. After eliminating redundant environmental variables, the direct and indirect effects of key environmental factors on the distribution of biological crusts and their spatial differences were analyzed by using structural equation modeling with geographic weighted regression. S22. Complete data cleaning, unify the units, construct and select feature processing, and construct biological crust distribution models using random forest, support vector machine and gradient boosting decision tree respectively. The formula for calculating the predicted values ​​of the random forest model is: ; in, These are the predicted values ​​from the random forest. For the number of trees, For the first The predicted value of each tree; S23. Model Validation: Screen relevant literature and extract biological crust parameters, environmental factors and spatial information. Standardize the data and randomly split all data into training and validation sets. Evaluate the model performance through evaluation indicators, including root mean square error, coefficient of determination and mean absolute percentage error.

[0012] Furthermore, in step S3, constructing the biological crust distribution and its prediction model includes the following steps: S31. Time slicing processing: Divide long-term environmental data into time periods according to seasons or years, train distribution models for each period, and generate probability grids of biological crust distribution for different time periods. S32. Spatiotemporal coupling mechanism: Quantify the impact of the spatiotemporal interaction of environmental variables on the probability distribution through panel data analysis or spatiotemporal kriging. S33. Dynamic Validation: Using the true values ​​of biological crust cover change retrieved from multiple periods of remote sensing images, the model's ability to predict spatiotemporal evolution is evaluated through cross-validation, and parameters are dynamically calibrated.

[0013] Further, in step S31, the calculation formula for the biological crust distribution model is: slope Soil type Vegetation cover ; in, This is a model for the distribution of biological crusts. For model expressions, The formula for calculating model uncertainty is: ; in, For the first The uncertainty of each grid For the first Quantiles of each grid cell, For the first Quantiles of each grid cell, For the first The average of multiple predictions for each grid cell.

[0014] The advantages of this invention are: This invention collects multi-dimensional environmental factor data through field surveys combined with UAVs and high-resolution remote sensing images. It then uses machine learning methods to construct a biological crust distribution prediction model to invert its spatiotemporal evolution. The biological crust factor is then incorporated as a sub-factor into the Chinese soil loss equation. The model is corrected and validated through gradient boosting decision trees, thereby quantifying the distribution pattern and influencing mechanism of biological crust. This effectively solves the problem of insufficient prediction accuracy of soil erosion models in arid and semi-arid regions and improves the applicability of the model in this area. Attached Figure Description

[0015] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application.

[0016] In the attached diagram: Figure 1 This is the design diagram of the field survey quadrat in Example 1.

[0017] Figure 2 This is a flowchart of the method for inverting the spatial distribution of biological crusts in Example 1.

[0018] Figure 3 This is a flowchart of factor selection and model construction in Example 1.

[0019] Figure 4 This is a flowchart of the construction of the biological crust distribution and its prediction model in Example 1. Detailed Implementation

[0020] The present invention will now be described in detail and specifically through specific embodiments to enable a better understanding of the invention. However, the following embodiments do not limit the scope of protection of the present invention.

[0021] Example 1 A method for clearly defining the spatial distribution of biocrusts is proposed. Taking biocrusts in typical small watersheds in arid and semi-arid regions as the research object, this method collects multidimensional environmental factor data affecting biocrust distribution through field surveys combined with UAV and high-resolution remote sensing imagery, and constructs a comprehensive environmental database. Machine learning methods are used to analyze the nonlinear relationships, interactions, and spatial differences among factors, identify key driving factors, and, combined with spatiotemporal dynamic analysis, construct a biocrust distribution prediction model applicable to different ecological environments to invert the spatiotemporal evolution patterns and main factor characteristics of biocrusts. The method includes the following steps: Step S1: Field survey: Set up repeating plots in woodland, shrubland and grassland, and record the slope aspect, slope, altitude, landform location and biological crust data and vegetation data of the plots. Obtain spectral data of algal crust, moss crust and bare soil through spectral measurement. Smooth the spectral data and calculate the first derivative spectrum to identify characteristic bands. Collect soil samples using the five-point sampling method and determine the soil physicochemical properties. Simultaneously monitor meteorological data and soil moisture content. In specific embodiments, such as Figure 1-2 As shown, the distribution inversion of biological crusts within a small watershed was carried out in the following manner: Field surveys focused on woodlands, shrublands, and grasslands, with 3-5 replicates in each area. Three 10×10 meter plots were established for each plot, and detailed information was recorded regarding slope aspect, slope gradient, elevation, geomorphic location, and data on biocrust type, cover, biomass, roughness, thickness, and vegetation cover, species, height, abundance, and evenness. An ASDFieldSpec 4 portable spectrometer (wavelength range 350-2500 nm) was used, with five spectral measurement points set up in each plot in the core area, avoiding vegetation shade. Measurements were taken in May, July, September, and November, obtaining 30 spectral curves each for algal crusts, moss crusts, and bare soil. Savitzky-Golay filtering was used to smooth the spectral curves, and first-order derivatives were calculated to identify characteristic bands. The chlorophyll variation zone of 680-750 nm was analyzed in detail, and five-point sampling was used to collect data from 0-20 cm depths. Soil samples were brought back to the laboratory to determine the soil's physical and chemical properties. Rain gauges and TDR probes were used to monitor meteorological data and soil moisture content.

[0022] Furthermore, such as Figure 3 As shown, step S2, factor screening and model building: key environmental factors related to biological crusts are screened through statistical tests and their effects and spatial differences are analyzed. After data preprocessing, a distribution model is built using machine learning methods. Combined with relevant data standardization processing and splitting of training and validation sets, the model is validated through evaluation indicators.

[0023] Factor selection and model construction include the following steps: S21. Factor Screening: Collinearity and independence tests were used to screen environmental factors that were highly correlated with biological crusts. After eliminating redundant environmental variables, the direct and indirect effects of key environmental factors on the distribution of biological crusts and their spatial differences were analyzed by using structural equation modeling with geographic weighted regression. S22. Complete data cleaning, unify the units, construct and select feature processing, and construct biological crust distribution models using random forest, support vector machine and gradient boosting decision tree respectively. The formula for calculating the predicted values ​​of the random forest model is: ; in, These are the predicted values ​​from the random forest. For the number of trees, For the first The predicted value of each tree; S23. Model Validation: Screen relevant literature and extract biological crust parameters, environmental factors and spatial information. Standardize the data and randomly split all data into training and validation sets. Evaluate the model performance through evaluation indicators, including root mean square error, coefficient of determination and mean absolute percentage error.

[0024] In a specific embodiment, during the factor selection and model building stages, statistical methods such as collinearity and independence tests are first used to screen environmental factors with high correlation to biological crust distribution, avoiding overfitting caused by multicollinearity among environmental covariates. If the absolute value of the Pearson correlation coefficient |r| between any two variables exceeds 0.80, redundant environmental variables are removed, retaining the factors most significantly correlated with the distribution of biological crust to determine key environmental factors. Subsequently, a structural equation model with embedded geographic weighted regression is used to explore the direct and indirect effects of key environmental factors on the distribution of biological crust and their spatial differences. Before model building, data cleaning, unification of dimensions, feature construction, and selection are completed to provide high-quality, standardized data input for subsequent models. Three models—random forest, support vector machine, and gradient boosting decision tree—are used to model the same dataset. The Random Forest package in R is used, and the caret package is used to optimize the model parameters mtry, netree, and nodesize to improve the accuracy and reliability of model predictions. Finally, their prediction accuracy and generalization ability are compared. The formula for calculating the predicted value of the random forest model is: ; in, These are the predicted values ​​from the random forest. For the number of trees, For the first The predicted value of each tree; During model validation, English and Chinese literature from 2000 to 2023 was screened on platforms such as Web of Science and CNKI using keywords such as "biological soil crusts," "loess plateau," and "environmental factors." Data such as biological crust parameters, environmental factors, and spatial information were extracted. Data with inconsistent units were standardized, and the correlation coefficients r in each literature were converted into Fisher's Z values ​​for integration. All data were mixed and randomly split, with 80% used as the training set and 20% as the validation set. The root mean square error, coefficient of determination, and mean absolute percentage error were used to evaluate the model performance.

[0025] Furthermore, such as Figure 4 As shown, step S3, distribution of biological crust and its prediction model: using the land use type map as the base map, combined with key environmental factor data and the biological crust distribution model, a biological crust distribution map and an uncertainty map are generated. For the seasonal / interannual variation characteristics of environmental variables, the key environmental variables are transformed into time dynamic parameters through time series decomposition technology, and a spatiotemporally coupled biological crust distribution probability model is constructed.

[0026] The steps involved in constructing a model for the distribution and prediction of biological crusts are as follows: S31. Time slicing processing: Divide long-term environmental data into time periods according to seasons or years, train distribution models for each period, and generate probability grids of biological crust distribution for different time periods. The formula for calculating the distribution model of biological crusts is: slope Soil type Vegetation cover ; in, This is a model for the distribution of biological crusts. For model expressions, The formula for calculating model uncertainty is: ; in, For the first The uncertainty of each grid For the first Quantiles of each grid cell, For the first Quantiles of each grid cell, For the first The average of multiple predictions for each grid cell.

[0027] S32. Spatiotemporal coupling mechanism: Quantify the impact of the spatiotemporal interaction of environmental variables on the probability distribution through panel data analysis or spatiotemporal kriging. S33. Dynamic Validation: Using the true values ​​of biological crust cover change retrieved from multiple periods of remote sensing images, the model's ability to predict spatiotemporal evolution is evaluated through cross-validation, and parameters are dynamically calibrated.

[0028] In a specific embodiment, during the construction of the biocrust distribution and prediction model, a land use type map is used as the base map. Using vector or raster data corresponding to key environmental factors and the biocrust distribution model, the prediction results are visualized using GIS software, generating a biocrust distribution map and an uncertainty map. This explores the distribution prediction patterns of biocrust under different environmental factors and their spatial variability. The bootstrap method is used to repeat the experiment 100 times to calculate the final prediction results and the uncertainties in the modeling process. 90% of the sample points are randomly selected from all sample points as the modeling set, and the remaining 10% are used as the independent validation set. After the model is run randomly 100 times, the average biocrust coverage of each raster is calculated as the final prediction result. Simultaneously, the upper and lower confidence intervals of the 95% confidence level for each raster are calculated. The expression for the biocrust distribution model is: slope Soil type Vegetation cover ; in, This is a model for the distribution of biological crusts. For model expressions, The formula for calculating model uncertainty is: ; in, For the first The uncertainty of each grid For the first The 95th percentile of each grid cell, For the first The 5th percentile of each grid cell, For the first The average of 100 predictions for each grid cell.

[0029] To address the seasonal and interannual variability of environmental variables, a time-series decomposition technique is introduced on top of the static distribution model. This involves noise removal through moving averages and extraction of periodic signals through Fourier transform, converting key environmental variables such as monthly average precipitation and quarterly accumulated temperature into dynamic time parameters to construct a spatiotemporally coupled probability model for biological crust distribution. Specific methods include time-slicing, dividing long-term environmental data into seasonal or interannual periods such as spring and El Niño years, training the distribution model separately for each period, and generating probability raster grids for biological crust distribution in different time periods; the spatiotemporal coupling mechanism quantifies the impact of spatiotemporal interactions of environmental variables on the distribution probability through panel data analysis or spatiotemporal kriging; and dynamic validation utilizes ground truth values ​​of biological crust cover changes retrieved from multi-period remote sensing images, evaluating the model's predictive ability for spatiotemporal evolution through cross-validation methods such as time block validation, and dynamically calibrating the parameters.

[0030] The specific embodiments of the present invention have been described in detail above, but they are merely examples, and the present invention is not equivalent to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions to the present invention are also within the scope of the present invention. Therefore, all equivalent transformations and modifications made without departing from the spirit and scope of the present invention should be covered within the scope of the present invention.

Claims

1. An inversion method for clearly determining the spatial distribution of biological crusts, characterized in that, This study focuses on biocrusts in typical small watersheds of arid and semi-arid regions. Through field surveys combined with UAV and high-resolution remote sensing imagery, multidimensional environmental factor data influencing biocrust distribution were collected. A comprehensive environmental database was constructed, and machine learning methods were used to analyze the nonlinear relationships, interactions, and spatial differences among these factors. Key driving factors were identified, and combined with spatiotemporal dynamic analysis, a biocrust distribution prediction model applicable to different ecological environments was constructed to invert the spatiotemporal evolution patterns and main factor characteristics of biocrusts. The study includes the following steps: Step S1: Field survey: Set up repeating plots in woodland, shrubland and grassland, and record the slope aspect, slope, altitude, landform location and biological crust data and vegetation data of the plots. Obtain spectral data of algal crust, moss crust and bare soil through spectral measurement. Smooth the spectral data and calculate the first derivative spectrum to identify characteristic bands. Collect soil samples using the five-point sampling method and determine the soil physicochemical properties. Simultaneously monitor meteorological data and soil moisture content. Step S2, Factor Screening and Model Building: Key environmental factors related to biological crusts are screened through statistical tests and their effects and spatial differences are analyzed. After data preprocessing, machine learning methods are used to build a distribution model. Combined with relevant data standardization and splitting of training and validation sets, the model is validated through evaluation indicators. Step S3, Distribution of biocrust and its prediction model: Using the land use type map as the base map, and combining key environmental factor data with the biocrust distribution model, a biocrust distribution map and an uncertainty map are generated. Based on the seasonal / interannual variation characteristics of environmental variables, the key environmental variables are transformed into time dynamic parameters through time series decomposition technology, and a spatiotemporally coupled biocrust distribution probability model is constructed.

2. The inversion method for clearly determining the spatial distribution of biological crusts according to claim 1, characterized in that, The parameters related to the distribution of biological crusts include coverage, type, and the ratio of algae to moss; the multidimensional environmental factors include rainfall, soil, and topography.

3. The inversion method for clearly determining the spatial distribution of biological crusts according to claim 2, characterized in that, In step S2, the factor screening and model construction include the following steps: S21. Factor Screening: Collinearity and independence tests were used to screen environmental factors that were highly correlated with biological crusts. After eliminating redundant environmental variables, the direct and indirect effects of key environmental factors on the distribution of biological crusts and their spatial differences were analyzed by using structural equation modeling with geographic weighted regression. S22. Complete data cleaning, unify the units, construct and select feature processing, and construct biological crust distribution models using random forest, support vector machine and gradient boosting decision tree respectively. The formula for calculating the predicted values ​​of the random forest model is: ; in, These are the predicted values ​​from the random forest. For the number of trees, For the first The predicted value of each tree; S23. Model Validation: Screen relevant literature and extract biological crust parameters, environmental factors and spatial information. Standardize the data and randomly split all data into training set and validation set. Evaluate the model performance through evaluation indicators, including root mean square error, coefficient of determination and mean absolute percentage error.

4. The inversion method for clearly determining the spatial distribution of biological crusts according to claim 3, characterized in that, In step S3, constructing the distribution and prediction model of biological crusts includes the following steps: S31. Time slicing processing: Divide long-term environmental data into time periods according to seasons or years, train distribution models for each period, and generate probability grids of biological crust distribution for different time periods. S32. Spatiotemporal coupling mechanism: Quantify the impact of the spatiotemporal interaction of environmental variables on the probability distribution through panel data analysis or spatiotemporal kriging. S33. Dynamic Validation: Using the true values ​​of biological crust cover change retrieved from multiple periods of remote sensing images, the model's ability to predict spatiotemporal evolution is evaluated through cross-validation, and parameters are dynamically calibrated.

5. The inversion method for clearly determining the spatial distribution of biological crusts according to claim 4, characterized in that, In step S31, the calculation formula for the biological crust distribution model is as follows: slope Soil type Vegetation cover ; in, This is a model for the distribution of biological crusts. For model expressions, The formula for calculating model uncertainty is: ; in, For the first The uncertainty of each grid For the first Quantiles of each grid cell, For the first Quantiles of each grid cell, For the first The average of multiple predictions for each grid cell.