A method for identifying early warning signals of forest and grassland vegetation degradation

By acquiring various data and combining vegetation indices with machine learning models, early warning signals of forest and grassland degradation are identified, solving the problem of the lack of early warning signal identification in existing technologies and enabling precise early warning and management strategy support for different ecosystems.

CN120541521BActive Publication Date: 2026-01-06INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510606780.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2026-01-06
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

Existing technologies lack effective methods for identifying early warning signals of forest and grassland vegetation degradation, resulting in the inability to provide timely warnings and risk simulations, which weakens the climate mitigation potential and ecosystem service functions of vegetation.

Method used

By acquiring vegetation, climate, soil, geographical environment, and socioeconomic data of the study area, and combining them with vegetation indices, the BFAST algorithm is used to identify negative mutation points in vegetation. Furthermore, a machine learning model is used to screen out key influencing factors of nature, human activities, and the environment, and to determine the critical point of vegetation change as an early warning signal.

Benefits of technology

It can accurately identify early warning signals of forest and grassland vegetation degradation, is applicable to different ecosystems and regions, provides scientific support for ecosystem management strategies, and fully considers the nonlinear characteristics of natural, human and environmental factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541521B_ABST
    Figure CN120541521B_ABST
Patent Text Reader

Abstract

The application relates to a forest and grass vegetation degradation early warning signal identification method, which comprises the following steps: acquiring geographical data of a research area; obtaining a vegetation significant degradation area according to a vegetation index based on the geographical data; identifying a vegetation negative mutation point based on the vegetation significant degradation area; screening out key influence factors affecting the vegetation negative mutation based on the identified vegetation negative mutation point and the vegetation index; taking the key influence factors as independent variables and the vegetation index as dependent variables to perform model training, calculating a vegetation change amount caused by the key influence factors through a SHAP method, determining a critical point of positive and negative changes of the vegetation change amount as an action threshold of the key influence factors, and obtaining a forest and grass vegetation degradation early warning signal. The application can accurately identify the early warning signal of forest and grass vegetation degradation by comprehensively considering the vegetation degradation factors through the acquired geographical data combined with the vegetation index, is suitable for different ecological systems and different regions, and provides scientific support for the development of management strategies for different ecological systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vegetation monitoring and analysis, and in particular to a method for identifying early warning signals of forest and grassland vegetation degradation. Background Technology

[0002] Despite global greening efforts, vegetation degradation persists in some regions, and even shows a trend of worsening, weakening vegetation's climate mitigation potential and leading to a decline in ecosystem services and human well-being. Existing research indicates the existence of early warning signals for vegetation degradation, which are detected by monitoring whether ecosystems are approaching their steady-state transition thresholds. While numerous studies have been conducted both domestically and internationally on the monitoring and driving mechanisms of vegetation degradation, a comprehensive methodology for identifying early warning signals and conducting risk simulations and early warning systems based on these signals remains lacking.

[0003] Therefore, there is an urgent need to provide a method for identifying early warning signals of forest and grassland vegetation degradation to solve the above-mentioned technical problems. Summary of the Invention

[0004] This application provides a method for identifying early warning signals of forest and grassland vegetation degradation. By acquiring vegetation data, climate data, soil data, geographical environment data, and socioeconomic data of the study area, and combining them with vegetation indices, the key natural, anthropogenic, and environmental factors affecting negative mutations in vegetation are obtained. By comprehensively considering the factors affecting vegetation degradation, the method can accurately identify early warning signals of forest and grassland vegetation degradation. It is applicable to different ecosystems and regions, and provides scientific support for the formulation of management strategies for different ecosystems.

[0005] In one aspect, this application provides a method for identifying early warning signals of forest and grassland vegetation degradation, comprising the following steps: acquiring geographical data of the study area; wherein, the geographical data includes vegetation data, climate data, soil data, geographical environment data, and socioeconomic data; analyzing the geographical data according to vegetation indices to obtain areas of significant vegetation degradation; wherein, the vegetation index is one of the important parameters reflecting crop growth and nutritional information, and the areas of significant vegetation degradation refer to areas where the vegetation index has shown a significant negative trend in the past specific time; identifying vegetation degradation areas based on the BFAST algorithm. Negative mutation points are identified; whereby negative mutation points refer to the time and location of negative mutations in areas of significant vegetation degradation. Based on the identified negative mutation points and the vegetation index, key natural, anthropogenic, and environmental factors influencing negative vegetation mutations are screened. Using the key factors as independent variables and the vegetation index as the dependent variable, a machine learning model is used for model training, and the vegetation change caused by each key factor is calculated using the SHAP method. The critical point of the positive-to-negative transition of the vegetation change is determined as the threshold of the key factors, thus obtaining an early warning signal for forest and grassland vegetation degradation.

[0006] Optionally, based on the geographic data, vegetation index analysis is used to identify areas of significant vegetation degradation, including: masking the study area by applying land use type data from the socio-economic data and desert distribution data from the geographic environment data; further classifying the study area's forest and grassland vegetation types into forest, shrub, and grassland types based on the land use type data; using the Theil-Sen trend algorithm, determining whether the vegetation index of the study area shows a negative trend over a specific historical period based on whether the slope of the vegetation index in the time series is negative; if so, determining whether the negative trend is significant based on the Mann-Kendall method; if significant, defining the corresponding area as a significantly degraded vegetation zone.

[0007] Optionally, based on the significantly degraded vegetation area, the BFAST algorithm is used to identify vegetation negative mutation points. Specifically, this includes: obtaining the temperature and precipitation series of the study area; using the moving average method based on the temperature and precipitation series to exclude vegetation negative mutations caused by climate factor variability and retaining vegetation negative mutations caused by climate change; wherein the climate change is trend-based, and climate variability is interannual fluctuation, excluding negative mutations caused by interannual climate fluctuations and retaining vegetation negative mutations caused by climate trend changes; using monthly NDVI time series data as input, and combining the BFAST model to perform mutation detection pixel by pixel, the time and location of negative mutations in the significantly degraded area are determined as vegetation negative mutation points.

[0008] Optionally, based on the identified vegetation negative mutation points and combined with the vegetation index analysis, key natural, anthropogenic and environmental factors affecting vegetation negative mutations are screened out, including: screening out natural, anthropogenic and environmental factors affecting vegetation negative mutations based on the identified vegetation negative mutation points.

[0009] A correlation heatmap is plotted to calculate the correlation values ​​between the influencing factors and compared with preset values ​​to screen out relevant influencing factors. The relevant influencing factors are used as independent variables and the monthly vegetation index is used as dependent variables. The random forest model is used to train the model for each type of forest and grassland vegetation to obtain the importance threshold of each relevant influencing factor for the occurrence of negative vegetation mutations in different vegetation types. Key influencing factors are screened out according to preset requirements.

[0010] Optionally, a correlation heatmap is plotted to calculate the correlation values ​​between the influencing factors and compared with preset values ​​to filter out relevant influencing factors. This includes: calculating the correlation of the influencing factors based on the heatmap function and plotting the correlation heatmap; obtaining the correlation values ​​between the influencing factors based on the correlation heatmap; comparing the correlation values ​​with the preset values; if the correlation value is greater than the preset value, then the corresponding influencing factor is excluded; otherwise, the influencing factor is retained to obtain relevant influencing factors.

[0011] Optionally, the relevant influencing factors are used as independent variables, and the monthly vegetation index is used as the dependent variable. A random forest model is trained for each forest and grassland vegetation type to obtain the importance threshold of each relevant influencing factor for negative mutations in different vegetation types. Key influencing factors are then selected according to preset requirements, including: for each forest and grassland vegetation type, using the monthly maximum NDVI as the dependent variable and the relevant influencing factors as independent variables to obtain training samples; dividing the training samples into training and validation sets; and using random forest to train and validate models for each forest and grassland vegetation type based on the training and validation sets respectively to obtain... A predictive model is constructed based on the characteristic relationship between the dependent and independent variables. The predicted values ​​obtained from the predictive model are compared with the actual values, and the prediction is reflected by the coefficient of determination (R²) and root mean square error (RMSE). The relative importance of each relevant influencing factor is obtained based on the results of random forest models for different forest and grassland vegetation types. The input of independent variables is controlled according to the importance values ​​of the relevant influencing factors, and the random forest model is trained. The importance thresholds for different vegetation types are determined by comparing the R² and RMSE of the model. When the importance of a relevant influencing factor is greater than the importance threshold, the corresponding relevant influencing factor is retained to obtain the key influencing factors.

[0012] Optionally, the relevant influencing factors are used as independent variables, and the monthly vegetation index is used as the dependent variable. A random forest model is trained for each forest and grassland vegetation type to obtain the importance threshold of each relevant influencing factor for negative mutations in different vegetation types. Key influencing factors are then selected according to preset requirements, including: for each forest and grassland vegetation type, using the monthly maximum NDVI as the dependent variable and the relevant influencing factors as independent variables to obtain training samples; dividing the training samples into training and validation sets; and using random forest to train and validate models for each forest and grassland vegetation type based on the training and validation sets respectively to obtain... A predictive model is constructed based on the characteristic relationship between the dependent and independent variables. The predicted values ​​obtained from the predictive model are compared with the actual values, and the prediction is reflected by the coefficient of determination (R²) and root mean square error (RMSE). The relative importance of each relevant influencing factor is obtained based on the results of random forest models for different forest and grassland vegetation types. The input of independent variables is controlled according to the importance values ​​of the relevant influencing factors, and the random forest model is trained. The importance thresholds for different vegetation types are determined by comparing the R² and RMSE of the model. When the importance of a relevant influencing factor is greater than the importance threshold, the corresponding relevant influencing factor is retained to obtain the key influencing factors.

[0013] Optionally, based on the identified vegetation negative mutation points and the vegetation index analysis, key natural, anthropogenic, and environmental influencing factors affecting vegetation negative mutations are screened out. This also includes: training a model for each vegetation negative mutation point using a machine learning model based on the key influencing factors; based on the lag effect of climate factors, during model training, other key influencing factors are kept constant using a controlled variable method, and climate influencing factors at different times are used as input to obtain predicted values; comparing the R² and RMSE of each model to determine the best-fit climate lag and cumulative time; wherein, the lag effect of climate factors refers to the climate factors at a specific time prior to the occurrence of vegetation negative mutations being most affected by climate factors.

[0014] Optionally, the key influencing factors are used as independent variables, and the vegetation index is used as the dependent variable. A machine learning model is used for model training, and the vegetation change caused by each key influencing factor is calculated using the SHAP method. The critical point of the positive-to-negative transition of vegetation change is determined as the effect threshold of the key influencing factors, thus obtaining an early warning signal for forest and grassland vegetation degradation. This includes: using the key influencing factors considering the lag effect of climate factors as independent variables and the annual maximum composite vegetation index NDVI as the dependent variable, training the model using a random forest model, and dividing the training set and validation set according to a specific ratio; decomposing the prediction results of the random forest model into the sum of the contributions of each key influencing factor, and calculating the vegetation change caused by each key influencing factor using the SHAP method; using each key influencing factor as the independent variable and the SHAP value corresponding to each key influencing factor as the dependent variable, fitting the model using GAM; and calculating the independent variable value corresponding to the dependent variable being 0 through reverse calculation to obtain the effect threshold of each key influencing factor, which serves as an early warning signal for forest and grassland vegetation degradation.

[0015] Optionally, a method for identifying early warning signals of forest and grassland vegetation degradation further includes: acquiring future scenario data; wherein the future scenario data includes climate data under various future scenarios provided by the Coupled Model Intercomparison Project Phase 6 (CMIP6), as well as socioeconomic data calculated based on population density, nighttime light index, and GDP growth rate, and unchanged soil data and geographical environment data; based on the future scenario data, the effect values ​​of each key influencing factor are obtained, and compared with the corresponding early warning signals to determine the degree of risk of forest and grassland vegetation degradation under the future scenarios, so as to realize risk warning of potential vegetation degradation areas under the future scenarios.

[0016] Optionally, based on future scenario data, the impact values ​​of each key influencing factor are obtained and compared with the corresponding early warning signals to determine the risk level of forest and grassland vegetation degradation under future scenarios, so as to realize risk warning of potential vegetation degradation areas under future scenarios. This also includes: extracting forest, shrub, grassland, and desert areas based on land use type data under different future scenarios; combining early warning signals of key influencing factors of different forest and grassland vegetation to calculate the risk level of forest and grassland vegetation degradation under future scenarios and obtaining the mean value of the degradation risk level for each type of vegetation; based on the mean value of the degradation risk level for each type of vegetation, the degradation risk level is divided into four levels: no degradation risk, low degradation risk, medium degradation risk, and high degradation risk, according to the natural breakpoint method; through statistical analysis of the risk level distribution of degradation warning areas for each type of forest and grassland vegetation under different future scenarios, the spatiotemporal distribution characteristics of degradation risk areas are obtained, and the dynamic change trend of the area of ​​degradation warning areas at each level is quantified.

[0017] Secondly, this application provides an early warning signal identification system for forest and grassland vegetation degradation. The system includes: an acquisition module for acquiring geographic data of the study area; wherein the geographic data includes vegetation data, climate data, soil data, geographic environmental data, and socio-economic data.

[0018] The module for identifying significantly degraded vegetation areas is divided into several modules. The first module analyzes vegetation indices based on geographic data to identify areas of significant vegetation degradation. Vegetation indices are important parameters reflecting crop growth and nutrient information; significantly degraded vegetation areas refer to regions where vegetation indices show a significant negative trend over a specific period. The second module identifies negative mutation points in vegetation based on these areas using the BFAST algorithm. Negative mutation points refer to the time and location of negative mutations within these areas. The third module identifies key influencing factors, using the identified negative mutation points and vegetation indices to screen for key natural, anthropogenic, and environmental factors influencing negative vegetation mutations. The fourth module identifies early warning signals, using key influencing factors as independent variables and vegetation indices as dependent variables. It trains a machine learning model and calculates the vegetation change caused by each key influencing factor using the SHAP method. The critical point of positive-to-negative transition in vegetation change is determined as the threshold for the key influencing factors, thus generating early warning signals for forest and grassland vegetation degradation.

[0019] Thirdly, this application also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor implements the steps of the method described above when executing the computer program.

[0020] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the method described above.

[0021] This application has at least the following advantages:

[0022] The above steps mainly involve acquiring vegetation, climate, soil, geographical environment, and socioeconomic data for the study area. Based on the acquired data, vegetation index data is used to identify areas of significant forest and grassland vegetation degradation, thereby obtaining negative mutation points. Based on these negative mutation points, key influencing factors affecting vegetation degradation are determined. Based on these key influencing factors, the threshold values ​​of each key influencing factor for negative mutations in forest and grassland vegetation are precisely determined, serving as early warning signals for forest and grassland vegetation degradation. This process fully considers the nonlinear characteristics of the effects of various environmental factors, including natural, anthropogenic, and environmental factors. It can accurately identify early warning signals for forest and grassland vegetation degradation and can be used in different ecosystems and regions, providing scientific support for the formulation of different ecosystem management strategies. Attached Figure Description

[0023] Figure 1 This is an example of an application environment diagram showing the method for identifying early warning signals of forest and grassland vegetation degradation;

[0024] Figure 2 This is a flowchart illustrating the steps of a method for identifying early warning signals of forest and grassland degradation in one embodiment.

[0025] Figure 3 Here is a flowchart illustrating the process of identifying significant negative mutation points in one embodiment;

[0026] Figure 4 This is a schematic diagram illustrating the process of screening key influencing factors in one embodiment;

[0027] Figure 5 This is a bar chart showing the importance of the characteristics of influencing factors related to different types of forest and grassland vegetation in one embodiment;

[0028] Figure 6 Here is a line graph showing the performance of the random forest model under different importance thresholds of relevant influencing factors in one embodiment;

[0029] Figure 7 This is a schematic diagram of a standard RNN structure in one embodiment;

[0030] Figure 8 This is a schematic diagram illustrating the principle of LSTM in one embodiment;

[0031] Figure 9 This is a schematic diagram illustrating the impact threshold of key influencing factors in one embodiment;

[0032] Figure 10 This is a block diagram illustrating the structure of an early warning signal identification system for forest and grassland vegetation degradation in one embodiment.

[0033] Figure 11 This is a schematic structural diagram of a computer device in one embodiment. Detailed Implementation

[0034] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present application and are not intended to limit the scope of the present application.

[0035] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application. As used herein, unless the context clearly indicates otherwise, the singular form also means...

[0036] Figures include plural forms. Furthermore, it should be understood that when used in this specification, the words “comprising” and / or “including” indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0037] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the embodiments of this application to facilitate a better understanding of the application. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The reference numerals in the following embodiments are for descriptive convenience and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.

[0038] For ease of understanding, the system to which this application applies will first be described. This application provides a method for identifying early warning signals of forest and grassland vegetation degradation, which can be applied to, for example... Figure 1 The system architecture shown includes a user-space file server 103 and a terminal device 101. The terminal device 101 communicates with the user-space file server 103 via a network. The user-space file server 103 can be a file server based on the NFSv3 / v4 protocol, running in a Linux environment. NFS (Network File System) is a network abstraction on top of a file system, allowing remote clients running on the terminal device 101 to access the file system over the network in a manner similar to a local file system. The terminal device 101 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, etc. The user-space file server 103 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0039] Figure 2 This application provides a flowchart illustrating a method for identifying early warning signals of forest and grassland vegetation degradation, which may include the following steps:

[0040] S201. Obtain geographical data of the study area; among which...

[0041] Geographic data includes vegetation data, climate data, soil data, geographic environment data, and socioeconomic data;

[0042] S202. Based on geographic data, vegetation index analysis is used to identify areas of significant vegetation degradation. Among them, vegetation index is one of the important parameters reflecting crop growth and nutritional information. Areas of significant vegetation degradation refer to areas where the vegetation index has shown a significant negative trend in the past.

[0043] S203. Identify negative mutation points of vegetation based on the BFAST algorithm in areas of significant vegetation degradation; where negative mutation points of vegetation refer to the time and location of negative mutations in areas of significant vegetation degradation.

[0044] S204. Based on the identified vegetation negative mutation points and vegetation index analysis, key natural, anthropogenic and environmental factors affecting vegetation negative mutations were screened out.

[0045] S205. Using key influencing factors as independent variables and vegetation index as dependent variable, a machine learning model is used for model training. The vegetation change caused by each key influencing factor is calculated using the SHAP method. The critical point of positive to negative change in vegetation change is determined as the threshold of the key influencing factor, thus obtaining an early warning signal for forest and grassland vegetation degradation.

[0046] In this embodiment, the acquired geographical data is analyzed and processed in conjunction with vegetation indices to identify areas of significant vegetation degradation. Then, the vegetation index analysis is used to screen out key natural, anthropogenic, and environmental factors that influence negative vegetation mutations. This fully considers the nonlinear characteristics of the effects of various environmental factors, including natural, anthropogenic, and environmental factors, and accurately identifies early warning signals of forest and grassland vegetation degradation. It provides early warning of vegetation degradation risks under different scenarios and is applicable to different ecosystems and regions. At the same time, it provides scientific support for the formulation of different ecosystem management strategies.

[0047] The following is a detailed explanation of each step:

[0048] Please refer to Figure 2 As shown, step S201: Obtain geographical data of the study area.

[0049] In this embodiment, it should be noted that the geographic data mainly includes vegetation data, climate data, soil data, geographic environment data, and socioeconomic data. According to the requirements of the calculation methods and model construction, relevant data products and statistical information of the study area are collected, and all data are converted to raster format. Specifically, the vegetation data comes from relevant vegetation index products such as MODIS (Moderate-resolution Imaging Spectroradiometer) and GIMMS (Global Inventory Modeling and Mapping Studies), including NDVI (Normalized Vegetation Index) and LAI (Leaf Area Index); the climate data includes future and historical climate data, such as temperature, precipitation, radiation, evapotranspiration, and water vapor pressure; the soil data consists of soil physicochemical properties, such as soil moisture and soil sand content; the geographic environment data includes topographic data and desert distribution data; and the socioeconomic data includes population density, land use type, nighttime light, and livestock numbers. The land use type data and desert distribution data primarily provide the data foundation for subsequent analysis of negative vegetation trend changes.

[0050] Please continue to refer to Figure 2 , Figure 3 As shown, step S202, based on geographic data and vegetation index analysis, identifies areas of significant vegetation degradation. This includes: using land use type data from socioeconomic data and desert distribution data in the geographic environment, masking is performed on the desert distribution data to remove native desert areas from the study area; based on the study area with native desert areas removed, the forest and grassland vegetation types in the study area are further divided into forest, shrub, and grassland types according to land use type data; based on the Theil-Sen trend algorithm, whether the vegetation index of the study area shows a negative trend in a specific historical period is determined by whether the slope of the vegetation index in the time series is negative; if so, the Mann-Kendall method is used to determine whether the negative trend is significant; if so, the corresponding area is defined as a significantly degraded vegetation area.

[0051] In this embodiment, it should be noted that vegetation index is one of the important parameters reflecting crop growth and nutritional information. Significantly degraded vegetation areas refer to regions where the vegetation index has shown a significant negative trend over a specific period. Furthermore, since desert areas also have some vegetation distribution, 1:100,000 desert distribution data is used to mask the original desert areas based on land use type data, resulting in desert vegetation distribution areas. The data is obtained by identifying vegetation degradation grid-by-grid at a regional scale based on remote sensing data, resulting in spatialized data. Theil-Sen trend analysis is used to define regions where the vegetation index has shown a significant negative trend over the past few decades as significantly degraded vegetation areas. The Mann-Kendall test is then used to analyze the trend significance of the vegetation index NDVI to determine the significantly degraded vegetation areas.

[0052] Specifically, the Theil-Sen trend method uses the median of all slopes of the time-series vegetation index NDVI as the overall trend of the series, and the calculation formula is as follows:

[0053]

[0054] In the formula, β is the median slope of all data pairs, and Median is the median calculation function. j and x i Let represent the values ​​of the j-th and i-th terms in the vegetation index time series, respectively. When β > 0, it indicates an upward trend; otherwise, it indicates a downward trend.

[0055] The Mann-Kendall method is used to determine the significance of trends. The specific calculation framework is as follows:

[0056] For the vegetation index sequence X i =(x1,x2,x3…x n When j > i, compare x. j With x i Let the magnitudes of the numbers be related and let S be the magnitude of this relationship. The test statistic S can be calculated using the following formula:

[0057]

[0058]

[0059] When the time series length of vegetation index is n≥10, the statistic S approximately follows a standard normal distribution. S can be standardized to obtain the test statistic Z. The formula for calculating the Z value is as follows:

[0060]

[0061] In the formula, n is the length of the vegetation index time series, and m is the number of repeated data sets in the series; ti Z is the number of repeated data in the i-th repeated data group. At a significance level of α = 0.05, Z... 1-α / 2 = 1.96. When |Z| < 1.96, the trend is not significant; when |Z| ≥ 1.96, the trend is significant. When the vegetation index sequence shows a significant downward trend, the corresponding area is defined as a significantly degraded area.

[0062] Please continue to refer to Figure 2 , Figure 3 As shown, step S203 involves identifying negative mutation points in vegetation based on the BFAST algorithm in areas of significant vegetation degradation. This includes: obtaining temperature and precipitation sequences for the study area; using the moving average method based on the temperature and precipitation sequences to exclude negative mutations in vegetation caused by variability in climate factors and retaining mutations caused by climate change; wherein, climate change is a trend, and climate variability is an interannual fluctuation, excluding negative mutations caused by interannual climate fluctuations and retaining negative mutations caused by trend changes in climate; using monthly NDVI time series data as input, combining the BFAST model to perform mutation checks pixel by pixel, and determining the time and location of negative mutations in significantly degraded areas as negative mutation points in vegetation.

[0063] In this embodiment, it should be noted that negative mutation points in vegetation refer to the time and location of negative mutations occurring in areas of significant vegetation degradation. Temperature and precipitation data can be downloaded from the National Science Data Center. The BFAST method is a time series decomposition model capable of capturing points where significant abrupt changes in the magnitude or direction of vegetation changes occur under external disturbances, and is widely used in nonlinear dynamic monitoring of vegetation. For the NDVI vegetation index sequence, the influence of seasonal variations on vegetation dynamics is first eliminated by separating the seasonal component, and vegetation mutations are identified only in the trend component, T. t and seasonal component S t They are represented as follows:

[0064] T t =a j +β j t,(t j-1 <t<t j ,j=1,...,m) (6),

[0065]

[0066] In the formula, for t1, ..., tm breakpoints, t is the location of the mutation point, and a j and β j These are the intercepts and slopes of the linear equations on either side of the mutation point, γ. j and δ jLet f represent the amplitude and phase of the piecewise equation, respectively. k is a positive integer, and f represents the frequency. The calculation is performed using 12 monthly NDVI value sequences per year, with f = 12.

[0067] Then, the trend and seasonal components obtained from the decomposition are used to identify breakpoints. If a continuous vegetation change trend is lacking, it is identified as a mutation point. Existing research shows that this method is overly sensitive to changes in dryland vegetation, and is significantly affected by interannual variations in climatic factors such as precipitation. Therefore, while using the moving average method to eliminate the influence of climatic factor variability, mutations caused by climate change are also considered. Based on areas of significant vegetation degradation in the study area, the timing and location of all negative mutations are further identified to determine negative vegetation mutation points. The moving average method, also known as the moving average method, is based on the simple average method. It calculates a moving average by sequentially adding and subtracting old and new data to eliminate random fluctuations, identify trends, and make predictions accordingly. The moving average method is a trend extrapolation technique. In essence, it involves curve fitting to a data sequence with a clear load change trend, and then using the new curve to predict the value at a future point. Here, the moving average method highlights low-frequency trends by eliminating high-frequency fluctuations, thus eliminating short-term variability in climatic factors such as temperature and precipitation, and highlighting long-term change patterns.

[0068] Specifically, in one example, such as Figure 3 As shown, for abrupt changes identified using the BFAST method, climate variability effects are further excluded. First, the temperature and precipitation series are processed using a moving average. The residual trend method establishes a linear relationship between temperature, precipitation, and vegetation indices, fitting a predicted NDVI value under specific climate conditions to represent the impact of natural factors on vegetation. The residuals between the actual and predicted NDVI values ​​are then calculated and trend analysis performed to represent the impact of human activities on vegetation. Based on this, the effects of natural factors are further separated. Climate variability effects are isolated using the temperature and precipitation data after moving average processing, and the long-term trend of vegetation change caused by climate variability is calculated. If the variability effect trend is not significant, the abrupt change point is considered. Conversely, if the variability effect of the abrupt change point identified by the BFAST method is significant, the abrupt change point is excluded.

[0069] Reference Figure 2 , Figure 4 As shown, in step S204, based on the identified vegetation negative mutation points and vegetation index analysis, key natural, anthropogenic, and environmental influencing factors affecting vegetation negative mutations are screened out, including:

[0070] Step S2041: Based on the identified vegetation negative mutation points, screen out the natural, anthropogenic and environmental factors that affect vegetation negative mutations;

[0071] In this embodiment, it should be noted that, specifically, to increase the reliability and robustness of the results, abrupt changes in the first and last three years of the vegetation index time series were removed, and the remaining negative abrupt changes were used for subsequent model training. For negative abrupt changes in forest and grassland vegetation, natural, anthropogenic, and environmental influencing factors affecting negative vegetation mutations were screened based on historical records. Table 1 lists the influencing factors of negative vegetation mutations:

[0072] Table 1 Factors affecting negative mutation points in vegetation

[0073]

[0074] Step S2042: Draw a correlation heatmap, calculate the correlation values ​​between influencing factors, and compare them with preset values ​​to screen out relevant influencing factors;

[0075] In this embodiment, it should be noted that the heatmap function can be used to calculate factor correlations and draw a correlation heatmap. The correlation values ​​between influencing factors are obtained from the heatmap; these values ​​are compared with preset values; if the correlation value is greater than the preset value, the corresponding influencing factor is excluded; otherwise, the influencing factor is retained to obtain the relevant influencing factors.

[0076] Specifically, in one example, a correlation heatmap is drawn for different forest and grassland vegetation types to calculate the correlation between factors. The correlation values ​​are compared with a preset value of 0.7, and variables with a correlation greater than 0.7 are excluded to reduce the impact of multicollinearity. Multicollinearity refers to a relationship where there is some correlation or high correlation among independent variables, where one independent variable can be explained by a linear combination of other independent variables. Therefore, variables with a correlation greater than 0.7 are excluded to reduce the impact of multicollinearity.

[0077] Reference Figure 4 , Figure 5 Step S2043: Using relevant influencing factors as independent variables and monthly vegetation index as dependent variables, train models for each type of forest and grassland vegetation using random forest to obtain the importance threshold of each relevant influencing factor for negative mutations in different vegetation types, and select key influencing factors according to preset requirements.

[0078] In this embodiment, it should be noted that random forests are widely used in prediction and classification problems. Compared with traditional decision tree methods, they are easier to mine data rules and have stronger generalization ability. They also have advantages such as non-parametric nature, low overfitting risk, high accuracy, and the ability to handle missing data and noise. For each forest and grassland vegetation type, the NDVI synthesized from the monthly maximum value is used as the dependent variable, and relevant influencing factors are used as independent variables to obtain training samples. The training samples are divided into training sets and validation sets. Based on the training sets and validation sets, the random forest model is trained and validated for each forest and grassland vegetation type to obtain the feature relationship between the dependent and independent variables and construct a prediction model. The predicted values ​​obtained by the prediction model are compared with the actual values, and the coefficient of determination R is used to determine the model. 2 The coefficient of determination and root mean square error (RMSE) are used to reflect the prediction results; based on the results of random forest models for different forest and grassland vegetation types, the relative importance values ​​of each relevant influencing factor are obtained; based on the magnitude of the characteristic importance values ​​of the relevant influencing factors, the input of independent variables is controlled, the random forest model is trained, and the R-squared values ​​of the models are compared. 2 The importance thresholds of different vegetation types are determined using RMSE; when the importance of a factor is greater than a specific threshold, the corresponding related influencing factors are taken into consideration to obtain the key influencing factors.

[0079] Specifically, using the selected negative mutation points of vegetation as training samples, relevant influencing factors and corresponding NDVIs were designed for each negative mutation point. A random forest machine learning method was used to build a model for each forest and grassland vegetation type, calculating the importance of each relevant influencing factor. According to the random forest algorithm, the original training dataset is denoted as T = {(x1, y1), (x2, y2), ..., (x...} m y m )}, where x is the set of independent variables, including natural, anthropogenic, and environmental factors; y is the dependent variable, i.e., NDVI. Then, the Bootstrap method is used to randomly sample from the dataset T to obtain N subsets y = {c1, c2, ... c...}. n Random forest trees are used to build regression models. It is a set of N subsets. Among them, a functional relationship between the NDVI of negative mutation points in vegetation and various influencing factors x is established using a random forest model, where f(x) represents the prediction result of the random forest regression model. N (x) represents the prediction made by a single decision tree. During model operation, the relative importance of each factor is calculated based on the Gini coefficient, resulting in the importance value of each relevant influencing factor. The higher the relative importance value, the greater the contribution of that factor to the model.

[0080] In one example, for each forest and grassland vegetation type, the monthly maximum synthesized NDVI is used as the dependent variable, and relevant influencing factors after eliminating correlations—that is, different natural, anthropogenic, and environmental factors—are used as independent variables. 70% of the samples are selected as the training set, and 30% as the validation set for training and parameter tuning. The characteristic relationships between the variables are obtained, and the coefficient of determination R is used to determine these relationships. 2 The root mean square error (RMSE) is used to reflect the prediction performance. 2 R-squared is used to measure the degree of linearity between two variables and to verify the relationship between the data and the simulated data. 2 The closer the value is to 1, the better the model accuracy. The root mean square error (RMSE) is the difference between the actual and predicted values. A smaller RMSE indicates that the predicted value is closer to the actual value, and the smaller the model error. It is described by the formula:

[0081]

[0082] In the formula, NDVI a These are measured values, NDVI m NDVI is the average of the measured values. p The values ​​are predicted, i = 1, 2, 3, ..., N, where i is the number of samples and N is the total number of samples. Based on the results of random forest models for different forest and grassland vegetation types, a ranking diagram of the relative importance of each factor is obtained.

[0083] Reference Figure 5 , Figure 6 As shown, based on the importance values ​​of relevant factor features, the input of independent variables is controlled, and a random forest model is trained. The R-squared of the model is then compared. 2 RMSE is used to determine the feature importance thresholds for different vegetation types, further optimizing the model's independent variable inputs and preventing overfitting while ensuring model training accuracy. When a factor's importance exceeds a specific threshold set according to preset requirements, the corresponding factor is considered to obtain key influencing factors for subsequent model training.

[0084] Table 2 Importance thresholds and independent variables of different forest and grassland vegetation characteristics

[0085]

[0086] In one example, there are 12 relevant influencing factors for random forests. First, they are ranked according to importance; the higher the importance value, the more important the factor. Factors with the lowest importance have little impact on the model; if they significantly affect the model's performance (R²), then... 2 If the RMSE has little impact, it can be excluded. Specifically, this can be achieved by controlling the input of independent variables and removing factors with importance below a certain value to obtain the desired result. Figure 6The results show that for the random forest model, when the importance of relevant influencing factors is less than 0.01, it has little impact on model performance; however, when it exceeds 0.01, it will cause R... 2 The rapid decrease in importance and the rapid increase in RMSE mean that the feature importance threshold for forests is 0.01. Removing relevant influencing factors with importance values ​​less than 0.01 yields key influencing factors, which saves training time and improves model training performance.

[0087] Reference Figure 2 , Figure 4 As shown, step S204, based on the identified vegetation negative mutation points and vegetation index analysis, screens out key natural, anthropogenic, and environmental influencing factors affecting vegetation negative mutations. This also includes: training a machine learning model for each vegetation negative mutation point based on these key influencing factors, considering the lag cumulative effect of each climate factor during model training, and finally comparing the R-values ​​of the machine learning models. 2 RMSE and CRSE determine the best-fit climate lag and cumulative time.

[0088] In this embodiment, it should be noted that the machine learning model can include three types: Random Forest, RNN (Recurrent Neural Network), and LSTM (Long Short-Term Memory). RNN is a deep neural network used for recognizing and processing time-series data. The core structure of an RNN consists of an input layer, hidden layers, and an output layer, with layers connected by weights. The core algorithm is to calculate the gradients of each parameter. The output of a node in the hidden layer at time t not only enters the output layer at time t but also becomes the input of the hidden layer at time t+1. In this way, RNN achieves temporal continuity within the network, giving it a significant advantage in processing time-series data. Figure 7 This is a three-layer RNN model, where x represents the input layer, h represents the hidden layer, o represents the output layer, L represents the loss function, and y represents the true value, i.e., the expected output value of the sample. U, V, and W represent the shared weights between the input layer and the hidden layer, the hidden layer and the output layer, and the hidden layer and the hidden layer, respectively. The output value of the hidden layer at time t and the model's predicted output value can be expressed by the following formula:

[0089] h t =f(Ux t +Wh t-1 +b) (10)

[0090]

[0091] In the formula, f and g are activation functions, f is the ReLU function, g is the softmax function, b and c are bias values, and h is the activation function. t and h t-1 These represent the values ​​of the hidden layer at the corresponding time points. This represents the calculated output value of the RNN. During the training process of a recurrent neural network, a loss function is used to represent the difference between the predicted and actual values. The specific function expression is:

[0092]

[0093] In the formula, n represents the number of selected nodes, L represents the loss function, and this study chooses the mean squared error loss function. t This indicates the expected output value.

[0094] LSTM is an improved form of RNN. In RNN, the gradient gradually approaches zero during the accumulation process over time. LSTM solves the "vanishing gradient" problem by introducing three mechanisms: a forget gate, an input gate, and an output gate. This makes LSTM more advantageous in processing long-term data series and further improves model accuracy. The principle of LSTM at a certain time t is as follows: Figure 8 As shown. The upper C represents the cell state introduced by LSTM, the lower h represents the hidden layer value, the red rectangle represents the learned neural network layer, the green circles represent mathematical operators (addition and multiplication), and tanh and σ represent activation functions with output values ​​between 0 and 1. 0 indicates rejection of all signals, 1 indicates acceptance of all signals, and numbers between 0 and 1 indicate how much signal is allowed. The forget gate controls whether information is retained; the amount of information forgotten depends on the new input X. t and the h passed from the previous layer t-1 The output is determined by the sigmoid function, and its formula is as follows:

[0095] f t =σ(W f *[h t-1 X t +b f (13)

[0096] In the formula, f t W represents the information retained after the information input to the previous layer of the network passes through the forget gate. f With b fThe weight information is then used. After the forget gate discards information, the input gate needs to replenish some information. First, the sigmoid function is used to determine which information values ​​from the previous layer need to be input and which need to be discarded. Then, a new activation function, tanh, is used to form a new output, which is multiplied by the result of the sigmoid function to update the input and is then compared with the result f from the forget gate. t The products are multiplied to form a new unit state. The formula is shown below:

[0097] i t =σ(W i *[h t-1 ,X t ]+b i (14),

[0098]

[0099] In the formula, i t Information preserved after the sigmoid function For the newly added input information, C t-1 For the old unit cell state information, C t This provides information for new unit cells.

[0100] When dealing with output values ​​and cell states, the sigmoid function is first run to determine the specific part of the output cell state. Then, the tanh function is used to manipulate the cell state, multiplying it by the output of the sigmoid function. Finally, the output is performed. The output gate formula is as follows:

[0101] o t =σ(W o *[h t-1 X t ]+b o (17),

[0102] h t =o t *tanh(C t (18),

[0103] In the formula, o t To output information about the state of the unit cell, h t This is the information part of the output.

[0104] The performance of machine learning models largely depends on their parameter configuration. To improve the simulation accuracy and applicability of the model and ensure the optimal combination of model parameters on the validation set, parameter optimization is essential. Traditional parameter optimization algorithms suffer from long computation times and susceptibility to local optima, while Bayesian optimization can address these issues to some extent. Bayesian optimization uses a probabilistic surrogate model to fit the true objective function and constructs a sampling function to measure the exploration level of each point, selecting reliable evaluation points for evaluation, avoiding unnecessary sampling, and ultimately obtaining the global optimum. Bayesian hyperparameter optimization methods are used to select and tune the parameters of random forests, RNNs, and LSTMs, with the optimization objective being to minimize the mean squared error loss function of the model. Referring to relevant research, the domain space for Bayesian optimization of random forests, RNNs, and LSTM models, i.e., the range of values ​​for each hyperparameter to be searched, is determined, as shown in Table 3.

[0105] Table 3 Bayesian optimization domain space for three machine learning models

[0106]

[0107] Furthermore, RNNs and LSTMs are often used in conjunction with the Adam (Adaptive Moment Estimation) optimization algorithm. This application utilizes the Adam optimizer to further optimize the algorithm process. The Adam optimization algorithm combines the advantages of adaptive gradient algorithms and root mean square propagation algorithms, assigning a separate learning rate to each weight and automatically adjusting it in a timely manner. This makes the model operation smoother and improves the training effect and generalization ability of the neural network model. By verifying and comparing the training effects of the three models, the reliability of the results is increased. If all three methods show the same cumulative lag effect, the credibility is high. Moreover, deep learning methods (RNN, LSTM) can consider temporal information, resulting in higher accuracy.

[0108] Specifically, during model training, other key influencing factors are kept constant by controlling for variables, and climate influencing factors at different times are used as inputs to obtain predicted values; then, the R-values ​​of each model are compared. 2 And RMSE, determine the best-fit climate lag and cumulative time. When R 2 The larger the value and the smaller the RMSE, the better the training effect, thus determining the corresponding lag cumulative effect of climate factors.

[0109] In one example, key influencing factors of different forest and grassland vegetation were used as independent variables, and NDVI as the dependent variable. Random forest, RNN, and LSTM were employed for model training at each vegetation negative mutation point, considering the lagged cumulative effects of various climate factors during model training. Based on relevant literature, a 1-3 month lagged cumulative effect was considered for temperature, precipitation, radiation, and potential evapotranspiration, while a 1-12 month lagged cumulative effect was considered for SPEI (Standardized Precipitation Evapotranspiration Index). The training and validation sets were divided in a 7:3 ratio, with data from 2000-2015 used for model training and data from 2016-2022 used for validation. Finally, the R-values ​​of the three machine learning models were compared. 2 RMSE and lag time determine the best fit.

[0110] Reference Figure 2 , Figure 9 As shown, step S205 involves using key influencing factors as independent variables and the annual vegetation index as the dependent variable, training the model using a machine learning model, calculating the vegetation change caused by each key influencing factor using the SHAP (SHapley Additive exPlanations) method, and determining the critical point of positive to negative change in vegetation change as the threshold of the key influencing factor's effect, thus obtaining an early warning signal for forest and grassland vegetation degradation. This includes: using key influencing factors considering the lag effect of climate factors as independent variables and the annual maximum composite vegetation index NDVI as the dependent variable, training the model using a random forest model, and dividing the training set and validation set according to a specific ratio; decomposing the model's prediction results into the sum of the contributions of each key influencing factor, and calculating the vegetation change caused by each key influencing factor using the SHAP method; using each key influencing factor as an independent variable and the corresponding SHAP value of each key influencing factor as the dependent variable, and fitting the model using GAM; and calculating the independent variable value corresponding to the dependent variable being 0 through reverse calculation to obtain the threshold of each key influencing factor, which serves as an early warning signal for forest and grassland vegetation degradation.

[0111] In this embodiment, it should be noted that when the ecosystem state approaches the steady-state transition threshold, its dynamic changes exhibit some common characteristics near the mutation point. These common characteristics constitute early warning signals, which can be characterized by a series of general indicators. To eliminate the influence of seasonal variations of influencing factors on forest and grassland vegetation, the growing season average of each climatic factor and the annual average of other influencing factors are further calculated. Annual data are used to identify and analyze early warning signals. The key influencing factors of different negative mutation points in forest and grassland vegetation, already identified, are used as independent variables, and the annual maximum composite vegetation index (NDVI) is used as the dependent variable. A random forest model is used for training, with the training and validation sets divided in a 7:3 ratio. The vegetation change caused by each influencing factor is calculated using the SHAP method. Furthermore, a generalized additive model is used to fit the nonlinear functional relationship between each influencing factor and the resulting vegetation change. The critical point of the positive-to-negative transition of the corresponding SHAP value is determined as the threshold of the influencing factor's effect, which can serve as an early warning signal for forest and grassland vegetation degradation.

[0112] SHAP is a game theory-based model interpretation method. In this application, an additive interpretation model is constructed to calculate the marginal contribution of each key natural, anthropogenic, and environmental influencing factor to the model, allocating the contribution value of each factor. SHAP analysis measures the magnitude of each influencing factor's impact on vegetation dynamics by calculating SHAP values. The results are additive, decomposing the prediction results of the machine learning model into the sum of the contributions of each factor and providing an intuitive interpretation. Specifically, when the SHAP value of a key influencing factor is greater than 0, it indicates that the key influencing factor has a positive impact on the predictor variable; when it is less than 0, it has a negative impact. When the SHAP value equals 0, it is the critical point where the direction of the key influencing factor's effect changes. The larger the absolute value of the SHAP value, the higher the importance of the factor and the greater its impact on the predictor variable. The SHAP value is calculated and exported using the shap package in Python, and the calculation formula is as follows:

[0113]

[0114] In the formula, φ i Let be the SHAP value of a key influencing factor i, F be the set of key influencing factors in the model, S be a subset of factors excluding factor i, |F| be the total number of all factors, |S| be the number of factors in subset S, v(S) be the contribution of factors in subset S to the model prediction, and ν(S∪{i})-v(S) be the marginal contribution of factor i to the model prediction. The weights represent the contribution of factor i to all feature combinations.

[0115] Using key influencing factors as independent variables and their corresponding SAP values ​​as dependent variables, a Generalized Additive Model (GAM) was employed for modeling. Related calculations were performed using the pygam package in Python. GAM is a non-parametric, highly flexible statistical model used to model non-linear relationships between variables. It does not require pre-specifying the function form, uses a smoothing function to establish regression relationships, and represents the contribution of each influencing factor as the sum of the smoothing functions. This allows for better fitting of complex data structures. The model fit was assessed using the significance level p-value and the coefficient of determination R-value. 2 This is reflected in the equation. Under the premise of an equation tolerance of 0.001, the independent variable value corresponding to the dependent variable being 0 is solved by reverse calculation, which is the threshold of each influencing factor. Combined with the action mechanism of each factor, early warning signals of different forest and grassland vegetation degradation are identified and determined.

[0116] like Figure 9 As shown, for four forest and grassland vegetation types—forest, shrubland, grassland, and desert—the environmental suitability of different vegetation types was fully considered. Points with significant negative mutations were selected as sample points for analysis. Based on the fitting results of the Gaussian Aspect-Oriented Model (GAM) for different forest and grassland vegetation types, the threshold values ​​of influencing factors on forest and grassland vegetation in the study area were compiled as early warning signals for the degradation of various vegetation types. The influencing factors and threshold values ​​for different forest and grassland vegetation types vary to some extent, making them more accurate as early warning signals for vegetation degradation. When the actual value of the corresponding environmental factor exceeds its threshold, the early warning mechanism is triggered. The more influencing factors in a certain area exceed their thresholds, the higher the risk of degradation in that area.

[0117] In one example, such as Figure 9 b. The SHAP value of temperature, a key influencing factor, decreases as the vertical axis (temperature) increases (the horizontal axis). When the temperature SHAP value is greater than 0, it indicates that temperature promotes vegetation growth; when the temperature SHAP value is less than 0, it indicates that temperature inhibits vegetation growth, leading to vegetation degradation. Therefore, when the dependent variable SHAP value is 0, it is the inflection point where the influence of temperature on vegetation shifts from promoting to inhibiting growth. Exceeding this threshold will lead to vegetation degradation. The same logic applies to other influencing factors.

[0118] After obtaining early warning signals of forest and grassland vegetation degradation, the process also includes step S206: acquiring future scenario data. The future scenario data includes climate data under various future scenarios provided by the Coupled Model Intercomparison Project Phase 6 (CMIP6), socioeconomic data calculated based on population density, nighttime light index, and GDP growth rate, as well as unchanged soil data and geographic environment data. Based on the future scenario data, the impact values ​​of each key influencing factor are obtained and compared with the corresponding early warning signals of forest and grassland vegetation to determine the degree of risk of forest and grassland vegetation degradation under the future scenarios, so as to realize the risk warning of potential vegetation degradation areas under the future scenarios.

[0119] In this embodiment, it should be noted that the future scenario data includes climate data, soil data, geographic environment data, and socioeconomic data under various future scenarios provided by the Coupled Model Intercomparison Project Phase 6 (CMIP6). Since the risk warning of forest and grassland vegetation degradation under future scenarios is based on calculations of key influencing factors and does not utilize vegetation data, vegetation data is not required for future scenario data. Specifically, based on the CMIP6 combined scenarios, three scenarios—SSP1-RCP2.6, SSP2-RCP4.5, and SSP5-RCP8.5—are selected to form a parameterized scheme for future scenarios, providing early warnings of forest and grassland vegetation degradation risks under different scenarios. Seven climate models—BCC-CSM2-MR, CAS-ESM2-0, CMCC-ESM2, GFDL-ESM4, INM-CM4-8, INM-CM5-0, and MRI-ESM2-0—are selected, and their monthly average precipitation, temperature, wind speed, surface radiation, and monthly maximum and minimum temperature data are averaged. To further improve the accuracy of various climate data under different scenarios, we used the global climate model dataset released by CMIP6 and the global high-resolution climate dataset released by WorldClim to downscale the future climate scenario data using the Delta method (Delta). We calculated potential evapotranspiration using the Hargreaves equation and, based on precipitation and potential evapotranspiration data, calculated the SPEI using the Python gma package. The formula for calculating potential evapotranspiration is as follows:

[0120] PET=0.0023×S0×(MaxT-MinT)×0.5×(MeanT+17.8) (20)

[0121] In the formula, PET represents potential evapotranspiration, MaxT, MinT, and MeanT represent the monthly maximum, minimum, and average temperatures, respectively; S0 represents the theoretical solar radiation reaching the top of the Earth's atmosphere, calculated based on the solar constant, Earth-Sun distance, Julian day, declination, etc.

[0122] Human disturbance intensity in socioeconomic data is related to population density; therefore, it is calculated by multiplying it by the ratio of future population density to current population density. The nighttime light index primarily reflects the region's economic development level; referencing the GDP growth rate set in the study area's planning, it is used as a proportionality coefficient to calculate future nighttime light data. Altitude, slope, soil organic carbon, and soil sand content in soil and geographical environmental data are environmental factors, and their values ​​are assumed to remain constant under future scenarios.

[0123] Based on future scenario data, the vegetation change caused by each key influencing factor under the future scenario is calculated, the characteristic relationship between the dependent and independent variables is obtained, and the effect value of each key influencing factor is identified. The effect value is compared with the corresponding effect threshold of each key influencing factor. When the effect value of the key influencing factor under the future scenario exceeds the effect threshold, it indicates that the forest and grassland vegetation is degraded, and early warning is carried out.

[0124] In addition, when the impact value of key influencing factors under future scenarios exceeds the impact threshold, it indicates that forest and grassland vegetation is degraded. Early warning is carried out, which also includes: extracting forest, shrub, grassland and desert areas based on land use type data under different future scenarios, combining early warning signals of key influencing factors of different forest and grassland vegetation, calculating the risk degree of forest and grassland vegetation degradation under future scenarios, and obtaining the mean value of the degradation risk degree of each type of vegetation; based on the mean value of the degradation risk degree of each type of vegetation, the degradation risk degree is divided into four levels: no degradation risk, low degradation risk, medium degradation risk and high degradation risk, according to the natural breakpoint method; by statistically analyzing the risk level distribution of degradation warning areas of each type of forest and grassland vegetation under different future scenarios, the spatiotemporal distribution characteristics of degradation risk areas are obtained, and the dynamic change trend of the area of ​​degradation warning areas at each level is quantified.

[0125] In this embodiment, it should be noted that, for classifying the degree of degradation risk into four levels—no degradation risk, low degradation risk, medium degradation risk, and high degradation risk—the degree of degradation risk can be calculated using a formula to determine the risk of forest and grassland vegetation degradation under future scenarios:

[0126]

[0127] In the formula, Value reflects the degree of risk of forest and grassland vegetation degradation under future scenarios, and n is the number of key influencing factors that exceed the early warning signal threshold.

[0128] The implementation principle of this embodiment is as follows: The above steps mainly involve acquiring vegetation, climate, soil, geographical environment, and socioeconomic data of the study area. Based on the acquired data, vegetation index data is used to identify areas of significant forest and grassland vegetation degradation and negative mutation points, thereby obtaining key influencing factors affecting vegetation degradation. Based on these key influencing factors and considering the lagged cumulative effect of climate factors, the threshold of each key influencing factor's effect on negative mutations in forest and grassland vegetation is accurately determined, serving as an early warning signal for forest and grassland vegetation degradation. This fully considers the nonlinear characteristics of the effects of various environmental factors, including natural, anthropogenic, and environmental factors, and can accurately identify early warning signals for forest and grassland vegetation degradation, applicable to different ecosystems and regions. Based on future scenario data, combined with early warning signals of different influencing factors on forest and grassland vegetation, the risk level of forest and grassland vegetation degradation under future scenarios is determined, providing a quantitative basis for early warning of vegetation degradation under future scenarios and providing scientific support for the formulation of management strategies for different ecosystems.

[0129] Reference Figure 10 As shown, this application also provides an early warning signal identification system for forest and grassland vegetation degradation. This system may include: an acquisition module 301, a significantly degraded vegetation area module 302, a vegetation negative mutation point identification module 303, a key influencing factor acquisition module 304, and an early warning signal identification module 305. The main functions of each component module are as follows:

[0130] The acquisition module 301 is used to acquire geographic data of the study area; the geographic data includes vegetation data, climate data, soil data, geographic environment data, and socio-economic data.

[0131] The vegetation significantly degraded area module 302 is used to obtain vegetation significantly degraded areas based on geographic data and vegetation index analysis; among them, vegetation index is one of the important parameters reflecting crop growth and nutrition information, and vegetation significantly degraded areas refer to areas where the vegetation index has shown a significant negative trend in a specific period of time in the past.

[0132] The vegetation negative mutation point identification module 303 is used to identify vegetation negative mutation points based on the BFAST algorithm in areas of significant vegetation degradation; wherein, vegetation negative mutation points refer to the time and location of negative mutations in areas of significant vegetation degradation.

[0133] The Key Influencing Factors Module 304 is used to screen out key natural, anthropogenic, and environmental influencing factors that affect negative vegetation mutations based on the identified vegetation negative mutation points and vegetation index analysis.

[0134] The early warning signal identification module 305 is used to train a model using a machine learning model with key influencing factors as independent variables and vegetation index as dependent variable, and to calculate the vegetation change caused by each key influencing factor using the SHAP method. The critical point of positive to negative change in vegetation change is determined as the threshold of the key influencing factor, thus obtaining the early warning signal of forest and grassland vegetation degradation.

[0135] like Figure 11 The diagram shown is a block diagram of a computer device according to an embodiment of this application. The term "computer device" is intended to represent various forms of digital computers or mobile devices. The digital computer may include a desktop computer, a portable computer, a workbench, a personal digital assistant, a server, a mainframe computer, and other suitable computers. The mobile device may include a tablet computer, a smartphone, a wearable device, etc.

[0136] like Figure 11 As shown, device 600 includes a computing unit 601, a ROM 602, a RAM 603, a bus 604, and an input / output (I / O) interface 605. The computing unit 601, ROM 602, and RAM 603 are interconnected via the bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0137] The computing unit 601 can execute various processes in the method embodiments of this application according to computer instructions stored in the read-only memory (ROM) 602 or computer instructions loaded from the storage unit 608 into the random access memory (RAM) 603. The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. The computing unit 601 can include, but is not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. In some embodiments, the methods provided in the embodiments of this application can be implemented as computer software programs, which are tangibly contained in a computer-readable storage medium, such as the storage unit 608.

[0138] RAM 603 may also store various programs and data required for the operation of device 600. Part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609.

[0139] The input unit 606, output unit 607, storage unit 608, and communication unit 609 in device 600 can be connected to I / O interface 605. The input unit 606 can be, for example, a keyboard, mouse, touchscreen, or microphone; the output unit 607 can be, for example, a display, speaker, or indicator light. Device 600 can exchange information and data with other devices through the communication unit 609.

[0140] It should be noted that the device may also include other components necessary for normal operation. It may also include only the components necessary for implementing the solution of this application, without necessarily including all the components shown in the figures.

[0141] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), payload programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof.

[0142] The computer instructions used to implement the methods of this application may be written in any combination of one or more programming languages. These computer instructions may be provided to the computing unit 601 such that when executed by the computing unit 601, such as a processor, the computer instructions cause the execution of the steps involved in the embodiments of the methods of this application.

[0143] The computer-readable storage medium provided in this application can be a tangible medium that can contain or store computer instructions for performing the steps involved in the method embodiments of this application. The computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, and other forms of storage media.

[0144] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for identifying early warning signals of degradation of forest and grass vegetation, characterized in that, The method comprises the following steps: Obtain geographical data of a study area; wherein, The geographical data comprises vegetation data, climate data, soil data, geographical environment data and socio-economic data; Based on the geographical data, a vegetation significantly degraded area is obtained according to vegetation index analysis; wherein, the vegetation index is one of important parameters reflecting crop growth and nutrition information, and the vegetation significantly degraded area refers to an area where the vegetation index presents a significant negative trend in a specific period of time in the past; Based on the vegetation significantly degraded area, a vegetation negative mutation point is identified according to a BFAST algorithm; wherein, the vegetation negative mutation point refers to the time and position of negative mutation of the vegetation significantly degraded area; based on the identified vegetation negative mutation point, key influencing factors of nature, human and environment affecting the vegetation negative mutation are screened out in combination with the vegetation index analysis; The key influencing factors are taken as independent variables, and the vegetation index is taken as a dependent variable; a machine learning model is used for model training; a vegetation change amount caused by each key influencing factor is calculated through a SHAP method; a critical point of positive and negative transformation of the vegetation change amount is determined as an action threshold of the key influencing factor; and an early warning signal of forest and grass vegetation degradation is obtained.

2. The method according to claim 1, wherein, The vegetation significantly degraded area obtained based on the geographical data according to the vegetation index analysis comprises: Based on land use type data in the socio-economic data and desert distribution data in the geographical environment data, a mask processing is performed to remove original desert areas of the study area; based on the study area from which the original desert areas are removed, the forest and grass vegetation types of the study area are further divided into forest, shrub and grassland types according to the land use type data; Based on a Theil-Sen trend algorithm, whether the vegetation index of the study area presents a negative trend in a specific period of time in the past is judged according to whether the slope of the vegetation index in the time series is negative; If yes, whether the negative trend is significant is judged based on a Mann-Kendall method on the basis of the negative trend; If significant, the corresponding area is defined as a vegetation significantly degraded area.

3. The method according to claim 2, wherein, The vegetation negative mutation point identified based on the vegetation significantly degraded area according to the BFAST algorithm specifically comprises: Obtain temperature and precipitation sequences of the study area; Based on the temperature and precipitation sequences, a moving average method is used to exclude the influence of climate factor variability on the vegetation negative mutation and retain the vegetation negative mutation caused by climate change; wherein, The climate change is a trend, and the climate variability is an interannual fluctuation; the negative mutation caused by the interannual climate fluctuation is excluded, and the vegetation negative mutation caused by the climate trend change is retained; Monthly NDVI time series data are taken as input, and a BFAST model is used for pixel-by-pixel mutation detection; the time and position of negative mutation of the significantly degraded area are determined as the vegetation negative mutation point.

4. The method according to claim 2, wherein, Based on the identified vegetation negative mutation point, key influencing factors of nature, human and environment affecting the vegetation negative mutation are screened out in combination with the vegetation index analysis, which comprises: Based on the identified vegetation negative mutation point, influencing factors of nature, human and environment affecting the vegetation negative mutation are screened out; The correlation heat map is drawn to calculate the correlation values between the influence factors, and the correlation influence factors are screened by comparing with preset values; The correlation influence factors are taken as independent variables, and the monthly vegetation index is taken as a dependent variable. Model training is performed for each forest and grass vegetation type according to a random forest model, so as to obtain the importance threshold of each correlation influence factor on the negative mutation of different vegetation types, and the key influence factors are screened according to preset requirements.

5. The forest and grass vegetation degradation early warning signal identification method according to claim 4, characterized in that, The correlation heat map is drawn to calculate the correlation values between the influence factors, and the correlation influence factors are screened by comparing with preset values, including: The correlation between the influence factors is calculated based on a heatmap function, and the correlation heat map is drawn; The correlation values between the influence factors are obtained according to the correlation heat map; and the correlation values are compared with the preset values; If the correlation values are greater than the preset values, the corresponding influence factors are excluded; Otherwise, the influence factors are retained to obtain the correlation influence factors.

6. The method according to claim 5, wherein, The correlation influence factors are taken as independent variables, and the monthly vegetation index is taken as a dependent variable. Model training is performed for each forest and grass vegetation type according to a random forest, so as to obtain the importance threshold of each correlation influence factor on the negative mutation of different vegetation types, and the key influence factors are screened according to preset requirements, including: For each forest and grass vegetation type, the NDVI synthesized by the monthly maximum value is taken as a dependent variable, and the correlation influence factors are taken as independent variables, so as to obtain training samples; The training samples are divided into a training set and a validation set; The training set and the validation set are used to train and verify the model for each forest and grass vegetation type according to a random forest, so as to obtain the feature relationship between the dependent variable and the independent variable, and construct a prediction model; The prediction value obtained according to the prediction model is compared with the true value, and the prediction situation is reflected by a determination coefficient R 2 and a root mean square error RMSE. According to the random forest model results of different forest and grass vegetation, the relative importance values of each correlation influence factor are obtained; According to the importance value of the correlation influence factor, the independent variable input is controlled, the random forest model is trained, and the importance threshold of different vegetation types is determined by comparing the R 2 and RMSE of the model; when the importance of the correlation influence factor is greater than the importance threshold, the corresponding correlation influence factor is retained to obtain the key influence factor.

7. The method according to claim 6, wherein, The key influence factors of nature, human and environment affecting the negative mutation of vegetation are screened based on the identified negative mutation points of vegetation and the vegetation index analysis, and the key influence factors are screened based on the key influence factors according to a machine learning model, and the model is trained for each negative mutation point of vegetation; The key influence factors are taken as independent variables, and the vegetation index is taken as a dependent variable. Model training is performed by using a machine learning model, the vegetation change amount caused by each key influence factor is calculated by using a SHAP method, and the critical point of positive and negative change of the vegetation change amount is determined as the action threshold of the key influence factor, so as to obtain the early warning signal of forest and grass vegetation degradation, including: Based on the lag effect of climate factors, during the model training, by controlling the variable method to control the invariability of other key influence factors, the climate influence factors at different times are taken as inputs to obtain prediction values; the R 2 and RMSE of each model are compared to determine the best fitting climate lag and cumulative time; wherein the lag effect of the climate factor refers to the time when the vegetation negative mutation is most affected by the climate factor, which is the climate factor at a specific time before the occurrence of the vegetation negative mutation.

8. The method according to claim 7, wherein, The key influence factors considering the lag effect of climate factors are taken as independent variables, and the annual maximum synthesized vegetation index NDVI is taken as a dependent variable. The random forest model is trained, and the training set and the validation set are divided according to a specific proportion; The prediction results of the random forest model are decomposed into the sum of the contributions of each key influence factor, and the vegetation change amount caused by each key influence factor is calculated by using a SHAP method. ​ Taking each key influencing factor as the independent variable and the SHAP value corresponding to each key influencing factor as the dependent variable, a GAM is used for fitting; By reverse calculation, the threshold value of each key influencing factor is obtained when the dependent variable is 0, which is used as an early warning signal of forest and grass vegetation degradation.

9. The method according to any one of claims 1-8, wherein, Also includes: Obtaining future scenario data; wherein, The future scenario data includes climate data under each future scenario provided by the International Coupled Model Comparison Program Phase VI (CMIP6) and social and economic data calculated according to population density, nighttime light index, GDP growth rate, as well as unchanged soil data and geographical environment data; based on the future scenario data, the action value of each key influencing factor is obtained, compared with the corresponding early warning signal, and the risk degree of forest and grass vegetation degradation under the future scenario is determined to realize the risk early warning of potential vegetation degradation area under the future scenario.

10. The method according to claim 9, wherein the method is characterized by, The future scenario data, the action value of each key influencing factor is obtained, compared with the corresponding early warning signal, and the risk degree of forest and grass vegetation degradation under the future scenario is determined to realize the risk early warning of potential vegetation degradation area under the future scenario, also includes: Based on the land use type data under different future scenarios, forest, shrub, grassland and desert areas are extracted, combined with the early warning signal of the key influencing factor of different forest and grass vegetation, the risk degree of forest and grass vegetation degradation under the future scenario is calculated, and the mean value of the risk degree of each type of vegetation degradation is obtained; Based on the mean value of the risk degree of each type of vegetation degradation, the grouping standard is obtained according to the natural breakpoint method to divide the risk degree into four grades of no degradation risk, low degradation risk, medium degradation risk and high degradation risk; Through statistical analysis of the degradation early warning area risk grade distribution of each forest and grass vegetation type under different future scenarios, the spatio-temporal distribution characteristics of the degradation risk area are obtained, and the dynamic change trend of the area of each grade of degradation early warning area is quantified.

Citation Information

Patent Citations

  • Forest fire early warning model construction method based on time decay precipitation algorithm

    CN114266392A

  • Black soil cultivated land organic carbon degradation diagnosis method based on geographic big data and AI

    CN119475116A