A gross primary productivity estimation method and system for arid regions

CN122735501APending Publication Date: 2026-09-11LANZHOU JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611091937.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种面向干旱区的总初级生产力估算方法,以解决现有技术中存在的以下问题:通量塔站点观测尺度与遥感像元尺度不匹配,导致训练样本代表性不足;传统植被指数难以准确表征干旱区植被对水分胁迫及极端气候事件的响应过程;不同气候区及不同植被类型之间生态异质性显著,单一模型难以兼顾区域差异;不同土层土壤湿度变量对GPP的指示能力存在差异,缺乏有效的优选机制

Benefits of technology

[0028] (1) Improve the spatial representativeness of training samples and reduce scale mismatch error.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122735501A_ABST
    Figure CN122735501A_ABST
Patent Text Reader

Abstract

The application discloses a total primary productivity estimation method for arid regions, and belongs to the technical field of remote sensing inversion and machine learning modeling. The method realizes high-precision estimation of total primary productivity by fusing a light energy utilization efficiency mechanism model and a random forest data-driven model. The method takes the process of vegetation photosynthesis as a theoretical basis, introduces multi-source remote sensing data and meteorological factors, and constructs a nonlinear mapping relationship under environmental constraints, so that the stability and accuracy of total primary productivity estimation under drought stress conditions are improved. The application can improve the accuracy and applicability of total primary productivity estimation in arid regions, and has good popularization and application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ecological remote sensing inversion and terrestrial ecosystem productivity estimation technology, and in particular to a method and system for estimating total primary productivity in arid regions. Background Technology

[0002] Gross Primary Productivity (GPP) is an important indicator of the photosynthetic carbon sequestration capacity of terrestrial ecosystems and a key parameter in global carbon cycle research. Arid and semi-arid regions cover vast areas, and their ecosystems are highly sensitive to changes in hydrothermal conditions. Therefore, accurate estimation of GPP in these regions is of great significance for ecological monitoring, carbon cycle research, and regional sustainable development.

[0003] Among the existing GPP estimation methods, remote sensing inversion methods based on vegetation indices such as NDVI, EVI, and NIRv are widely used. However, these methods mainly reflect vegetation structure or canopy reflectance characteristics and have limited ability to respond to changes in photosynthesis of vegetation in arid areas under short-term water stress, extreme high temperature, and rapid phenological fluctuations. They are difficult to accurately characterize the dynamic productivity changes of arid ecosystems.

[0004] Solar-induced chlorophyll fluorescence (SIF) can more directly reflect the photosynthetic process of vegetation and has a stronger physiological indicative significance compared with traditional vegetation indices. Although existing technologies have used products such as GOSIF for GPP estimation, there are still several shortcomings: First, the spatial representation range of flux tower sites is usually only a few meters to a few hundred meters, while the spatial resolution of SIF products can reach 0.05°, resulting in a significant scale mismatch and insufficient representativeness of training samples; second, existing methods often lack specific stratification strategies for the ecological heterogeneity of arid regions, making it difficult to take into account the differences between arid and non-arid regions and between different vegetation types; third, there are many variables related to water stress, and the indicative ability of soil moisture variables in different soil layers to GPP varies, and existing schemes lack optimization mechanisms for representative soil moisture factors.

[0005] Furthermore, existing GPP estimation products, such as MODISGPP, are mostly based on light energy use efficiency models, and GOSIFGPP is mostly based on semi-empirical SIF-GPP relationships. Under the complex environmental conditions of arid regions, their characterization of the nonlinear response under the combined effects of temperature, moisture, radiation, canopy structure, and extreme climate is still insufficient. Therefore, it is necessary to propose a GPP estimation scheme specifically for arid regions to improve sample representativeness, enhance the model's adaptability to ecological heterogeneity and water stress, and improve the accuracy and stability of GPP estimation in arid areas. Summary of the Invention

[0006] The purpose of this invention is to provide a method for estimating total primary productivity (GPP) in arid regions, addressing the following problems in existing technologies: the mismatch between flux tower site observation scale and remote sensing pixel scale leads to insufficient representativeness of training samples; traditional vegetation indices are difficult to accurately characterize the response of vegetation in arid regions to water stress and extreme climate events; significant ecological heterogeneity exists between different climate zones and different vegetation types, making it difficult for a single model to take into account regional differences; and soil moisture variables in different soil layers have varying indicative abilities for GPP, lacking an effective optimization mechanism.

[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0008] In a first aspect, the present invention provides a method for estimating total primary productivity in arid regions, comprising the following steps: S1, acquiring solar-induced chlorophyll fluorescence data, flux station total primary productivity observation data, vegetation structure parameter data, meteorological driving data, water status data, and topographic data for the target area; S2, preprocessing the multi-source data acquired in step S1, unifying various input variables to a monthly scale and Spatial resolution was achieved, and coordinate system unification, spatial registration, and temporal alignment were completed; S3, saturated vapor pressure difference (VPD) was calculated based on 2-meter air temperature and 2-meter dew point temperature, a heat wave index was constructed based on daily maximum temperature, and drought index (AI), standardized precipitation evapotranspiration index (SPEI), and soil moisture variables of different soil layers were extracted to form a set of water stress characteristics; S4, based on the location of flux stations... Using solar-induced chlorophyll fluorescence pixels as the analysis unit, spatial representativeness screening of site samples was performed by combining multi-period land cover data and monthly enhanced vegetation index data: the average coverage ratio of the dominant land cover type in multiple time phases was calculated, and the mean of the monthly EVI standard deviation was calculated; only when the average coverage ratio met the preset spatial homogeneity threshold and the mean of the monthly EVI standard deviation met the preset temporal stability threshold were the corresponding samples determined as valid training samples; S5, the valid training samples were climatically zoned according to the drought index AI. When the drought index AI was less than or equal to 0.5, it was classified as arid zone samples, and the rest were classified as non-arid zone samples; based on the climatic zoning, ecological stratification was performed according to vegetation type to construct training subsets of different ecological layers; S6, the surface soil moisture SWVL2 and the middle soil moisture S WVL3 and deep soil moisture (SWVL4) were used as candidate input features in the training. The candidate soil moisture variables were optimized through model feature importance analysis. Specifically, for a subset of arid zone samples, when the feature importance of middle soil moisture (SWVL3) was higher than that of surface soil moisture (SWVL2) and deep soil moisture (SWVL4), middle soil moisture (SWVL3) was selected as the representative soil moisture factor. In S7, total primary productivity estimation models were constructed for arid zone, non-arid zone, and ecological layers of each vegetation type. A random forest regressor was used as the estimator, and the hyperparameter combination was determined through grid search. Five-fold cross-validation was used for model training. In S8, the trained ecological layer models were applied to monthly grid data of the target area to output the monthly or annual total primary productivity estimation results of the target area.

[0009] Furthermore, the formula for calculating the saturated vapor pressure difference VPD in step S3 is as follows: ;

[0010] Where T is the air temperature (°C), Td is the dew point temperature (°C), and VPD is in kPa.

[0011] Furthermore, the formula for calculating the drought index AI in step S3 is as follows:

[0012] ;

[0013] For precipitation, Potential evapotranspiration. The drought index is used to characterize the long-term dryness and wetness of a region's climate; when the drought index AI ≤ 0.5, it indicates arid conditions.

[0014] Furthermore, the method for constructing the heatwave index in step S3 is as follows:

[0015] The number of consecutive days of high temperatures within a month is counted. If there is a continuous period of no less than 3 days with a daily maximum temperature of no less than 35 degrees Celsius, the heat wave index for that month is assigned a value of 1; otherwise, it is assigned a value of 0.

[0016] Furthermore, the formula for calculating the average coverage proportion P_dom of the dominant land cover type in multiple time phases in step S4 is as follows:

[0017] ;

[0018] in: The average coverage ratio of the dominant land cover type; For the first The proportion of the dominant type in each phase; The number of phases.

[0019] Furthermore, the formula for calculating the mean of the monthly EVI standard deviation, Sigma_EVI, in step S4 is as follows:

[0020] ;

[0021] in: For the first Standard deviation of monthly EVI; This refers to the number of months.

[0022] Furthermore, in step S6, the feature importance is calculated based on the mean squared error (MSE). The formula for calculating the mean squared error (MSE) is as follows: ;

[0023] in, To observe GPP, These are predicted values.

[0024] Furthermore, step S6 specifically includes:

[0025] In the training subset of samples from arid regions, the feature importance scores of surface soil moisture (SWVL2), mid-layer soil moisture (SWVL3), and deep soil moisture (SWVL4) were compared. When the feature importance of mid-layer soil moisture (SWVL3) was the highest, mid-layer soil moisture (SWVL3) was determined as the representative soil moisture factor to reflect the water availability of the root zone, the main root distribution layer of vegetation in arid regions.

[0026] Secondly, this invention provides a total primary productivity estimation system for arid regions, comprising: a data acquisition and preprocessing module for acquiring solar-induced chlorophyll fluorescence data, flux station total primary productivity observation data, vegetation structure parameter data, meteorological driving data, water status data, and topographic data, and performing time-scale unification, spatial resolution unification, coordinate system unification, spatial registration, and temporal alignment processing on the multi-source data; and a water stress feature construction module for calculating the saturated vapor pressure difference (VPD) based on 2-meter air temperature and 2-meter dew point temperature. A heat wave index was constructed based on daily maximum temperatures, and drought index (AI), standardized precipitation evapotranspiration index (SPEI), and soil moisture variables for different soil layers were extracted to construct a water stress feature set. A site matching and screening module was used to spatially screen site samples using the 0.05-degree solar-induced chlorophyll fluorescence pixel at the flux station as the analysis unit, combined with multi-period land cover data and monthly enhanced vegetation index (EVI) data. The module calculated the average coverage ratio of the dominant land cover type across multiple time phases and the mean of the monthly EVI standard deviation; only when the average coverage ratio met the preset criteria... When the spatial homogeneity threshold is met and the mean of the monthly EVI standard deviation meets the preset time stability threshold, the corresponding sample is determined as a valid training sample. The partitioning and stratification module is used to divide the valid training samples into arid and non-arid areas based on the drought index AI, and then to perform ecological stratification according to vegetation type to construct training subsets of different ecological layers. The feature selection module is used to use soil moisture variables of different soil layers as candidate input features to participate in model training, and to determine the representative soil moisture factor based on the importance evaluation of model features. Among them, for arid area samples, when the feature importance of the middle soil moisture SWVL3 is higher than that of the surface and deep soil moisture, the middle soil moisture SWVL3 is selected as the representative soil moisture factor. The model optimization and construction module is used to establish a nonlinear mapping relationship between input features and total primary productivity based on the random forest regression algorithm, and to optimize the model hyperparameters through grid search and five-fold cross-validation to construct a partitioning-stratification estimation model. The regression prediction and result output module is used to apply the trained model to the monthly grid data of the target area and output the monthly or annual total primary productivity results of the target area.

[0027] The beneficial effects of this invention are as follows: The total primary productivity estimation method and system for arid regions provided by this invention have the following advantages:

[0028] (1) Improve the spatial representativeness of training samples and reduce scale mismatch error.

[0029] This invention constructs a site selection mechanism based on land cover stability and vegetation index temporal stability, and imposes spatial consistency constraints on flux site observation data and remote sensing pixel data. This effectively solves the mismatch problem between flux site point-scale observation and remote sensing pixel scale, thereby improving the representativeness of training samples and enhancing the reliability of model training.

[0030] (2) Enhance the ability to characterize water stress processes in arid areas.

[0031] This invention introduces multiple water stress variables, such as vapor pressure deficit, drought index, standardized precipitation evapotranspiration index, soil moisture and heat wave index, to characterize the vegetation water limitation process from multiple aspects, including atmospheric drought, soil moisture supply and extreme climate events, so that the model can more accurately reflect the response characteristics of vegetation photosynthesis to changes in water conditions in arid areas.

[0032] (3) Improve the model’s ability to adapt to ecological heterogeneity.

[0033] This invention constructs sub-models for arid regions, non-arid regions, and different ecological layers by using a climate zoning strategy based on drought index and an ecological stratification strategy based on vegetation type. This effectively characterizes the differences in total primary productivity under different climatic conditions and ecosystem types, and improves the applicability and stability of the model at the regional scale.

[0034] (4) Achieve adaptive optimization of representative soil moisture factors.

[0035] This invention uses a feature importance evaluation mechanism based on a random forest model to screen soil moisture variables in different soil layers and determine the representative soil moisture factors that are most closely related to the water supply in the root zone of vegetation. This avoids the introduction of redundant variables and improves the effectiveness and physical rationality of the model input features.

[0036] (5) Enhance the ability to characterize complex nonlinear relationships.

[0037] This invention uses a random forest regression model to establish a nonlinear mapping relationship between multiple driving variables and total primary productivity, and combines grid search and cross-validation for model optimization. This can effectively characterize the complex nonlinear relationship under the combined effect of multiple factors such as light energy driving, water limitation and environmental regulation, thereby improving the estimation accuracy.

[0038] (6) Improve the accuracy and stability of total primary productivity estimation in arid areas.

[0039] This invention comprehensively considers multi-source remote sensing data, meteorological driving factors, and water stress mechanisms, and significantly improves the accuracy and stability of total primary productivity estimation results in arid areas through partitioned and hierarchical modeling and feature optimization methods, showing promising application prospects. Attached Figure Description

[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a general flowchart of a method for estimating total primary productivity in arid regions according to an embodiment of the present invention;

[0042] Figure 2 This is a flowchart of random forest modeling and model evaluation in a total primary productivity estimation method for arid regions according to an embodiment of the present invention. Detailed Implementation

[0043] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0044] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0045] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0046] This invention proposes a method for estimating total primary productivity (TPP) in arid regions. It constructs a multi-source data-driven TPP estimation model by integrating solar-induced chlorophyll fluorescence data, flux station GPP observation data, vegetation structure parameter data, meteorological data, water status data, and topographic data. The method achieves consistency across different data sources in the spatiotemporal dimensions by uniformly processing the multi-source data at monthly scales and a 0.05° spatial resolution. It improves the spatial matching between flux station samples and remote sensing pixels by introducing a station representativeness screening mechanism based on multi-period land cover consistency and monthly enhanced vegetation index (EVI) fluctuation characteristics. The method enhances the model's adaptability to the ecological heterogeneity of arid regions by dividing arid and non-arid areas based on the drought index AI and combining it with vegetation type for ecological stratification. It determines representative soil moisture factors by comparing candidates and analyzing the feature importance of soil moisture variables in different soil layers. Finally, it establishes a partitioned-stratified TPP estimation model using a random forest regression model combined with grid search and five-fold cross-validation, and applies this model to the target area to achieve monthly or annual-scale estimation of TPP in arid regions, thereby improving the accuracy and stability of the estimation results.

[0047] Reference Figures 1-2 This is one embodiment of the present invention, which provides a method for estimating total primary productivity (GPP) in arid regions, comprising the following steps: S1, acquiring solar-induced chlorophyll fluorescence (SIF) data, flux station GPP observation data, vegetation structure parameter data, meteorological driving data, water status data, and topographic data for the target area; S2, preprocessing the multi-source data acquired in step S1, unifying various input variables to a monthly scale and Spatial resolution was achieved, and coordinate system unification, spatial registration, and temporal alignment were completed; S3, saturated vapor pressure difference (VPD) was calculated based on 2-meter air temperature and 2-meter dew point temperature, a heat wave index was constructed based on daily maximum temperature, and drought index (AI), standardized precipitation evapotranspiration index (SPEI), and soil moisture variables of different soil layers were extracted to form a set of water stress characteristics; S4, based on the location of flux stations... Using solar-induced chlorophyll fluorescence pixels as the analysis unit, spatial representativeness screening of site samples was performed by combining multi-period land cover data and monthly enhanced vegetation index data: the average coverage ratio of the dominant land cover type in multiple time phases was calculated, and the mean of the monthly EVI standard deviation was calculated; only when the average coverage ratio met the preset spatial homogeneity threshold and the mean of the monthly EVI standard deviation met the preset temporal stability threshold were the corresponding samples determined as valid training samples; S5, the valid training samples were climatically zoned according to the drought index AI. When the drought index AI was less than or equal to 0.5, it was classified as arid zone samples, and the rest were classified as non-arid zone samples; based on the climatic zoning, ecological stratification was performed according to vegetation type to construct training subsets of different ecological layers; S6, the surface soil moisture SWVL2 and the middle soil moisture S WVL3 and deep soil moisture (SWVL4) were used as candidate input features in the training. The candidate soil moisture variables were optimized through model feature importance analysis. Specifically, for a subset of arid zone samples, when the feature importance of middle soil moisture (SWVL3) was higher than that of surface soil moisture (SWVL2) and deep soil moisture (SWVL4), middle soil moisture (SWVL3) was selected as the representative soil moisture factor. In S7, total primary productivity estimation models were constructed for arid zone, non-arid zone, and ecological layers of each vegetation type. A random forest regressor was used as the estimator, and the hyperparameter combination was determined through grid search. Five-fold cross-validation was used for model training. In S8, the trained ecological layer models were applied to monthly grid data of the target area to output the monthly or annual total primary productivity estimation results of the target area.

[0048] Furthermore, step S3 specifically includes:

[0049] Step S31: Calculate the vapor pressure deficit (VPD) based on the 2-meter air temperature and 2-meter dew point temperature to characterize the degree of atmospheric drought. When VPD increases, vegetation reduces stomatal conductance to minimize transpiration, leading to restricted carbon dioxide diffusion and thus inhibiting photosynthesis. Therefore, vapor pressure deficit, as an important variable connecting atmospheric water deficit and vegetation physiological responses, is used to characterize the stomatal regulation mechanism and its constraining effect on total primary productivity. Its calculation formula is as follows:

[0050] Where: T is the air temperature (°C), Td is the dew point temperature (°C), and VPD is in kPa;

[0051] Step S32: Extract the drought index AI, standardized precipitation evapotranspiration index SPEI, and soil moisture variables of different soil layers to construct a set of water stress characteristics.

[0052] Among them: (1) The drought index is defined as the ratio of precipitation to potential evapotranspiration:

[0053]

[0054] For precipitation, This represents potential evapotranspiration. The drought index is used to characterize the long-term dryness and wetness of the region's climate; when AI ≤ 0.5, it represents arid conditions.

[0055] (2) Standardized precipitation evapotranspiration index (SPEI) is used to characterize the difference between precipitation and evapotranspiration. It is preferred to use an SPEI series with a time scale of 1 to 6 months to reflect short- to medium-term abnormal changes in water content.

[0056] (3) Soil moisture variables include the volumetric water content of different soil layers, preferably including: SWVL2 (topsoil), SWVL3 (middle soil), and SWVL4 (deep soil). Each soil layer corresponds to a different depth range and is used to characterize the water supply capacity of the vegetation root zone.

[0057] The above variables characterize the water situation from three aspects: long-term climate background, short- and medium-term water fluctuations, and root zone water supply, so that the model can reflect the water transport and utilization process in the soil-vegetation-atmosphere continuum.

[0058] Step S33: Construct a heat wave index based on daily maximum temperatures. When there are consecutive months with no less than The day's highest temperature is not lower than the threshold. During periods of high temperatures, the heatwave index for that month is assigned a value of 1; otherwise, it is assigned a value of 0. The preferred duration threshold is 3 days. This is the temperature threshold, and its value range is... Preferred .

[0059] Heatwave indices are used to characterize the inhibitory effect of extreme high temperatures on photosynthesis. Under high-temperature conditions, vegetation transpiration increases and may lead to leaf water deficit, while the activity of photosynthetic enzymes decreases, thus negatively impacting total primary productivity.

[0060] Step S34: Combine vapor pressure deficit, drought index, standardized precipitation evapotranspiration index, soil moisture and heat wave index to construct a water stress feature set, and use it together with solar-induced chlorophyll fluorescence as model input variables.

[0061] Among them, solar-induced chlorophyll fluorescence is used to characterize the light energy distribution state during photosynthesis, while water stress variables constrain photosynthesis through stomatal regulation and water supply processes, enabling the model to express the coupling relationship between light energy drive and water limitation.

[0062] Furthermore, step S4 specifically includes:

[0063] Step S41: Based on the spatial location of the flux observation station, map the station observation data to the corresponding remote sensing pixels, preferably using a spatial resolution consistent with solar-induced chlorophyll fluorescence (SIF). As an analysis unit.

[0064] Let the location of the station be The corresponding remote sensing pixel is Then, establish the correspondence between the site observations and the pixel scale features:

[0065] in, These are flux station observations. These are the multi-source input features for the corresponding pixels.

[0066] Since flux observations have point-scale characteristics while remote sensing data have pixel-scale characteristics, there is an inconsistency between the two in terms of spatial scale. Therefore, it is necessary to screen the spatial representativeness of the station samples to reduce the errors caused by scale mismatch.

[0067] Step S42: Construct spatial homogeneity constraints based on land cover stability. Analyze the surface types within pixels using multi-period land cover data, and calculate the average coverage ratio of the dominant land cover type across multiple temporal phases, defined as:

[0068] in: The average coverage ratio of the dominant land cover type; For the first The proportion of the dominant type in each phase; The number of phases.

[0069] When satisfied At that time, it is considered that the pixel has a spatially stable dominant land surface type; among which, the threshold The value ranges from 0.4 to 0.7, with 0.5 being the preferred value.

[0070] This constraint is used to ensure that the ecosystem types observed by flux stations are consistent with the land surface types represented by remote sensing pixels, thereby improving the spatial consistency of the samples.

[0071] Step S43: Construct time-consistency constraints based on vegetation index stability. The dynamic changes in vegetation within a pixel are assessed using the Enhanced Vegetation Index (EVI) time series. The mean of the monthly EVI standard deviation is calculated and defined as:

[0072] in: For the first Standard deviation of monthly EVI; This refers to the number of months.

[0073] When satisfied At that time, the vegetation state of the pixel is considered to be stable over time; among which, the threshold

[0074] The value ranges from 0.05 to 0.15, with 0.1 being preferred. This constraint is used to exclude unstable pixels caused by land use changes, disturbances, or observation errors.

[0075] Step S44: Determine the effective training sample set. Only when a pixel simultaneously satisfies the spatial homogeneity constraint (step S42) and the temporal stability constraint (step S43) is the corresponding flux station sample determined as a valid training sample; otherwise, it is discarded.

[0076] Through the above screening process, a sample set that is spatially representative and temporally stable is obtained, thereby reducing the impact of scale mismatch and heterogeneity on model training and improving the reliability and generalization ability of the model.

[0077] Furthermore, step S6 specifically involves:

[0078] Step S61: Use soil moisture variables of different soil layers as candidate input features to participate in model training. The candidate input features include at least surface soil moisture SWVL2, middle soil moisture SWVL3 and deep soil moisture SWVL4.

[0079] Step S62: During the training of the random forest model, the candidate soil moisture variable is evaluated using a feature importance index. The feature importance is calculated based on the reduction in mean squared error during node splitting and is defined as follows:

[0080] in: Indicates the first Importance score of each feature; Indicates at node Features used The decrease in mean square error resulting from splitting; For the first The set of all split nodes in a decision tree; The number of decision trees.

[0081] By comparing the importance scores of SWVL2, SWVL3, and SWVL4, the soil moisture variable with the highest importance was selected as the representative input feature.

[0082] In step S63, when the characteristic importance of the intermediate soil moisture SWVL3 is higher than that of SWVL2 and SWVL4, SWVL3 is selected as the representative soil moisture factor to participate in the model construction.

[0083] The reason is that: the root system of vegetation is mainly distributed in the middle soil layer, and SWVL3 can more directly reflect the available water of plants; the surface soil moisture SWVL2 is easily affected by short-term precipitation and evaporation, and fluctuates greatly with low stability; although the deep soil moisture SWVL4 has strong stability, it is less correlated with the actual water absorption process of plants; therefore, the middle soil moisture is more representative in reflecting the relationship between root zone water supply and plant physiological response.

[0084] By combining data-driven feature importance analysis with ecological mechanisms through the above feature optimization process, the physical rationality of model input variables and prediction accuracy can be improved.

[0085] Furthermore, step S7 specifically includes:

[0086] Step S71: Construct sample representation and input feature vector. After completing sample selection and feature optimization, each sample is represented as:

[0087]

[0088] in, For the first The input feature vector of each sample, Indicates the first One characteristic variable, For the number of features, The corresponding total primary productivity observations are used. The input features include multiple source variables such as solar-induced chlorophyll fluorescence (SIF), vapor pressure deficit (VPD), soil moisture, standardized precipitation evapotranspiration index (SPEI), air temperature, precipitation, solar radiation, leaf area index (LAI), vegetation absorbance of photosynthetically active radiation (FAPAR), and topographic elevation, thereby constructing a multidimensional feature space for characterizing the photosynthetic process.

[0089] Step S72: Construct a random forest regression model and establish a nonlinear mapping relationship. Based on the above feature space, construct a random forest regression model to establish a nonlinear mapping relationship between input features and GPP:

[0090] in, No. A decision tree, The number of decision trees.

[0091] The Bootstrap sampling with replacement method is used in the model construction process to randomly sample samples from the original sample set to construct a sub-training set. For each decision tree, samples are randomly selected from all features when splitting at a node. Several features are involved in the partitioning, among which , Given the total number of features, the node splitting criterion is to minimize the mean squared error (MSE).

[0092] in, To observe GPP, These are predicted values.

[0093] By introducing sample perturbation (Bootstrap) and feature randomness, the model variance is reduced and the ability to approximate high-dimensional nonlinear relationships is enhanced, thereby improving the model's generalization performance.

[0094] Step S73 introduces the multi-source variable mechanism and establishes the coupling relationship. In the model construction process, SIF, as a directly observed variable of photosynthesis, provides core information on the light energy distribution process; water stress variables such as vapor pressure deficit and soil moisture affect the vegetation's ability to absorb carbon dioxide by regulating stomatal conductance, thus limiting photosynthesis; air temperature and radiation reflect the energy supply and enzymatic reaction conditions of photosynthesis, respectively. The combined effect of these variables enables the model to express the nonlinear coupling relationship between light energy driving, water limitation, and environmental regulation.

[0095] Step S74: Optimize model hyperparameters. Optimize the key hyperparameters of the random forest model, including the number of decision trees. The value ranges from 100 to 1000; maximum depth The value ranges from 5 to 30; the minimum number of samples required for node splitting. The value ranges from 2 to 10; minimum number of samples for leaf nodes. The value ranges from 1 to 5; the size of the feature subset. Preferred .

[0096] A grid search method is used to combine and traverse the above parameters, and five-fold cross-validation is used to evaluate the model performance under different parameter combinations, thereby determining the optimal hyperparameter configuration.

[0097] Step S75: Under optimal parameter configuration, train the random forest model and construct sub-models based on climate zoning and ecological stratification.

[0098] in: Climate zones (arid / non-arid); It refers to vegetation type (forest, grassland, shrubland, etc.).

[0099] By using zonal-layered modeling, the differences in photosynthesis under different water conditions and ecological structures are explicitly introduced into the model, thereby improving the model's ability to express regional heterogeneity.

[0100] This embodiment also provides a total primary productivity estimation system for arid regions, including: a data acquisition and preprocessing module for acquiring solar-induced chlorophyll fluorescence data, flux station total primary productivity observation data, vegetation structure parameter data, meteorological driving data, water status data, and topographic data, and performing time scale unification, spatial resolution unification, coordinate system unification, spatial registration, and temporal alignment processing on the multi-source data; and a water stress feature construction module for calculating the saturated vapor pressure difference (VPD) based on 2-meter air temperature and 2-meter dew point temperature, and based on... A heat wave index was constructed using daily maximum temperatures. Simultaneously, drought index (AI), standardized precipitation evapotranspiration index (SPEI), and soil moisture variables from different soil layers were extracted to construct a water stress characteristic set. A site matching and screening module was used to spatially screen site samples using the 0.05-degree solar-induced chlorophyll fluorescence pixel at the flux station location as the analysis unit, combined with multi-period land cover data and monthly enhanced vegetation index (EVI) data. This involved calculating the average coverage ratio of the dominant land cover type across multiple temporal phases and the mean of the monthly EVI standard deviation. Only when the average coverage ratio met the preset spatial... When the homogeneity threshold is met and the mean of the monthly EVI standard deviation meets the preset time stability threshold, the corresponding sample is determined as a valid training sample. The partitioning and stratification module is used to divide the valid training samples into arid and non-arid areas based on the drought index AI, and then to perform ecological stratification according to vegetation type to construct training subsets of different ecological layers. The feature selection module is used to use soil moisture variables of different soil layers as candidate input features to participate in model training, and to determine the representative soil moisture factor based on the importance evaluation of model features. Among them, for arid area samples, when the feature importance of the middle soil moisture SWVL3 is higher than that of the surface and deep soil moisture, the middle soil moisture SWVL3 is selected as the representative soil moisture factor. The model optimization and construction module is used to establish a nonlinear mapping relationship between input features and total primary productivity based on the random forest regression algorithm, and to optimize the model hyperparameters through grid search and five-fold cross-validation to construct a partitioned-stratified estimation model. The regression prediction and result output module is used to apply the trained model to the monthly grid data of the target area and output the monthly or annual total primary productivity results of the target area.

[0101] Example 1: Please refer to Figures 1 to 2 As shown in this embodiment, a method for estimating total primary productivity in arid regions includes the following steps:

[0102] S1: Obtain multi-source data and construct the raw dataset required for total primary productivity estimation.

[0103] Specifically, acquiring multi-source data and constructing the raw dataset required for total primary productivity estimation includes the following steps:

[0104] S11. Obtain solar-induced chlorophyll fluorescence data.

[0105] Preferably, the solar-induced chlorophyll fluorescence data uses the GOSIF product with a spatial resolution of 0.05° and a temporal resolution of 8 days, and is converted into monthly-scale data through temporal aggregation.

[0106] It should be noted that GOSIF data can more directly characterize the fluorescence signal of light energy distribution during vegetation photosynthesis. Compared with traditional vegetation indices such as NDVI and EVI, it has stronger physiological indicative significance and is more suitable for monitoring the rapid response process of vegetation in arid areas.

[0107] S12. Obtain GPP observation data from flux stations.

[0108] Preferably, the flux station GPP observation data uses monthly-scale GPP data from the FLUXNET2015, Drought-2018, and AmeriFluxFLUXNET datasets.

[0109] It should be noted that the GPP observations are preferably monthly-scale GPP products obtained based on the nighttime flux allocation method, and are used as label values ​​in model training.

[0110] S13. Acquire vegetation structure parameter data, meteorological driving data, moisture status data, and topographic data.

[0111] Among them, vegetation structure parameter data should include at least LAI and FAPAR; meteorological driving data should include at least 2-meter air temperature, precipitation, solar radiation and 2-meter dew point temperature; moisture status data should include at least drought index AI, standardized precipitation evapotranspiration index SPEI, heat wave index and soil moisture data of different soil layers; and topographic data should preferably be digital elevation model (DEM) data.

[0112] Preferably, the vegetation structure parameter data uses GLASS products, the meteorological driving data and soil moisture data use ERA5-Land reanalysis data, and the terrain data uses SRTM digital elevation model data.

[0113] S2: Preprocess the multi-source data to construct a standardized input dataset.

[0114] Specifically, preprocessing multi-source data to construct a standardized input dataset includes the following steps:

[0115] S21. Perform unified time-scale processing on data from different sources.

[0116] For 8-day or daily-scale data, monthly average, monthly cumulative, or time aggregation methods are used to unify the data to a monthly scale; for flux station data, monthly-scale observations corresponding to the time range of remote sensing data are extracted.

[0117] S22. Perform spatial resampling and coordinate system unification processing on data from different sources.

[0118] Data with different spatial resolutions were unified to a spatial resolution of 0.05° and converted to a unified geographic coordinate system.

[0119] It should be noted that the original spatial resolution of GOSIF and GLASS data is 0.05°, and the original spatial resolution of ERA5-Land is 0.1°. In this embodiment, the ERA5-Land data is resampled to a spatial resolution of 0.05° using bilinear interpolation.

[0120] S23. Perform spatial registration and time alignment processing on various types of data.

[0121] Using the grid containing the solar-induced chlorophyll fluorescence data as a reference, spatial alignment of various input variables is performed, ensuring that different variables correspond to the same spatial grid in the same month, thereby constructing a standardized input dataset.

[0122] S3: Construct a set of features related to water stress.

[0123] Specifically, constructing a set of water stress features includes the following steps:

[0124] S31. Calculate the vapor pressure deficit (VPD) based on the 2-meter air temperature and the 2-meter dew point temperature.

[0125] It should be noted that VPD is used to characterize the degree of atmospheric dryness. When VPD increases, vegetation stomatal conductance usually decreases, thereby limiting carbon dioxide entry into leaves and inhibiting photosynthesis. Therefore, VPD is a key constraint variable in estimating total primary productivity in arid areas.

[0126] S32. Construct a heat wave index based on daily maximum temperatures.

[0127] If a heatwave index for a given month is assigned a value of 1 when there is at least one instance in a month where the daily maximum temperature exceeds the threshold temperature T0 for three consecutive days, otherwise it is assigned a value of 0.

[0128] Preferably, the threshold temperature T0 is in the range of 30℃ to 40℃, and more preferably 35℃.

[0129] It should be added that heat waves not only directly affect the activity of photosynthetic enzymes, but also exacerbate transpiration and increase vegetation water deficit. Therefore, heat wave indicators can supplement the characterization of drought stress from the perspective of extreme high temperature events.

[0130] S33. Extract AI, SPEI and soil moisture variables for different soil layers.

[0131] Among them, the drought index AI is used to characterize the long-term dry and wet climate background, the standardized precipitation evapotranspiration index SPEI is used to characterize short- and medium-term abnormal changes in water, and soil moisture variables of different soil layers are used to characterize the water supply capacity of different depths of vegetation root zone.

[0132] Preferably, the soil moisture variables include surface soil moisture (SWVL2), middle soil moisture (SWVL3), and deep soil moisture (SWVL4).

[0133] It should be noted that AI, SPEI, and soil moisture variables characterize vegetation water stress from three aspects: long-term climate background, short-term dry-wet fluctuations, and root zone water supply capacity, respectively.

[0134] S34, Forming a set of water stress characteristics.

[0135] The VPD, heat wave index, AI, SPEI, and soil moisture variables of different soil layers are combined to form a water stress feature set, which is used as part of the subsequent model input.

[0136] S4: Spatial representativeness screening of flux site samples.

[0137] Specifically, the spatial representativeness screening of flux site samples includes the following steps:

[0138] S41. Using the 0.05° solar-induced chlorophyll fluorescence pixel where the flux station is located as the analysis unit, extract the multi-period land cover data and monthly enhanced vegetation index (EVI) data of the corresponding pixel.

[0139] S42. Calculate the dominant land cover consistency and vegetation stability indices for the pixels where the site samples are located.

[0140] This involves statistically analyzing the coverage proportions of dominant land cover types across multiple time phases and calculating the mean of the monthly EVI standard deviation.

[0141] S43. Determine the valid training samples based on the screening criteria.

[0142] When the pixel containing the sample meets the following conditions, the corresponding site sample is determined to be a valid training sample:

[0143] 1) The dominant land cover type has an average coverage rate of over 50% in at least four of the multiple time phases;

[0144] 2) The mean monthly EVI standard deviation is less than 0.1.

[0145] Otherwise, remove the sample from that site.

[0146] It should be noted that the above screening steps are used to reduce the scale mismatch error between flux station scale observations and remote sensing pixel scale features, and improve the spatial representativeness of the training samples.

[0147] S5: Based on AI, climate zones are created, and ecological stratification is performed in conjunction with vegetation types.

[0148] Specifically, the process of climate zoning based on AI and ecological stratification based on vegetation type includes the following steps:

[0149] S51. Based on the drought index AI, the effective training samples are divided into climate zones.

[0150] When AI ≤ 0.5, the samples are classified as arid zone samples; when AI > 0.5, the samples are classified as non-arid zone samples.

[0151] S52. After completing the climate zoning, ecological stratification is carried out according to the vegetation type of the sample, and training subsets of different ecological layers are constructed respectively.

[0152] S53. Select an ecological layer training subset with a sample size of no less than 50 for subsequent model training.

[0153] It should be added that by applying climate zoning and ecological stratification, the model's adaptability to different climate backgrounds and different ecosystem types can be enhanced.

[0154] S6: Select representative soil moisture factors.

[0155] Specifically, selecting representative soil moisture factors includes the following steps:

[0156] S61. Use soil moisture variables from different soil layers as candidate input features to participate in model training.

[0157] The candidate input features include at least SWVL2, SWVL3, and SWVL4.

[0158] S62. Calculate the feature importance of each candidate soil moisture variable based on the model training results.

[0159] Preferably, the feature importance evaluation results of the random forest model are used as the basis for judgment.

[0160] S63. The candidate soil moisture variables with the highest feature importance are identified as representative soil moisture factors.

[0161] When the characteristic importance of SWVL3 is higher than that of SWVL2 and SWVL4, SWVL3 is selected as the representative soil moisture factor.

[0162] It should be noted that SWVL3 corresponds to the moisture content of the middle soil layer, which is usually closer to the depth of the main root system of the vegetation, and therefore can better reflect the actual available water status of the vegetation.

[0163] S7: Construct a model for estimating total primary productivity.

[0164] Specifically, constructing a total primary productivity estimation model includes the following steps:

[0165] S71. Construct model input datasets for samples from arid areas, non-arid areas, and ecological layer training subsets of each vegetation type.

[0166] The solar-induced chlorophyll fluorescence characteristics, meteorological driving characteristics, water stress characteristics, vegetation structure characteristics, and topographic characteristics were used as input variables, and the GPP observation values ​​of flux stations were used as label values.

[0167] S72. A random forest regression model is used to establish a nonlinear mapping relationship between input features and total primary productivity.

[0168] The random forest regression model consists of multiple decision trees, is trained using Bootstrap sampling with replacement, and uses mean squared error as the node splitting criterion.

[0169] It should be added that the random forest model can effectively characterize the complex nonlinear relationship under the combined effect of multiple factors such as light energy drive, water limitation and environmental regulation.

[0170] S73. Optimize the hyperparameters of the random forest model through grid search, and evaluate the model's generalization ability using five-fold cross-validation.

[0171] Preferably, the number of decision trees n_estimators ranges from 100 to 1000, the maximum depth max_depth ranges from 5 to 20, and the minimum number of leaf samples min_samples_leaf ranges from 1 to 5.

[0172] S74. Determine the optimal estimation model for each ecological layer.

[0173] The optimal combination of hyperparameters is determined based on the cross-validation results, and the final model training is completed.

[0174] S8: Output the total primary productivity estimate.

[0175] Specifically, outputting the total primary productivity estimate includes the following steps:

[0176] S81. Apply the trained optimal models of each ecological layer to the monthly grid data of the target area.

[0177] S82. Output the monthly or annual total primary productivity estimate results for the target region.

[0178] S83. Output the feature importance ranking and performance evaluation index corresponding to the model.

[0179] It should be noted that the performance evaluation indicators preferably include the coefficient of determination R2, root mean square error RMSE, and mean absolute error MAE.

[0180] Example 2

[0181] This embodiment is basically the same as Embodiment 1, except that it expands on the selection of data sources and feature variables.

[0182] S1: Acquire multi-source data. Among them, solar-induced chlorophyll fluorescence data can be obtained from GOSIF, TROPOMI SIF, or OCO-2 SIF products; vegetation structure parameter data can be obtained from GLASS or MODIS LAI products; and meteorological driving data can be obtained from other reanalysis datasets besides ERA5-Land.

[0183] S2: Construct an extended feature set. Based on Example 1, in addition to AI and SPEI, standardized soil moisture index SSI, evapotranspiration variables, or vegetation transpiration-related variables can be further introduced to enhance the characterization of water-limiting processes in arid regions.

[0184] S3: Construct the estimation model and output the results. The same modeling process as in Example 1 is used to complete the training and prediction, and the remaining steps are the same as in Example 1.

[0185] The beneficial effects of this invention are as follows:

[0186] 1. This invention constructs a multi-source data-driven method for estimating total primary productivity, which integrates solar-induced chlorophyll fluorescence data, flux station observation data, vegetation structure parameters, meteorological driving data, water status data, and topographic data, and unifies them to the same time scale and spatial resolution, thereby achieving consistent expression of different data sources in the spatiotemporal dimension and improving the consistency and reliability of input data.

[0187] 2. This invention effectively alleviates the mismatch between flux station point-scale observation and remote sensing pixel scale by introducing a station representativeness screening mechanism based on multi-period land cover consistency and vegetation index fluctuation characteristics, thereby improving the spatial representativeness of training samples and enhancing the effectiveness of model training.

[0188] 3. This invention constructs a zoning-layer modeling system by using a climate zoning based on drought index and an ecological stratification strategy based on vegetation type, combined with the characteristic importance analysis and optimization of soil moisture variables in different soil layers. This system effectively characterizes the ecological heterogeneity of arid areas and improves the adaptability and stability of the model under different climatic conditions and ecosystem types.

[0189] 4. This invention introduces water stress variables such as vapor pressure deficit, standardized precipitation evapotranspiration index, and heat wave index, and uses a random forest regression model to establish a nonlinear mapping relationship between multi-source driving variables and total primary productivity. Combined with grid search and cross-validation for model optimization, it can effectively characterize the coupling effect of light energy driving, water limitation and environmental regulation, thereby significantly improving the accuracy and stability of total primary productivity estimation.

[0190] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for estimating total primary productivity in arid regions, characterized in that, Includes the following steps: S1: Acquire solar-induced chlorophyll fluorescence data, total primary productivity observation data from flux stations, vegetation structure parameter data, meteorological driving data, water status data, and topographic data for the target area; S2: Preprocess the multi-source data acquired in step S1, unifying various input variables to a monthly scale. Spatial resolution was achieved, and coordinate system unification, spatial registration, and temporal alignment were completed; S3, saturated vapor pressure difference (VPD) was calculated based on 2-meter air temperature and 2-meter dew point temperature, a heat wave index was constructed based on daily maximum temperature, and drought index (AI), standardized precipitation evapotranspiration index (SPEI), and soil moisture variables of different soil layers were extracted to form a set of water stress characteristics; S4, based on the location of flux stations... Using solar-induced chlorophyll fluorescence pixels as the analysis unit, spatial representativeness screening of site samples is performed by combining multi-period land cover data and monthly enhanced vegetation index data: the average coverage ratio of the dominant land cover type in multiple time phases is calculated, and the mean of the monthly EVI standard deviation is calculated; only when the average coverage ratio meets the preset spatial homogeneity threshold and the mean of the monthly EVI standard deviation meets the preset temporal stability threshold, the corresponding sample is determined as a valid training sample; S5, the valid training samples are divided into climate zones according to the drought index AI. When the drought index AI is less than or equal to 0.5, it is classified as arid zone sample, and the rest are classified as non-arid zone sample; based on the climate zone, ecological stratification is performed according to vegetation type to construct training subsets of different ecological layers; S6, the surface soil moisture SWVL2 and the middle soil moisture are... SWVL3 and deep soil moisture SWVL4 are used as candidate input features in training, and the candidate soil moisture variables are optimized through model feature importance analysis. Specifically, for the subset of samples from the arid area, when the feature importance of middle soil moisture SWVL3 is higher than that of surface soil moisture SWVL2 and deep soil moisture SWVL4, middle soil moisture SWVL3 is selected as the representative soil moisture factor. S7, Total Primary Productivity estimation models are constructed for the arid area, non-arid area, and ecological layers of each vegetation type. A random forest regressor is used as the estimator, the hyperparameter combination is determined through grid search, and the model is trained using five-fold cross-validation. S8, the trained ecological layer models are applied to the monthly grid data of the target area, and the monthly or annual total primary productivity estimation results of the target area are output.

2. The method according to claim 1, characterized in that, The formula for calculating the saturated vapor pressure difference (VPD) in step S3 is as follows: ; Where T is the air temperature (°C), Td is the dew point temperature (°C), and VPD is in kPa.

3. The method according to claim 1, characterized in that, The formula for calculating the drought index AI in step S3 is as follows: ; For precipitation, Potential evapotranspiration. The drought index is used to characterize the long-term dry and wet conditions of the region, and when the drought index AI ≤ 0.5, it characterizes drought conditions.

4. The method according to claim 1, characterized in that, The method for constructing the heatwave index in step S3 is as follows: The number of consecutive days of high temperatures within a month is counted. If there is a continuous period of no less than 3 days with a daily maximum temperature of no less than 35 degrees Celsius, the heat wave index for that month is assigned a value of 1; otherwise, it is assigned a value of 0.

5. The method according to claim 1, characterized in that, The formula for calculating the average coverage proportion P_dom of the dominant land cover type in multiple time phases in step S4 is as follows: ; in: The average coverage ratio of the dominant land cover type; For the first The proportion of the dominant type in each phase; The number of phases.

6. The method according to claim 1, characterized in that, The formula for calculating the mean of the monthly EVI standard deviation, Sigma_EVI, in step S4 is as follows: ; in: For the first Standard deviation of monthly EVI; This refers to the number of months.

7. The method according to claim 1, characterized in that, The feature importance mentioned in step S6 is calculated based on the mean squared error (MSE), and the formula for calculating the mean squared error (MSE) is as follows: ; in, To observe GPP, These are predicted values.

8. The method according to claim 1, characterized in that, Step S6 specifically includes: In the training subset of samples from arid regions, the feature importance scores of surface soil moisture (SWVL2), mid-layer soil moisture (SWVL3), and deep soil moisture (SWVL4) were compared. When the feature importance of mid-layer soil moisture (SWVL3) was the highest, mid-layer soil moisture (SWVL3) was determined as the representative soil moisture factor to reflect the water availability of the root zone, the main root distribution layer of vegetation in arid regions.

9. A total primary productivity estimation system for arid regions, characterized in that, include: The data acquisition and preprocessing module is used to acquire solar-induced chlorophyll fluorescence data, total primary productivity observation data from flux stations, vegetation structure parameter data, meteorological driving data, water status data, and topographic data. It performs time-scale unification to a monthly scale, spatial resolution unification to 0.05 degrees, coordinate system unification, spatial registration, and temporal alignment processing on the multi-source data. The water stress feature construction module is used to calculate the saturated vapor pressure difference (VPD) based on 2-meter air temperature and 2-meter dew point temperature, and to construct a heat wave index based on daily maximum temperatures. It also extracts the drought index (AI), standardized precipitation evapotranspiration index (SPEI), and other parameters. Soil moisture variables within the same soil layer are used to construct a set of water stress characteristics. A site matching and screening module is used to spatially screen site samples based on the 0.05-degree solar-induced chlorophyll fluorescence pixel where the flux station is located, combined with multi-period land cover data and monthly enhanced vegetation index data. This involves calculating the average coverage ratio of the dominant land cover type across multiple time phases and the mean of the monthly EVI standard deviation. Only when the average coverage ratio meets a preset spatial homogeneity threshold and the mean of the monthly EVI standard deviation meets a preset temporal stability threshold is the corresponding sample considered a valid training sample. The module is divided into four parts: a partitioning and stratification module, which uses the drought index AI to divide the effective training samples into arid and non-arid areas, and then performs ecological stratification according to vegetation type to construct training subsets for different ecological layers; a feature selection module, which uses soil moisture variables of different soil layers as candidate input features to participate in model training, and determines representative soil moisture factors based on the importance evaluation of model features; specifically, for arid area samples, when the feature importance of the middle soil moisture SWVL3 is higher than that of the surface and deep soil moisture, the middle soil moisture SWVL3 is selected as the representative soil moisture factor; a model optimization and construction module, which uses the random forest regression algorithm to establish a nonlinear mapping relationship between input features and total primary productivity, and optimizes the model hyperparameters through grid search and five-fold cross-validation to construct a partitioned-stratified estimation model; and a regression prediction and result output module, which applies the trained model to monthly grid data of the target area and outputs the monthly or annual total primary productivity results of the target area.