A method and system for simulating net primary productivity using a CASA model

CN122653017APending Publication Date: 2026-08-28CHONGQING INST OF GEOLOGY & MINERAL RESOURCES +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610778034.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

然而,传统CASA模型在模拟NPP时,高度依赖大量经验性静态参数(如最大光能利用率、温度及水分胁迫系数等),这些参数通常依据全球或区域平均水平设定,或依赖少量实测数据人工赋值,难以真实反映不同植被类型在复杂地理环境下的空间异质性与生态适应性

Benefits of technology

[0008] Compared with existing technologies, the present invention provides a method for simulating net primary productivity using a CASA model, which can achieve intelligent prediction and adaptive configuration of static parameters of the CASA model in large-scale, multi-vegetation-type regions, thereby improving the accuracy, automation, and cross-regional generalization ability of net primary productivity simulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653017A_ABST
    Figure CN122653017A_ABST
Patent Text Reader

Abstract

The application discloses a CASA model net primary productivity simulation method and system, and the method comprises the following steps: collecting real-time position data of a side mold, pressure data of a hydraulic system and temperature data of a mold contact surface through a multi-sensor fusion module; based on the real-time position data, the pressure data and the temperature data, a dynamic friction and thermal deformation coupling model in a side mold movement process is constructed, and a position deviation trend in a current stroke stage is predicted; according to the position deviation trend and a preset mold closing accuracy requirement, a dynamic compensation strategy of the side mold stroke based on a fuzzy adaptive PID is generated; the dynamic compensation strategy is executed, the hydraulic cylinder flow and pressure are adjusted in real time through a high-frequency servo valve group, and position closed-loop correction is completed within one stroke cycle. By using the embodiment of the application, intelligent prediction and adaptive configuration of CASA model static parameters in a large range and a multi-vegetation type region can be realized, and the precision, the automation degree and the cross-region generalization capability of net primary productivity simulation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of ecological remote sensing technology, and in particular to a method and system for simulating net primary productivity using the CASA model. Background Technology

[0002] Net primary productivity (NPP) is a core indicator for assessing the carbon budget and climate regulation function of ecosystems. Accurate simulation of NPP is crucial for global change research, ecological resource management, and the achievement of dual-carbon goals. Currently, CASA models based on remote sensing processes are widely used due to their clear mechanisms and readily available input data. However, traditional CASA models rely heavily on a large number of empirical static parameters (such as maximum light use efficiency, temperature, and water stress coefficients) when simulating NPP. These parameters are typically set based on global or regional averages or manually assigned values ​​based on limited measured data, making it difficult to accurately reflect the spatial heterogeneity and ecological adaptability of different vegetation types in complex geographical environments. Existing methods suffer from significant problems when handling parameter configurations across large areas with multiple vegetation types, including strong subjectivity, weak generalization ability, and difficulty in dynamic updates. This results in low NPP simulation accuracy and limits the reliable application of these models in different ecological scenarios. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for simulating net primary productivity using a CASA model, in order to overcome the shortcomings of the prior art. This method enables intelligent prediction and adaptive configuration of static parameters of the CASA model for large-scale, multi-vegetation-type regions, thereby improving the accuracy, automation, and cross-regional generalization ability of net primary productivity simulation.

[0004] One embodiment of this application provides a method for simulating net primary productivity using a CASA model, the method comprising: Acquire basic data for the study area, including at least remote sensing data, vegetation type data, meteorological data, and sample plot observation data; The basic data is preprocessed to obtain a standardized input dataset with uniform spatial resolution, uniform coordinate system and uniform time scale; A vegetation type static parameter sample library is constructed based on the standardized input dataset. The vegetation type static parameter sample library includes at least a vegetation type identifier, sample feature variables corresponding to the vegetation type, and target static parameter values. An artificial intelligence model is used to establish a mapping relationship between the sample feature variables and the target static parameter values. Based on the trained artificial intelligence model, the static parameters corresponding to different vegetation types in the study area are automatically predicted to obtain the vegetation type static parameter configuration results. The static parameter configuration results of the vegetation type are input into the CASA model, and the net primary productivity at the pixel scale or regional scale is calculated by combining the temperature stress factor and the water stress factor to obtain the net primary productivity simulation results.

[0005] Another embodiment of this application provides a CASA model net primary productivity simulation system, the system comprising: The acquisition module is used to acquire basic data of the study area, which includes at least remote sensing data, vegetation type data, meteorological data, and sample plot observation data. The processing module is used to preprocess the basic data to obtain a standardized input dataset with uniform spatial resolution, uniform coordinate system and uniform time scale; The construction module is used to construct a vegetation type static parameter sample library based on the standardized input dataset. The vegetation type static parameter sample library includes at least a vegetation type identifier, sample feature variables corresponding to the vegetation type, and target static parameter values. The prediction module is used to establish the mapping relationship between the sample feature variables and the target static parameter values ​​using an artificial intelligence model, and to automatically predict the static parameters corresponding to different vegetation types in the study area based on the trained artificial intelligence model, so as to obtain the vegetation type static parameter configuration results. The simulation module is used to input the static parameter configuration results of the vegetation type into the CASA model, and calculate the net primary productivity at the pixel scale or the regional scale by combining the temperature stress factor and the water stress factor, so as to obtain the net primary productivity simulation results.

[0006] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.

[0007] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.

[0008] Compared with existing technologies, the present invention provides a method for simulating net primary productivity using a CASA model, which can achieve intelligent prediction and adaptive configuration of static parameters of the CASA model in large-scale, multi-vegetation-type regions, thereby improving the accuracy, automation, and cross-regional generalization ability of net primary productivity simulation. Attached Figure Description

[0009] Figure 1 A hardware structure block diagram of a computer terminal for a CASA model net primary productivity simulation method provided in an embodiment of the present invention; Figure 2 A flowchart illustrating a method for simulating net primary productivity using a CASA model, as provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a CASA model net primary productivity simulation system provided in an embodiment of the present invention. Detailed Implementation

[0010] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0011] This invention first provides a method for simulating net primary productivity using a CASA model. This method can be applied to electronic devices, such as computer terminals, specifically ordinary computers.

[0012] The following detailed explanation uses a computer terminal as an example. Figure 1 This is a hardware structure block diagram of a computer terminal for a CASA model net primary productivity simulation method provided in an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0013] See Figure 2 The embodiments of the present invention provide a method for simulating net primary productivity using a CASA model, which may include the following steps: S201, Obtain basic data of the study area, including at least remote sensing data, vegetation type data, meteorological data, and sample plot observation data; Specifically, remote sensing image data covering the study area can be obtained through a remote sensing data service platform, including normalized vegetation index, enhanced vegetation index, leaf area index, surface reflectance and surface temperature data, to generate a raw remote sensing image dataset. The core of this step is to acquire remote sensing data representing the vegetation growth status and land surface features within the study area, providing a basic data source for subsequent sample feature extraction and static parameter inversion, ensuring complete data coverage and comprehensive parameter types. The specific implementation method is as follows: The remote sensing data service platform is the core channel for acquiring remote sensing image data. It can provide multi-source, multi-resolution remote sensing data. The platform adopts standardized data transmission protocols to ensure the stability and integrity of data downloads, and the data transmission rate is controlled above 10MB / s to avoid data loss due to transmission interruptions. Before acquisition, the boundary of the study area must be clearly defined using latitude and longitude coordinates.

[0014] The acquired remote sensing image data contains five core parameters. The meaning, acquisition requirements, and examples of each parameter are as follows: Normalized Difference Vegetation Index (NDVI) is a core indicator reflecting vegetation cover and growth status. The calculation formula is NDVI=(NIR-R) / (NIR+R), where NIR is the near-infrared reflectance and R is the red reflectance, with a value range of -1 to 1. A larger value indicates higher vegetation cover. The acquired NDVI data has a spatial resolution of 30m and a temporal resolution of 16 days, covering all 12 months of the study area to ensure the capture of vegetation growth dynamics throughout the year. Enhanced Vegetation Index (EVI) is a vegetation index optimized based on NDVI, which can reduce interference from atmospheric aerosols and soil background. The calculation formula is EVI=2.5×(NIR-R) / (NIR+6×R-7.5×B+1), where B is the blue reflectance, with a value range of -1 to 1. The acquired EVI data has the same resolution and time scale as the NDVI data and is used to supplement the characterization of vegetation growth status.

[0015] Leaf area index (LAI) is the ratio of the total leaf area of ​​plants per unit land area to the land area, measured in m² / m², and ranging from 0 to 10. A higher LAI value indicates denser vegetation. The acquired LAI data has a spatial resolution of 30m and a temporal resolution of one month, generated using a remote sensing inversion algorithm with an inversion accuracy controlled within ±0.5 to ensure data accuracy. Surface reflectance is the ratio of solar radiation reflected by a surface object to incident solar radiation, ranging from 0 to 1. The acquired data includes red, green, blue, near-infrared, and shortwave infrared data. The data consists of several bands, each with wavelength ranges of 0.63-0.69 μm, 0.52-0.58 μm, 0.45-0.52 μm, 0.85-0.87 μm, and 1.55-1.75 μm, with a spatial resolution of 30 m and a temporal resolution of 16 days, used to extract surface feature information. Surface temperature (LST) is the actual temperature of the surface, measured in °C. The acquired data has a spatial resolution of 100 m and a temporal resolution of 1 day, generated through thermal infrared remote sensing inversion, with an inversion accuracy controlled within ±1 °C, used to characterize the surface thermal environment.

[0016] After all remote sensing image data is acquired, it is categorized and organized according to parameter type. Key information such as parameter name, spatial resolution, temporal resolution, wavelength range (if applicable), and inversion accuracy of each data is labeled. The data is then uniformly saved in TIFF format to form the original remote sensing image dataset. This ensures that various types of data can be quickly identified and retrieved during subsequent preprocessing. The dataset size is controlled within 50GB to facilitate storage and processing.

[0017] Extract spatial distribution data of vegetation types in the study area from vegetation classification thematic maps or land use cover products to generate the original vegetation type dataset; The core of this step is to obtain spatial distribution information of different vegetation types within the study area, which serves as the basis for classifying the static parameter sample database of vegetation types, ensuring accurate vegetation type division and clear spatial distribution. The specific implementation method is as follows: Thematic vegetation classification maps and land use cover products are the main sources of spatial distribution data for vegetation types. Thematic vegetation classification maps are distribution maps of vegetation types drawn for specific areas with high classification accuracy. Land use cover products, on the other hand, are large-scale, standardized land use classification data. The appropriate data source can be selected according to the scale and accuracy requirements of the study area. In the example, a 1:100,000 thematic vegetation classification map of the study area is selected. This thematic map uses the Chinese vegetation classification system, which divides vegetation into 5 major categories and 12 subcategories: arbor forests, shrub forests, grasslands, cultivated land, and wetlands. The classification accuracy reaches more than 90%, which can meet the needs of sample database construction.

[0018] The extraction process requires initial spatial matching of vegetation classification thematic maps or land use cover products to ensure that their coordinate systems are initially consistent with remote sensing image data, avoiding spatial misalignment. A spatial cropping tool is then used, with the study area boundary as the cropping range, to remove vegetation type data outside the area while retaining vegetation type information within the study area. The cropping precision is controlled within one pixel to ensure the integrity of the spatial distribution of vegetation types.

[0019] The cropped vegetation type data is encoded, and a unique code is assigned to each vegetation type to facilitate subsequent sample classification and identification. The code uses two digits. The major category code is 01-05, which corresponds to arbor forest, shrub forest, grassland, cultivated land and wetland respectively. The sub-category code adds one digit to the major category code. In the example, arbor forest (01) includes three sub-categories: evergreen broad-leaved forest (011), deciduous broad-leaved forest (012) and coniferous forest (013). Shrub forest (02) includes two sub-categories: evergreen shrub (021) and deciduous shrub (022), and so on.

[0020] After coding is completed, supplement the detailed information of vegetation type, including vegetation type name, main dominant species, growth environment characteristics, etc. In the example, the main dominant species of evergreen broad-leaved forest (011) are camphor tree and nanmu tree, and the growth environment is mountainous area with an altitude of 200-800m and an annual precipitation of more than 1200mm; the main dominant species of deciduous broad-leaved forest (012) are poplar and willow tree, and the growth environment is plains and hills with an altitude of 100-500m.

[0021] Finally, the processed vegetation type spatial distribution data, coding information, and detailed descriptions are organized and integrated, and saved as vector format to generate the original vegetation type dataset. The dataset includes a vegetation type spatial vector layer, a coding lookup table, and a vegetation type description document, ensuring that vegetation type identifiers can be quickly extracted in the future, providing a classification basis for the construction of the sample library.

[0022] Monthly solar radiation, average temperature, maximum temperature, minimum temperature, precipitation, and evapotranspiration data for the study area were obtained through meteorological reanalysis data and ground meteorological stations, generating a raw meteorological element dataset. The core of this step is to acquire meteorological factor data affecting vegetation growth within the study area, providing a foundation for static parameter inversion and subsequent calculation of stress factors in the CASA model, and ensuring the temporal continuity and spatial integrity of meteorological data. The specific implementation method is as follows: Meteorological reanalysis data and ground meteorological station data are the core data sources for obtaining meteorological elements. Meteorological reanalysis data is standardized data generated by integrating multi-source meteorological observation data through numerical simulation methods. It has wide spatial coverage and strong temporal continuity. Ground meteorological station data is accurate meteorological data from on-site observations and can be used to correct the bias of reanalysis data. The combination of the two can ensure the accuracy of meteorological data.

[0023] The acquired meteorological elements include six core parameters. The meanings, units, and acquisition requirements of each parameter are as follows: Solar radiation (SOL) refers to the total solar radiation received per unit area; average temperature (T_avg) refers to the average monthly temperature in °C, reflecting the temperature environment for vegetation growth; precipitation (P) refers to the monthly rainfall.

[0024] After all meteorological data was acquired, the time scale was standardized, and all data was organized into monthly data. The acquisition time, data source, accuracy, and unit of each data point were labeled. Data with deviations were corrected, and missing data were filled in using linear interpolation methods to ensure the temporal continuity of the data. Subsequently, the six meteorological element data were integrated by month and spatial location to generate a raster format raw meteorological element dataset. Each raster cell corresponds to one meteorological element value, and the spatial resolution is consistent with the remote sensing data (30m), which facilitates subsequent spatial matching and fusion with other data.

[0025] Data from plot surveys and flux tower observations within the study area were collected, including plot biomass, plot productivity observations, leaf photosynthetic parameters, and flux tower carbon flux observations, to generate the original plot observation dataset.

[0026] The core of this step is to obtain measured data on vegetation growth within the study area for static parameter inversion and subsequent model accuracy verification. This ensures the representativeness and accuracy of the measured data and provides the foundation for constructing the sample library based on the target static parameter values. The specific implementation method is as follows: Plot survey data and flux tower observation data are the core means of obtaining measured information. Plot survey data is used to obtain vegetation biomass, productivity and leaf photosynthetic parameters, while flux tower observation data is used to obtain vegetation carbon flux information. The combination of the two can comprehensively reflect the growth status and physiological characteristics of vegetation, and provide accurate measured basis for static parameter inversion.

[0027] The typical sampling method was used for the sample plot survey. Based on the distribution of vegetation types in the study area, typical sample plots were selected in different growth environments (altitude, slope, soil type) for each vegetation type to ensure the representativeness of the sample plots. A total of 20 sample plots were selected in the study area, including 4 arbor forests, 4 shrub forests, 4 grasslands, 4 cultivated lands, and 4 wetlands. The area of ​​each sample plot was 10m×10m, and the distance between sample plots was not less than 1km to avoid mutual interference between sample plots.

[0028] During the sample plot survey, the core parameters measured included: sample plot biomass, which refers to the dry matter weight of vegetation per unit area, measured in kg / m², and was determined using the harvest method. All vegetation in the sample plot was harvested separately for aboveground and belowground parts, dried to constant weight, and then weighed with a weighing accuracy of 0.01 kg. Each sample plot was measured three times, and the average value was taken as the final biomass data. Sample plot productivity observation value, which refers to the increase in biomass of vegetation per unit time, measured in kg / (m²・a), was obtained through two consecutive years of sampling. Soil biomass was measured and the annual increase was calculated. Leaf photosynthetic parameters, including net photosynthetic rate, stomatal conductance, and intercellular CO2 concentration, were measured in μmol / (m²·s), mol / (m²·s), and μmol / mol. The measurements were taken in the field using a photosynthesis meter between 9:00 and 11:00 am on a sunny day. Five typical plants were selected from each plot, and three functional leaves from each plant were measured. The average value was taken as the leaf photosynthetic parameters for that plot.

[0029] Flux tower observation data were obtained by deploying flux towers within the study area. In this example, three flux towers were selected and deployed in three typical vegetation types: forest, grassland, and wetland. The observation height of the flux towers was 10-15m to ensure that the observation range covered the vegetation within a 100m radius. The observation parameter was the carbon flux observation value, expressed in gC / (m²·d), which reflects the net carbon absorption capacity of the vegetation. The observation frequency was once every 30 minutes, and the daily average was summarized and the monthly average was summarized to ensure the temporal continuity of the data. The observation accuracy was controlled within ±0.1gC / (m²·d).

[0030] After all sample plot survey data and flux tower observation data were collected, they were organized and integrated. The sample plot location (latitude and longitude coordinates), vegetation type, observation time, observation method, and accuracy of each data point were labeled. Outlier data were removed, and missing data were supplemented to ensure data completeness and accuracy. Subsequently, the data was categorized and organized according to data type to generate the original sample plot observation dataset, including sample plot survey data tables and flux tower observation data tables, providing a measured basis for the subsequent inversion of target static parameter values.

[0031] S202, preprocess the basic data to obtain a standardized input dataset with unified spatial resolution, unified coordinate system and unified time scale; Specifically, the original remote sensing image dataset, original vegetation type dataset, original meteorological element dataset, and original sample plot observation dataset can be processed using coordinate system one, and all data can be transformed to the same coordinate reference system using Albers equal area projection or universal transverse Mercator projection to generate a multi-source dataset with unified coordinates. The core of this step is to eliminate coordinate deviations from different data sources, ensuring that all data are aligned within the same spatial reference frame. This lays the foundation for subsequent spatial overlay, feature extraction, and parameter matching. The specific implementation method is as follows: Coordinate system processing is a prerequisite for multi-source data fusion. Different original datasets may have different coordinate systems. If these are not unified, spatial misalignment will occur, affecting the accuracy of subsequent data processing. This study uses the Universal Transverse Mercator (UTM) projection for coordinate transformation. This projection is an conformal transverse cylindrical projection, which has the advantages of no angular distortion and minimal length distortion, making it suitable for spatial data processing in mid-latitude regions.

[0032] During the conversion process, four types of original datasets were processed separately: the original remote sensing image dataset originally used the WGS84 geographic coordinate system (latitude and longitude coordinates), which was converted to UTM50N plane coordinates (unit: degrees) using a projection conversion algorithm, with the error controlled within one pixel during the conversion process; the original vegetation type dataset was in vector format and originally used the Beijing 54 coordinate system, which was first converted to the WGS84 geographic coordinate system and then further converted to UTM50N projected coordinates to ensure that the vector boundaries were accurately aligned with the remote sensing images; the original meteorological element dataset was in raster format, and the same conversion method as the remote sensing images was used to complete the coordinate conversion; the original sample plot observation dataset contained the latitude and longitude coordinates of the sample plots, which was directly converted to UTM50N plane coordinates, and the plane coordinate values ​​of each sample plot were labeled to facilitate subsequent spatial matching with the raster data.

[0033] After the transformation is completed, coordinate verification is performed on all datasets. Twenty verification points are randomly selected, and the coordinate deviation before and after the transformation is compared to ensure that the deviation does not exceed 0.05m. After the verification is qualified, the four types of datasets are integrated to generate a multi-source dataset with unified coordinates. All data are based on the UTM50N projection coordinate system to ensure that the spatial position of each data source is consistent and there is no misalignment in subsequent spatial operations.

[0034] Resampling is performed on the spatial data in a multi-source dataset with unified coordinates. The other data are then unified to the same spatial resolution based on the highest resolution data, generating a raster dataset with unified resolution. The core of this step is to eliminate the resolution differences between different spatial data and ensure that the pixel size of various raster data is consistent, which facilitates subsequent pixel-by-pixel feature extraction and parameter calculation. The specific implementation method is as follows: The core principle of resampling is "using the highest resolution data as a reference." This means selecting the dataset with the highest spatial resolution from the multi-source dataset as a reference, and resampling other lower-resolution datasets to that resolution to ensure a uniform pixel size across all raster data after resampling. In the example, the Normalized Difference Vegetation Index (NDVI) and Enhanced Vegetation Index (EDI) data in the original remote sensing image dataset have a spatial resolution of 30m, the highest among all spatial data. Therefore, 30m is used as the uniform resolution for resampling other spatial data.

[0035] The resampling method selects an appropriate algorithm based on the data type: For remote sensing image data (NDVI, EVI, surface reflectance, etc.), a bilinear interpolation algorithm is used. This algorithm calculates the target pixel value by weighted averaging of the values ​​of four adjacent pixels, resulting in high interpolation accuracy. It can effectively preserve the spatial continuity of the data, avoid jagged distortion after resampling, and control the interpolation error within 5%. Vegetation type data is in vector format and is first converted to raster format (30m resolution). The nearest neighbor interpolation algorithm is used, which can ensure the integrity of vegetation types, avoid type confusion, and ensure that the vegetation type of each pixel is consistent with the original vector data. The meteorological element dataset originally had a resolution of 1km. It is resampled to 30m using a bilinear interpolation algorithm, refining the large-scale meteorological data to each pixel, ensuring the spatial continuity of meteorological elements, and preserving the accuracy of the data.

[0036] During resampling, pixel alignment is strictly controlled to ensure that the pixel positions of various data types correspond one-to-one after resampling, avoiding pixel misalignment. After resampling, the data resolution is verified. Ten pixels are randomly selected, and the numerical deviations before and after resampling are compared to ensure that the deviation of remote sensing image data does not exceed 5%, the deviation of meteorological data does not exceed 10%, and there are no type errors in vegetation data. After verification, all raster data are integrated to generate a raster dataset with uniform resolution. The pixel size of all data is 30m×30m, providing a unified spatial basis for subsequent time-scale matching and feature extraction.

[0037] Time scale matching is performed on multi-temporal data in a raster dataset with uniform resolution, daily data are aggregated into monthly scale data, and meteorological interpolation data with different time frequencies are unified to the same time node to generate a time-matched multi-source dataset. The core of this step is to unify the time frequency of various multi-temporal data to ensure that different data sources are aligned in the time dimension, which facilitates subsequent sample matching and parameter calculation by time node. The specific implementation method is as follows: The core of time scale matching is to unify data with different time frequencies to the same time scale, which in this case is the monthly scale. At the same time, data from different time points are aligned to the same time point each month to ensure the temporal consistency of the data.

[0038] For daily meteorological data (such as daily precipitation and daily temperature), the arithmetic mean method is used to aggregate the data into monthly scale data. That is, the average value of all daily data within a month is calculated as the monthly scale data. For example, the average value of the daily average temperature within a month is the average temperature of that month, and the sum of the daily precipitation within a month is the precipitation of that month. Outliers in the daily data are removed during the aggregation process to ensure the accuracy of the monthly scale data. For 16-day resolution remote sensing image data, the maximum value synthesis method is used to aggregate the data into monthly scale data. That is, the maximum value among all 16-day image data within a month is selected as the remote sensing parameter value for that month. This can effectively highlight the best state of vegetation growth in that month and avoid data deviations caused by factors such as cloud cover.

[0039] Meteorological interpolation data at different time frequencies are unified to the 15th of each month as the time node. That is, all monthly scale data are identified by the 15th of each month. For sample plot observation data (quarterly observation), it is matched to the 15th of the middle month of the corresponding quarter. For example, the sample plot observation data of the second quarter (April-June) is matched to the time node of May 15th to ensure that the time nodes of the sample plot data and the raster data are consistent.

[0040] After time scale matching is completed, all data are time-verified to ensure that all types of data in each month have corresponding time nodes, with no missing or misaligned time. After integration, a multi-source dataset with time matching is generated. All data are on a monthly scale, and the time node is uniformly set to the 15th of each month, which facilitates subsequent sample matching based on time and spatial location.

[0041] Outlier removal, missing value imputation, continuous variable normalization, and vegetation type category variable encoding are performed on time-matched multi-source datasets to generate a standardized input dataset.

[0042] The core of this step is to clean and standardize the multi-source data, eliminating data noise and format differences, ensuring data quality, and providing qualified input data for subsequent sample library construction and model training. The specific implementation method is as follows: Outlier removal follows the 3σ principle: first, the mean (μ) and standard deviation (σ) of each continuous variable (such as NDVI, temperature, precipitation, etc.) are calculated; data exceeding the range of μ±3σ are identified as outliers and removed; after removal, neighboring pixel values ​​are used to fill the gaps to ensure data continuity; for category data such as vegetation type, there are no outliers, only data with spatially incorrect locations are removed.

[0043] Missing value imputation uses the appropriate method based on the data type: for continuous variables (such as temperature, precipitation, LAI, etc.), linear interpolation is used, which calculates the missing value by linearly fitting the valid data adjacent to it; for spatial data, the nearest neighbor interpolation method is used, which selects the average of the four neighboring pixels as the missing value to ensure the integrity of the spatial data. Min-max normalization is used for continuous variables, and the normalized data avoids the influence of unit differences on the weight allocation during subsequent model training.

[0044] The vegetation type categorical variable is encoded using a numerical coding method, assigning a unique numerical code to each vegetation type. The coding rules are consistent with the original vegetation type dataset, namely, forest 01, shrub forest 02, grassland 03, cultivated land 04, and wetland 05. Subcategories are extended based on the major category codes, such as evergreen broad-leaved forest 011 and deciduous broad-leaved forest 012. The encoded data can be directly identified and processed by the subsequent model, avoiding the problem that categorical variables cannot be directly input into the model.

[0045] After all processing is completed, the data is standardized and validated to ensure that outliers are completely removed, missing values ​​are properly imputed, continuous variables are properly normalized, and categorical variables are correctly coded. All processed data are then integrated to generate a standardized input dataset with a spatial resolution of 30m, a coordinate system of UTM50N, and a time scale of months. This dataset can be directly used to construct a sample library of static parameters for vegetation types.

[0046] S203, construct a vegetation type static parameter sample library based on the standardized input dataset. The vegetation type static parameter sample library includes at least a vegetation type identifier, sample feature variables corresponding to the vegetation type, and target static parameter values. Specifically, vegetation type identifiers can be extracted from standardized input datasets as a basis for sample classification, vegetation index features, surface reflectance features, surface temperature features and leaf area index features can be extracted from remote sensing data, temperature features, precipitation features and solar radiation features can be extracted from meteorological data, and elevation, slope and soil type features can be extracted from environmental factors to generate a sample feature variable matrix. The core of this step is to systematically extract various feature variables related to vegetation static parameters from the standardized input dataset, integrate them into a sample feature variable matrix according to unified rules, provide comprehensive and accurate input features for subsequent static parameter prediction, and ensure that the feature variables have a strong correlation with the target static parameters. The specific implementation method is as follows: The core basis for sample classification is the vegetation type identifier, which is extracted from the vegetation type data in the standardized input dataset. It adopts the digital coding rules set in the previous preprocessing stage, namely, forest 01, shrub forest 02, grassland 03, cultivated land 04, wetland 05. The subclass coding is extended based on the major categories, such as evergreen broad-leaved forest 011, deciduous broad-leaved forest 012. Each pixel or sample plot corresponds to a unique vegetation type identifier. During the extraction process, it is ensured that the identifier corresponds one-to-one with the vegetation type, without coding errors or confusion. This serves as the core basis for subsequent sample classification and grouping, ensuring that samples of different vegetation types are not mixed together.

[0047] Feature extraction from remote sensing data is a core component of sample feature variables. The four types of features extracted are key parameters characterizing vegetation growth status and the land surface environment. Specific extraction methods and examples are as follows: Vegetation index features primarily extract the Normalized Difference Vegetation Index (NDVI) and the Enhanced Vegetation Index (EVI). Monthly data from the 15th of each month are selected for extraction. The monthly mean, maximum value, and standard deviation are calculated for each pixel or plot, resulting in three feature values. In the example, the monthly mean NDVI of a certain arbor forest plot is 0.62, the maximum value is 0.78, and the standard deviation is 0.08; the monthly mean EVI is 0.55, the maximum value is 0.71, and the standard deviation is 0.07. These features reflect vegetation cover. The study extracted reflectance values ​​from five bands: red, green, blue, near-infrared, and short-wave infrared. Monthly average and maximum values ​​were extracted for each band, resulting in a total of 10 feature values. The wavelength ranges for these bands were 0.63-0.69 μm (red), 0.52-0.58 μm (green), 0.45-0.52 μm (blue), 0.85-0.87 μm (near-infrared), and 1.55-1.75 μm (short-wave infrared). In the example, the monthly average reflectance of a grassland plot was 0.18 with a maximum of 0.22, and the monthly average reflectance of the near-infrared band was 0.65 with a maximum of 0.72, reflecting the reflectance characteristics of the surface vegetation and soil.

[0048] Land surface temperature features are extracted from the monthly average, maximum, and minimum land surface temperature (LST) values ​​on the 15th of each month, totaling three features in °C, with an accuracy controlled within ±1 °C. In the example, the monthly average LST value of a wetland plot is 22.5 °C, the maximum value is 28.3 °C, and the minimum value is 17.8 °C, which can characterize the impact of the surface thermal environment on vegetation growth. Leaf area index (LAI) features are extracted from the monthly average and maximum values ​​on the 15th of each month, totaling two features in m² / m², with an accuracy controlled within ±0.5. In the example, the monthly average LAI value of a cultivated land plot is 2.8, and the maximum value is 3.5, which can reflect the density of vegetation leaves and is directly related to the photosynthetic capacity of vegetation.

[0049] Meteorological data feature extraction focuses on key meteorological factors affecting vegetation photosynthesis and growth, specifically including: temperature features: monthly average temperature, maximum temperature, and minimum temperature (3 features in total, in °C, with an accuracy controlled within ±0.5 °C; in the example, the monthly average temperature of a shrubland plot is 18.6 °C, the monthly average maximum temperature is 25.3 °C, and the monthly average minimum temperature is 11.9 °C); precipitation features: monthly average precipitation and cumulative precipitation (2 features in total, in mm, with an accuracy controlled within ±1 mm; in the example, the monthly average precipitation of the shrubland plot is 85.2 mm, and the cumulative precipitation is 1023 mm); solar radiation features: monthly average total solar radiation (MJ / m², with an accuracy controlled within ±5 MJ / m²; in the example, the monthly average solar radiation of the plot is 485 MJ / m²). These features directly determine the energy and water supply of the vegetation.

[0050] Environmental factor feature extraction focuses on spatial environmental conditions that affect vegetation growth, specifically including: elevation feature extraction of the altitude of each pixel or plot, in meters, with an accuracy controlled within ±1m. In the example, the elevation of a coniferous forest plot is 1250m; slope feature extraction of the slope angle of each pixel or plot, in degrees (°), with an accuracy controlled within ±0.1°. In the example, the slope of the coniferous forest plot is 15.3°. Slope can affect water retention and soil erosion, thus affecting vegetation growth; soil type features are extracted using digital coding, assigning unique codes to different soil types, such as brown soil 01, brown earth 02, and alluvial soil 03. In the example, the soil type of the coniferous forest plot is brown soil, coded as 01. Soil type directly affects the nutrient supply capacity of vegetation.

[0051] After all features are extracted, they are integrated according to a unified rule to generate a sample feature variable matrix. Each row of the matrix corresponds to a sample unit (a single pixel or a single plot), and each column corresponds to a feature variable. The first column of the matrix is ​​the vegetation type identifier, and the subsequent columns are remote sensing features, meteorological features, and environmental factor features, respectively. It contains a total of 1 (vegetation type identifier) ​​+ 3 + 10 + 3 + 2 (remote sensing features) + 3 + 2 + 1 (meteorological features) + 1 + 1 + 1 (environmental factor features) = 27 feature columns. In the example, the data in a certain row of the matrix is: 011 (evergreen broad-leaved forest), 0.62 (monthly average NDVI), 0.78 (maximum NDVI), 0.08 (NDVI standard deviation) ... 1250 (elevation), 15.3 (slope), 01 (soil type). The matrix format is standardized to ensure that the subsequent artificial intelligence model can directly read and process it.

[0052] The maximum light energy utilization rate was obtained by inverting the leaf photosynthesis measurement data in the sample plot observation dataset. The upper limit parameter and lower limit parameter of the simple ratio were obtained by fitting the flux tower observation data with the CASA model, and a set of target static parameter values ​​was generated. The core of this step is to obtain the true values ​​of vegetation static parameters through inversion of measured data and model fitting. These values ​​serve as target values ​​for subsequent training of the artificial intelligence model, ensuring the accuracy and reliability of the target static parameter values ​​and providing a foundation for establishing the mapping relationship between feature variables and target parameters. The specific implementation method is as follows: Maximum light energy utilization efficiency (ε_max) is the core static parameter of the CASA model. It refers to the amount of organic carbon that vegetation can convert per unit of photosynthetically active radiation under ideal conditions (no temperature or water stress), with the unit being gC / MJ. Its value is obtained by inverting leaf photosynthetic measurement data from the sample plot observation dataset. The inversion adopts the response model of leaf photosynthetic rate and photosynthetically active radiation, namely the rectangular hyperbola model. The calculation formula is: P_n=ε_max×PAR×(1-e^(-α×PAR / P_n_max)), where P_n is the net photosynthetic rate of leaves (unit: μmol / (m²・s)), PAR is the photosynthetically active radiation (unit: μmol / (m²・s)), α is the initial light energy utilization efficiency (unit: μmolCO2 / μmolPAR), and P_n_max is the maximum net photosynthetic rate (unit: μmol / (m²・s)).

[0053] During the inversion process, leaf photosynthesis measurement data from 9:00-11:00 AM on sunny days were selected from the sample plot observation data. Five typical plants were selected from each sample plot, and three functional leaves were measured from each plant, for a total of 15 sets of data. After removing outliers, the value of ε_max was solved by nonlinear fitting method. The fitting accuracy was controlled at R²≥0.85 to ensure the reliability of the inversion results. In the example, the photosynthetic data of leaves in an evergreen broad-leaved forest plot showed that when PAR was 1200 μmol / (m²·s), P_n was 18.5 μmol / (m²·s); when PAR was 1800 μmol / (m²·s), P_n was 25.3 μmol / (m²·s). The ε_max of this plot was obtained by fitting, which was 1.85 gC / MJ. The ε_max varied among different vegetation types: the ε_max range for arbor forest was 1.70-1.90 gC / MJ, for grassland it was 1.40-1.60 gC / MJ, and for cultivated land it was 1.60-1.80 gC / MJ.

[0054] The upper limit parameter of the simple ratio (SR_max) and the lower limit parameter of the simple ratio (SR_min) are key static parameters in the CASA model for calculating the proportion of photosynthetically active radiation absorbed by vegetation. SR_max is the maximum value of the simple ratio during the vigorous growth period of vegetation, and SR_min is the minimum value of the simple ratio during the stagnant growth period of vegetation. The formula for calculating the simple ratio (SR) is SR=NIR / R, where NIR is the reflectance in the near-infrared band and R is the reflectance in the red band.

[0055] The two parameters were obtained by fitting the flux tower observation data with the CASA model. The core logic of the fitting was to adjust the values ​​of SR_max and SR_min to make the carbon flux value simulated by the CASA model as close as possible to the carbon flux value observed by the flux tower. The least squares method was used for fitting, with the goal of minimizing the root mean square error (RMSE) between the simulated and observed values ​​and maximizing the coefficient of determination (R²). The fitting accuracy was controlled at R² ≥ 0.80 and RMSE ≤ 0.5 gC / (m²・d). In the example, the monthly average carbon flux observed by the flux tower of a certain grassland was 8.2 gC / (m²・d). By adjusting the values ​​of SR_max and SR_min through fitting, when SR_max = 5.8 and SR_min = 1.2, the carbon flux value simulated by the CASA model was 8.1 gC / (m²・d). At this time, RMSE = 0.08 gC / (m²・d) and R² = 0.86, which showed the best fitting effect. Therefore, SR_max = 5.8 and SR_min = 1.2 were determined for this grassland.

[0056] Significant differences exist in SR_max and SR_min among different vegetation types. For arbor forests, SR_max ranges from 6.0 to 7.0, and SR_min from 1.3 to 1.5; for shrub forests, SR_max ranges from 5.5 to 6.5, and SR_min from 1.2 to 1.4; and for wetlands, SR_max ranges from 5.0 to 6.0, and SR_min from 1.1 to 1.3. After calculating all static parameter values, a target static parameter value set is generated. Each element in the set corresponds to the target value of a sample unit, containing three parameters: ε_max, SR_max, and SR_min. Each parameter is labeled with its corresponding vegetation type identifier, sample plot location, and calculation method to ensure the traceability of parameter values. In the example, one element in the set is: vegetation type identifier 011 (evergreen broad-leaved forest), ε_max = 1.85 gC / MJ, SR_max = 6.5, and SR_min = 1.4.

[0057] The sample feature variable matrix is ​​matched and aligned with the target static parameter value set according to spatial location and time node. Each sample unit contains vegetation type identifier, feature variables and target static parameter values ​​to generate a sample unit set. The core of this step is to achieve precise matching between feature variables and target static parameter values, ensuring a one-to-one correspondence between the features of each sample unit and the target value. This provides complete sample data for subsequent AI model training and avoids training bias due to mismatches. The specific implementation method is as follows: The core principle of matching and alignment is "spatial consistency and temporal uniformity," meaning that the feature variables and target static parameter values ​​of each sample unit must correspond to the same spatial location and the same time point, ensuring a direct correlation between the two and accurately reflecting the correspondence between vegetation features and static parameters at that spatial location and time point. Spatial location matching employs a coordinate matching method, using the UTM50N plane coordinates in the standardized input dataset to compare the coordinates of each sample unit in the sample feature variable matrix with the coordinates of the corresponding plot or pixel in the target static parameter value set. The coordinate deviation is controlled within 0.05m to ensure complete spatial consistency.

[0058] The time node matching is based on monthly scale data on the 15th of each month. The feature values ​​of each month in the sample feature variable matrix are matched with the parameter values ​​of the corresponding month in the target static parameter value set. For sample plot observation data (quarterly observation), it is matched to the 15th of the middle month of the corresponding quarter to ensure that the time nodes are consistent. In the example, the sample plot static parameter values ​​of the second quarter (April-June) are matched to the feature variable values ​​of May 15 to avoid invalid samples due to time misalignment.

[0059] During the matching process, vegetation type identifiers are used as the initial screening criteria. The feature variable matrix and target parameter value set are grouped by vegetation type, and samples of the same vegetation type are matched one-to-one. Then, spatial coordinates and time nodes are used for further precise alignment, and samples with mismatched spatial locations or time nodes are eliminated. Each matched sample unit contains three core pieces of information: vegetation type identifier (e.g., 011 evergreen broad-leaved forest), complete feature variables (specific values ​​of 27 feature columns), and corresponding target static parameter values ​​(ε_max, SR_max, SR_min). These three are stored together to ensure that the information of each sample unit is complete and accurate.

[0060] The complete information of a sample unit in the example is as follows: vegetation type identifier 011, characteristic variables (monthly average NDVI 0.62, monthly average EVI 0.55, monthly average red reflectance 0.18, monthly average surface temperature 22.5℃, monthly average air temperature 18.6℃, monthly average precipitation 85.2mm, monthly average solar radiation 485MJ / m², elevation 1250m, slope 15.3°, soil type code 01, etc.), and target static parameter values ​​(ε_max=1.85gC / MJ, SR_max=6.5, SR_min=1.4). This sample unit fully reflects the correspondence between vegetation characteristics and static parameters of evergreen broad-leaved forest at a specific spatial location and time point.

[0061] After all sample units are matched, they are integrated to form a sample unit set. Each sample unit in the set undergoes both spatial and temporal verification to ensure that there are no matching errors and no missing information. The number of sample units is determined according to the scale of the study area. In the example, a total of 1200 sample units were generated in the study area, including 300 arbor forests, 280 shrub forests, 260 grasslands, 240 cultivated lands, and 120 wetlands. The sample number is evenly distributed, which can fully cover different vegetation types and growth environments, and provide sufficient sample support for subsequent model training.

[0062] The sample unit set is quality-screened, abnormal samples are removed and representative samples are added. The sample units are organized and stored according to vegetation type and ecological zone to generate a static parameter sample library of vegetation type.

[0063] The core of this step is to optimize sample quality, ensure the representativeness and reliability of the samples, and organize and store them according to standards to form a standardized static parameter sample library for vegetation types. This provides high-quality sample data for subsequent artificial intelligence model training. The specific implementation method is as follows: Sample quality screening consists of two parts: outlier removal and representative sample supplementation. Outlier removal employs a comprehensive screening method, considering both characteristic variables and target static parameter values ​​to ensure the removal of all invalid samples. For characteristic variables, the 3σ principle is used to remove outliers. This involves calculating the mean (μ) and standard deviation (σ) of each characteristic variable, and identifying samples exceeding the range of μ ± 3σ as outliers. In the example, the NDVI characteristic has μ = 0.45 and σ = 0.12; samples with NDVI < 0.09 or NDVI > 0.81 are removed. For target static parameter values, samples exceeding reasonable ranges are removed based on the static parameter ranges for different vegetation types. In the example, the reasonable range for ε_max in arbor forests is 1.70-1.90 gC / MJ; samples with ε_max = 1.50 gC / MJ or ε_max = 2.00 gC / MJ are removed. Samples with missing characteristic variables or missing target parameter values ​​are also removed to ensure the completeness of information in each sample unit.

[0064] After outlier removal, a representativeness analysis is performed on the sample unit set to check for issues such as incomplete vegetation cover or missing growth environments. If the number of samples for a certain vegetation type is too small, or samples for a certain growth environment (such as high altitude or steep slope) are missing, representative samples are added. Supplementing samples uses a combination of field surveys and data interpolation. For vegetation types with insufficient samples, new plots of the corresponding vegetation type are surveyed to determine characteristic variables and target static parameter values. For missing growth environments, interpolation is performed using sample data from nearby similar environments. The supplemented samples must conform to the characteristics of the vegetation type and growth environment to ensure consistency with existing samples. For example, if the number of wetland samples is insufficient, 20 new wetland plots are added, and corresponding characteristic variables and target parameter values ​​are supplemented, bringing the total number of wetland samples to 140, matching the number of samples from other vegetation types.

[0065] After sample quality optimization, sample units are organized and stored according to vegetation type and ecological zone. The core of this organization is to facilitate subsequent model training by grouping by vegetation type, thereby improving model prediction accuracy. First, primary classification is performed according to major vegetation types (tree forest, shrub forest, grassland, cultivated land, wetland). Within each major category, secondary classification is performed according to subcategories (e.g., tree forest is divided into evergreen broad-leaved forest, deciduous broad-leaved forest, coniferous forest). Then, within each subcategory, tertiary classification is performed according to ecological zone (e.g., mountain, plain, hill), ensuring clear and hierarchical classification of sample units.

[0066] The storage format adopts a standardized format. Each sample unit is stored in the order of "vegetation type identifier - ecological zone - spatial coordinates - time node - feature variable - target static parameter value". At the same time, a sample index is established to mark the relevant information of each sample (such as collection time and calculation method) to facilitate subsequent quick query, retrieval and update. The storage path of a sample unit in the example is: arbor forest - evergreen broad-leaved forest - mountain - (523120m, 3856210m) - May 15 - feature variable - target parameter value. The index information includes the collection time of this sample as May 2024, ε_max is obtained by inversion of leaf photosynthetic data, and SR_max and SR_min are obtained by fitting flux tower data.

[0067] After all sample units are organized and stored, a static parameter sample library for vegetation types is generated. This sample library consists of three parts: a set of sample units, a sample classification index, and a sample quality verification report. The sample library contains 1280 samples, covering 5 major categories and 12 subcategories of vegetation types and 3 ecological zones. The sample quality verification report shows that the abnormal sample removal rate is 6.25%, 80 representative samples were added, and the sample qualification rate reached over 98%. It can be directly used for the training and verification of subsequent artificial intelligence models, providing a high-quality sample foundation for the automatic prediction of static parameters.

[0068] S204, using an artificial intelligence model to establish a mapping relationship between the sample feature variables and the target static parameter values, and based on the trained artificial intelligence model, automatically predict the static parameters corresponding to different vegetation types in the study area to obtain the vegetation type static parameter configuration results. Specifically, the static parameter sample library of vegetation type can be divided into a training sample set and a validation sample set. The cross-validation method is used to determine the sample division ratio and generate the model training dataset and validation dataset. The core of this step is to rationally divide the sample set and ensure the scientific and reasonable nature of the sample division through cross-validation. This avoids overfitting or underfitting of the model due to sample allocation bias, and provides reliable dataset support for subsequent model training and accuracy verification. The specific implementation method is as follows: The core premise of sample partitioning is to maintain the representativeness of the samples, ensuring that both the training and validation sets cover samples from different vegetation types and ecological zones, and that the sample distribution is consistent with the original sample database, avoiding the concentration of samples from a certain type of vegetation or ecological environment in a single dataset. Cross-validation is used to determine the sample partitioning ratio; in this study, 5-fold cross-validation was chosen. This method randomly divides the original sample database into five subsets of similar size. Four subsets are selected as training samples and one subset as validation samples each time, repeated five times. The optimal partitioning ratio is determined by the average accuracy of the five validations. This effectively reduces the bias caused by the randomness of sample partitioning and ensures the rationality of the partitioning ratio.

[0069] After five-fold cross-validation, the ratio of the training sample set to the validation sample set was determined to be 7:3. The training sample set comprises 70% of the total samples and is used for model training, fitting the mapping relationship between feature variables and target static parameters. The validation sample set comprises 30% of the total samples and is used for model accuracy verification, testing the model's generalization ability. In the example, the vegetation type static parameter sample library contains 1280 sample units. After being divided in a 7:3 ratio, the training sample set contains 896 samples, and the validation sample set contains 384 samples. The sample proportions for each vegetation type are consistent with the original sample library: 210 samples of arbor forest, 196 samples of shrub forest, 182 samples of grassland, 168 samples of cultivated land, and 98 samples of wetland, ensuring the representativeness of both training and validation samples.

[0070] During the partitioning process, samples were selected using random sampling to avoid duplicate or missing samples. The vegetation type identifier, spatial location, and time point of each sample were recorded to facilitate subsequent accuracy verification and result traceability. After partitioning, the integrity of the two sample sets was checked to ensure that neither the training nor the validation set contained missing information or outliers. These sets were then integrated to generate a model training dataset and a validation dataset. The training dataset included a matrix of sample feature variables and corresponding target static parameter values, while the validation dataset also contained complete feature variables and target values ​​in the same format as the training dataset, ensuring the consistency of subsequent model training and validation.

[0071] Construct an artificial intelligence model and train it using a training sample set. Use a random forest model, gradient boosting tree model, or extreme gradient boosting model to establish a nonlinear mapping relationship between sample feature variables and target static parameter values, and generate a trained parameter prediction model. The core of this step is to select a suitable artificial intelligence model, fit the nonlinear relationship between feature variables and target static parameters through a training sample set, and generate a parameter prediction model with predictive capabilities. This ensures that the model can accurately capture the intrinsic relationship between feature variables and static parameters. The specific implementation method is as follows: This study uses the Random Forest model to construct the parameter prediction model. This model belongs to the ensemble learning model category, consisting of multiple decision trees. It has advantages such as resistance to overfitting, strong ability to handle high-dimensional data, and no need for feature normalization (normalization has been performed in this study to further improve accuracy). It is suitable for situations where sample feature variables have multiple dimensions and the static parameters and feature variables exhibit a non-linear relationship. The core logic of the Random Forest model is to extract multiple sample subsets from the training sample set using the bootstrap sampling method. Each subset is used to train a decision tree, and the prediction results are finally output through voting or averaging, reducing the prediction error of a single decision tree and improving the model's stability and prediction accuracy.

[0072] During model construction, key hyperparameters are set to ensure effective model training. The specific hyperparameters and their meanings are as follows: The number of decision trees is set to 100. Too many decision trees increase training time, while too few reduce prediction accuracy. 100 trees achieve a balance between accuracy and efficiency. The maximum decision tree depth is set to 8. The maximum depth controls the complexity of the decision tree. Excessive depth can lead to overfitting, while insufficient depth can lead to underfitting. 8 layers effectively avoid both problems. The minimum number of samples per leaf node is set to 5, meaning each leaf node contains at least 5 samples to ensure representativeness and avoid prediction bias due to insufficient samples. The feature sampling ratio is set to 0.8, meaning 80% of the feature variables are randomly selected during training of each decision tree to enhance the model's generalization ability.

[0073] During model training, the sample feature variable matrix of the training sample set is used as input, and the target static parameter values ​​(maximum light energy utilization, upper limit parameter of simple ratio, lower limit parameter of simple ratio) are used as output. A parameter-by-parameter training approach is adopted to establish the mapping relationship between the feature variables and each static parameter. The model's fitting accuracy is monitored in real time during training. When the model's fitting accuracy on the training sample set stabilizes (accuracy change less than 0.01 over 10 consecutive iterations), training is stopped to avoid overfitting due to overtraining. In the example, by the 35th iteration, the fitting accuracy stabilized, the coefficient of determination R² on the training sample set reached 0.88, and the root mean square error RMSE was 0.07 gC / MJ. At this point, the model has fully captured the nonlinear relationship between the feature variables and the target parameters, generating a completed parameter prediction model.

[0074] The validation dataset is input into the trained parameter prediction model to verify its accuracy. The coefficient of determination and root mean square error between the predicted and true values ​​are calculated. If the validation accuracy does not reach the threshold, the model hyperparameters are adjusted and the model is retrained to generate an optimized parameter prediction model. The core of this step is to verify the generalization ability of the trained model. Accuracy metrics are used to determine if the model meets the requirements. If the accuracy is insufficient, the hyperparameters are optimized and the model is retrained to ensure the prediction accuracy and reliability of the final model. The specific implementation method is as follows: The core of accuracy validation lies in calculating two key metrics between the predicted and true values ​​of the validation dataset: the coefficient of determination (R²) and the root mean square error (RMSE). The R² measures the goodness of fit between the model's predicted and true values, ranging from 0 to 1. A closer R² to 1 indicates a better fit and higher prediction accuracy. The RMSE measures the average deviation between the predicted and true values, with units consistent with the target static parameter. A smaller RMSE indicates less prediction bias and higher reliability. The calculation logic for these two metrics is as follows: R² is obtained by comparing the sum of squared deviations between the true and predicted values ​​with the sum of squared deviations between the true and the mean true values; RMSE is obtained by taking the square root of the average of the squared deviations between the predicted and true values.

[0075] The accuracy thresholds set for this study are: R² ≥ 0.85, RMSE ≤ 0.10 gC / MJ (for maximum light energy utilization), and RMSE ≤ 0.3 (for the upper and lower limits of the simple ratio). These thresholds have been verified through multiple experiments and meet the accuracy requirements for simulating net primary productivity using the CASA model. The validation dataset is input into the trained parameter prediction model, which outputs the predicted static parameters for each sample. The predicted values ​​are compared with the actual values ​​in the validation dataset, yielding R² = 0.86, RMSE = 0.09 gC / MJ (maximum light energy utilization), RMSE = 0.25 (upper limit of the simple ratio), and RMSE = 0.22 (lower limit of the simple ratio). All indicators meet the accuracy thresholds, and the model does not require hyperparameter adjustment.

[0076] If the validation accuracy does not reach the threshold, the model hyperparameters need to be adjusted and the model retrained. The adjustment logic is as follows: If R² is low and RMSE is high, it indicates that the model is underfitting. In this case, the number of decision trees can be increased, the maximum decision tree depth can be increased, and the minimum number of samples per leaf node can be decreased. If the training sample set accuracy is high but the validation sample set accuracy is low, it indicates that the model is overfitting. In this case, the number of decision trees can be reduced, the maximum decision tree depth can be decreased, the minimum number of samples per leaf node can be increased, and the feature sampling ratio can be decreased. In the example, if the initial validation yields R²=0.82 and RMSE=0.12gC / MJ, which does not reach the threshold, the number of decision trees will be adjusted to 120, the maximum decision tree depth will be adjusted to 10, and the model will be retrained until the validation accuracy meets the threshold, ultimately generating the optimized parameter prediction model.

[0077] The standardized input dataset is input into the optimized parameter prediction model, which predicts the maximum light energy utilization rate and other static parameters on a pixel-by-pixel or vegetation type-by-vegetation-type basis, generating the vegetation type static parameter configuration results.

[0078] The core of this step is to use the optimized model to predict the static parameters of all pixels or vegetation types within the study area in batches, generating complete static parameter configuration results for vegetation types. This provides accurate parameter support for subsequent CASA model input. The specific implementation method is as follows: The prediction process employs a pixel-by-pixel prediction method, adapting to the raster format of the standardized input dataset to ensure that each pixel receives a corresponding static parameter prediction value. It also considers the consistency of vegetation types, maintaining the predicted parameters for pixels of the same vegetation type within a reasonable range to avoid parameter anomalies. During prediction, the sample feature variable matrix of the standardized input dataset (containing feature variables from all pixels) is input into the optimized parameter prediction model. The model outputs the predicted values ​​of three core static parameters pixel-by-pixel: maximum light energy utilization (ε_max), upper limit parameter of simple ratio (SR_max), and lower limit parameter of simple ratio (SR_min). The prediction accuracy is consistent with that of the validation dataset, with R² ≥ 0.85 and RMSE meeting the threshold requirements.

[0079] During the prediction process, the reasonableness of the predicted parameters for each pixel is verified. Based on the reasonable range of static parameters for different vegetation types, outlier predicted values ​​are eliminated. In the example, the reasonable range for ε_max of arbor forest is 1.70-1.90 gC / MJ. If the predicted ε_max value of a certain arbor forest pixel is 1.55 gC / MJ, it is determined to be an outlier. The average predicted value of surrounding pixels of the same vegetation type is used for correction to ensure the reasonableness of the parameters. Simultaneously, the predicted parameters are statistically analyzed by vegetation type, and the average, maximum, and minimum static parameters for each type are calculated for subsequent parameter reasonableness analysis and model optimization reference.

[0080] After prediction, the predicted static parameters of all pixels are integrated to generate static parameter configuration results for vegetation types. These results are in raster format with a spatial resolution consistent with the standardized input dataset (30m × 30m). Each raster pixel corresponds to three static parameter values. A parameter configuration statistics table is also generated, containing statistical information on the static parameters for various vegetation types. In the example, the predicted average ε_max for arbor forest is 1.82 gC / MJ, the average SR_max is 6.3, and the average SR_min is 1.4; the predicted average ε_max for grassland is 1.52 gC / MJ, the average SR_max is 5.6, and the average SR_min is 1.2. The configuration results clearly present the distribution of static parameters for different vegetation types and spatial locations within the study area and can be directly input into the CASA model for net primary productivity simulation.

[0081] S205, input the static parameter configuration results of the vegetation type into the CASA model, and calculate the net primary productivity at the pixel scale or regional scale by combining the temperature stress factor and the water stress factor to obtain the net primary productivity simulation results.

[0082] Specifically, total solar radiation data and the proportion of photosynthetically active radiation absorbed by vegetation can be extracted from the standardized input dataset, the photosynthetically active radiation absorption component can be calculated, and photosynthetically active radiation raster data can be generated. The core of this step is to obtain the essential energy input parameters required for the CASA model to calculate net primary productivity, namely the photosynthetically active radiation absorption component. This lays the foundation for subsequent calculations of actual light energy utilization and simulations of net primary productivity. The specific implementation method is as follows: The total solar radiation data in the standardized input dataset is preprocessed monthly-scale data derived from meteorological reanalysis data and observations from surface meteorological stations. The unit is MJ / m², with an accuracy controlled within ±5 MJ / m². During extraction, monthly total solar radiation values ​​for all pixels within the study area are screened to ensure accurate correspondence between the radiation data of each pixel and its spatial location and time point, avoiding data loss or misalignment. The vegetation absorbance of photosynthetically active radiation (FPAR) refers to the proportion of photosynthetically active radiation absorbed by the vegetation canopy to the total photosynthetically active radiation. Its value is derived from the vegetation index characteristics in the standardized input dataset, calculated through the fitting relationship between the normalized difference in vegetation index (NDVI) and FPAR. The fitting formula is FPAR = a × NDVI + b, where a and b are fitting coefficients. Experimental verification showed that a is set to 0.85 and b to 0.05 to ensure the accuracy of FPAR calculation.

[0083] The formula for calculating the photosynthetically active radiation (APAR) absorption component is: APAR = PAR_total × FPAR, where APAR is the photosynthetically active radiation absorption component (unit: MJ / m²), PAR_total is the photosynthetically active radiation portion of the total solar radiation (unit: MJ / m²), and FPAR is the proportion of photosynthetically active radiation absorbed by vegetation (unitless). PAR_total is calculated by multiplying the total solar radiation by the proportion of photosynthetically active radiation. Typically, photosynthetically active radiation accounts for 45%-50% of the total solar radiation; in this case, it is set to 48%, i.e., PAR_total = total solar radiation × 0.48.

[0084] In the example, the total solar radiation of a certain arbor forest pixel is 500 MJ / m², and the calculated PAR_total = 500 × 0.48 = 240 MJ / m². The FPAR of this pixel is 0.75, so APAR = 240 × 0.75 = 180 MJ / m². This value directly reflects the actual energy absorbed by the vegetation canopy and is the core energy basis for calculating net primary productivity. After the calculation is completed, the APAR values ​​of all pixels are integrated to generate raster data of absorbed photosynthetically active radiation. The raster spatial resolution is consistent with the standardized input dataset (30m × 30m), and each raster corresponds to an APAR value. At the same time, the corresponding vegetation type is labeled to ensure that the raster data accurately matches the vegetation type and spatial location, providing data support for subsequent calculations of actual light energy utilization.

[0085] Based on the maximum light energy utilization rate in the static parameter configuration results of vegetation type, combined with the low temperature stress factor, high temperature stress factor and water stress factor calculated from meteorological data, the actual light energy utilization rate of the pixel is calculated, and actual light energy utilization rate raster data is generated. The core of this step is to correct the maximum light energy utilization rate through stress factors to obtain the actual available light energy utilization rate of the pixel, reflecting the impact of environmental factors on vegetation photosynthesis and ensuring the accuracy of net primary productivity calculation. The specific implementation method is as follows: The maximum light energy utilization efficiency (ε_max) is derived from the static parameter configuration results of vegetation type. Different vegetation types correspond to different ε_max values. In the example, the ε_max of arbor forest is 2.8 gC / MJ, the ε_max of grassland is 2.2 gC / MJ, and the ε_max of shrub forest is 2.5 gC / MJ. This value is the ideal light energy utilization efficiency of vegetation under no environmental stress. It needs to be corrected in combination with the three major stress factors to obtain the actual light energy utilization efficiency (ε_actual). The correction formula is: ε_actual = ε_max × T1 × T2 × W, where T1 is the low temperature stress factor, T2 is the high temperature stress factor, and W is the water stress factor.

[0086] The low-temperature stress factor (T1) is used to correct the inhibitory effect of low temperature on the activity of photosynthetic enzymes in vegetation. The calculation logic is as follows: when the average temperature is below 0℃, T1=0, at which point the photosynthetic process of vegetation stops; when the temperature is between 0℃ and 10℃, T1=0.05×average temperature, reflecting the gradual increase in photosynthetic enzyme activity under low temperature; when the temperature is between 10℃ and 30℃, T1=1, at which point the temperature is suitable and there is no low-temperature stress; when the temperature is above 30℃, T1 gradually decreases with increasing temperature, T1=1-0.02×(average temperature-30), until the average temperature reaches 40℃ and T1=0.8. In the example, the average temperature of a certain pixel is 18℃, which is between 10℃ and 30℃, so T1=1; if the average temperature is 5℃, then T1=0.05×5=0.25.

[0087] The high-temperature stress factor (T2) is used to correct for the damage caused by high temperatures to the vegetation photosynthetic system. The calculation logic is as follows: when the average temperature is between 10℃ and 30℃, T2 = 1, indicating no high-temperature stress; when the temperature is above 30℃, T2 = 1 - 0.03 × (average temperature - 30℃); when the temperature reaches 40℃, T2 = 0.7; when the temperature is below 10℃, T2 gradually decreases with decreasing temperature, T2 = 0.05 × average temperature + 0.5; when the temperature is 0℃, T2 = 0.5. In the example, the average temperature of a certain pixel is 35℃, so T2 = 1 - 0.03 × (35 - 30) = 0.85.

[0088] The water stress factor (W) is used to correct the impact of water supply on vegetation photosynthesis, reflecting the water sufficiency of vegetation growth. The calculation logic is based on the ratio of precipitation to evapotranspiration, with the formula: W = 1 - 0.01 × (PET - P), where PET is potential evapotranspiration (unit: mm) and P is precipitation (unit: mm). When P ≥ PET, W = 1, indicating no water stress; when P < PET, W decreases with increasing water deficit, reaching a minimum of 0.3, ensuring that the vegetation photosynthetic process does not completely stop. In the example, for a certain pixel with PET of 80 mm and P of 60 mm, W = 1 - 0.01 × (80 - 60) = 0.8, indicating mild water stress, which will inhibit light energy utilization to some extent.

[0089] Combining the three stress factors mentioned above, the actual light energy utilization rate of each pixel is calculated. In the example, for the arbor forest pixel, ε_max = 2.8 gC / MJ, T1 = 1, T2 = 0.85, W = 0.8, then ε_actual = 2.8 × 1 × 0.85 × 0.8 = 1.87 gC / MJ. After the calculation, the actual light energy utilization rates of all pixels are integrated to generate actual light energy utilization rate raster data. The raster resolution is consistent with the absorbed photosynthetically active radiation raster data. Each raster corresponds to an actual light energy utilization rate value, and the vegetation type is labeled to ensure that the data can be directly used for subsequent net primary productivity calculations.

[0090] The net primary productivity of each pixel is obtained by performing a pixel-by-pixel product operation on the absorbed photosynthetically active radiation grid data and the actual light energy utilization grid data, thus generating preliminary net primary productivity simulation results. This step is the core computational step in the net primary productivity simulation. The core logic is to obtain the net primary productivity (NPP) of each pixel by multiplying the energy input (APAR) and the utilization efficiency (ε_actual), thus realizing the transformation from energy input to organic matter accumulation. The specific implementation method is as follows: The formula for calculating net primary productivity is: NPP = APAR × ε_actual × 0.45, where 0.45 is the carbon conversion factor, used to convert the energy unit of photosynthetically active radiation (MJ / m²) into the organic carbon mass unit (gC / m²). This factor comes from the proportion of carbon in the photosynthetic products of vegetation. After experimental verification, it can accurately realize the conversion between energy and carbon mass, ensuring the accuracy of NPP calculation results.

[0091] Pixel-by-pixel product operations must strictly adhere to the principle of spatial correspondence to ensure that the APAR value and ε_actual value of the same pixel are precisely matched, avoiding calculation errors caused by spatial misalignment. During the operation, the absorbed photosynthetically active radiation (APAR) raster data and the actual light energy utilization rate (EPR) raster data are first spatially aligned to ensure a one-to-one correspondence between the pixel positions of the two rasteres. Then, the two values ​​for each pixel are multiplied, and combined with the carbon conversion coefficient, the NPP value for that pixel is obtained, in gC / m²·month (given time scale of month).

[0092] In the example, a grassland pixel has an APAR of 150 MJ / m² and an ε_actual of 1.6 gC / MJ. Therefore, the NPP of this pixel is 150 × 1.6 × 0.45 = 108 gC / m²·month. This value represents the amount of organic carbon accumulated by this pixel through photosynthesis in one month, directly reflecting the vegetation's productivity. After the calculation is completed, the NPP values ​​of all pixels are integrated to generate preliminary net primary productivity simulation results. These results are in raster format and include the NPP value, vegetation type identifier, spatial coordinates, and other information for each pixel. The APAR, ε_actual, and carbon conversion coefficient used in the calculation are also labeled to ensure traceability of the results.

[0093] After the initial simulation results are generated, a simple rationality check needs to be performed to remove outliers. For example, when APAR is 0 or ε_actual is 0, NPP should be recorded as 0. If abnormally high values ​​(such as exceeding 300 gC / m²·month) or negative values ​​appear, the corresponding APAR and ε_actual data need to be checked, corrected, and recalculated to ensure the reliability of the initial net primary productivity simulation results and lay the foundation for subsequent scale aggregation and result output.

[0094] The preliminary net primary productivity simulation results are spatially distributed and scaled, and the spatial distribution maps and statistical summary data of net primary productivity at the pixel scale or region scale are output as the net primary productivity simulation results.

[0095] The core of this step is to optimize and integrate the preliminary simulation results to form standardized simulation outcomes that meet the application needs at different scales. The specific implementation method is as follows: Spatial layout involves covering the entire study area with the preliminary net primary productivity (NPP) simulation results, ensuring that each pixel has a corresponding NPP value. Simultaneously, it combines vegetation type distribution, presenting NPP values ​​for different vegetation types in separate zones. For example, NPP values ​​for different vegetation types such as arbor forests, shrub forests, and grasslands are labeled with different colors, generating a spatial distribution map of NPP. This map clearly shows the spatial distribution differences of NPP within the study area, intuitively reflecting the vegetation productivity of different regions. For instance, NPP values ​​in mountainous areas are generally higher than in plains, and NPP values ​​in wetland areas are higher than in arid areas.

[0096] Scale aggregation processing involves combining pixel-scale NPP data into data at different regional scales based on application requirements. This mainly includes two aspects: First, pixel-scale aggregation, which combines 30m×30m pixel NPP data into administrative regional scales such as townships and counties, calculating the average, maximum, minimum, and total NPP for each region for regional vegetation productivity assessment. Second, time-scale aggregation, which combines monthly NPP data into annual data, calculating the annual total and average NPP within the study area for analyzing interannual variation patterns of vegetation productivity.

[0097] The core contents of the statistical summary data include: total NPP (unit: gC), average NPP (unit: gC / m²·year) in the study area, average and total NPP of different vegetation types, and NPP distribution in different ecological zones. In the example, the total annual NPP of the study area is 125,000 gC / km², and the average NPP is 180 gC / m²·year. Among them, the average NPP of arbor forest is 220 gC / m²·year, the average NPP of grassland is 150 gC / m²·year, and the average NPP of wetland is 200 gC / m²·year.

[0098] The output consists of two parts: first, a spatial distribution map of net primary productivity (NPP), presented as a raster map, with spatial coordinates, vegetation type boundaries, and NPP value ranges marked to ensure the map is clear, standardized, and directly usable for spatial analysis; second, statistical summary data, presented in tabular form, including region, vegetation type, NPP-related statistical indicators, and annotations of calculation methods and data sources to ensure the traceability and usability of statistical data.

[0099] The final net primary productivity simulation results contain both precise data at the pixel scale and statistical data at the regional scale, which can meet the application requirements of the CASA model and provide a scientific basis for vegetation ecological assessment, carbon storage estimation and ecological environment governance in the study area.

[0100] Another embodiment of the present invention provides a CASA model net primary productivity simulation system, see [link to relevant documentation]. Figure 3The system may include: The acquisition module 301 is used to acquire basic data of the study area, which includes at least remote sensing data, vegetation type data, meteorological data, and sample plot observation data; Processing module 302 is used to preprocess the basic data to obtain a standardized input dataset with unified spatial resolution, unified coordinate system and unified time scale; Construction module 303 is used to construct a vegetation type static parameter sample library based on the standardized input dataset. The vegetation type static parameter sample library includes at least a vegetation type identifier, sample feature variables corresponding to the vegetation type, and target static parameter values. The prediction module 304 is used to establish a mapping relationship between the sample feature variables and the target static parameter values ​​using an artificial intelligence model, and to automatically predict the static parameters corresponding to different vegetation types in the study area based on the trained artificial intelligence model, so as to obtain the vegetation type static parameter configuration results. The simulation module 305 is used to input the static parameter configuration results of the vegetation type into the CASA model, and calculate the net primary productivity at the pixel scale or the regional scale by combining the temperature stress factor and the water stress factor, so as to obtain the net primary productivity simulation results.

[0101] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.

[0102] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0103] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.

Claims

1. A method for simulating net primary productivity using a CASA model, characterized in that, The method includes: Acquire basic data for the study area, including at least remote sensing data, vegetation type data, meteorological data, and sample plot observation data; The basic data is preprocessed to obtain a standardized input dataset with uniform spatial resolution, uniform coordinate system and uniform time scale; A vegetation type static parameter sample library is constructed based on the standardized input dataset. The vegetation type static parameter sample library includes at least a vegetation type identifier, sample feature variables corresponding to the vegetation type, and target static parameter values. An artificial intelligence model is used to establish a mapping relationship between the sample feature variables and the target static parameter values. Based on the trained artificial intelligence model, the static parameters corresponding to different vegetation types in the study area are automatically predicted to obtain the vegetation type static parameter configuration results. The static parameter configuration results of the vegetation type are input into the CASA model, and the net primary productivity at the pixel scale or regional scale is calculated by combining the temperature stress factor and the water stress factor to obtain the net primary productivity simulation results.

2. The method according to claim 1, characterized in that, The acquisition of basic data for the study area includes at least remote sensing data, vegetation type data, meteorological data, and sample plot observation data, including: Remote sensing image data covering the study area is acquired through a remote sensing data service platform, including normalized vegetation index, enhanced vegetation index, leaf area index, surface reflectance and surface temperature data, to generate a raw remote sensing image dataset. Extract spatial distribution data of vegetation types in the study area from vegetation classification thematic maps or land use cover products to generate the original vegetation type dataset; Monthly solar radiation, average temperature, maximum temperature, minimum temperature, precipitation, and evapotranspiration data for the study area were obtained through meteorological reanalysis data and ground meteorological stations, generating a raw meteorological element dataset. Data from plot surveys and flux tower observations within the study area were collected, including plot biomass, plot productivity observations, leaf photosynthetic parameters, and flux tower carbon flux observations, to generate the original plot observation dataset.

3. The method according to claim 2, characterized in that, The preprocessing of the basic data to obtain a standardized input dataset with unified spatial resolution, unified coordinate system, and unified time scale includes: The original remote sensing image dataset, original vegetation type dataset, original meteorological element dataset, and original sample plot observation dataset are processed by coordinate system one. All data are transformed to the same coordinate reference system using Albers equal area projection or universal transverse Mercator projection to generate a multi-source dataset with unified coordinates. Resampling is performed on the spatial data in a multi-source dataset with unified coordinates. The other data are then unified to the same spatial resolution based on the highest resolution data, generating a raster dataset with unified resolution. Time scale matching is performed on multi-temporal data in a raster dataset with uniform resolution, daily data are aggregated into monthly scale data, and meteorological interpolation data with different time frequencies are unified to the same time node to generate a time-matched multi-source dataset. Outlier removal, missing value imputation, continuous variable normalization, and vegetation type category variable encoding are performed on time-matched multi-source datasets to generate a standardized input dataset.

4. The method according to claim 3, characterized in that, The vegetation type static parameter sample library is constructed based on the standardized input dataset. This sample library includes at least a vegetation type identifier, sample feature variables corresponding to the vegetation type, and target static parameter values, including: Vegetation type identifiers are extracted from the standardized input dataset as the basis for sample classification. Vegetation index features, surface reflectance features, surface temperature features and leaf area index features are extracted from remote sensing data. Temperature features, precipitation features and solar radiation features are extracted from meteorological data. Elevation, slope and soil type features are extracted from environmental factors to generate a sample feature variable matrix. The maximum light energy utilization rate was obtained by inverting the leaf photosynthesis measurement data in the sample plot observation dataset. The upper limit parameter and lower limit parameter of the simple ratio were obtained by fitting the flux tower observation data with the CASA model, and a set of target static parameter values ​​was generated. The sample feature variable matrix is ​​matched and aligned with the target static parameter value set according to spatial location and time node. Each sample unit contains vegetation type identifier, feature variables and target static parameter values ​​to generate a sample unit set. The sample unit set is quality-screened, abnormal samples are removed and representative samples are added. The sample units are organized and stored according to vegetation type and ecological zone to generate a static parameter sample library of vegetation type.

5. The method according to claim 4, characterized in that, The process involves establishing a mapping relationship between the sample feature variables and the target static parameter values ​​using an artificial intelligence model, and automatically predicting the static parameters corresponding to different vegetation types within the study area based on the trained artificial intelligence model, to obtain the vegetation type static parameter configuration results, including: The static parameter sample library of vegetation type is divided into training sample set and validation sample set. The cross-validation method is used to determine the sample division ratio and generate model training dataset and validation dataset. Construct an artificial intelligence model and train it using a training sample set. Use a random forest model, gradient boosting tree model, or extreme gradient boosting model to establish a nonlinear mapping relationship between sample feature variables and target static parameter values, and generate a trained parameter prediction model. The validation dataset is input into the trained parameter prediction model to verify its accuracy. The coefficient of determination and root mean square error between the predicted and true values ​​are calculated. If the validation accuracy does not reach the threshold, the model hyperparameters are adjusted and the model is retrained to generate an optimized parameter prediction model. The standardized input dataset is input into the optimized parameter prediction model, which predicts the maximum light energy utilization rate and other static parameters on a pixel-by-pixel or vegetation type-by-vegetation-type basis, generating the vegetation type static parameter configuration results.

6. The method according to claim 5, characterized in that, The process involves inputting the static parameter configuration results of the vegetation type into the CASA model, and combining temperature stress factors and water stress factors to calculate net primary productivity at the pixel or regional scale, thereby obtaining the net primary productivity simulation results, including: Extract total solar radiation data and the proportion of photosynthetically active radiation absorbed by vegetation from the standardized input dataset, calculate the photosynthetically active radiation absorption component, and generate photosynthetically active radiation raster data. Based on the maximum light energy utilization rate in the static parameter configuration results of vegetation type, combined with the low temperature stress factor, high temperature stress factor and water stress factor calculated from meteorological data, the actual light energy utilization rate of the pixel is calculated, and actual light energy utilization rate raster data is generated. The net primary productivity of each pixel is obtained by performing a pixel-by-pixel product operation on the absorbed photosynthetically active radiation grid data and the actual light energy utilization grid data, thus generating preliminary net primary productivity simulation results. The preliminary net primary productivity simulation results are spatially distributed and scaled, and the spatial distribution maps and statistical summary data of net primary productivity at the pixel scale or region scale are output as the net primary productivity simulation results.

7. A CASA model net primary productivity simulation system, characterized in that, The system includes: The acquisition module is used to acquire basic data of the study area, which includes at least remote sensing data, vegetation type data, meteorological data, and sample plot observation data. The processing module is used to preprocess the basic data to obtain a standardized input dataset with uniform spatial resolution, uniform coordinate system and uniform time scale; The construction module is used to construct a vegetation type static parameter sample library based on the standardized input dataset. The vegetation type static parameter sample library includes at least a vegetation type identifier, sample feature variables corresponding to the vegetation type, and target static parameter values. The prediction module is used to establish the mapping relationship between the sample feature variables and the target static parameter values ​​using an artificial intelligence model, and to automatically predict the static parameters corresponding to different vegetation types in the study area based on the trained artificial intelligence model, so as to obtain the vegetation type static parameter configuration results. The simulation module is used to input the static parameter configuration results of the vegetation type into the CASA model, and calculate the net primary productivity at the pixel scale or the regional scale by combining the temperature stress factor and the water stress factor, so as to obtain the net primary productivity simulation results.

8. The system according to claim 7, characterized in that, The acquisition module is specifically used for: Remote sensing image data covering the study area is acquired through a remote sensing data service platform, including normalized vegetation index, enhanced vegetation index, leaf area index, surface reflectance and surface temperature data, to generate a raw remote sensing image dataset. Extract spatial distribution data of vegetation types in the study area from vegetation classification thematic maps or land use cover products to generate the original vegetation type dataset; Monthly solar radiation, average temperature, maximum temperature, minimum temperature, precipitation, and evapotranspiration data for the study area were obtained through meteorological reanalysis data and ground meteorological stations, generating a raw meteorological element dataset. Data from plot surveys and flux tower observations within the study area were collected, including plot biomass, plot productivity observations, leaf photosynthetic parameters, and flux tower carbon flux observations, to generate the original plot observation dataset.

9. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1-6 when it is run.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1-6.