Ecological system model parameter optimization system and method based on parallel scheduling

By constructing a plant species library and clustering algorithm to identify PFT, combining multiple sensitivity analysis and optimization algorithms, the Dask parallel computing framework is used to solve the inefficiency of PFT classification and parameter optimization in the ecosystem model, and an efficient and automated ecosystem simulation is achieved.

CN120373543AInactive Publication Date: 2025-07-25SHAANXI NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510454791.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing ecosystem models lack spatial dynamic representation capabilities during PFT classification, the parameter optimization process has algorithmic limitations and low computational efficiency, the generality and scalability of the model are difficult to meet the needs of rapid evaluation, and the utilization of computing resources is inefficient.

Method used

A plant species library was constructed, a clustering algorithm was used to identify species and map PFT, and sensitivity analysis was performed in combination with Sobol’, Morris and EFAST methods, a multi-algorithm collaborative optimization framework was constructed, and a Dask parallel computing framework was used to achieve task parallelization and dynamic resource scheduling, and a unified API interface was designed to adapt to multiple ecological models.

Benefits of technology

It realizes automatic division of PFT types and high-precision identification, significantly reduces manual intervention, improves model prediction accuracy and adaptability, shortens computing time, improves computing efficiency and resource utilization, and reduces system integration and user operation complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373543A_ABST
    Figure CN120373543A_ABST
Patent Text Reader

Abstract

The invention provides an ecological system model parameter optimization system and method based on parallel scheduling, and belongs to the technical field of resources and environments. By constructing a species feature library, adopting a clustering method for species identification, and combining various sensitivity analysis algorithms and an automatic parameter optimization strategy, full-process automatic configuration of remote sensing data preprocessing, feature extraction, species identification and PFT mapping is realized, and the workload of manual intervention and manual parameter adjustment is remarkably reduced; an unsupervised / semi-supervised learning method is introduced, so that high-precision automatic identification of main vegetation species in the region is ensured; according to the method, a multi-method fusion sensitivity analysis of Sobol ', Morris, EFAST and the like is adopted, a multi-algorithm collaborative optimization framework is combined, accurate and automatic adjustment of key physiological parameters is achieved, so that the precision and adaptability of model prediction are greatly improved, a Dask parallel computing framework and an SLURM scheduler are utilized, efficient fragmentation and dynamic resource scheduling of tasks are achieved, and the computing efficiency is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of resources and environment, and particularly relates to an ecosystem model parameter optimization system and method based on parallel scheduling. Background Art

[0002] Modern ecosystem simulation takes mathematical modeling as the core and integrates automated simulation analysis technologies of multi-source data and computational science. Its standardized process usually covers three key links: construction of the basic data layer, architecture of the model operation layer, and simulation optimization and verification.

[0003] At the basic data level, a multi-dimensional data acquisition system plays a key role. Through the ground observation network, relying on ecological stations for species census, soil profile sampling, and meteorological monitoring, high-precision in-situ data can be obtained; at the same time, through field surveys and sampling and measurement of key vegetation species in the area, key plant functional type information of the area can be obtained. In addition, by integrating spatialized data such as climate reanalysis data (e.g., ERA5) and digital elevation model (DEM), a detailed historical climate dataset can be constructed. Vegetation parameterization modeling simplifies complex biological communities into a limited number of categories (e.g., the LPJ-GUESS model divides global vegetation into 12 PFTs) through the classification of plant functional types (PFTs), and establishes a PFT characteristic parameter library including physiological process parameters such as photosynthetic response curves (Farquhar model) and stomatal conductance (Ball-Berry equation).

[0004] In the architecture of the model operation layer, first, a suitable process mechanism model needs to be selected, such as a carbon cycle model (e.g., BIOME-BGC), a dynamic global vegetation model (DGVM), or a land surface process model (e.g., CLM), and grid settings such as 0.5° resolution and spatio-temporal resolution at the monthly scale are set, and the atmospheric forcing field is coupled to drive the model operation. For the uncertainty of parameters in the model, methods such as OTA, Morris, Sobol’ etc. can be used to determine the key parameter subset (e.g., maximum carboxylation rate Vcmax, leaf nitrogen content, etc.), so as to improve the optimization efficiency and accuracy of the model.

[0005] In the simulation optimization and verification stage, GPP, ET observed by flux towers and leaf area index (LAI) retrieved by remote sensing are often used as verification benchmarks, and the simulation accuracy is evaluated through evaluation indexes such as Nash-Sutcliffe efficiency coefficient (NSE), root mean square error (RMSE), and coefficient of determination (R 2 ) etc. At the same time, applying a certain optimization algorithm (such as genetic algorithm, simulated annealing algorithm, or Bayesian optimization algorithm, etc.), constructing a cost function to minimize the deviation between the simulated value and the observed value, and continuously bringing the key parameter subset determined by sensitivity analysis into the model to achieve continuous optimization of the model.

[0006] Ecosystem simulation has a wide range of scenarios in practical applications. In the assessment of climate change responses, models can be used to predict the impact of increased CO2 concentration on net primary productivity (NPP) of vegetation and to quantify the weakening of carbon sink functions by extreme drought events, such as the analysis of tipping points in the Amazon rainforest; in the quantification of the benefits of ecological engineering, models can simulate the impact of the Three-North Shelter Forest Project on surface albedo and its effect on regional climate, and at the same time evaluate the potential of mangrove restoration projects in carbon sequestration and storm surge reduction; in the field of resource management decision support, by coupling crop models with ecosystem models, irrigation and fertilization regimes can be optimized, and the dynamic relationship between wildfire spread models and vegetation flammability parameters can also be studied.

[0007] Although the existing technical framework has established a complete ecosystem simulation chain, it still faces three core bottlenecks in practical applications: First, the spatio-temporal dynamic representation ability of PFT classification is insufficient; second, there are algorithm limitations and low computational efficiency in the parameter optimization process; third, the model generality and scalability are difficult to meet the rapid assessment requirements. The specific manifestations are as follows: 1. There is uncertainty in the selection of plant functional types (PFTs) in a specific region Plant Functional Types (PFTs) is a method of classifying different plants into several categories based on their physiological and ecological characteristics, morphological structures, and ecological functions. PFT classification not only reflects the similarities of plants in growth, competition, resource utilization, etc., but also reveals their response mechanisms to the cycling of substances such as carbon, energy, and water in the ecosystem. In current ecosystem models, vegetation is usually parameterized in the form of PFTs, and by calibrating the physiological and ecological parameters (such as photosynthetic rate, respiration rate, etc.) of each PFT type, the simulation of processes such as carbon cycling, energy balance, and climate change response can be achieved. Therefore, accurate PFT classification and parameter calibration play a crucial role in improving the accuracy and reliability of ecological models.

[0008] Currently, PFT data mainly rely on ground surveys or are obtained from existing references, but due to high labor costs and limited sample coverage, it is difficult to meet the needs of large-scale and dynamic simulations. Remote sensing technology can provide continuous spatial observation data, providing an important data source for large-area ecosystem monitoring. Currently, remote sensing data are mainly reflected in the following aspects: agricultural precision management, forestry and carbon sink monitoring, environmental and ecological protection, disaster emergency response, climate change research, urban planning and management, water resource management, etc. There is still a lack of research on using remote sensing data to obtain PFT information in a certain region. How to make full use of remote sensing data to achieve automatic classification of PFT types has become an important technical challenge and research hotspot in current ecological model research and parameter calibration.

[0009] The calibration of PFT parameters in existing ecosystem models highly relies on manual experience. The traditional method requires first conducting on-site investigations of the main vegetation species information in the research area, then determining the main PFT types in this area based on the species information, and then sampling the key vegetation species and bringing them back to the laboratory to measure the key parameters of PFT. After that, the parameters measured in the laboratory need to be manually edited to configure files, binding each PFT category with the measured parameters one by one. The entire process has a very large time span and requires manual intervention throughout. It is almost a very difficult thing to complete a quick and accurate simulation of a certain area.

[0010] 2. Limitations of Sensitivity Analysis and Optimization Methods After obtaining the types of PFT, it is necessary to determine the parameters of each PFT. However, due to the rich types of PFT and the large number of parameters for each PFT, if parameter optimization is directly carried out, the time and resources consumed will be unimaginable. For ecosystem models, the simulation results are generally sensitive to a limited number of parameter counts. Therefore, only the parameters sensitive to the output of interest need to be found, and then only these sensitive parameters need to be optimized.

[0011] However, in current research, single methods (such as the Sobol’ algorithm, efast method, or Morris method) are mostly used for parameter sensitivity analysis to screen key parameters. However, different methods may lead to conflicting results due to differences in mathematical principles (such as global / local sensitivity): for example, the Sobol’ method has extremely high computational costs for high-dimensional parameter spaces, while the Morris method cannot quantify parameter interaction effects; existing technologies have not established a multi-method cross-validation mechanism, resulting in an unstable sensitive parameter set.

[0012] The parameter optimization link also faces challenges: many ecosystem models only provide the source code of the model, and the relevant technologies for sensitivity analysis and its optimization are designed by users. However, designing an automated parameter tuning scheme is a very time-consuming operation, and this process requires mastering a lot of programming techniques, which greatly hinders the research progress of ecosystem simulation. Moreover, the optimizations between different models are separated, and there is no general framework to provide a general API interface to complete the call and optimization of different models.

[0013] 3. Insufficient Computational Efficiency and Parallelization The single run of an ecosystem model can take several hours to several days (depending on the spatial resolution and time span). Sensitivity analysis and parameter optimization require multiple iterative calls to the model. The traditional serial computing mode has extremely low efficiency, with the CPU utilization rate of a single task less than 30%, and it cannot dynamically allocate computing nodes. Even many researchers manually adjust parameters and run the model again and again, which undoubtedly consumes a lot of time and causes great uncertainty in the results.

[0014] Based on the above content analysis, the following core defects exist in the prior art: 1) The determination of PFT in a specific area requires a large time span and complicated manual operations; 2) The degree of automation of parameter sensitivity analysis and calibration is low. The optimization of many studies relies on manually running the model repeatedly for analysis, which has great uncertainty; 4) The sensitivity analysis and optimization methods are single. Due to the lack of a general call interface, only single methods are used for sensitivity analysis and optimization, resulting in low reliability of parameter importance ranking; 5) There is no general framework to complete the call and optimization of different models, which causes a great obstacle to the comparative research and analysis between models; 6) The utilization of computing resources is inefficient. The model operation, parameter sensitivity analysis and optimization are executed serially, and the multi-core CPU parallel acceleration is not utilized. Summary of the Invention

[0015] The purpose of the present invention is to overcome the above deficiencies and provide an ecosystem model parameter optimization system and method based on parallel scheduling.

[0016] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, the present invention provides an ecosystem model parameter optimization system based on parallel scheduling, including: A plant species library construction module for constructing a plant species library and establishing a mapping relationship between species and PFT categories; A PFT automatic mapping module for selecting remote sensing image data of a target area from the constructed plant species library, and based on the remote sensing image data of the target area, combining a clustering algorithm to realize the identification of species types in the target area and the automatic mapping of PFT to obtain a PFT parameter set; A sensitivity analysis module for calculating the global sensitivity index of the PFT parameter set by using the Sobol’ method, calculating the local sensitivity index of the PFT parameter set by using the Morris method, and evaluating the interaction relationship between parameters in the PFT parameter set by using the EFAST method, and screening out the common sensitive parameters; A multi-algorithm optimization module for constructing a multi-algorithm collaborative optimization framework, inputting the common sensitive parameters into the constructed multi-algorithm collaborative optimization framework, and calling a genetic algorithm, Bayesian optimization or particle swarm algorithm to perform iterative update of the sensitive parameters to obtain the updated common sensitive parameters; A feedback calibration module for feeding back the updated common sensitive parameters to the PFT parameter set; A parallel computing framework for realizing the parallelization of the automatic mapping, sensitivity analysis and optimization tasks of the PFT parameter set based on the Dask framework, and parallelly distributing the tasks to cluster nodes.

[0017] The Dask framework integrates the SLURM scheduler through Dask-Jobqueue to automatically allocate CP resources according to the task load.

[0018] The Dask framework designs a general API interface, and the API interface adapts to multiple ecological models.

[0019] In a second aspect, the present invention provides an ecological system model parameter optimization method based on parallel scheduling, including the following steps: Construct a plant species library and establish a mapping relationship between species and PFT categories; Select remote sensing image data of the target area from the constructed plant species library. Based on the remote sensing image data of the target area and combined with a clustering algorithm, realize the species type recognition of the target area and the automatic mapping of PFT, and obtain a PFT parameter set; Use the Sobol’ method to calculate the global sensitivity index of the PFT parameter set, use the Morris method to calculate the local sensitivity index of the PFT parameter set, use the EFAST method to evaluate the interaction relationship between parameters in the PFT parameter set, and screen out the common sensitive parameters; Construct a multi-algorithm collaborative optimization framework, input the common sensitive parameters into the constructed multi-algorithm collaborative optimization framework, and call a genetic algorithm, Bayesian optimization or particle swarm algorithm to perform iterative update of the sensitive parameters to obtain the updated common sensitive parameters; Feed the updated common sensitive parameters back to the PFT parameter set; Based on the Dask framework, realize the automatic mapping of the PFT parameter set, sensitivity analysis and parallelization of the optimization task, and parallelly allocate the tasks to the cluster nodes.

[0020] In the step of selecting remote sensing image data of the target area from the constructed plant species library, based on the remote sensing image data of the target area and combined with a clustering algorithm, realizing the species type recognition of the target area and the automatic mapping of PFT, and obtaining a PFT parameter set, the specific method is as follows: Select remote sensing image data of the target area from the constructed plant species library, preprocess the obtained remote sensing image data, and extract features from the preprocessed remote sensing image data; Cluster the extracted features of the remote sensing image data to generate multiple candidate species clusters, match the candidate species clusters with the species in the plant species library, and identify the species composition of the species clusters; Based on the established mapping relationship between species and PFT categories, classify the species composition of the identified species clusters into the corresponding PFT categories and output the PFT parameter set.

[0021] In the step of clustering the features of the extracted remote sensing image data to generate multiple candidate species clusters and matching the candidate species clusters with the species in the plant species library to identify the species composition of the species clusters, the specific method is as follows: Use an unsupervised or semi-supervised clustering algorithm to preliminarily classify the features of the extracted remote sensing image data, and determine the appropriate number of clusters through the elbow method or silhouette coefficient, so as to form multiple candidate species clusters; Match the candidate species clusters with the species records in the plant species library, and use the Euclidean distance index to identify the species composition of the species clusters.

[0022] In the step of calculating the global sensitivity index of the PFT parameter set by the Sobol’ method, calculating the local sensitivity index of the PFT parameter set by the Morris method, evaluating the interaction relationship between the parameters in the PFT parameter set by the EFAST method, and screening out the common sensitive parameters, the specific method is as follows: Use the Sobol’ method to calculate the global sensitivity index of the PFT parameter set and obtain the total sensitivity index of each parameter to the output result; Use the Morris method to calculate the local sensitivity index of the PFT parameter set to obtain the mean value μ*; Use the EFAST method to evaluate the interaction relationship between the parameters in the PFT parameter set to obtain the interaction effect index; Screen out the common sensitive parameters according to the preset standard as the parameter optimization target parameter set.

[0023] The common sensitive parameters are the parameters that simultaneously satisfy the total sensitivity index > 0.1, the mean value μ* ranks in the top 30%, and the proportion of the interaction effect index > 15%.

[0024] In a third aspect, the present invention provides an electronic device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the ecosystem model parameter optimization method based on parallel scheduling are implemented.

[0025] In a fourth aspect, the present invention provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the ecosystem model parameter optimization method based on parallel scheduling are implemented.

[0026] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides an optimized scheduling method for ecosystem model parameters based on parallel scheduling, including the following steps: constructing a plant species library and establishing a mapping relationship between species and PFT categories; selecting remote sensing image data of the target area from the constructed plant species library, and based on the remote sensing image data of the target area, combining a clustering algorithm to achieve species type identification and PFT automatic mapping of the target area, obtaining a PFT parameter set; using the Sobol’ method to calculate the global sensitivity index of the PFT parameter set, using the Morris method to calculate the local sensitivity index of the PFT parameter set, using the EFAST method to evaluate the interaction relationship between parameters in the PFT parameter set, and screening out common sensitive parameters according to a preset standard; constructing a multi-algorithm collaborative optimization framework, inputting the common sensitive parameters into the constructed multi-algorithm collaborative optimization framework, and calling a genetic algorithm, Bayesian optimization or particle swarm algorithm to perform iterative update of the sensitive parameters, obtaining the updated common sensitive parameters; feeding back the updated common sensitive parameters to the PFT parameter set; realizing the parallelization of the sensitivity analysis of the PFT parameter set and the optimization task based on the Dask framework, and parallelly allocating the tasks to the cluster nodes. By constructing a species feature library, using a clustering method for species identification, combining multiple sensitivity analysis algorithms and an automated parameter optimization strategy, the full-process automatic configuration of remote sensing data preprocessing, feature extraction, species identification and PFT mapping is realized, significantly reducing the workload of manual intervention and manual parameter adjustment; by introducing a clustering method, high-precision automatic identification of the main vegetation species in the area is ensured; the sensitivity analysis integrating multiple methods such as Sobol’, Morris and EFAST is adopted, combined with a multi-algorithm collaborative optimization framework such as a genetic algorithm, Bayesian optimization, particle swarm algorithm, etc., to realize the precise automatic adjustment of key physiological parameters, thereby greatly improving the accuracy and adaptability of model prediction; the Dask framework is used to realize the efficient sharding and dynamic resource scheduling of tasks, so that the running time of complex models is shortened several times compared with traditional serial computing, significantly improving the computing efficiency.

[0027] Furthermore, a unified API interface is designed to solve the problem of inconsistent interfaces of different ecosystem models, reducing the complexity of system integration and user operation.

[0028] Furthermore, based on a high-performance parallel computing platform and dynamic resource scheduling technology, the full-process automation and efficient operation are ensured, providing an intelligent, flexible and end-to-end comprehensive solution for ecosystem simulation and parameter calibration. Description of the Drawings

[0029] Figure 1 It is the system diagram of the present invention; Figure 2 It is the method flow diagram of the present invention; Figure 3 It is the system diagram of Embodiment 3 of the present invention. Detailed Implementation Modes

[0030] To further understand the content of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.

[0031] Embodiment 1 As Figure 1 shown, an ecosystem model parameter optimization system based on parallel scheduling, characterized by comprising: A plant species library construction module, used to construct a plant species library and establish a mapping relationship between species and PFT categories; A PFT automatic mapping module, used to select remote sensing image data of a target area from the constructed plant species library, and based on the remote sensing image data of the target area, combine a clustering algorithm to realize species type recognition and PFT automatic mapping of the target area, and obtain a PFT parameter set; A sensitivity analysis module, used to calculate the global sensitivity index of the PFT parameter set by using the Sobol’ method, calculate the local sensitivity index of the PFT parameter set by using the Morris method, evaluate the interaction relationship between parameters in the PFT parameter set by using the EFAST method, and screen out the common sensitive parameters according to a preset standard; A multi-algorithm optimization module, used to construct a multi-algorithm collaborative optimization framework, input the common sensitive parameters into the constructed multi-algorithm collaborative optimization framework, and call a genetic algorithm, Bayesian optimization or particle swarm algorithm to perform iterative update of the sensitive parameters, and obtain the updated common sensitive parameters; A feedback calibration module, used to feedback the updated common sensitive parameters to the PFT parameter set; A parallel computing framework, used to realize the parallelization of the automatic mapping, sensitivity analysis and optimization tasks of the PFT parameter set based on the Dask framework, and distribute the tasks in parallel to the cluster nodes.

[0032] Further, the Dask framework integrates the SLURM scheduler through Dask-Jobqueue to automatically allocate CP resources according to the task load.

[0033] Further, the Dask framework designs a general API interface, and the API interface adapts to a variety of ecological models.

[0034] Embodiment 2 As Figure 2 shown, an ecosystem model parameter optimization method based on parallel scheduling includes the following steps: S1: Construct a plant species library and establish a mapping relationship between species and PFT categories; S2: Select the remote sensing image data of the target area from the constructed plant species library. Based on the remote sensing image data of the target area, combine the clustering algorithm to realize the species type recognition and PFT automatic mapping of the target area, and obtain the PFT parameter set; S3: Use the Sobol’ method to calculate the global sensitivity index of the PFT parameter set, use the Morris method to calculate the local sensitivity index of the PFT parameter set, use the EFAST method to evaluate the interaction relationship between the parameters in the PFT parameter set, and screen out the common sensitive parameters according to the preset standard; S4: Construct a multi-algorithm collaborative optimization framework, input the common sensitive parameters into the constructed multi-algorithm collaborative optimization framework, and call the genetic algorithm, Bayesian optimization or particle swarm algorithm to perform iterative update of the sensitive parameters to obtain the updated common sensitive parameters; S5: Feed back the updated common sensitive parameters to the PFT parameter set; S6: Based on the Dask framework, realize the parallelization of the automatic mapping, sensitivity analysis and optimization tasks of the PFT parameter set, and allocate the tasks in parallel to the cluster nodes.

[0035] Specifically, in S1, build a plant species library and establish the mapping relationship between species and PFT categories. The specific method is as follows: Integrate the data in the existing literature, flora and public databases to construct a comprehensive plant species library. In this stage, it is necessary to not only include the common tree species (such as poplar, willow, cypress, etc.) and other vegetation species in the region, but also record in detail the key characteristics of each species, such as spectrum, texture, temporal variation and growth habits. At the same time, according to the criteria such as growth form, leaf type, shade tolerance, leaf type and climate distribution, clarify the PFT category corresponding to each species, and ensure the accuracy and reliability of the data through expert review and cross-validation.

[0036] Specifically, in S2, select the remote sensing image data of the target area from the constructed plant species library. Based on the remote sensing image data of the target area, combine the clustering algorithm to realize the species type recognition and PFT automatic mapping of the target area, and obtain the PFT parameter set. The specific method is as follows: S21: Select the remote sensing image data of the target area from the constructed plant species library, preprocess the obtained remote sensing image data, and extract the features of the preprocessed remote sensing image data; Select high-resolution and hyperspectral remote sensing image data (such as MODIS, Sentinel series data) of the target area from the constructed plant species library and preprocess it, and then extract features from the preprocessed remote sensing image data. The preprocessing steps include geometric correction, radiometric correction, and spatio-temporal registration, aiming to ensure the accuracy and consistency of the image data. Subsequently, by extracting vegetation indices such as NDVI and EVI, as well as spectral features reflecting subtle species differences in hyperspectral data, combined with texture and shape information, and temporal information capturing the dynamic changes of the vegetation growth season, a solid foundation is laid for subsequent species identification.

[0037] S22: Cluster the features of the extracted remote sensing image data to generate multiple candidate species clusters, match the candidate species clusters with the species in the plant species library, and identify the species composition of the species clusters; Use unsupervised or semi-supervised clustering algorithms (such as K-means or DBSCAN) to preliminarily classify the features of the extracted remote sensing image data, determine the appropriate number of clusters through the elbow method or silhouette coefficient, and thus form multiple candidate species clusters. Then, match the typical features of these candidate species clusters with the species records in the previously constructed plant species library, and use similarity metrics such as Euclidean distance to preliminarily identify the specific species that may exist in the region. In this way, even in the absence of a large amount of labeled data, the species composition in the target area can be determined relatively accurately.

[0038] S23: Based on the established mapping relationship between species and PFT categories, classify the species composition of the identified species clusters into the corresponding PFT categories and output the PFT parameter set; After identifying the specific species, according to the previously established mapping relationship between species and PFT, automatically classify each species into the corresponding PFT category. This process does not involve further determination of various physiological parameters, but directly divides the PFT types in the region through species. Verify the identification results through on-site investigation data or historical remote sensing temporal data to ensure the accuracy of species identification and PFT mapping. Continuously optimize the species library, adjust the clustering algorithm and matching rules according to the verification feedback to improve the adaptability and accuracy of the overall automated process.

[0039] Specifically, in S3, use the Sobol’ method to calculate the global sensitivity index of the PFT parameter set, use the Morris method to calculate the local sensitivity index of the PFT parameter set, use the EFAST method to evaluate the interaction relationship between parameters in the PFT parameter set, and screen out the common sensitive parameters according to the preset standard; the specific method is as follows: According to the preliminary model operation results, use the Sobol’ method to calculate the global sensitivity index of the PFT parameter set and obtain the total sensitivity index of each parameter to the output results (such as GPP, ET, LAI, etc.); The local sensitivity index of the PFT parameter set is calculated using the Morris method to obtain the mean value μ*; The EFAST method is used to evaluate the interaction relationship between parameters in the PFT parameter set to obtain the interaction effect index; The common sensitive parameters are screened out according to the preset criteria to provide a stable target parameter set for parameter optimization; The common sensitive parameters are those that simultaneously satisfy the Sobol total sensitivity index > 0.1, the mean value μ* calculated by Morris ranks in the top 30%, and the proportion of the EFAST interaction effect index > 15%. These parameters are used as the core objects for subsequent automated parameter optimization.

[0040] Specifically, in S4, a multi-algorithm collaborative optimization framework is constructed. The common sensitive parameters are input into the constructed multi-algorithm collaborative optimization framework, and genetic algorithms, Bayesian optimization, or particle swarm algorithms are called to iteratively update the sensitive parameters to obtain the updated common sensitive parameters. The specific method is as follows: Construct a multi-algorithm collaborative optimization framework that supports tuning methods such as genetic algorithms, Bayesian optimization, and particle swarm algorithms, and allows users to customize the cost function (such as RMSE, R2, etc.) according to specific application scenarios. Input the common sensitive parameters into the constructed multi-algorithm collaborative optimization framework, call the algorithms therein, and use automated scripts to complete multiple iterative runs of the model. Each iteration updates the values of the sensitive parameters according to the principle of minimizing the objective function until convergence to the optimal solution.

[0041] Specifically, in S5, the optimization results are automatically fed back to the PFT parameter library to form a closed-loop parameter calibration mechanism, realizing full-process automated parameter tuning.

[0042] Specifically, in S6, based on the Dask parallel computing framework, the overall task is automatically decomposed into multiple subtasks, including remote sensing data classification processing, calculating sensitivity indicators, and performing parameter optimization. Through the task graph management function of Dask, these subtasks are parallelly allocated to each node of the cluster for execution, significantly improving the computing efficiency. To achieve dynamic resource scheduling, the framework integrates the SLURM scheduler and uses the Dask-Jobqueue module to automatically allocate CP resources according to the task load to ensure the efficient use of cluster resources.

[0043] Furthermore, the Dask parallel computing framework designs a general API interface. By defining the base class EcosystemModel, users can implement methods for model-specific parameter settings (set_params), model running (run), and output parsing (get_outputs). For example, for the LPJ-GUESS model, users can set parameters by updating the configuration file (such as param.ins), then call the executable file to run the simulation, and the framework automatically parses the metrics (such as GPP, ET, LAI, etc.) in the output file. This API shields the technical differences between models, unifies the calling method, and solves the problem of fragmented optimization processes. Compared with the complexity of manually configuring script files in traditional MPI frameworks (such as MPICH), this framework automatically manages dependencies and environments through a configuration file, reduces the technical threshold, enables non-professional users to get started quickly, and promotes the wide application of ecosystem simulation technology.

[0044] Furthermore, the framework realizes the automatic management of configuration files through a model wrapper. During the parameter optimization process, after the optimization algorithm (such as Bayesian optimization) generates parameter combinations, the Dask master process distributes tasks to each worker node. Each node generates a local configuration file according to the parameters, then runs the model executable file, and collects the output results. After the operation is completed, the output data is returned to the master process through a shared file system (such as NFS) or distributed storage (such as Amazon S3) for the next round of sensitivity analysis or optimization iteration. This process requires no manual intervention, ensuring seamless connection between parameter updates and model running. For tasks that take a long time (such as multi-scenario simulations), the parallel design can shorten the computing time from several hours to dozens of minutes, greatly improving efficiency.

[0045] Through the combination of Dask and SLURM, efficient parallelization of tasks and dynamic resource scheduling are achieved; the unified API design adapts to multiple ecosystem models, solves the problem of inconsistent interfaces, and reduces the usage complexity. Compared with existing methods (such as using MPI alone or manually configuring models), this framework significantly improves the computing efficiency and applicability, providing an end-to-end solution for ecosystem simulation.

[0046] Embodiment 3 As Figure 3 shown, the present invention also provides an electronic device 100 for an ecosystem model parameter optimization method based on parallel scheduling; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.

[0047] The memory 101 can be used to store the computer program 103. The processor 102 realizes the steps of the method for optimizing the ecosystem model parameters based on parallel scheduling described in Embodiment 2 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 may include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0048] The at least one processor 102 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or the processor 102 may also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100 and connects various parts of the entire electronic device 100 through various interfaces and lines.

[0049] The memory 101 in the electronic device 100 stores multiple instructions to implement a method for optimizing the ecosystem model parameters based on parallel scheduling. The processor 102 can execute the multiple instructions to achieve: Construct a plant species library and establish a mapping relationship between species and PFT categories; Select remote sensing image data of the target area from the constructed plant species library, and based on the remote sensing image data of the target area, combine a clustering algorithm to realize the species type recognition of the target area and the automatic mapping of PFT, and obtain a PFT parameter set; The global sensitivity indices of the PFT parameter set are calculated using the Sobol’ method, the local sensitivity indices of the PFT parameter set are calculated using the Morris method, and the EFAST method is used to evaluate the interaction relationships among the parameters in the PFT parameter set to screen out the common sensitive parameters; A multi-algorithm collaborative optimization framework is constructed. The common sensitive parameters are input into the constructed multi-algorithm collaborative optimization framework, and the genetic algorithm, Bayesian optimization, or particle swarm algorithm is called to perform iterative updates of the sensitive parameters to obtain the updated common sensitive parameters; The updated common sensitive parameters are fed back to the PFT parameter set; Based on the Dask framework, the sensitivity analysis of the PFT parameter set and the parallelization of the optimization task are realized, and the tasks are parallelly allocated to the cluster nodes.

[0050] Example 4 If the modules / units integrated in the electronic device 100 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, and read-only memory (ROM, Read-Only Memory).

[0051] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0052] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0053] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0054] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. An ecosystem model parameter optimization system based on parallel scheduling, characterized in that, Including: A plant species library construction module, which is used to construct a plant species library and establish a mapping relationship between species and PFT categories; A PFT automatic mapping module, which is used to select remote sensing image data of a target area from the constructed plant species library, and based on the remote sensing image data of the target area, combine a clustering algorithm to realize the identification of species types in the target area and the automatic mapping of PFT, and obtain a PFT parameter set; A sensitivity analysis module, which is used to calculate the global sensitivity index of the PFT parameter set by using the Sobol’ method, calculate the local sensitivity index of the PFT parameter set by using the Morris method, evaluate the interaction relationship between parameters in the PFT parameter set by using the EFAST method, and screen out the common sensitive parameters; A multi-algorithm optimization module, which is used to construct a multi-algorithm collaborative optimization framework, input the common sensitive parameters into the constructed multi-algorithm collaborative optimization framework, and call a genetic algorithm, Bayesian optimization or particle swarm algorithm to perform iterative update of the sensitive parameters, and obtain the updated common sensitive parameters; A feedback calibration module, which is used to feedback the updated common sensitive parameters to the PFT parameter set; A parallel computing framework, which is used to realize the parallelization of the automatic mapping, sensitivity analysis and optimization tasks of the PFT parameter set based on the Dask framework, and parallelly allocate tasks to cluster nodes.

2. The parameter optimization system for an ecosystem model based on parallel scheduling according to claim 1, wherein The Dask framework integrates the SLURM scheduler through Dask-Jobqueue to automatically allocate CP resources according to the task load.

3. The parameter optimization system for an ecosystem model based on parallel scheduling according to claim 1, characterized in that, The Dask framework designs a general API interface, and the API interface adapts to a variety of ecological models.

4. An ecological system model parameter optimization method based on parallel scheduling, characterized in that, Including the following steps: Construct a plant species library and establish a mapping relationship between species and PFT categories; Select remote sensing image data of a target area from the constructed plant species library, and based on the remote sensing image data of the target area, combine a clustering algorithm to realize the identification of species types in the target area and the automatic mapping of PFT, and obtain a PFT parameter set; Calculate the global sensitivity index of the PFT parameter set by using the Sobol’ method, calculate the local sensitivity index of the PFT parameter set by using the Morris method, evaluate the interaction relationship between parameters in the PFT parameter set by using the EFAST method, and screen out the common sensitive parameters; Construct a multi-algorithm collaborative optimization framework, input the common sensitive parameters into the constructed multi-algorithm collaborative optimization framework, and call a genetic algorithm, Bayesian optimization or particle swarm algorithm to perform iterative update of the sensitive parameters, and obtain the updated common sensitive parameters; Feedback the updated common sensitive parameters to the PFT parameter set; Realize the parallelization of the automatic mapping, sensitivity analysis and optimization tasks of the PFT parameter set based on the Dask framework, and parallelly allocate tasks to cluster nodes.

5. The method for optimizing the parameters of an ecosystem model based on parallel scheduling according to claim 4, wherein In the step of selecting remote sensing image data of a target area from the constructed plant species library, and based on the remote sensing image data of the target area, combining a clustering algorithm to realize the identification of species types in the target area and the automatic mapping of PFT, and obtaining a PFT parameter set, the specific method is as follows: Select remote sensing image data of a target area from the constructed plant species library, preprocess the obtained remote sensing image data, and extract features from the preprocessed remote sensing image data; Cluster the features of the extracted remote sensing image data to generate multiple candidate species clusters, match the candidate species clusters with the species in the plant species library, and identify the species composition of the species clusters; Based on the established mapping relationship between species and PFT categories, classify the species composition of the identified species clusters into the corresponding PFT categories and output the PFT parameter set.

6. The method for optimizing the parameters of an ecosystem model based on parallel scheduling according to claim 4, wherein In the step of clustering the features of the extracted remote sensing image data to generate multiple candidate species clusters, matching the candidate species clusters with the species in the plant species library, and identifying the species composition of the species clusters, the specific method is as follows: Use an unsupervised or semi-supervised clustering algorithm to preliminarily classify the features of the extracted remote sensing image data, and determine the appropriate number of clusters through the elbow method or silhouette coefficient to form multiple candidate species clusters; Match the candidate species clusters with the species records in the plant species library, and use the Euclidean distance index to identify the species composition of the species clusters.

7. A method for optimizing the parameters of an ecosystem model based on parallel scheduling according to claim 4, characterized in that In the step of calculating the global sensitivity index of the PFT parameter set using the Sobol’ method, calculating the local sensitivity index of the PFT parameter set using the Morris method, evaluating the interaction relationship between the parameters in the PFT parameter set using the EFAST method, and screening out the common sensitive parameters, the specific method is as follows: Use the Sobol’ method to calculate the global sensitivity index of the PFT parameter set and obtain the total sensitivity index of each parameter to the output result; Use the Morris method to calculate the local sensitivity index of the PFT parameter set to obtain the mean value μ*; Use the EFAST method to evaluate the interaction relationship between the parameters in the PFT parameter set to obtain the interaction effect index; Screen out the common sensitive parameters according to the preset criteria as the target parameter set for parameter optimization.

8. The method for optimizing the parameters of an ecosystem model based on parallel scheduling according to claim 7, characterized in that The common sensitive parameters are the parameters that simultaneously satisfy the total sensitivity index > 0.1, the mean value μ* ranks in the top 30%, and the proportion of the interaction effect index > 15%.

9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for optimizing the parameters of the ecosystem model based on parallel scheduling according to any one of claims 4 to 8.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for optimizing the parameters of the ecosystem model based on parallel scheduling according to any one of claims 4 to 8.

Citation Information

Cited By

  • Key factor identification method for influencing process industrial load to participate in demand response and computer equipment thereof

    CN120995056A