A data assimilation scheme based on twin model

Through the data assimilation scheme based on the twin model, the problems of incomplete mathematical equations and insufficient spatial and temporal representativeness of parameters in the land surface process model simulation are solved, the data assimilation process is simplified, the model simulation accuracy and forecasting ability are improved, and efficient data assimilation effects are achieved.

CN114692512BActive Publication Date: 2025-09-05INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210476944.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-02
Publication Date
2025-09-05
Estimated Expiration
2042-05-02

AI Technical Summary

Technical Problem

Existing land surface process model simulations have problems such as incomplete mathematical equations, large differences from actual conditions due to simplification, insufficient temporal and spatial representativeness of parameters, and unclear biophysical responses at different spatial scales. In addition, traditional data assimilation methods are complex, time-consuming and labor-intensive to operate, and have poor cross-platform portability, making it difficult to effectively improve model simulation accuracy.

Method used

A data assimilation scheme based on twin models is adopted. By constructing twin models of process models and observation models, and using methods such as multivariate regression and deep learning, the time evolution and physical connection of state variables are reconstructed, and the twin assimilation amount is calculated for data assimilation.

Benefits of technology

It simplifies the data assimilation process, improves the simulation accuracy and forecasting capability of land surface process models, reduces the workload of professional and technical personnel, and improves the cost-effectiveness of data assimilation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114692512B_ABST
    Figure CN114692512B_ABST
Patent Text Reader

Abstract

The present invention relates to a twin model-based data assimilation scheme, comprising the following steps: 1) constructing process twin models of the process models at different times and observation twin models of the observation models at different times using process model simulations and observations of the study area at different times; 2) resimulating the process twin models at different times, taking into account simulation errors, to obtain twin simulations of each state variable at different times in the study area, and converting them into observation space using the observation twin model to obtain twin observations; 3) calculating the twin assimilation amount of each state variable in each grid in the study area using the twin observations of each state variable at different times and actual observations; and 4) assimilating each state variable of the process model at the assimilation time using the twin assimilation amount. The present invention makes data assimilation convenient, easy to implement, flexible, and efficient, freeing technical personnel from the tedious task of modifying assimilation code and allowing them to focus on improving the effectiveness of process model simulation through data assimilation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data assimilation scheme for a land surface process model, and in particular to a scheme for data assimilation using a twin model of a process model. Background Art

[0002] Land surface processes are the sum of the physical, chemical, plant physiological, human, and microbial processes that occur within and between the various systems that comprise the land surface. They are a crucial component of the Earth's macrosystem. Land surface process models are important methods for studying land surface processes and the interplay of matter, energy, radiation, moisture, and momentum. They are also crucial for simulating land surface conditions such as surface temperature, soil temperature, and soil moisture. They are also crucial tools for forecasting land surface fluxes such as surface runoff, nitrogen, phosphorus, and carbon, surface evapotranspiration, and sensible heat. Furthermore, they are crucial techniques for assessing the effects of human activities on natural forcing.

[0003] Currently, the simulation of land surface process models still needs improvement. The main reasons are: 1) incomplete mathematical equations: Land surface process models are composed of a large number of mathematical equations, which are abstract descriptions of various land surface processes. During the description process, researchers ignore certain characteristics of land surface processes due to their limited understanding and knowledge; 2) oversimplification of mathematical equations: Due to simplifications and assumptions in the derivation process, the mathematical equations constructed are basically descriptive expressions of relatively ideal conditions, which differ from the actual situation; 3) the parameters in land surface process models are not temporally and spatially representative: Some parameters have strong heterogeneity in both time and space, but for practical reasons, only experimental observation parameters can be used, such as the thermal diffusivity and water diffusivity of soil and the relative emissivity of the surface, which invisibly reduces the simulation accuracy of land surface process models; 4) the underlying biophysical responses of land surface process models at different spatial scales are unclear: the response unit line of the more representative runoff generation process has strong spatial heterogeneity. This parameter is very important in the study of watershed hydrological processes, but it is currently used as the default value in hydrological process models.

[0004] With the enrichment of surface observation data, using data assimilation to improve land surface process model simulation has become an important method. Currently, data assimilation methods are divided into sequential data assimilation, continuous data assimilation and a combination of the two. However, improving land surface process model simulations through data assimilation is not an easy task. 1) Data assimilation requires multiple datasets, including soil properties, vegetation biophysical characteristics, meteorological forcing, river network data, human activity data, and dam data. These datasets must also have high spatial and temporal resolution, making them challenging to prepare. Currently, the temporal resolution of all data, except for meteorological forcing, needs to be improved. 2) Data preparation is time-consuming and labor-intensive. Data preparation requires significant effort and time for preprocessing, such as format conversion, cropping, and temporal and spatial interpolation. Changes in the study area often require repetition of similar tasks. 3) Data assimilation involves modifying the land surface process model framework. Land surface process models are typically written in Fortran. Modifying land surface process models requires both familiarity with Fortran and an understanding of the overall land surface process model framework to integrate the assimilation method. Implementing assimilation in other programming languages ​​is even more challenging. 4) Assimilation methods are difficult to port across different land surface process models or platforms. Porting requires the installation of various software packages, which are sometimes incompatible and can easily interfere with each other, leading to failures in land surface process model assimilation.

[0005] In view of the current shortcomings of the difficulty in assimilating land surface process model data, it is urgent to develop a data assimilation solution that is convenient to operate, easy to implement, flexible and efficient, so as to solve the three-step problem of changing the model, changing the algorithm and changing the framework in the traditional data assimilation method, get rid of the three-step dilemma of the traditional data assimilation method, which is time-consuming, labor-intensive and mentally demanding, and overcome the shortcomings of the traditional data assimilation method such as poor portability, weak cross-platform and low inheritance. The physical connection between the state variables and diagnostic variables of the land surface process model is used to construct an equivalent model of the land surface process model and the observation model to achieve the purpose of data assimilation, so as to effectively improve the cost-effectiveness of data assimilation. Summary of the Invention

[0006] Aiming at the difficulty of executing existing data assimilation methods, a data assimilation scheme based on twin model is provided.

[0007] To achieve the above object, the technical solution of the present invention is as follows:

[0008] A data assimilation scheme based on a twin model includes the following steps:

[0009] S1: Using the process model simulation and observation at different times in the study area, a process twin model of the process model at different times and an observation twin model of the observation model at different times are constructed;

[0010] S2: Using the process twin model at different times to consider the simulation error and re-simulate the twin simulation of each state variable at different times in the study area, and use the observation twin model to convert it into the observation space to obtain the twin observation;

[0011] S3: Calculate the twin assimilation amount of each state variable of each grid in the study area using the twin observations of each state variable at different times and the actual observations;

[0012] S4: Use the twin assimilation quantity to assimilate the state variables of the process model at the assimilation moment.

[0013] The data assimilation scheme based on twin models, wherein S1 includes: using the process models of the study area at different times to simulate and construct process twin models of the process models at different times, specifically:

[0014]

[0015] Where t represents the assimilation time, I represents the number of grids simulated by the process model in the study area, T represents the assimilation length, i represents the grid number, PM represents the process model, and X t+τi 、X t+τ-1i They represent the simulation of the state vector of the i-th grid at time t+τ and time t+τ-1 in the study area, respectively. t+τ-1i represents the auxiliary data set required by PM to simulate the state vector of the i-th grid in the study area at time t+τ-1, TPM t+τ-1 To use {X t+τ-1i} 1≤i≤I 、{X t+τi} 1≤i≤I 、{P t+τ-1i} 1≤i≤I The PM process twin model constructed at time t+τ-1, ε t+τ-1i Indicates TPM t+τ-1 The error when simulating the state vector of the i-th grid in the study area;

[0016] The observation twin model of the observation model at different times is constructed by using the process model simulation and observation at different times in the study area, specifically:

[0017]

[0018] In the formula, {O t+τi} 1≤i≤I represents the actual observation of each grid in the study area at time t+τ, TOM t+τ To use {X t+τi} 1≤i≤I 、{O t+τi} 1≤i≤I 、{P t+τi} 1≤i≤IThe constructed observation twin model at time t+τ, ξ t+τi Represents TOM t+τ The state vector of the i-th grid in the study area is simulated as X t+τi Error in transforming to observation space.

[0019] The data assimilation solution based on the twin model, wherein S2 includes:

[0020] The twin simulations of the state variables at different times in the study area are obtained by resimulating the process twin models at different times taking into account the simulation errors. Specifically:

[0021]

[0022] In the formula, {ζ t+τ-1in} 1≤i≤I Indicates the grids {X t+τ-1i} 1≤i≤I The nth case simulation error, N represents the number of cases of simulation error considered, Represents the use of the t+τ-1 time process twin model TPM t+τ-1 The twin simulation of each state variable in each grid of the study area obtained by simulation under the n-th case simulation error;

[0023] The observation twin model is used to convert the twin simulation of each state variable of each grid at different times in the study area into the observation space to obtain twin observations, specifically:

[0024]

[0025] Where, Indicates the use of t+τ time observation twin model TOM t+τ Will Twin observations converted to observation space.

[0026] The data assimilation scheme based on the twin model, wherein S3 is specifically:

[0027]

[0028]

[0029]

[0030]

[0031] Where, κ ti To utilize the twin observations of each state variable at different times And actual observation t+τi} 1≤τ≤TThe calculated twin assimilation coefficient of the i-th grid state vector at the assimilation time t, R t+τi is the covariance matrix of the i-th grid observation vector at time t+τ, α ti is the calculated twin assimilation amount of the i-th grid at the assimilation time t, and ′ represents the matrix or vector transpose.

[0032] The data assimilation scheme based on the twin model, wherein S4 is specifically:

[0033]

[0034]

[0035] Where, X ti is the state vector of the i-th grid at the assimilation time t, is the twin assimilation amount α of the i-th grid at assimilation time t ti The assimilated state vector.

[0036] Compared with existing data assimilation technology, the beneficial effects of the present invention are:

[0037] The present invention provides a data assimilation scheme based on a twin model. Compared with traditional data assimilation methods, the present invention uses land surface process model simulation and observation data to construct twin models of the land surface process model and the observation model at different simulation times to describe the time evolution of the state vector and the physical connection between the state vector and the observation. The data assimilation scheme based on the twin model can make the forecast improvement of the land surface process state, model simulation, complementary integration with observations, and in-depth mining of model simulation information very convenient, avoiding the tedious work of model framework adjustment, assimilation method integration, and model loop operation in traditional data assimilation to improve land surface process model simulation. It can free professional and technical personnel from the heavy assimilation code modification and better focus their energy on the effect of data assimilation to improve process model simulation. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 A flow chart of the method in the embodiment provided by the present invention; DETAILED DESCRIPTION

[0039] The following will be combined with the accompanying drawings used in the embodiments of the present invention to fully and clearly explain the technical solutions and basic principles of the embodiments of the present invention. Obviously, the embodiments described are only representative embodiments of the present invention, not all embodiments. The embodiments cited are only used to explain the present invention, not to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] like Figure 1 As shown in the figure, a data assimilation scheme based on the twin model includes the following steps:

[0041] Step 1: Assimilation preparation, including assimilation length, land surface process model simulation and observation dataset;

[0042] Step 2: Twin model construction, using process model simulation and observation at different times in the study area to construct process twin models of process models at different times and observation twin models of observation models at different times;

[0043] Step 3: Generate twin simulations and twin observations. Use the process twin models at different times to re-simulate the state variables of the study area at different times by taking into account the simulation error, and use the observation twin model to convert them into the observation space to obtain twin observations.

[0044] Step 4: Generate twin assimilation quantities, and use the twin observations of each state variable at different times and actual observations to calculate the twin assimilation quantities of each state variable in each grid in the study area;

[0045] Step 5: Twin assimilation, using the twin assimilation amount to assimilate each state variable of the process model at the assimilation moment;

[0046] Step 6: Use the assimilated improved process model state vector to perform subsequent period optimization according to Step 2-Step 5.

[0047] In an exemplary embodiment, in step S1, the assimilation preparation includes assimilation length, land surface process model simulation and observation data set; one of the land surface process model simulation options may be the ERA5 hourly land surface data set from 1950 to the present, which covers the entire world with a spatial resolution of 0.1°×0.1° and includes simulations of 50 surface variables, including surface temperature, surface evapotranspiration, and soil water; or it may be the Second Modern Retrospective Analysis of Research and Applications (MERRA2) 3-hourly data set from January 1, 1980 to December 31, 2016, which covers the entire world with a spatial resolution of 0.5°×0.625° and includes 46 surface fluxes and diagnostic variables; the observation data set may be an observation vector composed of surface temperature products, soil moisture products, surface evapotranspiration products, etc.

[0048] In an exemplary embodiment, in step S2, the twin model construction uses process model simulation and observation at different times of the study area to construct process twin models of the process models at different times and observation twin models of the observation models at different times, including:

[0049] The process twin models of the process models at different times in the study area are simulated to construct the process twin models of the process models at different times, specifically:

[0050]

[0051] Where t represents the assimilation time, I represents the number of grids simulated by the process model in the study area, T represents the assimilation length, i represents the grid number, PM represents the process model, and X t+τi 、X t+τ-1i represents the simulation of the state vector of the i-th grid at time t+τ and time t+τ-1 in the study area, which consists of the simulation of state variables from the ERA5 or MERRA-2 simulation data set, P t+τ-1i It represents the auxiliary data set required by PM to simulate the state vector of the i-th grid in the study area at time t+τ-1, corresponding to the parameter simulation in the ERA5 or MERRA-2 simulation data set, TPM t+τ-1 To use {X t+τ-1i} 1≤i≤I 、{X t+τi} 1≤i≤I 、{P t+τ-1i} 1≤i≤I The PM process twin model constructed at time t+τ-1, ε t+τ-1i Indicates TPM t+τ-1 The error when simulating the state vector of the i-th grid in the study area; the process twin model is composed of {X t+τ-1i} 1≤i≤I 、{X t+τi} 1≤i≤I 、{P t+τ-1i} 1≤i≤I I data pairs composed of {(X t+τ-1i ,P t+τ-1i ),X t+τi} 1≤i≤I Constructed through multiple regression equations, deep learning, random forests, neural networks, support vector regression, or ridge regression;

[0052] The observation twin model of the observation model at different times is constructed by using the process model simulation and observation at different times in the study area, specifically:

[0053]

[0054] In the formula, {O t+τi} 1≤i≤I represents the actual observation of each grid in the study area at time t+τ, TOM t+τ To use {X t+τi} 1≤i≤I 、{O t+τi} 1≤i≤I 、{P t+τi} 1≤i≤I The constructed observation twin model at time t+τ, ξ t+τi Represents TOM t+τ The state vector of the i-th grid in the study area is simulated as Xt+τi Error in conversion to observation space; t+τi} 1≤i≤I The observation vector consists of grid data in the study area of ​​surface temperature products, soil moisture products, and surface evapotranspiration products; TOM t+τ By {X t+τi} 1≤i≤I 、{O t+τi} 1≤i≤I 、{P t+τi} 1≤i≤I I data pairs composed of {(X t+τi ,P t+τi ),O t+τi} 1≤i≤I Constructed through multiple regression equations, deep learning, random forests, neural networks, support vector regression, or ridge regression.

[0055] In an exemplary embodiment, in step S3, the generation of twin simulations and twin observations includes the following steps: using the process twin models at different times to re-simulate the state variables of the study area at different times by taking into account the simulation error, and converting the observation twin models into the observation space to obtain the twin observations.

[0056] The twin simulations of the state variables at different times in the study area are obtained by resimulating the process twin models at different times taking into account the simulation errors. Specifically:

[0057]

[0058] In the formula, {ζ t+τ-1in} 1≤i≤I Indicates the grids {X t+τ-1i} 1≤i≤I Obey (0, Δ t+τ-1i ) The nth case of the white noise simulation error of the multidimensional normal distribution, Δ t+τ-1i For X t+τ-1i The covariance matrix of , N represents the number of cases of simulation error considered, Represents the use of the t+τ-1 time process twin model TPM t+τ-1 The twin simulation of each state variable in each grid of the study area obtained by simulation under the n-th case simulation error;

[0059] The twin simulations at different times are converted to the observation space through the observation twin model to obtain twin observations, specifically:

[0060]

[0061] Where, Indicates the use of t+τ time observation twin model TOM t+τ Will Twin observations converted to observation space.

[0062] In an exemplary embodiment, in step S4, the generation of twin assimilation amounts is performed by using twin observations of each state variable at different times and actual observations to calculate the twin assimilation amounts of each state variable of each grid in the study area, specifically:

[0063]

[0064]

[0065]

[0066]

[0067] Where, κ ti To utilize the twin observations of each state variable at different times And actual observation t+τi} 1≤τ≤T The calculated twin assimilation coefficient of the i-th grid state vector at the assimilation time t, R t+τi is the covariance matrix of the i-th grid observation vector at time t+τ, α ti is the calculated twin assimilation amount of the i-th grid at the assimilation time t, and ′ represents the matrix or vector transpose.

[0068] In an exemplary embodiment, in step S5, the twin assimilation uses the twin assimilation amount to assimilate each state variable of the process model at the assimilation moment, specifically:

[0069]

[0070]

[0071] Where, X ti is the state vector of the i-th grid at the assimilation time t, is the state vector X of the i-th grid at assimilation time t ti Twin assimilation α ti The assimilated state vector.

[0072] The present invention is described in detail in the examples in this specification. The above embodiments are only used to illustrate the core idea of ​​the present invention. For those skilled in the art, the technical solution of the present invention can be improved based on the principles of the present invention, and these improvements should be within the scope of protection of the present invention.

Claims

1. A data assimilation scheme based on twin models, characterized in that: The following steps are involved: S1: Using the process model simulation and observation at different times in the study area, a process twin model of the process model at different times and an observation twin model of the observation model at different times are constructed; S2: Using the process twin model at different times to consider the simulation error and re-simulate the twin simulation of each state variable at different times in the study area, and use the observation twin model to convert it into the observation space to obtain the twin observation; S3: Use the twin observations of each state variable at different times and the actual observations to calculate the twin assimilation amount of each state variable in each grid in the study area, specifically: Where, κ ti To utilize the twin observations of each state variable at different times And actual observation t+τi } 1≤τ≤T The calculated twin assimilation coefficient of the i-th grid state vector at the assimilation time t, R t+τi is the covariance matrix of the i-th grid observation vector at time t+τ, α ti is the calculated twin assimilation amount of the i-th grid at the assimilation time t, ′ represents the matrix or vector transpose; S4: Use the twin assimilation quantity to assimilate the state variables of the process model at the assimilation moment.

2. A data assimilation solution based on a twin model as claimed in claim 1, characterized in that: Said S1 comprises: The process twin models of the process models at different times in the study area are simulated to construct the process twin models of the process models at different times, specifically: Where t represents the assimilation time, I represents the number of grids simulated by the process model in the study area, T represents the assimilation length, i represents the grid number, PM represents the process model, and X t+τi 、X t+τ-1i They represent the simulation of the state vector of the i-th grid at time t+τ and time t+τ-1 in the study area, respectively. t+τ-1i represents the auxiliary data set required by PM to simulate the state vector of the i-th grid in the study area at time t+τ-1, TPM t+τ-1 To use {X t+τ-1i } 1≤i≤I 、{X t+τi } 1≤i≤I 、{P t+τ-1i } 1≤i≤I The PM process twin model constructed at time t+τ-1, ε t+τ-1i Indicates TPM t+τ-1 The error when simulating the state vector of the i-th grid in the study area; The observation twin model of the observation model at different times is constructed by using the process model simulation and observation at different times in the study area, specifically: In the formula, {O t+τi } 1≤i≤I represents the actual observation of each grid in the study area at time t+τ, TOM t+τ To use {X t+τi } 1≤i≤I 、{O t+τi } 1≤i≤I 、{P t+τi } 1≤i≤I The constructed observation twin model at time t+τ, ξ t+τi Represents TOM t+τ The state vector of the i-th grid in the study area is simulated as X t+τi Error in transforming to observation space.

3. The data assimilation solution based on the twin model according to claim 1, characterized in that: Said S2 comprises: The twin simulations of the state variables at different times in the study area are obtained by resimulating the process twin models at different times taking into account the simulation errors. Specifically: In the formula, {ζ t+τ-1in } 1≤i≤I Indicates the grids {X t+τ-1i } 1≤i≤I The nth case simulation error, N represents the number of cases of simulation error considered, Represents the use of the t+τ-1 time process twin model TPM t+τ-1 The twin simulation of each state variable in each grid of the study area obtained by simulation under the n-th case simulation error; The observation twin model is used to convert the twin simulation of each state variable of each grid at different times in the study area into the observation space to obtain twin observations, specifically: Where, Indicates the use of t+τ time observation twin model TOM t+τ Will Twin observations converted to observation space.

4. The data assimilation solution based on the twin model according to claim 1, characterized in that: Said S4 is specifically: Where, X ti is the state vector of the i-th grid at the assimilation time t, is the state vector X of the i-th grid at assimilation time t ti Twin assimilation α ti The assimilated state vector.

Citation Information

Patent Citations

  • Twin body model construction method and device and computer equipment

    CN109684727A

  • Drainage basin hydrological model parameter dynamic estimation method based on digital twinborn technology

    CN114357716A