Flux data interpolation method, device and equipment based on vortex motion covariance and medium

By using a flux data interpolation method based on eddy covariance, and employing a preset strategy and residual prediction model to perform weighted processing and three-level residual correction on the flux data, the problem of poor versatility and low accuracy of the eddy covariance flux data interpolation method in cross-site and cross-ecological type application scenarios is solved, and high-precision and automated interpolation processing is achieved.

CN121834152APending Publication Date: 2026-04-10RAINROOT SCI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing eddy covariance flux data interpolation methods have poor versatility in cross-site and cross-ecological application scenarios, low accuracy when samples are insufficient, and low degree of automation.

Method used

A flux data interpolation method based on eddy covariance is adopted. The flux dataset is weighted and interpolated by a set of preset strategies. The method combines a few-shot learning model and an automatic machine learning model built by a meta-learning framework. The method uses historical flux data to predict residuals, performs three-level residual prediction and correction, and ensures the accuracy and reliability of the interpolation results through multi-dimensional quality assessment.

Benefits of technology

It improves the accuracy and versatility of high-throughput data interpolation, especially in scenarios with few samples, achieving high-precision automated interpolation processing, reducing manual intervention, and improving the efficiency and reliability of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834152A_ABST
    Figure CN121834152A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vortex motion covariance flux data processing, in particular to a flux data interpolation method and device based on vortex motion covariance, equipment and a medium. The method comprises the following steps: acquiring a flux data set to be interpolated; respectively performing weighted interpolation processing on the flux data set through strategies in a preset strategy set to obtain a basic interpolation result; determining a residual prediction model for performing prediction according to the basic interpolation result according to the relationship between the number of samples in the flux data set and a preset number threshold value; and based on the determined residual prediction model, predicting a residual value of the basic interpolation result to obtain a residual prediction value, and completing interpolation processing according to the residual prediction value. Compared with the prior art, on one hand, a mode which takes a result of a traditional method as a basis and is combined with a model to carry out recorrection is constructed, so that the interpolation precision is improved, and on the other hand, the universality for a few-sample scene and the interpolation precision under the few-sample scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of eddy covariance flux data processing technology, and in particular to a flux data interpolation method, apparatus, equipment and medium based on eddy covariance. Background Technology

[0002] In meteorology and fluid mechanics, eddies refer to the irregular, chaotic, and random motion of air or any fluid. Eddy covariance is a core method for observing carbon flux in terrestrial ecosystems. It directly calculates exchange rates by measuring turbulence, and is used to measure the exchange fluxes of CO2, H2O, and energy between terrestrial ecosystems and the atmosphere. Currently, eddy covariance flux data interpolation is mainly based on tools such as REddyProc (which uses eddy correlation methods to calculate fluxes of water vapor, carbon dioxide, methane, other trace gases, and energy). Commonly used methods include lookup table (LUT), diurnal mean variation (MDC), and marginal distribution sampling (MDS).

[0003] These methods rely on historical data statistics or similar environment matching, and have a certain degree of reliability when the data quality is high and the sample is sufficient. However, in cross-site and cross-ecological application scenarios, the above methods expose obvious limitations: on the one hand, the fixed parameters lead to poor adaptability, and different sites require repeated manual parameter tuning; on the other hand, the imputation error increases significantly when the sample is insufficient, and these methods have low automation and rely on human experience, which restricts the efficiency of batch processing.

[0004] Therefore, designing a throughput data interpolation method that can adapt to different data scales, improve the versatility of scenarios, and ensure data accuracy has become an urgent problem to be solved. Summary of the Invention

[0005] In view of the above-mentioned shortcomings and deficiencies of the prior art, this application provides a flux data interpolation method, apparatus, equipment and medium based on eddy covariance. The main purpose is to solve the problems of poor versatility, low accuracy when the sample size is small and low degree of automation of the current eddy covariance flux data interpolation scheme.

[0006] To achieve the above objectives, the main technical solutions adopted in this application include:

[0007] In a first aspect, embodiments of this application provide a flux data interpolation method based on eddy covariance, including:

[0008] Obtain the flux dataset to be interpolated;

[0009] The throughput dataset is weighted and interpolated using strategies from a preset strategy set to obtain the basic interpolation result.

[0010] Based on the relationship between the number of samples in the throughput dataset and a preset threshold, a residual prediction model for making predictions based on the basic interpolation results is determined; wherein, the residual prediction model includes a first prediction model or a second prediction model, the first prediction model is a few-shot learning model built on a meta-learning framework, which is configured to learn from historical throughput data of different ecological type sites as a cross-domain dataset to obtain transferable residual prediction capabilities as meta-knowledge.

[0011] The first prediction model and the second prediction model are trained by using historical flux dataset, environmental features of historical flux dataset and historical basic interpolation results as training inputs, and using the historical residuals calculated from the historical basic interpolation results and the real observations as training labels.

[0012] Based on the determined residual prediction model, the residual values ​​of the basic interpolation results are predicted to obtain the residual prediction values, and the interpolation process is completed according to the residual prediction values.

[0013] Optionally, based on the determined residual prediction model, the residual values ​​of the basic interpolation results are predicted to obtain the predicted residual values, including:

[0014] The environmental features corresponding to the basic interpolation results and the throughput dataset to be interpolated are input into the determined residual prediction model to obtain the residual prediction values; wherein, the environmental features include climate features, ecological features and spatiotemporal features;

[0015] The step of inputting the basic interpolation results and the environmental features corresponding to the flux dataset to be interpolated into the determined residual prediction model to obtain residual prediction values ​​specifically includes: performing a first-level residual prediction and correction on the basic interpolation results based on the climate features to obtain a first correction result; performing a second-level residual prediction and correction based on the ecological features and the first correction result to obtain a second correction result; and performing a third-level residual prediction and correction based on the spatiotemporal features and the second correction result to obtain the residual prediction values.

[0016] Optionally, the step of predicting the residual value of the basic interpolation result based on the determined residual prediction model to obtain the residual prediction value, and performing interpolation processing based on the residual prediction value, includes: adding the residual prediction value to the basic interpolation result to obtain the optimized interpolation result; performing uncertainty assessment and quality confidence assessment on the optimized interpolation result to obtain a quality confidence score; the uncertainty assessment and quality confidence assessment include at least statistical consistency test, spatiotemporal continuity constraint test, energy balance test, and model consistency test; if the quality confidence score is less than or equal to a preset quality confidence score threshold, adjusting the parameters of the determined residual prediction model, recalculating the optimized interpolation result and the corresponding quality confidence score, until the quality confidence score is greater than the preset quality confidence score threshold, and outputting the optimized interpolation result and the corresponding quality confidence score.

[0017] Optionally, determining the residual prediction model for prediction based on the basic imputation result according to the relationship between the number of samples in the throughput dataset and a preset number threshold includes: determining the first prediction model as the residual prediction model when the number of samples in the throughput dataset is less than or equal to the preset number threshold; obtaining the environmental features of the throughput dataset when the number of samples in the throughput dataset is greater than the preset number threshold; and determining the second prediction model in a preset model set based on the environmental features of the throughput dataset; wherein the preset model set includes several prediction models with different specialization algorithms, and each prediction model with a specialization algorithm has a mapping relationship with the environmental features.

[0018] Optionally, after determining the residual prediction model for prediction based on the basic imputation results according to the relationship between the number of samples in the throughput dataset and a preset number threshold, the method further includes: calculating a switching demand score based on the current number of samples, sample quality, and the historical performance of the currently determined residual prediction model; and if the switching demand score is higher than a preset score threshold, switching the currently determined residual prediction model to another one of the first prediction model or the second prediction model.

[0019] Optionally, the preset strategy set includes at least the lookup table method, the average daily variation method, and the marginal distribution sampling method; the step of performing weighted interpolation processing on the flux dataset using strategies from the preset strategy set to obtain the basic interpolation result includes: performing weighted interpolation processing on the flux dataset using the lookup table method, the average daily variation method, and the marginal distribution sampling method to obtain the lookup table interpolation result, the average daily variation method interpolation result, and the marginal distribution sampling method interpolation result; and performing weighted calculation on the lookup table interpolation result, the average daily variation method interpolation result, and the marginal distribution sampling method interpolation result to obtain the basic interpolation result.

[0020] Optionally, obtaining the throughput dataset to be imputed includes: determining initial throughput data; performing dimensional unification processing on the initial throughput data; performing outlier removal and missing value identification on the dimensional unification data to determine the data that needs to be imputed, so as to obtain the throughput dataset to be imputed.

[0021] Secondly, embodiments of this application provide a flux data interpolation device based on eddy covariance, comprising:

[0022] The acquisition unit is configured to acquire the throughput dataset to be interpolated.

[0023] The weighting unit is configured to perform weighted interpolation processing on the throughput dataset using strategies from a preset strategy set to obtain the basic interpolation result;

[0024] The determining unit is configured to determine a residual prediction model for making predictions based on the basic imputation results, based on the relationship between the number of samples in the throughput dataset and a preset threshold number; wherein the residual prediction model includes a first prediction model or a second prediction model, the first prediction model being a few-shot learning model built on a meta-learning framework, configured to learn from historical throughput data of different ecological type sites as a cross-domain dataset to obtain transferable residual prediction capabilities as meta-knowledge;

[0025] The first prediction model and the second prediction model are trained by using historical flux dataset, environmental features of historical flux dataset and historical basic interpolation results as training inputs, and using the historical residuals calculated from the historical basic interpolation results and the real observations as training labels.

[0026] The processing unit is configured to predict the residual values ​​of the basic interpolation results based on the determined residual prediction model, obtain the residual prediction values, and complete the interpolation processing based on the residual prediction values.

[0027] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the flux data interpolation method based on eddy covariance described in the first aspect.

[0028] Fourthly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the flux data interpolation method based on eddy covariance described in the first aspect.

[0029] Using the above technical solution, this application provides a flux data interpolation method based on eddy covariance. First, the flux dataset to be interpolated is acquired, and then weighted interpolation is performed on the flux dataset using strategies from a preset strategy set to obtain basic interpolation results. Next, based on the relationship between the number of samples in the flux dataset and a preset threshold, a residual prediction model is determined for prediction based on the basic interpolation results. The residual prediction model includes a first prediction model or a second prediction model. The first prediction model is a few-shot learning model built based on a meta-learning framework, configured to learn using historical flux data from different ecological type sites as a cross-domain dataset to obtain transferable residual prediction capabilities as meta-knowledge. Furthermore, both the first and second prediction models are trained using historical flux datasets, environmental features of historical flux datasets, and historical basic interpolation results as training inputs, and using historical residuals calculated from historical basic interpolation results and actual observations as training labels. Finally, based on the determined residual prediction model, the residual values ​​of the basic interpolation results are predicted to obtain residual prediction values, and interpolation is completed based on these residual prediction values. Compared to conventional methods in related technologies that rely solely on traditional interpolation methods to obtain interpolation results, this application, on the one hand, uses the basic interpolation results obtained through traditional methods as the data foundation. It then uses a first or second prediction model, pre-trained based on historical basic interpolation results and environmental features, to predict the residual values ​​of the basic interpolation results. This constructs a model that uses the results of traditional methods as a basis and further refines them using the model, thereby improving interpolation accuracy. On the other hand, using the number of samples as a criterion, it employs a first prediction model with transferable residual prediction capabilities to predict in scenarios with few samples, thus improving general applicability and interpolation accuracy in scenarios with few samples. Attached Figure Description

[0030] Figure 1 A flowchart illustrating a flux data interpolation method based on eddy covariance provided in this application embodiment;

[0031] Figure 2 A schematic diagram of the technical architecture of a flux data interpolation method based on eddy covariance provided in this application embodiment;

[0032] Figure 3 This is a schematic diagram of a flux data interpolation device based on eddy covariance provided in an embodiment of this application. Detailed Implementation

[0033] To better understand the above technical solutions, exemplary embodiments of this application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application can be understood more clearly and thoroughly, and that the scope of this application can be fully conveyed to those skilled in the art.

[0034] To address the shortcomings of traditional eddy covariance flux data interpolation schemes, such as poor versatility, low accuracy with small sample sizes, and low automation, this application proposes an eddy covariance-based flux data interpolation method. This method can be applied to environmental monitoring systems in fields such as agricultural management, forestry resource assessment, and carbon resource assessment. During runtime, it can execute any of the eddy covariance-based flux data interpolation methods mentioned below. Figure 1 As shown, the method includes:

[0035] S101, Obtain the throughput dataset to be interpolated.

[0036] The flux dataset to be interpolated refers to a dataset of flux data arranged in chronological order, obtained through data preprocessing processes such as format standardization, missing value identification, and outlier removal. This dataset contains several missing positions due to equipment, environmental, or other reasons. Flux data is the net rate of vertical transport of a substance or energy through a specific interface per unit area per unit time. For example, to observe the carbon dioxide flux in a certain area of ​​a forest (such as above a grassland), it is necessary to first collect the vertical wind speed and carbon dioxide concentration above the grassland, and then calculate the flux by taking the covariance of these two data points. The specific calculation process can be found in industry-standard flux calculation methods, which will not be elaborated upon here.

[0037] S102, weighted interpolation processing is performed on the throughput dataset using strategies from the preset strategy set to obtain the basic interpolation result.

[0038] Specifically, the preset strategy set includes at least the lookup table method, the average daily variation method, and the marginal distribution sampling method. These three strategies are relatively universal for interpolation calculations of flux data, and traditional methods use one of these three strategies for interpolation calculations. In this embodiment, instead of directly using a universal strategy, the interpolation results of the three universal strategies are calculated separately and then weighted and averaged to obtain the basic interpolation result as the data basis for subsequent processing.

[0039] S103, Based on the relationship between the number of samples in the throughput dataset and the preset quantity threshold, determine the residual prediction model used for prediction based on the basic interpolation results.

[0040] Before determining the residual prediction model, a process is introduced to assess the sample size of the throughput dataset, allowing for the selection of a specific residual prediction model when the sample size is insufficient. The sample size refers to the number of valid data points in the throughput dataset.

[0041] In S103, the residual prediction model includes either a first prediction model or a second prediction model. The first prediction model is a few-shot learning model built on a meta-learning framework, which can predict residuals better when the number of samples is insufficient. That is, the first prediction model is configured to learn from historical flux data of sites of different ecological types as a cross-domain dataset to obtain transferable residual prediction capabilities as meta-knowledge.

[0042] Specifically, in the meta-training phase, the first prediction model uses multiple imputation tasks constructed from historical flux data of various ecological types of sites as a meta-training set. The model is trained on this meta-training set, enabling it to learn the general mapping rules from environmental features and basic imputation results to residuals. Then, in the application phase, when applied to sites with a small sample size, a residual prediction model suitable for that site can be obtained by fine-tuning the parameters of the first prediction model obtained in the meta-training phase using only the limited samples provided by that site.

[0043] Furthermore, the above only illustrates the conceptual design of the first prediction model; the training process will be explained in detail below. The first and second prediction models are trained by using historical flux datasets, environmental features of the historical flux datasets, and historical basic interpolation results as training inputs, and by using the historical residuals calculated from the historical basic interpolation results and the actual observations as training labels. Environmental features include climate features (air temperature, soil moisture, and periodic features such as seasonal variations and diurnal variations), ecological features (photosynthetically active radiation, etc.), and spatiotemporal features (geographical location information, spatial altitude, time, etc.).

[0044] This section provides a specific implementation process to illustrate the training process. First, a historical flux dataset is acquired, such as the flux data of a specific gas arranged chronologically within three hours before and after the event. Corresponding environmental features are then obtained, including 10 features such as air temperature 25℃, soil moisture 0.3 units, and photosynthetically active radiation 1500 units. Next, a specific time point is considered missing (e.g., 12:30). Using three common imputation strategies, the imputation result for 12:30 is predicted, yielding three imputation results of -11.8, -12.3, and -11.5. These are then weighted to obtain the historical baseline imputation result of -11.9. Next, based on the actual observation value of 12:30 (-12.5) and the historical imputation result (-11.9), the historical residual of -0.6 is calculated. Following the method shown in the example above, the training input and training labels for the model are obtained, and the model is then trained to obtain the first and second prediction models.

[0045] Furthermore, it should be noted that although the training process described above is consistent for both the first and second prediction models, the models and datasets used differ. Specifically, the first prediction model is a meta-learning model that collects historical data from multiple different stations, such as station A (northern forest), station B (temperate farmland), and station C (tropical savanna). Each station provides more than two years of flux observation data, primarily aimed at quickly learning transferable residual prediction capabilities to adapt to different types and scenarios with a limited sample size. The second prediction model, on the other hand, can use historical flux data from the same station over the past five years.

[0046] In summary, the training process of the first and second prediction models is quite consistent. Both models are trained by using historical throughput data (historical basic interpolation results processed by three preset strategies) and corresponding environmental features as model inputs and historical residuals as model training labels. The difference lies in that the first prediction model adopts a few-shot learning model built with a meta-learning framework. During the training process, it also focuses on cultivating the ability to quickly learn transferable residual prediction capabilities when the sample size is small, so as to ensure the accuracy of residual prediction when the sample size is small in the application process.

[0047] S104. Based on the determined residual prediction model, the residual values ​​of the basic interpolation results are predicted to obtain the residual prediction values, and the interpolation process is completed based on the residual prediction values.

[0048] Referring to the training process example in S103, a specific application example is also provided in the application phase to illustrate the process of the technical solution. However, it should be noted that this example is only a simplified implementation method for understanding the solution; more detailed information is not shown, and it is not intended to limit the creativity of the technical solution.

[0049] One feasible application scenario is that when observing the carbon dioxide flux of a forest, data is missing at 12:30 due to equipment failure, and it is necessary to interpolate the flux observation values ​​collected at 12:30.

[0050] First, the flux dataset to be interpolated is obtained (S101). For example, the actual flux observation value at 12:00 is -10.5, 12:30 is missing, and 13:00 is -8.2. Here, {-10.5, missing, -8.2} is the flux dataset to be interpolated. Then, the 12:30 position is predicted using three general methods (S102), resulting in three values: -9.8, -10.1, and -9.5. The three values ​​are then weighted (the weights can be determined based on the historical performance of the data; in this embodiment, the weight ratio is 4:3.5:2.5), resulting in the basic interpolation result of -9.83.

[0051] The sample size is then compared with a preset threshold (S103). The preset threshold can be set to 50 to ensure a sufficient number of samples, thereby determining whether the residual prediction model is the first or second prediction model. Due to the small sample size in this example, this step will not be described in detail. After determining the residual prediction model (S104), the basic interpolation result -9.83 and the environmental features are input into the residual prediction model, resulting in a residual of +0.7 at 12:30. The final optimized interpolation result -9.13 is then used as the interpolation value at that point.

[0052] This example demonstrates that sample size is introduced as a criterion in S103, and a method using the first prediction model is proposed to address insufficient sample size, thus solving the problem of low accuracy when the sample size is small. Furthermore, this embodiment does not employ the conventional approach of "direct end-to-end flux prediction," but instead constructs a hybrid architecture combining traditional methods as a foundation with model residual correction. Even if the model corrections in S103 and S104 are completely ineffective, its accuracy will not be lower than related technologies, thus achieving higher accuracy.

[0053] In this embodiment, the flux dataset to be imputed is first acquired, and then weighted imputation is performed on the flux dataset using strategies from a preset strategy set to obtain basic imputation results. Next, based on the relationship between the number of samples in the flux dataset and a preset threshold, a residual prediction model is determined for prediction based on the basic imputation results. This residual prediction model includes either a first prediction model or a second prediction model. The first prediction model is a few-shot learning model built on a meta-learning framework, configured to learn using historical flux data from different ecological types of sites as a cross-domain dataset to obtain transferable residual prediction capabilities as meta-knowledge. Both the first and second prediction models are trained using the historical flux dataset, its environmental features, and historical basic imputation results as training inputs, and the historical residuals calculated from the historical basic imputation results and actual observations as training labels. Finally, based on the determined residual prediction model, the residual values ​​of the basic imputation results are predicted to obtain residual prediction values, and imputation is completed based on these residual prediction values. Compared to conventional methods in related technologies that rely solely on traditional interpolation methods to obtain interpolation results, this embodiment, on the one hand, uses the basic interpolation results obtained by traditional methods as the data foundation. It then uses a first or second prediction model, pre-trained based on historical basic interpolation results and environmental features, to predict the residual values ​​of the basic interpolation results. This constructs a model that uses the results of traditional methods as a basis and further refines them using the model, thereby improving interpolation accuracy. On the other hand, using the number of samples as a criterion, it employs a first prediction model with transferable residual prediction capabilities to predict in scenarios with few samples, thus improving versatility and interpolation accuracy in scenarios with few samples.

[0054] Furthermore, this embodiment also proposes a three-level residual prediction, interpolation result quality assessment, determination of the second prediction model through the AutoML model set, and an intelligent switching strategy between the first and second prediction models, which will be described in each embodiment below.

[0055] Optionally, based on a determined residual prediction model, the residual values ​​of the basic interpolation results are predicted to obtain residual prediction values, including: inputting the environmental features corresponding to the basic interpolation results and the flux dataset to be interpolated into the determined residual prediction model to obtain residual prediction values; wherein, the environmental features include climate features, ecological features and spatiotemporal features.

[0056] The environmental characteristics corresponding to the basic interpolation results and the flux dataset to be interpolated are input into the established residual prediction model to obtain residual prediction values. Specifically, this includes: performing a first-level residual prediction and correction on the basic interpolation results based on climate characteristics to obtain a first-level correction result; performing a second-level residual prediction and correction based on ecological characteristics and the first-level correction result to obtain a second-level correction result; and performing a third-level residual prediction and correction based on spatiotemporal characteristics and the second-level correction result to obtain residual prediction values.

[0057] In this embodiment, the first-level residual in the three-level residual system aims to uncover the difference between the basic interpolation results and the true values, focusing on climatic characteristics (such as seasons, diurnal variation, and large-scale weather patterns) and correcting systematic biases caused by long-term climatic patterns. For example, if the system identifies the current pattern as "high afternoon temperatures in summer," it will automatically call the correction parameters for this type of pattern to perform an initial coarse adjustment to the basic interpolation results. The second-level residual aims to uncover systematic biases influenced by ecological characteristics, fine-tuning the first-level correction results based on ecological characteristics (such as vegetation type, leaf area index, and soil moisture) to reflect the specific differences in the responses of different ecosystems to similar climatic conditions. The third-level residual introduces spatiotemporal characteristics, utilizing the continuity constraint of interpolation results at adjacent time points to locally smooth and fine-tune the second-level results, eliminating possible non-physical jumps or noise and ensuring the temporal smoothness and rationality of the interpolation sequence. By proposing a hierarchical three-level residual prediction and correction mechanism, the aim is to decompose the complex residual prediction task into a progressive optimization process from macro to micro and from system to local, so as to more accurately capture and correct interpolation errors from different sources and significantly improve interpolation accuracy.

[0058] Optionally, based on a determined residual prediction model, the residual values ​​of the basic interpolation results are predicted to obtain residual prediction values, and interpolation processing is performed based on the residual prediction values, including: adding the residual prediction values ​​to the basic interpolation results to obtain optimized interpolation results; performing uncertainty assessment and quality confidence assessment on the optimized interpolation results to obtain a quality confidence score; the uncertainty assessment and quality confidence assessment include at least statistical consistency tests, spatiotemporal continuity constraint tests, energy balance tests, and model consistency tests; if the quality confidence score is less than or equal to a preset quality confidence score threshold, the parameters of the determined residual prediction model are adjusted, and the optimized interpolation results and corresponding quality confidence scores are recalculated until the quality confidence score is greater than the preset quality confidence score threshold, and the optimized interpolation results and corresponding quality confidence scores are output.

[0059] In this embodiment, after generating the optimized interpolation results, a multi-dimensional automated quality assessment process is implemented. Statistical consistency testing ensures the results fall within a reasonable confidence interval based on historical error distributions, while spatiotemporal continuity constraint testing checks for abrupt changes in data at adjacent time points that violate physical laws. Energy balance testing verifies the rationality of the results from the perspective of ecosystem principles, targeting energy flux or carbon balance. Model consistency testing assesses the prediction dispersion of each sub-model in the model set, determining the degree of consensus (the model set, i.e., the AutoML model, will be explained later). These tests are combined into a quality confidence score of 0-100. If the score is below a preset threshold (e.g., 85 points), iterative optimization is automatically triggered, slightly adjusting the internal parameters of the residual prediction model, and then the interpolation and evaluation process is re-executed until a high-confidence result is produced. This improves the reliability and usability of the data product, significantly reduces the workload of subsequent manual quality checks, and achieves full automation and intelligence in the processing.

[0060] Optionally, based on the relationship between the number of samples in the throughput dataset and a preset threshold, a residual prediction model for prediction based on the basic interpolation results is determined, including: when the number of samples in the throughput dataset is less than or equal to the preset threshold, determining a first prediction model as a residual prediction model; when the number of samples in the throughput dataset is greater than the preset threshold, obtaining environmental features of the throughput dataset; and determining a second prediction model from a preset model set based on the environmental features of the throughput dataset; wherein the preset model set includes several prediction models with different specialization algorithms, and each prediction model with a specialization algorithm has a mapping relationship with the environmental features.

[0061] In this embodiment, when the number of samples is less than or equal to a preset threshold (e.g., 50), the first prediction model is selected for residual prediction. This embodiment focuses on the second prediction model, which is dynamically constructed based on the Automated Machine Learning (AutoML) paradigm. When the number of samples exceeds the preset threshold, the second prediction model is selected from a preset model set based on the environmental characteristics of the throughput dataset—that is, by AutoML automatically discovering effective features in the data, such as periodic patterns in time series and interactive effects of environmental factors. The preset model set includes several prediction models with different algorithmic strengths, such as time series models (LSTM), tree models (XGBoost), deep learning models (CNN), Transformer, LightGBM, Random Forest, ExtraTrees, CatBoost, neural ensemble models, and Stacking ensemble models. As a feasible implementation method, the AutoML mode can automatically tune parameters, using a Bayesian optimization algorithm to intelligently search for the optimal parameter combination of the model, avoiding the blindness and time-consuming nature of manual parameter tuning. Through the above methods, the prediction performance of the second prediction model is significantly improved.

[0062] Optionally, after determining the residual prediction model for prediction based on the basic imputation results according to the relationship between the number of samples in the throughput dataset and the preset number threshold, the method further includes: calculating a switching demand score based on the current number of samples, sample quality, and the historical performance of the currently determined residual prediction model; if the switching demand score is higher than the preset score threshold, switching the currently determined residual prediction model to another one of the first prediction model or the second prediction model.

[0063] In this embodiment, a dynamic intelligent switching strategy is proposed, which enables switching between a first prediction model and a second prediction model when the current prediction level is unsatisfactory after the residual prediction model has been determined. The switching depends on the current sample size, sample quality, and a switching requirement score determined by the historical performance of the currently determined residual prediction model. Specifically, the switching requirement score is calculated as score = α × sample_score + β × quality_score + γ × performance_score, where α, β, and γ are preset parameters, sample_score is the current sample size score, quality_score is the sample quality score, and performance_score is the historical performance score of the determined residual prediction model, and α + β + γ = 1. In other words, if the score is higher than the preset scoring threshold, the system switches to another prediction model for residual prediction. For example, if the sample size of a site just exceeds the threshold (e.g., 55 is greater than the threshold of 50), but the new batch of data is noisy and of poor quality, and historical records show that the first prediction model performs more robustly under this quality, then a high switching demand score is calculated, and the system decides to temporarily switch back to the first prediction model for processing. This ensures that the optimal interpolation strategy can be executed in various application scenarios, improving the level of intelligence and automation.

[0064] Optionally, the preset strategy set includes at least the lookup table method, the average daily variation method, and the marginal distribution sampling method. The flux dataset is weighted and imputed using the strategies in the preset strategy set to obtain the basic imputed result, including: weighting the flux dataset using the lookup table method, the average daily variation method, and the marginal distribution sampling method to obtain the lookup table imputed result, the average daily variation method imputed result, and the marginal distribution sampling method imputed result; and then weighting the lookup table imputed result, the average daily variation method imputed result, and the marginal distribution sampling method imputed result to obtain the basic imputed result.

[0065] In this embodiment, three algorithms—LUT, MDC, and MDS—are run. Crucially, the fusion weights are not fixed but dynamically calculated based on the long-term performance of each method on the site's historical validation dataset. Furthermore, the base imputation result serves as a baseline, exhibiting lower variance and higher stability. This dynamic weighted fusion strategy integrates the advantages of multiple traditional methods, offsetting the limitations of a single method.

[0066] Optionally, the throughput dataset to be imputed is obtained, including: determining the initial throughput data; performing dimension unification processing on the initial throughput data; removing outliers and identifying missing values ​​on the dimension-unified data to determine the data that needs to be imputed, so as to obtain the throughput dataset to be imputed.

[0067] In this embodiment, dimension unification means unifying and standardizing the dimensions of environmental factors (temperature, radiation, humidity, etc.) from different sensors to eliminate differences in numerical ranges; outlier removal means removing physically impossible values; missing value identification means marking the locations of uncollected values ​​and removed values ​​for subsequent interpolation processing, which is the basis for ensuring the quality of subsequent data processing.

[0068] The following section provides supplementary explanations regarding the technical architecture, case data, and execution of some steps in the specific applications described in the above embodiments. For example... Figure 2 As shown, a specific technical architecture is illustrated, including an input layer for integrating multi-site flux observation data and environmental factor data, a preprocessing layer for data cleaning, feature standardization and enhancement through intelligent algorithms, an adaptive learning layer for intelligently switching AI processing modes based on sample size, a residual learning layer for multi-level deep learning to optimize imputation residuals, and an output layer for providing high-precision imputation results and reliability assessment.

[0069] Specifically, the input layer acquires initial throughput data, including NEE, H, LE, and environmental variables. The preprocessing layer then performs multidimensional variable unification, outlier detection and handling, and periodic pattern extraction. Furthermore, it performs interpolation using Lookup Table (LUT), Mean Daily Variation (MDC), and Marginal Distribution Sampling (MDS) methods, respectively, and then fuses them according to weights to obtain the basic interpolation result. The weights are automatically assigned based on the performance of each method on historical validation data, with better-performing methods receiving higher weights. Finally, a weighted average is used to generate the baseline interpolation result. This method is more stable and reliable than relying on a single method.

[0070] Then, it is determined whether the sample size is greater than the preset threshold of 50, so as to determine the first prediction model or the second prediction model as the residual prediction model.

[0071] The first prediction model employs a few-shot meta-learning strategy, rapidly adapting and optimizing for current low-shot scenarios through cross-site knowledge transfer. The second prediction model uses AutoML, incorporating over 10 models. It automatically selects the optimal model based on environmental characteristics and performs automatic parameter tuning via hyperparameter Bayesian optimization (including the number of few-shot samples, attention heads, hidden layer dimensions, and ensemble model count). Climate, ecology, and spatial environmental features, along with the basic imputation results, are used as input to the residual prediction model. This model utilizes a spatiotemporal sequence encoding LSTM combined with an attention mechanism and a ResNet residual prediction network to ultimately output the predicted residual values.

[0072] In this process, an adaptive dual-mode AI learning framework is proposed, employing differentiated AI learning strategies based on varying data sample sizes. In cases of small samples, it leverages cross-domain knowledge transfer; in cases of sufficient samples, it capitalizes on the advantages of large-scale data modeling, achieving intelligent mode switching. This includes:

[0073] Small sample mode: When the number of valid samples is less than 50, the system starts the meta-learning mechanism, draws on the mature experience of other sites, and quickly adapts to the characteristics of the new site.

[0074] AutoML mode: When the number of valid samples reaches or exceeds 50, the system automatically searches for the optimal machine learning model, including various architectures such as neural networks and ensemble learning.

[0075] Intelligent switching mechanism: Based on dynamic evaluation of data quality, sample distribution and historical performance, it ensures the accuracy of mode selection.

[0076] After obtaining the residual values, the output layer undergoes quality control, specifically including statistical consistency tests, spatiotemporal continuity constraints, and ecological constraints, to obtain the final interpolation results and their confidence levels. Uncertainty assessment helps researchers determine the reliability of the interpolation results; low-confidence interpolated values ​​can be marked as requiring manual review, while high-confidence results can be directly used for scientific research analysis.

[0077] The flux data interpolation method based on eddy covariance provided by any of the above embodiments, in addition to using the basic interpolation results obtained by traditional interpolation methods as the data basis and constructing a mode that uses the results of traditional methods as the basis and combines them with the model for further correction, thereby improving the interpolation accuracy; and using the number of samples as the criterion, performing predictions through a first prediction model with transferable residual prediction capabilities for low-sample scenarios, thereby improving versatility and interpolation accuracy in low-sample scenarios, also has at least the following technical effects:

[0078] (1) A multi-level residual analysis mechanism was constructed, which decomposed the residual into global systematic deviation, environmental condition-related deviation, and spatiotemporal continuity fine-tuning residual. Through hierarchical modeling and progressive correction, the interpolation errors from different sources were systematically identified and corrected, which significantly improved the interpretability and final accuracy of the interpolation model under complex environmental conditions.

[0079] (2) Adaptive dual-mode AI learning framework. Different model strategies are adopted according to different sample sizes, including small sample mode and AutoML mode. AutoML can automatically analyze data characteristics, select the most suitable model architecture, and perform fine-tuning of parameters. The optimal second prediction model can be determined without manual intervention.

[0080] (3) Intelligent switching mechanism: After mode selection, the system continuously performs dynamic evaluation based on sample quality, data distribution and real-time model performance. When the evaluation indicators indicate that the current mode is not optimal, it can automatically and seamlessly switch between the first prediction model and the second prediction model, ensuring that the system can adopt the most suitable algorithm strategy under any data conditions, thereby enhancing the robustness and scenario adaptability of the overall solution.

[0081] (4) Based on the attention mechanism of environmental characteristics, the model can automatically assign dynamic weights to different environmental characteristics (such as climate, ecology, and spatiotemporal characteristics) during the multi-level residual analysis process, so that the network focuses on the most critical influencing factors under the current conditions. This significantly improves the interpolation accuracy and consistency of sites across different ecological types.

[0082] (5) Quality control and output: The optimized interpolation results are rigorously verified through statistical consistency tests, spatiotemporal continuity constraint tests, energy balance tests, and model consistency tests, and quality scores and confidence intervals are generated. Not only are the interpolated values ​​output, but also their credible quantitative indicators are output, providing a reliable data quality basis for subsequent scientific analysis, and the reliability of low-quality results can be automatically improved through iterative optimization closed loop.

[0083] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides a flux data interpolation device based on eddy covariance, such as... Figure 3 As shown, the device includes:

[0084] Acquisition unit 301 is configured to acquire the throughput dataset to be interpolated.

[0085] Weighting unit 302 is configured to perform weighted interpolation processing on the throughput dataset using strategies from a preset strategy set to obtain basic interpolation results;

[0086] The determining unit 303 is configured to determine a residual prediction model for making predictions based on the basic imputation results, based on the relationship between the number of samples in the throughput dataset and a preset threshold number; wherein, the residual prediction model includes a first prediction model or a second prediction model, the first prediction model is a few-shot learning model built based on a meta-learning framework, and is configured to learn using historical throughput data from different ecological type sites as a cross-domain dataset to obtain transferable residual prediction capabilities as meta-knowledge;

[0087] The first prediction model and the second prediction model are trained by using historical flux dataset, environmental features of historical flux dataset and historical basic interpolation results as training inputs, and using the historical residuals calculated from the historical basic interpolation results and the real observations as training labels.

[0088] The processing unit 304 is configured to predict the residual value of the basic interpolation result based on the determined residual prediction model, obtain the residual prediction value, and complete the interpolation process according to the residual prediction value.

[0089] In specific application scenarios, the processing unit 304 is further configured to input the basic interpolation result and the environmental features corresponding to the throughput dataset to be interpolated into the determined residual prediction model to obtain the residual prediction value; wherein, the environmental features include climate features, ecological features and spatiotemporal features.

[0090] In a specific application scenario, the processing unit 304 is further configured to perform a first-level residual prediction and correction on the basic interpolation result based on the climate characteristics to obtain a first correction result; perform a second-level residual prediction and correction based on the ecological characteristics and the first correction result to obtain a second correction result; and perform a third-level residual prediction and correction based on the spatiotemporal characteristics and the second correction result to obtain the residual prediction value.

[0091] In a specific application scenario, the processing unit 304 is further configured to add the residual prediction value to the basic interpolation result to obtain an optimized interpolation result; perform uncertainty assessment and quality confidence assessment on the optimized interpolation result to obtain a quality confidence score; the uncertainty assessment and quality confidence assessment include at least statistical consistency test, spatiotemporal continuity constraint test, energy balance test, and model consistency test; if the quality confidence score is less than or equal to a preset quality confidence score threshold, adjust the parameters of the determined residual prediction model, recalculate the optimized interpolation result and the corresponding quality confidence score, until the quality confidence score is greater than the preset quality confidence score threshold, and output the optimized interpolation result and the corresponding quality confidence score.

[0092] In a specific application scenario, the determining unit 303 is further configured to: determine the first prediction model as the residual prediction model when the number of samples in the throughput dataset is less than or equal to the preset number threshold; acquire the environmental features of the throughput dataset when the number of samples in the throughput dataset is greater than the preset number threshold; and determine the second prediction model from a preset model set based on the environmental features of the throughput dataset. The preset model set includes several prediction models with different specialization algorithms, and each specialization algorithm prediction model has a mapping relationship with the environmental features.

[0093] In a specific application scenario, the determining unit 303 is further configured to calculate a switching requirement score based on the current number of samples, sample quality, and the historical performance of the currently determined residual prediction model; if the switching requirement score is higher than a preset score threshold, the currently determined residual prediction model is switched to another one of the first prediction model or the second prediction model.

[0094] In specific application scenarios, the weighting unit 302 is further configured to perform weighted interpolation processing on the throughput dataset using the lookup table method, the average daily variation method, and the marginal distribution sampling method, respectively, to obtain the interpolation results of the lookup table method, the average daily variation method, and the marginal distribution sampling method; and to perform weighted calculation on the interpolation results of the lookup table method, the average daily variation method, and the marginal distribution sampling method to obtain the basic interpolation result.

[0095] In specific application scenarios, the acquisition unit 301 is further configured to: determine initial throughput data; perform dimensional unification processing on the initial throughput data; perform outlier removal and missing value identification on the data after dimensional unification processing; determine the data that needs to be imputed; and obtain the throughput dataset to be imputed.

[0096] It should be noted that other corresponding descriptions of the functional units involved in the flux data interpolation device based on eddy covariance provided in this embodiment can be found in [reference]. Figure 1 and Figure 2 The corresponding description in [the document] will not be repeated here.

[0097] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method shown.

[0098] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0099] Based on the above, Figure 1 and Figure 2 The method shown, and Figure 3To achieve the above objectives, this application also provides an electronic device, which can be configured on a computer side, etc. The device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to achieve the above-described objectives. Figure 1 and Figure 2 The method shown.

[0100] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.

[0101] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0102] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0103] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented using software plus necessary general-purpose hardware platforms, or it can be implemented using hardware. By applying the scheme of this embodiment, compared with the conventional methods in related technologies that only use traditional interpolation methods to obtain interpolation results, this embodiment, on the one hand, uses the basic interpolation results obtained by the traditional interpolation method as the data basis, and predicts the residual values ​​of the basic interpolation results using a first prediction model or a second prediction model pre-trained based on historical basic interpolation results and environmental characteristics. This constructs a pattern based on the results of the traditional method, combined with the model for further correction, thereby improving the interpolation accuracy. On the other hand, using the sample size as a criterion, a first prediction model with transferable residual prediction capabilities is used to predict in low-sample scenarios, thereby improving versatility and interpolation accuracy in low-sample scenarios.

[0104] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0105] In the description of this specification, the terms "one embodiment," "some embodiments," "embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0106] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make modifications, alterations, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A method for flux data interpolation based on eddy covariance covariance, characterized in that, The method comprises: obtaining a flux data set to be interpolated; performing weighted interpolation on the flux data set by a strategy in a preset strategy set to obtain a basic interpolation result; determining a residual prediction model for predicting based on the basic interpolation result according to a relationship between a sample quantity in the flux data set and a preset quantity threshold; wherein the residual prediction model comprises a first prediction model or a second prediction model, the first prediction model is a small sample learning model constructed based on a meta-learning framework, and is configured to learn with historical flux data of different ecological type sites as a cross-domain data set to obtain a transferable residual prediction capability as meta-knowledge; the first prediction model and the second prediction model are trained by taking the historical flux data set, environmental characteristics of the historical flux data set, and historical basic interpolation results as training inputs, and taking historical residuals calculated from the historical basic interpolation results and true observation values as training labels; based on the determined residual prediction model, predicting a residual value of the basic interpolation result to obtain a residual prediction value, and completing interpolation processing according to the residual prediction value.

2. The method of claim 1, wherein, based on the determined residual prediction model, predicting a residual value of the basic interpolation result to obtain a residual prediction value, comprising: inputting the basic interpolation result and the environmental characteristics corresponding to the flux data set to be interpolated into the determined residual prediction model to obtain a residual prediction value; wherein the environmental characteristics include climate characteristics, ecological characteristics, and spatio-temporal characteristics; the inputting the basic interpolation result and the environmental characteristics corresponding to the flux data set to be interpolated into the determined residual prediction model to obtain a residual prediction value specifically comprises: based on the climate characteristics, performing first-level residual prediction and correction on the basic interpolation result to obtain a first correction result; based on the ecological characteristics and the first correction result, performing second-level residual prediction and correction to obtain a second correction result; based on the spatio-temporal characteristics and the second correction result, performing third-level residual prediction and correction to obtain the residual prediction value.

3. The method of claim 1, wherein, the predicting a residual value of the basic interpolation result based on the determined residual prediction model to obtain a residual prediction value, and completing interpolation processing according to the residual prediction value, comprising: adding the residual prediction value to the basic interpolation result to obtain an optimized interpolation result; performing uncertainty evaluation and quality confidence evaluation on the optimized interpolation result to obtain a quality confidence score; the uncertainty evaluation and quality confidence evaluation at least include statistical consistency test, spatio-temporal continuity constraint test, energy balance test, and model consistency test; in the case that the quality confidence score is less than or equal to a preset quality confidence score threshold, adjusting parameters of the determined residual prediction model, recalculating the optimized interpolation result and the corresponding quality confidence score, until the quality confidence score is greater than the preset quality confidence score threshold, and outputting the optimized interpolation result and the corresponding quality confidence score.

4. The method of claim 1, wherein, The residual prediction model used for prediction according to the basic interpolation result is determined according to a relationship between a sample quantity in the flux data set and a preset quantity threshold, and the residual prediction model comprises: In a case where the sample quantity in the flux data set is less than or equal to the preset quantity threshold, the first prediction model is determined as the residual prediction model; In a case where the sample quantity in the flux data set is greater than the preset quantity threshold, an environmental feature of the flux data set is obtained; The second prediction model is determined from a preset model set according to the environmental feature of the flux data set, and the preset model set comprises prediction models of several different feature algorithms, and each prediction model of a feature algorithm has a mapping relationship with an environmental feature.

5. The method of claim 1, wherein, After the residual prediction model used for prediction according to the basic interpolation result is determined according to the relationship between the sample quantity in the flux data set and the preset quantity threshold, the method further comprises: A switching demand score is calculated based on a current sample quantity, a sample quality and a historical performance of the currently determined residual prediction model; In a case where the switching demand score is higher than a preset score threshold, the currently determined residual prediction model is switched to another one of the first prediction model or the second prediction model.

6. The method of claim 1, wherein, The preset strategy set at least comprises a lookup table method, an average daily variation method and a marginal distribution sampling method; The flux data set is weighted and interpolated by using a strategy in the preset strategy set to obtain a basic interpolation result, and the method comprises: The flux data set is weighted and interpolated by using the lookup table method, the average daily variation method and the marginal distribution sampling method to obtain a lookup table interpolation result, an average daily variation interpolation result and a marginal distribution sampling interpolation result; The lookup table interpolation result, the average daily variation interpolation result and the marginal distribution sampling interpolation result are weighted and calculated to obtain the basic interpolation result.

7. The method of claim 1, wherein, The flux data set to be interpolated comprises: Initial flux data is determined; Dimensional uniformity processing is performed on the initial flux data; After the dimensionally uniform data is processed, outliers are removed and missing values are identified to determine data that needs to be interpolated to obtain the flux data set to be interpolated.

8. A flux data interpolation device based on eddy covariance covariance, characterized by, Comprise: An acquisition unit is configured to acquire a flux data set to be interpolated; A weighting unit is configured to perform weighted interpolation on the flux data set by using a strategy in a preset strategy set to obtain a basic interpolation result; A determination unit is configured to determine a residual prediction model used for prediction according to the basic interpolation result according to a relationship between a sample quantity in the flux data set and a preset quantity threshold, and the residual prediction model comprises a first prediction model or a second prediction model, the first prediction model is a small sample learning model constructed based on a meta-learning framework, and is configured to learn historical flux data of different ecological type sites as a cross-domain data set to obtain a transferable residual prediction capability as meta-knowledge. The first prediction model and the second prediction model are trained by taking a historical flux dataset, environmental features of the historical flux dataset, and a historical base interpolation result as training inputs, and taking a historical residual calculated from the historical base interpolation result and a real observation value as a training label. The processing unit is configured to predict a residual value of the base interpolation result based on the determined residual prediction model to obtain a residual prediction value, and complete the interpolation processing according to the residual prediction value.

9. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, The processor executes the computer program to implement the method in any one of claims 1 to 7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method in any one of claims 1 to 7.