Method for processing carbon flux data obtained by vortex correlation method

The systematic processing method of carbon flux data obtained through the vortex correlation method solves the problems of many outliers, many missing data and complex processing processes, and realizes efficient and accurate processing of data, supporting scientific research and technical applications.

CN120196874APending Publication Date: 2025-06-24NATIONAL MARINE ENVIRONMENTAL MONITORING CENTRE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510252991.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the prior art, carbon flux data processing has problems such as many data outliers, many missing data, and complex processing processes, which makes it difficult to guarantee the accuracy and reliability of the data.

Method used

It provides a method for processing carbon flux data obtained by the vortex correlation method, including data acquisition, Eddypro software preprocessing, TOVI software postprocessing, Excel processing and R language interpolation, etc. Through comprehensive and systematic preprocessing and postprocessing, data cleaning, outlier value removal, missing data interpolation, etc. are realized.

Benefits of technology

Through this method, a complete, accurate and reliable carbon flux data set can be obtained, which improves the accuracy and reliability of data, optimizes data processing processes, reduces costs, improves efficiency, and supports scientific research in the fields of ecology, meteorology, etc.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196874A_ABST
    Figure CN120196874A_ABST
Patent Text Reader

Abstract

The invention discloses a method for processing carbon flux data obtained by a vortex correlation method. The method comprises the steps of data acquisition, original data preprocessing, data post-processing, data interpolation, data storage and additional description. The method comprises the following steps: capturing a signal by adopting an eddy correlation method, and recording original data in real time through a Smart Flux collector; in the preprocessing stage, invalid data, abnormal values, rainfall influence data and the like are removed by using Eddypro software, and necessary instrument error and inclination correction is carried out; in the post-processing stage, Eddypro and TOVI software are combined for further processing, including energy balance correction, quality control and the like; for missing data, adopting a linear interpolation method and other interpolation methods; the processed data can be stored in an Excel mode, an R language mode or a programming mode. According to the method, the carbon flux data quality can be effectively improved, and a solid foundation is provided for subsequent scientific research and data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of ecology, meteorology, and data processing. Specifically, it is a method for processing carbon flux data obtained by the eddy covariance method. This method is mainly applied to fields such as forestry carbon sink research, ecosystem carbon cycle monitoring, and urban carbon emission assessment. By directly measuring the carbon flux between the terrestrial ecosystem and the atmosphere through the eddy covariance method, it provides data support for evaluating the carbon sequestration capacity of ecosystems such as forests, wetlands, and urban green spaces, revealing the carbon cycle process, and monitoring the carbon emission status. Background Art

[0002] In the context of the increasingly in-depth research on global climate change and the urgent need for ecosystem carbon cycle management, accurately measuring and evaluating the carbon exchange amount between the ecosystem and the atmosphere has become a key area of scientific research. The eddy covariance method, as a direct, continuous, and non-destructive measurement technique, has become the core method in ecosystem carbon cycle research because it can provide high-precision surface-atmosphere gas exchange flux data. The core of this method lies in using high-precision sensors to measure the wind speed pulsation and gas concentration pulsation in real time, and calculating the carbon flux data through complex algorithms. These data are of crucial significance for understanding the carbon budget balance of the ecosystem, evaluating the carbon sequestration function, and formulating carbon management strategies.

[0003] However, the eddy covariance method faces many challenges in practical applications. First of all, the original carbon flux data is often disturbed by various noises and errors, including but not limited to the systematic errors of the instrument itself, the random interference of environmental factors, and data loss during data transmission. The existence of these noises and errors seriously reduces the accuracy and reliability of the data, making it risky to directly use the original data for scientific research.

[0004] To solve this problem, it is particularly important to effectively preprocess and postprocess the original carbon flux data. The preprocessing stage mainly includes data cleaning and quality control, aiming to eliminate invalid or abnormal data to ensure the accuracy of subsequent analysis. The postprocessing stage involves multiple links such as missing data imputation, data smoothing and filtering, time synchronization and calibration, aiming to further improve the integrity and reliability of the data. These processing steps are crucial for extracting a high-quality carbon flux data set that can be used for scientific research.

[0005] However, most traditional data processing methods rely on manual operations, which are not only time-consuming and laborious, but also vulnerable to human factors, making it difficult to guarantee data processing efficiency and accuracy. In addition, with the continuous progress of observation technology and the continuous expansion of the observation network, the amount of data obtained by the eddy covariance method has increased sharply, and traditional methods have been difficult to meet the needs of large-scale data processing. Therefore, it has become an urgent task to develop an automated and systematic carbon flux data processing method for the eddy covariance method. Summary of the Invention

[0006] The purpose of the present invention is to provide a processing method for carbon flux data obtained by the eddy covariance method, which is accurate and efficient, in order to solve the problems of many data outliers, many missing data, and complex processing process in the existing carbon flux data processing. Through the method of the present invention, comprehensive and systematic preprocessing and postprocessing can be realized for the original carbon flux data obtained by the eddy covariance method, including steps such as data cleaning, outlier removal, and missing data imputation, so as to obtain a complete, accurate, and reliable carbon flux data set.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A processing method for carbon flux data obtained by the eddy covariance method, comprising the following steps:

[0009] a) Data acquisition:

[0010] Use eddy covariance method equipment to capture signals to obtain original carbon flux data, the data including timestamp, air temperature, humidity, wind speed, wind direction, and carbon flux information;

[0011] b) Eddypro software preprocessing:

[0012] Import the original data into Eddypro software for preprocessing, specifically including:

[0013] Remove data with flag = 2;

[0014] Remove outliers and "wild points";

[0015] Remove data for half an hour before and after rainfall;

[0016] Remove data after the signal relative to the ultrasonic time band;

[0017] Perform delay time correction and coordinate axis rotation;

[0018] Export the 30-minute average data as a ".csv" format file;

[0019] c) TOVI software postprocessing:

[0020] The data processed by Eddypro software is imported into TOVI software for further processing, including:

[0021] Check the integrity of the data and remove obviously abnormal data based on the visual interface;

[0022] Fill in the missing time points in the biomet and full files in the ".csv" file to ensure that the time series of the two files meet the requirements;

[0023] Record data anomalies or missing data, including data source, data type, time period, reasons for exclusion, and other explanations;

[0024] Perform meteorological data interpolation and export processed data to Excel;

[0025] d) Excel processing:

[0026] The data processed by TOVI software was pre-processed in Excel to meet the requirements of R language interpolation, including:

[0027] Conduct flux data quality control based on different ecosystem types such as natural wetlands, de-aquaculture restoration areas, and aquaculture ponds, and remove unreasonable data points;

[0028] Check and quality control the indicators required for R language interpolation;

[0029] Create R language input files, sort the indicators, and fill in the missing data;

[0030] e) R language interpolation:

[0031] Use R language to interpolate missing data after Excel processing to obtain complete annual data, including:

[0032] Selection of interpolation method based on ecosystem type, including night-split or day-split methods;

[0033] After the imputation, the data status was checked, including NEE, GPP, and Re. If there were obvious abnormalities, the data were re-interpolated;

[0034] Confirm the validity of the data after interpolation.

[0035] Furthermore, the specific requirements for recording data anomalies or missing data in step c) include:

[0036] The data source must clearly identify the origin or ecosystem type of the data;

[0037] The data type needs to specify whether it is air temperature, humidity, wind speed, wind direction or carbon flux;

[0038] The time period shall cover data records for at least 5 consecutive days to analyze the long-term trends and abnormal patterns of the data;

[0039] The reasons for exclusion shall detail the specific problems that led to the data being excluded, such as sensor failures and abnormal data fluctuations;

[0040] The other description section can be used to record any additional information related to data anomalies or missing data, such as weather conditions and equipment maintenance records.

[0041] Furthermore, in step d), the specific operations for flux data quality control include:

[0042] For natural wetland ecosystems, during the growing season from May to October, negative carbon fluxes at night and significantly unreasonable positive carbon fluxes during the day are excluded because natural wetlands, as carbon sinks, absorb carbon dioxide during the day and release it at night;

[0043] For the restored wetland areas from fishpond reclamation, since it is an ecosystem in restoration, there may be carbon release during the day, but it must be a release at night. Therefore, only negative carbon fluxes at night are excluded;

[0044] For areas affected by human activities such as fishponds, since the data may be affected by multiple factors, data for both night and day are retained without exclusion;

[0045] During the processing, data points that deviate significantly from the overall trend shall also be excluded according to the annual data change trend to ensure the accuracy and representativeness of the data.

[0046] Furthermore, when using R language for missing data imputation in step e), the selection of the specific imputation method is based on the following:

[0047] For natural wetlands and restored wetland areas from fishpond reclamation, since the carbon flux data at night is relatively stable, the night splitting method is used for imputation;

[0048] For areas affected by human activities such as fishponds, since the carbon flux data during the day varies greatly, the day splitting method is used for imputation;

[0049] During the imputation process, seasonal variations and weather condition factors of the data shall also be considered to ensure the accuracy and rationality of the imputation results;

[0050] After imputation, the data shall be carefully checked, and if there are obvious anomalies, the imputation process shall be redone.

[0051] Furthermore, it also includes data visualization and analysis steps, specifically including:

[0052] Use charts and images to display the key steps and results in the data processing process, including data integrity checks, the effects of outlier removal, and interpolation results;

[0053] Through visualization analysis means, help users better understand data characteristics and change trends for quality control and result verification;

[0054] The visualization tools can include Excel, R language plotting packages, and professional data visualization software.

[0055] Furthermore, the eddy covariance method device includes an ultrasonic anemometer, an infrared gas analyzer, and related data acquisition and transmission equipment, with the specific configuration as follows:

[0056] The ultrasonic anemometer is used to measure three-dimensional wind speed and direction in real time to ensure the accuracy and timeliness of the data;

[0057] The infrared gas analyzer is used to measure the concentration changes of gases such as carbon dioxide to calculate the carbon flux;

[0058] The data acquisition and transmission equipment is responsible for transmitting the measured data to the data processing system in real time for subsequent analysis and processing;

[0059] The equipment also needs to be calibrated and maintained regularly to ensure the accuracy and stability of the measured data.

[0060] The processing method of the carbon flux data obtained by the eddy covariance method of the present invention has the following beneficial effects:

[0061] It is an accurate and efficient processing method for carbon flux data obtained by the eddy covariance method to solve problems such as many data outliers, many missing data, and complex processing processes in the existing carbon flux data processing. Through the method of the present invention, it is possible to comprehensively and systematically preprocess and postprocess the original carbon flux data obtained by the eddy covariance method, including steps such as data cleaning, outlier removal, and missing data interpolation, so as to obtain a complete, accurate, and reliable carbon flux data set;

[0062] Furthermore, it also includes:

[0063] Improve the accuracy and reliability of carbon flux data and provide high-quality data support for scientific research in fields such as ecology and meteorology;

[0064] Optimize the data processing process, reduce data processing costs, and improve data processing efficiency;

[0065] Promote the application and development of the eddy covariance method in fields such as forestry carbon sinks, ecosystem carbon cycle monitoring, and urban carbon emission assessment, and provide a scientific basis and technical support for addressing climate change, etc.

[0066] In summary, the present invention provides an efficient and accurate method for processing carbon flux data obtained by the eddy covariance method to meet the urgent needs of data processing in fields such as ecology and meteorology, and to promote scientific research and technological applications in related fields. Description of the Drawings

[0067] Figure 1 It is a programming example diagram of a method for processing carbon flux data obtained by the eddy covariance method of the present invention. Detailed Embodiments

[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0069] Embodiment:

[0070] I. Embodiment Background

[0071] The Eddy Covariance Method (EC) is an internationally recognized standard method for flux observation. It uses micrometeorological principles to estimate the covariance of vertical wind speed and the pulsation of substances or energy to directly measure the carbon flux between terrestrial ecosystems and the atmosphere. However, the original carbon flux data often contains a large number of outliers and missing data, and preprocessing and postprocessing are required to obtain reliable analysis results.

[0072] II. Specific Implementation Steps

[0073] 1. Data Collection and Preliminary Arrangement

[0074] Use eddy covariance method equipment (ultrasonic anemometer and infrared gas analyzer) for continuous observation, and collect original data containing information such as timestamp, air temperature, humidity, wind speed, wind direction, and carbon flux (Net Ecosystem Exchange, NEE);

[0075] Arrange the original data according to the time series to form a dataset containing all observed variables.

[0076] Specific implementation process:

[0077] Step 1.1: Use eddy covariance method equipment for continuous observation

[0078] Step 1.1.1: Equipment preparation:

[0079] Ensure that the eddy covariance method equipment (ultrasonic anemometer and infrared gas analyzer) is in good working condition and has been calibrated and maintained as necessary;

[0080] Install the device in a suitable location to avoid being affected by topographic, building, and vegetation interference factors.

[0081] Step 1.1.2: Observation parameter setting:

[0082] Set the observation parameters according to the research requirements, including the observation frequency, data recording interval, and data storage format;

[0083] Ensure that the device can record raw data containing timestamp, air temperature, humidity, wind speed, wind direction, and carbon flux (net ecosystem exchange NEE) in real time and accurately.

[0084] Step 1.1.3: Continuous observation:

[0085] Start the device for continuous observation. The observation time should be determined according to the research requirements. Usually, long-term (several months or years) continuous observation is required;

[0086] During the observation process, regularly check the device operation status and data recording situation to ensure the integrity and accuracy of the data.

[0087] Step 1.2: Organize the raw data according to the time series

[0088] Step 1.2.1: Data export:

[0089] Export the raw data from the eddy covariance method device. The data is usually stored in the form of binary files, text files, or database files.

[0090] Step 1.2.2: Data format conversion:

[0091] Convert the raw data into a suitable format according to the requirements of the subsequent data analysis software. For example, convert the data into CSV, Excel, or NetCDF format.

[0092] Step 1.2.3: Time series organization:

[0093] Organize the raw data according to the time series to ensure that each observed variable has a corresponding timestamp for subsequent time series analysis;

[0094] Check whether the time intervals in the data are uniform. If necessary, interpolate the data to fill in the missing time points.

[0095] Step 1.2.4: Dataset construction:

[0096] Merge the sorted data into a dataset that includes all observed variables, which should contain all observed variables such as timestamp, air temperature, humidity, wind speed, wind direction, and carbon flux;

[0097] Add necessary metadata to the dataset, such as the observation location, observation time range, and equipment model, for subsequent data analysis and applications.

[0098] Original data and dataset types

[0099] Original data type:

[0100] Timestamp: The time point when the observation data is recorded, usually represented in the form of year, month, day, hour, minute, and second;

[0101] Air temperature: The air temperature at the observation location, usually measured in degrees Celsius (°C);

[0102] Humidity: The air humidity at the observation location, usually expressed as relative humidity (%).

[0103] Wind speed: The wind speed at the observation location, usually measured in meters per second (m / s);

[0104] Wind direction: The wind direction at the observation location, usually expressed in degrees (°), representing the angle of the wind direction relative to the north;

[0105] Carbon flux: The carbon flux at the observation location, usually measured in micromoles per square meter per second (μmol / m 2 / s), representing the net ecosystem exchange (NEE).

[0106] Dataset type:

[0107] Time series dataset: A dataset that includes all observed variables arranged in a time series, which can be used for time series analysis and trend prediction studies;

[0108] Multidimensional dataset: A dataset that includes multiple observed variables, which can be used for multivariate statistical analysis and correlation analysis studies;

[0109] Metadata dataset: A dataset that includes metadata such as the observation location, observation time range, and equipment model, which can be used for data quality control and data traceability studies.

[0110] Through the above data collection and preliminary sorting process, complete, accurate, and reliable original data and datasets can be obtained, providing a solid foundation for subsequent data analysis and applications.

[0111] 2. Data preprocessing

[0112] Data cleaning and format conversion: Use Eddypro software for data preprocessing, including removing data with flag = 2 (i.e., unreliable data with a quality flag of 2), removing outliers and "wild points", removing data for half an hour before and after rainfall, and at the same time performing data format conversion to meet the requirements of subsequent analysis software;

[0113] Delay time correction and coordinate axis rotation: Synchronize the data of the ultrasonic anemometer and the infrared gas analyzer in time, and perform tilt correction (i.e., coordinate axis rotation) to eliminate the influence of terrain and instrument installation angle on the data.

[0114] Specific implementation process:

[0115] Step 2.1: Data cleaning and format conversion

[0116] Step 2.1.1: Selection of data preprocessing software

[0117] Select the professional data processing software Eddypro, which is designed specifically for data preprocessing of the eddy covariance method and has powerful data cleaning, format conversion and quality control functions.

[0118] Step 2.1.2: Removal of unreliable data

[0119] In the Eddypro software, filter according to the data quality flag (flag), especially remove the data with flag = 2, which are marked as unreliable due to instrument failure, poor environmental conditions or data processing errors.

[0120] Step 2.1.3: Removal of outliers and "wild points"

[0121] Use the outlier detection function in the Eddypro software to automatically or manually identify and remove outliers and "wild points" in the data. Outliers refer to data that deviate significantly from surrounding data points, while "wild points" refer to individual abnormal data points.

[0122] Step 2.1.4: Removal of data for half an hour before and after rainfall

[0123] According to rainfall records or meteorological data, identify rainfall events and remove the data within half an hour before and after rainfall, because rainfall will affect air flow and gas exchange, resulting in inaccurate carbon flux data measured by the eddy covariance method.

[0124] Step 2.1.5: Data format conversion

[0125] After completing data cleaning, the data is converted into the format required by subsequent analysis software, which usually includes exporting the data from binary files, text files, or database files in a specific format and converting it into a common format such as CSV, Excel, or NetCDF.

[0126] Step 2.2: Delay time correction and coordinate axis rotation

[0127] Step 2.2.1: Time synchronization

[0128] Perform time synchronization on the data of the ultrasonic anemometer and the infrared gas analyzer, which usually involves adjusting the clock settings of the two devices to ensure that the timestamps they record are exactly the same. If there is a slight time difference between the devices, corresponding time offset adjustments need to be made during data analysis.

[0129] Step 2.2.2: Tilt correction (coordinate axis rotation)

[0130] Perform tilt correction in the Eddypro software, which eliminates the influence of terrain and instrument installation angle on the data by rotating the coordinate axes. Tilt correction usually includes horizontal plane correction and vertical plane correction to ensure that the measured wind speed and direction are parallel or perpendicular to the ground.

[0131] Specific method of tilt correction: In Eddypro, users can input corresponding parameters according to the instrument installation angle and terrain features, and the software will automatically perform coordinate axis rotation.

[0132] Step 2.2.3: Verify the correction effect

[0133] After completing the delay time correction and coordinate axis rotation, it is necessary to verify the corrected data, which usually involves comparing the data differences before and after correction and checking whether the corrected data conforms to the expected physical laws (such as the rationality of wind speed and direction).

[0134] If problems are still found in the corrected data, it is necessary to recheck the correction parameters or use other methods for data preprocessing.

[0135] Through the implementation process of the above data preprocessing, it can ensure that the carbon flux data obtained by the eddy covariance method is more accurate and reliable in subsequent analysis. At the same time, this also provides a solid foundation for subsequent steps such as data quality control, outlier removal, and missing data imputation.

[0136] 3. Data quality control and outlier removal

[0137] In the TOVI post-processing software, obvious abnormal data such as data jumps and unreasonable extreme values are removed according to the visualization interface.

[0138] Further quality control is carried out according to the carbon cycle characteristics of the ecosystem, such as excluding negative values at night and positive values during the day in the growing season of natural wetlands, and excluding negative values at night in the restored areas of wetland restoration through returning farmland to wetland.

[0139] Specific implementation process:

[0140] Step 3.1: Exclude significantly abnormal data in post-processing software such as TOVI

[0141] Step 3.1.1: Software selection and data import

[0142] Select TOVI post-processing software as the tool for data quality control and outlier exclusion;

[0143] Import the carbon flux data obtained by the eddy covariance method after preliminary preprocessing into TOVI software.

[0144] Step 3.1.2: Visual interface inspection

[0145] Use the visual interface of TOVI software to visually display the data in the form of graphs, such as time series graphs and scatter plots;

[0146] Carefully check for significantly abnormal data such as jump points and unreasonable extreme values in the data through the visual interface.

[0147] Step 3.1.3: Outlier exclusion

[0148] According to the inspection results of the visual interface, manually or automatically exclude significantly abnormal data such as jump points and unreasonable extreme values in the data;

[0149] When excluding abnormal data, ensure that normal data is not accidentally deleted, and at the same time pay attention to maintaining the integrity and continuity of the data.

[0150] Step 3.1.4: Data saving and backup

[0151] Save the data set after excluding abnormal data as a new file for subsequent analysis;

[0152] At the same time, back up the original data and the data set after excluding abnormal data to prevent data loss or damage.

[0153] Step 3.2: Conduct further quality control according to the carbon cycle characteristics of the ecosystem

[0154] Step 3.2.1: Understand the carbon cycle characteristics of the ecosystem

[0155] Before conducting further quality control, it is necessary to deeply understand the carbon cycle characteristics of the ecosystem under study;

[0156] For example, during the growing season of natural wetlands, carbon absorption (negative value) is usually observed at night, while carbon emission (positive value) occurs during the day; in the restored area of wetland restoration, carbon absorption (negative value) may also be observed at night, but the value may vary due to the influence of restoration measures.

[0157] Step 3.2.2: Set quality control rules

[0158] According to the carbon cycle characteristics of the ecosystem, set reasonable quality control rules;

[0159] For example, for the data during the growing season of natural wetlands, the rule can be set as follows: the data at night should be negative, and the data during the day should be positive; if the data at night shows a positive value or the data during the day shows a negative value, it is regarded as abnormal data and should be excluded.

[0160] Step 3.2.3: Apply quality control rules

[0161] Apply the set quality control rules to the dataset, and automatically or manually exclude the data that does not conform to the rules;

[0162] When applying the rules, pay attention to the influence of factors such as the seasonality of the data and weather conditions on the data, and avoid mistakenly deleting normal data.

[0163] Step 3.2.4: Result verification and adjustment

[0164] Verify the dataset after excluding abnormal data to ensure that the data conforms to the carbon cycle characteristics of the ecosystem;

[0165] If it is found that there are still abnormal data, it is necessary to recheck the quality control rules and make necessary adjustments.

[0166] Step 3.2.5: Data saving and reporting

[0167] Save the dataset after further quality control as the final analysis dataset;

[0168] Prepare a data quality control report, which details the process, method, results of quality control, and any necessary adjustment instructions.

[0169] Through the above implementation process of data quality control and outlier exclusion, it can be ensured that the carbon flux data obtained by the eddy covariance method is more accurate and reliable, providing a solid foundation for subsequent scientific research and analysis.

[0170] 4. Missing data imputation

[0171] Process the data in Excel to meet the requirements of R language imputation, including filling in the missing time points and creating an R language input file;

[0172] Use R language for missing data imputation, and select appropriate imputation methods according to the ecosystem type, such as night splitting method, day splitting method, average daily variation curve method, and look-up table method. During the imputation process, factors such as seasonal variations and weather conditions of the data need to be considered.

[0173] Specific implementation process:

[0174] Step 4.1: Process the data in Excel to meet the requirements of R language imputation

[0175] Step 4.1.1: Fill in the missing time points

[0176] In Excel, first check whether the time series of the data set is complete and identify the missing time points;

[0177] For the missing time points, decide whether to fill them in according to the sampling frequency of the data (such as every hour, every day) and the characteristics of the ecosystem;

[0178] If filling in is required, corresponding rows or columns can be inserted in Excel and the missing timestamps can be filled; for the missing data values, they can be temporarily marked as NA or empty values.

[0179] Step 4.1.2: Create an R language input file

[0180] Save the processed data as an Excel file, ensuring that the file format matches the input requirements of R language;

[0181] If the R language script requires a specific data format (such as CSV, TXT), then the data needs to be exported to the corresponding format in Excel;

[0182] When exporting the data, pay attention to retaining all columns in the data (including timestamps, ecosystem types, carbon flux values), and ensure that the column names are clear and accurate.

[0183] Step 4.2: Use R language for missing data imputation

[0184] Step 4.2.1: Load and check the data

[0185] In R language, use functions such as read.csv(), read.table(), or readxl::read_excel() to load the data in the Excel file;

[0186] Use functions such as summary() and str() to check the structure and missing value situation of the data.

[0187] Step 4.2.2: Select an appropriate imputation method

[0188] Select an appropriate interpolation method according to the type of ecosystem (natural wetland, wetland restoration area after aquaculture withdrawal) and the characteristics of the data (seasonal variation, weather conditions);

[0189] Common interpolation methods include the night splitting method, the day splitting method, the average daily variation curve method, and the look-up table method. Which method to choose specifically needs to be determined according to the actual situation of the data and the characteristics of the ecosystem.

[0190] Step 4.2.3: Implement the interpolation process

[0191] Night splitting method: If the dataset contains data for both night and day, and the night data has a specific pattern (such as negative values), then the night data can be interpolated separately. For example, the average value of adjacent night data or linear interpolation can be used to fill in the missing night data;

[0192] Day splitting method: Similar to the night splitting method, but for day data. If the day data has a specific pattern (such as positive values), then corresponding methods can be used for interpolation;

[0193] Average daily variation curve method: Utilize the daily variation pattern of the data to calculate the average change value at each time point and fill in the missing data based on this change value. This method is applicable when the daily variation pattern of the data is relatively stable;

[0194] Look-up table method: If the dataset contains some known or predictable values (such as carbon flux values under specific weather conditions), then a look-up table can be established and used to fill in the missing data. This method is applicable when there are obvious patterns or regularities in the dataset;

[0195] In R language, functions or packages such as na.omit(), zoo::na.fill(), mice(), and missForest() can be used to implement the above interpolation methods. The specific implementation method needs to be determined according to the selected method and the characteristics of the data.

[0196] Step 4.2.4: Verify the interpolation results

[0197] Use a visualization tool (ggplot2 package) to compare the data before and after interpolation and check whether the interpolation results are reasonable;

[0198] Compare the data distribution and change trends before and after interpolation to ensure that the interpolation results do not introduce new outliers or change the original pattern of the data.

[0199] Step 4.2.5: Save the interpolated data

[0200] Save the interpolated data as a new file for subsequent analysis;

[0201] When saving data, pay attention to retaining all columns and column names of the data, and ensure that the file format matches the input requirements of subsequent analysis tools.

[0202] Through the above implementation process of missing data imputation, missing values in the dataset can be effectively filled, improving the integrity and availability of the data. At the same time, by selecting an appropriate imputation method according to the characteristics of the ecosystem and the actual situation of the data, the accuracy and rationality of the imputation results can be ensured.

[0203] 5. Data Visualization and Analysis

[0204] Use professional tools such as Excel and R language plotting packages to draw time-varying graphs and seasonal variation graphs of different flux components to visually display the data change trends and characteristics;

[0205] Conduct time series correlation analysis and regression analysis to explore the relationship between eddy covariance flux and meteorological factors (such as temperature, humidity, wind speed);

[0206] Calculate the light response curve parameters using daytime flux and radiation data to evaluate the ecosystem's response to light; calculate the temperature sensitivity parameters using nighttime flux and temperature data to understand the ecosystem's sensitivity to temperature changes.

[0207] Specific implementation process:

[0208] Step 5.1: Data Visualization

[0209] Step 5.1.1: Plotting with Excel

[0210] Collect data: Collect eddy covariance flux component data and corresponding time information from data sources (databases, files, API interfaces);

[0211] Data cleaning: Clean the collected data to remove duplicates, missing values, and outliers to ensure the accuracy and integrity of the data.

[0212] Insert a chart:

[0213] In Excel, select the table area containing the data;

[0214] Click on the "Insert" tab and select the chart type such as bar chart, line chart, pie chart, or scatter plot according to needs. For time-varying graphs and seasonal variation graphs, a line chart is usually selected because it can clearly show the data change trend over time;

[0215] Adjust the format settings of the axes, tick values, titles, and labels of the chart to make the chart more beautiful and easy to understand.

[0216] Customize the chart:

[0217] Set the starting angle of the sector (e.g., when creating a pie chart), and adjust the size and separation of the chart to meet the requirements of aesthetic and professional display;

[0218] Enhance the readability and attractiveness of the chart by changing colors and line styles;

[0219] Step 5.1.2: Use R language plotting packages

[0220] Install and load the necessary packages: including ggplot2 or shiny;

[0221] Data preparation: Similar to Excel, first collect, clean, and organize the data;

[0222] Plotting:

[0223] Use functions in the ggplot2 package (such as geom_line() to draw line charts and geom_bar() to draw bar charts) to create charts;

[0224] Customize the chart by setting aesthetic mappings for layers, colors, shapes, and sizes;

[0225] Use the shiny package to create interactive charts that allow users to interact with the data through sliders, dropdown menu controls, and update the chart in real time;

[0226] Export the chart: Export the created chart as an image or PDF format for easy use in reports or presentations.

[0227] Step 5.2: Time series related analysis and regression analysis

[0228] Step 5.2.1: Time series related analysis

[0229] Determine the type of time series: Based on the characteristics of the data, determine whether it is an absolute number time series, relative number time series, or average number time series;

[0230] Stationarity test: Conduct a stationarity test on the time series. If it is not stationary, then difference or other transformations are required to make it stationary;

[0231] Autocorrelation and partial autocorrelation analysis: Calculate the autocorrelation function (ACF) and partial autocorrelation function (PACF) of the time series to determine the appropriate autoregressive (AR) or moving average (MA) order;

[0232] Build a model: Based on the results of ACF and PACF, select an appropriate ARIMA model or other time series models for fitting;

[0233] Model Diagnosis and Optimization: Diagnose the model, check the normality, independence, and homoscedasticity of the residuals, adjust the model parameters as needed, and optimize the model performance.

[0234] Step 5.2.2: Regression Analysis

[0235] Determine the regression type: Select a simple regression or multiple regression model according to the characteristics of the data;

[0236] Select variables: Determine the independent variables (such as meteorological factors like temperature, humidity, and wind speed) and the dependent variable (such as eddy covariance);

[0237] Establish a regression equation: Use statistical software (such as R, SPSS) to perform regression analysis and obtain the regression equation;

[0238] Test the regression equation: Conduct an economic meaning test, a goodness-of-fit test, and a hypothesis test on the regression equation, and use the coefficient of determination R 2 to evaluate the goodness of fit of the model;

[0239] Prediction and Application: Use the regression equation for prediction, analyze the relationship between eddy covariance and meteorological factors, and provide a scientific basis for ecosystem management.

[0240] Step 5.3: Calculation of Light Response Curve Parameters and Temperature Sensitivity Parameters

[0241] Step 5.3.1: Calculation of Light Response Curve Parameters

[0242] Data Preparation: Collect daytime flux and radiation data;

[0243] Calculate parameters: Use the collected data to calculate light response curve parameters (such as light saturation point, light compensation point), which help evaluate the ecosystem's response ability to light;

[0244] Result Analysis: Analyze the performance of the ecosystem under different light conditions based on the calculated parameter values, and provide suggestions for light resource management and ecosystem optimization.

[0245] Step 5.3.2: Calculation of Temperature Sensitivity Parameters

[0246] Data Preparation: Collect nighttime flux and temperature data;

[0247] Calculate parameters: Use the collected data to calculate temperature sensitivity parameters (such as Q10 value), which can reflect the sensitivity of the ecosystem to temperature changes;

[0248] Result Analysis: Analyze the performance of the ecosystem under different temperature conditions based on the calculated parameter values, and provide a scientific basis for temperature regulation and ecosystem protection.

[0249] In summary, data visualization and analysis is a systematic and meticulous process, involving multiple aspects such as data collection, organization, visualization, and statistical analysis. Through scientific and reasonable implementation steps and professional analysis tools, the information and patterns behind the data can be deeply explored, providing strong support for ecosystem management and decision-making.

[0250] 6. Data Output and Application

[0251] Output the processed carbon flux data in the formats of Excel and CSV for subsequent data analysis and application.

[0252] Apply the processed data to scientific research in fields such as ecology and meteorology, such as evaluating the carbon sequestration capacity of forest and wetland ecosystems, revealing the carbon cycle process, and monitoring carbon emission status.

[0253] Specific implementation process:

[0254] Step 6.1: Data Output

[0255] Step 6.1.1: Implementation process:

[0256] Data preparation:

[0257] Ensure that all processed carbon flux data has been verified, including missing data imputation, outlier handling, and quality control steps.

[0258] The data should include all necessary metadata, such as timestamps, ecosystem types, geographical locations, and meteorological data.

[0259] Format selection:

[0260] Select an appropriate output format according to the requirements of subsequent analysis and application. The Excel format is suitable for cases where manual viewing and editing of data are needed, while the CSV format is more convenient for machine reading and processing.

[0261] Data export:

[0262] Use data processing software (such as Excel, R language) to export the data in the selected format. In Excel, the data table can be directly saved as an.xlsx or.csv file; in R language, the write.csv() and write.xlsx() functions can be used to export the data.

[0263] File naming and storage:

[0264] Select a clear and meaningful name for the exported file, usually including the dataset name, processing date, and format information.

[0265] Store the files in a safe and easily accessible location to ensure they can be easily found for subsequent analysis and application.

[0266] Data verification:

[0267] After exporting the data, the integrity and accuracy of the data should be verified again to ensure there are no omissions or errors.

[0268] Step 6.2: Data application

[0269] Step 6.2.1: Implementation process:

[0270] Determine the application field:

[0271] Based on the characteristics of the data and the research purpose, determine which field the data will be applied to, such as ecology, meteorology, environmental science.

[0272] Select research methods:

[0273] According to the application field and research objectives, select appropriate research methods. For example, when evaluating the carbon sequestration capacity of ecosystems such as forests and wetlands, methods such as carbon cycle models and ecosystem productivity models can be used; when revealing the carbon cycle process, techniques such as isotope labeling and remote sensing monitoring can be used; when monitoring carbon emission status, methods such as emission factor method and carbon emission inventory can be used.

[0274] Data preprocessing:

[0275] According to the requirements of the selected research method, further preprocess the data. For example, operations such as normalizing, standardizing, and smoothing the data are required to improve the accuracy and reliability of data analysis.

[0276] Model construction and analysis:

[0277] Use appropriate statistical software or programming tools (such as R language, Python) to construct an analysis model and input the processed carbon flux data for calculation and analysis;

[0278] Based on the model results, reveal the carbon cycle process, evaluate the carbon sequestration capacity, monitor the carbon emission status, and draw corresponding conclusions and recommendations.

[0279] Result presentation and reporting:

[0280] Present the analysis results in the form of charts and tables for a more intuitive understanding of the data and analysis results;

[0281] Write a research report or paper to elaborate in detail on the research methods, data sources, analysis process, results, and conclusion content for peer review and reference.

[0282] Data sharing and publication:

[0283] Share the processed carbon flux data and analysis results with peers or the public to promote the development of scientific research and environmental protection. This can be achieved through channels such as academic journals, online databases, and open data platforms for publication and sharing.

[0284] Through the above steps, the processed carbon flux data can be effectively output and applied to scientific research in fields such as ecology and meteorology, providing strong support for revealing the carbon cycle process, evaluating carbon sink capacity, and monitoring carbon emission status, etc.

[0285] III. Implementation Effects

[0286] Significant improvement in data processing efficiency and accuracy

[0287] Through the method of the present invention, a comprehensive and systematic preprocessing and postprocessing of the original carbon flux data obtained by the eddy covariance method has been successfully achieved. This process includes multiple key steps such as data cleaning, quality control, missing data imputation, outlier detection and processing, data smoothing and filtering, time synchronization and calibration, etc. Each step has been carefully designed and optimized to ensure the integrity and accuracy of the data;

[0288] After implementing this method, the efficiency of data processing has been significantly improved. A large amount of data cleaning and quality control work that originally needed to be done manually can now be quickly completed through automated scripts and algorithms. This not only greatly shortens the data processing time but also reduces the possibility of human errors, thereby improving the accuracy of the data.

[0289] Reduction of data processing costs

[0290] With the improvement of data processing efficiency, implementing the method of the present invention has also brought a significant reduction in data processing costs. The automated and semi-automated processing flow reduces the dependence on human resources and lowers the labor cost. At the same time, due to the improvement of the accuracy and efficiency of data processing, the additional costs caused by data errors or repeated processing are also reduced.

[0291] Provide high-quality data support for scientific research

[0292] The carbon flux data set processed by the method of the present invention is complete, accurate, and reliable. These data provide valuable data support for scientific research in fields such as ecology and meteorology. Scientists can use these data sets to deeply study the carbon cycle process of ecosystems, evaluate the carbon sink capacity of different ecosystems, and monitor and analyze urban carbon emission status, etc. These data not only help to reveal the internal mechanisms of ecosystems but also provide an important basis for formulating scientific and reasonable environmental protection policies.

[0293] Promote the application and development of the eddy covariance method

[0294] The method of the present invention not only improves the effect of data processing, but also further promotes the application and development of the eddy covariance method in multiple fields. In the aspect of forestry carbon sink, this method makes the eddy covariance method become one of the important means to evaluate the carbon sink capacity of forest ecosystems. In the monitoring of ecosystem carbon cycle, this method provides strong support for the real-time monitoring and analysis of the ecosystem carbon cycle process. In addition, in the field of urban carbon emission assessment, this method also provides a reliable data basis for accurately estimating urban carbon emissions and formulating emission reduction strategies.

[0295] In summary, by implementing the method of the present invention, the comprehensive and systematic preprocessing and postprocessing of the original carbon flux data obtained by the eddy covariance method have been successfully realized, significantly improving the efficiency and accuracy of data processing, reducing the data processing cost, providing high-quality data support for scientific research, and promoting the application and development of the eddy covariance method in multiple fields. These implementation effects fully demonstrate the feasibility and effectiveness of the method of the present invention.

[0296] Specific Application Example 1

[0297] Background Information

[0298] This application example aims to demonstrate how to use a carbon flux data processing method obtained by the eddy covariance method to comprehensively and systematically preprocess and postprocess actual observation data. Suppose there is a set of eddy covariance method observation data from a certain forest ecosystem, including original carbon flux data, meteorological data (such as temperature, humidity, wind speed, etc.) and time information. The following will detail the specific steps and results of data processing.

[0299] I. Data Collection

[0300] 1. Collection Equipment and Environmental Settings

[0301] Equipment: Use a Smart Flux field collector equipped with an eddy covariance sensor. This collector is built-in with a high-precision ultrasonic anemometer and an infrared gas analyzer, which can measure three-dimensional wind speed fluctuations, temperature and gas concentration fluctuations in real time;

[0302] Environmental Settings: Install the collector in an open area of the forest ecosystem to ensure that there are no tall obstacles in the observation area to reduce air flow disturbance. At the same time, ensure that the collector is at a certain height from the ground to obtain representative boundary layer air flow information.

[0303] 2. Data Recording

[0304] Sampling Frequency: Set the collector to record the original data at a high frequency (such as 10Hz), including parameters such as wind speed, temperature, humidity and gas concentration;

[0305] Time synchronization: Ensure that the built-in clock of the collector is synchronized with the standard time source for subsequent time synchronization processing of data;

[0306] Data storage: The original data is stored in real-time in the built-in memory of the collector and exported to the computer regularly for subsequent processing.

[0307] II. Specific implementation of the data processing method

[0308] 1. Data preprocessing

[0309] Step 2.1 Data cleaning

[0310] Implementation with Python script: Write a Python script to read the original data file, remove the outliers beyond the preset threshold range (such as wind speed range, temperature range, gas concentration range, etc.), and at the same time, record the removed data points and reasons to generate a data cleaning report.

[0311] Step 2.2 Quality control

[0312] Continuity check: Check the continuity of the data time series, identify and mark the missing or abnormally jumping data points;

[0313] Consistency check: Compare the data of different sensors (such as anemometer and gas analyzer) to ensure the consistency between the data;

[0314] Comparison with other observation data: Compare the processed data with the observation data of the meteorological station during the same period to further verify the reliability of the data;

[0315] Generate a quality control report: Record the problems and solutions found during the quality control process to generate a detailed quality control report.

[0316] Step 2.3 Imputation of missing data

[0317] Linear interpolation method: For the data points with a short missing time (such as less than 2 hours), use the linear interpolation method for imputation;

[0318] Nearest neighbor average method: For the data points with a long missing time or poor linear interpolation effect, use the nearest neighbor average method for imputation. At the same time, consider the time series characteristics and daily variation rules of the data to improve the accuracy of imputation.

[0319] 2. Data postprocessing

[0320] Step 3.1 Data smoothing and filtering

[0321] Moving average filter: Apply a moving average filter to smooth the data and select an appropriate window size to balance the smoothing effect and detail retention;

[0322] Savitzky-Golay filter: For data that requires higher smoothing accuracy, the Savitzky-Golay filter is used for processing. This filter can retain the shape characteristics of the data while smoothing the data.

[0323] Step 3.2 Time Synchronization and Calibration

[0324] Time synchronization: All observed data are processed for time synchronization according to timestamps to ensure that all data are consistent in time.

[0325] Calibration: Standard meteorological parameters (such as temperature, humidity, air pressure, etc.) are used to calibrate the data to eliminate the influence of instrument errors and environmental disturbances on the data.

[0326] Step 3.3 Flux Component Calculation

[0327] NEE calculation: The net ecosystem exchange (NEE), that is, the net carbon exchange between the ecosystem and the atmosphere, is calculated using the principle of the eddy covariance method.

[0328] Re and GPP calculations: The ecosystem respiration (Re) and gross primary productivity (GPP) are calculated by combining the ecosystem energy balance equation and NEE data. Specific methods include the night respiration model method, light response curve fitting method, etc.

[0329] 3. Data Output and Application

[0330] Step 4.1 Data Output

[0331] The processed data are saved as Excel or CSV format files for subsequent analysis and application. At the same time, detailed data description documents and data quality assessment reports are provided.

[0332] Step 4.2 Data Application

[0333] The processed data are applied to fields such as ecosystem carbon cycle research and forestry carbon sink assessment. Key information such as the carbon budget balance status, carbon sink function and its influencing factors of the ecosystem are revealed through data analysis.

[0334] Based on the research results, scientific research papers are published or relevant environmental protection policies are formulated to provide a scientific basis for ecosystem management and climate change response.

[0335] 5. Specific Numerical Values and Output Result Table

[0336] The following is a simplified table showing some processed carbon flux data (taking the NET ecosystem exchange NEE as an example) and the corresponding meteorological data:

[0337] Timestamp NEE (μmolm-1) Temperature (°C) Humidity (%) Wind speed (m / s) 2023-04-01T00:00 -2.5 15.0 80.0 1.2 2023-04-01T01:00 -2.3 14.8 78.5 1.1 ... ... ... ... ... 2023-04-30T22:00 -1.8 16.5 79.0 1.3 2023-04-30T23:00 -1.9 16.3 78.8 1.2

[0338] Note: The data in the above table are only examples. The actual data will include more time points and more detailed observational data. In addition, the values in the table have been processed through preprocessing and postprocessing steps, including data cleaning, quality control, missing data imputation, data smoothing and filtering, time synchronization and calibration.

[0339] Through this application example, it has been successfully demonstrated how to comprehensively and systematically preprocess and postprocess actual observational data using a carbon flux data processing method obtained by the eddy covariance method, and a complete, accurate, and reliable carbon flux data set has been obtained. These data sets provide strong support for subsequent scientific research and decision-making support.

[0340] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of the present invention, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.

Claims

1. A method for processing carbon flux data obtained by eddy covariance method, characterized in that: The following steps are involved: a) Data collection: Using eddy covariance method equipment to capture signals to obtain raw carbon flux data, the data including timestamp, air temperature, humidity, wind speed, wind direction and carbon flux information; b) Eddypro software preprocessing: Import the raw data into Eddypro software for preprocessing, including: Remove the data with flag=2; Remove outliers and "wild points"; Remove the data half an hour before and after the rainfall; Data after removing the signal relative to the ultrasound machine time band; Perform delay time correction and coordinate axis rotation; Export 30 minutes of average data as a ".csv" format file; c)TOVI software post-processing: The data processed by Eddypro software is imported into TOVI software for further processing, including: Check the integrity of the data and remove obviously abnormal data based on the visual interface; Fill in the missing time points in the ".csv" file biomet and full files to ensure that the time series of the two files meet the requirements; Record data anomalies or missing data, including data source, data type, time period, reasons for exclusion, and other explanations; Perform meteorological data interpolation and export processed data to Excel; d) Excel processing: The data processed by TOVI software was pre-processed in Excel to meet the requirements of R language interpolation, including: Conduct flux data quality control based on different ecosystem types such as natural wetlands, de-aquaculture restoration areas, and aquaculture ponds, and remove unreasonable data points; Check and quality control the indicators required for R language interpolation; Create R language input files, sort the indicators, and fill in the missing data; e) R language interpolation: Use R language to interpolate missing data after Excel processing to obtain complete annual data, including: Selection of interpolation method based on ecosystem type, including night-split or day-split methods; After the imputation, the data status was checked, including NEE, GPP, and Re. If there were obvious abnormalities, the data were re-interpolated; Confirm the validity of the data after interpolation.

2. The method for processing carbon flux data obtained by eddy covariance method according to claim 1, characterized in that: Specific requirements for recording data anomalies or missing data in step c) include: The data source must clearly identify the origin or ecosystem type of the data; The data type needs to specify whether it is air temperature, humidity, wind speed, wind direction or carbon flux; The time period must cover at least 5 consecutive days of data recording in order to analyze long-term trends and abnormal patterns in the data; The reasons for exclusion should include detailed records of the specific issues that caused the data to be excluded, including sensor failure or abnormal data fluctuations; The Other Notes section can be used to record any additional information related to data anomalies or omissions, including weather conditions and equipment maintenance records.

3. The method for processing carbon flux data obtained by eddy covariance method according to claim 1, characterized in that: The specific operations of flux data quality control in step d) include: For natural wetland ecosystems, during the growing season from May to October, negative carbon fluxes at night and obviously unreasonable positive carbon fluxes during the day were eliminated; For the reforestation and restoration area, only the negative carbon flux at night was eliminated; For areas affected by human activities such as aquaculture ponds, both night and day data were retained and not eliminated; During the processing, data points that deviate significantly from the overall trend are eliminated based on the data change trend throughout the year to ensure the accuracy and representativeness of the data.

4. The method for processing carbon flux data obtained by eddy covariance method according to claim 1, characterized in that: When using R language to interpolate missing data in step e), the specific interpolation method is selected based on the following: For natural wetlands and restoration areas, the night split method is used for interpolation; For areas that are heavily affected by human activities, such as aquaculture ponds, the day split method was used for interpolation; During the interpolation process, seasonal changes in data and weather conditions are taken into account to ensure the accuracy and rationality of the interpolation results; After the interpolation is completed, the data is checked in detail. If there are obvious abnormalities, the interpolation process needs to be repeated.

5. A method for processing carbon flux data obtained by eddy covariance method according to any one of claims 1 to 4, characterized in that: It also includes data visualization and analysis steps, including: Use charts and images to present the steps and results of data processing, including data integrity checks, outlier removal effects, and interpolation results; Through visual analysis, users can better understand data characteristics and changing trends to facilitate quality control and result verification; Visualization tools may include Excel, R language drawing packages, and professional data visualization software.

6. The method for processing carbon flux data obtained by eddy covariance method according to claim 5, characterized in that: The eddy covariance method equipment includes an ultrasonic anemometer, an infrared gas analyzer, and data acquisition and transmission equipment, and the specific configuration is as follows: Ultrasonic anemometer is used to measure three-dimensional wind speed and direction in real time to ensure the accuracy and real-time nature of the data; Infrared gas analyzers are used to measure changes in the concentration of gases such as carbon dioxide to calculate carbon flux; The data acquisition and transmission equipment is responsible for transmitting the measurement data to the data processing system in real time for subsequent analysis and processing; The equipment also needs to be calibrated and maintained regularly to ensure the accuracy and stability of the measurement data.

Citation Information

Cited By

  • Reflected radiation data interpolation method, system and application

    CN120974084A

  • A method, system and application for interpolating reflected radiance data

    CN120974084B