A method and device for estimating the total amount of near-infrared co

By establishing an ERT machine learning model and combining ground and satellite data, the problem of insufficient accuracy in existing CO2 total column estimation was solved, achieving high-precision CO2 total column calculation and a simplified calculation process.

CN115759280BActive Publication Date: 2026-01-06AEROSPACE INFORMATION RES INST CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211311113.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2026-01-06
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

Existing methods and equipment for estimating total CO2 column volume are not accurate enough and cannot effectively meet the requirements for high-precision total CO2 column volume measurement.

Method used

An ERT model was established using machine learning methods. A training dataset was constructed using ground monitoring data, satellite data, and meteorological data. The total CO column was estimated using the ERT machine learning model. Input parameters included time, longitude, latitude, meteorological data, and satellite observation geometric parameters. The model training was optimized to improve accuracy.

Benefits of technology

It achieves high-precision calculation of total CO2 column volume, improves the accuracy of total CO2 column volume estimation, simplifies the calculation process, and facilitates engineering implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115759280B_ABST
    Figure CN115759280B_ABST
Patent Text Reader

Abstract

The application discloses a near-infrared CO column total amount estimation method and equipment, and solves the problem of poor CO column total amount measurement precision. The method comprises the following steps: establishing a machine learning database, wherein the machine learning database comprises CO column total amount monitored on the ground, monitoring time, longitude, latitude, meteorological data corresponding to a ground monitoring point, and digital elevation model data; establishing an extreme random tree machine learning model, constructing a training data set for the data in the machine learning database, setting model parameters for model training, and terminating training until the training condition is met to obtain a CO estimation model; input parameters of the extreme random tree machine learning model comprise the time, longitude, latitude, meteorological data and digital elevation model data, and the output parameter is the CO column total amount. The equipment uses the method. The application can realize high-precision prediction of the CO column total amount.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of atmospheric environment remote sensing technology, and in particular to a method and equipment for estimating the total amount of near-infrared CO column. Background Technology

[0002] Currently, satellite sensors capable of detecting CO include AIRS (Atmospheric Infrared Sounder), IASI (Infrared Atmospheric Sounding Interferometer), TES (Tropospheric Emission Spectrometer), Tropomi (TROPOspherical Monitoring Instrument), and SCIAMACHY (Atmospheric Mapping Scanning Imaging Absorption Spectrometer). Near-infrared spectroscopy is highly sensitive to the integral of CO (carbon monoxide) along the optical path, making this spectral range particularly suitable for CO detection. Current methods for retrieving total CO column concentrations mainly include weighted function modified-differential absorption spectroscopy (WFM-DOAS) algorithms, maximum likelihood estimation algorithms, and maximum a posteriori methods—all fully physical inversion methods. However, all these algorithms require a robust atmospheric background database and their CO column concentration estimates are not always accurate. Summary of the Invention

[0003] This invention provides a method and device for estimating the total CO column volume in near-infrared radiation, which solves the problem of poor measurement accuracy of the total CO column volume in existing methods and devices.

[0004] To solve the above problems, the present invention is implemented as follows:

[0005] This invention provides a method for estimating the total CO2 column in near-infrared radiation, comprising the following steps: establishing a machine learning database, which includes: the total CO2 column volume, time, longitude, latitude, meteorological data corresponding to the ground monitoring stations, and digital elevation model (DEM) data. Establishing an ERT machine learning model, constructing a training dataset from the data in the machine learning database, setting model parameters, and training the model until the model accuracy is optimal, then terminating the training to obtain the CO2 estimation model. The input parameters of the ERT machine learning model include: the time, longitude, latitude, meteorological data corresponding to the ground monitoring stations, and DEM data; the output parameter is the total CO2 column volume monitored on the ground.

[0006] Preferably, the machine learning database further includes: near-infrared satellite incident and emission spectral data and satellite observation geometric parameters; when training the ERT machine learning model, the input parameters also include the near-infrared satellite incident and emission spectral data and satellite observation geometric parameters.

[0007] Preferably, the machine learning database further includes at least one of aerosol optical thickness, band surface reflectance, and normalized vegetation index; correspondingly, when training the ERT machine learning model, the input parameters further include at least one of aerosol optical thickness, band surface reflectance, and normalized vegetation index.

[0008] Furthermore, the meteorological data corresponding to the ground monitoring stations includes: 2-meter temperature, eastward component of 10-meter wind, northerly component of 10-meter wind, relative humidity, precipitation, total water vapor column, boundary layer height, cloud base height, and cloud coverage.

[0009] Furthermore, when setting model parameters for model training, the set model parameters include: the number of decision trees (n_estimators), the maximum depth of the decision trees (max_depth), the minimum number of split samples (min_samples_split), and the minimum number of leaves (min_samples_leaf).

[0010] Furthermore, the optimal model accuracy metric is: the loss of the training dataset is minimized and the loss of the validation dataset is minimized, wherein the validation dataset is a dataset randomly selected from the machine learning database.

[0011] Furthermore, the method for constructing a training dataset from the data in the machine learning database is as follows: a predetermined proportion of the data in the machine learning database is selected as the training dataset, and the remaining data is used as the validation dataset.

[0012] Furthermore, the near-infrared satellite incident and emission spectral data include the radiance of incident solar radiation and emitted Earth radiation in a specific selected band, and the satellite observation geometric parameters include solar zenith angle, solar azimuth angle, satellite zenith angle, and satellite azimuth angle.

[0013] Furthermore, the method also includes processing the satellite observation geometric parameters, including processing the solar zenith angle and the satellite zenith angle into secant values, and processing the angle difference between the solar azimuth angle and the satellite azimuth angle into a sine value.

[0014] Furthermore, the method is characterized by comprising: performing spatiotemporal matching and data preprocessing on the data in the machine learning database, and removing outlier and invalid data exceeding three standard deviations.

[0015] This invention also proposes a near-infrared CO column total estimation device, using the method described in any embodiment, comprising: a data processing module and a model determination module; the data processing module is used to acquire the machine learning database and establish a training dataset; the model determination module is used to train the ERT machine learning model according to the training dataset until the training conditions are met to obtain a CO estimation model.

[0016] This application also proposes a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the methods described in any embodiment of this application.

[0017] Furthermore, this application also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in any embodiment of this application.

[0018] The at least one technical solution adopted in the embodiments of this application can achieve the following beneficial effects: This application can train a high-precision CO vertical column concentration inversion model based on near-infrared satellite spectral data, meteorological data, surface parameters, and aerosol data using machine learning. The total CO column concentration calculated by the method of this application has high accuracy, simple calculation method, strong operability, and is easy to implement in engineering. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0020] Figure 1 This is a flowchart illustrating an embodiment of the method of the present invention.

[0021] Figure 2 The average CO column concentration in the embodiments of the method of the present invention;

[0022] Figure 3 This is an embodiment of the device of the present invention;

[0023] Figure 4 This is another embodiment of the apparatus of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0025] CO is an important trace gas in the atmosphere and a major air pollutant in some urban areas. Understanding the total CO column can improve our understanding of tropospheric chemistry and long-range atmospheric transport. The main sources of CO are fossil fuel combustion, biomass combustion, and the atmospheric oxidation of methane and other hydrocarbons. In the mid-latitude northern regions, fossil fuel combustion is the primary source of CO, while in the tropics, the oxidation of isopentenylene and biomass combustion play important roles. The most important sink for CO is its reaction with hydroxyl radicals (OH). The lifetime of CO ranges from several weeks to several months, making it an excellent tracer for studying long-range transport processes. Furthermore, CO is one of the highest priority gases measured by the Global Environment and Security Monitoring Programme (GMES) Atmospheric Composition and Climate Monitoring Project (MACC).

[0026] As research into atmospheric composition deepens, studies on CO have also progressed. CO is colorless, odorless, and tasteless, and its carbon content in the atmosphere is second only to CO2 and CH4. It is mainly produced by the combustion of fossil fuels and biomass, and its concentration in the atmosphere is very low. However, as a highly reactive gas, CO can significantly affect the oxidation properties of the near-surface atmosphere, thus impacting the composition of the near-surface atmosphere, especially greenhouse gases such as CO2, CH4, and O3. Therefore, CO is also known as an indirect greenhouse gas. Furthermore, CO participates in near-surface photochemical reactions, serving as a precursor to near-surface O3 production, and is a cause of photochemical smog pollution. Besides its impact on global climate and environment, as a major near-surface pollutant, CO can affect human health. Inhaling high concentrations of CO can induce symptoms such as memory loss, myocardial infarction, coronary heart disease, and asphyxiation due to oxygen deficiency. Because of CO's significant role in global climate and human health, research on atmospheric CO has become extremely important. Since the Industrial Revolution, with global warming, the warming effect of greenhouse gases has gradually become apparent. As a significant reactive gas in the atmosphere, CO concentration changes are closely related to global industrial development and greenhouse gas emissions, leading to climate effects and environmental pollution that have gradually attracted widespread attention. With in-depth research, understanding the chemical reaction mechanisms of CO in the atmosphere; analyzing its source and sink characteristics and their impact on global climate change; and understanding the spatiotemporal distribution characteristics of CO to analyze its impact on the environment and human health as a pollutant are of great importance.

[0027] Within the 2.3 μm spectral range of the solar shortwave infrared portion, the total CO column, sensitive to the tropospheric boundary layer, can be derived from sunlight reflected from Earth's atmosphere. The CO absorption band lies between 2305 nm and 2385 nm. For clear-sky measurements, this spectral range is minimally affected by atmospheric scattering. Therefore, the near-infrared region is highly sensitive to the CO integral along the optical path, making this spectral range particularly suitable for CO detection.

[0028] Currently, methods for inverting total CO column concentrations based on spaceborne near-infrared spectral data mainly include the weighted function modified-differential absorption spectroscopy (WFM-DOAS) algorithm, the maximum likelihood estimation algorithm, and the maximum a posteriori method. Buchwitz et al. used the weighted function modified WFM-DOAS algorithm to invert tropospheric CO in the SCIAMACHY near-infrared spectral region, and the inversion results were consistent with data products from the MOPITT (Tropospheric Pollution Measurement Instrument) on the Terra satellite. The increase in atmospheric CO concentration caused by biomass burning events observed by MOPITT could also be monitored by SCIAMACHY. The linear correlation coefficient between the CO column concentrations inverted by SCIAMACHY and MOPITT was between 0.4 and 0.7. Therefore, the accuracy of existing techniques for estimating total CO column concentrations is relatively low.

[0029] Extremely Random Trees (ERT) is an important ensemble learning method based on Bagging, used to handle classification and regression problems. Random Forests generate new training sets by repeatedly and randomly sampling k samples with replacement from the original training sample set N. Then, k classification trees are generated from the bootstrap sample set to form a random forest. The classification result of the new data is determined by the voting results of the classification trees. ERT is also a classifier that integrates multiple decision trees. Compared with Random Forest, it differs in two main ways: First, ERT generally does not use random sampling; each decision tree uses the original training set. Second, the features of the end-random trees are randomly selected. Because the splitting is random, it can sometimes yield better results than Random Forest. The variation in total CO2 column emissions is influenced by numerous factors (biomass combustion emissions, meteorological factors, aerosols, atmospheric oxidation, etc.). Currently, no research has proven a linear or nonlinear relationship, or a polynomial relationship, between these factors and total CO2 column emissions. Therefore, machine learning methods are used to learn and model complex nonlinear relationships, constructing relationships between different components to establish a model for predicting total CO2 column emissions.

[0030] The innovations of this invention are as follows: First, this application creatively applies a machine learning model to near-infrared CO column concentration estimation, achieving higher estimation accuracy compared to traditional linear or nonlinear models. Second, during model training, this invention fully considers the influence of atmospheric and surface parameters involved in the scattering and reflection processes of incident sunlight between the atmosphere and the land surface. Satellite data, AOD data, surface reflectance data, and meteorological parameters are incorporated into the model input parameters, further optimizing the model and improving the accuracy of CO column total estimation.

[0031] The technical solutions provided by the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0032] Figure 1 This is a flowchart of a method embodiment of the present invention. The total CO column volume is estimated using a machine learning model. As an embodiment of the present invention, the near-infrared CO column volume estimation method includes the following steps 101-102:

[0033] Step 101: Establish a machine learning database.

[0034] In step 101, the machine learning database includes: the total CO column data, time, longitude, latitude, meteorological data that is spatiotemporally matched with the ground monitoring data (meteorological data corresponding to the ground monitoring stations), and digital elevation model data.

[0035] In step 101, the data in the machine learning database further includes: near-infrared satellite incident and exit spectral data, and satellite observation geometric parameters. Further, it may also include at least one of the following parameters: aerosol optical thickness, band surface reflectance, and normalized difference vegetation index. It should be noted that the band surface reflectance refers to the surface reflectance corresponding to the near-infrared band.

[0036] Preferably, the meteorological data that is spatiotemporally matched with the ground monitoring data includes: 2-meter temperature, eastward component of 10-meter wind, northerly component of 10-meter wind, relative humidity, precipitation, total water vapor column, boundary layer height, cloud base height, and cloud coverage.

[0037] Preferably, the near-infrared satellite incident and emission spectral data include: the radiance (incident radiation intensity) of incident solar radiation in a specific selected band and the radiance (emission radiation intensity) of emitted Earth radiation in a specific selected band. The satellite observation geometric parameters include: satellite zenith angle, satellite azimuth angle, solar zenith angle, and solar azimuth angle. The solar zenith angle and satellite zenith angle are processed into secant values, and the angle difference between the solar azimuth angle and satellite azimuth angle is processed into a sine value.

[0038] It should be noted that atmospheric CO column concentration is affected by factors such as temperature, water vapor, wind, clouds, solar radiation, atmospheric oxidation and biomass combustion emissions, and land cover type. This concentration is also reflected in the near-infrared spectrum, which is further influenced by surface reflectance and atmospheric aerosol optical thickness. The total CO column data selected in this embodiment is from TCCON (Total Carbon Column Observing Network) measurements, which include CO column measurements from near-infrared Fourier spectrometers distributed across various global sites. Other data include meteorological data from ECMWF's fifth-generation meteorological reanalysis product ERA5; near-infrared TROPOMI satellite data (including geographic information, incident and emitted radiation intensities, and observational geometric parameters); digital elevation model (DEM) data; and MODIS data on 550 nm aerosol optical thickness and 2.3 μm surface reflectance.

[0039] The data selection summary of the machine learning database is shown in Table 1 below. It should be noted that the machine learning database in this embodiment is not limited to data obtained from the sensors listed in Table 1, but can also be data obtained from other sensors; no particular limitation is made here.

[0040] Table 1 Machine Learning Databases

[0041]

[0042]

[0043]

[0044]

[0045]

[0046]

[0047] A machine learning database is established by matching and fusing CO data with other data, such as meteorological data and digital elevation model data.

[0048] It should be noted that the Radiance data selected by the TROPOMI sensor is mainly the Radiance and Iradiance of the absorption channel near CO 2.3 micrometers.

[0049] In step 101, the method further includes: performing data matching and data preprocessing on the data in the machine learning database, and removing abnormal data and invalid data. Abnormal data refers to data exceeding three standard deviations, or abnormal data identified by other methods, which are not specifically limited here.

[0050] Based on the characteristics of the above data, all data are matched and fused, and then the data is filtered to remove invalid values ​​and normalized. Specifically, the selected ERA5 data, including reflectance, rainfall intensity, surface temperature, and cloud cover, are all one-dimensional data; the total CO column in TCCON is one-dimensional data; the selected TROPOMI incident and emitted radiation intensity data are CO data near the absorption wavelength of 2.3 micrometers, because data in these bands are sensitive to the total CO; the 550nm aerosol optical thickness (AOD) data, normalized difference vegetation index (NDVI) data, and surface reflectance data provided by MODIS are one-dimensional data; and the digital elevation model (DEM) data is one-dimensional data.

[0051] During data matching, the latitude and longitude of TCCON's global stations were used as the spatial reference, and the satellite transit time was used as the time standard. ERA5 data, TROPOMI data, and MODIS data were matched according to latitude and longitude and satellite observation time. Spatially, data within 7 kilometers of the TCCON station were selected for matching, and temporally, data within 1 hour before and after the satellite transit were selected. Then, outliers exceeding three standard deviations were removed from these data, and the mean was taken.

[0052] When preprocessing TCCON data, invalid values ​​(i.e., null values ​​for CO content) are removed, and outliers exceeding three standard deviations are eliminated; then the data are normalized.

[0053] In step 101, the method for constructing a training dataset from the machine learning database is as follows: A predetermined proportion of data from the machine learning database is randomly selected as the training dataset, and the remaining data is used as the validation dataset. For example, the predetermined proportion of the training dataset is less than or equal to 80%, and correspondingly, the predetermined proportion of the validation dataset is greater than or equal to 20%. It should be noted that the predetermined proportion can also be other values, and no particular limitation is made here.

[0054] Step 102: Establish an ERT machine learning model. Construct a training dataset from the data in the machine learning database, set the model parameters, and train the model until the training conditions are met. Then, terminate the training to obtain the CO estimation model.

[0055] In step 102, the input parameters of the ERT machine learning model include: the year, month, day, longitude, latitude, meteorological data and digital elevation model data that are matched with the ground observation time and space, and the output parameter is the total CO column of the ground monitoring.

[0056] Preferably, the input parameters of the ERT machine learning model also include near-infrared satellite incident radiation intensity and outgoing radiation intensity. Further, it may also include at least one of the following parameters: observation geometric parameters, aerosol optical thickness data, surface reflectance, and normalized vegetation index.

[0057] In step 102, based on the characteristics of the data in step 101, the ERT machine learning model is selected to train the data in the training dataset to obtain the CO estimation model.

[0058] Here, we choose the ERT model. Extreme Random Trees (ERT) is an ensemble learning machine learning model. It's also a classifier that integrates multiple decision trees. Compared to Random Forest classifiers, it differs in two main ways: First, Extreme Random Trees generally do not use random sampling; each decision tree uses the original training set. Second, the features of the end-random trees are randomly selected. Because the splits are random, it can sometimes produce better results than Random Forest.

[0059] The data characteristics of this invention perfectly match the characteristics of the events processed by ERT. The training input parameters include: time (year, month, day), longitude, latitude, 2-meter temperature, eastward component of 10-meter wind, northerly component of 10-meter wind, relative humidity, precipitation, total water vapor column, boundary layer height, cloud base height and cloud cover, aerosol optical thickness, surface reflectivity, normalized difference vegetation index, incident and emitted radiation intensities at 2.3 micrometers using satellite observations in 20 bands, satellite zenith secant, solar azimuth secant, and the tangent of the difference between the satellite azimuth and the solar azimuth. The training output parameter is the total CO2 column.

[0060] When setting model parameters for training, the set model parameters include: the number of iterations being greater than or equal to a first threshold and the training batch size being greater than or equal to a second threshold. The training condition is that the loss of the training dataset is minimized and the loss of the validation dataset is minimized. It should be noted that the first and second thresholds are set values; for example, the first threshold is 600 and the second threshold is 72.

[0061] For example, 80% of the data is randomly selected as the training dataset, and the remaining 20% ​​is used as the validation dataset to supervise the training results. The parameters set during model training are: 600 iterations, 72 training batches, and training terminates when the loss of the training and validation data reaches its minimum, thus obtaining the CO estimation model.

[0062] Figure 2 The average CO column concentration in the method embodiment of the present invention is... Figure 2 In this study, an ETR inversion model was developed based on 2020 data. A total of 4165 data records from 25 sites were used, with 80% randomly selected as the training set and 20% as the test set (833 records in total). The ETR inversion results showed an adjusted correlation coefficient of 0.8560 squared and a root mean square error of 4.8690 ppbv, indicating that the average CO concentration in the column could be accurately estimated.

[0063] The CO2 total column calculation method proposed in this invention constructs a training dataset using historical data and employs machine learning methods to improve the accuracy of CO2 total column calculation. Furthermore, the input parameters can include near-infrared satellite spectral data and satellite observation geometric parameters, aerosol optical thickness, surface reflectance, normalized vegetation index, and meteorological data, which can further improve the calculation accuracy.

[0064] Figure 3 As an embodiment of the device of the present invention, using any embodiment of the method of the present invention, as an embodiment of the device of the present invention, a near-infrared CO column total estimation device 300 includes: a data processing module 301 and a model determination module 302.

[0065] The data processing module is used to acquire the machine learning database and establish a training dataset.

[0066] The model determination module is used to train the ERT machine learning model based on the training dataset until the training conditions are met, thereby obtaining the CO estimation model.

[0067] It should be noted that the specific methods for implementing the functions of the data processing module and the model determination module are as described in the steps of the method embodiments of this application, and will not be repeated here.

[0068] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0069] Therefore, this application also proposes a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the methods described in any embodiment of this application.

[0070] Furthermore, this application also proposes an electronic device (or computing device) including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in any embodiment of this application.

[0071] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0072] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0073] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0074] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, a network interface, and memory. Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0075] Specifically, such as Figure 4 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application. The shown electronic device 400 is merely an example and should not be construed as limiting the functionality or scope of use of the embodiments of this application. It includes: one or more processors 420; and a storage device 410 for storing one or more programs. When the one or more programs are executed by the one or more processors 420, the one or more processors implement the near-infrared CO column total estimation method provided in any embodiment of this application. For example, the method includes steps 101-102 of the embodiments of this application.

[0076] The electronic device also includes an input device 430 and an output device 440; the processor, storage device, input device and output device in the electronic device can be connected by a bus or other means, as shown in the figure, which is connected by a bus 450.

[0077] Storage device 410, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and module units, such as the program instructions corresponding to the cloud bottom height determination method in the embodiments of this application. The storage device may mainly include a program storage area and a data storage area, wherein the program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on terminal usage, etc. Furthermore, the storage device may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the storage device may further include memory remotely located relative to processor 420, and these remote memories can be connected via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0078] Input device 430 can be used to receive input digital, character, or voice information, and to generate key signal inputs related to user settings and function control of the electronic device. Output device 440 may include electronic devices such as a display screen and a speaker.

[0079] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0080] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A method for estimating the total amount of CO in the near infrared, characterized by, The method comprises the following steps: a machine learning database is established, which comprises: CO column amount monitored on the ground, time, longitude, latitude, meteorological data corresponding to the ground monitoring station, digital elevation model data, near-infrared satellite incident and outgoing spectral data, and satellite observation geometric parameters; an ERT machine learning model is established, training data set is constructed from the data in the machine learning database, model parameters are set for model training, until the model precision index is optimal, the training is terminated to obtain a CO estimation model, the input parameters of the ERT machine learning model comprise: the time, longitude, latitude, meteorological data corresponding to the ground monitoring station, digital elevation model data, near-infrared satellite incident and outgoing spectral data, and satellite observation geometric parameters, and the output parameter is the CO column amount monitored on the ground; The machine learning database further comprises at least one of aerosol optical depth, band surface reflectivity, and normalized vegetation index. Accordingly, when the ERT machine learning model is trained, the input parameters further comprise at least one of the aerosol optical depth, the band surface reflectivity, and the normalized vegetation index.

2. The method of claim 1, wherein the near infrared CO column total is estimated by: The meteorological data corresponding to the ground monitoring station comprises: 2-meter temperature, 10-meter wind east component, 10-meter wind north component, relative humidity, rainfall, water vapor column amount, boundary layer height, cloud base height, and cloud coverage.

3. The near-infrared CO column total volume estimation method as described in claim 1, characterized in that, When the model parameters are set for model training, the set model parameters comprise: the number of decision trees, the maximum depth of the decision tree, the minimum split sample, and the minimum number of leaves.

4. The near-infrared CO column total volume estimation method as described in claim 1, characterized in that, The optimal model precision index is that the loss of the training data set reaches the minimum and the loss of the validation data set reaches the minimum, and the validation data set is a randomly selected data set in the machine learning database.

5. The near-infrared CO column total volume estimation method as described in claim 1, characterized in that, The method for constructing the training data set from the data in the machine learning database is to set a proportion of data in the machine learning database as the training data set and the remaining data as the validation data set.

6. The method of claim 1, wherein the near infrared CO column total is estimated by: The near-infrared satellite incident and outgoing spectral data include the radiance of incident solar radiation and outgoing earth radiation of a specific selected band, and the satellite observation geometric parameters include the solar zenith angle, the solar azimuth angle, the satellite zenith angle, and the satellite azimuth angle.

7. The near-infrared CO column total volume estimation method as described in claim 6, characterized in that, The method further comprises processing the satellite observation geometric parameters, including processing the solar zenith angle and the satellite zenith angle into secant values and processing the angle difference between the solar azimuth angle and the satellite azimuth angle into sine values.

8. The method of claim 1 to 7, wherein, The method further comprises performing data spatiotemporal matching and data preprocessing on the data in the machine learning database to eliminate abnormal data and invalid data exceeding three times the standard deviation.

9. A near infrared CO column total amount estimation apparatus using the method according to any one of claims 1 to 8, characterized by The method comprises: a data processing module and a model determination module; The data processing module is configured to obtain the machine learning database and establish a training data set; The model determination module is configured to train an ERT machine learning model according to the training data set until a training condition is met to obtain a CO estimation model.

Citation Information

Patent Citations

  • PM2.5 concentration estimation method based on stationary orbit satellite

    CN110595968A

  • Atmospheric carbon dioxide column concentration high coverage reconstruction method

    CN114974453A