Double-branch aerosol characteristic parameter estimation method based on convolutional neural network
By using the DAeroNet model, combined with CNN and FCNN, images and one-dimensional data are processed, and pixels of clouds, ice, and snow are removed. This solves the uncertainty problem in aerosol remote sensing inversion, achieves high-precision estimation of aerosol characteristic parameters, and improves monitoring efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2026-03-13
AI Technical Summary
Existing aerosol remote sensing inversion methods have uncertainties in cloud pixel interference and surface reflectance estimation, resulting in low accuracy of aerosol characteristic parameter inversion and making it difficult to achieve efficient and accurate large-scale monitoring.
A dual-branch model based on convolutional neural networks (DAeroNet) is adopted, combining convolutional neural networks (CNN) and fully connected neural networks (FCNN). Through the feature extraction module (FEM) and channel attention module (CAM), images and one-dimensional data are processed. The K-means algorithm and multi-band thresholding are used to remove cloud, ice and snow pixels, thereby improving the estimation accuracy and robustness.
It achieves high-precision estimation of aerosol characteristic parameters, improves the efficiency and accuracy of aerosol monitoring, and has better robustness and generalization ability.
Smart Images

Figure CN121659697A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a dual-branch aerosol characteristic parameter estimation method (DAeroNet) based on a convolutional neural network. Background Technology
[0002] Atmospheric aerosols refer to solid, liquid, and solid-liquid mixture particles with aerodynamic diameters between 0.001 and 100 micrometers suspended in the atmosphere. They are generated by human and natural activities and are mainly distributed in the stratosphere and troposphere, forming an important component of the Earth's atmosphere. [1–3] Aerosols can absorb and scatter both short-wave and long-wave solar radiation, directly affecting the Earth's radiation balance. This reduces surface temperature by decreasing the amount of solar radiation reaching the Earth; this is known as the direct effect of aerosols. [4] Aerosols absorb large amounts of solar radiation, which warms the troposphere, affects relative humidity and stability, and reduces cloud formation and lifespan; this is known as the semi-direct effect of aerosols. [5] Aerosols can also act as cloud condensation nuclei, altering the microscopic physical properties of clouds such as absorption and scattering coefficients, liquid water content, and cloud droplet particle distribution, thereby affecting the spatiotemporal distribution of solar radiation energy. This is known as the indirect effect of aerosols. [6] Air pollution is a serious environmental problem that humanity has faced since the beginning of the 21st century. [7] Aerosol particles are a major pollutant affecting air quality, and high concentrations of aerosol particles can seriously impact people's health, lives, and daily activities. [8,9] With the continuous development of human society, industrial pollution, urban traffic, biomass burning, and other anthropogenic factors have released large amounts of aerosols, while natural factors such as soil dust, forest fires, and volcanic eruptions also produce aerosols. The sources of atmospheric aerosols are mainly divided into natural and anthropogenic sources, including dust aerosols, black carbon aerosols, sulfate aerosols, sea salt aerosols, and organic carbon aerosols. Dust aerosols, as the largest component, have a particularly significant impact on the environment and climate.
[10] The harmfulness of aerosol particles varies depending on their size. Coarse particles with a diameter of 10-100 micrometers, such as dust aerosols during sandstorms, not only affect atmospheric visibility but also harm the environment, transportation, and human health. Long-term exposure to dust pollution can easily lead to skin, respiratory, and cardiovascular diseases, and in severe cases, even death. Particulate matter with a diameter of 10 micrometers or less (PM10) mainly comes from various industrial pollution emissions. Once inhaled, it can cause various cardiopulmonary diseases. In recent years, PM10 has become the primary pollutant in large and medium-sized cities in my country. Particulate matter with a diameter of 2.5 micrometers or less (PM2.5) can remain suspended in the air for a relatively long time. The smaller the particle size, the deeper it can penetrate into the human respiratory tract, affecting the cardiovascular, cerebrovascular, and respiratory systems.
[11] Aerosols affect the atmospheric environment, radiation balance, and land-atmosphere system of the entire ecosystem, and are closely related to human production and life. Related research and monitoring have received widespread attention from many scholars at home and abroad.
[0003] Atmospheric aerosol monitoring mainly includes aerosol concentration and distribution, physical properties, optical properties, and evolution processes.
[12] Remote sensing monitoring of aerosols can be categorized into ground-based remote sensing and satellite remote sensing, depending on the monitoring platform on which the instrument is located. Ground-based remote sensing measures direct and scattered solar radiation using a solar photometer, and then calculates relevant aerosol characteristic parameters such as AOD, AE, and spectral distribution based on the absorption and scattering of solar radiation by aerosols. Ground-based remote sensing allows for continuous real-time observation and provides high-precision inversion results, but its observation coverage is small, and the distribution and number of stations are uneven. The construction and maintenance costs of these stations are high, requiring significant human and material resources, making it difficult to achieve continuous monitoring over large areas. With the development and maturation of satellite remote sensing technology and the continuous iteration and upgrading of high-precision sensors, satellite remote sensing can provide detailed information on the spatial variations of aerosols, possessing long-term and large-scale aerosol detection capabilities. This has significant implications and broad research prospects for global aerosol research.
[13] On the other hand, the rise and continuous development of artificial intelligence algorithms in recent years have led to data-driven machine learning methods demonstrating good performance in the inversion and prediction of aerosol optical thickness and particulate matter concentration. Machine learning methods have a strong nonlinear fitting ability for massive amounts of data and can be used as statistical methods to extract remote sensing information. Since the physical meaning of some inversion parameters is often difficult to describe accurately, this limits the development of physical model-based inversion methods, while machine learning methods can better solve this type of quantitative parameter inversion problem. Among them, neural network models have a significant advantage over other machine learning methods in terms of the accuracy of remote sensing parameter estimation.
[14] By utilizing satellite remote sensing data and auxiliary data, and employing a neural network model, this study aims to achieve high-precision estimation of three parameters: aerosol optical thickness, angstrom index, and fine mode proportion. This research is significant in promoting the monitoring and research of atmospheric aerosols at regional and global scales, and has important research value and practical significance.
[0004] Traditional aerosol remote sensing inversion methods are mostly based on radiative transfer processes, involving numerous parameters such as surface reflectance and observation geometry. Before inversion, all possible scenarios for the target area need to be considered to establish a lookup table (LUT). Based on the radiance simulated by the LUT, the apparent reflectance observed by the satellite is compared to retrieve the aerosol characteristic parameters. The challenges of aerosol satellite remote sensing inversion lie in uncertainties such as cloud pixel interference, aerosol model assumptions, and surface reflectance estimation.
[15] Among these, effectively decoupling the surface and atmosphere to separate the surface contribution from the satellite-received signal is crucial. The accuracy of the surface reflectance estimation significantly affects the final inversion results of aerosol satellite remote sensing. Studies have shown that a surface reflectance error of 0.01 can lead to an AOD error of 0.1, resulting in an error of approximately 10 times.
[16] Due to the inherent heterogeneity and anisotropy of the Earth's surface, and the spatial variation in surface reflectivity caused by different types of surface cover, surface reflectivity also changes over time due to factors such as solar radiation conditions and vegetation growth.
[17] For single-view, non-polarized satellite sensors such as the Moderate Resolution Imaging Spectroradiometer (MODIS), the lack of multi-angle observation information and the influence of spectral band selection make accurate estimation of surface reflectance a challenge. Building a Local Underlying Data Set (LUT) is time-consuming, and the lack of multi-angle observations makes it difficult to pre-define reasonable scenarios, set accurate parameters, and construct a more precise LUT, thus affecting the accuracy of aerosol characteristic parameter inversion results. Therefore, designing an aerosol characteristic parameter estimation model that balances efficiency and accuracy is currently a research hotspot.
[0005] Currently, there are still many uncertainties in estimating Earth's climate forcing, and aerosols are one of the biggest factors among them.
[18] Aerosol optical depth (AOD) is considered one of the main predictors of ground-level particulate matter concentrations (such as PM10 and PM2.5).
[19] By multiplying the fine modal proportion (FMF) by the aerosol optical depth (AOD), the fine modal optical depth (FAOD) can be obtained. Fine particulate matter typically dominates the contribution of AOD, therefore FAOD correlates better with PM2.5 than AOD.
[20] The AE index contains information about aerosol types. Combined with AOD and FMF using different thresholds, it can be used to determine aerosol types. Large-scale observations of the optical, microphysical, and spatiotemporal distribution of aerosols are of great significance for studying environmental and climate effects. Most satellite-derived FMF products are highly unreliable for land areas. MODIS's DT algorithm has been used to produce global land and ocean FMF products, but its land-based products are not recommended for use due to their unreliability. [21,22] Levy et al.
[23] The results of MODIS-retrieved atmospheric radiation (AE) products over land are poor in quantitative analysis, and users are advised to derive and evaluate AE parameters themselves. Satellite remote sensing inversion of aerosol characteristic parameters is essentially a nonlinear regression problem. Compared to classical machine learning methods, deep learning methods are superior in fitting nonlinear relationships. LUT-based inversion methods require prior knowledge and assumptions about reasonable aerosol models and surface parameters. However, deep learning methods, utilizing multi-band spectral reflectance data and other auxiliary data, can train models to directly fit atmospheric radiative transfer processes, improving both computational efficiency and accuracy.
[24] However, most current research focuses on inputting multi-band information from a single pixel to perform point-to-point characteristic parameter inversion, failing to fully utilize the spatial information surrounding the target pixel. Therefore, a convolutional neural network (CNN) is chosen as the backbone network. An image containing multi-band data is input, and convolutional kernels are used to capture information from the local receptive field. Features are extracted layer by layer to obtain the aerosol information hidden between multi-band reflectances, thereby fitting their complex nonlinear relationships. The aim is to achieve high-precision daily estimation of three aerosol characteristic parameters with a resolution of 1 kilometer, enabling quantitative applications. Summary of the Invention
[0006] This invention provides a dual-branch aerosol characteristic parameter estimation model based on a convolutional neural network. The input is a 12×12 image data containing 10 band channels centered on the target pixel. By utilizing its automatic feature extraction feature, spatial and channel information is extracted and compressed. The one-dimensional data is constrained by a fully connected neural network branch to avoid the adverse effects caused by interference with the reflectivity of a single pixel band, thereby improving the estimation accuracy and robustness.
[0007] The technical solution adopted in this invention is as follows: A dual-branch aerosol characteristic parameter estimation method based on convolutional neural networks is proposed. This method consists of two parts: the establishment of the dual-branch Aerosol Characteristic Parameters Estimation Network (DAeroNet) and the preparation of the dataset. DAeroNet comprises a main branch convolutional neural network (CNN) and a second branch fully connected neural network (FCNN). The main branch convolutional neural network includes three convolutional blocks. The first two convolutional blocks each consist of a Feature Extraction Module (FEM), a 3×3 standard convolutional layer, and a 2×2 max-pooling layer. The last convolutional block is a 3×3 standard convolutional layer. The FEM incorporates a Channel Attention Module. CAM replaces standard convolution; in the first convolutional block, except for the channel attention module CAM, the number of convolutional kernels for all other convolutions is set to 512; in the second convolutional block, similarly, except for CAM, the number of convolutional kernels for all other convolutions is set to 256; the last standard convolution has 128 kernels; the convolutional neural network branch is used to process 12×12 image data containing 10 channels; through layer-by-layer convolutional feature extraction and pooling compression, a 128×3×3 feature map is obtained; the feature map is regularized to 0.3 and Dropout is applied, and after flattening, it is passed through a fully connected layer to obtain a one-dimensional data with a length of 500; The fully connected neural network FCNN has four fully connected layers with hidden nodes of 512, 256, 128, and 50 respectively. The 15 input features of the center pixel are first increased to 512 dimensionality in the first fully connected layer, then progressively reduced in dimensionality to learn the non-linear relationships, finally yielding a one-dimensional data point of length 50. The one-dimensional data outputs from the two branches are concatenated and then passed through a fully connected layer to output a single target parameter. In the convolutional neural network (CNN) branch, each convolutional layer uses the ReLU activation function, while in the fully connected neural network (FCNN) branch, the Leaky ReLU activation function is used. The dataset is obtained by extracting data from three products (MOD021KM, MOD09, and MOD03), performing corresponding pixel matching, and then spatiotemporally matching it with AERONET site data to filter out valid data. K-Means algorithm and multi-band thresholding are used to remove interfering pixels related to clouds, ice, and snow.
[0008] Further: The structure of the Feature Extraction Module (FEM) is as follows: the input original feature map is divided into two parts according to the channel dimension, each containing half of the channel data; the two parts of the feature map are subtracted to obtain the difference feature map between the channel information, and the three feature maps are then reassembled to obtain a new feature map; after simple recombination, the difference information between channels is added while retaining the original channels, which facilitates the network to learn the correlation between channels; especially for the input original image, the difference between Top of Atmosphere Reflectance (TOA) and Land Surface Reflectance (LSR) can be obtained, and this difference information contains rich aerosol information; after channel recombination, the learning and nonlinear fitting of the network are accelerated; the recombined feature map is passed through two-dimensional convolutional layers with 3×3 and 1×1 convolutional kernels respectively to extract feature information from different scales, and then passed through CAM to determine the importance weight of each channel, highlighting channels containing important feature information and suppressing channels containing noise or irrelevant information, and finally the two parts of the feature map are added as the output.
[0009] The design principle of the Channel Attention Module (CAM) is as follows: After performing max pooling along the horizontal and vertical coordinate directions respectively, the two feature vectors are concatenated together. After dimensionality reduction through a 1×1 two-dimensional convolution, they are split back into two feature vectors. After passing through the ReLU activation function, they are then subjected to 1×1 two-dimensional convolution to restore the original number of channels. The channel dimensionality reduction ratio is uniformly set to 2. In addition, global max pooling is used to extract global spatial information of each channel for supplementation. The pooled features are then processed by a 3×3 one-dimensional convolution. These feature vectors are then processed by the Sigmoid activation function to obtain the corresponding channel weights. The channel weights are multiplied by the original feature map to obtain the final output feature map.
[0010] The specific method for preparing the dataset is as follows: Within half an hour before and after the satellite's transit, the average of the ground station data for that period is taken. A 6-kilometer radius around the ground station is used as the sampling window. Assuming uniform aerosol distribution, the satellite product is matched with the ground data based on the nearest pixel. A cloud masking algorithm combining K-means and multi-band thresholding is used to remove cloud, ice, and snow pixels, resulting in 12×12 pixel images. Each image contains data from 10 channels, namely the apparent reflectance and surface reflectance of Bands 1, 2, 3, 4, and 7. Furthermore, the 10 band data corresponding to the center pixel of the image are extracted separately, along with water vapor and four geometric angles, to form a total of 15 one-dimensional input features.
[0011] The main innovative points of this invention are as follows: (1) A dual-branch neural network model is designed by combining CNN and FCNN. CNN processes image data and FCNN processes one-dimensional data. Feature information is learned from the two scales and finally aggregated to achieve high-precision estimation of target parameters.
[0012] (2) A channel attention module (CAM) is proposed and integrated into the feature extraction module (FEM) to enhance the model’s feature extraction and nonlinear fitting capabilities.
[0013] (3) The cloud mask algorithm obtained by combining the K-means algorithm with multi-band threshold can effectively improve the recognition accuracy of interference pixels such as clouds, ice, and snow that may affect the estimation accuracy.
[0014] Compared with existing data products, experimental results demonstrate that DAeroNet has better robustness, generalization ability, and potential for quantitative estimation of aerosol characteristic parameters compared with mainstream algorithm models. Attached Figure Description
[0015] Figure 1 Comparison charts showing the effects of several cloud masking technologies; Figure 2 12×12km 2 Schematic diagram of image matching with ground stations; Figure 3 This is a diagram showing the overall structure of the DAeroNet model of this invention; Figure 4 This is a structural diagram of the FEM of the present invention; Figure 5 This is a detailed structural diagram of the CAM of the present invention; Figure 6 Validation results for independent sites in the Northwest region for each model to estimate AOD; Figure 7 Validation results for independent sites in the Northwest region for estimating AE for each model; Figure 8 Validation results for FMF at independent sites in the Northwest region were estimated for each model; Figure 9 Spatial distribution of AOD as of August 1, 2022; Figure 10 Spatial distribution of AE as of August 1, 2022; Figure 11 Spatial distribution of FMF as of August 1, 2022; Figure 12 Spatial distribution of AOD as of September 1, 2022; Figure 13 Spatial distribution of AE as of September 1, 2022; Figure 14The spatial distribution of FMF as of September 1, 2022. Detailed Implementation
[0016] The invention will be further explained below with reference to the accompanying drawings, both theoretically and experimentally.
[0017] Reference Figure 1-5A dual-branch aerosol characteristic parameter estimation method based on convolutional neural networks is proposed. This method consists of two parts: the establishment of the dual-branch Aerosol Characteristic Parameters Estimation Network (DAeroNet) and the preparation of the dataset. DAeroNet comprises a main branch convolutional neural network (CNN) and a second branch fully connected neural network (FCNN). The main branch convolutional neural network includes three convolutional blocks. The first two convolutional blocks each consist of a Feature Extraction Module (FEM), a 3×3 standard convolutional layer, and a 2×2 max-pooling layer. The last convolutional block is a 3×3 standard convolutional layer. Channel attention modules are integrated into the FEM module. The CAM module replaces the standard convolution; in the first convolutional block, except for the channel attention module CAM, the number of convolutional kernels for all other convolutions is set to 512; in the second convolutional block, similarly, except for CAM, the number of convolutional kernels for all other convolutions is set to 256; the last standard convolution has 128 kernels; the convolutional neural network branch is used to process 12×12 image data containing 10 channels; through layer-by-layer convolutional feature extraction and pooling compression, a 128×3×3 feature map is obtained; the feature map is regularized to 0.3 and Dropout is applied, and after flattening, it is passed through a fully connected layer to obtain a one-dimensional image of length 500. The data; the fully connected neural network FCNN has four fully connected layers with hidden nodes of 512, 256, 128, and 50 respectively; the 15 input features of the center pixel of the fully connected neural network are first increased to 512 dimensions through the first fully connected layer, and then the dimensions are reduced layer by layer to learn the non-linear relationship, finally obtaining a one-dimensional data of length 50; the one-dimensional data outputs of the two branches are concatenated, and finally a single target parameter is output through the fully connected layer; in the convolutional neural network (CNN) branch, the ReLU activation function is used after each convolutional layer, while the Leaky ReLU activation function is used in the fully connected neural network (FCNN) branch; the dataset is obtained by extracting data from three products MOD021KM, MOD09, and MOD03, matching corresponding pixels, and then matching them spatiotemporally with AERONET site data to select effective data, and using the K-Means algorithm and multi-band thresholding to remove cloud, ice, and snow interference pixels.
[0018] Reference Figure 4The structure of the Feature Extraction Module (FEM) is as follows: the input original feature map is divided into two parts according to the channel dimension, each containing half of the channel data; the two parts of the feature map are subtracted to obtain the difference feature map between the channel information, and the three feature maps are then reassembled to obtain a new feature map; after simple recombination, the difference information between channels is added while retaining the original channels, which facilitates the network to learn the correlation between channels; especially for the input original image, the difference between Top of Atmosphere Reflectance (TOA) and Land Surface Reflectance (LSR) can be obtained, and this difference information contains rich aerosol information; after channel recombination, the learning and nonlinear fitting of the network are accelerated; the recombined feature map is passed through two-dimensional convolutional layers with 3×3 and 1×1 convolutional kernels respectively to extract feature information from different scales, and then passed through CAM to determine the importance weight of each channel, highlighting the channels containing important feature information and suppressing the channels containing noise or irrelevant information, and finally the two parts of the feature map are added as the output.
[0019] To improve cross-channel learning, the weights of important features are selectively enhanced, focusing on effective pixel information and reducing interference from invalid pixels. A Channel Attention Module (CAM) is designed, such as... Figure 5 As shown, its design principle is as follows: After performing max pooling along the horizontal and vertical coordinate directions respectively, the two feature vectors are concatenated together. After dimensionality reduction through a 1×1 two-dimensional convolution, they are split back into two feature vectors. After passing through the ReLU activation function, they are then subjected to 1×1 two-dimensional convolution to restore the original number of channels. The channel dimensionality reduction ratio is uniformly set to 2. In addition, global max pooling is used to extract global spatial information of each channel for supplementation. The pooled features are then processed by a 3×3 one-dimensional convolution. These feature vectors are then processed by the Sigmoid activation function to obtain the corresponding channel weights. The channel weights are multiplied by the original feature map to obtain the final output feature map.
[0020] The specific method for preparing the dataset is as follows: Within half an hour before and after the satellite's transit, the average of the ground station data for that period is taken. A 6-kilometer radius around the ground station is used as the sampling window. Assuming uniform aerosol distribution, the satellite product is matched with the ground data based on the nearest pixel. A cloud masking algorithm combining K-means and multi-band thresholding is used to remove cloud, ice, and snow pixels, resulting in 12×12 pixel images. Each image contains data from 10 channels, namely the apparent reflectance and surface reflectance of Bands 1, 2, 3, 4, and 7. Furthermore, the 10 band data corresponding to the center pixel of the image are extracted separately, along with water vapor and four geometric angles, to form a total of 15 one-dimensional input features.
[0021] Experimental verification of the present invention: 1. Ground stations and verification areas Northwest my country is located in the mid-latitude arid and semi-arid zone, characterized by drought and water scarcity. In recent years, influenced by global warming, precipitation has increased in some areas, and the climate is gradually shifting from "warm and dry" to "warm and humid." [25–28] Northwest China possesses vast deserts and Gobi, experiencing frequent sandstorms. With the region's economic development, increased human activities such as industrial emissions, vehicle exhaust emissions, and energy consumption have also generated substantial aerosol production. This invention focuses on Northwest my country, at 76.7°E. ◦ Up to 106.1 ◦ 34.1°N ◦ Up to 41.3 ◦ The selected region was chosen as the validation area. This region encompasses southern Xinjiang Uygur Autonomous Region, northern Tibet Autonomous Region, northern Qinghai Province, and central Gansu Province, including parts of several deserts such as the Taklamakan Desert, Tengger Desert, and Badain Jaran Desert, characterized by high altitudes and complex terrain. Due to the scarcity of stations in Northwest China and the relatively short timeframe and insufficient data sample size of currently available publicly available data, 50 AERONET ground stations in East, Central, and South Asia were selected to obtain long-term ground observation data for model training and evaluation. Furthermore, observational data from two ground stations—Qinghai Lake-Waliguan Mountain (Mt_WLG) in Qinghai Province and the Lanzhou University Semi-Arid Climate and Environment Observatory (SACOL) in Gansu Province—were selected for independent site validation of the model.
[0022] 2. Data source 2.1 AERONET ground observation data The AERONET global automatic aerosol observation network, based on the CIMEL CE318 automatic sun photometer, provides various aerosol-related parameters, including AOD, across eight channels: 1640, 1020, 870, 675, 500, 440, 380, and 340 nanometers. These parameters are used for satellite inversion verification and synergy with other databases. This invention utilizes observation data from 52 ground stations, as detailed in Table 1. The data from 50 stations covers the period from 2020 to 2022, while the data from two independent stations covers the period from 2009 to 2013. The AERONET observation parameters required for this invention include AOD at 550 nanometers, AE in the 440-870 nanometer range, and FMF, which are used as ground truth values for model training, testing, and verification.
[0023] Since AERONET does not directly provide AOD data at a wavelength of 550 nm, band interpolation using the angstrom index is required to obtain the AOD value at 550 nm. The specific calculation process is as follows: (1) (2) (3) The AOD value refers to the wavelength; represents the angstrom index; and is the turbidity coefficient.
[0024] In this invention, the AOD value at 550 nm is calculated using the AOD values at the 440 nm and 670 nm wavelengths. The data used can be downloaded from the official website (https: / / aeronet.gsfc.nasa.gov).
[0025] 2.2 MODIS Data The Moderate Resolution Imaging Spectroradiometer (MODIS), carried by the Terra and Aqua satellites, provides 36 measurement bands from visible light to thermal infrared (0.405–14.385 μm), suitable for remote sensing monitoring of atmospheric aerosols. Table 2 shows the characteristic attributes of MODIS bands 1 to 7. Terra crosses the equator from north to south at approximately 10:30 AM local solar time, while Aqua crosses the equator from south to north at approximately 1:30 PM local solar time. Each observation covers a width of approximately 2330 km, allowing observation of the entire Earth's surface every 1 to 2 days, and repeats its orbit every sixteen days. This invention uses data from six products from the Terra satellite: MOD021KM, MOD09, MOD35_L2, MOD03, MOD04_L2, and MOD04_3K. Data can be downloaded from the MODIS website (https: / / modis.gsfc.nasa.gov).
[0026] 2.3 MERRA-2 Data MERRA-2 (Modern-Era Retrospective Analysis for Research and Applications, Version 2) is an atmospheric reanalysis dataset developed by NASA's Goddard Space Flight Center. It provides high-resolution atmospheric data from 1980 to the present and is widely used in climate research, environmental monitoring, and weather forecasting.
[29] MERRA-2 aerosol reanalysis provides long-term aerosol parameter simulations from 1980 to the present. This invention will use the MERRA-2 tavg1_2d_aer_Nx product, which provides a global, hourly spatial resolution of 0.5. ◦ ×0.625◦ The column mass density, surface mass concentration, and other parameters of the aerosol components (black carbon, dust, sea salt, sulfate, and organic carbon) were used. The products used were downloaded from the GES DISC website (https: / / disc.gsfc.nasa.gov). The 550 nm AOD and 470-870 nm AE of the product were selected. The AE of the 470-870 nm and 440-870 nm bands are very close and the difference is negligible, so the two can be directly compared. MERRA-2 does not provide FMF parameters, which can be obtained by conversion using formula (4).
[30] : (4)
[0027] Table 1. AERONET ground sites used in this invention
[0028] Table 2 Characteristic attributes of MODIS Bands 1-7
[0029] 3. Cloud mask design Thick clouds can block reflected signals from the ground, while thin clouds can increase the reflectivity of incident satellite signals. In addition, surface cover types such as ice and snow can also increase reflectivity, thus affecting the accuracy of aerosol characteristic parameter estimation.
[31] The MODIS-provided 1km resolution cloud mask product MOD35_L2 can accurately identify cloud-covered pixels, but it still suffers from misclassification and missed detection of cloud pixels. Therefore, a cloud mask algorithm needs to be designed for the verification area and all ground stations used to improve the accuracy of cloud, ice, and snow pixel identification. In the cloud mask design process, the K-means algorithm was first attempted, using Band 1 and Band 2 from the MOD021KM product to distinguish between cloud-covered and cloudless pixels. The results showed significant missed detection issues. Then, a multi-band threshold algorithm was used, with continuous adjustment of the threshold parameters, resulting in better cloud identification accuracy. Combining the identification results of the two methods revealed that they could complement each other, further improving the cloud pixel identification rate. Based on this, the Normalized Difference Snow Index (NDSI) was added to identify pixels with ice or snow cover types, eliminating pixels that would affect estimation accuracy for subsequent dataset preparation. The calculation process of NDSI is shown in Formula 5. The cloud masking algorithm of this invention requires multiple band data from the MOD021KM product. Specific parameters and thresholds are shown in Table 3. Here, BT11.2 refers to the brightness temperature at a wavelength of 11.2 micrometers. This invention selects Band 31 with a wavelength range of 10.78 micrometers to 11.28 micrometers; the conversion process is shown in Formula 6. Cloud masking data identified as clouds in the MOD35_L2 product is compared with the cloud masking algorithm of this invention. Through observation and comparison, it can be found that the MOD35_L2 product exhibits obvious misidentification and omission of cloud pixels. The cloud masking obtained using the cloud masking algorithm of this invention shows significant improvement in these areas. Figure 1 As shown in the middle red circle.
[0030] Table 3. Pixel Filtering Thresholds for Clouds, Ice, and Snow
[0031] (5) (6) In the formula, R represents radiance.
[0032] 4. Dataset Preparation When creating the dataset, it is necessary to select feature factors that are highly correlated with the estimated target parameters while avoiding parameter redundancy. Since the apparent reflectance (TOA) of the MODIS 2.12-micron band (Band 7) exhibits a good linear relationship with surface reflectance under different aerosol optical thickness conditions, it is less affected by aerosols. Furthermore, the surface reflectance of the 0.47-micron (Band 3) and 0.66-micron (Band 1) bands can be converted from the surface reflectance of the 2.12-micron band using empirical formulas. These three bands are also commonly used in radiative transfer model inversion methods such as the Deep Blue algorithm and the dark target method, and are therefore quite important. (Chen et al.)
[21] Band 5 and Band 6 band losses were found in the MOD02HKM product; therefore, this invention does not consider using data from these two bands to avoid interference. Due to the difficulty in separating aerosols from surface components, especially in complex surface areas, surface reflectance has been considered for aerosol inversion.
[32] To fully utilize the aerosol information contained in multi-band reflectance, and referring to the feature factor selection of the physical model-based inversion method, and considering the data redundancy problem, the apparent reflectance of Band1, Band2, Band3, Band4, and Band7 and their corresponding surface reflectance were selected, along with water vapor and four geometric angles, to form a total of 15-dimensional input features, as detailed in Table 4.
[0033] Table 4. Input features of the dataset used in this invention
[0034] Data extraction and pixel matching were performed on three products: MOD021KM, MOD09, and MOD03. The average of ground station data was taken within half an hour before and after the satellite's transit. A 6-kilometer radius centered on the ground station was used as the sampling window. Assuming uniform aerosol distribution, the satellite product was matched with the ground data based on the nearest pixel. The cloud mask algorithm of this invention was used to remove cloud, ice, and snow pixels. To facilitate pooling operations in the subsequent CNN network, the width and height of the dataset images were chosen to be even numbers. Finally, 12×12 pixel images centered on the station were created, each containing data from 10 channels: apparent reflectance and surface reflectance of Bands 1, 2, 3, 4, and 7. Each pixel represents 1 km. 2 The obtained image spatial range is 12×12km. 2 This invention uniformly uses the pixels in the seventh row and seventh column as the center pixel for matching with ground stations. For example... Figure 2 As shown, red dots represent ground stations, and green squares represent center pixels.
[0035] There is research[21, 33] It was found that when the proportion of cloudless pixels in the image data used to train the model was above 90%, the impact on the accuracy of the estimation results was minimal, while the accuracy decreased when the proportion dropped below 90%. To ensure accuracy and obtain more data samples, only images in the dataset with more than 90% of the total pixels being valid (i.e., cloudy pixels not exceeding 10%) were retained. A one-dimensional list dataset corresponding to the center pixels of the images was generated, containing the 15 input features shown in Tables 3-4. Matching was performed on 50 AERONET sites over the three years from 2020 to 2022, resulting in 12,266 valid data entries used for model training and evaluation.
[0036] 5. Experimental Setup The experiment used observational data from 50 AERONET sites and MODIS satellite product data from 2020-2022, preprocessed and spatiotemporally matched to create a dataset. Ten-fold cross-validation based on site, time, and sample was performed to evaluate the model. Independent site validation was conducted using sites within the validation area that were not used in the training phase to further examine the model's generalization and robustness. For the site-based ten-fold cross-validation, the data was randomly divided into 10 groups. Each group's training set contained data from 45 sites, and the validation set contained data from the remaining 5 sites. The specific groupings are shown in Table 5. Table 5. Validation set partitioning based on site-specific 10-fold cross-validation
[0037] Cross-validation, as described above, is a statistical method used to evaluate the performance and generalization ability of machine learning models, and it is also a widely used evaluation technique in the field of remote sensing.
[31] This method involves dividing the dataset into multiple subsets and then performing multiple training and validation iterations on these subsets to obtain a more reliable performance evaluation. There are various partitioning criteria; this invention performs cross-validation on the dataset based on three partitioning methods: ground station, time, and sample, to evaluate model performance. Specifically: using ground station observation data as the target ground truth, the dataset is randomly divided into 10 training and validation sets based on the location of the ground stations used, for 10-fold site cross-validation of the model; the dataset is arranged chronologically and divided into 10 training and validation sets in chronological order for 10-fold time cross-validation of the model; finally, the dataset is randomly divided into 10 different training and validation sets in an 8:2 ratio for 10-fold sample cross-validation of the model. Furthermore, the model's generalization and robustness are further verified by training the model using all datasets as training samples and testing it with observation data from independent stations not previously used for training. This invention uses several statistical equations, including Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), Mean Bias Error (MBE), Pearson Correlation Coefficient (R), Index of Agreement (IOA), and Expected Error (EE), to evaluate the accuracy of each model's estimation results in terms of accuracy, bias, and correlation.
[0038] This experiment was conducted using Python 3.8.16 and the open-source deep learning framework PyTorch. The GPU was an NVIDIA 4090 with 24GB of VRAM, accelerated using cuDNN 11.8. The CPU was an Intel i9-13900K with a frequency of 3.0GHz. In the experiment, the initial learning rate of the neural network model was set to 0.00001, and the Adam optimizer was used to adjust the learning rate. The number of training iterations on the dataset was set to 800, and the batch size was set to 500. For the ANN network, the input was 15-dimensional data of the center pixel; for the CNN network, the input was image data. For the machine learning model, the input was also 15-dimensional data of the center pixel, and Bayesian optimization was used to select appropriate hyperparameters. Through experiments, the mean absolute error (MAE) and mean squared error (MSE) were compared, and MAE, which has lower sensitivity to outliers, was ultimately chosen as the loss function.
[0039] 6. Experimental Results 6.1 Ablation Experiment To evaluate the effectiveness of each branch and module of the model in this invention, site-based 10-fold cross-validation was used to conduct comprehensive ablation experiments. Tables 6, 7, and 8 show the experimental results of the relevant modules estimating the three aerosol characteristic parameters AOD, AE, and FMF, respectively. As can be seen from the tables, the individual estimation results of the CNN and FCNN branches are not ideal. Combining the two branches or combining the CNN branch with the FEM module can improve the overall estimation accuracy of the model. The best results are achieved when all relevant modules are combined to form DAeroNet, where all aerosol characteristic parameter estimation results are the best except for MBE, demonstrating the effectiveness and importance of each module.
[0040]
[0041]
[0042]
[0043] 6.2 Model Comparison To verify the superiority of the proposed method in estimating aerosol characteristic parameters, a comparative experiment was conducted under the same experimental environment, comparing it with seven mainstream algorithm models for the three parameters AOD, AE, and FMF, using three types of 10-fold cross-validation and independent site validation.
[0044] 6.2.1 Cross-validation results Tables 9, 10, and 11 show the 10-fold cross-validation results for estimating AOD; Tables 12, 13, and 14 show the 10-fold cross-validation results for estimating AE; and Tables 15, 16, and 17 show the 10-fold cross-validation results for estimating FMF.
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052] Table 16 Time-based 10-fold cross-validation of FMF estimates for each model
[0053]
[0054] As shown in Tables 9 to 17, all models performed best in sample-based cross-validation, followed by time-based cross-validation. Site-based cross-validation, clearly the most stringent of the three, best tests model generalization, and therefore yielded the worst results. A comparison reveals that DAeroNet performed best in estimating the three characteristic parameters across almost all cross-validation metrics, except for MBE (only in the site-based 10-fold cross-validation for FMF estimation, where RMSE was second only to XGBoost). Machine learning models performed well in MBE estimation. Neural network models (DenseNetReg, ResNetReg, NNAero) generally outperformed machine learning models (RF, LGBM, XGBoost, CatBoost) in AOD estimation, but the opposite was true for AE and FMF estimation. DAeroNet achieved EEs exceeding 68% in all estimations except for 63.29% in the site-based 10-fold cross-validation for AOD estimation, where the percentage falling within the expected error (EE). 0.68 is the probability of one standard deviation of a normal distribution; exceeding this indicates that the estimation results are likely to be used quantitatively.
[31] .
[0055] 6.2.2 Independent Site Verification Results To further evaluate the effectiveness and generalization of the models, independent site validation was performed on sites within the validation region that had never been trained on. Matching data from these independent sites from 2009 to 2013, totaling 177 valid data points, was used for validation. All models were trained using 12,266 valid data points from 50 sites from 2020 to 2022. Figure 6 , 7 Figures 8 and 9 show the estimation results of the three characteristic parameters, respectively.
[0056] Depend on Figures 6 to 8It can be seen that, except for AOD estimation where EE is second only to NNAero and R is second only to ResNetReg, DAeroNet performs best in all other metrics when estimating the three target parameters (in the figure, EE represents the percentage of points falling within the expected error envelope, and the number of points within the expected error is in parentheses; the red line is the 1:1 line of the ground measurement; the blue line is the expected error envelope; the color of the scatter points indicates the degree of clustering, the denser the clustering, the darker the red of the corresponding scatter points, and vice versa). All models underestimated AOD to varying degrees, especially when AOD > 0.2, the underestimation was more obvious. Conversely, all models overestimated AE and FMF, especially in the low-value portion. DAeroNet's validation results for AOD, AE, and FMF showed EE metrics of 53.11%, 76.84%, and 61.58%, respectively, with only AE exceeding 68%, reaching the point of quantitative quantification. Clearly, the significant temporal and spatial differences between training and independent validation data impact the model's estimation accuracy. Based on the validation results described above, DAeroNet exhibits good robustness and generalization ability, demonstrating superior performance. The number of parameters (Params) and floating-point operations (FLOPs) are commonly used metrics for measuring model structural complexity and computational load. Table 3-18 shows the complexity of each neural network model. Due to differences in complexity measurement methods between classic machine learning methods like RF and neural networks, relevant metrics for the four machine learning models are not listed in the table. Generally, under standard hyperparameter settings, the complexity of classic machine learning models is usually lower than that of neural network models, and this is also true for the machine learning method used in this paper. As shown in Table 18, the overall Params and FLOPs of all models are at a low level. Although DAeroNet's model complexity is relatively high, it still has an advantage in performance.
[0057]
[0058] 6.3 Spatial Distribution Display To evaluate the spatial distribution of aerosol characteristic parameters estimated by DAeroNet over the entire region, AOD, AE, and FMF data from MERRA-2, AOD data from the Dark Target (DT) method of MOD04_3K, and AOD data from DTAOD, Deep Blue (DB) algorithm, and combined AOD data from DT and DB of MOD04_L2 were compared. Figure 9 , 12 The images show the spatial distribution of AOD within the verification area on August 1st and September 1st, 2022, respectively. Figure 10 ,13 The spatial distribution of AE over the past two days, Figure 11 , 14 The spatial distribution of FMF over these two days is shown. The spatial resolution of the MOD04_3K product is 3 km, the MOD04_L2 product is 10 km, the MERRA-2 product is 0.5° × 0.625° (approximately 50 km), and the DAeroNet estimation result has a spatial resolution of 1 km. The spatial distribution of AOD estimated by DAeroNet is highly similar to the data from these products, especially to the DBAOD data of MOD04_L2. The similarity between the AE and FMF estimated by DAeroNet and the MERRA-2 product is also good. High values are evident in the region east of 100°E and south of 40°N, while the DAeroNet estimation results are significantly lower in the region west of 81°E and south of 36°N. Overall, the spatial distribution of the three aerosol characteristic parameters estimated by DAeroNet is in good agreement with existing products.
[0059] Table 19 shows the results of Wei et al. [30,34,35,36,37] Performance evaluations of various products, including MOD04_3K and MOD04_L2, were conducted in East Asia, Southeast Asia, and South Asia. A comparison with DAeroNet's sample-based 10-cross-validation results shows that DAeroNet achieves a performance level comparable to existing products in estimating three aerosol characteristic parameters, and even demonstrates superior performance in the EE (Extreme Emissions) metric.
[0060] Table 19 Performance Evaluation of Data Products
[0061] References [1] Lee KH, Li Z, Kim YJ, et al. Atmospheric aerosolmonitoring from satellite observations: A history of three decades[J]. Atmospheric and Biological Environmental Monitoring, 2009. 13–38. [2] Kang Fugui, Li Yaohui. A review of dust aerosol research in Northwest China over the past 10 years [J]. Arid Meteorology, 2011, 29(2):144–150. [3] Tang Yuming, Deng Ruru, Xu Minduan, et al. Diurnal variation of aerosol optical properties in Guangzhou in autumn [J]. Journal of Sun Yat-sen University (Natural Science Edition), 2019, 58(2):58–67. [4] Rap A, Scott CE, Spracklen DV, et al. Naturalaerosol direct and indirect radiative effects[J]. Geophysical Research Letters, 2013, 40(12):3297–3301. [5] Procopio A, Artaxo P. Direct and semi-direct aerosol effects: amodeled study for biomass burning aerosol radiative forcing in the amazonregion[C]. Proceedings of AIP Conference Proceedings, volume 1100, 649–652. American Institute of Physics, 2009. [6] Liao Li, Lou Sijia, Fu Yu, et al. Radiative forcing of aerosols in eastern China on a synoptic scale and their impact on surface air temperature [J]. Atmospheric Sciences, 2015, 39(1):68–82. [7] Jion MMMF, Islam ARMT, Shahrier M, et al. Acritical review of no2 and aod in major asian cities: challenges, mitigation approaches and way forwards[J]. Air Quality, Atmosphere&Health, 2024. 1–17. [8] She Lu. Research on aerosol optical thickness inversion and dust monitoring based on himawari-8 / ahi data [D]. Doctoral dissertation. Beijing: University of Chinese Academy of Sciences (Institute of Remote Sensing and Digital Earth, Chinese Academy of Sciences), 2018. [9] Wang Y, Ali MA, Bilal M, et al. Identification of no2 and so2pollution hotspots and sources in jiangsu province of china[J]. RemoteSensing, 2021, 13(18):3742.
[10]
[10] Zhao H, Gui K, Wang Y, et al. Long-term distribution and evolution trends of absorption aerosol optical depth withdifferent chemical components in global and typical regions[J]. AtmosphericResearch, 2025, 314:107819.
[11] Su Danfeng, Zeng Sanwu. Study on the hazards of fine particulate matter PM2.5 to various systems of the human body [J]. Medical Information, 2019, 32(18):32–34.
[12] Zhang Chunyang. Study on the inversion of aerosol characteristic parameters in Kunming area using CE318 observation data [D]. Kunming: Yunnan University, 2017.
[13] Tang Yuming, Deng Ruru, Liu Yongming, et al. A review of research on remote sensing inversion of atmospheric aerosols [J]. Remote Sensing Technology and Application, 2018, 33(1):25–34.
[14] Fan Y, Sun L. Satellite aerosol optical depth retrieval based on fully connected neural network (fcnn) and a combine algorithm of simplifiedaerosol retrievalalgorithm and simplified and robust surface reflectanceestimation (sremara)[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2023, 16:4947–4962.
[15] Li Z, Zhao X, Kahn R, et al.Uncertainties in satellite remotesensing of aerosols and impact on monitoring its long-term trend: a reviewand perspective[C]. Proceedings of Annales Geophysicae, volume 27, 2755–2770.CopernicusGmbH, 2009.
[16] Wang Y, Xue Y, Guang J, et al.Simultaneously retrieval ofaerosol optical depth and surface albedo with fy-2 geostationary data[C].Proceedings of 2011 IEEE International Geoscience and Remote SensingSymposium, 2912–2914. IEEE, 2011.
[17] He T, Liang S, Wang D, et al.Estimation of surface albedo anddirectional reflectance from moderate resolution imaging spectroradiometer(modis) observations[J]. Remote Sensing of Environment, 2012, 119:286–300.
[18] Masson-Delmotte V, Zhai P, PiraniA, et al. Climate change 2021:the physical science basis[J]. Contribution ofWorking Group I to the SixthAssessment Report of the Intergovernmental Panel on Climate Change, 2021, 2(1):2391.
[19] Park S, Lee J, Im J, et al.Estimation of spatially continuousdaytime particulate matter concentrations under all sky conditions throughthe synergistic use of satellite-based aod andnumerical models[J]. Science ofthe Total Environment, 2020, 713:136516.
[20] Zhang Y, Li Z. Remote sensing ofatmospheric fine particulatematter (pm2. 5) mass concentration near the ground from satellite observation[J]. Remote Sensing of Environment, 2015, 160:252–262.
[21] Chen X, de Leeuw G, Arola A, etal. Joint retrieval of theaerosol fine mode fraction and optical depth using modis spectral reflectanceover northern and eastern china: Artificial neural network method[J]. RemoteSensing of Environment, 2020, 249:112006.
[22] Yan X, Zang Z, Li Z, et al. Aglobal land aerosol fine-modefraction dataset (2001–2020) retrieved from modis using hybridphysical anddeep learning approaches[J]. Earth System Science DataDiscussions, 2021,2021:1–27.
[23] Levy R C, Mattoo S, Munchak L, etal. The collection 6 modisaerosol products over land and ocean[J]. AtmosphericMeasurement Techniques,2013, 6(11):2989–3034.
[24] Nanda S, De Graaf M, Veefkind J P,et al. A neural networkradiative transfer model approach applied to the tropospheric monitoringinstrument aerosol height algorithm[J]. Atmospheric Measurement Techniques,2019, 12(12):6619–6634.
[25] Yihui D, Yanju L, Ying X, et al.Regional responses to globalclimate change: progress and prospects for trend, causes, and projection ofclimatic warming-wetting in northwest china[J]. Advances in Earth Science,2023, 38(6):551.
[26] Yao X, Zhang M, Zhang Y, et al.New insights into climatetransition in northwest china[J]. J. Arid. Land,2022. 1–15.
[27] Wang C, Zhang S, Li K, et al.Change characteristics ofprecipitation in northwest china from 1961 to 2018[J]. Chin. J. Atmos. Sci,2021, 45(4):713–724.
[28] Yang J, Jiang Z, Liu X, et al. Influence research on springvegetation of eurasia tosummer drought-wetness over the northwest china[J].Arid Land Geography, 2012,35(1):10–22.
[29] Randles C, Da Silva A, Buchard V,et al. The merra-2 aerosolreanalysis, 1980 onward. part i: System description and data assimilationevaluation[J]. Journal of Climate, 2017, 30(17):6823–6850.
[30] Su X, Huang Y, Wang L, et al.Validation and diurnal variationevaluation of merra-2 multiple aerosol properties on a global scale[J].Atmospheric Environment, 2023, 311:120019.
[31] Cao M, Zhang M, Su X, et al. Atwo-stage machine learningalgorithm for retrieving multiple aerosol properties over land: Developmentand validation[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023.
[32] Kang E, Park S, Kim M, et al.Direct aerosol optical depthretrievals using modis reflectance data and machine learning over east asia[J]. Atmospheric Environment, 2023, 309:119951.
[33] Cai H, Zhong B, Liu H, et al. Animproved deep learning networkfor aod retrieving from remote sensing imagery focusing on sub-pixel cloud[J]. GIScience&Remote Sensing, 2023, 60(1):2262836.
[34] Wei J, Li Z, Sun L, et al. Modiscollection 6.1 3 km resolutionaerosol optical depth product: Global evaluation and uncertainty analysis[J].Atmospheric Environment, 2020, 240:117768.
[35] Wei J, Li Z, Peng Y, et al. Modiscollection 6.1 aerosol opticaldepth products over land and ocean: validation and comparison[J]. AtmosphericEnvironment, 2019, 201:428–440.
[36] Huang G, Su X, Wang L, et al.Evaluation and analysis of long-term modis maiac aerosol products in china[J].Science of The TotalEnvironment, 2024, 948:174983.
[37] Ansari K, Ramachandran S. Opticaland physical characteristics ofaerosols over asia: Aeronet, merra-2 and cams[J]. Atmospheric Environment,2024, 326:120470。
Claims
1. A method for estimating the characteristic parameters of aerosols based on a convolutional neural network, characterized in that, The estimation method consists of two parts: the establishment of the dual-branch AerosolCharacteristic Parameters Estimation Network (DAeroNet) and the preparation of the dataset. DAeroNet comprises a main branch of a Convolutional Neural Network (CNN) and a second branch of a Fully Connected Neural Network (FCNN). The main branch of the CNN includes three convolutional blocks: the first two blocks each consist of a Feature Extraction Module (FEM), a 3×3 standard convolutional layer, and a 2×2 max-pooling layer; the last convolutional block is a 3×3 standard convolutional layer. The FEM incorporates a Channel Attention module. The CAM module replaces the standard convolution; in the first convolutional block, except for the channel attention module CAM, the number of convolutional kernels for all other convolutions is set to 512; in the second convolutional block, similarly, except for CAM, the number of convolutional kernels for all other convolutions is set to 256; the last standard convolution has 128 kernels; the convolutional neural network branch is used to process 12×12 image data containing 10 channels; through layer-by-layer convolutional feature extraction and pooling compression, a 128×3×3 feature map is obtained; the feature map is regularized to 0.3 and Dropout is applied, and after flattening, it is passed through a fully connected layer to obtain a one-dimensional image of length 500. The data; the fully connected neural network FCNN has four fully connected layers with hidden nodes of 512, 256, 128, and 50 respectively; the 15 input features of the center pixel of the fully connected neural network are first increased to 512 dimensions through the first fully connected layer, and then the dimensions are reduced layer by layer to learn the non-linear relationship, finally obtaining a one-dimensional data of length 50; the one-dimensional data outputs of the two branches are concatenated, and finally a single target parameter is output through the fully connected layer; in the convolutional neural network (CNN) branch, the ReLU activation function is used after each convolutional layer, while the Leaky ReLU activation function is used in the fully connected neural network (FCNN) branch; the dataset is obtained by extracting data from three products MOD021KM, MOD09, and MOD03, matching corresponding pixels, and then matching them spatiotemporally with AERONET site data to select effective data, and using the K-Means algorithm and multi-band thresholding to remove cloud, ice, and snow interference pixels.
2. The dual-branch aerosol characteristic parameter estimation model based on convolutional neural networks according to claim 1, characterized in that, The structure of the Feature Extraction Module (FEM) is as follows: the input original feature map is divided into two parts according to the channel dimension, each containing half of the channel data; the feature maps of the two parts are subtracted to obtain the difference feature map between the channel information; the three feature maps are then reassembled to obtain a new feature map. After simple reorganization, while retaining the original channels, the difference information between channels is added to facilitate the network's learning of the correlation between channels; In particular, for the original input image, it can obtain the difference between the apparent reflectance (Top of Atmosphere Reflectance, TOA) and the land surface reflectance (LSR), which contains rich aerosol information. After channel reconstruction, the learning and nonlinear fitting of the network are accelerated. The reconstructed feature map is passed through two-dimensional convolutional layers with 3×3 and 1×1 kernels respectively to extract feature information from different scales. Then, it is passed through CAM to determine the importance weight of each channel, highlighting the channels containing important feature information and suppressing the channels containing noise or irrelevant information. Finally, the two feature maps are added together as the output.
3. The dual-branch aerosol characteristic parameter estimation model based on convolutional neural networks according to claim 1, characterized in that, The design principle of the Channel Attention Module (CAM) is as follows: after performing max pooling along the horizontal and vertical coordinate directions respectively, the two feature vectors are concatenated together. After dimensionality reduction through a 1×1 two-dimensional convolution, they are split back into two feature vectors. After passing through the ReLU activation function, they are then dimensionality-upped by a 1×1 two-dimensional convolution to restore the original number of channels. The channel dimensionality reduction ratio is uniformly set to 2. In addition, global max pooling is used to extract global spatial information of each channel for supplementation. The pooled features are then processed by a 3×3 one-dimensional convolution. These feature vectors are processed by the Sigmoid activation function to obtain corresponding channel weights; the channel weights are then multiplied by the original feature map to obtain the final output feature map.
4. The dual-branch aerosol characteristic parameter estimation model based on convolutional neural networks according to claim 1, characterized in that, The specific method for preparing the dataset is as follows: Within half an hour before and after the satellite's transit, the average of the ground station data for that period is taken. A 6-kilometer radius around the ground station is used as the sampling window. Assuming uniform aerosol distribution, the satellite product is matched with the ground data based on the nearest pixel. A cloud masking algorithm combining K-means and multi-band thresholding is used to remove cloud, ice, and snow pixels, resulting in 12×12 pixel images. Each image contains data from 10 channels, namely the apparent reflectance and surface reflectance of Bands 1, 2, 3, 4, and 7. Furthermore, the 10 band data corresponding to the center pixel of the image are extracted separately, along with water vapor and four geometric angles, to form a total of 15 one-dimensional input features.