A method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data
By acquiring high-resolution LAI data and MODIS AOD aerosol data, combined with the segmented quantile regression method, the uncertainty problem of the relationship between vegetation coverage and aerosol concentration in the existing technology is solved, and a more accurate analysis of the impact of vegetation coverage on aerosol concentration is achieved, supporting ecological protection and sustainable development.
Patent Information
- Application Number
- CN202411916762.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-12-24
AI Technical Summary
The lack of high-resolution LAI products in the prior art makes it difficult to conduct in-depth research on the relationship between vegetation coverage and aerosol concentration, especially in climate change monitoring and simulation, resulting in high misjudgment and uncertainty.
By acquiring high-resolution LAI data, long-term, high-consistent space-time continuous LAI products were produced, and combined with MODIS AOD aerosol data, the impact of vegetation coverage on aerosol concentration was analyzed by segmented quantile regression method, and the data was processed and verified using a random forest model and neural network model.
Improves the accuracy and reliability of LAI data, enables more accurate description of the complex relationship between vegetation coverage and aerosol concentration, reduces errors, provides more accurate parameter estimation and model prediction, and supports ecological conservation and sustainable development.
Smart Images

Figure CN119692201B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental climate protection, and particularly to a method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data. Background Art
[0002] At present, the research on vegetation coverage and particulate matter pollution based on remote sensing inversion mainly focuses on two-dimensional green biomass. Although some scholars have studied the quantitative relationship between vegetation coverage and AOD based on remote sensing inversion of particulate matter concentration, there is a lack of high-resolution LAI products at home and abroad, and few in-depth studies have been conducted on the relationship between leaf area index and aerosol pollution represented by AOD. Summary of the Invention
[0003] Based on this, in view of the problem of studying the relationship between vegetation coverage and aerosols due to the lack of high-resolution LAI products, it is necessary to provide a method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data.
[0004] A method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data, characterized in that the method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data includes the following steps:
[0005] S1. Obtain data and produce a long-time series, highly consistent, spatio-temporally continuous LAI product;
[0006] S2. Obtain MODIS AOD aerosol data and perform data processing together with the LAI of the same period;
[0007] S3. Adopt a piecewise quantile regression method to study the relationship between LAI and aerosol sections.
[0008] This application discloses a method for analyzing the impact of vegetation coverage on aerosol concentration using LAI data. Currently, long-term vegetation observation data is particularly important for monitoring and simulating the response and feedback of the vegetation system to climate change. The leaf area index (LAI) has been identified as a key climate variable by the Global Climate Observation System. Current LAI products have uncertainties and it is difficult to capture the LAI changes caused by land cover type changes. There is an urgent need for LAI products with higher spatial resolution in applications such as agriculture and hydrology. Therefore, this application uses high-resolution LAI data to evaluate and analyze the method of the impact of vegetation coverage on aerosol concentration in the atmosphere. First, a large amount of data is obtained from various platforms to produce long-term, highly consistent, spatio-temporally continuous LAI products. Through continuous multi-year data, the change trends of the entire life cycle of vegetation growth, development, and senescence can be captured. The long-term and high consistency of the data make the LAI values obtained at different time points comparable, avoiding misjudgment of vegetation change trends due to data differences, and greatly increasing the depth and breadth of LAI product applications. AOD aerosol data is collected using the MODIS sensor, and at the same time, high-resolution LAI data of the same period is used. By processing the two together, based on the reliability of the high-resolution LAI data, by analyzing the differences in AOD aerosols between vegetated areas and non-vegetated areas, it is possible to intuitively infer whether the aerosol concentration is related to the vegetation coverage. Finally, the segmented quantile regression method is used to analyze and study the LAI data and aerosols. Since the relationship between LAI and aerosols may not be a simple linear relationship, the segmented quantile regression method can capture this non-linear feature, more accurately describe the complex relationship between the two, better adapt to the characteristics of the data, reduce errors and uncertainties, and provide more accurate parameter estimation and model prediction. The method of using LAI data to analyze the impact of vegetation coverage on aerosol concentration can accurately quantify the growth state and coverage of vegetation, and further deeply understand whether vegetation affects the aerosol concentration in the atmosphere, revealing the complex interaction between the ecosystem and the atmospheric environment.
[0009] In one of the embodiments, the specific steps of S1 are as follows:
[0010] S11. Collection and preprocessing of data;
[0011] S12. Data purification, optimization and strict screening, including elements purity screening, outlier cleaning, and thin cloud removal steps;
[0012] S13. Due to the registration errors caused by the projection differences and spatial resolution differences between MODIS and Landsat, a method of expanding the spatial neighborhood window is proposed to select homogeneous samples to improve the uncertainty brought by the spatial registration error;
[0013] S14. Train a random forest model based on pure samples to generate a high-score LAI reference dataset covering the preset time and prediction range;
[0014] S15. Use the reference dataset in S14 and the PKU GIMMS NDVI product to construct a BP neural network model that takes into account the feature expression of rapid forest expansion, and use the neural network model to generate long-term LAI products for the preset area;
[0015] S16. Perform repair processing and verification release on the long-term LAI products produced in S15.
[0016] Through the collection and preprocessing of data, the first step in the study of LAI data and aerosols is achieved. The various data collected provide the basic materials for subsequent analysis and modeling. Without sufficient data, it is impossible to deeply understand the relationship between LAI and aerosols and their impacts on the environment and climate. Since there may be some values in the data that deviate significantly from the normal range, these outliers may be caused by measurement errors, emergencies, or interference during the data collection process. Therefore, the collected data needs to be purified and optimized. For example, through steps such as meta-purity screening, outlier cleaning, and thin cloud removal, these errors and mistakes can be identified and corrected, reducing the error level in the data, improving the accuracy of LAI data, and making the data closer to the real situation. Due to the registration errors caused by the projection differences and spatial resolution differences between MODIS and Landsat, the method of expanding the spatial neighborhood window is used to select homogeneous samples to improve the uncertainty brought by the spatial registration error and enhance the accuracy of data processing. Because the optimized dataset can better represent the target population or phenomenon, the random forest model trained based on these data has stronger generalization ability, can perform well on new and unseen data, and can generate a high-score LAI reference dataset covering the preset time and prediction range. Combining these reference datasets with the PKU GIMMS NDVI product can construct a BP neural network model that takes into account the feature expression of rapid forest expansion. According to the BP neural network model, long-term products for China can be generated. By performing repair processing and verification on this product, it can be released for use, providing strong support for ecological protection and sustainable development and making contributions to global ecological protection.
[0017] In one of the embodiments, the specific steps of S11 are as follows:
[0018] S111. Rely on the GEE platform to download the Landsat reflectance data and MODIS LAI data required for the project;
[0019] S112. Perform preprocessing steps on the data in S111, including quality control, sample purification, and outlier cleaning;
[0020] S113. Obtain remote sensing data, including GLASS GLC land cover and PKU GIMMS NDVI;
[0021] S114. Perform preprocessing steps such as projection conversion, temporal synthesis, and spatial aggregation on the remote sensing data in S113.
[0022] By downloading Landsat reflectance data and MODIS LAI data on the Google Earth Engine (GEE) platform and performing preprocessing steps such as quality control, sample purification, and outlier cleaning on them, it provides crucial, correct, and useful data for the production of LAI products. Since Landsat reflectance data can be used to calculate vegetation indices, thereby quantitatively evaluating the growth status, coverage, biomass, etc. of vegetation, this is very helpful for studying aspects such as the productivity of ecosystems, seasonal changes in vegetation, the health status of ecosystems, and the impact of climate change on vegetation. And MODIS LAI data provides large-area and continuous LAI information, which can reflect the growth state and health degree of vegetation. By analyzing LAI data, the growth stage, growth rate, biomass accumulation, etc. of vegetation can be understood, and the long-term dynamic change trend of vegetation can be understood. These data provide great help for the production of LAI products. In addition to the above data, remote sensing data including GLASS GLC land cover and PKU GIMMS NDVI, etc. are also obtained, and preprocessing steps such as projection conversion, temporal synthesis, and spatial aggregation are performed on these remote sensing data. Remote sensing data can monitor the distribution and coverage of vegetation over a large area quickly, providing data support for vegetation resource inventory and management, and greatly improving the integrity of the collected data.
[0023] In one of the embodiments, the specific steps of "training a random forest model based on pure samples" in S14 are as follows:
[0024] S141. Evaluate the spatio-temporal representativeness of the random forest model training samples and supplement samples in special periods and special regions;
[0025] S142. Introduce MODIS saturated LAI pixels to train the random forest model and test the impact of sample balance on the model accuracy to achieve the purpose of improving the accuracy of the random forest model.
[0026] By evaluating the spatio-temporal representativeness of the training samples of the random forest model, it is possible to determine whether the training samples can comprehensively and accurately reflect the characteristics and changes of the research object at different times and in different spaces, avoid inaccurate prediction results, and supplement the samples in special periods and special regions, which can enable the model to better capture the patterns and laws in various situations, improve the accuracy of the model. Reasonable sample selection and supplementation can reduce the computational cost of the model, improve the training and prediction efficiency of the model, and samples with good spatio-temporal representativeness can make the model more stable, reducing the risks of overfitting and underfitting. By introducing MODIS saturated LAI pixels to train the random forest model, the diversity of the training data is increased, and the prediction accuracy of the model in high vegetation coverage areas is improved. The impact of sample balance on the model accuracy cannot be ignored either. When dealing with unbalanced data sets, the random forest, due to its integrated characteristics based on multiple decision trees, has better robustness compared to other models. However, when the class distribution of the data set is extremely unbalanced, even the random forest may tend to the majority class, thus ignoring the characteristics of the minority class. Therefore, adopting sample balance techniques can help the model better learn the characteristics of all classes, improve the prediction accuracy of the model for the minority class, and thus achieve the purpose of improving the accuracy of the random forest model.
[0027] In one of the embodiments, the specific steps of S16 are as follows:
[0028] S161. Develop a long short-term memory neural network algorithm to perform spatio-temporal repair on the long-time series LAI products in the preset area, and use a variety of product verification methods to systematically evaluate the uncertainty of the produced LAI products;
[0029] S162. Release the long-time series, high-consistency spatio-temporally continuous LAI products in the preset area and their quality control documents after product verification and quality assessment, and at the same time release the 30m Landsat LAI and 10m Sentinel LAI inversion models in the preset area online on the GEE platform.
[0030] By developing the long short-term memory neural network algorithm, it is possible to predict the LAI value in the future time period based on the existing LAI time series data according to the spatio-temporal correlation of the data, and to perform spatio-temporal repair of the long-time series LAI products within the preset area. Developing the long short-term memory neural network algorithm can smooth the LAI time series, reduce data fluctuations caused by sensor noise or environmental factors, improve data quality, and provide support for further data analysis and ecological research. Then, using a variety of product verification methods to systematically evaluate the uncertainty of the produced LAI products can identify errors and uncertainties in the LAI products, provide a basis for further optimizing the algorithm and improving data processing. At the same time, comprehensively evaluating the uncertainty helps to enhance users' trust in the LAI products and ensure the reliability of the data in scientific research and practical applications. Release the long-time series, high-consistency spatio-temporally continuous LAI products within the preset area and their quality control documents after product verification and quality assessment. At the same time, release the 30m Landsat LAI and 10m Sentinel LAI inversion models within the preset area on the GEE platform. Using the powerful computing power and user interface of GEE, global users can conveniently access and process LAI data. These data and tools can not only support the research on aerosols, but also provide help in many fields such as ecology, agriculture, environmental science, and climate change. At the same time, release the inversion model code to the GEE platform to encourage other researchers to use, modify, and expand these models, promoting scientific collaboration and knowledge sharing.
[0031] In one of the embodiments, the specific steps of S2 are as follows:
[0032] S21. Obtain MODIS AOD aerosol data. After operations such as cropping and calibration using ENVI software, use ENVI modeler and LUTS table to invert aerosol AOD, and use machine learning methods to improve the accuracy.
[0033] S22. Select the LAI data in the same period as the aerosol data. Based on the urban planning management unit, use ArcGIS software to perform fishnet segmentation on the LAI and AOD data respectively.
[0034] S23. Statistically analyze the attribute data of each urban planning management unit of the LAI and AOD data, match the data of the same management unit in pairs, and then export it to Excel and save it as a CSV format file.
[0035] By obtaining MODIS AOD aerosol data and using ENVI software to perform operations such as cropping and calibration on it to remove unnecessary areas, the dataset becomes more targeted and refined. Cropping is used to extract data for the study area, while calibration ensures the geographical positioning accuracy of the data, providing a reliable basis for subsequent analysis. Then, by using ENVI modeler and LUTS table to invert aerosol AOD, the aerosol optical properties under different conditions can be more accurately simulated, thereby improving the accuracy of the inverted AOD. Finally, the accuracy of MODIS AOD aerosol data is further improved through machine learning methods. Based on the processed MODIS AOD aerosol data, LAI data for the same period is obtained. For these two types of data based on urban planning management units, using ArcGIS software to perform fishnet segmentation on LAI and AOD data respectively can divide these large-area datasets into smaller, more manageable and processable units, facilitating users to conduct spatial statistics and analysis. Then, the attribute data of each urban planning management unit of LAI and AOD data is statistically analyzed, and the data of the same management unit is pairwise matched, enabling a more intuitive observation of the impact of LAI data on AOD aerosol data. Finally, these data are exported to Excel and saved as a CSV format file, making it easy to transfer and use among multiple operating systems and different programs, facilitating data sharing among different users or systems.
[0036] In one embodiment, the specific steps of S3 are as follows:
[0037] S31. Confirm the number of segments in the segmented quantile regression method and screen the segments;
[0038] S32. Segment the LAI and establish a quantile regression model for each segment based on the segmented quantile regression method. Select appropriate quantiles and use the quantile regression method to estimate the parameters;
[0039] S33. Extract the boundary line around the "scatter cloud", and determine whether the boundary line belongs to the upper bound or lower bound type according to the shape of the scatter cloud;
[0040] S34. To eliminate the influence of outliers, use the MATLAB program to obtain the 95% quantiles of each part, and try 10%, 25%, 50%, 75%, and 90% as boundary points respectively, and use Origin 8 software to fit these boundary points to obtain the constraint line;
[0041] S35. Identify the type of the constraint line in S34 according to the fitted R2 and analyze, study and utilize it;
[0042] S36. Obtain the quantile regression model according to the type of the constraint line in S35 and estimate and evaluate this model.
[0043] By segmenting the LAI and determining the number of segments in the segmented quantile regression method, the non-linear relationships and complex patterns in the data can be better captured. Reasonably selecting the number of segments can avoid the overfitting problem caused by an overly complex model. At the same time, the number of segments also affects the interpretability and prediction accuracy of the quantile regression model. A quantile regression model is established for each LAI segment, appropriate quantiles are selected, and the parameters are estimated using the quantile regression method. Firstly, quantile regression provides a robust method to estimate the central tendency and dispersion of the data. Secondly, different quantiles can help reveal the influencing factors of the LAI data changes under different environmental conditions. By comparing the regression results of different quantiles, the robustness of the model to different data subsets can be evaluated. In the LAI data segmentation, the boundary line around the "scatter cloud" is extracted. According to the shape of the scatter cloud, it is determined whether the boundary line belongs to the upper bound or lower bound type, which is convenient for better confirming the mathematical equation and helping to establish the quantile regression model. The MATLAB program is used to obtain the 95% quantiles of each part to eliminate the influence of outliers, and then the Origin 8 software is used to fit these boundary points to obtain the constraint line. The obtained constraint line is analyzed and studied, and a quantile regression model is established. Finally, the model is estimated and evaluated, and the reasons for the impact of LAI data on the reduction of aerosol concentration are found.
[0044] In one of the embodiments, the specific steps of S31 are as follows:
[0045] S311. Preset the number of LAI segments;
[0046] S312. Data exploration: Search for outliers or natural groupings of the data based on box plot analysis, or observe the distribution characteristics of the data using histograms and cumulative distribution functions;
[0047] S313. Determine the number of segments based on the existing knowledge and experience;
[0048] S314. Model selection criterion: Use statistical criteria to assist in determining the number of segments;
[0049] S315. Cross-validation: Perform cross-validation at different numbers of segments and select the number of segments with the best prediction performance;
[0050] S316. Model complexity and interpretability trade-off: Increasing the number of segments can improve the complexity of the model and may obtain a better fit, but it may also lead to an overly complex model that is difficult to interpret. It is necessary to balance between the complexity and interpretability of the model;
[0051] S317. Stepwise regression: Determine the optimal number of segments through the stepwise regression method. By gradually adding or deleting segments, evaluate the performance changes of the model;
[0052] S318, Model Diagnosis: Diagnose models with different numbers of segments to check for violations of model assumptions, including the distribution of residuals and heteroscedasticity.
[0053] Segmenting the LAI data is a very important step. Segmenting helps analyze the LAI data at different scales, from local to global, to better understand the scale effect. The following are the steps and methods for segmentation. First, preset the number of segments based on existing knowledge and experience. Furthermore, the number of segments can be confirmed through data exploration, model selection criteria, and cross-validation. During the segmentation process, it should be noted that the number of segments should not be too large. Although a large number of segments can increase the complexity of the model and may obtain a better fit, it may also lead to an overly complex model that is difficult to interpret. A balance needs to be struck between the complexity and interpretability of the model. The optimal number of segments can also be confirmed through stepwise regression. Finally, diagnose models with different numbers of segments to check for violations of model assumptions, including the distribution of residuals and heteroscedasticity, to ensure that the selected model is applicable to the specific segment of the data and to check whether the model can reasonably capture the characteristics of the data.
[0054] In one embodiment, the specific steps of S35 are as follows:
[0055] S351, Study the shape and characteristics of the constraint line, including convex wave type, hump type, or exponential type. Analyze the specific constraint relationships for aerosol reduction in different LAI sections through the characteristics, use it as an evaluation tool for LAI aerosol reduction, and extract the threshold when there is a threshold.
[0056] S352, Identify and analyze the main sections that affect the constraint line and understand how these sections affect the interaction and constraint effect of aerosols.
[0057] S353, Use the constraint line method to simulate the potential changes of aerosols and predict the response of aerosol concentration under different section scenarios.
[0058] S354, Evaluate the scale dependence of the constraint line method to ensure the effectiveness and applicability of the analysis results at different scales.
[0059] By studying the shape and characteristics of the constraint line, and analyzing the specific constraint relationships for aerosol reduction in different LAI sections based on these characteristics, identifying the constraint relationships for aerosol reduction in different LAI sections helps to formulate targeted environmental improvement measures, and can also evaluate the potential impact of vegetation coverage on aerosol concentration, thereby understanding the ecological effect of vegetation on air quality. Using the constraint line as an evaluation tool for LAI to reduce aerosols helps to confirm the threshold between vegetation coverage and aerosol reduction, that is, the aerosol concentration may change significantly when reaching a certain LAI value. By identifying and analyzing the main sections that affect the constraint line, and using the constraint line method to simulate the potential changes in aerosols, it can help us understand which factors have a significant impact on aerosol concentration and distribution, and understand how these sections affect the interaction and constraint effect of aerosols, which helps to reveal how to affect aerosol concentration and distribution through ecological processes. Identifying the aerosol concentration in different sections also helps to evaluate the air quality risks in different regions and provide a basis for formulating risk mitigation measures, which helps to achieve environmental protection and sustainable development goals. However, when using the constraint line, attention should also be paid to the scale dependence of the constraint line method, which helps to understand the differences in ecological processes and patterns at different scales and ensure the effectiveness and applicability of the analysis results at different scales.
[0060] In one embodiment, the specific steps of "estimating and evaluating the model" in S36 are as follows:
[0061] S361. Model estimation: Use MATLAB to estimate the parameters of the quantile regression model for each segment;
[0062] S362. Model evaluation: Evaluate the goodness of fit and predictive ability of each quantile regression model, including residual inspection and significance testing of the model.
[0063] By using MATLAB to estimate the parameters of the quantile regression model for each segment, MATLAB can efficiently reveal the conditional distribution characteristics of LAI at different quantile levels, provide more comprehensive information, and can also accurately estimate the parameters of the quantile regression model, reduce the influence of outliers and extreme values, and correctly reflect the conditional distribution characteristics at different quantiles. At the same time, using MATLAB can quickly segment and perform regression analysis on a large amount of data, improving work efficiency. Evaluating the goodness of fit and predictive ability of each quantile regression model can diagnose whether the model meets the basic assumptions through evaluation, ensure that the model can accurately capture the relationship between LAI and aerosol concentration, and help select the most suitable model for LAI data. Brief Description of the Drawings
[0064] Figure 1 It is a schematic flow chart of a method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data;
[0065] Figure 2 Schematic diagram of the process for obtaining data and producing long-time-series, highly consistent spatio-temporally continuous LAI products;
[0066] Figure 3 Schematic diagram of the process for data collection and preprocessing;
[0067] Figure 4 Schematic diagram of the process for training a random forest model based on pure samples;
[0068] Figure 5 Schematic diagram of the process for patching and validating / releasing the produced long-time-series LAI products;
[0069] Figure 6 Schematic diagram of the process for obtaining MODIS AOD aerosol data and processing the data together with the contemporaneous LAI;
[0070] Figure 7 Schematic diagram of the process for studying the relationship between LAI and aerosol segments using the piecewise quantile regression method;
[0071] Figure 8 Schematic diagram of the process for confirming the number of segments in the piecewise quantile regression method and screening the segments;
[0072] Figure 9 For fitting R 2 Schematic diagram of the process for identifying the type of the constraint line of S34 and analyzing, studying, and utilizing it;
[0073] Figure 10 Schematic diagram of the process for estimating and evaluating the quantile regression model. Detailed implementation manners
[0074] In order to more clearly understand the above objects, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0075] Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0076] The following describes a method for analyzing the influence of vegetation coverage on aerosol concentration by LAI data according to some embodiments of the present invention with reference to the accompanying drawings.
[0077] Embodiment
[0078] AsFigures 1 to 10 As shown in the figure, this embodiment discloses a method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data, including the following steps:
[0079] S1. Obtain data and produce a long-time-series, highly consistent, spatio-temporally continuous LAI product;
[0080] S2. Obtain MODIS AOD aerosol data and perform data processing together with the LAI of the same period;
[0081] S3. Adopt the segmented quantile regression method to study the relationship between LAI and aerosol segments.
[0082] This application discloses a method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data. Currently, long-time-series vegetation observation data is particularly important for monitoring and simulating the response and feedback of the vegetation system to climate change. The leaf area index (LAI) has been determined as a key climate variable by the Global Climate Observation System. The current LAI products have uncertainties and it is difficult to capture the LAI changes caused by land cover type changes. There is an urgent need for LAI products with higher spatial resolution in applications such as agriculture and hydrology. Therefore, this application uses high-resolution LAI data to evaluate and analyze the method of the influence of vegetation coverage on the aerosol concentration in the atmosphere. First, a large amount of data is obtained from various platforms to produce a long-time-series, highly consistent, spatio-temporally continuous LAI product. Through continuous multi-year data, the change trends of the entire life cycle of vegetation growth, development, and senescence can be captured. The long-time-series and high consistency of the data make the LAI values obtained at different time points comparable, avoiding misjudgment of the vegetation change trend due to data differences, and greatly increasing the depth and breadth of the application of LAI products. The AOD aerosol data is collected by using the MODIS sensor, and at the same time, the high-resolution LAI data of the same period is used. The two are processed together. According to the reliability of the high-resolution LAI data, by analyzing the differences in AOD aerosols between vegetated areas and non-vegetated areas, it can be intuitively inferred whether the aerosol concentration is related to the vegetation coverage degree. Finally, the segmented quantile regression method is used to analyze and study the LAI data and aerosols. Since the relationship between LAI and aerosols may not be a simple linear relationship, the segmented quantile regression method can capture this non-linear characteristic, more accurately describe the complex relationship between the two, better adapt to the characteristics of the data, reduce errors and uncertainties, and provide more accurate parameter estimation and model prediction. The method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data can accurately quantify the growth state and coverage degree of vegetation, and further deeply understand whether vegetation affects the aerosol concentration in the atmosphere, revealing the complex interaction relationship between the ecosystem and the atmospheric environment.
[0083] Such as Figure 2As shown, in addition to the features of the above embodiments, this embodiment further defines that the specific steps of S1 are as follows:
[0084] S11. Data collection and preprocessing work;
[0085] S12. Data purification and optimization, and strict screening, including the steps of primary purity screening, outlier cleaning, and thin cloud removal;
[0086] S13. Due to the projection differences and spatial resolution differences between MODIS and Landsat, resulting in registration errors, a method of expanding the spatial neighborhood window is proposed to select homogeneous samples to improve the uncertainty brought by the spatial registration errors;
[0087] S14. Training a random forest model based on pure samples to produce a high - resolution LAI reference dataset covering the preset time and prediction range;
[0088] S15. Using the reference dataset in S14 and the PKU GIMMS NDVI product to construct a BP neural network model considering the feature expression of rapid forest expansion, and using the neural network model to produce long - time - series LAI products for the preset area;
[0089] S16. Repairing and validating the long - time - series LAI products produced in S15 and then releasing them.
[0090] Through the collection and preprocessing of data, the first step in the research on LAI data and aerosols is achieved. The various data collected provide the basic materials for subsequent analysis and modeling. Without sufficient data, it is impossible to deeply understand the relationship between LAI and aerosols and their impact on the environment and climate. Since there may be some values in the data that deviate significantly from the normal range, these outliers may be caused by measurement errors, emergencies, or interference during the data collection process. Therefore, it is necessary to purify and optimize the collected data. For example, through steps such as meta-purity screening, outlier cleaning, and thin cloud removal, these errors and mistakes can be identified and corrected, reducing the error level in the data, improving the accuracy of LAI data, and making the data closer to the real situation. Due to the registration errors caused by the projection differences and spatial resolution differences between MODIS and Landsat, the method of expanding the spatial neighborhood window is used to select homogeneous samples to improve the uncertainty caused by spatial registration errors and enhance the accuracy of data processing. Since the optimized dataset can better represent the target population or phenomenon, the random forest model trained based on these data has stronger generalization ability, can perform well on new and unseen data, and can produce a high-quality LAI reference dataset covering the preset time and prediction range. Combining these reference datasets with the PKU GIMMS NDVI product can construct a BP neural network model that takes into account the characteristics of rapid forest expansion. According to the BP neural network model, a long-term time series product for China can be produced. After repairing and validating this product, it can be released for use, providing strong support for ecological protection and sustainable development and making contributions to global ecological protection.
[0091] As Figure 3 shown, in addition to the features of the above embodiments, this embodiment further defines that the specific steps of S11 are as follows:
[0092] S111. Rely on the GEE platform to download the Landsat reflectance data and MODIS LAI data required for the project;
[0093] S112. Perform preprocessing steps including quality control, sample purification, and outlier cleaning on the data in S111;
[0094] S113. Obtain remote sensing data, including GLASS GLC land cover and PKU GIMMS NDVI;
[0095] S114. Perform preprocessing steps such as projection conversion, time synthesis, and spatial aggregation on the remote sensing data in S113.
[0096] By downloading Landsat reflectance data and MODIS LAI data on the Google Earth Engine (GEE) platform and performing preprocessing steps such as quality control, sample purification, and outlier cleaning, it provides crucial, correct, and useful data for the production of LAI products. Since Landsat reflectance data can be used to calculate vegetation indices, thereby quantitatively evaluating the growth status, coverage, biomass, etc. of vegetation, this is very helpful for studying aspects such as the productivity of ecosystems, seasonal changes in vegetation, the health status of ecosystems, and the impact of climate change on vegetation. MODIS LAI data provides large-area and continuous LAI information, which can reflect the growth state and health of vegetation. By analyzing the LAI data, it is possible to understand the growth stage, growth rate, biomass accumulation, etc. of vegetation, and the long-term dynamic change trend of vegetation. These data provide great help for the production of LAI products. In addition to the above data, remote sensing data such as GLASS GLC land cover and PKUGIMMS NDVI were also obtained, and preprocessing steps such as projection conversion, temporal synthesis, and spatial aggregation were performed on these remote sensing data. Remote sensing data can rapidly monitor the distribution and coverage of vegetation over a large area, providing data support for vegetation resource inventory and management, and greatly improving the integrity of the collected data.
[0097] As Figure 4 shown, in addition to the features of the above embodiments, this embodiment further defines that the specific steps of "training a random forest model based on pure samples" in S14 are as follows:
[0098] S141. Evaluate the spatio-temporal representativeness of the random forest model training samples and supplement samples for special periods and special regions;
[0099] S142. Introduce MODIS saturated LAI pixels to train the random forest model and test the impact of sample balance on the model accuracy to achieve the purpose of improving the accuracy of the random forest model.
[0100] By evaluating the spatio-temporal representativeness of the training samples of the random forest model, it is possible to determine whether the training samples can comprehensively and accurately reflect the characteristics and changes of the research object at different times and in different spaces, avoid inaccurate prediction results, and supplement samples in special periods and special regions, which can enable the model to better capture patterns and regularities in various situations, improve the accuracy of the model. Reasonable sample selection and supplementation can reduce the computational cost of the model, improve the training and prediction efficiency of the model, and samples with good spatio-temporal representativeness can make the model more stable and reduce the risks of overfitting and underfitting. By introducing MODIS saturated LAI pixels to train the random forest model, the diversity of training data is increased, and the prediction accuracy of the model in high vegetation coverage areas is improved. The impact of sample balance on model accuracy cannot be ignored either. When dealing with unbalanced datasets, due to its integrated characteristics based on multiple decision trees, the random forest has better robustness compared to other models. However, when the class distribution of the dataset is extremely unbalanced, even the random forest may tend to the majority class and thus ignore the characteristics of the minority class. Therefore, adopting sample balance techniques can help the model better learn the characteristics of all classes, improve the prediction accuracy of the model for the minority class, and thus achieve the purpose of improving the accuracy of the random forest model.
[0101] As Figure 5 shown, in addition to the features of the above embodiments, this embodiment further defines that the specific steps of S16 are as follows:
[0102] S161. Develop a long short-term memory neural network algorithm to perform spatio-temporal repair on long-time series LAI products in a preset area, and use a variety of product verification methods to systematically evaluate the uncertainty of the produced LAI products;
[0103] S162. Release long-time series, high-consistency spatio-temporally continuous LAI products in the preset area and their quality control documents after product verification and quality assessment, and at the same time release the inversion models of 30m Landsat LAI and 10m Sentinel LAI in the preset area online on the GEE platform.
[0104] By developing the long short-term memory neural network algorithm, based on the existing LAI time series data, it is possible to predict the LAI values in the future time period according to the spatio-temporal correlation of the data, and to perform spatio-temporal repair of the long-time series LAI products within the preset area. Developing the long short-term memory neural network algorithm can smooth the LAI time series, reduce data fluctuations caused by sensor noise or environmental factors, improve data quality, and provide support for further data analysis and ecological research. Then, using a variety of product verification methods to systematically evaluate the uncertainty of the produced LAI products, it is possible to identify errors and uncertainties in the LAI products, provide a basis for further optimizing the algorithm and improving data processing. At the same time, comprehensively evaluating the uncertainty helps to enhance users' trust in the LAI products and ensure the reliability of the data in scientific research and practical applications. Release long-time series, highly consistent spatio-temporally continuous LAI products within the preset area and their quality control documents after product verification and quality assessment. At the same time, release the 30m Landsat LAI and 10m Sentinel LAI inversion models within the preset area on the GEE platform. Using the powerful computing power and user interface of GEE, global users can conveniently access and process LAI data. These data and tools can not only support the research on aerosols, but also provide help for multiple fields such as ecology, agriculture, environmental science, and climate change. At the same time, release the inversion model code to the GEE platform to encourage other researchers to use, modify, and expand these models, promoting scientific collaboration and knowledge sharing.
[0105] As Figure 6 shown, in addition to the features of the above embodiments, this embodiment further defines: The specific steps of S2 are as follows:
[0106] S21. Obtain MODIS AOD aerosol data. After operations such as cropping and calibration using ENVI software, use ENVI modeler and LUTS table to invert aerosol AOD, and use machine learning methods to improve the accuracy;
[0107] S22. Select LAI data in the same period as the aerosol data. Based on the urban planning management unit, use ArcGIS software to perform fishnet segmentation on LAI and AOD data respectively;
[0108] S23. Statistically analyze the attribute data of each urban planning management unit of LAI and AOD data, match the data of the same management unit in pairs, and then export it to Excel and save it as a CSV format file.
[0109] By obtaining MODIS AOD aerosol data and using ENVI software to perform operations such as cropping and calibration on it to remove unnecessary areas, the dataset becomes more targeted and refined. The cropping is used to extract data for the study area, while the calibration ensures the geographical positioning accuracy of the data, providing a reliable basis for subsequent analysis. Then, by using ENVI modeler and LUTS table to invert aerosol AOD, the aerosol optical properties under different conditions can be more accurately simulated, thereby improving the accuracy of the inverted AOD. Finally, the accuracy of MODIS AOD aerosol data is further enhanced through machine learning methods. Based on the processed MODIS AOD aerosol data, LAI data of the same period is obtained. For these two types of data based on urban planning management units, using ArcGIS software to perform fishnet segmentation on LAI and AOD data respectively, these large-area datasets can be divided into smaller, more manageable and processable units, facilitating users to conduct spatial statistics and analysis. Then, the attribute data of each urban planning management unit of LAI and AOD data is statistically analyzed, and the data of the same management unit is pairwise matched, enabling a more intuitive observation of the impact of LAI data on AOD aerosol data. Finally, these data are exported to Excel and saved as a CSV format file, making it easy to transfer and use among various operating systems and different programs, facilitating data sharing among different users or systems.
[0110] As Figure 7 shown, in addition to the features of the above embodiments, this embodiment further defines that the specific steps of S3 are as follows:
[0111] S31. Confirm the number of segments in the piecewise quantile regression method and screen the segments;
[0112] S32. Segment LAI and establish a quantile regression model for each segment based on the piecewise quantile regression method, select appropriate quantiles, and use the quantile regression method to estimate parameters;
[0113] S33. Extract the boundary line around the "scatter cloud", and determine whether the boundary line belongs to the upper bound or lower bound type according to the shape of the scatter cloud;
[0114] S34. To eliminate the influence of outliers, use the MATLAB program to obtain the 95% quantiles of each part, respectively try 10%, 25%, 50%, 75%, 90% as boundary points, and use Origin 8 software to fit these boundary points to obtain the constraint line;
[0115] S35. According to the fitted R 2 Identify the type of the constraint line in S34 and analyze, study and utilize it;
[0116] S36. Obtain a quantile regression model based on the constraint line type in S35, and estimate and evaluate this model.
[0117] By segmenting LAI and determining the number of segments in the segmented quantile regression method, the non - linear relationships and complex patterns in the data can be better captured. Reasonably choosing the number of segments can avoid the over - fitting problem caused by an overly complex model. At the same time, the number of segments also affects the interpretability and prediction accuracy of the quantile regression model. Establish a quantile regression model for each LAI segment, select appropriate quantiles, and use the quantile regression method to estimate the parameters. First, quantile regression provides a robust method to estimate the central tendency and dispersion of the data. Second, different quantiles can help reveal the influencing factors of LAI data changes under different environmental conditions. By comparing the regression results of different quantiles, the robustness of the model to different data subsets can be evaluated. Extract the boundary lines around the "scatter cloud" in the LAI data segmentation. According to the shape of the scatter cloud, determine whether the boundary line belongs to the upper - bound or lower - bound type, which is convenient for better confirming the mathematical equation and helps to establish a quantile regression model. Use the MATLAB program to obtain the 95% quantiles of each part to eliminate the influence of outliers, and then use the Origin 8 software to fit these boundary points to obtain the constraint line. Analyze and study the obtained constraint line and establish a quantile regression model. Finally, estimate and evaluate this model, and find the reasons for the influence of LAI data on aerosol concentration reduction.
[0118] As Figure 8 shown, in addition to the features of the above - mentioned embodiments, this embodiment further defines that the specific steps of S31 are as follows:
[0119] S311. Preset the number of LAI segments;
[0120] S312. Data exploration: Analyze the box plot to find outliers or natural groupings of the data, or use the histogram and cumulative distribution function to observe the distribution characteristics of the data;
[0121] S313. Determine the number of segments according to the existing knowledge and experience;
[0122] S314. Model selection criterion: Use statistical criteria to assist in determining the number of segments;
[0123] S315. Cross - validation: Perform cross - validation on different numbers of segments and select the number of segments with the best prediction performance;
[0124] S316. Trade - off between model complexity and interpretability: Increasing the number of segments can improve the complexity of the model and may obtain a better fit, but it may also lead to an overly complex model that is difficult to interpret. It is necessary to balance between the complexity and interpretability of the model;
[0125] S317, Stepwise regression: Determine the optimal number of segments through stepwise regression. By gradually adding or deleting segments, evaluate the performance changes of the model.
[0126] S318, Model diagnosis: Diagnose models with different numbers of segments to check for violations of model assumptions, including the distribution of residuals and heteroscedasticity.
[0127] Segmenting the LAI data is a very important step. Segmenting helps analyze the LAI data at different scales, from local to global, and better understand the scale effect. The following are the steps and methods for segmentation. First, preset the number of segments based on existing knowledge and experience. Furthermore, the number of segments can be confirmed through data exploration, model selection criteria, and cross-validation. During the segmentation process, attention should be paid that the number of segments should not be too large. Although too many segments can increase the complexity of the model and may obtain a better fitting degree, it may also lead to the model being too complex to explain. A balance needs to be struck between the complexity and interpretability of the model. The optimal number of segments can also be confirmed through stepwise regression. Finally, diagnose models with different numbers of segments to check for violations of model assumptions, including the distribution of residuals and heteroscedasticity, to ensure that the selected model is applicable to the specific segments of the data and to check whether the model can reasonably capture the characteristics of the data.
[0128] As Figure 9 shown, in addition to the features of the above embodiments, this embodiment further defines that the specific steps of S35 are as follows:
[0129] S351, Study the shape and characteristics of the constraint line, including convex wave type, hump type, or exponential type. Analyze the specific constraint relationships for reducing aerosols in different LAI sections through the characteristics, use it as an evaluation tool for LAI to reduce aerosols, and extract the threshold when there is a threshold.
[0130] S352, Identify and analyze the main sections that affect the constraint line, and understand how these sections affect the interaction and constraint effect of aerosols.
[0131] S353, Use the constraint line method to simulate the potential changes of aerosols and predict the response of aerosol concentration in different section scenarios.
[0132] S354, Evaluate the scale dependence of the constraint line method to ensure the effectiveness and applicability of the analysis results at different scales.
[0133] By studying the shape and characteristics of the constraint line, and analyzing the specific constraint relationships for aerosol reduction in different LAI sections based on these characteristics, identifying the constraint relationships for aerosol reduction in different LAI sections helps to formulate targeted environmental improvement measures, and can also evaluate the potential impact of vegetation coverage on aerosol concentration, thereby understanding the ecological effect of vegetation on air quality. Using the constraint line as an evaluation tool for LAI to reduce aerosols helps to confirm the threshold between vegetation coverage and aerosol reduction, that is, the aerosol concentration may change significantly when reaching a certain LAI value. By identifying and analyzing the main sections that affect the constraint line, and using the constraint line method to simulate the potential changes in aerosols, it can help us understand which factors have a significant impact on aerosol concentration and distribution, understand how these sections affect the interaction and constraint effect of aerosols, which helps to reveal how to affect aerosol concentration and distribution through ecological processes. Identifying the aerosol concentration in different sections also helps to evaluate the air quality risks in different regions, and provides a basis for formulating risk mitigation measures, which helps to achieve environmental protection and sustainable development goals. However, when using the constraint line, attention should also be paid to the scale dependence of the constraint line method, which helps to understand the differences in ecological processes and patterns at different scales, and ensure the effectiveness and applicability of the analysis results at different scales.
[0134] As Figure 10 shown, in addition to the features of the above embodiments, this embodiment further defines that the specific steps of "estimating and evaluating the model" in S36 are as follows:
[0135] S361. Model estimation: Use MATLAB to estimate the parameters of the quantile regression model for each segment;
[0136] S362. Model evaluation: Evaluate the goodness of fit and predictive ability of each quantile regression model, including residual inspection and significance testing of the model.
[0137] By using MATLAB to estimate the parameters of the quantile regression model for each segment, MATLAB can efficiently reveal the conditional distribution characteristics of LAI at different quantile levels, provide more comprehensive information, and can also accurately estimate the parameters of the quantile regression model, reduce the influence of outliers and extreme values, and correctly reflect the conditional distribution characteristics at different quantiles. At the same time, using MATLAB can quickly segment and perform regression analysis on a large amount of data, improving work efficiency. Evaluate the goodness of fit and predictive ability of each quantile regression model. Through evaluation, it can be diagnosed whether the model meets the basic assumptions, ensure that the model can accurately capture the relationship between LAI and aerosol concentration, and help select the most suitable model for LAI data.
[0138] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0139] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data, characterized in that, The method for analyzing the impact of vegetation coverage on aerosol concentration using LAI data includes the following steps: S1. Obtain data and produce a long-term, highly consistent, spatio-temporally continuous LAI product; S2. Obtain MODIS AOD aerosol data and perform data processing together with the LAI of the same period; S3. Use the segmented quantile regression method to study the relationship between LAI and aerosol segments; The specific steps of S3 are as follows: S31. Confirm the number of segments in the segmented quantile regression method and screen the segments; S32. Segment the LAI, and establish a quantile regression model for each segment based on the segmented quantile regression method, select appropriate quantiles, and use the quantile regression method to estimate the parameters; S33. Extract the boundary line around the "scatter cloud", and determine whether the boundary line belongs to the upper bound or lower bound type according to the shape of the scatter cloud; S34. To eliminate the influence of outliers, use the MATLAB program to obtain the 95% quantiles of each part, try 10%, 25%, 50%, 75%, and 90% as boundary points respectively, and use Origin 8 software to fit these boundary points to obtain the constraint line; S35. According to the fitted R 2 Identify the type of the constraint line of S34 and analyze, study, and utilize it; S36. Obtain the quantile regression model according to the constraint line type of S35, and estimate and evaluate the model; The specific steps of S35 are as follows: S351. Study the shape and characteristics of the constraint line, including convex wave type, hump type or exponential type, analyze the specific constraint relationship between different LAI segments for reducing aerosols through the characteristics, use it as an evaluation tool for LAI to reduce aerosols, and extract the threshold when there is a threshold; S352. Identify and analyze the main segments affecting the constraint line; S353. Use the constraint line method to simulate the potential changes of aerosols and predict the response of aerosol concentration under different segment scenarios; S354. Evaluate the scale dependence of the constraint line method.
2. The method for analyzing the influence of vegetation coverage on aerosol concentration by LAI data according to claim 1, wherein The specific steps of S1 are as follows: S11. Data collection and preprocessing work; S12. Data purification and optimization, and strict screening, including elements purity screening, outlier cleaning, and thin cloud removal steps; S13. Due to the projection difference and spatial resolution difference between MODIS and Landsat resulting in registration errors, the method of expanding the spatial neighborhood window is proposed to select homogeneous samples to improve the uncertainty caused by spatial registration errors; S14. Train a random forest model based on pure samples to produce a high-resolution LAI reference dataset covering the preset time and prediction range; S15. Use the reference dataset of S14 and the PKU GIMMS NDVI product to construct a BP neural network model considering the expression of the rapid expansion characteristics of forests, and use the neural network model to produce a long-term LAI product for the preset area; S16. Repair and verify the long-term LAI product produced in S15 and release it.
3. The method for analyzing the influence of vegetation coverage on aerosol concentration based on LAI data according to claim 2, wherein The specific steps of S11 are as follows: S111. Download the Landsat reflectance data and MODIS LAI data required for the project relying on the GEE platform; S112. Perform preprocessing steps on the data of S111, including quality control, sample purification, and outlier cleaning; S113. Obtain remote sensing data, including GLASS GLC land cover and PKU GIMMS NDVI; S114. Perform preprocessing steps of projection conversion, temporal synthesis, and spatial aggregation on the remote sensing data in S113.
4. The method for analyzing the influence of vegetation coverage on aerosol concentration based on LAI data according to claim 2, wherein The specific steps of "training a random forest model based on pure samples" in S14 are as follows: S141. Evaluate the spatio-temporal representativeness of the training samples of the random forest model and supplement the samples in special periods and special regions; S142. Introduce MODIS saturated LAI pixels to train the random forest model and test the impact of sample balance on the model accuracy to achieve the purpose of improving the accuracy of the random forest model.
5. The method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data according to claim 2, characterized in that, The specific steps of S16 are as follows: S161. Develop a long short-term memory neural network algorithm to perform spatio-temporal repair on long-time series LAI products in a preset area, and systematically evaluate the uncertainty of the produced LAI products using various product verification methods; S162. Release long-time series, highly consistent spatio-temporally continuous LAI products and their quality control documents in the preset area after product verification and quality assessment, and simultaneously release the inversion models of 30m Landsat LAI and 10m Sentinel LAI in the preset area online on the GEE platform.
6. The method for analyzing the influence of vegetation coverage on aerosol concentration by using LAI data according to claim 1, characterized in that, The specific steps of S2 are as follows: S21. Obtain MODIS AOD aerosol data. After cropping and calibration operations using ENVI software, invert aerosol AOD using ENVImodeler and LUTS tables, and use machine learning methods to improve the accuracy; S22. Select LAI data in the same period as the aerosol data. Based on urban planning management units, use ArcGIS software to perform fishnet segmentation on LAI and AOD data respectively; S23. Statistically analyze the attribute data of each urban planning management unit of LAI and AOD data, match the data of the same management unit in pairs, and then export it to Excel and save it as a CSV format file.
7. The method for analyzing the influence of vegetation coverage on aerosol concentration by LAI data according to claim 1, wherein The specific steps of S31 are as follows: S311. Preset the number of LAI segments; S312. Data exploration: Look for outliers or natural groupings of data according to box plot analysis, or observe the distribution characteristics of data using histograms and cumulative distribution functions; S313. Determine the number of segments based on existing knowledge and experience; S314. Model selection criterion: Use statistical criteria to assist in determining the number of segments; S315. Cross-validation: Perform cross-validation on different numbers of segments and select the number of segments with the best prediction performance; S316. Model complexity and interpretability trade-off: Increasing the number of segments can improve the complexity of the model and may obtain a better fit, but it may also make the model too complex to interpret. It is necessary to balance between the complexity and interpretability of the model; S317. Stepwise regression: Determine the best number of segments through stepwise regression methods. By gradually adding or deleting segments, evaluate the performance changes of the model; S318. Model diagnosis: Diagnose the models with different numbers of segments to check whether there are violations of model assumptions, including the distribution and heteroscedasticity of residuals.
8. The method for analyzing the influence of vegetation coverage on aerosol concentration according to the LAI data as claimed in claim 1, wherein The specific steps of "estimating and evaluating the model" in S36 are as follows: S361. Model estimation: Use MATLAB to estimate the parameters of the quantile regression model for each segment; S362. Model evaluation: Evaluate the goodness of fit and predictive ability of each quantile regression model, including residual inspection and significance testing of the model.
Citation Information
Patent Citations
Vegetation coverage change attribution method
CN115018127A
Forest disturbance monitoring method and device, computer equipment and storage medium
CN118038287A