Mountain area accumulated snow area reconstruction method based on machine learning
Through a machine learning-based method, the optical remote sensing data and terrain data are processed using the XGBoost model, which solves the limitations of obtaining high spatial and temporal resolution of snow area data in the prior art, and realizes the reconstruction of 30m high-resolution snow area, improving the accuracy and efficiency of monitoring.
Patent Information
- Application Number
- CN202510178291.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The prior art has limitations in obtaining high spatial and temporal resolution of snow area data, especially in mountainous areas, it is difficult to accurately monitor the distribution and changes of snow.
Using a machine learning-based approach, a sample data set is constructed by acquiring optical remote sensing data and terrain data for preprocessing, and training is performed using the XGBoost model to achieve reconstruction of 30m high-resolution snow cover area.
The reconstruction of snow-covered area with high spatial and temporal resolution is achieved, which improves the accuracy and efficiency of snow-covered monitoring, and can meet the needs of highly spatial heterogeneous snow distribution in mountainous areas and the time series continuity requirements of hydrological simulation.
Smart Images

Figure CN120125641A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of snow cover monitoring, and relates to, but is not limited to, a method for reconstructing the snow area in mountainous areas based on machine learning. Background Art
[0002] Terrestrial snow cover is the most widely distributed component of the cryosphere, the second largest component of the cryosphere, one of the most active elements on the Earth's surface, and also the variable that changes fastest seasonally in the cryosphere. Accurately monitoring the spatio-temporal changes of snow cover is of great significance for research on climate change, water resource management, disaster prevention, etc.
[0003] The snow area is one of the important variables of snow cover characteristics, which can reflect the spatio-temporal distribution of snow cover and is also an important input data for some hydrological models. There are mainly two means to obtain the snow cover area. One is the traditional snow observation means, which monitors the snow cover area through meteorological stations and snow measurement lines. However, restricted by the distribution of stations, it is impossible to accurately monitor the distribution and changes of snow cover in inaccessible mountainous areas. The other is satellite remote sensing monitoring. Satellite remote sensing, with its characteristics of fast speed, all-weather, large spatial scale, etc., has become the first choice for retrieving snow parameters on a large scale. Satellite remote sensing can obtain parameters such as the snow cover area in large areas and inaccessible places, greatly making up for the deficiencies of traditional observation means. Retrieving the snow cover area by remote sensing can be divided into optical remote sensing and microwave remote sensing. However, due to the low spatial resolution of microwave remote sensing, it limits its application in mountainous areas and other areas with complex terrain. Therefore, in the research of retrieving the snow cover area by remote sensing, most are through satellite optical remote sensing data. Hyperspectral and high spatio-temporal resolution optical remote sensing sensors can effectively identify the distribution range and temporal change trend of snow cover and have been widely used in snow area monitoring.
[0004] With the continuous progress and development of remote sensing technology, the inversion algorithms and products of the snow cover area have become mature. The forms of the snow cover area retrieved by remote sensing mainly include two types: binary snow area (Snow Cover Extent, SCE) and fractional snow cover (Fractional Snow Cover, FSC). Among them, the binary snow area is represented by 0 or 1 to indicate whether the pixel is a snow pixel, and the fractional snow cover uses a numerical range from 0 to 1 to represent the proportion of snow cover within the pixel.
[0005] Snow cover has strong spatio-temporal heterogeneity and is one of the large-scale features that change most rapidly on Earth. Traditional monitoring methods are difficult to effectively monitor the spatio-temporal changes of large-scale snow cover and cannot provide sufficient snow cover range data for watershed-scale hydrological models. With the development of remote sensing technology, it provides an effective solution for quickly and accurately obtaining large-scale snow cover. However, existing snow cover area products, including the Moderate-Resolution Imaging Spectroradiometer (MODIS), still have their inherent limitations. For example, for the MOD10A1 daily product, its relatively coarse spatial resolution (500m) cannot accurately depict the spatial details of mountain snow cover; for the Landsat series with better spatial resolution in the United States, its 16-day revisit cycle is difficult to capture the temporal phase characteristics of snowmelt changes. Limited by the sensor performance, it is often impossible to achieve both high temporal and spatial resolutions. It is stretched when facing the distribution of patchy snow cover in mountainous areas with high heterogeneity and the modeling and simulation of runoff snowmelt and snow water equivalent in alpine areas of mountain hydrological models, and cannot simultaneously meet the spatial requirements of the highly spatially heterogeneous snow cover distribution in mountainous areas and the temporal series continuity requirements of hydrological simulations.
[0006] Therefore, developing a high-precision snow cover area downscaling inversion algorithm with high spatio-temporal resolution to produce daily 30m snow cover area data has important research significance and value for mountain snow cover monitoring, hydrological process simulation, and ecosystem evolution. Summary of the Invention
[0007] In view of this, the embodiments of the present invention provide a method for reconstructing mountain snow cover area based on machine learning, which at least overcomes the limitations of the prior art in obtaining high spatio-temporal resolution snow cover area data.
[0008] The technical solution of the embodiments of the present invention is specifically as follows:
[0009] The embodiments of the present invention provide a method for reconstructing mountain snow cover area based on machine learning, and the method includes:
[0010] Obtain optical remote sensing data and terrain data and perform preprocessing; the optical remote sensing data includes MOD09GA data and Landsat8 OLI data; the preprocessing includes terrain correction and snow cover identification; use the snow cover area generated by the snow cover identification algorithm for Landsat8 OLI data as the true value label, and construct a sample data set based on the preprocessed MOD09GA data and terrain data; divide the sample data set into a training set and a validation set at a ratio of 4:1 to train and validate the XGBoost model; the XGBoost model improves the prediction performance by combining multiple decision trees; use the trained XGBoost model to predict the snow cover area of new input data.
[0011] In some embodiments, the MOD09GA data includes 7 surface reflectance bands, as well as 4 bands related to the observation state, namely the sensor zenith angle, the sensor azimuth angle, the solar zenith angle, and the solar azimuth angle; the terrain data includes elevation images with a resolution of 30 meters, slope data, and aspect data.
[0012] The acquisition of optical remote sensing data and terrain data and the preprocessing thereof include: performing terrain correction on the MOD09GA data and the Landsat8 OLI data; using a snow cover identification algorithm that combines three remote sensing indices, namely the Normalized Difference Snow Index (NDSI), the Normalized Difference Vegetation Index (NDVI), and the Normalized Difference Forest Snow Index (NDFSI), to identify snow cover on the Landsat8 OLI surface reflectance image, and obtaining a snow cover area image.
[0013] In some embodiments, the following terrain correction model is used during the terrain correction process:
[0014] L H (λ) = L I (λ) - a(λ) * (IC - IC H )
[0015]
[0016] In the formula, Z is the solar zenith angle, S is the slope, φ z is the solar azimuth angle, is the aspect, IC is defined as a function of the solar zenith angle, the slope, the solar azimuth angle, and the aspect, and its value range is from -1 to 1; L H (λ) is the pixel reflectance value of the λ band after correction, L I (λ) is the pixel reflectance value of the λ band before correction, a(λ) is obtained from the relationship between the original image radiance value L I (λ) and its corresponding IC, L I (λ) = a(λ) * IC + b(λ), and IC H is the IC of the horizontal plane, that is, the cosine of the solar zenith angle.
[0017] In some embodiments, the snow-covered area generated from Landsat8 OLI data using a snow recognition algorithm is used as the ground truth label. Based on the preprocessed MOD09GA data and terrain data, a sample dataset is constructed, including: cropping the MOD09GA image and elevation image based on the range of the Landsat8 OLI surface reflectance image at the same time of the day;
[0018] Interpolating and resampling the cropped images to the same 30-meter spatial resolution as the Landsat8 OLI image using the nearest neighbor interpolation method, and matching the number of rows and columns of all the images used; using the 30-meter snow-covered area image obtained by the snow recognition algorithm from the Landsat8 OLI surface reflectance image at the same time of the day as the ground truth label, and using the QA band of Landsat8 OLI to exclude cloud pixels; performing stitching, splitting, missing value filling, and data normalization operations on all the data unified to the 30-meter spatial resolution, and saving the processed data as an npy file.
[0019] In some embodiments, the method further includes: implementing the construction of the XGBoost model using the Jupyter Notebook software platform; during the model training process, separating the input data for training according to the row and column numbers of Landsat8 OLI; wherein, snow-covered area reconstruction rules applicable to different regions are established for training at the positions with different row and column numbers, and a preset number of valid pixels are set for training at the position of a single row and column number.
[0020] In some embodiments, the snow recognition algorithm that combines the three remote sensing indices of NDSI, NDVI, and NDFSI is used to recognize snow on the Landsat8 OLI surface reflectance image, including: in non-forest areas and forest areas, when NDSI is greater than 0.4 and the reflectance of the near-infrared band NIR is greater than 0.11, it is recognized as snow; in forest areas, when NDSI is between 0 and 0.4, and when NDFSI is greater than 0.4 and NDVI is less than 0.6, it is recognized as snow.
[0021] In some embodiments, the method further includes: selecting 12 scenes of Landsat8 OLI data that cover the entire study area as the test dataset; using MOD09GA and elevation data as inputs, and reconstructing the snow-covered area data through the trained XGBoost model; using the Landsat8 OLI surface reflectance of cloud-free images to generate the snow-covered area as the true data; the cloud amount of the cloud-free images is less than a specific value; comparing the reconstructed snow-covered area data with the true data to evaluate the performance of the trained XGBoost model.
[0022] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0023] In the embodiments of the present invention, MOD09GA and Landsat8 OLI surface reflectance data and terrain data are acquired. After data preprocessing, a sample data set is made, and a machine learning model based on the XGBoost algorithm is trained to directly reconstruct the snow cover area with a high resolution of 30 meters (m). This method uses optical remote sensing data with high temporal resolution and high spatial resolution as well as terrain data. Combining with the XGBoost algorithm can better learn the complex patterns and relationships in these data, and achieve more efficient and accurate prediction of the snow cover area and reconstruction of snow cover. Compared with the traditional downscaling method, the new method has higher accuracy, can achieve the downscaling effect of the snow cover area of 30m per day, and improves the ability of snow cover area monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where:
[0025] Figure 1 It is a schematic flow chart of the method for reconstructing the snow cover area in mountainous areas based on machine learning provided by the embodiments of the present invention;
[0026] Figure 2 It is a comparison chart of the changes in NDSI calculated from the Landsat8 OLI surface reflectance before and after terrain correction provided by the embodiments of the present invention;
[0027] Figure 3 It is a schematic diagram of the training process of the XGBoost model provided by the embodiments of the present invention;
[0028] Figure 4 It is a comparison chart of the true color synthesis and the snow cover reconstruction effect of XGBoost at row and column number 135033 provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0030] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0031] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present invention described here can be implemented in an order other than that illustrated or described here.
[0032] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the field to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0033] Machine learning can process high-dimensional, multi-feature data and can mine implicit relationships between different features in complex features. The present invention uses the Extreme Gradient Boosting (XGBoost) model with better performance to reconstruct 30m high-resolution snow area data.
[0034] Figure 1 A flowchart of a method for reconstructing snow area in a mountainous area based on machine learning is provided in an embodiment of the present invention, such as Figure 1 As shown, the method comprises at least the following steps:
[0035] Step S110, acquiring optical remote sensing data and terrain data and performing preprocessing; the optical remote sensing data includes MOD09GA data and Landsat8 OLI data; the preprocessing includes terrain correction and snow recognition.
[0036] Here, MOD09GA data is a standard product of MODIS, which provides reflectivity information of the earth's surface in multiple bands, and is of great significance for studying the radiation characteristics of the earth's surface, vegetation cover, land use change, climate change, etc. These data can be obtained through NASA's official website or through other data providers.
[0037] The data characteristics of MOD09GA data are as follows: 1) Band range: The MOD09GA product contains surface reflectance data in bands 1 - 7, which cover the visible and near-infrared spectral ranges; 2) Spatial resolution: The MOD09GA product has two spatial resolutions, namely 500 meters and 1000 meters. Among them, the reflectance data in bands 1 - 7 are provided at a resolution of 500 meters, while some additional geographic information data (such as the number of observations, quality assessment level, etc.) are provided at a resolution of 1000 meters. 3) Temporal resolution: The MOD09GA product is daily data, that is, surface reflectance observations are provided once a day. 4) Data format: MOD09GA data is usually stored in the HDF4 (Hierarchical Data Format version 4) format, which is a file format for storing and organizing large-scale data.
[0038] Landsat8 OLI data is provided by NASA and includes 9 bands with a spatial resolution of 30 meters, including a panchromatic band with a resolution of 15 meters. These data can be obtained through the Geospatial Data Cloud website or the official website of the United States Geological Survey (USGS).
[0039] Terrain data uses the SRTM V3 digital elevation model (DEM) data provided by NASA to extract elevation data for the study area, and calculates slope and aspect terrain factors based on the elevation data. Elevation (DEM), slope, and aspect data are used as auxiliary spatial feature data for the algorithm and as important input parameters for terrain correction.
[0040] Preprocessing the acquired data includes: performing terrain correction on MOD09GA data and Landsat8 OLI data, and using terrain features such as DEM, slope, and aspect as auxiliary spatial feature data for the algorithm; performing snow cover identification processing on Landsat8 OLI data. Among them, the purpose of terrain correction is to eliminate the influence of terrain effects on the algorithm. By performing geometric transformation on the reflectance, the reflectance of pixels is converted to the same reference plane, eliminating the change in radiance values caused by terrain changes, so that the pixels can reflect the true spectral characteristics of the ground objects; in the snow cover identification processing, appropriate snow cover identification algorithms (such as threshold method, classification method, etc.) are used to extract the snow cover area from Landsat8 OLI data.
[0041] In other ways, the preprocessing operations also include: based on the cloud mask and water body mask provided by the quality control QA band, performing masking processing on the MOD09GA band data to remove cloud and water body pixels in the surface reflectance, retaining the information of land pixels under clear sky, and avoiding the influence of clouds and water bodies on the algorithm inversion.
[0042] Step S120: Use the snow cover area generated from Landsat8 OLI data by the snow cover recognition algorithm as the ground truth label, and construct a sample dataset based on the preprocessed MOD09GA data and terrain data.
[0043] Here, the snow cover area generated from Landsat8 OLI data by the snow cover recognition algorithm is used as the ground truth input during the model training process, and cloud pixels are removed using the quality control QA band. Then, the terrain-corrected MOD09GA data, Landsat 8 OLI snow cover area data, and terrain data (elevation, slope, aspect) are integrated into a sample dataset.
[0044] Step S130: Divide the sample dataset into a training set and a validation set at a ratio of 4:1 to train and validate the XGBoost model; the XGBoost model improves the prediction performance by combining multiple decision trees.
[0045] Here, XGBoost is an ensemble algorithm based on gradient-boosted decision trees that can handle classification and regression problems in weakly supervised learning. In snow cover area reconstruction, the XGBoost algorithm combines multiple weak classifiers into a single strong classifier to obtain better classification performance.
[0046] The XGBoost algorithm is as follows: Given a dataset D = {(x i ,y i )}(|D| = n, x i ∈R m , y i ∈R n ), the sample prediction result can be expressed as:
[0047]
[0048] In the formula, is the predicted result, K is the number of trees, x i represents the sample, f k corresponds to the tree structure of the prediction score, represents the spatial function integrated by K trees, T is the number of nodes of the tree, s is the node label corresponding to the sample, and w s (x) represents the score at node q.
[0049] XGBoost performs feature splitting by continuously adding trees. Each time a tree is added, a new function is learned. Each decision tree is trained based on the previous round of prediction results to minimize the value of the objective function. This step-by-step iterative training method enables the model to have good generalization ability.
[0050] Step S140: Use the trained XGBoost model to predict the snow cover area for new input data.
[0051] Here, use the trained XGBoost model to predict the new input data. The input data includes the same feature data as the training sample data. According to the prediction results, generate a 30m high-resolution mountain snow cover area reconstruction map. Reconstruct the 30m high-resolution snow cover area. The prediction results can be used in fields such as snow cover monitoring and climate change research. In other embodiments, the data in the validation set can also be used to validate the trained XGBoost model to evaluate the prediction performance of the model.
[0052] The embodiments of the present invention can use high-resolution Landsat8 OLI data to generate real data of the snow cover area. At the same time, combined with MOD09GA data and DEM data, the accuracy and efficiency of snow cover area reconstruction are improved through the XGBoost algorithm. This machine learning-based snow cover area reconstruction method has advantages such as high efficiency, accuracy, and scalability, and is of great significance for fields such as snow cover monitoring and climate change research.
[0053] In some embodiments, the MOD09GA data includes 7 surface reflectance bands, and 4 bands related to the observation state, namely the sensor zenith angle, sensor azimuth angle, solar zenith angle, and solar azimuth angle; the terrain data includes a 30-meter resolution elevation image, slope data, and aspect data.
[0054] The above "obtain optical remote sensing data and terrain data and perform preprocessing" in step S110 includes the following steps: perform terrain correction on the MOD09GA data and Landsat8 OLI data; use a snow cover identification algorithm that combines three remote sensing indices, NDSI, NDVI, and NDFSI, to identify the snow cover on the Landsat8 OLI surface reflectance image to obtain a snow cover area image.
[0055] Here, in mountainous areas, due to the undulating terrain and relatively complex topography, the aspect and slope change significantly. The irradiance received by pixels in different aspects varies greatly. For example, the irradiance on the shady slope is weak, the brightness value is too low, and it has a low reflectance value, while the irradiance on the sunny slope is strong, the brightness value is too high, and it has a high reflectance value. This makes the same ground object show different spectral information due to different aspects, and different ground objects show the same spectral information in different aspects. Therefore, when the inversion and estimation method based on the assumption of a flat surface is applied to mountainous areas, large errors will occur. To eliminate the influence brought by this terrain effect, terrain correction has become an essential and important step in the preprocessing of mountain remote sensing images.
[0056] In some embodiments, the following terrain correction model is used during the terrain correction process:
[0057] L H (λ)=L I (λ)-a(λ) * (IC - IC H );
[0058]
[0059] Wherein, Z is the solar zenith angle, S is the slope, φ z is the solar azimuth angle, is the slope aspect, IC is defined as a function of the solar zenith angle, slope, solar azimuth angle and slope aspect, and its value range is from -1 to 1; L H (λ) is the pixel reflectance value of the corrected λ band, L I (λ) is the pixel reflectance value of the λ band before correction, a(λ) is the relationship between the radiance value L I (λ) of the original image and its corresponding IC L I (λ)=a(λ)*IC + b(λ) is obtained, IC H is the IC of the horizontal plane, that is, the cosine of the solar zenith angle.
[0060] In the above embodiment, a semi - empirical model based on an empirical rotation model is used to perform topographic correction on the image of the study area. This model converts the reflectance of pixels to the same reference plane by introducing a surface empirical rotation model that does not assume a Lambertian body, eliminating the change in radiance value caused by topographic changes. The model formula considers the relationship between the pixel reflectance values before and after correction, as well as the influence of the solar incidence angle (IC) on the reflectance. The value range of IC is from -1 to 1, which is jointly determined by the solar zenith angle, slope, solar azimuth angle and slope aspect. Among them, the slope and slope aspect are calculated according to the DEM.
[0061] This topographic correction model can better eliminate the influence brought by the topographic effect, Figure 2 showing a comparison chart of the changes in NDSI calculated from the Landsat8 OLI surface reflectance before and after topographic correction. As shown in Figure 2, an image with row and column numbers 136033 on December 20, 2020 is intercepted, and a comparison chart of the mountainous undulating area before and after topographic correction. Before topographic correction, it can be seen from Figure 2 that the image is affected by the mountainous terrain undulation. At the position of the shaded slope of the mountain that should be snow, the pixel radiance is low, resulting in a small NDSI value calculated, which cannot reflect the information of the real ground object. After topographic correction, the low radiance value at the shaded area of the slope of the mountain is increased. For example, the NDSI in the low - value area in the northeast part of the intercepted area has increased compared with before. Thus, it can be seen that the image after topographic correction can effectively weaken the mountain effect caused by topographic undulation.
[0062] In some embodiments, the above step S120 is further implemented through the following steps: Based on the range of the Landsat8 OLI surface reflectance image at the same time of the day, crop the MOD09GA image and the elevation image; use the nearest neighbor interpolation method to interpolate and resample the cropped images to the same 30-meter spatial resolution as the Landsat8 OLI image, and match the number of rows and columns of all the images used; use the 30-meter snow cover area image obtained by the snow cover recognition algorithm using the Landsat 8 OLI surface reflectance image at the same time of the day as the ground truth label, and use the QA band of Landsat8 OLI to exclude cloud pixels; perform splicing, splitting, missing value filling, and data normalization operations on all the data unified to the 30-meter spatial resolution, and save the processed data as an npy file.
[0063] Here, the cropping operation is used to crop the feature data such as the MODIS image and the DEM image to the same range as the Landsat8 OLI image. The resampling operation uses the nearest neighbor interpolation method to interpolate and resample all the images to a 30m resolution, ensuring that the number of rows and columns of all the images match. The cloud pixel exclusion operation uses the quality control QA band of Landsat8 OLI to exclude cloud pixels. Then, perform operations such as splicing, splitting, and missing value filling on the data unified to the 30m resolution. Normalize the data to improve the efficiency and effect of model training. Save the processed data in the npy format for convenient use in subsequent model predictions.
[0064] When inputting into the model, use the Landsat 8 OLI surface reflectance image to generate a 30m resolution snow cover area image using the snow cover recognition algorithm as the ground truth label. Each sample should include multiple features and corresponding labels (i.e., the snow cover area or snow cover range identified and generated at the current moment), where the multiple features include seven reflectance bands in the MOD09GA data, four bands related to the observation state (sensor zenith angle, sensor azimuth angle, solar zenith angle, and solar azimuth angle), and three terrain features (DEM, slope, and aspect), such as surface reflectance bands, terrain factors, etc., for a total of 14 feature data. The process of using the constructed sample dataset for model training is as Figure 3 shown, input the 14 feature data of each sample into the XGBoost model to reconstruct the 30m high-resolution snow cover area.
[0065] In some embodiments, the method further includes: during the model training process, separate the input data for training according to the row and column numbers of Landsat8 OLI; among them, at the positions with different row and column numbers, train to establish snow cover area reconstruction rules applicable to different regions, and set a preset number of valid pixels at the position of a single row and column number for training.
[0066] Here, in the embodiments of the present invention, the XGBoost model is implemented on the Jupyter Notebook software platform, which is a snow cover area downscaling model. The preprocessed data is input into the model for training.
[0067] Since there may be different geographical features at different row and column numbers, the data of different Landsat 8 OLI row and column numbers are separately trained to establish a snow cover area reconstruction model for different geographical regions, improving the applicability and accuracy of the model. There are approximately 360 million effective pixels for training at the location of a single row and column number, which provides a large amount of training data for the model.
[0068] In some embodiments, the snow cover identification algorithm that combines the three remote sensing indices of NDSI, NDVI, and NDFSI is used to identify snow cover in the Landsat8 OLI surface reflectance image, including: in non-forest areas and forest areas, when NDSI is greater than 0.4 and the reflectance of the near-infrared band NIR is greater than 0.11, it is identified as snow cover; in forest areas, when NDSI is between 0 and 0.4, and when NDFSI is greater than 0.4 and NDVI is less than 0.6, it is identified as snow cover.
[0069] Here, the currently widely used SNOMAP snow cover identification algorithm mainly relies on the threshold of NDSI to divide snow cover and non-snow cover ground objects. However, research shows that the threshold for snow cover discrimination will change under different vegetation cover conditions. The multi-remote sensing index dynamic identification algorithm that combines NDSI, NDVI, and NDFSI in the embodiments of the present invention can more accurately identify snow cover on different land cover types, improving the accuracy of snow cover identification.
[0070] In some embodiments, the method further includes: selecting 12 scenes of Landsat8 OLI data that cover the entire study area as the test data set; using MOD09GA and elevation data as inputs, and reconstructing the snow cover area data through the trained XGBoost model; using the Landsat8 OLI surface reflectance of cloud-free images to generate the snow cover area as the true data; the cloud amount of the cloud-free image is less than a specific value; comparing the reconstructed snow cover area data with the true data to evaluate the performance of the trained XGBoost model.
[0071] Here, Figure 4 shows the comparison of the reconstructed snow cover area at row and column number 135033. Figure 4 It can be seen that the terrain at row and column number 135033 is complex, with many fragmented patchy snow covers, and the distribution of snow cover is very patchy. The snow cover area reconstructed by the XGBoost algorithm can better show the distribution of snow cover.
[0072] The present invention uses terrain correction to effectively weaken the terrain effect brought by mountainous terrain. Based on MODIS and Landsat8 OLI data and terrain data, the XGBoost algorithm is used to directly reconstruct the snow area with high spatiotemporal resolution. The XGBoost model is selected as the core algorithm. Compared with the traditional random forest algorithm, the extreme gradient boosting XGBoost model with better performance can achieve significant efficiency improvement when processing large-scale data. The generalization ability of the model is also significantly improved. It can process data with higher dimensions and more features, and can mine the implicit relationship between different features in complex features, which is crucial for the reconstruction of snow area in complex terrain areas.
[0073] The algorithm of the present invention can effectively improve the accuracy of snow area reconstruction, and the reconstruction effect is good. The reconstructed snow pixels have high accuracy, and the probability of misclassifying snow pixels is small. The reconstructed snow area can more clearly depict the spatial details of snow distribution, and show more accurate snow distribution. It has high stability in different indicators such as OA, PA and Kappa coefficient values. Machine learning is combined with remote sensing data to achieve the reconstruction of snow area with a high temporal and spatial resolution of 30m per day in mountainous areas.
[0074] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the size of the serial number of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention. The serial numbers of the above-mentioned embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0075] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0076] In several embodiments provided by the present invention, it should be understood that the disclosed methods can be implemented in other ways. The methods disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments. The features disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0077] As described above, it is only the implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for reconstructing snow area in mountainous areas based on machine learning, characterized in that: include: Acquire optical remote sensing data and terrain data and perform preprocessing; The optical remote sensing data includes MOD09GA data and Landsat8 OLI data; the preprocessing includes terrain correction and snow recognition; The snow area generated by the snow recognition algorithm in the preprocessed Landsat8 OLI data is used as the true value label, and a sample dataset is constructed based on the preprocessed MOD09GA data and terrain data. The sample data set is divided into a training set and a validation set in a ratio of 4:1 to train and validate the XGBoost model; the XGBoost model improves prediction performance by combining multiple decision trees; Use the trained XGBoost model to predict snow cover area on new input data.
2. The method according to claim 1, characterized in that The MOD09GA data includes 7 bands of surface reflectance, and 4 bands related to the observation status, namely, sensor zenith angle, sensor azimuth angle, solar zenith angle and solar azimuth angle; the terrain data includes 30-meter resolution elevation images, slope data and slope aspect data; The step of obtaining optical remote sensing data and terrain data and performing preprocessing includes: Perform terrain correction on MOD09GA data and Landsat8 OLI data; The snow recognition algorithm combining the three remote sensing indices of NDSI, NDVI and NDFSI is used to identify snow on the Landsat 8OLI surface reflectivity image to obtain the snow area image.
3. The method according to claim 2, characterized in that The following terrain correction models are used during the terrain correction process: L H (λ)=L I (λ)-a(λ) * (IC-IC H ); Where Z is the solar zenith angle, S is the slope, φ z is the solar azimuth, is the slope direction, IC is defined as the function of solar zenith angle, slope, solar azimuth and slope direction, and its value range is -1 to 1; L H (λ) is the pixel reflectance value of the λ band after correction, L I (λ) is the pixel reflectance value of the λ band before correction, and a(λ) is the radiance value L of the original image. I The relationship between (λ) and its corresponding IC L I (λ)=a(λ)*IC+b(λ) to obtain, IC H IC is the cosine of the solar zenith angle in the horizontal plane.
4. The method according to any one of claims 1 to 3, characterized in that: The snow area generated by the Landsat8 OLI data using the snow recognition algorithm is used as the true value label, and a sample data set is constructed based on the preprocessed MOD09GA data and terrain data, including: Based on the range of the Landsat8 OLI surface reflectance image at the same time of the day, the MOD09GA image and elevation image are cropped; The cropped images were resampled to the same 30-meter spatial resolution as the Landsat 8OLI images using the nearest neighbor interpolation method, and the number of rows and columns of all images used were matched; The 30-meter snow area image obtained by the snow recognition algorithm using the Landsat 8OLI surface reflectance image at the same time of the day is used as the true value label, and the QA band of Landsat8 OLI is used to remove cloud pixels; All data unified to a spatial resolution of 30 meters are spliced, segmented, missing value filled, and data normalized, and the processed data are saved as npy files.
5. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: The XGBoost model was constructed using the Jupyter Notebook software platform; During the model training process, the input data are trained separately according to the row and column numbers of Landsat8 OLI. Among them, the positions with different row and column numbers are trained to establish snow area reconstruction rules applicable to different regions, and the positions with a single row and column number are trained with a preset number of valid pixels.
6. The method according to claim 2, characterized in that The snow recognition algorithm using the three remote sensing indices of normalized snow cover index NDSI, normalized vegetation index NDVI and normalized forest snow cover index NDFSI is used to identify snow on the Landsat8OLI surface reflectivity image, including: In non-forested and forested areas, when the NDSI is greater than 0.4 and the near-infrared band NIR reflectance is greater than 0.11, it is identified as snow accumulation; In forested areas, snow cover is identified when the NDSI is between 0 and 0.4 and when the NDFSI is greater than 0.4 and the NDVI is less than 0.
6.
7. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Twelve scenes of Landsat8 OLI data covering the entire study area were selected as the test dataset; Using MOD09GA and elevation data as input, the snow area data was reconstructed through the trained XGBoost model; The snow area is generated as the real data using the Landsat8 OLI surface reflectance of a cloud-free image; the cloud amount of the cloud-free image is less than a specific value; The reconstructed snow area data is compared with the real data, and the performance of the trained XGBoost model is evaluated.
Citation Information
Patent Citations
Universal terrain correction optimization method based on remote sensing image segmentation unit
CN111257854A
Glacier albedo inversion method and system
CN115631429A
High mountain accumulated snow environment simulation method and system based on SCyclGAN
CN116958468A
Random forest accumulated snow coverage identification method for high-resolution series optical satellites
CN118015447A
XGBoost-based downscaling snow depth inversion method
CN118446092A
Cited By
Sub-pixel accumulated snow mapping method based on medium-resolution remote sensing data
CN121505439A
A sub-pixel snow filling method based on medium-resolution remote sensing data
CN121505439B
Snow depth detection method and device based on machine learning, medium and product
CN121582798A
Outdoor snowfield area identification method
CN122198341A