Multi-source data fusion-based Yangtze River source region snow melting data set construction method
Through multi-source data fusion and deep learning methods, a snow melting data set in the Yangtze River source area was constructed, which solved the problem of inaccurate snow melting data in the existing technology, achieved high-precision snow melting monitoring, and supported water resource management and ecological protection.
Patent Information
- Application Number
- CN202510765350.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing snow accumulation data is difficult to accurately reflect the snow melting process in the source area of the Yangtze River. Satellite observation data is greatly affected by weather conditions. The data on the ground observation site lacks snow melting information. The accuracy of the data is relied on the quality of the model, making it difficult to provide comprehensive and accurate snow melting information.
Multi-source data fusion method is adopted, including snow-water equivalent data, snow melting data, ground site data, potential evaporation data and hydrological site traffic data. Through deep learning methods such as long-term short-term memory network models and adaptive standardization methods, effective information of multi-source data is extracted to form a high-precision snow melting data set.
The spatial and temporal resolution and accuracy of snow melting data can be improved, and the spatial distribution and temporal dynamic changes of snow melting in the Yangtze River source area can be captured more accurately, and the dry season runoff prediction and flood warning are supported, which is of great ecological protection significance.
Smart Images

Figure CN120296678A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of application of hydrometeorological data, and particularly to a method for constructing a snowmelt dataset in the source area of the Yangtze River based on multi-source data fusion. Background Art
[0002] Snowmelt is an important source of spring water resources in many regions and is of great significance for water resource management. The snowmelt process in the source area of the Yangtze River affects the runoff volume of the Jinsha River Basin, and thus has a significant impact on hydropower generation and the ecosystem. Although the importance of snowmelt monitoring is self-evident, due to the complexity and variability of the snowmelt process, there are few observation means, and currently, many challenges are still faced, especially in the high-altitude or complex terrain source area of the Yangtze River.
[0003] Currently, the mainstream snow cover data includes satellite observation data, ground station observation data, and reanalysis data. Satellite observation data provides a wide range of spatial coverage and can monitor the distribution and changes of snow cover in real time, but it is greatly affected by weather conditions, and cloud cover will cause missing observation data; the data of ground observation stations provides accurate snow depth data but no snowmelt data; reanalysis data is generated by a numerical weather prediction model combined with various observation data and provides snow cover information with a long time series, but its accuracy depends on the accuracy of the model and the quality of the input data. In short, the existing snow cover data is difficult to accurately reflect the snowmelt process. Satellite observation and reanalysis data each have limitations and are difficult to provide comprehensive and accurate snowmelt information alone. Therefore, there is an urgent need for a more comprehensive method to improve the accuracy of snowmelt data in the source area of the Yangtze River. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for constructing a snowmelt dataset in the source area of the Yangtze River based on multi-source data fusion in view of the above-mentioned deficiencies of the existing technology. The data with different spatio-temporal resolutions are effectively fused, and the effective information of multi-source data is extracted through a deep learning method to improve the accuracy of snowmelt data.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: The present invention provides a method for constructing a snowmelt dataset in the source area of the Yangtze River based on multi-source data fusion, including the following steps; S1. Collect data of the study area: Collect snow water equivalent data, snowmelt data, ground station data, potential evapotranspiration data, and hydrological station flow data; The snow water equivalent data is daily snow water equivalent product data; The snowmelt data is reanalysis daily snowmelt data; The ground station data is measured snow depth data, temperature data, and snowfall data on a daily scale; The potential evapotranspiration data is a 1km monthly potential evapotranspiration dataset in the source area of the Yangtze River; The flow data of the hydrological station are the measured flow data on a daily scale; S2. Data preprocessing: including spatial matching of ground stations and calculation of snowmelt data; the spatial matching of ground stations is to match the snow cover data and potential evapotranspiration data with unified resolution with the ground station data through geographical position coordinates to generate a sample data sequence and a time series sample data; The calculation of snowmelt data is to calculate the satellite remote sensing snowmelt sequence and the station snowmelt sequence respectively by using the satellite remote sensing snow water equivalent data and the ground station data; S3. Spatiotemporal feature fusion of multi-source snowmelt data: including spatial feature fusion of snowmelt data and time series feature fusion of snowmelt data; The spatial feature fusion of snowmelt data refers to fusing the snowmelt data sets based on ground stations and satellite remote sensing in S2 by using a two-layer long short-term memory network model, and then restoring them by using an activation function and an inverse normalization method to form the monthly-scale snowmelt data after fusion; The time series feature fusion of snowmelt data refers to further fusing the reanalysis daily snowmelt data in S1 with the monthly-scale snowmelt data after fusion in S2 by using the adaptive instance normalization method to finally obtain the daily snowmelt data; S4. Evaluation of the fusion effect of snowmelt data: using snowmelt observation substitute data and accuracy evaluation indicators to evaluate the accuracy of the monthly-scale snowmelt data after fusion.
[0006] Further, in S2, the spatial matching of ground stations is the daily snow water equivalent product data. Based on the grid center point, the grid value closest to the longitude and latitude of the meteorological station is selected as the snow water equivalent data corresponding to the meteorological station for calculating the ground station snowmelt data based on remote sensing data.
[0007] Further, in S2, the extraction of time series sample data is specifically: based on the reanalysis daily snowmelt data with a time resolution, extracting the grid daily snowmelt time series sample data closest to the longitude and latitude of the meteorological station; If a certain month has 31 days, a 31×1 time series is formed.
[0008] Further, the calculation of snowmelt data in S2 is specifically: Using the temperature index model based on the relationship between snowmelt and temperature to calculate the monthly snowmelt data, generating the monthly snowmelt data based on ground station data and remote sensing data respectively. The specific formula is: ; where, is the station snowmelt data of the th month; is the temperature coefficient, with the unit of mm / °C / day, which is an empirical coefficient connecting the snowmelt rate and air temperature; is the cumulative positive air temperature, with the unit of °C; is the monthly ground station snow cover data. The daily snow water equivalent data can be used, and it is the total amount of snowmelt; When using the ground station snowfall data, it is necessary to substitute and calculate the monthly ground station snow cover data : ; Among them, is the ground station snowfall, with the unit of mm; is the ground station snow sublimation, with the unit of mm; The formula for ground station snow sublimation is: ; Among them, is the minimum value function, and the smaller value between the snowmelt amount calculated by the function and the total amount of snow currently used for snowmelt is taken as the final snowmelt amount; is the total amount of snow cover for sublimation; is the potential evapotranspiration. The potential evapotranspiration is the sample data generated by matching the monthly potential evapotranspiration dataset of the Yangtze River source area at 1 km with the ground station data through geographic position coordinates.
[0009] Furthermore, in the step S3, the specific process of fusing the ground station and remote sensing snowmelt sample data using the two-layer long short-term memory network model is as follows: First, the monthly-scale snowmelt sequence sample based on the station data is input into the memory neurons of the long short-term memory network model with multiple hidden layers for cyclic connection. The output eigenvalue and the monthly-scale snowmelt sequence sample based on the remote sensing data are input into the second-layer long short-term memory network model. Finally, the fused feature variables are output as the fused monthly-scale snowmelt data using the activation function and inverse normalization. The specific formula is: ; Among them, is the th station, and the fused monthly-scale snowmelt data for the th month; is the th ground station, and the snowmelt feature vector for the th month; is the standard deviation function; is the arithmetic mean function.
[0010] Further, in step S3, the time series feature fusion of the snowmelt data is specifically as follows: The daily snowmelt data is further fused with the monthly-scale snowmelt data after being fused by a long short-term memory network model using the adaptive standardization method, and finally the daily snowmelt data is obtained. The specific formula is: ; where is the snowmelt data of the th station on the th day of the th month; is the reanalysis data of the snowmelt of the th station on the th day of the th month.
[0011] Further, the Pearson correlation coefficient is used to evaluate the reliability of the fused daily-scale snowmelt data; Due to the lack of snowmelt observation data, alternative data will be used to assist in proving the feasibility of the snowmelt data. Therefore, the snowmelt observation alternative products in step S4 include the snow depth change data of the ground stations and the measured flow data of the hydrological stations; the snow depth change of the ground stations is obtained by subtracting the snow depth of the previous day from the snow depth of the current day as the alternative data for the snowmelt observation of the previous day; the measured flow data of the hydrological stations is the dry-season flow data of the nearby hydrological stations selected as the alternative data for the dry-season snowmelt observation.
[0012] The beneficial effects of the present invention are as follows: The dataset construction method can effectively utilize remote sensing data, reanalysis data and station data, calculate the snowmelt at the stations through the temperature coefficient algorithm, and realize the efficient fusion of the spatio-temporal information of the snowmelt by using deep learning and other methods. It overcomes the drawbacks of a single data source, further improves the spatio-temporal resolution of the snowmelt data, more accurately captures the spatial distribution and temporal dynamic changes of the snowmelt in the Yangtze River source area, and is of great significance for the prediction of dry-season runoff, flood warning and ecological protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a flowchart of a method for constructing a snowmelt dataset in the Yangtze River source area based on multi-source data fusion; Figure 2 is a flowchart of an embodiment; Figure 3 is the daily snowmelt data of a certain year and month in the study area. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0015] Please refer to Figure 1 and Figure 2 , a method for constructing a snowmelt dataset in the source area of the Yangtze River based on multi-source data fusion, comprising the following steps: S1. Collect data in the study area: Collect snow water equivalent data, snowmelt data, ground station data, potential evapotranspiration data, and hydrological station flow data; The snow water equivalent data is daily snow water equivalent product data with a spatial resolution of 0.25°; Among them, the daily snow water equivalent product data is satellite remote sensing snow water equivalent product data developed by the Japan Aerospace Exploration Agency; The snowmelt data is reanalysis daily snowmelt data with a spatial resolution of 0.25°; The snowmelt data is the analyzed daily snowmelt data of the fifth-generation global climate reanalysis product provided by the European Centre for Medium-Range Weather Forecasts.
[0016] The ground station data is measured snow depth data, temperature data, and snowfall data at the daily scale; The potential evapotranspiration data is a 1-km monthly potential evapotranspiration dataset in the source area of the Yangtze River; The hydrological station flow data is measured flow data at the daily scale; S2. Data preprocessing: Including spatial matching of ground stations and calculation of snowmelt data; the spatial matching of ground stations is to match the snow cover data and potential evapotranspiration data with unified resolution with the ground station data through geographic position coordinates to generate a sample data sequence and a time series sample data; The calculation of snowmelt data is to calculate the satellite remote sensing snowmelt sequence and the station snowmelt sequence respectively using the satellite remote sensing snow water equivalent data and the ground station data; S3. Spatiotemporal feature fusion of multi-source snowmelt data: Including spatial feature fusion of snowmelt data and time series feature fusion of snowmelt data; The spatial feature fusion of snowmelt data refers to fusing the snowmelt datasets based on ground stations and satellite remote sensing in S2, and then using an activation function and an inverse normalization method for reduction to form monthly-scale snowmelt data after fusion; The time series feature fusion of snowmelt data refers to further fusing the ERA5 reanalysis daily snowmelt data in S1 with the monthly-scale snowmelt data after fusion in S2 using the adaptive instance normalization method to finally obtain daily snowmelt data; S4. Evaluation of the fusion effect of snowmelt data: Use snowmelt observation substitute data and accuracy evaluation indicators to evaluate the accuracy of the fused snowmelt data.
[0017] In S2, the spatial matching of ground stations is the daily snow water equivalent data. Based on the grid center point, the grid value closest to the longitude and latitude of the meteorological station is selected as the snow water equivalent data corresponding to the meteorological station, which is used to calculate the snowmelt data of the station based on remote sensing data.
[0018] In S2, the extraction of time series sample data is specifically as follows: Based on the reanalysis daily snowmelt data with a time resolution, the grid daily snowmelt time series sample data closest to the longitude and latitude of the meteorological station is extracted; If a certain month has 31 days, a 31×1 time series is formed.
[0019] The calculation of the snowmelt data in S2 is specifically as follows: The temperature index model based on the relationship between snowmelt and temperature is used to calculate the monthly snowmelt data, and the monthly snowmelt data based on station data and remote sensing data are generated. The specific formula is: ; Among them, is the station snowmelt data for the th month; is the temperature coefficient, with the unit of mm / °C / day, which is an empirical coefficient relating the snowmelt rate to the air temperature; is the cumulative positive air temperature, with the unit of °C; is the ground station snow cover data for the th month. The daily snow water equivalent data can be used, which is the total amount of snowmelt; When using the station snowfall data, it is necessary to substitute and calculate the ground station snow cover data for the th month : ; Among them, is the ground station snowfall, with the unit of mm; is the ground station snow sublimation, with the unit of mm; The formula for ground station snow sublimation is: ; Among them, is the function to find the minimum value. The smaller value between the snowmelt amount calculated by the function and the total amount of snow currently used for snowmelt is taken as the final snowmelt amount; is the total amount of snow cover for sublimation; is the potential evapotranspiration, and the empirical coefficient is queried according to Table 1; The potential evapotranspiration is the sample data generated by matching the 1km monthly potential evapotranspiration dataset in the Yangtze River source area with the ground station data through geographical position coordinates.
[0020] Table 1 Empirical parameters for estimating snow sublimation of different climate zones and different snow types Value
[0021] In S3, the specific process of using the two-layer long short-term memory network model to fuse the station and remote sensing snowmelt sample data is as follows: First, the monthly-scale snowmelt sequence samples based on station data are input into the memory neurons of the long short-term memory network model with multiple hidden layers for cyclic connection. The output eigenvalue and the monthly-scale snowmelt sequence samples based on remote sensing data are input into the second-layer long short-term memory network model. The specific formula of the long short-term memory model memory neuron is: ; Among them, the superscript 1 is the first-layer long short-term memory network model; is the forgetting gate; is the activation function; is the weight matrix; and are the hidden states of the current station and the previous station respectively. The data of each unit in the first-layer model is , and in the second-layer long short-term memory network model, the remote sensing snowmelt sample data and the output of the first-layer model are used as the input of the second-layer model; is the bias term; is the input of the station snowmelt sample data; is the hyperbolic tangent function; is the cell state of the current station; The structure of the long short-term memory network model is used to fuse the data features of the station and remote sensing, and the specific formula is as follows: ; Among them, the superscript 2 refers to the second-layer long short-term memory network model; is the input of the remote sensing snowmelt sample data.
[0022] The fused feature variables are output as the fused monthly-scale snowmelt data by using the activation function and inverse normalization, and the specific formula is: ; Among them, is the th station, and the monthly-scale snowmelt data after fusion in the th month; is the snowmelt feature vector of the th station in the th month; is the standard deviation function; is the arithmetic mean function.
[0023] In S3, the specific feature fusion of the snowmelt data time series is as follows: Using the adaptive standardization method, the daily snowmelt data is further fused with the monthly-scale snowmelt data after being fused by the long short-term memory network model, and finally the daily snowmelt data is obtained. The specific formula is: ; where, is the snowmelt data of the th station on the th day of the th month; is the reanalysis data of the snowmelt of the th station on the th day of the th month.
[0024] Due to the lack of observed data on snowmelt, alternative data will be used to assist in proving the feasibility of the snowmelt data. Therefore, the snowmelt observation alternative products in S4 include the station snow depth change data and the measured flow data of the hydrological stations; the station snow depth change is obtained by subtracting the snow depth of the previous day from the snow depth of the current day as the alternative data for the snowmelt observation of the previous day; the measured flow data of the hydrological stations is the dry-season flow data of the nearest hydrological station as the alternative data for the dry-season snowmelt observation.
[0025] Among them, the specific formula of the Pearson correlation coefficient is: ; where, is the number of meteorological stations in the study area; is the measured snow depth change data of the ground station or the dry-season flow data of the nearest hydrological station, and the average value is ; is the daily snowmelt data product after fusion, and the average value is .
[0026] After the data fusion, the accuracy of the snowmelt data fusion method for the multi-source data in this embodiment is evaluated. Among them, the daily snowmelt data of a certain year and month in the study area is as Figure 3 shown. It can be seen from Figure 3 that the fused snowmelt data obtained by the multi-source snowmelt data fusion method proposed in the present invention has a relatively consistent variation law with the measured snow depth change data of the ground stations in the study area, and is closer to the station observation results than the reanalysis data.
[0027] Further, the precision evaluation results based on the Pearson correlation coefficient are shown in Table 2. It can be seen from Table 2 that the correlation coefficients between the fused snowmelt data obtained by the multi-source snowmelt data fusion method proposed in the present invention and the measured snow depth change data of the ground stations and the measured flow of the nearby hydrological stations in the study area are higher than those of the reanalysis data products, reaching 0.8 and 0.68 respectively, indicating that the multi-source snowmelt data fusion method proposed in the present invention can obtain snowmelt data with high precision, fully considering multiple data characteristics of ground stations, satellite remote sensing and reanalysis data, thereby improving the precision of the snowmelt data obtained after the multi-source snowmelt data fusion.
[0028] Table 2 Precision Evaluation of the Fusion Results of Multi-Source Snowmelt Data in a Certain Study Area in the Embodiment
[0029] The above-described embodiments merely represent the implementation manners of the present invention, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A method for constructing a snowmelt dataset in the source area of the Yangtze River based on multi-source data fusion, characterized in that, Including the following steps; S1. Collect data in the study area: Collect snow water equivalent data, snowmelt data, ground station data, potential evapotranspiration data, and hydrological station flow data; The snow water equivalent data is daily snow water equivalent product data; The snowmelt data is reanalyzed daily snowmelt data; The ground station data is measured snow depth data, temperature data, and snowfall data at the daily scale; The potential evapotranspiration data is a monthly potential evapotranspiration dataset with a resolution of 1 km in the Yangtze River source area; The hydrological station flow data is measured flow data at the daily scale; S2. Data preprocessing: including spatial matching of ground stations and calculation of snowmelt data; The spatial matching of ground stations is to match the snow cover data and potential evapotranspiration data with unified resolution with the ground station data through geographical point coordinates to generate a sample data sequence and a time series sample data; The calculation of snowmelt data is to calculate the snowmelt sequence based on satellite remote sensing and the snowmelt sequence of stations respectively using satellite remote sensing snow water equivalent data and ground station data; S3. Spatiotemporal feature fusion of multi-source snowmelt data: including spatial feature fusion of snowmelt data and temporal series feature fusion of snowmelt data; The spatial feature fusion of snowmelt data refers to fusing the snowmelt datasets based on ground stations and satellite remote sensing in S2 using a two-layer long short-term memory network model, and then restoring them using an activation function and an inverse normalization method to form monthly-scale snowmelt data after fusion; The temporal series feature fusion of snowmelt data refers to further fusing the reanalyzed daily snowmelt data in S1 with the monthly-scale snowmelt data after fusion in S2 using the adaptive instance normalization method to finally obtain daily snowmelt data; S4. Evaluation of the fusion effect of snowmelt data: Evaluate the accuracy of the monthly-scale snowmelt data after fusion using snowmelt observation surrogate data and accuracy evaluation indicators.
2. A method for constructing a snowmelt dataset in the source area of the Yangtze River based on multi-source data fusion according to claim 1, characterized in that: In S2, the spatial matching of ground stations is the daily snow water equivalent product data. Based on the grid center point, select the grid value closest to the longitude and latitude of the meteorological station as the snow water equivalent data corresponding to the meteorological station for calculating the ground station snowmelt data based on remote sensing data.
3. A method for constructing a snowmelt dataset in the source area of the Yangtze River based on multi-source data fusion according to claim 2, characterized in that: In S2, the extraction of time series sample data is specifically as follows: Based on the time resolution of the reanalyzed daily snowmelt data, extract the grid daily snowmelt time series sample data closest to the longitude and latitude of the meteorological station; If a certain month has 31 days, then a 31×1 time series is formed.
4. A method for constructing a snowmelt dataset in the source area of the Yangtze River based on multi-source data fusion according to claim 3, characterized in that, The calculation of snowmelt data in S2 is specifically as follows: Calculate the monthly snowmelt data using a temperature index model based on the relationship between snowmelt and temperature, and generate monthly snowmelt data based on ground station data and remote sensing data respectively. The specific formula is: ; Among them, is the site snowmelt data for the month; is the temperature coefficient, in units of millimeters per degree Celsius per day, which is an empirical coefficient relating the snowmelt rate to the air temperature; is the cumulative positive air temperature, in units of degrees Celsius; is the ground site snow cover data for the month, which can use the daily snow water equivalent data and is the total amount of snowmelt; When using the snowfall data of ground stations, it is necessary to substitute and calculate the snow cover data of ground stations for the month : ; wherein, is the snowfall at the ground station, in millimeters; is the snow sublimation at the ground station, in millimeters; The formula for snow sublimation at the ground station is: ; Among them, is the function for obtaining the minimum value. The smaller value between the snowmelt amount calculated by the function and the total amount currently used for snowmelt is taken as the final snowmelt amount; is the total amount of snow cover for sublimation; is the potential evapotranspiration. The potential evapotranspiration is the sample data generated by matching the monthly potential evapotranspiration dataset of 1 km in the Yangtze River source area with the ground station data through geographical position coordinates.
5. A method for constructing a snowmelt dataset in the source region of the Yangtze River based on multi-source data fusion according to claim 4, characterized in that In S3, the specific process of using a two-layer long short-term memory network model to fuse ground station and remote sensing snowmelt sample data is as follows: First, the monthly-scale snowmelt sequence samples based on station data are input into the memory neurons of the long short-term memory network model with multiple hidden layers for cyclic connection. The output eigenvalue and the monthly-scale snowmelt sequence samples based on remote sensing data are input into the second-layer long short-term memory network model. Finally, the fused feature variables are output as the fused monthly-scale snowmelt data using the activation function and inverse normalization. The specific formula is as follows: ; Among them, is the th site, and the monthly-scale snowmelt data after fusion in the th month; is the th ground site, and the snowmelt feature vector in the th month; is the standard deviation function; is the arithmetic mean function.
6. A method for constructing a snowmelt dataset in the source area of the Yangtze River based on multi-source data fusion according to claim 5, characterized in that, In S3, the specific fusion of the time series features of snowmelt data is as follows: The daily snowmelt data is further fused with the monthly-scale snowmelt data fused by the long short-term memory network model using the adaptive standardization method, and finally the daily snowmelt data is obtained. The specific formula is as follows: ; Among them, is the th station, and the snowmelt data on the th month and the th day; is the th station, and the snowmelt reanalysis data on the th month and the th day.
7. A method for constructing a snowmelt dataset in the source area of the Yangtze River based on multi-source data fusion according to claim 6, characterized in that: Using the Pearson correlation coefficient Conduct a reliability assessment on the daily-scale snowmelt data after fusion; Due to the lack of observed data on snowmelt, alternative data will be used to assist in proving the feasibility of the snowmelt data. Therefore, the snowmelt observation alternative products in S4 include the ground station snow depth change data and the measured flow data of hydrological stations. The ground station snow depth change is obtained by subtracting the snow depth of the previous day from the snow depth of the current day as the alternative data for snowmelt observation of the previous day. The measured flow data of hydrological stations is the dry-season flow data of the nearest hydrological station selected as the alternative data for dry-season snowmelt observation.
Citation Information
Patent Citations
Accumulated snow phenological information fusion method and system based on multi-source remote sensing data
CN106780433A
Long-time sequence accumulated snow remote sensing data set construction method based on remote sensing data
CN114494859A
Multi-source remote sensing rainfall data fusion method based on multi-modal deep learning
CN114970743A
Small watershed rainfall estimation method and device based on multi-source information fusion and storage device
CN116338821A
MODIS accumulated snow area product missing information space-time reconstruction method
CN117635820A