Surface temperature and downscaling temperature inversion method based on remote sensing data
Through multi-source remote sensing data fusion and random forest model downscale technology, the problems of insufficient resolution and poor robustness of surface temperature inversion in traditional methods are solved, and high-precision urban heat island effect analysis and environmental management are achieved.
Patent Information
- Application Number
- CN202510492880.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-26
AI Technical Summary
In the study of urban heat island effect, traditional methods have problems such as insufficient data resolution, large spatial and temporal matching error, outlier interference and specific emissivity calculation deviation, resulting in low surface temperature inversion accuracy and poor model robustness.
Multi-source remote sensing data fusion technology is used to optimize spatiotemporal matching through cubic spline interpolation, outliers are eliminated using the Cook distance method, and temperature drop scale is used to calculate the specific emissivity with the NDVI threshold method to generate high-precision surface temperature data.
High-precision urban surface temperature inversion is achieved, the robustness and downscale accuracy of the model are improved, and the problems of low resolution and large verification errors exist in traditional methods are solved.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing data processing, and specifically relates to a high-precision surface temperature and downscaled temperature inversion method based on multi-source remote sensing data, which is suitable for urban heat island effect analysis, refined climate research and environmental management. Background Art
[0002] The urban heat island effect significantly affects urban ecological security and residents' health, and accurate inversion of land surface temperature (LST) is the key to studying the urban heat island effect. Traditional methods have the following limitations: 1. Insufficient data resolution: Relying on the inversion results of a single satellite (such as Landsat) is limited to a spatial resolution of 30 meters, which makes it difficult to capture subtle differences in urban surface temperature; 2. Spatiotemporal matching error: Meteorological station data is sparse and has a low temporal resolution (usually 3 hours), which is inconsistent with the satellite transit time, resulting in reduced verification accuracy; 3. Outlier interference: Existing methods lack a systematic outlier screening mechanism, which affects the robustness of the model; 4. Emissivity calculation bias: The NDVI threshold method parameters are empirically based and not optimized for complex urban surfaces, resulting in errors in emissivity estimation. Summary of the Invention
[0003] 1. Purpose of the Invention The present invention proposes a high-precision temperature inversion method that integrates multi-source remote sensing data with meteorological station observations. It optimizes spatiotemporal matching through cubic spline interpolation, eliminates outliers through the Cook distance method, and implements temperature downscaling through the random forest model, thus solving the problems of low resolution, large verification error, and poor model robustness of traditional methods.
[0004] 2. Technical Solution A method for inverting land surface temperature and downscaled temperature based on remote sensing data comprises the following steps: Step 1: Verification of fitting between measured surface temperature and inverted temperature (1) Cubic spline interpolation: The measured temperature data at meteorological stations at intervals of 3 hours throughout the day are interpolated into a continuous all-day temperature curve, and the Landsat 8 / 9 imaging time and the station's geographical location are superimposed to generate time-space matching LST inversion data; (2) Fitting accuracy verification: The coefficient of determination (R²) and root mean square error (RMSE) are used to compare the inverted LST with the measured temperature. The calculation formula is as follows:
[0005] Where: , temperature values observed by ground meteorological stations; , LST value retrieved from Landsat; n, total number of records of observation and retrieved data; , the average temperature value observed by ground meteorological stations.
[0006] (3) Outlier elimination: Cook's distance method is used to evaluate the impact of data records on the regression model. The formula for calculating Cook's distance is as follows:
[0007] Where: Indicates the Cook distance value corresponding to the i-th data record; It represents the LST value of the jth record obtained by Landsat inversion; Indicates to remove the i The first record obtained by refitting j The recorded Landsat-retrieved LST values; Indicates the total number of records of observation and inversion data; Then it is the mean square error, that is, the square of RMSE. >4 / ( n − p −1) are considered as outliers and are removed.
[0008] Step 2: Surface temperature downscaling calculation (1) Calculation of emissivity: The vegetation coverage is determined based on the NDVI threshold method, and then the surface emissivity ε is calculated:
[0009] Wherein, NDVI is the normalized difference vegetation index, which is an important indicator for evaluating vegetation coverage; NDVI values representing bare soil or areas without vegetation cover; It specifically refers to the NDVI value presented by the pixel completely covered by vegetation, that is, the value of the pure vegetation pixel. Take the empirical value and , that is, when the NDVI of a pixel exceeds 0.70, The value is 1; when NDVI is less than 0.05, The value is 0.
[0010] (2) Random forest downscaling: In the Google Earth Engine (GEE) platform, Landsat 8 / 9 (30 m) and Sentinel-2 (10 m) data were combined to build a random forest model with NDVI, NDBI, and NDWI as input features and LST as output variable. The model was then migrated to high-resolution Sentinel-2 data to generate 10 m resolution LST. (3) Residual optimization: perform regression residual calculation and spatial smoothing on the downscaling results to reduce high-frequency noise.
[0011] Step 3: Verify downscaling results The consistency of the downscaled LST (10 m) and the original Landsat LST (30 m) was assessed by a linear regression model, requiring the regression slope to be 0.95–1.05, the absolute value of the intercept to be ≤ 0.5 °C, and R² to be ≥ 0.90.
[0012] 3. Advantages of the present invention (1) Multi-source data fusion: combining Landsat, Sentinel-2, and meteorological station data, taking into account both spatial resolution and temporal continuity; (2) Space-time matching optimization: cubic spline interpolation solves the problem of mismatch between satellite transit time and meteorological observation time; (3) Improved robustness: Cook's distance method automatically removes outliers to avoid model overfitting; (4) High downscaling accuracy: The random forest model captures the nonlinear relationship between multispectral features and LST, and residual smoothing further optimizes details. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is a flow chart of the surface temperature and downscaled temperature inversion method based on remote sensing data; Figure 2 is the fitting degree between the measured temperature at the meteorological station and the inverted temperature; Figure 3 The Cook distance calculation result between the measured temperature and the inverted temperature at the meteorological station; Figure 4 Scatter plot of 10-meter surface temperature and 30-meter downscaled surface temperature. DETAILED DESCRIPTION
[0014] The features of the present invention are further described below by implementing examples, but without limiting the claims of the present invention.
[0015] Nanning City, Guangxi Zhuang Autonomous Region was selected as the study area. QGIS software was used to access the Google Earth tile map service to obtain high-definition image TIF data of Nanning City (0.3-meter resolution in 2022). Geographic information data were obtained using the National Tibetan Plateau Science Data Center, the National Earth System Science Data Center, the remote sensing image data platform of the United States Geological Survey, and the Wuhan University Resource Sharing Platform. The China Meteorological Data Network was used to obtain ground meteorological station data in the study area at three-hour intervals throughout the day.
[0016] 1. According to the method described in step (1), the measured data of the meteorological stations at three-hour intervals throughout the day are interpolated into a full-day temperature curve using the cubic spline interpolation method. Finally, the LST 30m obtained by inversion is obtained by combining the position of the meteorological station in the Landsat image and superimposing the imaging time of Landsat L8 / L9.
[0017] 2. The LST 30m inverted in step 1 was verified with the three-hour interpolated temperature data of the meteorological stations in Guangxi at the corresponding coordinates of the imaging time, and the degree of deviation was tested using the root mean square error (RMSE) and the coefficient of determination (R2). Figure 2 R² indicates the linear correlation between the model's predictions and ground observations; RMSE indicates the model's prediction error; MPE indicates whether there is a systematic deviation between the overall predictions and the observed values; and RPE represents the relative prediction error as a percentage of the observed values. The results indicate a strong linear correlation between the model's predictions and ground observations, a relatively small prediction error, and a relatively accurate overall prediction. However, there is a certain degree of systematic deviation between the overall predictions and the observed values, suggesting a tendency for overestimation.
[0018] 3. Use the Cook distance method to filter outliers in the data set obtained in step 1 and calculate the Cook distance of the data ( Figure 2 After applying the Cook's distance method, we observe that these two outliers have large Cook's distances, indicating that they significantly affect the fit of the regression model. These two data points are identified as outliers and removed. After removing these outliers, we recalculate the regression model parameters and find that the model fit is significantly improved compared to when the outliers were not removed, with a decrease in the RMSE.
[0019] 4. Downscale the data obtained in step 1 to obtain LST data at 10 m resolution. Use the linear regression equation to evaluate the relationship between the downscaled LST and the LST at 30 m resolution ( Figure 4The slope of the linear regression equation was used to analyze the changing trends among the variables in the dataset. The slope of the linear regression equation for each date ranged from 0.6449 to 0.8818, indicating that the changing trends among the variables in the dataset were similar and relatively consistent. A strong relationship was observed between the 30-meter resolution LST and the LST downscaled to 10 meters, indicating that the downscaling process allowed for a relatively accurate prediction of the 10-meter resolution LST. Furthermore, the intercept of the linear regression equation, reflecting the relationship between the expected downscaled LST and the inverted temperature, was positive for each date. This indicates that when the 30-meter resolution LST is zero, the expected 10-meter resolution LST is larger than the 30-meter resolution LST, indicating a certain degree of overestimation of the inverted surface temperature. Furthermore, the R-squared value of the linear regression equation for each date indicates the reliability of the downscaled LST in explaining the 30-meter resolution LST. These results demonstrate that the model has strong explanatory power and fits the data well. Finally, the impact of temperature changes on the model fitting effect does not show obvious regularity, which indicates that the model has a relatively robust fitting ability for data under different temperature ranges and can adapt to data changes under different temperature conditions.
[0020] It should be noted that the embodiments described above are only some of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field and related fields without making creative work should fall within the scope of protection of the present invention. Modifications or equivalent replacements of the technical solutions of the present invention do not depart from the spirit and scope of the technical solutions of the present invention, and should all be included in the scope of the claims of the present invention.
Claims
1. Claim 1: A surface temperature and downscaled temperature inversion method based on remote sensing data, characterized in that: The following steps are involved: (1) Verification of fitting between measured surface temperature and inverted temperature: a. Using cubic spline interpolation, interpolate three-hourly temperature data from weather stations throughout the day into a continuous all-day temperature curve. b. Invert the Land Surface Temperature (LST) by combining the imaging time of Landsat 8 / 9 satellite images with the location of the weather stations. c. The coefficient of determination (R²), mean prediction error (MPE), root mean square prediction error (RMSE), and relative prediction error (RPE) were used to verify the fitting accuracy of the inverted temperature and the measured temperature. d. Eliminate outliers in the dataset based on the Cook distance method and optimize the inversion model; (2) Surface temperature downscaling calculation: a. Calculate the surface emissivity based on the NDVI threshold method, which is determined by the mapping relationship between vegetation cover (Pv) and NDVI value; b. Retrieve initial land surface temperature (LST) using Landsat 8 / 9 satellite data; c. Using a random forest model, we combined the NDVI, NDBI, and NDWI indices from Landsat 8 / 9 and Sentinel-2 satellite data to construct a regression relationship between surface temperature and multi-source remote sensing data. d. Apply the regression model to high-resolution Sentinel-2 data to generate downscaled refined surface temperature; (3) Verification of downscaling results: a. Use a linear regression model to evaluate the consistency of downscaled surface temperature with the original Landsat inverted temperature; b. Optimize the downscaling results by smoothing the regression residuals.
2. Claim 2: The method according to claim 1, characterized in that The interpolation interval of the cubic spline interpolation method described in step (1) is three hours, and the fitting accuracy meets the following conditions: R² ≥ 0.85, RMSE ≤ 1.5°C, RPE ≤ 5%.
3. Claim 3: The method according to claim 1, characterized in that The outlier rejection threshold of the Cook distance method described in step (1) for: Where: Indicates the Cook distance value corresponding to the i-th data record; It represents the LST value of the jth record obtained by Landsat inversion; Indicates to remove the i The first record obtained by refitting j The recorded Landsat-retrieved LST values; Indicates the total number of records of observation and inversion data; The Cook distance value of the data record is When it exceeds 4 / (np-1), it is determined to be an outlier, where n is the total number of records and p is the number of model parameters.
4. Claim 4: The method according to claim 1, characterized in that The surface emissivity described in step (2) The calculation formula is: .in, The calculation formula of vegetation coverage Pv is: NDVI_soil=0.05, NDVI_eg=0.70; when NDVI≥0.70, Pv=1; when NDVI≤0.05, Pv=0.
5. Claim 5: The method according to claim 1, characterized in that The input data of the random forest model in step (2) include: NDVI, NDBI, and NDWI indices from Landsat 8 / 9, multispectral reflectance from Sentinel-2, and surface emissivity ε.
6. Claim 6: The method according to claim 1, characterized in that The evaluation indicators of the linear regression model in step (3) include: The slope range is 0.95~1.05, the absolute value of the intercept is ≤0.5℃, and R² is ≥0.90.