City canopy meteorological simulation data correction method based on combined ML and modified WRF-BEM model

By combining machine learning and a modified WRF-BEM model, the errors and inhomogeneities of the WRF model in urban canopy meteorological data simulation were resolved, achieving high-precision correction of meteorological data and accurate prediction of extreme events, thus improving the scientific nature of urban environmental understanding and planning.

CN121980894APending Publication Date: 2026-05-05XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XI AN JIAOTONG UNIV
Filing Date
2025-11-25
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In existing technologies, WRF models suffer from large and uneven errors in simulating urban canopy meteorological data, making it difficult to perform differential correction using traditional methods, thus affecting the accuracy and reliability of meteorological data.

Method used

By combining machine learning (ML) and a modified WRF-BEM model, and through a comparative analysis framework of the integrated model and the base model, the optimal machine learning model is selected. Combined with the characteristics of local climate zones (LCZ), meteorological data is corrected, including accurate predictions of temperature, humidity, and wind speed.

Benefits of technology

It significantly reduces the error of meteorological data, improves the accuracy of temperature, humidity and wind speed simulations, enhances the robustness of machine learning models and their ability to capture extreme events, and provides a more accurate understanding of the urban canopy environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121980894A_ABST
    Figure CN121980894A_ABST
Patent Text Reader

Abstract

The invention discloses an urban canopy meteorological simulation data correction method based on a combined ML and a modified WRF-BEM model. The method is specifically implemented according to the following steps: step 1, obtaining observation data; step 2, acquiring analog data; 3, setting a machine learning model; 4, evaluating an initial simulation result; and 5, obtaining a machine learning result. The problems of correction effect error and large local deviation of meteorological data in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of urban canopy simulation meteorological data correction technology, specifically involving a correction method for urban canopy meteorological simulation data based on a combination of ML and modified WRF-BEM models. Background Technology

[0002] With the continuous increase in global urban populations, cities face a series of climate and environmental problems, such as urban heat islands, air pollution, and extreme weather. These environmental problems significantly impact the health of urban residents, increase energy and resource consumption, and pose challenges to sustainable urban development. Therefore, a deep understanding of urban climate and environmental characteristics, and the development of effective mitigation strategies and scientific urban planning recommendations, have become key research areas. However, achieving this goal requires accurate canopy meteorological simulation data at the urban scale as a foundation.

[0003] The development of the mesoscale model WRF (Weather Research and Forecasting Model) provides powerful tools and scientific support for this purpose. WRF can consider complex urban topography and land use types, simulating meteorological conditions in urban areas at high resolution (e.g., 1 km or smaller), including temperature, humidity, wind speed, and precipitation. These high-precision simulation data provide important evidence for studying urban climate phenomena such as the urban heat island effect and local circulation, and also lay a solid foundation for formulating scientific environmental policies and urban planning. Over the past decade, researchers have done much work to improve the accuracy of WRF simulations. For example, to more accurately describe the physical processes involved in heat, momentum, and water vapor exchange in the urban environment, researchers developed the Urban Canopy Model (UCM) and coupled it to the WRF model in 2004. UCM considers the influence of complex urban underlying surface morphology, including the shadows of urban canopy buildings, reflections of shortwave and longwave radiation, wind profiles in the canopy, and multi-layered heat transfer equations for roofs, walls, and road surfaces. Extensive research has confirmed that incorporating UCM can improve the correlation between WRF simulation data and observational data, and significantly reduce the RMSE of temperature and humidity, providing a more accurate reflection of the climate and environment in urban areas.

[0004] However, due to the current difficulty in effectively characterizing and describing the inherent complexity and uncertainty of the climate system, coupled with a lack of in-depth understanding of the dynamic processes of climate, the output variables of WRF models based on physical processes still have significant deviations compared to observational data, limiting their reliability in practical applications. To address this, some researchers have attempted to reduce simulation errors by artificially correcting the simulation data later. For example, Bhati et al., based on their earlier research, artificially reduced the temperature obtained from WRF simulations by 2.78°C and increased the relative humidity by 11.23%, reducing the simulation error by nearly half. However, this correction method has significant drawbacks: WRF simulation errors exhibit temporal and spatial variability, making it unreasonable to apply uniform amplitude correction to meteorological data at all times and locations within a city. Differentiated correction for meteorological data at different times and locations requires processing large amounts of simulation and observational data and establishing the mapping relationship between them, which exceeds the processing capabilities of traditional methods.

[0005] In recent years, machine learning (ML) has demonstrated significant advantages in meteorology and environmental science. Its powerful data processing and pattern recognition capabilities enable it to automatically learn and extract features from massive datasets, discovering complex mapping relationships between simulated and observed data. Furthermore, machine learning's ability to handle nonlinear relationships and its high computational efficiency help overcome the limitations of manual correction methods, achieving more accurate and efficient data correction. Recently, research combining ML models with wind-rendering simulations (WRF) has successfully corrected simulation results for rainfall, snowfall, ozone concentration, wind energy, and solar energy, providing more accurate guidance for local production and daily life. Although the application of ML in meteorology and environmental science is becoming increasingly widespread, research on ML correction of temperature, humidity, and wind speed within the urban canopy to reduce WRF simulation errors remains lacking. Case studies are urgently needed to explore the feasibility of this application.

[0006] Furthermore, previous studies have typically focused on improving prediction accuracy by optimizing machine learning (ML) models, with less attention paid to enhancing the correction effect of meteorological data by improving the quality of wind-rendered simulation data. More accurate WRF simulations exhibit lower systematic errors and provide more reliable spatiotemporal distributions of temperature, humidity, and wind speed. This helps ML models better learn the relationship between these characteristics and observed data, allowing them to focus more on correcting random errors and local biases, thereby improving correction effectiveness. Simultaneously, high-quality simulation data can enhance the robustness of ML models, enabling them to maintain good correction performance across different temporal and spatial scales. Therefore, the impact of WRF simulation quality on the prediction accuracy of ML models urgently needs further investigation. Summary of the Invention

[0007] The purpose of this invention is to provide a correction method for urban canopy meteorological simulation data based on a combination of ML and modified WRF-BEM models, which solves the problems of large errors in the correction effect and large local deviations of meteorological data in the prior art.

[0008] The technical solution adopted in this invention is a correction method based on urban canopy meteorological simulation data combining ML and modified WRF-BEM models, characterized by the following steps:

[0009] Step 1: Acquisition of observation data; Step 2: Acquiring simulation data; Step 3: Machine learning model setup; Step 4: Evaluation of initial simulation results; Step 5: Machine learning results.

[0010] The invention is further characterized in that, Step 1 is implemented in the following steps: The observation data consisted of hourly data from several meteorological stations distributed throughout the selected area throughout the year, including temperature and humidity measured at 2 meters and wind speed measured at 10 meters. Preprocessing of observation data: The preprocessing of the observation data included outlier detection and removal, as well as missing value imputation. Wind speed and rainfall did not show significant diurnal variation patterns, so linear imputation was used. Temperature and relative humidity had simple diurnal variation patterns, that is, the highest temperature and the lowest relative humidity usually occurred around noon, and the lowest temperature and the highest relative humidity occurred before sunrise. Therefore, while performing linear imputation on temperature and relative humidity, the diurnal variation pattern was also considered. That is, when the missing value was an extreme point, imputation was performed based on the diurnal variation pattern of the nearby days under the same weather conditions at that station. Finally, a total of 122,640 hourly meteorological observation data from several meteorological stations throughout the year were obtained.

[0011] Step 2 is implemented in the following steps: Based on simulated meteorological data acquired using WRF-ARW 4.5, the center of the WRF nested computational domain was determined. The model employs a triple-nested grid. The parent domain D1 contains the center of the first nested domain D2, aiming to capture the background field characteristics of large-scale weather systems. The parent domain D1 has a horizontal spatial resolution of 9 km and a grid size of 120×120 in the horizontal direction. The first nested domain D2 contains the center of D3, aiming to focus on characterizing the interaction between regional-scale meteorological fields and topographic effects. D2 has a resolution of 3 km and a grid size of 118×118 in the region. The second nested domain D3 is the observation area, i.e., the study domain, with a spatial resolution of 1 km and a grid size of 118×118 in the region. It aims to finely simulate the distribution characteristics of wind, humidity, and temperature fields at the urban scale. Vertically, all three nested domains use 45-layer grids, extending from the ground to a height of 5000 Pa. The vertical grid is relatively dense near the surface and gradually becomes sparser with increasing height.

[0012] In the WRF model, the physical parameterization process selected was as follows: Thompson scheme for cloud microphysics, RRTMG scheme for longwave / shortwave radiation, modified MM5 Monin-Obukhov scheme for near-surface layer processes, Noah-MP scheme for land surface processes, YSU scheme for boundary layer processes, and Kain-Fritsch scheme for cumulus convection. Simulation data were generated using the standard WRF model and the modified WRF-BEM model, respectively. The canopy parameters for each LCZ were set based on relevant architectural design documents and field survey results for the selected area.

[0013] Step 3 is implemented in the following steps: The optimal machine learning model was selected by constructing a comparative analysis framework between the ensemble model and the basic models. First, three basic models were trained: Random Forest (RF), Extreme Gradient Boosting (XGB), and Gradient Boosting Decision Tree (GBDT). Bayesian optimization was used to obtain the optimal predictive performance of each basic model under specific parameter configurations, i.e., predictive canopy meteorological data (temperature, humidity, and wind speed) that more closely approximate actual observations. Then, linear regression was performed based on the results of each basic model to obtain a final ensemble model. Machine learning was then performed separately for the three meteorological elements, and the mean absolute error (MAE), root mean square error (RMSE), and correlation coefficient (R²) were used. 2 As an evaluation metric for model performance, the optimal model is obtained through metric evaluation; We introduce the SHAP method from interpretable machine learning. Through SHAP values, we can intuitively see which features have a greater impact on the model's prediction results and how that impact trends are. All meteorological information was divided into a training set and a test set, with the training set accounting for 80% and the test set accounting for 20% of the total data, respectively. The input data for the machine learning model consists of nine features: WRF simulation temperature, humidity, wind speed, rainfall, station codes, and local LCZ feature variables, namely maximum building height, building proportion, grassland proportion, and tree proportion. Mapping relationships are established based on the temperature, humidity, and wind speed observed at the meteorological station to obtain the prediction results of the machine learning model.

[0014] Step 3, the specific parameter configuration, refers to the optimal combination of hyperparameters determined for each base model using the Bayesian optimization algorithm. For the Random Forest model, the optimization parameters include the number of trees (n_estimators), maximum depth (max_depth), minimum number of split samples (min_samples_split), minimum number of leaf samples (min_samples_leaf), and random feature selection strategy (max_features). For the XGBoost model, the focus is on optimizing the learning rate (learning_rate), boosting epochs (n_estimators), maximum tree depth (max_depth), sample and feature sampling ratios (subsample, colsample_bytree), and regularization parameters (reg_alpha, reg_lambda). For the GBDT model, the tuning parameters include boosting epochs (n_estimators), learning rate (learning_rate), maximum tree depth (max_depth), and sample sampling ratio (subsample). Bayesian optimization effectively searches the parameter space by constructing surrogate models, finding better parameter combinations with fewer evaluations compared to grid search and random search.

[0015] In step 3, the final ensemble model is constructed using the Stacking generalization method. The specific process consists of three steps: First, three basic models—Random Forest, XGBoost, and GBDT—are trained using their respective optimal parameter configurations. Then, the three trained basic models are used to predict the validation set, resulting in three sets of predictions as meta-features. Finally, using the predictions of these three basic models as input features and the true labels of the validation set as the target variable, a linear regression model is trained as a meta-learner. The linear regression model automatically learns the optimal weight coefficients of the predictions of each basic model. The final ensemble prediction result is obtained by weighted linearly combining the predictions of the three basic models.

[0016] Step 4 is implemented in the following steps: Using statistical parameters such as mean absolute error (MA), root mean square error (RMS), and Pearson correlation R02 The simulation results for air temperature, specific humidity, and wind speed were evaluated. Compared to the standard WRF model, WRF+BEM significantly reduced the simulation errors of each meteorological element and improved their correlation. The simulated air temperature and specific humidity showed a strong correlation with the measured values, while the correlation for wind speed was weak. Additionally, the accuracy of the simulated wind speed in wind class classification was evaluated using the TS index. The TS index is based on binary classification, and the TS calculation method is as follows:

[0017] In the calculation of wind speed TS, wind speeds falling within the wind level range are classified as positive, and those falling outside the range are classified as negative. According to the wind speed level classification, 0.1m / s to 0.2m / s is level 0 wind, 0.3m / s to 1.5m / s is level 1 wind, 1.6m / s to 3.3m / s is level 2 wind, and 3.4m / s to 5.4m / s is level 3 wind. If the TS value is higher than 0.6, it indicates that the model's prediction performance is relatively good; if the TS value is between 0.2 and 0.4, it indicates that the model's prediction performance is average; and if the TS value is lower than 0.2, it means that the accuracy is poor.

[0018] Step 5 is implemented in the following steps: 1) Temperature training and prediction errors Using ML respectively WRF and ML WRF-UCM This represents the machine learning results obtained using simulated data from the standard WRF model and the modified WRF-BEM model as training data; The temperature error was significantly reduced after processing with a machine learning statistical model. After correction, the MAE of the standard WRF simulated temperature results decreased by 1.08℃, the RMSE decreased by 1.35℃, and the R... 2 The value was increased to 0.97. Furthermore, the improved quality of WRF simulation data also led to improved temperature results. WRF-UCM Compared to ML WRF The MAE of the temperature was further reduced by 0.43°C, and the RMSE was reduced by 0.56°C. 2 Increased to 0.98; 2) Specific humidity training and prediction error After processing with a machine learning statistical model, the specific humidity error was significantly reduced. The MAE of the standard WRF simulated specific humidity results decreased by 0.72 g / kg, and the RMSE decreased by 1.03 g / kg. 2 The value was increased to 0.95. Furthermore, improvements in the quality of WRF simulation data also led to improvements in specific humidity results. WRF-UCM Compared to ML WRF The specific wet MAE was further reduced by 0.14 g / kg, RMSE was reduced by 0.21 g / kg, and R 2Increased to 0.97.

[0019] 3) Wind speed training and prediction error The wind speed error was significantly reduced after processing with a machine learning statistical model. The MAE of the standard WRF simulated wind speed results decreased by 1.47 m / s, the RMSE decreased by 1.93 m / s, and the R... 2 The value increased from 0.06 to 0.38. Furthermore, the improved quality of the WRF simulation data also led to improved wind speed results, although ML... WRF-UCM Compared to ML WRF The wind speed MAE and RMSE were only slightly reduced by 0.01 m / s and 0.02 m / s, respectively, but the correlation R of the data was significantly improved. 2 It has been significantly improved to 0.48.

[0020] The accuracy trends of the corrected simulated wind speed predictions for each wind level are consistent with the initial trends, with the highest TS score for level 1 winds. Specifically, compared to the standard WRF simulation results, machine learning correction significantly improved the prediction accuracy for level 1 winds, achieving a TS score as high as 0.57, demonstrating excellent accuracy. The TS scores for level 0 and level 2 winds also showed significant improvement, approaching 0.2. The TS score for level 3 winds remained very low after correction, indicating that the standard WRF model's simulation results for level 3 winds were poor, and the machine learning model's results were also poor. Furthermore, compared to ML... WRF ,ML WRF-UCM This resulted in a significant improvement in TS scores for all wind levels, especially for level 2 and level 3 winds.

[0021] The beneficial effects of this invention are: 1) In order to improve the quality of training results, this invention not only selects the optimal machine learning model by constructing a comparative analysis framework between the integrated model and the basic model, but also improves the quality of WRF simulation data by incorporating the modified BEM model, thereby improving the ML results.

[0022] Physical models, based on atmospheric dynamics, thermodynamics, and microphysics, yield results that conform to physical laws and possess clear physical meaning. They can explain the mechanisms behind meteorological phenomena. Therefore, improving physical models can reduce systematic errors, making WRF simulation results closer to real observations. The final errors may be simpler (e.g., primarily random errors), allowing machine learning models to be trained on more accurate data, thus more easily learning the patterns of error and improving the accuracy of ML results. Furthermore, since extreme weather events (e.g., strong winds, heavy rains) constitute a relatively small proportion of training data, machine learning typically performs poorly on these events. Researchers usually analyze extreme weather events by creating more accurate and reliable artificial intelligence models; however, this invention demonstrates that improving physical models can better reflect extreme events in simulation data, thereby improving the ability to capture extreme events in ML results. Figure 11 As shown, relying solely on machine learning (ML) correction did not improve the prediction accuracy of Force 3 winds. However, by incorporating a modified BEM model into the wind flow prediction (WRF) to more reasonably account for the varying obstruction effects of canopies at different heights, the TS score for Force 3 winds was significantly improved. Consequently, based on better simulated wind speed results, the ML correction results also demonstrated a better predictive ability for Force 3 winds. Therefore, in future research combining machine learning and WRF, it is recommended to simultaneously improve the ML model and the WRF physical processes to achieve more accurate meteorological data predictions.

[0023] 2) This invention, through interpretable machine learning, discovered that the physical characteristics of urban underlying surfaces under the LCZ framework have a certain impact on both temperature and humidity, and the trend of this impact is physically interpretable. This indicates that the LCZ concept is also applicable to ML research. More importantly, combining the LCZ concept facilitates researchers' deeper understanding of urban climate differences and their causes, and even provides new insights. For example, according to the SHAP results of this invention, building characteristic variables (building proportion, maximum building height) have a significantly greater impact on temperature and humidity than vegetation characteristic variables (grassland proportion, tree proportion). This indicates that the formation of unique urban canopy environments is closely related to urban buildings (e.g., local UHI), therefore, the solution to urban environmental problems should also start with urban buildings (e.g., mitigating the thermal effects of urban buildings to achieve thermal mitigation).

[0024] This invention demonstrates an effective combination of physical model improvement and ML model improvement in correcting urban canopy meteorological data, which is worth further reference in future research to help provide more accurate weather forecasts. Attached Figure Description

[0025] Figure 1This invention relates to the WRF nested computational domain in the correction method for urban canopy meteorological simulation data based on the combination of ML and modified WRF-BEM models. Figure 2 This is a flowchart of the machine learning process in the correction method for urban canopy meteorological simulation data based on the combination of ML and modified WRF-BEM model in this invention; Figure 3 It is the correlation coefficient of the initial simulation results in the correction method of urban canopy meteorological simulation data based on the combination of ML and modified WRF-BEM model in this invention; Figure 4 This is a comparison chart of TS scores for each wind level in the correction method of urban canopy meteorological simulation data based on the combination of ML and modified WRF-BEM models of this invention; Figure 5 This is a fitting graph (temperature) of predicted and observed data in the correction method of urban canopy meteorological simulation data based on the combination of ML and modified WRF-BEM model in this invention. Figure 6(a) is a pie chart showing the contribution of temperature prediction in the correction method based on urban canopy meteorological simulation data combining ML and modified WRF-BEM models in this invention. Figure 6(b) is a honeycomb diagram showing the contribution of temperature prediction in the correction method of urban canopy meteorological simulation data based on the combination of ML and modified WRF-BEM models in this invention. Figure 7 This invention relates to the correction method for urban canopy meteorological simulation data based on a combination of ML and modified WRF-BEM models, which uses the PDP of building proportion and maximum building height with respect to temperature. Figure 8 This is a fitting graph (specific humidity) of predicted and observed data in the correction method of urban canopy meteorological simulation data based on the combination of ML and modified WRF-BEM model in this invention. Figure 9(a) is a pie chart showing the contribution of specific humidity prediction in the correction method of urban canopy meteorological simulation data based on the combination of ML and modified WRF-BEM models in this invention. Figure 9(b) is a honeycomb diagram showing the contribution of specific humidity prediction in the correction method of urban canopy meteorological simulation data based on the combination of ML and modified WRF-BEM models in this invention. Figure 10(a) shows the building proportion and maximum building height in the correction method of urban canopy meteorological simulation data based on the combination of ML and modified WRF-BEM model in this invention; Figure 10(b) shows the PDP of grassland and tree proportions in the correction method of urban canopy meteorological simulation data based on the combination of ML and modified WRF-BEM model in this invention; Figure 11This invention relates to the correction method for urban canopy meteorological simulation data based on a combination of ML and modified WRF-BEM models, specifically the variation of the corrected wind speed result TS score with wind level. Detailed Implementation

[0026] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0027] This invention is a correction method based on urban canopy meteorological simulation data combining ML and modified WRF-BEM models, and is implemented according to the following steps: Step 1: Acquisition of observation data; Step 1 is implemented in the following steps: The observation data consisted of hourly data from several meteorological stations distributed throughout the selected area throughout the year, including temperature and humidity measured at 2 meters and wind speed measured at 10 meters. Preprocessing of observation data: The preprocessing of the observation data included outlier detection and removal, as well as missing value imputation. Wind speed and rainfall did not show significant diurnal variation patterns, so linear imputation was used. Temperature and relative humidity had simple diurnal variation patterns, that is, the highest temperature and the lowest relative humidity usually occurred around noon, and the lowest temperature and the highest relative humidity occurred before sunrise. Therefore, while performing linear imputation on temperature and relative humidity, the diurnal variation pattern was also considered. That is, when the missing value was an extreme point, imputation was performed based on the diurnal variation pattern of the nearby days under the same weather conditions at that station. Finally, a total of 122,640 hourly meteorological observation data from several meteorological stations throughout the year were obtained.

[0028] Step 2: Acquiring simulation data; Step 2 is implemented in the following steps: Based on simulated meteorological data acquired using WRF-ARW 4.5, the center of the WRF nested computational domain was determined. The model adopts a triple-nested grid. The parent domain D1 contains the center of the first nested domain D2, aiming to capture the background field characteristics of large-scale weather systems. The parent domain D1 has a horizontal spatial resolution of 9 km and a grid size of 120×120 in the horizontal direction. The first nested domain D2 contains the center of D3, aiming to focus on characterizing the interaction between the regional-scale meteorological field and topographic effects. D2 has a resolution of 3 km and a grid size of 118×118 in the region. The second nested domain D3 is the observation area, i.e., the study domain, with a spatial resolution of 1 km and a grid size of 118×118 in the region. It aims to finely simulate the distribution characteristics of wind, humidity, and temperature fields at the urban scale. In the vertical direction, all three nested domains use 45-layer grids, extending from the ground to a height of 5000 Pa. The vertical grid is relatively dense near the ground surface and gradually becomes sparser with increasing height.

[0029] In the WRF model, the physical parameterization process selected was as follows: Thompson scheme for cloud microphysics, RRTMG scheme for longwave / shortwave radiation, modified MM5 Monin-Obukhov scheme for near-surface layer processes, Noah-MP scheme for land surface processes, YSU scheme for boundary layer processes, and Kain-Fritsch scheme for cumulus convection. Simulation data were generated using the standard WRF model and the modified WRF-BEM model, respectively. The canopy parameters for each LCZ were set based on relevant architectural design documents and field survey results for the selected area.

[0030] Step 3: Machine learning model setup; Step 3 is implemented in the following steps: The optimal machine learning model was selected by constructing a comparative analysis framework between the ensemble model and the basic models. First, three basic models were trained: Random Forest (RF), Extreme Gradient Boosting (XGB), and Gradient Boosting Decision Tree (GBDT). Bayesian optimization was used to obtain the optimal predictive performance of each basic model under specific parameter configurations, i.e., predictive canopy meteorological data (temperature, humidity, and wind speed) that more closely approximate actual observations. Then, linear regression was performed based on the results of each basic model to obtain a final ensemble model. Machine learning training was then conducted separately for the three meteorological elements (temperature, specific humidity, and wind speed), using the mean absolute error (MAE), root mean square error (RMSE), and correlation coefficient (R²). 2 As an evaluation metric for model performance, the optimal model is obtained through metric evaluation; After completing model training and performance evaluation, in order to further understand the impact of each feature on the model prediction results, this invention introduces the SHAP method in interpretable machine learning. Through the SHAP value, we can intuitively see which features have a greater impact on the model prediction results and what the trend of the impact is. (Taking Figure 9(b) as an example, the larger the SHAP value, the greater the degree of impact. In addition, the red dot to the right of the central axis / the blue dot to the left of the central axis indicates that the feature has a positive correlation with the prediction results, and vice versa. If there is no such trend, it indicates that the trend of the impact is unclear.)

[0031] All meteorological information was divided into a training set and a test set, with the training set accounting for 80% and the test set accounting for 20% of the total data, respectively. The input data for the machine learning model consists of nine features: WRF simulation temperature, humidity, wind speed, rainfall, station codes, and local LCZ feature variables, namely maximum building height, building proportion, grassland proportion, and tree proportion. Mapping relationships are established based on the temperature, humidity, and wind speed observed at the meteorological station to obtain the prediction results of the machine learning model.

[0032] Step 3, the specific parameter configuration, refers to the optimal combination of hyperparameters determined for each base model using the Bayesian optimization algorithm. For the Random Forest model, the optimization parameters include the number of trees (n_estimators), maximum depth (max_depth), minimum number of split samples (min_samples_split), minimum number of leaf samples (min_samples_leaf), and random feature selection strategy (max_features). For the XGBoost model, the focus is on optimizing the learning rate (learning_rate), boosting epochs (n_estimators), maximum tree depth (max_depth), sample and feature sampling ratios (subsample, colsample_bytree), and regularization parameters (reg_alpha, reg_lambda). For the GBDT model, the tuning parameters include boosting epochs (n_estimators), learning rate (learning_rate), maximum tree depth (max_depth), and sample sampling ratio (subsample). Bayesian optimization effectively searches the parameter space by constructing surrogate models, finding better parameter combinations with fewer evaluations compared to grid search and random search.

[0033] In step 3, the final ensemble model is constructed using the Stacking generalization method. The process involves three steps: First, three base models—Random Forest, XGBoost, and GBDT—are trained separately using their optimal parameter configurations. Then, each of the three trained base models is used to predict the validation set, resulting in three sets of predictions as meta-features. Finally, using these three sets of predictions as input features and the true labels of the validation set as the target variable, a linear regression model is trained as the meta-learner. The linear regression model automatically learns the optimal weight coefficients of each base model's predictions. The final ensemble prediction result is obtained by weighted linearly combining the predictions of the three base models. This method fully leverages the advantages of different algorithms, improving overall prediction performance by learning the optimal combination weights. Compared to simple average ensembles, it has stronger adaptability and better generalization ability. The ensemble model aims to utilize the diversity and complementarity of different models to obtain a more robust and accurate prediction model. Because the ensemble model can integrate the prediction information of multiple base models, it captures more data features and patterns, exhibiting higher prediction accuracy and stability in most cases. However, in certain specific scenarios, due to the combined effects of factors such as dataset characteristics, model complexity, and computational resources, the base model may exhibit performance comparable to or even better than the ensemble model. Therefore, it is necessary to select the best-performing model.

[0034] Step 4: Evaluation of initial simulation results; Step 4 is implemented in the following steps: Using statistical parameters such as mean absolute error (MA), root mean square error (RMS), and Pearson correlation R0 2 The simulation results for air temperature, specific humidity, and wind speed were evaluated. Compared to the standard WRF model, WRF+BEM significantly reduced the simulation errors of each meteorological element and improved their correlation. The simulated air temperature and specific humidity showed a strong correlation with the measured values, while the correlation for wind speed was weak. Additionally, the accuracy of the simulated wind speed in wind class classification was evaluated using the TS index. The TS index is based on binary classification, and the TS calculation method is as follows:

[0035] In the calculation of wind speed TS, wind speeds falling within the wind level range are classified as positive, and those falling outside the range are classified as negative. According to the wind speed level classification, 0.1m / s to 0.2m / s is level 0 wind, 0.3m / s to 1.5m / s is level 1 wind, 1.6m / s to 3.3m / s is level 2 wind, and 3.4m / s to 5.4m / s is level 3 wind. If the TS value is higher than 0.6, it indicates that the model's prediction performance is relatively good; if the TS value is between 0.2 and 0.4, it indicates that the model's prediction performance is average; and if the TS value is lower than 0.2, it means that the accuracy is poor.

[0036] Step 5: Machine learning results.

[0037] Step 5 is implemented in the following steps: 1) Temperature training and prediction errors Using ML respectively WRF and ML WRF-UCM This represents the machine learning results obtained using simulated data from the standard WRF model and the modified WRF-BEM model as training data; The temperature error was significantly reduced after processing with a machine learning statistical model. After correction, the MAE of the standard WRF simulated temperature results decreased by 1.08℃, the RMSE decreased by 1.35℃, and the R... 2 The value was increased to 0.97. Furthermore, the improved quality of WRF simulation data also led to improved temperature results. WRF-UCM Compared to ML WRF The MAE of the temperature was further reduced by 0.43°C, and the RMSE was reduced by 0.56°C. 2 Increased to 0.98.

[0038] 2) Specific humidity training and prediction error After processing with a machine learning statistical model, the specific humidity error was significantly reduced. The MAE of the standard WRF simulated specific humidity results decreased by 0.72 g / kg, and the RMSE decreased by 1.03 g / kg. 2 The value was increased to 0.95. Furthermore, improvements in the quality of WRF simulation data also led to improvements in specific humidity results. WRF-UCM Compared to ML WRF The specific wet MAE was further reduced by 0.14 g / kg, RMSE was reduced by 0.21 g / kg, and R 2 Increased to 0.97.

[0039] 3) Wind speed training and prediction error The wind speed error was significantly reduced after processing with a machine learning statistical model. The MAE of the standard WRF simulated wind speed results decreased by 1.47 m / s, the RMSE decreased by 1.93 m / s, and the R... 2 The value increased from 0.06 to 0.38. Furthermore, the improved quality of the WRF simulation data also led to improved wind speed results, although ML... WRF-UCM Compared to ML WRF The wind speed MAE and RMSE were only slightly reduced by 0.01 m / s and 0.02 m / s, respectively, but the correlation R of the data was significantly improved. 2It has been significantly improved to 0.48.

[0040] The accuracy trends of the corrected simulated wind speed predictions for each wind level are consistent with the initial trends, with the highest TS score for level 1 winds. Specifically, compared to the standard WRF simulation results, machine learning correction significantly improved the prediction accuracy for level 1 winds, achieving a TS score as high as 0.57, demonstrating excellent accuracy. The TS scores for level 0 and level 2 winds also showed significant improvement, approaching 0.2. The TS score for level 3 winds remained very low after correction, indicating that the standard WRF model's simulation results for level 3 winds were poor, and the machine learning model's results were also poor. Furthermore, compared to ML... WRF ,ML WRF-UCM This resulted in a significant improvement in TS scores for all wind levels, especially for level 2 and level 3 winds.

[0041] Example 1 To fill the aforementioned research gap, this invention: (1) Based on hourly WRF data of Xi'an city in 2023 and data from 14 meteorological stations, this study explored the feasibility of using ML models to correct the urban canopy WRF simulation of temperature, humidity and wind speed under different time and space conditions. In addition, in order to better characterize the influence of urban underlying surface features on canopy meteorological elements, this invention incorporates urban underlying surface physical features based on the local climate zone (LCZ) framework into the ML model for the first time, in order to improve the accuracy, robustness and generalizability of machine learning models in predicting dynamic and complex meteorological elements, and to facilitate the understanding and interpretation of data; (2) Using simulation data generated by different WRF models (standard WRF model and modified WRF-BEM) as training data for machine learning, the influence of accurate UCM on the correction of WRF simulation results by ML model was investigated, which confirmed the importance of improving WRF physical models in applying machine learning to correct meteorological data.

[0042] (3) The contribution of each feature to the model prediction results was quantified by interpretable machine learning, and its influence trend was visualized with PDP to further understand the influence of each feature on the model prediction results, laying a scientific foundation for understanding the causes of the unique canopy environment of the city.

[0043] This invention aims to provide more accurate wind-humidity-thermal data for urban canopies by combining ML and WRF, thereby providing scientific support for understanding the urban canopy environment and formulating scientific environmental policies and urban planning.

[0044] 1) Acquisition of observation data: The research of this invention was conducted in Xi'an, a mega-city in Northwest China. The observation data consisted of hourly data from 14 meteorological stations distributed throughout Xi'an city throughout 2023, including temperature and humidity measured at 2 meters and wind speed measured at 10 meters.

[0045] The preprocessing of observation data in this invention includes outlier detection and removal, as well as missing value imputation. Outliers (data value 999999) typically occur for two reasons: one is that the data was not recorded at a certain moment for some reason, in which case only the data at that moment is abnormal, while the data near that moment is normal; the other is that the sensor malfunctioned for some reason, recording an outlier at that moment, in which case the data near that moment may also be abnormal, requiring manual judgment based on the diurnal and seasonal variation patterns of the data. For missing value imputation, since missing values ​​are usually a small number, one to a few data points, and the total amount of missing data throughout the year is small, only about 0.3% of the total data volume, this invention performs missing value imputation in a simple way: wind speed and rainfall do not have significant diurnal variation patterns, so a linear imputation method is used. Temperature and relative humidity have simple diurnal variation patterns, that is, the highest temperature and lowest relative humidity usually occur around noon, and the lowest temperature and highest relative humidity occur before sunrise. Therefore, while performing linear interpolation, the diurnal variation pattern was also considered. That is, when the missing value was an extreme point, interpolation was performed based on the diurnal variation pattern of the nearby days under the same weather conditions at that station. Finally, a total of 122,640 hourly meteorological observation data from 14 meteorological stations from 0:00 on January 1, 2023 to 0:00 on January 1, 2024 were obtained.

[0046] Example 2 2) Acquisition of simulation data: This invention acquires simulated meteorological data based on WRF-ARW 4.5. The WRF nested computational domain is as follows: Figure 1As shown, the computational domain is centered at the Bell Tower in Xi'an (34.26º N, 108.94º E). The model employs a triple-nested grid. The parent domain (D1) covers parts of Shaanxi, the Qinling Mountains, and the Qinghai-Tibet Plateau, with a horizontal spatial resolution of 9 km and a grid size of 120×120, aiming to capture the background field characteristics of large-scale weather systems. The first nested domain (D2) has a resolution of 3 km and a grid size of 118×118, focusing on characterizing the interaction between regional-scale meteorological fields and topographic effects. The second nested domain (D3) is the research area of ​​this invention, encompassing Xi'an and its surrounding areas, with a spatial resolution of 1 km and a grid size of 118×118, aiming to finely simulate the distribution characteristics of wind, humidity, and temperature fields at the urban scale. Vertically, all three nested domains use a 45-layer grid, extending from the ground to a height of 5000 Pa. The vertical grid is denser near the surface, gradually thinning out with increasing altitude. The red dots in the figure indicate the locations of meteorological stations.

[0047] The LCZ map is a 100-meter resolution global land cover dataset created by Demuzere et al. The simulation period was from 00:00 on January 1, 2023 to 00:00 on January 1, 2024, ultimately yielding 122,640 hourly meteorological data points (temperature, specific humidity, wind speed, and precipitation) from 14 meteorological stations, corresponding one-to-one with the observed data. The physical parameterization process in the WRF model selected currently widely accepted and adopted schemes: the Thompson scheme for cloud microphysics, the RRTMG scheme for longwave / shortwave radiation, the modified MM5 Monin-Obukhov scheme for near-surface processes, the Noah-MP scheme for land surface processes, the YSU scheme for boundary layer processes, and the Kain-Fritsch scheme for cumulus convection. Since the quality of training data is closely related to the quality of machine learning training results, this invention uses a standard WRF model (without canopy parameterization scheme) and a modified WRF-BEM model to generate simulation data to explore the impact of the quality of WRF simulation data on machine learning training results. The canopy parameters of each LCZ are set according to relevant documents of Xi'an City's architectural design and field survey results. The specific parameter settings are shown in Tables 1 and 2.

[0048] The governing equations of the large eddy simulation model include:

[0049] Momentum equation:

[0050] Temperature convection-diffusion equation:

[0051] Water vapor convection-diffusion equation:

[0052] In the formula: t Indicates time in seconds; These are respectively horizontal, longitudinal, and vertical; These are the filtered transverse, spanwise, and vertical velocity components (m·s). -1 ; Air density is expressed in kg·m³. -3 ; The pressure after filtration is expressed as kg·m. -1 ·s -2 ; Indicates the temperature after filtration in K; Represents the thermal buoyancy term, where T 0 represents the reference temperature; g represents the gravitational acceleration in m·s². -2 ; The kinematic viscosity of air / m 2 ·s -1 ; Specific humidity after filtration / g·kg -1 ; Pr Represent Prandtl numbers; Sc Represents the Schmitt number; Represents the resistance source terms of buildings and vegetation; Indicates heat source / sink; Indicates water vapor source / sink; Let represent the subgrid stress tensor, heat flux, and water vapor flux, respectively; Subgrid stress tensor Modeling was performed using the universal scale Smagorinsky model (UMSM); the modeling formulas for heat flux and water vapor flux are as follows:

[0053]

[0054] In the formula, This represents the subgrid eddy viscosity coefficient, determined by UMSM; Denotes the sub-lattice Prandtl number; Denotes the sub-grid Schmitt number; Correction of natural underlying surface wind profile parameters in land surface models; In the Noah-Mp land surface model, near-surface wind speeds are represented by logarithmic wind profiles. To take control:

[0055] In the formula: Represents frictional speed / m·s -1 ; represents the Karman constant; Indicates zero-plane displacement height in meters; Indicates roughness length in meters (m). The most important parameter in wind profiles is the vegetation layer. and , for natural LCZ and Quantification and correction are required. Determined by average momentum loss :

[0056] Subtract the height In a logarithmic height coordinate system, the straight line fitted to the wind profile can be extrapolated to the height of the zero wind speed point to obtain the result. .

[0057] WRF model optimization is divided into BEM model optimization and land surface model optimization.

[0058] The specific optimizations for BEM mode are as follows: After reasonably considering the dynamic-thermal effects of vegetation layer using the integrated method, the average simulated air temperature on sunny summer days decreased by 0.28℃, and the RMSE decreased by 0.44℃. Specifically, the average air temperature in the built LCZ decreased by 0.34℃, and the average air temperature in the natural LCZ decreased by 0.16℃. This indicates that although modifying the heat source term in the BEM model only affects the temperature-vapor source / sink of the built LCZ, the air temperature throughout the study area decreased due to air advection. The average specific humidity increased by 0.36 g / kg, and the RMSE decreased by 0.45 g / kg. The average wind speed decreased by 0.05 m / s, and the wind speed RMSE decreased by 0.09 m / s. The corrections for the latent / sensible heat of vegetation layer in WRF-BEM under other weather conditions and seasons are as follows: Weather conditions affect the Bonnby and solar radiation of vegetation, which in turn affect the sensible and latent heat fluxes of the LCZ vegetation layer. Compared with the results under sunny conditions, the vegetation Bonn ratio did not change significantly under other weather conditions. Therefore, it is assumed that weather conditions do not affect the distribution of sensible and latent heat of vegetation by influencing the vegetation Bonn ratio. The amount of solar radiation reaching the Earth's surface is affected by cloud cover / thickness, which in turn affects the sensible / latent heat flux of the vegetation layer. Since the solar radiation reaching the Earth's surface is directly accessed in WRF, we can simply parameterize the sensible / latent heat flux of vegetation under different weather conditions by measuring the changes in the amount of solar radiation reaching the Earth's surface. The sensible heat / latent heat of vegetation under any weather conditions can be simply expressed as:

[0059]

[0060] In the formula: This represents the sensible and latent heat of vegetation on any given date / W·m -2 ; These represent the sensible heat and latent heat of vegetation under sunny conditions, respectively, in W·m. -2 ; This represents the solar radiation reaching the Earth's surface on a sunny day / W·m -2 ; This represents the solar radiation reaching the Earth's surface on any given date / W·m -2 ; The effects of seasonal variation on vegetation sensible / latent heat are assessed through three aspects: shade area, Born ratio, and leaf area index. June to August is defined as summer, March to May and September to November as spring and autumn, and December to February as winter. It is assumed that these three effects are consistent within the same season, differing only between different seasons. The shade area, Born ratio, and leaf area index for each season are obtained or set as follows: 1) The shadow area increases as the solar altitude angle decreases. Therefore, the shadow area is greater in winter than in spring and autumn than in summer. Modeling software was used to obtain the shadow area for different seasons. The average values ​​of April 15 and October 15 were used for spring and autumn, and the average value of January 15 was used for winter. 2) The Born ratio is related to vegetation growth. In summer, vegetation grows vigorously, evapotranspiration reaches its peak, latent heat flux dominates, and the Born ratio is small. In spring and autumn, vegetation grows slowly, latent heat flux decreases compared to summer, and the Born ratio increases. In winter, vegetation withers, evaporation almost stops, sensible heat flux dominates, and the Born ratio is large. The Born ratio for different seasons is obtained through observation. 3) Vegetation leaf area index: LAI is set to 3 in spring and autumn, and 0.3 in winter.

[0061] The specific optimizations to the land surface model are as follows: Modify vegetation parameters in land surface model and Afterwards, the average wind speed and wind speed error on sunny summer days decreased significantly. In addition, the change in the wind field slightly affected the temperature field, and the average temperature and error decreased slightly, but the specific humidity did not change. After correction, the difference between the average wind speed and the measured value of the natural LCZ was greatly reduced, and the wind speed RMSE decreased significantly. In addition, the wind speed of the built LCZ was also slightly affected, and the average value and error were reduced to a certain extent. After correction, there was no significant difference between the simulated wind speed and the measured wind speed of the built LCZ and the natural LCZ. Similarly, the corrections to the land surface model regarding the dynamic parameters of the natural LCZ vegetation layer for other seasons are as follows: By correcting the vegetation parameters of the BEM and Noah-mp land surface models in the WRF model, the dynamic-thermal effects of the urban vegetation layer are considered more reasonably, significantly improving the simulation accuracy of the WRF model. The impact of seasonal variation on vegetation dynamics parameters during land surface processes is mainly reflected in the change of leaf area index (LAI). As leaves wither, the LAI decreases, and the obstructive effect of vegetation on the wind field weakens. Similarly, using a large eddy model to simulate various scenarios, wind profiles under different conditions are obtained, yielding the natural LCZ values. and The range.

[0062] Table 1. Parameter settings for each LCZ canopy

[0063] Table 2 Building Height Settings for Each LCZ

[0064] Example 3 3) Machine learning model setup: This invention selects the optimal machine learning model by constructing a comparative analysis framework between an ensemble model (Stacking Regressor) and a base model. The flowchart is as follows: Figure 2 As shown.

[0065] First, this invention trained three basic models: Random Forest Regressor (RF), Extreme Gradient Boosting Regressor (XGB), and Gradient Boosting Regressor (GBDT). Bayesian optimization was used to obtain the optimal predictive performance of each basic model under specific parameter configurations. Then, linear regression was performed based on the results of the three basic models to obtain a final ensemble model. The ensemble model aims to leverage the diversity and complementarity of different models to obtain a more robust and accurate predictive model. Because the ensemble model can integrate the predictive information of multiple basic models, it captures more data features and patterns, and in most cases exhibits higher predictive accuracy and stability. However, in certain specific scenarios, due to the combined influence of factors such as dataset characteristics, model complexity, and computational resources, the basic models may exhibit performance comparable to or even better than the ensemble model. To select the best-performing model, machine learning training was performed on three meteorological elements, and the mean absolute error (MAE), root mean square error (RMSE), and correlation coefficient (R²) were used. 2This metric serves as an evaluation indicator for model performance. Through metric evaluation, the selection of the optimal model is derived.

[0066] After completing model training and performance evaluation, to further understand the impact of each feature on the model's prediction results, this invention introduces the SHAP (SHapley Additive exPlanations) method from interpretable machine learning. Based on the Shapley value concept in cooperative game theory, the SHAP method can quantify the contribution of each feature to the model's prediction results, providing global and local interpretability. That is, through SHAP values, one can intuitively see which features have a greater impact on the model's prediction results and what the trend of this impact is.

[0067] In this invention, all meteorological information is divided into two parts: a training set and a test set. The training set is used to iterate the machine learning model using simulated data and actual measurement data, adjusting its parameters to derive statistical mapping relationships. The test set is used to apply these statistical relationships obtained through machine learning to correct other simulated data that were not involved in the training, and to compare them with actual measurement data to verify the accuracy of the statistical relationships. The data allocation ratio for the training set and the test set is 80% and 20% of the total data, respectively.

[0068] In addition to WRF simulations of temperature, humidity, and wind speed, this invention also selects rainfall as a covariate in the machine learning model to improve the correction effect. Furthermore, this invention incorporates physical feature variables from the LCZ framework to improve the accuracy, robustness, and generalizability of the machine learning model in predicting dynamic and complex canopy meteorological elements, and to facilitate data understanding and interpretation. To reduce the complexity of the ML model, this invention only selects LCZ features that have a significant impact on temperature, humidity, and wind speed, including building proportion, maximum building height, grassland proportion, and tree proportion. In addition, the influence of meteorological station location on meteorological parameters is considered. This is because meteorological parameters at a meteorological station are affected not only by the underlying surface structure but also by the background meteorological conditions and the surrounding area. However, since the number of meteorological stations is insufficient to further classify and discuss the influence of their geographical location based on background meteorological conditions and the surrounding area, this invention simply sets up a column of feature values ​​called "station codes," assigning values ​​from 0 to 13 to the 14 meteorological stations to express the influence of geographical location on meteorological parameters.

[0069] Ultimately, the input data for the machine learning model of this invention consists of nine features: WRF simulated temperature, humidity, wind speed, rainfall, station code, and local LCZ feature variables (maximum building height, building proportion, grassland proportion, and tree proportion). Mapping relationships are established based on the temperature, humidity, and wind speed observed at the meteorological station to obtain the prediction results of the machine learning.

[0070] Example 4 Evaluation of initial simulation results: To evaluate the effectiveness of machine learning in correcting simulated data, this invention first analyzes the errors in the initial WRF simulation data. This invention uses statistical parameters such as mean absolute error (MAE), root mean square error (RMSE), and Pearson correlation coefficient (R²). 2 This was used to evaluate the simulation results for air temperature, specific humidity, and wind speed. See Table 3 and... Figure 3 As shown Table 3 Errors in the initial WRF simulation results

[0071] Compared to the results of the standard WRF model, WRF+BEM significantly reduced the simulation errors of various meteorological elements and improved their correlation. Furthermore, it can be seen that there is a strong correlation between the simulated air temperature and specific humidity and the measured values, while the correlation of wind speed is very weak. This is because neither observed nor simulated wind speed exhibits significant diurnal or seasonal variations, and because they are instantaneous values, the fluctuations are large and random, making it difficult to correlate them. Therefore, this invention additionally uses the TS (ThreatScore) index to evaluate the accuracy of simulated wind speed in wind class classification, as accurate prediction of wind speed levels is crucial for guiding people's daily outdoor activities and many applications (such as wind energy assessment, disaster assessment, etc.). The TS index is based on binary classification, as shown in Table 4. Table 4. Binary Classification Confusion Matrix

[0072] The TS calculation method is as follows:

[0073] In the calculation of wind speed TS, wind speeds falling within the wind level range are classified as positive, and those falling outside the range are classified as negative. According to the wind speed classification, 0.1 m / s to 0.2 m / s is level 0 wind, 0.3 m / s to 1.5 m / s is level 1 wind, 1.6 m / s to 3.3 m / s is level 2 wind, and 3.4 m / s to 5.4 m / s is level 3 wind (Xi'an is located in a calm zone; based on observed wind speeds, this invention only classifies winds up to level 3). Figure 4 The TS scores are the initial simulation results for different wind levels.

[0074] A good TS value is typically between 0.4 and 0.6; a TS value higher than 0.6 indicates excellent model forecasting performance; a TS value between 0.2 and 0.4 indicates average model forecasting performance; and a TS value below 0.2 indicates poor accuracy. It can be seen that the standard WRF model's simulation results only have average accuracy in predicting level 1 winds, while the accuracy in predicting level 0 and level 3 winds is extremely poor. By reasonably considering the drag effect of urban canopy buildings and vegetation, the inclusion of UCM significantly improves the prediction accuracy for all wind levels, especially level 1 winds.

[0075] Example 5 Machine learning results: 1) Temperature training and prediction errors Table 5 shows the test set temperature prediction error indices for each basic model and the ensemble model. This invention uses ML... WRF and ML WRF-UCM This represents the machine learning results obtained using simulated data from the standard WRF model and the modified WRF-BEM model as training data.

[0076] Table 5. Temperature values ​​MLWRF and MLWRF-UCM for each machine model

[0077] Table 5 shows that the ensemble model performs best across all metrics in the test set, making it the optimal model among the four. Compared to Table 3, the temperature error is significantly reduced after processing with the machine learning statistical model. After correction, the MAE of the standard WRF simulated temperature results decreased by 1.08℃, and the RMSE decreased by 1.35℃ (error reduced by approximately 43%). 2 The value was increased to 0.97. Furthermore, the improved quality of WRF simulation data also led to improved temperature results, ML... WRF-UCM Compared to ML WRF The MAE of the temperature was further reduced by 0.43°C, and the RMSE was reduced by 0.56°C (the error was further reduced by approximately 18%). 2 Increased to 0.98.

[0078] Figure 5 A fitted graph of temperature prediction data and observation data obtained through machine learning (ML) WRF-UCM (Results of the integrated model).

[0079] The top and right sides of the fitting plot show the number of observed and predicted data points in different temperature ranges, respectively. Blue data represents the training set, and orange data represents the test set. The fitting plot shows that the performance of the ensemble model on the training and test sets is not significantly different, indicating no overfitting (overfitting refers to a model performing well on the training set but poorly on the test set), confirming the model's accuracy in processing new meteorological data. Furthermore, the data points in the test set, whether for summer high temperatures or winter low temperatures, are generally distributed around the Y=X line, meaning that the predicted and observed data fit well across different times and spaces. This demonstrates the feasibility and accuracy of using machine learning for temperature data correction in various temporal and spatial contexts.

[0080] Figures 6(a) and 6(b) show the weights and trends (SHAP values) of various meteorological elements, LCZ feature values, and geographical location on temperature prediction obtained using the interpretable machine learning SHAP method during the temperature prediction process.

[0081] It can be seen that LCZ-related features have a significant impact on temperature prediction, confirming the necessity of incorporating LCZ feature values ​​into meteorological parameter prediction in this invention. Among them, the proportion of buildings has the greatest impact, followed by the maximum building height, while the proportion of trees and grassland has a very small impact on temperature prediction. It is evident that the increase / decrease of the natural underlying surface on regional temperature is negligible compared to the increase / decrease of urban buildings. Figure 7 A Partial Dependence Plot (PDP) is generated to show the marginal effect of features on the prediction results of machine learning models, representing the proportion of buildings and their maximum height as a function of temperature.

[0082] A higher number in PDP indicates a higher value for that parameter (here, temperature). Consistent with the trend in the honeycomb plot in Figure 6, the building proportion has a positive correlation with temperature. This is easily explained because an increased building proportion increases the capture of radiation within the street valley and also increases anthropogenic heat, leading to higher temperatures. However, the maximum building height does not show a significant trend in its impact on temperature. This is likely because building height exhibits strong spatial heterogeneity; for the same maximum building height, the distribution of internal building heights may differ, making it difficult to reflect its influence on temperature.

[0083] Example 6 Specific humidity training and prediction error Table 6 shows the test set specific humidity prediction error indices for each basic model and the integrated model.

[0084] Table 6. Specific humidity MLWRF and MLWRF-UCM for each machine model

[0085] Table 6 shows that the integrated model performs best in all metrics on the test set, making it the optimal model among the four. Compared to Table 3, the specific humidity error is significantly reduced after processing with the machine learning statistical model. The MAE of the standard WRF simulated specific humidity results decreased by 0.72 g / kg, and the RMSE decreased by 1.03 g / kg (error reduction of approximately 48%). 2 The value was increased to 0.95. Furthermore, the improved quality of the WRF simulation data also led to improvements in the specific wetted water results. WRF-UCM Compared to ML WRF The specific wet MAE was further reduced by 0.14 g / kg, and the RMSE was reduced by 0.21 g / kg (the error was further reduced by approximately 10%). 2 Increased to 0.97.

[0086] Figure 8 A fitted graph of specific humidity prediction data and observation data obtained through machine learning (ML) WRF-UCM (Results of the integrated model).

[0087] As can be seen from the data scattering of the test set in the fitting graph, whether it is the high specific humidity data in summer or the low specific humidity data in winter, they are basically distributed around the oblique line Y=X. This shows that using machine learning to correct specific humidity data is feasible and accurate in different times and spaces.

[0088] Figures 9(a) and 9(b) show the weights and trends (SHAP values) of various meteorological elements, LCZ feature values, and geographical location influence on specific humidity prediction, obtained using the interpretable machine learning SHAP method during the specific humidity prediction process.

[0089] As can be seen, LCZ-related features also have a significant impact on the prediction of specific humidity. Among them, the building proportion has the greatest impact, followed by the maximum building height, while the tree proportion and grassland proportion have very little impact on the specific humidity prediction. It is evident that the impact of the increase / decrease of the natural underlying surface on the regional specific humidity is negligible compared to the impact of the increase / decrease of urban buildings. Figure 10 shows the PDP of building proportion and maximum building height, as well as grassland proportion and tree proportion, on specific humidity.

[0090] Consistent with the trend shown in the honeycomb plot in Figure 9(b), specific humidity decreases with increasing building proportion. This is because impermeable buildings allow only a very limited amount of water vapor evaporation; therefore, increasing the building proportion relatively reduces the proportion of the natural underlying surface and the water vapor evaporation from the natural underlying surface, leading to a decrease in specific humidity. The effect of maximum building height is relatively insignificant. Grassland allows for significant water evaporation, increasing specific humidity; therefore, specific humidity increases with increasing grassland proportion. Trees do not show a significant trend in their effect on specific humidity because evaporation from trees occurs at a certain height; therefore, compared to grassland, trees have a less significant effect on specific humidity at pedestrian height (2 meters).

[0091] Example 7 3) Wind speed training and prediction error Table 7 shows the test set wind speed prediction error indices for each basic model and the integrated model.

[0092] Table 7 shows the wind speeds MLWRF and MLWRF-UCM for each machine model.

[0093] As shown in Table 7, the integrated model exhibits the best performance across all metrics in the test set, making it the optimal model among the four. Compared to Table 3, the wind speed error is significantly reduced after processing with the machine learning statistical model. The MAE of the standard WRF simulated wind speed results decreased by 1.47 m / s, and the RMSE decreased by 1.93 m / s (error reduction of approximately 76%). 2 The value increased from 0.06 to 0.38. Furthermore, the improved quality of the WRF simulation data also led to improved wind speed results, although ML... WRF-UCM Compared to ML WRF The wind speed MAE and RMSE were only slightly reduced by 0.01 m / s and 0.02 m / s, respectively, but the correlation R of the data was significantly improved. 2 It has been significantly improved to 0.48.

[0094] As can be seen, the accuracy trend of the corrected simulated wind speed results for each wind level is consistent with the trend of the initial results, with the highest TS score for level 1 wind. Specifically, compared to the standard WRF simulation results, machine learning correction significantly improved the prediction accuracy for level 1 wind, achieving a TS score as high as 0.57, demonstrating excellent accuracy. The TS scores for level 0 and level 2 winds also showed significant improvement, approaching 0.2. The TS score for level 3 wind remained at a very low level after correction, indicating that the standard WRF model's simulation results for level 3 wind were poor, and the machine learning model's results were also poor. Furthermore, compared to ML... WRF ML WRF-UCMThis resulted in a significant improvement in TS scores for all wind levels, especially for level 2 winds (TS score increased by 0.09) and level 3 winds (TS score increased to 0.10).

Claims

1. A correction method based on urban canopy meteorological simulation data combining ML and modified WRF-BEM models, characterized in that, The specific steps are as follows: Step 1: Acquisition of observation data; Step 2: Acquiring simulation data; Step 3: Machine learning model setup; Step 4: Evaluation of initial simulation results; Step 5: Machine learning results.

2. The correction method for urban canopy meteorological simulation data based on a combination of ML and modified WRF-BEM models according to claim 1, characterized in that, Step 1 is implemented in the following steps: The observation data consisted of hourly data from several meteorological stations distributed throughout the selected area throughout the year, including temperature and humidity measured at 2 meters and wind speed measured at 10 meters. Preprocessing of observation data: The preprocessing of the observation data included outlier detection and removal, as well as missing value imputation. Wind speed and rainfall did not show significant diurnal variation patterns, so linear imputation was used. Temperature and relative humidity had simple diurnal variation patterns, that is, the highest temperature and the lowest relative humidity usually occurred around noon, and the lowest temperature and the highest relative humidity occurred before sunrise. Therefore, while performing linear imputation on temperature and relative humidity, the diurnal variation pattern was also considered. That is, when the missing value was an extreme point, imputation was performed based on the diurnal variation pattern of the nearby days under the same weather conditions at that station. Finally, a total of 122,640 hourly meteorological observation data from several meteorological stations throughout the year were obtained.

3. The correction method for urban canopy meteorological simulation data based on a combination of ML and modified WRF-BEM models according to claim 2, characterized in that, Step 2 is implemented in the following steps: Based on simulated meteorological data acquired using WRF-ARW 4.5, the center of the WRF nested computational domain was determined. The model employs a triple-nested grid. The parent domain D1 contains the center of the first nested domain D2, aiming to capture the background field characteristics of large-scale weather systems. The parent domain D1 has a horizontal spatial resolution of 9 km and a grid size of 120×120 in the horizontal direction. The first nested domain D2 contains the center of D3, aiming to focus on characterizing the interaction between regional-scale meteorological fields and topographic effects. D2 has a resolution of 3 km and a grid size of 118×118 in the region. The second nested domain D3 is the observation area, i.e., the study domain, with a spatial resolution of 1 km and a grid size of 118×118 in the region. It aims to finely simulate the distribution characteristics of wind, humidity, and temperature fields at the urban scale. Vertically, all three nested domains use 45-layer grids, extending from the ground to a height of 5000 Pa. The vertical grid is relatively dense near the surface and gradually becomes sparser with increasing height. In the WRF model, the physical parameterization process selected was as follows: Thompson scheme for cloud microphysics, RRTMG scheme for longwave / shortwave radiation, modified MM5 Monin-Obukhov scheme for near-surface layer processes, Noah-MP scheme for land surface processes, YSU scheme for boundary layer processes, and Kain-Fritsch scheme for cumulus convection. Simulation data were generated using the standard WRF model and the modified WRF-BEM model, respectively. The canopy parameters for each LCZ were set based on relevant architectural design documents and field survey results for the selected area.

4. The correction method for urban canopy meteorological simulation data based on a combination of ML and modified WRF-BEM models according to claim 3, characterized in that, Step 3 is implemented in the following steps: The optimal machine learning model was selected by constructing a comparative analysis framework between the ensemble model and the basic models. First, three basic models were trained: Random Forest (RF), Extreme Gradient Boosting (XGB), and Gradient Boosting Decision Tree (GBDT). Bayesian optimization was used to obtain the optimal predictive performance of each basic model under specific parameter configurations, i.e., predictive canopy meteorological data (temperature, humidity, and wind speed) that more closely approximate actual observations. Then, linear regression was performed based on the results of each basic model to obtain a final ensemble model. Machine learning was then performed separately for the three meteorological elements, and the mean absolute error (MAE), root mean square error (RMSE), and correlation coefficient (R²) were used. 2 As an evaluation metric for model performance, the optimal model is obtained through metric evaluation; We introduce the SHAP method from interpretable machine learning. Through SHAP values, we can intuitively see which features have a greater impact on the model's prediction results and how that impact trends are. All meteorological information was divided into a training set and a test set, with the training set accounting for 80% and the test set accounting for 20% of the total data, respectively. The input data for the machine learning model consists of nine features: WRF simulation temperature, humidity, wind speed, rainfall, station codes, and local LCZ feature variables, namely maximum building height, building proportion, grassland proportion, and tree proportion. Mapping relationships are established based on the temperature, humidity, and wind speed observed at the meteorological station to obtain the prediction results of the machine learning model.

5. The correction method for urban canopy meteorological simulation data based on a combination of ML and modified WRF-BEM models according to claim 4, characterized in that, The specific parameter configuration in step 3 refers to the optimal combination of hyperparameters determined for each base model using the Bayesian optimization algorithm. For the random forest model, the optimization parameters include the number of trees (n_estimators), maximum depth (max_depth), minimum number of split samples (min_samples_split), minimum number of leaf samples (min_samples_leaf), and random feature selection strategy (max_features). For the XGBoost model, the focus is on optimizing the learning rate (learning_rate), boosting epochs (n_estimators), maximum tree depth (max_depth), sample and feature sampling ratios (subsample, colsample_bytree), and regularization parameters (reg_alpha, reg_lambda). For the GBDT model, the adjustment parameters include boosting epochs (n_estimators), learning rate (learning_rate), maximum tree depth (max_depth), and sample sampling ratio (subsample). Bayesian optimization effectively searches the parameter space by constructing surrogate models, finding better parameter combinations with fewer evaluations compared to grid search and random search.

6. The correction method for urban canopy meteorological simulation data based on a combination of ML and modified WRF-BEM models according to claim 5, characterized in that, In step 3, the final ensemble model is constructed using the Stacking generalization method. The specific process consists of three steps: First, three basic models—Random Forest, XGBoost, and GBDT—are trained using their respective optimal parameter configurations. Then, the three trained basic models are used to predict the validation set, resulting in three sets of predictions as meta-features. Finally, using the predictions of these three basic models as input features and the true labels of the validation set as target variables, a linear regression model is trained as a meta-learner. The linear regression model automatically learns the optimal weight coefficients of the predictions of each basic model. The final ensemble prediction result is obtained by weighted linearly combining the predictions of the three basic models.

7. The correction method for urban canopy meteorological simulation data based on a combination of ML and modified WRF-BEM models according to claim 6, characterized in that, Step 4 is implemented in the following steps: Using statistical parameters such as mean absolute error (MA), root mean square error (RMS), and Pearson correlation R0 2 The simulation results for air temperature, specific humidity, and wind speed were evaluated. Compared to the standard WRF model, WRF+BEM significantly reduced the simulation errors of each meteorological element and improved their correlation. The simulated air temperature and specific humidity showed a strong correlation with the measured values, while the correlation for wind speed was weak. Additionally, the accuracy of the simulated wind speed in wind class classification was evaluated using the TS index. The TS index is based on binary classification, and the TS calculation method is as follows: In the calculation of wind speed TS, wind speeds falling within the wind level range are classified as positive, and those falling outside the range are classified as negative. According to the wind speed level classification, 0.1m / s to 0.2m / s is level 0 wind, 0.3m / s to 1.5m / s is level 1 wind, 1.6m / s to 3.3m / s is level 2 wind, and 3.4m / s to 5.4m / s is level 3 wind. If the TS value is higher than 0.6, it indicates that the model's prediction performance is relatively good; if the TS value is between 0.2 and 0.4, it indicates that the model's prediction performance is average; and if the TS value is lower than 0.2, it means that the accuracy is poor.

8. The correction method for urban canopy meteorological simulation data based on a combination of ML and modified WRF-BEM models according to claim 7, characterized in that, Step 5 is implemented in the following steps: 1) Temperature training and prediction errors Using ML respectively WRF and ML WRF-UCM This represents the machine learning results obtained using simulated data from the standard WRF model and the modified WRF-BEM model as training data; 2) Specific humidity training and prediction error The specific humidity error is reduced after processing with a machine learning statistical model; 3) Wind speed training and prediction error The wind speed error was reduced after processing by a machine learning statistical model.