A night heat wave prediction method based on precursor underlying surface factor and regression optimization

By combining image recognition and various machine learning models with a method based on precursory underlying surface factors and regression optimization, the shortcomings of existing heat wave prediction systems in predicting nighttime heat waves have been addressed, achieving efficient and stable nighttime heat wave prediction and climate variable prediction.

CN120850803BActive Publication Date: 2025-12-05NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511339760.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-12-05
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Existing heat wave forecasting systems are unable to effectively capture abnormally high nighttime minimum temperatures and lack accurate early warning methods. Furthermore, traditional factor selection relies on expert experience or black-box algorithms, resulting in high model complexity, high computational costs, and poor interpretability of results. The application of multiple models in nighttime heat wave forecasting is not effective.

Method used

A method based on precursory underlying surface factors and regression optimization was adopted. By acquiring global monthly factor data and human comfort index, nighttime heat wave events were screened, interannual increments were calculated, key areas were extracted using image recognition technology, and multiple machine learning models were combined for optimal combination to construct an efficient and stable nighttime heat wave prediction model.

Benefits of technology

It achieves efficient and robust prediction of nighttime heat waves, reduces reliance on expert knowledge, improves the objectivity of factor selection and the interpretability of the model, reduces computational complexity and the risk of overfitting, and is applicable to the prediction of nighttime heat wave frequency and other climate variables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850803B_ABST
    Figure CN120850803B_ABST
Patent Text Reader

Abstract

The application discloses a kind of night heat wave prediction methods based on precursor underlying surface factor and regression optimization, comprising: the regional average of the frequency of night heat wave of each year is calculated, and the target sequence of heat wave prediction is generated;The spatial correlation field and the significance level field of the interannual increment data of each factor obtained are calculated with the interannual increment of target sequence, the key area is extracted and the regional average of the interannual increment of each factor in the range is calculated, and the candidate factor set and model input independent variable subset are generated;With the one-dimensional candidate feature sequence in subset as independent variable, with the interannual increment of target sequence as dependent variable, respectively adopt multiple machine learning methods to construct prediction model and carry out fitting, obtain the prediction model corresponding to the optimal combination, obtain the heat wave prediction result of target year.The present application can make full use of precursor underlying surface factor information, automatically identify sensitive area highly correlated with target variable, and combine different model structures to systematically evaluate and screen the prediction effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of meteorological disaster prediction, and particularly relates to a night heat wave prediction method based on precursor underlying surface factors and regression optimization. BACKGROUND

[0002] In the past few decades, the frequency, intensity and duration of heat wave events have significantly increased, and the harm of night heat wave (abnormally high daily minimum temperature) is particularly prominent. It exacerbates the risk of cardiovascular and respiratory diseases by hindering the body's heat dissipation recovery at night, and also increases energy consumption and power grid load. However, the existing heat wave prediction system mainly focuses on daily maximum temperature, and it is difficult to effectively capture the climate risks brought by abnormally high minimum temperature at night. There is a lack of precise early warning means for night heat wave. In view of the above reasons, it is of great significance to develop a prediction method for night heat wave.

[0003] In the process of constructing the heat wave prediction method, the selection of precursor factors is crucial. Although the traditional static factors (such as sea ice coverage and snow depth) commonly used in previous studies have high signal stability, they often respond insufficiently to rapid abnormal processes of the climate system. In contrast, dynamic factors such as the monthly variability of sea ice and snow depth at mid-high latitudes can more sensitively capture the non-equilibrium state and rapid adjustment mechanism within the climate system, providing key new information sources for the prediction system. Therefore, it is necessary to include such dynamic factors in the prediction system to make up for the shortcomings of static factors and improve the prediction ability.

[0004] In terms of prediction area selection and factor identification methods, existing work mainly falls into two categories: one is the "manual annotation" method based on expert experience, which extracts potential prediction areas and factors by identifying relevant distribution maps or regression significant regions. However, this method relies on professional knowledge, is highly subjective and time-consuming. The second is the "automatic fitting" method based on deep learning, which automatically learns high-order features and factors by inputting a large amount of meteorological data. However, it has high computational cost, poor result interpretability, and complex model structure.

[0005] In addition, multi-model ensemble methods (such as random forest, gradient boosting, LSTM, etc.) have gradually attracted attention in the field of heat wave prediction, but their application in the specific scenario of night heat wave prediction is still immature and has a series of practical problems. First, the number of samples of night heat wave events is relatively small, and directly introducing multiple models will significantly increase the demand for data, which may cause some models to overfit or unstable training due to insufficient samples. Second, different models have different response mechanisms to the same input factors, which may lead to inconsistencies between the prediction results, even mutual interference, thereby masking the potential physical driving signals. In addition, multi-model methods usually involve a large number of parameter adjustments and model fusion processes, which significantly increase the computational overhead and system complexity, and also reduce the interpretability and transferability of the prediction results. Therefore, at the current stage, blindly introducing multi-models in the task of night heat wave prediction cannot fundamentally improve the prediction performance, but may weaken the overall effect due to improper feature selection or input redundancy.

[0006] The traditional factor selection relies on expert experience or black box algorithms, which further exacerbates the defects of the model. Therefore, how to intelligently select high-explainability factors and efficiently combine the selected factors and models to more comprehensively depict the occurrence mechanism of night heat wave and improve the accuracy and adaptability of heat wave prediction has become a technical problem to be solved.

[0007] Therefore, it is of great scientific value and application prospect to develop a new heat wave prediction method that takes into account the diversity of precursor factors, the intelligent extraction of factors, and the complexity of model structures. SUMMARY

[0008] The present application proposes a night heat wave prediction method based on precursor underlying surface factors and regression optimization to address the key problems in current heat wave prediction technology, such as the need for independent modeling of night heat wave, the dependence of precursor factor selection on expert experience or black box algorithms, and the insufficient description of nonlinear and coupled characteristics in the model construction process. The method can automatically identify sensitive areas highly correlated with the target variable based on the full use of precursor underlying surface factor information, and combine different model structures to systematically evaluate and select the prediction effect. The present application aims to improve the scientificity, stability, and interpretability of the prediction system by intelligently extracting key prediction factor regions and flexibly constructing and evaluating multiple model combinations, thereby enhancing the technical capabilities in dealing with climate extreme events.

[0009] To achieve the above technical purposes, the technical solutions adopted by the present application are as follows:

[0010] A night heat wave prediction method based on precursor underlying surface factors and regression optimization, the method comprising the following steps:

[0011] S1, obtaining global monthly precursor underlying surface factor data in a set period and hourly human comfort index data in a research area; based on the obtained hourly human comfort index data in the research area, screening all night heat wave events; calculating regional average of night heat wave frequency of each year to generate a target sequence for heat wave prediction;

[0012] S2, transforming data structure of the precursor underlying surface factor data, aggregating it monthly into an annual x month x latitude x longitude format; combining the transformed precursor underlying surface factor data and the regional average of night heat wave frequency of each year, calculating the difference between adjacent years to obtain annual increment data of each factor and annual increment of the target sequence;

[0013] S3, calculating the spatial correlation field and the significance level field of the annual increment data of each factor and the annual increment of the target sequence, extracting connected regions, selecting connected regions with an area greater than a preset area threshold as key regions, recording the latitude and longitude range of each key region, and calculating the regional average of the annual increment of each factor in the range to generate N one-dimensional candidate feature sequences with the same length as the candidate factor set;

[0014] S4, in the candidate factor set, according to the principle of minimal sufficiency, under the constraint of meeting the prediction performance, selecting a set with the smallest number of factors for combination to generate a subset of model input variables; for each subset, taking the one-dimensional candidate feature sequence in the subset as the independent variable, taking the annual increment of the target sequence as the dependent variable, respectively adopting multiple machine learning methods to construct and fit the prediction model, and setting the parameters of each model to optimize the model in the fitting process; using the explained variance R² and the Pearson correlation coefficient corr to score each subset and the combination of machine learning methods to obtain the prediction model corresponding to the optimal combination as the optimal prediction model;

[0015] S5, using the selected optimal prediction model, inputting the latest annual night observation value or predicted value of the precursor underlying surface factor corresponding to the model to predict the annual increment of night heat wave in the future year, and superimposing the predicted annual increment of night heat wave on the night heat wave frequency value of the previous year to obtain the night heat wave prediction result of the target year.

[0016] Further, in step S1, the precursor underlying surface factor data includes sea surface temperature, land surface temperature, sea ice coverage, snow depth, soil moisture, sea surface temperature monthly variability, land surface temperature monthly variability, sea ice coverage monthly variability, snow depth monthly variability, and soil moisture monthly variability.

[0017] Further, in step S1, based on the obtained hourly human comfort index data in the research area, all night heat wave events are screened by using a double threshold method, which specifically includes:

[0018] a relative air temperature threshold and an absolute air temperature threshold are set, wherein the relative air temperature threshold is the 90th percentile of the 15-day moving average of the corresponding calendar day of the reference period;

[0019] In the study period, any day when the minimum human comfort index exceeds both the relative air temperature threshold and the absolute air temperature threshold is determined to be an independent heat night. When independent heat nights appear continuously for not less than 3 days, a night heat wave event is screened out.

[0020] Further, in step S3, the process of extracting connected regions includes the following steps:

[0021] S31, taking the interannual increment data of each factor and the interannual increment of the target sequence as input, calculating the Pearson correlation coefficient at each spatial grid point to obtain the correlation field R(x, y) and the significance level field p(x, y) of the spatial grid coordinate (x, y), and constructing a significance mask M(x, y) based on the correlation threshold and the significance test, and generating a binary image I(x, y) by combining the two:

[0022] ;

[0023] wherein is the correlation threshold, is the significance level;

[0024] S32, performing connectivity analysis on the regions in the image I(x, y) with pixel value 1, constructing a path using the 4-neighborhood standard, and identifying all mutually connected pixel sets:

[0025] Divide the connected pixel sets into K regions For any pixel , there is a path such that adjacent pixels and are connected in the 4-neighborhood, and the pixel value of the pixel ;

[0026] Calculate the area of the connected region :

[0027] ;

[0028] S33, for all connected regions, only keep the regions that meet the area greater than the minimum area threshold T as key regions, and record their latitude and longitude ranges; the minimum area threshold T is set according to the spatial resolution of the study area and the climate significance;

[0029] S34. The annual factor increment data in the key area is averaged regionally to extract a one-dimensional time series for model training, which is then used as the input of candidate features.

[0030] Step S4 further includes:

[0031] Equipment selection factor set is In the candidate factor set, select subsets of all factors with a value of 2. and all subsets with a factor of 3 And merge them into a subset library of input independent variables for the model. Its scale is ,in, This represents the number of combinations of choosing two elements from set S. This represents the number of combinations of selecting 3 elements from set S;

[0032] For the m-th subset Using one-dimensional candidate feature sequences in a subset as independent variables and the interannual increment of the target sequence as the dependent variable, four machine learning methods—linear regression, random forest, gradient boosting, and long short-term memory neural network—were employed to construct prediction models.

[0033] ;

[0034] In the formula, These correspond to the linear regression model, random forest model, gradient boosting model, and long short-term memory neural network model, respectively. They are respectively The functions of these four models; the subscript i represents the number of the element in the set of independent variables of the model input, and the subscript t is the number of the time dimension;

[0035] During the fitting process, key parameters of the random forest model, gradient boosting model, and long short-term memory neural network model were set to optimize the model. The random forest model includes three key parameters: the number of decision trees, the random seed, and the maximum depth of the trees. The gradient boosting model includes three key parameters: the number of weak learners, the learning rate, and the random seed. The long short-term memory neural network model includes four key parameters: the learning rate, the number of tolerance rounds, the maximum number of training rounds, and the validation set split ratio.

[0036] Furthermore, in step S4, the explained variance R² and Pearson correlation coefficient corr are calculated for each subset and machine learning method combination. The combination results of subsets and machine learning methods are automatically sorted and filtered by the optimal subset regression based on the comprehensive score (0.5×R²+0.5×|corr|) to obtain the optimal prediction model.

[0037] Further, in step S4, the combined parameters of the precursor underlying surface factor in the subset include the factor name, month, and spatial range.

[0038] Compared with the prior art, the present application has the following advantages:

[0039] First, the night heat wave prediction method based on precursor underlying surface factors and regression optimization of the present application realizes efficient and stable prediction of one-dimensional interannual variables by introducing the technical route of interannual increment modeling, image recognition key area extraction, and optimal subset regression. Unlike the traditional method of directly using the original sequence, the present application first performs difference processing on the adjacent years of the factor and the target sequence, removes the interference of long-term trends and interdecadal drift, and enables the model to focus on the key signals that truly affect the interannual variation, thereby significantly improving the stability and accuracy of the prediction.

[0040] Second, the night heat wave prediction method based on precursor underlying surface factors and regression optimization of the present application uses image recognition technology to automatically identify key areas from the precursor underlying surface factor related field as candidate factors, which ensures the comprehensiveness and objectivity of factor selection, improves the selection efficiency and reduces the dependence on expert knowledge, guarantees the comprehensiveness and objectivity of factor selection, and speeds up the model construction speed. Its feasibility and effectiveness have been verified in climate seasonal prediction.

[0041] Third, the night heat wave prediction method based on precursor underlying surface factors and regression optimization of the present application combines factor screening and sensitive area recognition technology based on candidate factors to construct a high-dimensional feature library of tens of thousands of candidate subsets. Through spatial correlation and significance analysis, this method effectively compresses the feature space and improves the representativeness and consistency of the input factors. Subsequently, optimal subset regression is used to automatically select the optimal combination, eliminating the tedious steps of manual parameter tuning and empirical judgment, and realizing the full-process automation of the model from construction to deployment. Compared with the input redundancy and noise sensitivity problems that the conventional multi-model scheme may face, this method has completed high-quality factor selection in the preprocessing stage, providing a stable and unified input basis for the model, so that the subsequent machine learning model not only has nonlinear fitting capability but also effectively reduces the risk of overfitting, thereby ensuring high prediction performance in different years and different climate states. This method is not only suitable for the prediction of night heat wave frequency, but also can be extended to winter precipitation, drought days and other interannual climate variables, providing strong technical support for climate monitoring, disaster warning and resource scheduling. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 The flowchart of the night heat wave prediction method based on precursor underlying surface factors and regression optimization of the present application;

[0043] Figure 2 Cumulative night heat wave frequency distribution diagram (take a region in South China as an example) from 1981 to 2024;

[0044] Figure 3 Interannual variation diagram of regional average night heat wave frequency in South China in summer;

[0045] Figure 4 Optimal prediction model diagram, wherein (a) is model input data, (b) is model fitting result in training period, and (c) is model fitting result in test period. DETAILED DESCRIPTION

[0046] The embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0047] A night heat wave prediction method based on precursor underlying surface factors and regression optimization, the method comprising the following steps:

[0048] S1, obtaining monthly precursor underlying surface factor data in a global range and hourly human comfort index data in a research region within a set period; based on the obtained hourly human comfort index data in the research region, all night heat wave events are screened out; the regional average value of the night heat wave frequency of each year is calculated, and a target sequence for heat wave prediction is generated;

[0049] S2, performing data structure transformation on the precursor underlying surface factor data, and aggregating it into an annual x month x latitude x longitude format according to month; combining the precursor underlying surface factor data after format transformation and the regional average value of the night heat wave frequency of each year, calculating the difference between adjacent years to obtain the annual increment data of each factor and the annual increment of the target sequence;

[0050] S3, calculating the spatial correlation field and the significance level field of the annual increment data of each factor and the annual increment of the target sequence, extracting connected regions, selecting connected regions with an area greater than a preset area threshold as key regions, recording the latitude and longitude range of each key region, and calculating the regional average value of the annual increment of each factor in the range to generate N one-dimensional candidate feature sequences with consistent length as the candidate factor set;

[0051] S4, in the alternative factor set, according to the principle of minimum sufficiency, under the constraint of meeting the prediction performance, select the set of the minimum number of factors to combine to generate the model input independent variable subset; for each subset, a one-dimensional candidate feature sequence in the subset is taken as the independent variable, and the interannual increment of the target sequence is taken as the dependent variable, and a plurality of machine learning methods are used to construct and fit the prediction model, and in the fitting process, the model parameters are set to optimize the model; the explained variance R² and the Pearson correlation coefficient corr are used to score each subset and the combination of machine learning methods, and the optimal combination corresponding to the prediction model is obtained as the optimal prediction model;

[0052] S5, using the selected optimal prediction model, inputting the observed value or predicted value of the precursor underlying surface factor of the latest year corresponding to the model, predicting the interannual increment of the night heat wave in the future year, and superimposing the predicted interannual increment of the night heat wave on the night heat wave frequency value of the last year to obtain the prediction result of the night heat wave in the target year.

[0053] Referring to Figure 1 The specific technical solutions of the present application are as follows:

[0054] 1) Data acquisition

[0055] Global monthly precursor underlying surface factor data and hourly human comfort index data in the study area in a set period are obtained, and the precursor underlying surface factor data specifically includes: sea surface temperature, land surface temperature, sea ice coverage, snow depth, soil moisture, sea surface temperature inter-monthly variability, land surface temperature inter-monthly variability, sea ice coverage inter-monthly variability, snow depth inter-monthly variability, and soil moisture inter-monthly variability.

[0056] 2) Calculate the frequency of night heat wave

[0057] Based on the hourly human comfort index data in the study area obtained in step 1), a double threshold method is used to define a heat wave event. First, all dates in the study period are investigated, and when the minimum human comfort index of a day simultaneously exceeds the 90th percentile of the 15-day moving average of the corresponding calendar day in the reference period (1981-2010) (relative temperature threshold) and 26℃ (absolute temperature threshold), it is determined as an "independent hot night". When the independent hot night appears continuously for not less than 3 days, it is defined as a night heat wave event. Further, the regional average value of the frequency of night heat wave in each year is calculated as the target sequence for heat wave prediction.

[0058] 3) Data preprocessing based on monthly aggregation and interannual increment

[0059] 3.1) The data structure of the precursor underlying surface factor data is transformed from the original (month scale time axis x latitude x longitude) to the (year x month x latitude x longitude) format.

[0060] 3.2) Calculate the difference between the processed precursor underlying surface factor data and the regional mean sequence of the night heat wave frequency in step 2) to obtain the interannual increment data The formula is:

[0061] (1) ;

[0062] In the formula, and are the regional mean values of the night heat wave frequency in year and year-1, respectively.

[0063] 4) Construct a set of alternative factors

[0064] Calculate the spatial correlation field and the significance level field of the interannual increment data of each factor obtained in step 3.2) and the interannual increment of the target sequence, and automatically extract the significant connected regions by image recognition technology. In the identification process, the parameters (corr, p, T) of different factors are adjusted to optimize the selection of regions. Record the latitude and longitude range of each key region, and calculate the regional mean value of the interannual increment of the factor in the range to generate N one-dimensional candidate feature sequences of the same length as the set of alternative factors.

[0065] The principle of image recognition technology is:

[0066] Let the binary image be I: Ω→{0,1}, where Ω is the image domain. For the image , the connectedness of the region with pixel value 1 is analyzed, and the path is constructed using the 4-neighborhood standard to identify all mutually connected pixel sets. Specifically, the connected pixel sets are divided into K regions , first, the connectedness condition is screened: for any pixel , there exists a path , such that the adjacent pixels and are connected in the 4-neighborhood, and the pixel value of the pixel ; next, the area of each connected region is calculated, where the area of the connected region is: . For all connected regions, only the regions that meet the area greater than the minimum area threshold T are kept as key regions, and their latitude and longitude ranges are recorded; the minimum area threshold T is set according to the spatial resolution of the study area and the climate significance.

[0067] 5) Optimize and select the optimal model through regression

[0068] In high-dimensional small sample data (for example, one-dimensional data with a length of more than forty), too many independent variables will increase the variance of the model and lead to overfitting, although it will improve the performance. It has been tested that with the increase of the number of independent variables, the performance improvement value will also be greatly reduced. Therefore, the present application proposes that, in the alternative factor set, according to the principle of minimal sufficiency, a set of the minimum number of factors is selected for combination to generate a subset of model input independent variables under the premise of meeting the prediction performance.

[0069] In practical applications, the relationship curves of prediction performance and factor number and the relationship curves of factor number and model variance can be drawn respectively, and the minimum factor number can be set according to actual needs. For example, in the seasonal scale climate prediction of the present embodiment, the R² value of the model with 2-3 independent variables (i.e. 2-3 precursor factor combinations) has reached a high level, and in this example, the R² value reaches 0.85, which meets the conventional prediction performance requirements, indicating that combining 2-3 key factors can have good explanatory ability for the interannual variation of night heat waves while effectively avoiding the risk of overfitting.

[0070] Selecting all subsets of factor number 2 and all subsets of factor number 3 from the alternative factor set obtained in step 4) respectively to generate a subset of model input independent variables, the method being as follows: let the original independent variable set be and the generated input independent variable subset be with a size of wherein represents the number of combinations of 2 elements selected from the set S, represents the number of combinations of 3 elements selected from the set S.

[0071] For each subset, taking the one-dimensional candidate feature sequence in the subset as the independent variable and the interannual increment of the target sequence obtained in step 3.2) as the dependent variable, four machine learning methods, linear regression ( ), random forest ( ), gradient boosting ( ) and long short-term memory neural network ( ) are used to construct prediction models. The mathematical framework of this step is as follows:

[0072] Let the feature subset be and for each subset construct a prediction model:

[0073] .

[0074] During the fitting process, the key parameters of the random forest model, the gradient boosting model, and the long short-term memory neural network model are set to optimize the model. Specifically, the random forest model includes three key parameters: n_estimators (number of decision trees), random_state (random seed), and max_depth (maximum depth of the tree); the gradient boosting model includes three key parameters: n_estimators (number of weak learners), learning_rate (learning rate), and random_state (random seed); the long short-term memory neural network model includes four key parameters: learning_rate (learning rate), patience (tolerance number of rounds), max_epochs (maximum number of training rounds), and validation split ratio (validation set division ratio).

[0075] The following is an explanation of the meanings and effects of the key parameters of each model.

[0076] For the random forest model, n_estimators (number of decision trees) indicates the number of decision trees that make up the random forest, and the more trees, the more stable the model usually is; max_depth (maximum depth of the tree) is used to control the complexity of each tree to prevent overfitting; random_state (random seed) is used to set the initial state of random number generation to ensure that the model results are repeatable.

[0077] For the gradient boosting model, n_estimators (number of weak learners) indicates the total number of regression trees iteratively constructed; learning_rate (learning rate) is used to control the contribution of each iteration to the final model, and the smaller the learning rate, the more robust the model; random_state (random seed) is used to control the random process in model training to ensure the reproducibility of the results.

[0078] For long short-term memory neural networks, learning_rate (learning rate) is used to control the size of the step of the model at each weight update, which is an important factor affecting the training efficiency and effect; early stopping (early stopping mechanism) is used to stop training when the loss of the validation set does not decrease continuously for several rounds, in order to prevent overfitting, including two parameters: patience (tolerance number of rounds) and max_epochs (maximum number of training rounds), wherein patience (tolerance number of rounds) is the patience value set when using the early stopping mechanism, indicating how many rounds of improvement can be tolerated before continuing training; max_epochs (maximum number of training rounds) is used to limit the maximum number of training rounds of the model; validation split ratio (validation set division ratio) divides a part of the original training set into a validation set, which is used to evaluate the performance of the model during training.

[0079] Calculate the explained variance R for each combination of subset and machine learning method 2 and the Pearson correlation coefficient corr, and automatically rank and screen the combination results of the subset and the machine learning method according to the comprehensive score (0.5xR²+0.5x|corr|) by the optimal subset regression. Finally, the optimal prediction model is obtained according to the score.

[0080] 6) Use the model to make predictions

[0081] Using the optimal prediction model selected in step 5), input the observed or predicted values of the key factors in the latest year to predict the interannual increment of nighttime heat waves in the next year. The process includes: extracting the numerical values of the key precursor factors corresponding to the optimal prediction model (including factor name, month, and spatial range) in the latest year, which are used as the input independent variables of the model, using the trained optimal machine learning model to infer, and outputting the predicted value of the interannual increment of nighttime heat wave frequency in the target year. Finally, the predicted interannual increment of nighttime heat waves is superimposed on the nighttime heat wave frequency value of the previous year to obtain the prediction result of the nighttime heat wave in the target year, realizing the deployment and application of the model in the actual climate service scenario.

[0082] Examples

[0083] The following is an example of predicting the regional average of summer night heat wave frequency in South China defined based on the global thermal climate index UTCI. The specific implementation steps and parameter settings of the present application are further clarified in combination with the drawings to ensure that a person skilled in the art can reproduce them completely. 1) Select the South China region (106°E-120°E, 20°N-27°N) as the research region, download the hourly human comfort index data of the fifth generation atmospheric reanalysis dataset of the European Centre for Medium-Range Weather Forecasts (ECMWF) for the summer (June-August) of 1981-2024, and the monthly global range of underlying surface factor data for 1980-2024. The underlying surface factors include sea surface temperature, land surface temperature, sea ice coverage, snow depth and soil moisture. Considering that the underlying surface state in the previous winter may have a lagging effect on the summer heat wave of the current year, the underlying surface factor data is taken from 1980. Then, the monthly variability of the downloaded underlying surface factor data is calculated using the formula: delta_A(time) = A(time)-A(time-1) (time represents the monthly time step), which includes the monthly variability of sea surface temperature, land surface temperature, sea ice coverage, snow depth and soil moisture.

[0084] 2) Extract the human comfort index data of the summer (June-August) of each year in South China. Use the double-threshold method described in step 2) of the technical solution: calculate the 90th percentile value of the fifteen-day moving average of the same period calendar day from 1981 to 2010 at each grid point as the relative threshold of air temperature, and set 26℃ as the absolute threshold of air temperature. Traverse all dates of the summer of 1981-2024, if the minimum human comfort index of a day exceeds both the relative threshold of air temperature and the absolute threshold of air temperature, then the day is determined as an "independent hot night"; when "independent hot night" appears continuously for at least three days, it is defined as a "night heat wave" event. Through the above determination, the frequency of night heat wave occurrence in each year of summer in the research region can be counted. Take a certain region in South China as an example, its spatial distribution is shown in Figure 2 . Then average in the spatial dimension, finally get the regional average sequence of summer night heat wave frequency in South China from 1981 to 2024, as shown in Figure 3 .

[0085] 3) Aggregate the underlying surface factor data into (year x month x latitude x longitude) format by month, i.e. form annual blocks in the time dimension on the original data structure. On this basis, the regional average value sequence of summer night heat wave frequency in South China and the monthly aggregated data of each underlying surface factor are calculated according to formula (1) in technical scheme step 3) to obtain the annual increment data of each factor and the annual increment sequence of night heat wave frequency. This difference operation explicitly removes the influence of long-term trend and interdecadal drift, so that the subsequent model can focus on the key signals that truly affect the interannual variation.

[0086] 4) Calculate the Pearson correlation coefficient and its t-test significance level of each factor interannual increment field and the night heat wave frequency interannual increment sequence in the spatial dimension to obtain the two-dimensional correlation field and the significance level field. The pixels that meet the preset conditions (for sea surface temperature and its increment, set corr>0.3, p<0.01; for other factors, set corr>0.2, p<0.05) are regarded as "correlation points", and they are binarized: correlation points are white, and other points are black. Then use the four-neighbor connected rule to label the connected domain of white pixels, and count the number of pixels in each connected region; the connected domain with a pixel number reaching the minimum area threshold T (for sea surface temperature and its increment, set T>100; for other factors, set T>20) is regarded as a "key area". Finally, record the latitude and longitude range of each key area, and spatially average the factor interannual increment field in these areas to obtain a one-dimensional candidate feature sequence consistent with the length of the target sequence. In this example, a total of 74 one-dimensional candidate feature sequences are obtained. Summarize them and organize them in list form (column name format: 'factor name_month_latitude and longitude range'). A complete candidate factor library is constructed.

[0087] 5) Next, according to the technical scheme step 5), traverse all combinations of factors with a number of 2 and 3 from the candidate factor library to construct an input independent variable subset library. In this example, the elements of the input independent variable subset library have a total of C(74, 2)+C(74, 3)=67525. For each feature subset, use the one-dimensional candidate feature sequence it contains as the independent variable, and the interannual increment sequence of night heat wave frequency as the dependent variable, respectively using four different machine learning models: linear regression ( ), random forest ( ), gradient boosting ( ) and long short-term memory neural network ( The model was trained using the following methods: Random Forest (n_estimators=100, random_state=42, max_depth=5), Gradient Boosting (n_estimators=100, learning_rate=0.05, random_state=42), and LSTM (learning_rate=0.01, early stopping (patience=10, max_epochs=300), validation split ratio=20%). A total of 67,525 (number of feature sequence subsets) * 4 (number of machine learning models) = 270,100 models were generated. After each model was trained, the explained variance R² and Pearson correlation coefficient corr were calculated using test set data. All methods were then ranked according to a comprehensive score (0.5 × R² + 0.5 × |corr|), and the highest-scoring optimal prediction model was selected.

[0088] 6) In this example, the optimal prediction model with the highest overall score is determined as follows: Figure 4 As shown, the one-dimensional factor sequence includes: the regional average soil moisture content in May each year from 1981 to 2024 within the range of (94°E–110°E, 8°N–25°N); the regional average sea surface temperature in November each year from 1980 to 2023 within the range of (160°W–80°W, 5°S–7.5°N); and the regional average snow depth in May each year from 1981 to 2024 within the range of (67.5°E–110°E, 62.5°N–75°N). The machine learning model used is an LSTM model. Its overall score is 0.90, and the correlation coefficient between the model-fitted sequence and the observed sequence is 0.94. During the testing period, the explained variance R² reached 0.85, demonstrating the model's good predictive performance. Furthermore, the three factors automatically selected by this method are highly correlated with key mechanisms such as ENSO, upstream water vapor transport, and mid-to-high latitude wave trains, further demonstrating the reliability of the method.

[0089] 7) The actual prediction stage of the heatwave frequency. First, the interannual increment values of the factors corresponding to the optimal prediction model are extracted as the input variables of the model. Specifically, they include: the increment value of the soil moisture in the Indo-China Peninsula (94°E-110°E, 8°N-25°N) in May relative to the previous year; the increment value of the sea surface temperature in the equatorial region of the eastern Pacific (160°W-80°W, 5°S-7.5°N) in November relative to the previous two years; and the increment value of the snow depth in the Eurasian high-latitude region (67.5°E-110°E, 62.5°N-75°N) in May relative to the previous year. The above three independent variables are input into the optimal prediction model that has been trained, and the interannual increment value prediction result of the current annual nighttime heatwave frequency is output. Subsequently, the increment value is added to the measured value of the nighttime heatwave frequency in the previous year, and the regional average prediction result of the current annual nighttime heatwave frequency is obtained.

[0090] While the preferred embodiments of the application have been described, it should be apparent that various modifications and changes can be made to the embodiments without departing from the spirit and scope of the application. Accordingly, it is intended that all such modifications and changes be included within the scope of the application as claimed.

[0091] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. A method for predicting night heat wave based on precursor underlying surface factors and regression optimization, characterized in that, The method comprises the following steps: S1, obtaining global monthly precursor underlying surface factor data in a set period and hourly human comfort index data in a research area; based on the obtained hourly human comfort index data in the research area, all night heat wave events are screened out; the regional average of the night heat wave frequency of each year is calculated to generate a target sequence for heat wave prediction; S2, the data structure of the precursor underlying surface factor data is transformed, which is aggregated monthly into the format of year x month x latitude x longitude; combined with the precursor underlying surface factor data after format transformation and the regional average of the night heat wave frequency of each year, the difference between adjacent years is calculated to obtain the annual increment data of each factor and the annual increment of the target sequence; S3, the spatial correlation field and the significance level field of the annual increment data of each factor and the annual increment of the target sequence are calculated, the connected regions are extracted, the connected regions with an area greater than a preset area threshold are selected as key regions, the latitude and longitude range of each key region is recorded, and the regional average of the annual increment of each factor in the range is calculated to generate N one-dimensional candidate feature sequences with the same length as the candidate factor set; S4, in the candidate factor set, according to the principle of minimal sufficiency, a set of minimal factors is selected for combination under the constraint of meeting the prediction performance to generate a subset of model input variables; for each subset, the one-dimensional candidate feature sequence in the subset is taken as the independent variable, the annual increment of the target sequence is taken as the dependent variable, and a plurality of machine learning methods are used to construct and fit prediction models; in the fitting process, the model parameters are set to optimize the model; the explained variance R² and the Pearson correlation coefficient corr are used to score each subset and the combination of machine learning methods to obtain the prediction model corresponding to the optimal combination as the optimal prediction model; S5, using the selected optimal prediction model, inputting the observed value or predicted value of the precursor underlying surface factor of the latest year corresponding to the model, predicting the annual increment of the night heat wave of the next year, and superimposing the predicted annual increment of the night heat wave on the night heat wave frequency value of the last year to obtain the night heat wave prediction result of the target year.

2. The method of predicting night heat wave based on precursory underlying surface factors and regression optimization according to claim 1, characterized in that, In step S1, the precursor underlying surface factor data includes sea surface temperature, land surface temperature, sea ice coverage, snow depth, soil moisture, sea surface temperature monthly variability, land surface temperature monthly variability, sea ice coverage monthly variability, snow depth monthly variability, and soil moisture monthly variability.

3. The method for predicting the night heat wave based on the precursor underlying surface factor and regression optimization according to claim 1, characterized in that, In step S1, based on the obtained hourly human comfort index data in the research area, all night heat wave events are screened out by using a double threshold method, which specifically includes: Setting a relative temperature threshold and an absolute temperature threshold, wherein the relative temperature threshold is the 90th percentile of the 15-day moving average of the corresponding calendar day; Examine all dates in the research period, when the minimum human comfort index of any day exceeds the relative temperature threshold and the absolute temperature threshold at the same time, it is determined as an independent hot night; when the independent hot night appears continuously for not less than 3 days, a night heat wave event is screened out.

4. The method for predicting the night heat wave based on the precursor underlying surface factor and regression optimization according to claim 1, characterized in that, In step S3, the process of extracting the connected region comprises the following steps: S31, taking the interannual increment data of each factor and the interannual increment of the target sequence as input, calculating the Pearson correlation coefficient at each spatial grid point to obtain the correlation field R(x, y) and the significance level field p(x, y) of the spatial grid coordinate (x, y), and constructing a significance mask M(x, y) based on the correlation threshold and the significance test, and generating a binary image I(x, y) by combining the two: ; wherein is a threshold of relevance, is a level of significance; S32, performing connectivity analysis on the region with pixel value 1 in the image I(x, y), constructing a path using the 4-neighborhood standard, and identifying all mutually connected pixel sets: dividing the connected pixel set into K regions , for any pixel , there exists a path such that adjacent pixels are connected under 4-neighborhood, and the pixel value of pixel is ; Computing the area of a connected region :​ ; S33, for all connected regions, only keep the region that meets the area greater than the minimum area threshold T as the key area, and record its latitude and longitude range; the minimum area threshold T is set according to the spatial resolution and climate significance of the study area; S34, performing regional averaging on the annual factor increment data in the key area to extract a one-dimensional time series for model training as a candidate feature input.

5. The method for predicting the night heat wave based on the precursor underlying surface factor and regression optimization according to claim 1, characterized in that, Step S4 further comprises: The device selection factor set is In the alternative factor set, all subsets of factor number 2 are selected respectively And all subsets of factor number 3 And after merging, it is used as a model input independent variable subset library The size is Wherein, Indicates the number of combinations of selecting 2 elements from the set S, Indicates the number of combinations of selecting 3 elements from the set S; For the mth subset With the one-dimensional candidate feature sequence in the subset as the independent variable and the interannual increment of the target sequence as the dependent variable, four machine learning methods, including linear regression model, random forest model, gradient boosting model, and long short-term memory neural network model, were used to construct prediction models. ; In the formula, corresponding to the linear regression model, the random forest model, the gradient boosting model, and the long short-term memory neural network model, respectively; are respectively the functions of the four models; the subscript i represents the number of elements in the independent variable set of the model input, and the subscript t is the number of time latitude. During the fitting process, the key parameters of the random forest model, the gradient boosting model and the long short-term memory neural network model are set to optimize the model; wherein the random forest model includes three key parameters: the number of decision trees, the random seed and the maximum depth of the tree; the gradient boosting model includes three key parameters: the number of weak learners, the learning rate and the random seed; the long short-term memory neural network model includes four key parameters: the learning rate, the tolerance number of rounds, the maximum number of training rounds and the proportion of the validation set.

6. The method for predicting the night heat wave based on the precursor underlying surface factor and regression optimization according to claim 1, characterized in that, In step S4, the explained variance R² and the Pearson correlation coefficient corr are calculated for each subset and the combination of machine learning methods, and the combination results of the subset and the machine learning method are automatically sorted and screened according to the comprehensive score (0.5×R²+0.5×|corr|) by the optimal subset regression to obtain the optimal prediction model.

7. The method for predicting the night heat wave based on the precursor underlying surface factor and regression optimization according to claim 1, characterized in that, In step S4, the combination parameters of the precursor underlying surface factors in the subset include the factor name, the month and the spatial range.

Citation Information

Patent Citations

  • Chinese short-term climate prediction method and system based on artificial intelligence

    CN116128099A

  • Sub-season heat wave prediction method and storage medium

    CN119670524A