Rime forming and maintaining time collaborative forecasting method based on random forest-Transform fusion

Through the integration of random forests and Transformer models, a coordinated forecast method for rime formation and maintenance time was constructed, which solved the problems of low time resolution and uncaptured coupling effect of meteorological elements in the existing rime forecast method, and achieved accurate forecasting of the entire life cycle of rime, improving the accuracy and practicality of prediction.

CN120428358AActive Publication Date: 2025-08-05ANHUI PROVINCIAL PUBLIC METEOROLOGICAL SERVICE CENT

Patent Information

Application Number
CN202510548163.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-05
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing rime forecasting methods have low time resolution, and cannot accurately capture the rime formation period and meteorological changes, and fail to effectively consider the complex coupling effect between meteorological elements, so they cannot accurately predict the formation and maintenance time of rime.

Method used

Random forest and Transformer model are fused, and a collaborative forecast method for rime formation and maintenance time is constructed through hourly meteorological data, and the SMOTE method is used for upsampling. The rime formation forecast model is optimized by combining cross-verification and grid search, and a rime maintenance forecast model is constructed through dynamic window sampling strategy and self-attention mechanism to output the hourly presence and duration of rime.

Benefits of technology

It realizes accurate forecasting of the entire life cycle of rime, improves the prediction accuracy of rime formation and maintenance time, solves the problems of low time resolution and uncaptured coupling effect of meteorological elements in traditional methods, and significantly improves the practicality of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428358A_ABST
    Figure CN120428358A_ABST
Patent Text Reader

Abstract

The invention discloses a rime formation and maintenance time collaborative forecasting method based on random forest-Transform fusion, and relates to the technical field of weather forecast, the method comprises the following steps: collecting meteorological elements of a target area from multi-source data, and sorting the data into an hour-by-hour data set; a rime formation forecasting factor is extracted from the hour-by-hour data set, a random forest method is adopted to construct a rime formation forecasting model, samples in the hour-by-hour data set are used for training the rime formation forecasting model, and cross validation and grid search are used for evaluating and optimizing the model; rime maintenance forecasting factors are extracted from the hour-by-hour data set, a rime maintenance forecasting model is constructed by adopting a Transform method, and the rime maintenance forecasting model is trained by using samples in the hour-by-hour data set; and the meteorological data of the target area in the forecasting time period is processed by combining the rime forming forecasting model and the rime maintaining forecasting model, and finally forecasting whether rime exists per hour or not and the duration of the whole rime process are output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of weather forecasting, and in particular to a method for collaboratively forecasting the formation and maintenance time of rime based on random forest-Transformer fusion. Background Art

[0002] Rime refers to the condensation of water vapor into frost at low temperatures, typically found in mountainous areas and river valleys in cold regions. Rime landscapes are not only highly ornamental but also have a significant impact on transportation, safety, and tourism.

[0003] Currently, the analysis and forecast of meteorological conditions for rime mainly relies on daily statistical data from meteorological observatories, such as daily minimum temperature and daily average humidity. However, this method has low temporal resolution and cannot capture the precise period of rime formation and meteorological changes, making it difficult to accurately determine the specific time of rime formation.

[0004] Furthermore, many current rime prediction methods use static threshold methods, which fail to account for the complex coupling effects of meteorological factors. For example, the interaction between temperature, humidity, and wind speed has a significant impact on rime formation, but static threshold methods fail to effectively capture the synergistic effects of these meteorological factors. Furthermore, static threshold methods often fail to capture the sudden changes in meteorological factors in real time, resulting in inaccurate predictions.

[0005] Most existing rime forecasting methods focus solely on whether rime will form, neglecting to predict how long it will persist once it forms. Traditional methods can only determine whether rime will appear, but are unable to track and predict its duration. As rime forms and melts, meteorological conditions change continuously and complexly, requiring a full-lifecycle forecasting system, rather than a single "rime present" or "no rime" judgment.

[0006] Therefore, accurately predicting the formation and duration of rime is crucial to related industries. Summary of the Invention

[0007] The purpose of the present invention is to achieve accurate forecast of rime formation and duration by integrating random forest and Transformer models and utilizing hourly meteorological data, thereby improving the accuracy and practicality of rime life cycle prediction.

[0008] The technical solution of the present invention is to provide a collaborative forecasting method for rime formation and maintenance time based on random forest-Transformer fusion, which includes:

[0009] S1. Collect meteorological elements of the target area from the data of ground meteorological elements, meteorological elements of various layers in space, and records of the presence or absence of rime in the target area, and organize these data into an hourly dataset containing ground meteorological, space meteorological, and rime occurrence records;

[0010] S2. Extract rime formation prediction factors from the hourly dataset, upsample the data using the SMOTE method to expand the sample size, construct a rime formation prediction model using the random forest method, train the rime formation prediction model using samples from the hourly dataset, and evaluate and optimize the rime formation prediction model using a combination of cross-validation and grid search.

[0011] S3. Extract rime maintenance prediction factors from the hourly dataset, construct a rime maintenance prediction model using the Transformer method, increase the number of samples with rime in the training set through a dynamic window sampling strategy, and train the rime maintenance prediction model using samples from the hourly dataset;

[0012] S4. Processing meteorological data of the target area within the forecast period in combination with the rime formation forecast model and the rime maintenance forecast model, ultimately outputting hourly rime presence or absence forecasts and the duration of the entire rime process;

[0013] Specifically, step S4 includes:

[0014] S41. Obtain hourly ground meteorological forecast data for the target area based on the model forecast product or the meteorological element deterministic forecast product, and use the data to construct a feature vector for the forecast period;

[0015] S42, input the characteristic vector into the rime formation prediction model to obtain the rime formation probability P form , calculate the forecast probability of rime formation T dynamic ;

[0016] S43, continuously monitor the rime and determine the dynamic threshold value T dynamic , when the dynamic threshold value T of rime is determined for 6 consecutive hours dynamic When the value is >0.65, the characteristic vector is input into the rime maintenance forecast model to obtain the rime maintenance forecast probability P for the next 6 hours. maintain If the rime maintains the forecast probability P in the next 6 hours maintain >0.65, the feature vector is input into the rime maintenance forecast model again to obtain the rime maintenance forecast probability in a longer period of time until its value is less than 0.65;

[0017] In the above process, when there is a forecast probability of rime formation at a certain time t, T dynamic and rime maintenance forecast probability P maintain When there are two forecast results, the final probability of rime at that time is Pfinal(t) T dynamic 、T maintain The maximum value of the two is the final probability of rime at that time, P final(t) When the value is >0.65, it is determined that there is rime at that time, otherwise there is no rime, thus obtaining the hourly rime forecast for the forecast period; intervals no more than 2 hours are regarded as the same rime process, and the total duration from rime formation to melting is calculated, and the rime duration forecast is output.

[0018] In any of the above technical solutions, further, the rime formation forecast probability T dynamic The calculation formula is:

[0019] T dynamic =P form +0.05×RH min -0.02×(T current +2);

[0020] Among them, RH min T is the lowest humidity value in the past 6 hours at the forecast time. current The temperature at the forecast time.

[0021] In any of the above technical solutions, further, step S1 specifically includes:

[0022] S11, selecting a target area and obtaining hourly ground meteorological observation data from a meteorological observation station in the target area, wherein the hourly ground meteorological observation data includes records of air temperature, dew point temperature, relative humidity, wind speed, precipitation, and the presence or absence of rime, where the presence of rime is recorded as 1 and the absence of rime is recorded as 0;

[0023] S12. Clean the acquired hourly ground meteorological observation data. Data cleaning includes removing values that exceed the reasonable range of meteorological elements and filling in missing data at meteorological observation stations. Regression interpolation is used to fill in missing data.

[0024] S13. Based on the reanalysis dataset of the same meteorological type as the target area, obtain the meteorological element data of each spatial layer through interpolation method. The data includes the temperature, relative humidity, wind speed of the upper layer, the current layer, and the lower layer, and obtain the low cloud cover and total cloud cover data;

[0025] S14, performing format unification processing on the ground hourly meteorological observation data cleaned in step S12 and the spatial meteorological element data of each layer obtained in step S13;

[0026] S15. Arrange the data in the unified format in chronological order to ensure that the ground meteorological element and space meteorological element data at each moment match the record of the presence or absence of rime. The arranged data form a complete hourly data set for a time period.

[0027] S16. Retrieve the time periods that meet the conditional indicators in the hourly data set and mark them as "formation events". The conditional indicators include: rime, relative humidity > 98%, and temperature < 0°C.

[0028] In any of the above technical solutions, further, step S2 specifically includes:

[0029] S21, ground temperature T and dew point temperature T d The data is processed to calculate the temperature dew point difference and temperature stratification. The calculation processes of the two are as follows:

[0030] Calculate the temperature dew point difference TT d :TT d =T 2m -T d2m ; Among them, T 2m is the temperature at 2m above the ground, T d2m It is the dew point temperature 2m above the ground.

[0031] Calculate the temperature stratification T s :T s =T 下层 -T 当前层 ;

[0032] S22. Add the temperature dew point difference and temperature stratification to the hourly data set to extract the prediction factors of rime formation;

[0033] S23. Use the SMOTE method to upsample the data, specifically including:

[0034] For each minority class sample, select the other minority class samples of its nearest neighbors and generate new sample points between them. The generation formula is: X new =X minority +λ×(X neighbor -X minority ); where λ is a random number in the range [0,1], X minority Represents the feature vector of a sample in the minority class, X neighbor Indicates that there are more samples in the minority class than in X minority The feature vector of another sample of the nearest neighbor, X new According to X minority and X neighbor The new sample is generated, and its feature vector is obtained by linear interpolation between these two samples;

[0035] S24. Divide the samples into a training set and a test set in a ratio of 7:3. Use the random forest algorithm to construct a rime formation prediction model on the training set. The output of the model is the rime formation prediction probability. The random forest model consists of multiple decision trees, each of which is independently constructed on a randomly selected sample subset and feature subset. Use the K-fold cross-validation method to train and validate the training set multiple times. The training set is evenly divided into K subsets. Each time, one of the subsets is selected as the validation set, and the remaining K-1 subsets are used as the training set. Repeat this process K times to obtain the average performance of the model in different data partitions under each parameter configuration.

[0036] During the cross-validation process, a comprehensive search was conducted on the key hyperparameters in the rime formation prediction model, including the number of trees, maximum depth, and minimum number of split samples. Specifically, the candidate value ranges of each hyperparameter were pre-set to form a parameter grid. Then, K-fold cross-validation was performed on each hyperparameter combination in the grid, and its performance indicators on the validation set were recorded. The performance indicators include accuracy and TS score. The TS score calculation formula is: TS = number of hits / (number of hits + number of null reports + number of missed reports), where the number of hits is the number of events that were predicted to occur and actually occurred, the number of null reports is the number of events that were predicted to occur but did not actually occur, and the number of missed reports is the number of events that were predicted not to occur but actually occurred.

[0037] After traversing all candidate parameter combinations, a set of hyperparameter configurations with the best performance indicators is selected, and the rime formation prediction model is retrained on the entire training set using this optimal configuration.

[0038] In any of the above technical solutions, further, the prediction factors for rime formation include: month, day, ground temperature, temperature dew point difference, relative humidity, wind speed, precipitation, total cloud cover, low cloud cover, temperature stratification and temperature, relative humidity and wind speed of each layer, a total of 19 factors.

[0039] In any of the above technical solutions, further, step S3 specifically includes:

[0040] S31. Extracting rime maintenance prediction factors from the hourly data set. The rime maintenance prediction basic factors include: time, temperature, relative humidity, wind speed, precipitation, and rime record, a total of 6 factors;

[0041] S32. Constructing derived factors based on the basic factors for rime maintenance forecasting, calculating the moving averages and variances of the temperature and humidity in the 24 hours preceding the current time to capture the statistical characteristics of short-term meteorological data; the rime maintenance forecasting factors include all basic factors and derived factors for rime maintenance forecasting, namely, time, temperature, relative humidity, wind speed, precipitation, mean temperature, temperature variance, mean relative humidity, relative humidity variance, and rime record, a total of 10 factors;

[0042] S33. Sine / cosine transformation is performed on the time variable to retain the periodic characteristics and enhance the model's ability to learn the law of the circadian cycle. The transformed time Enc(t) is expressed as:

[0043]

[0044] Numerical features other than rime records were normalized using the Z-Score method to ensure consistency in the numerical scale of different features. The processing formula is:

[0045]

[0046] Where Z is the result of Z-Score standardization, x represents the processed feature, μ represents the mean of the training set, and σ represents the standard deviation of the training set;

[0047] S34. A dynamic window sampling strategy is used to solve the class imbalance problem. The positive sample window is defined as the period from 24 hours before the rime event to the end of the event. Dense sliding sampling is performed in this area with a step size of 1 hour. Sparse sampling is used in the negative sample area with a step size of 6 hours to effectively supplement the number of rime samples and improve the imbalance of data class distribution.

[0048] S35. The dataset is divided into training set, validation set, and test set in chronological order to ensure temporal integrity. The ratio of training set, validation set, and test set is 7:2:1. The Transformer method is used to construct a rime maintenance forecast model and it is trained using the training set. The output of the model is the rime maintenance forecast probability.

[0049] In any of the above technical solutions, further, the working method of the rime maintenance forecast model includes:

[0050] First, the feature vectors of 24 consecutive hours are stacked in chronological order to form a two-dimensional feature matrix with 24 rows and 10 columns as the model input, where each row corresponds to one hour of meteorological data and each column represents a different forecast factor;

[0051] A learnable positional encoding matrix is introduced into the input layer and added to the linearly transformed input features, enabling the model to explicitly perceive the temporal order of meteorological data. Multiple parallel attention heads are set up, each of which independently calculates the attention weights of the query, key, and value vectors, and calculates the association strength of each time step through scaled dot product calculations. The outputs of multiple heads are concatenated and linearly transformed to fuse the feature representations of different subspaces. A two-layer fully connected network is connected after the self-attention layer, using ReLU as the activation function. Residual connections are implemented after each self-attention layer and feedforward network layer, and the sub-layer output is added to the original input, followed by layer normalization.

[0052] The 24-hour meteorological feature sequence of the training set is used to capture the long-range dependencies between spatiotemporal features through a neural network model with a self-attention mechanism. The last layer of hidden state is mapped into the hourly probability of rime maintenance in the next 6 hours through a fully connected layer.

[0053] In any of the above technical solutions, further, the rime maintenance forecast model implements a phased parameter optimization strategy during the training process:

[0054] In the first stage, a global search algorithm is used to determine the structural parameter combination of the attention mechanism within the preset parameter space. In the second stage, the learning rate parameters and gradient constraints are dynamically adjusted based on the validation set feedback.

[0055] The model's generalization ability for unknown time series data is continuously monitored through the validation set. When the sliding standard deviation of the validation set's TS score falls below the set threshold for N consecutive training cycles, the early stopping mechanism is triggered. The periodic learning rate scheduling algorithm is then combined to balance the model convergence process.

[0056] From multiple model snapshots saved during training, the optimal model architecture is selected based on comprehensive evaluation indicators of the validation set.

[0057] The beneficial effects of the present invention are:

[0058] The technical solution in this invention accurately captures the critical meteorological conditions for the formation of rime for the first time through hourly data modeling; a method for predicting the maintenance time of rime is proposed to solve the problem of predicting the duration of the rime process; the hybrid model (random forest + Transformer) takes into account the feature extraction in the formation stage and the temporal dependency in the maintenance stage, and outputs a forecast product for the entire life cycle of rime (appearance-maintenance-ablation), which improves the accuracy of the rime process forecast and better serves scenarios such as scenic area opening time planning and tourist viewing arrangements.

[0059] During the implementation of the technical solution of the present invention, the SMOTE algorithm is introduced to upsample minority samples, which alleviates the sample imbalance problem caused by the scarcity of rime data and significantly improves the model's ability to recognize rime events. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The advantages of the above and additional aspects of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:

[0061] Figure 1 This is a schematic flow chart of a method for collaboratively predicting the formation and maintenance time of rime based on random forest-Transformer fusion according to an embodiment of the present invention;

[0062] Figure 2This is a model fusion architecture diagram of a collaborative forecasting method for rime formation and maintenance time based on random forest-Transformer fusion according to an embodiment of the present invention. DETAILED DESCRIPTION

[0063] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present invention and the features therein can be combined with each other without conflict.

[0064] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0065] like Figure 1 As shown, this embodiment provides a method for collaboratively predicting the formation and maintenance time of rime based on random forest-Transformer fusion, which includes:

[0066] S1. Collect meteorological elements of the target area from multi-source data and organize these data into hourly datasets containing ground meteorology, space meteorology, and rime occurrence records. By judging the meteorological conditions in the data, mark the time periods that meet specific conditions as rime "formation events", laying the data foundation for subsequent model training.

[0067] The multi-source data include ground meteorological elements in the target area, meteorological elements in various spatial layers, and records of the presence or absence of rime.

[0068] Step S1 specifically includes:

[0069] S11. Select a target area and obtain hourly ground meteorological observation data from a meteorological observation station in the target area. The hourly ground meteorological observation data includes records of air temperature, dew point temperature, relative humidity, wind speed, precipitation, and the presence or absence of rime. The presence of rime is recorded as 1, and the absence of rime is recorded as 0.

[0070] S12. Clean the acquired hourly ground meteorological observation data. Data cleaning includes removing values that exceed the reasonable range of meteorological elements and filling in the missing data of meteorological observation stations. The regression interpolation method is used to fill in the missing data.

[0071] S13. Based on the reanalysis data set of the same meteorological type as the target area, obtain the meteorological element data of each spatial layer through interpolation method. The data includes the temperature, relative humidity, wind speed of the upper layer, the current layer, and the lower layer, and obtain the low cloud cover and total cloud cover data.

[0072] For example, based on the Global Atmospheric Reanalysis Dataset (ERA5) dataset, when analyzing the Golden Summit of Mount Emei, the upper layer can be selected as 600hPa or above, the layer where it is located is 700hPa, and the lower layer can be selected as 850hPa.

[0073] S14: Formatting the hourly ground meteorological observation data cleaned in step S12 and the spatial meteorological element data of each layer obtained in step S13 is unified, and converting these meteorological data into a unified standard format includes but is not limited to:

[0074] Unified time format: All data times are unified to Beijing time to ensure time consistency between different data sources.

[0075] Unified unit formats: For example, the temperature unit should be unified as degrees Celsius, the wind speed should be unified as meters per second, and the units of other meteorological elements such as humidity also need to be unified.

[0076] S15. Arrange the data in a unified format in chronological order to ensure that the ground meteorological element and space meteorological element data at each moment match the records of the presence or absence of rime. The arranged data form a complete hourly data set within a time period.

[0077] S16. Retrieve the time periods that meet the conditional indicators in the hourly data set and mark them as "formation events". The conditional indicators include: rime, relative humidity > 98%, and temperature < 0°C.

[0078] S2. Extract predictors of rime formation from the hourly dataset, construct a rime formation prediction model using the random forest method, train the rime formation prediction model using samples from the hourly dataset, and combine cross-validation and grid search to achieve more comprehensive model evaluation and hyperparameter optimization, thereby improving the reliability and generalization ability of the model.

[0079] Step S2 specifically includes:

[0080] S21, ground temperature T and dew point temperature T d The data is processed to calculate the temperature dew point difference and temperature stratification. The calculation processes of the two are as follows:

[0081] Calculate the temperature dew point difference (TT d ):TT d =T 2m -T d2m ; Among them, T 2m is the temperature at 2m above the ground, T d2m It is the dew point temperature at 2m above the ground.

[0082] Calculate the temperature stratification T s :T s =T 下层 -T当前层 ; Take the Golden Summit of Mount Emei as an example, that is, T s =T 850 -T 700 , T 850 The temperature at 850hPa, T 700 The temperature is 700hPa.

[0083] S22. Add temperature dew point difference and temperature stratification to the hourly data set to extract the prediction factors of rime formation. The prediction factors of rime formation include: monthly, daily, ground temperature, temperature dew point difference, relative humidity, wind speed, precipitation, total cloud cover, low cloud cover, temperature stratification and temperature, relative humidity and wind speed of each layer, a total of 19 factors.

[0084] S23. Since rime formation events are relatively rare in actual data, in order to improve the model's learning ability for minority class samples, the SMOTE (Synthetic Minority Oversampling Technique) method is used to upsample the data. The specific method is:

[0085] For each minority class sample, select the other minority class samples of its nearest neighbors and generate new sample points between them. The generation formula is: X new =X minority +λ×(X neighbor -X minority ); where λ is a random number in the range [0,1], X minority Represents the feature vector of a sample in the minority class, X neighbor Indicates that there are more samples in the minority class than in X minority The feature vector of another sample of the nearest neighbor, X new According to X minority and X neighbor The new sample is generated, and its feature vector is obtained by linear interpolation between these two samples.

[0086] Through this method, the number of samples of rime formation events is effectively supplemented, thereby improving the imbalance problem of data category distribution.

[0087] S24. Divide the samples into training set and test set in a ratio of 7:3, and use the random forest algorithm to construct a rime formation prediction model on the training set. The output of the model is the rime formation prediction probability. The random forest model consists of multiple decision trees, and each decision tree is independently constructed on a randomly selected sample subset and feature subset. The K-fold cross-validation method is used to train and validate the training set multiple times. The training set is evenly divided into K subsets, and one of them is selected as the validation set each time, and the remaining K-1 subsets are used as the training set. This is repeated K times to obtain the average performance of the model in different data partitions under each set of parameter configurations.

[0088] During the cross-validation process, a comprehensive search was conducted on the key hyperparameters in the rime formation prediction model, including the number of trees, maximum depth, minimum number of split samples, etc. Specifically, the candidate value range of each hyperparameter was set in advance to form a parameter grid. Then, K-fold cross-validation was performed on each hyperparameter combination in the grid, and its performance indicators on the validation set were recorded. The performance indicators include accuracy and TS (Threat Score) score. The TS score calculation formula is: TS = number of hits / (number of hits + number of null reports + number of missed reports), where the number of hits is the number of events that were predicted to occur and actually occurred, the number of null reports is the number of events that were predicted to occur but did not actually occur, and the number of missed reports is the number of events that were predicted not to occur but actually occurred.

[0089] After traversing all candidate parameter combinations, a set of hyperparameter configurations with the best performance indicators is selected, and the rime formation prediction model is retrained on the entire training set using this optimal configuration.

[0090] Finally, the rime formation prediction model optimized by cross-validation and grid search not only has high prediction accuracy on the training set, but also shows good generalization ability and robustness on the test set. It can accurately capture the mutation characteristics in hourly meteorological data and achieve accurate forecast of rime formation events.

[0091] S3. Extract the rime maintenance prediction factors from the hourly dataset, use the Transformer method to build a rime maintenance prediction model, and use the samples in the hourly dataset to train the rime maintenance prediction model.

[0092] Step S3 specifically includes:

[0093] S31. Extract rime maintenance prediction factors from the hourly data set. The basic factors for rime maintenance prediction include: time, temperature, relative humidity, wind speed, precipitation, and rime record, a total of 6 factors.

[0094] S32. Construct derivative factors based on the basic factors of rime maintenance forecast, calculate the moving average and variance of temperature and humidity in the 24 hours before the current moment, so as to capture the statistical characteristics of meteorological data in the short term; the rime maintenance forecast factors include all basic factors and derivative factors of rime maintenance forecast, namely: time, temperature, relative humidity, wind speed, precipitation, mean temperature, temperature variance, mean relative humidity, relative humidity variance and rime record, a total of 10 factors.

[0095] S33. Sine / cosine transformation is performed on the time variable to retain the periodic characteristics and enhance the model's ability to learn the law of the circadian cycle. The transformed time Enc(t) is expressed as:

[0096]

[0097] Numerical features other than rime records were normalized using the Z-Score method to ensure consistency in the numerical scale of different features. The processing formula is:

[0098]

[0099] Among them, Z is the result of Z-Score standardization processing, x represents the processed feature, μ represents the mean of the training set, and σ represents the standard deviation of the training set.

[0100] S34. Since rime samples are relatively rare in actual data, in order to improve the model's learning ability for minority class samples, a dynamic window sampling strategy is used to solve the class imbalance problem. The specific method is:

[0101] The positive sample window is defined as the period from 24 hours before the occurrence of the rime event to the end of the event, and dense sliding sampling is performed in this area with a step size of 1 hour; the negative sample area is sparsely sampled with a step size of 6 hours. Through this method, the number of rime samples is effectively supplemented, thereby improving the imbalance problem of data category distribution.

[0102] S35. The dataset is divided into training set, validation set, and test set in chronological order to ensure temporal integrity. The ratio of training set, validation set, and test set is 7:2:1. The Transformer method is used to construct a rime maintenance forecast model and it is trained using the training set. The output of the model is the rime maintenance forecast probability. The Transformer model has a strong ability to capture temporal dependencies and is suitable for processing continuous meteorological data and its complex dynamic change characteristics.

[0103] Specifically, the working methods of the rime maintenance forecast model include:

[0104] First, the feature vectors of 24 consecutive hours are stacked in chronological order to form a two-dimensional feature matrix with 24 rows and 10 columns as the model input, where each row corresponds to one hour of meteorological data and each column represents a different forecast factor.

[0105] A learnable positional encoding matrix is introduced into the input layer and added to the linearly transformed input features, enabling the model to explicitly perceive the temporal order of meteorological data. Multiple parallel attention heads are set up, each of which independently calculates the attention weights for the query, key, and value vectors. The association strength at each time step is calculated through scaled dot products. The outputs of multiple heads are concatenated and linearly transformed, fusing the feature representations of different subspaces to enhance the model's ability to capture key patterns such as sudden temperature changes and sustained humidity changes. Two fully connected layers are connected after the self-attention layer, and the ReLU activation function is used to introduce nonlinear transformations, improving the model's ability to represent complex combinations of meteorological features. Residual connections are implemented after each self-attention layer and feedforward network layer, adding the sub-layer output to the original input. Layer normalization is then performed to effectively alleviate the vanishing gradient problem and accelerate model convergence.

[0106] Using a 24-hour meteorological feature sequence as training input, a neural network model incorporating a self-attention mechanism captures long-range dependencies between spatiotemporal features. The last hidden state is mapped via a fully connected layer to the hourly probability of rime maintenance for the next six hours. A phased parameter optimization strategy is implemented during training. In the first phase, a global search algorithm is used to determine the structural parameter combination of the attention mechanism within a preset parameter space. In the second phase, the learning rate parameter and gradient constraints are dynamically adjusted based on validation set feedback. The model's generalization ability to unknown time series data is continuously monitored using the validation set. An early stopping mechanism is triggered when the sliding standard deviation of the validation set's TS score falls below a set threshold for N consecutive training cycles. A periodic learning rate scheduling algorithm is then used to balance the model's convergence. From multiple model snapshots saved during training, the optimal model architecture is selected based on comprehensive validation set evaluation metrics. Ultimately, the rime maintenance forecast model demonstrates not only high prediction accuracy on the validation set but also good generalization on the test set.

[0107] S4. Combine the rime formation forecast model and the rime maintenance forecast model to process the meteorological data of the target area within the forecast period, and finally output the hourly forecast of whether rime will appear and the duration of the entire rime process.

[0108] Step S4 specifically includes:

[0109] S41. Based on the model forecast product or the deterministic forecast product of meteorological elements, obtain the hourly ground meteorological forecast data of the target area, including air temperature, dew point temperature, relative humidity, wind speed, and precipitation, and obtain the air temperature, relative humidity, wind speed, low cloud cover, and total cloud cover data of the upper layer, the current layer, and the lower layer; at the same time, calculate the temperature dew point difference and the temperature stratification characteristic vector, and use all the above data to construct the forecast period characteristic vector.

[0110] S42, input the characteristic vector into the rime formation prediction model to obtain the rime formation probability P form, calculate the forecast probability of rime formation T dynamic , the calculation formula is:

[0111] T dynamic =P form +0.05×RH min -0.02×(T current +2);

[0112] Among them, RH min T is the lowest humidity value in the past 6 hours at the forecast time. current The temperature at the forecast time.

[0113] S43, continuously monitor the rime and determine the dynamic threshold value T dynamic , when the dynamic threshold value T of rime is determined for 6 consecutive hours dynamic When the value is >0.65, the characteristic vector is input into the rime maintenance forecast model to obtain the rime maintenance forecast probability P for the next 6 hours. maintain If the rime maintains the forecast probability P in the next 6 hours maintain >0.65, the feature vector is input into the rime maintenance forecast model again to obtain the rime maintenance forecast probability in a longer period of time until its value is less than 0.65.

[0114] In the above process, when there is a forecast probability of rime formation at a certain time t, T dynamic and rime maintenance forecast probability P maintain When there are two forecast results, the final probability of rime at that time is P final(t) T dynamic 、P maintain The maximum value of the two is the final probability of rime at that time, P final(t) When the value is >0.65, it is determined that there is rime at that time, otherwise there is no rime, thus obtaining the hourly rime forecast for the forecast period; intervals no more than 2 hours are regarded as the same rime process, and the total duration from rime formation to melting is calculated, and the rime duration forecast is output.

[0115] In summary, the present invention proposes a collaborative prediction method for rime formation and maintenance time based on random forest-Transformer fusion, including:

[0116] S1. Collect meteorological elements of the target area from the data of ground meteorological elements, meteorological elements of various layers in space, and records of the presence or absence of rime in the target area, and organize these data into an hourly data set containing ground meteorological, space meteorological, and rime occurrence records.

[0117] S2. Extract the prediction factors of rime formation from the hourly dataset, use the random forest method to build a rime formation prediction model, use the samples in the hourly dataset to train the rime formation prediction model, and use a combination of cross-validation and grid search to evaluate and optimize the model.

[0118] S3. Extract the rime maintenance prediction factors from the hourly dataset, use the Transformer method to build a rime maintenance prediction model, and use the samples in the hourly dataset to train the rime maintenance prediction model.

[0119] S4. Combine the rime formation forecast model and the rime maintenance forecast model to process the meteorological data of the target area within the forecast period, and finally output the hourly forecast of whether rime will appear and the duration of the entire rime process.

[0120] The steps in the present invention can be adjusted in sequence, combined, or deleted according to actual needs.

[0121] The units in the device of the present invention can be combined, divided and deleted according to actual needs.

[0122] Although the present invention has been disclosed in detail with reference to the accompanying drawings, it should be understood that these descriptions are merely illustrative and are not intended to limit the application of the present invention. The scope of the present invention is defined by the appended claims and includes various modifications, variations, and equivalents made to the invention without departing from the scope and spirit of the present invention.

Claims

1. A collaborative prediction method for rime formation and duration based on random forest-Transformer fusion, characterized by: The method comprises: S1. Collect meteorological elements of the target area from the data of ground meteorological elements, meteorological elements of various layers in space, and records of the presence or absence of rime in the target area, and organize these data into an hourly dataset containing ground meteorological, space meteorological, and rime occurrence records; S2. Extract rime formation prediction factors from the hourly dataset, upsample the data using the SMOTE method to expand the sample size, construct a rime formation prediction model using the random forest method, train the rime formation prediction model using samples from the hourly dataset, and evaluate and optimize the rime formation prediction model using a combination of cross-validation and grid search. S3. Extract rime maintenance prediction factors from the hourly dataset, construct a rime maintenance prediction model using the Transformer method, increase the number of samples with rime in the training set through a dynamic window sampling strategy, and train the rime maintenance prediction model using samples from the hourly dataset; S4. Processing meteorological data of the target area within the forecast period in combination with the rime formation forecast model and the rime maintenance forecast model, ultimately outputting hourly rime presence or absence forecasts and the duration of the entire rime process; Specifically, step S4 includes: S41. Obtain hourly ground meteorological forecast data for the target area based on the model forecast product or the meteorological element deterministic forecast product, and use the data to construct a feature vector for the forecast period; S42, input the characteristic vector into the rime formation prediction model to obtain the rime formation probability P form , calculate the forecast probability of rime formation T dynamic ; S43, continuously monitor rime and determine the dynamic threshold value T dynamic , when the dynamic threshold value T of rime is determined for 6 consecutive hours dynamic When the value is >0.65, the characteristic vector is input into the rime maintenance forecast model to obtain the rime maintenance forecast probability P for the next 6 hours. maintain If the rime maintains the forecast probability P in the next 6 hours maintain >0.65, the feature vector is input into the rime maintenance forecast model again to obtain the rime maintenance forecast probability in a longer period of time until its value is less than 0.65; In the above process, when there is rime formation at a certain time t, the forecast probability T dynamic and rime maintenance forecast probability P maintain When there are two forecast results, the final probability of rime at that time is P final(t) T dynamic 、P maintain The maximum value of the two is the final probability of rime at that time, P final(t) When the value is >0.65, it is determined that there is rime at that time, otherwise there is no rime, thus obtaining the hourly rime forecast for the forecast period; intervals no more than 2 hours are regarded as the same rime process, and the total duration from rime formation to melting is calculated, and the rime duration forecast is output.

2. The method for collaboratively predicting the formation and duration of rime based on random forest-Transformer fusion according to claim 1, characterized in that: The rime formation prediction probability T dynamic The calculation formula is: T dynamic =P form +0.05×RH min -0.02×(T current +2); Among them, RH min T is the lowest humidity value in the past 6 hours at the forecast time. current The temperature at the forecast time.

3. The method for collaboratively predicting the formation and duration of rime based on random forest-Transformer fusion according to claim 1, characterized in that: The step S1 specifically includes: S11, selecting a target area and obtaining hourly ground meteorological observation data from a meteorological observation station in the target area, wherein the hourly ground meteorological observation data includes records of air temperature, dew point temperature, relative humidity, wind speed, precipitation, and the presence or absence of rime, where the presence of rime is recorded as 1 and the absence of rime is recorded as 0; S12. Clean the acquired hourly ground meteorological observation data. Data cleaning includes removing values that exceed the reasonable range of meteorological elements and filling in missing data at meteorological observation stations. Regression interpolation is used to fill in missing data. S13. Based on the reanalysis dataset of the same meteorological type as the target area, obtain the meteorological element data of each spatial layer through interpolation method. The data includes the temperature, relative humidity, wind speed of the upper layer, the current layer, and the lower layer, and obtain the low cloud cover and total cloud cover data; S14, performing format unification processing on the ground hourly meteorological observation data cleaned in step S12 and the spatial meteorological element data of each layer obtained in step S13; S15. Arrange the data in the unified format in chronological order to ensure that the ground meteorological element and space meteorological element data at each moment match the record of the presence or absence of rime. The arranged data form a complete hourly data set for a time period. S16. Retrieve the hourly data set for times that meet the conditional indicators and mark them as "formation events." The conditional indicators include: rime, relative humidity > 98%, and temperature < 0°C.

4. The method for collaboratively predicting the formation and duration of rime based on random forest-Transformer fusion according to claim 1, characterized in that: The step S2 specifically includes: S21, ground temperature T and dew point temperature T d The data is processed to calculate the temperature dew point difference and temperature stratification. The calculation processes of the two are as follows: Calculate the temperature dew point difference TT d :TT d =T 2m -T d2m ; Among them, T 2m is the temperature at 2m above the ground, T d2m It is the dew point temperature at 2m above the ground. Calculate the temperature stratification T s :T s =T 下层 -T 当前层 ; S22. Add the temperature dew point difference and temperature stratification to the hourly data set to extract the prediction factors of rime formation; S23. Use the SMOTE method to upsample the data, specifically including: For each minority class sample, select the other minority class samples of its nearest neighbors and generate new sample points between them. The generation formula is: X new =X minority +λ×(X neighbor -X minority ); where λ is a random number in the range [0,1], X minoryity Represents the feature vector of a sample in the minority class, X neighbor Indicates that there are more samples in the minority class than in X minority The feature vector of another sample of the nearest neighbor, X new According to X minority and X neighbor The new sample is generated, and its feature vector is obtained by linear interpolation between these two samples; S24. Divide the samples into a training set and a test set in a ratio of 7:

3. Use the random forest algorithm to construct a rime formation prediction model on the training set. The output of the model is the rime formation prediction probability. The random forest model consists of multiple decision trees, each of which is independently constructed on a randomly selected sample subset and feature subset. Use the K-fold cross-validation method to train and validate the training set multiple times. The training set is evenly divided into K subsets. Each time, one of the subsets is selected as the validation set, and the remaining K-1 subsets are used as the training set. Repeat this process K times to obtain the average performance of the model in different data partitions under each parameter configuration. During the cross-validation process, a comprehensive search was conducted on the key hyperparameters in the rime formation prediction model, including the number of trees, maximum depth, and minimum number of split samples. Specifically, the candidate value ranges of each hyperparameter were pre-set to form a parameter grid. Then, K-fold cross-validation was performed on each hyperparameter combination in the grid, and its performance indicators on the validation set were recorded. The performance indicators include accuracy and TS score. The TS score calculation formula is: TS = number of hits / (number of hits + number of null reports + number of missed reports), where the number of hits is the number of events that were predicted to occur and actually occurred, the number of null reports is the number of events that were predicted to occur but did not actually occur, and the number of missed reports is the number of events that were predicted not to occur but actually occurred. After traversing all candidate parameter combinations, a set of hyperparameter configurations with the best performance indicators is selected, and the rime formation prediction model is retrained on the entire training set using this optimal configuration.

5. The method for collaboratively predicting the formation and duration of rime based on random forest-Transformer fusion according to claim 4, characterized in that: The prediction factors for rime formation include: month, day, ground temperature, temperature dew point difference, relative humidity, wind speed, precipitation, total cloud cover, low cloud cover, temperature stratification and temperature, relative humidity and wind speed of each layer, a total of 19 factors.

6. The method for collaboratively predicting the formation and duration of rime based on random forest-Transformer fusion according to claim 1, characterized in that: The step S3 specifically includes: S31. Extracting rime maintenance prediction factors from the hourly data set. The rime maintenance prediction basic factors include: time, temperature, relative humidity, wind speed, precipitation, and rime record, a total of 6 factors; S32. Constructing derived factors based on the basic factors for rime maintenance forecasting, calculating the moving averages and variances of the temperature and humidity in the 24 hours preceding the current time to capture the statistical characteristics of short-term meteorological data; the rime maintenance forecasting factors include all basic factors and derived factors for rime maintenance forecasting, namely, time, temperature, relative humidity, wind speed, precipitation, mean temperature, temperature variance, mean relative humidity, relative humidity variance, and rime record, a total of 10 factors; S33. Sine / cosine transformation is performed on the time variable to retain the periodic characteristics and enhance the model's ability to learn the law of the circadian cycle. The transformed time Enc(t) is expressed as: Numerical features other than rime records were normalized using the Z-Score method to ensure consistency in the numerical scale of different features. The processing formula is: Where Z is the result of Z-Score standardization, x represents the processed feature, μ represents the mean of the training set, and σ represents the standard deviation of the training set; S34. A dynamic window sampling strategy is used to solve the class imbalance problem. The positive sample window is defined as the period from 24 hours before the rime event to the end of the event. Dense sliding sampling is performed in this area with a step size of 1 hour. Sparse sampling is used in the negative sample area with a step size of 6 hours to effectively supplement the number of rime samples and improve the imbalance of data class distribution. S35. The dataset is divided into training set, validation set, and test set in chronological order to ensure temporal integrity. The ratio of training set, validation set, and test set is 7:2:

1. The Transformer method is used to construct a rime maintenance forecast model and it is trained using the training set. The output of the model is the rime maintenance forecast probability.

7. The method for collaboratively predicting the formation and duration of rime based on random forest-Transformer fusion according to claim 6, characterized in that: The working method of the rime maintenance forecast model comprises: First, the feature vectors of 24 consecutive hours are stacked in chronological order to form a two-dimensional feature matrix with 24 rows and 10 columns as the model input, where each row corresponds to one hour of meteorological data and each column represents a different forecast factor; A learnable positional encoding matrix is introduced into the input layer and added to the linearly transformed input features, enabling the model to explicitly perceive the temporal order of meteorological data. Multiple parallel attention heads are set up, each of which independently calculates the attention weights of the query, key, and value vectors, and calculates the association strength of each time step through scaled dot product calculations. The outputs of multiple heads are concatenated and linearly transformed to fuse the feature representations of different subspaces. A two-layer fully connected network is connected after the self-attention layer, using ReLU as the activation function. Residual connections are implemented after each self-attention layer and feedforward network layer, and the sub-layer output is added to the original input, followed by layer normalization. The 24-hour meteorological feature sequence of the training set is used to capture the long-range dependencies between spatiotemporal features through a neural network model with a self-attention mechanism. The last layer of hidden state is mapped into the hourly probability of rime maintenance in the next 6 hours through a fully connected layer.

8. The method for collaboratively predicting the formation and duration of rime based on random forest-Transformer fusion according to claim 7, characterized in that: The rime maintenance forecast model implements a phased parameter optimization strategy during the training process: In the first stage, a global search algorithm is used to determine the structural parameter combination of the attention mechanism within the preset parameter space. In the second stage, the learning rate parameters and gradient constraints are dynamically adjusted based on the validation set feedback. The model's generalization ability for unknown time series data is continuously monitored through the validation set. When the sliding standard deviation of the validation set's TS score falls below the set threshold for N consecutive training cycles, the early stopping mechanism is triggered. The periodic learning rate scheduling algorithm is then combined to balance the model convergence process. From multiple model snapshots saved during training, the optimal model architecture is selected based on comprehensive evaluation indicators of the validation set.

Citation Information

Patent Citations

  • Local cloud and fog prediction method and system

    CN114926744A

  • Rime meteorological landscape forecasting method and system, electronic equipment and medium

    CN117518296A

  • Urban multi-source short-duration rainfall data fusion correction method considering neighborhood rainfall

    CN118348616A

Cited By

  • Video image and weather fusion-based cloud sea landscape short-time forecasting method and system

    CN121505454A