A heavy fog time series prediction method and system for small sample scenarios
By splitting the fog time series forecast model into single-time point samples and constructing a multi-dimensional feature set, combined with data augmentation and model optimization, the problems of sample scarcity and loss of time series information in fog forecasting are solved, achieving high-precision hourly visibility forecasts, which are applicable to airports, highway traffic and other fields.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI PROVINCIAL PUBLIC METEOROLOGICAL SERVICE CENT
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-19
AI Technical Summary
In fog forecasting, due to the scarcity of high-quality time series samples, the accuracy of time series forecasts is low, and time series information is easily lost, making it difficult for existing technologies to meet operational needs.
By decoupling time series, long-time series samples are split into independent single-time point samples, and a multi-dimensional feature set integrating physical mechanisms and temporal features is constructed. Combined with hierarchical sampling, oversampling and downsampling data augmentation strategies, the LightGBM model is adopted and the training process is optimized by a custom scoring function to restore the temporal consistency of the initial visibility.
It achieves accurate hourly visibility forecasts under small sample conditions, and exhibits excellent forecasting performance, especially in low visibility and dense fog scenarios. The model is lightweight and highly interpretable, and has good prospects for engineering applications.
Smart Images

Figure CN121682749B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of weather time series forecasting technology, specifically to a method and system for forecasting heavy fog in small sample scenarios. Background Technology
[0002] Dense fog significantly reduces visibility, posing a serious threat to the safe operation of aviation takeoffs and landings, road traffic, and other sectors. Accurate hourly forecasts of airport visibility (including dominant visibility and runway visual range) are crucial for ensuring on-time flight scheduling and preventing flight delays and safety incidents. Traditional fog forecasting methods largely rely on numerical weather prediction models or empirical statistical models; however, numerical forecasts have limited ability to capture small-scale meteorological processes, and empirical models lack generalization ability.
[0003] In recent years, deep learning models, such as LSTM and Transformer, have demonstrated advantages in time series forecasting tasks. However, these models are highly dependent on massive amounts of high-quality time series data. In actual fog forecasting scenarios, complete, continuous, high-quality time series observation data are extremely scarce. After data matching and quality control, only a small number of effective samples are often obtained, leading to overfitting in deep learning models and making it difficult to meet operational needs. Therefore, developing high-precision fog time series forecasting technology suitable for small sample scenarios has become an important research topic in the meteorological service field. Current fog forecasting technology mainly faces the following problems:
[0004] The scarcity and uneven distribution of samples: There are very few high-quality fog time series samples, and the data exhibits an extreme "long-tail distribution" characteristic. Low visibility fog samples account for less than 3%, but such scenarios are the core focus of forecasting.
[0005] The physical mechanism is complex: visibility changes are affected by multiple meteorological factors such as temperature, humidity, wind speed, and inversion layer, exhibiting strong nonlinear characteristics, and the correlation between a single factor and visibility is extremely low.
[0006] Challenges in preserving temporal information: Directly using temporal models in small sample scenarios can easily lead to overfitting, while splitting temporal data into single-point samples can result in the loss of temporal continuity information, making it difficult to capture the impact of periodic changes such as sunrise and sunset on fog formation and dissipation.
[0007] The invention patent with publication number CN120871298A discloses a method and device for quantitative forecasting of sea fog using a periodically enhanced Transformer. The paper describes a multi-source approach to acquire sea fog data for the target area. While multi-source data can indeed alleviate the problems of insufficient effective samples and incomplete information to some extent by expanding observation coverage, supplementing missing data, and providing more influencing factors, multi-source data does not necessarily equate to an "increased availability of effective samples." It usually comes with additional processing costs such as spatiotemporal matching, format standardization, error consistency, and label alignment, and may introduce noise. Furthermore, the sample amplification method in this patent involves temporal resampling and spatial combined sampling. This amplification method is highly dependent on high-resolution data sources and does not consider the joint temporal-spatial dependence. The scarcity of fog samples is not subjectively caused by data collection methods, but rather stems from the low probability of fog events themselves, which is an objectively existing statistical law.
[0008] The invention patent with publication number CN112037906A discloses a method and system for expanding sample data of long-term physiological signal time series. Although this patent involves sample expansion, it is based on feature / indicator layer expansion using multi-timescale analysis. The patent uses indicators from different time scales as indicators for different samples to expand the samples. These "new samples" are highly correlated with the original samples, essentially representing an amplification at the feature / representation level. This patent requires screening and verifying indicators that are not significantly correlated with the time scale and have differences before expanding different scales as different samples; otherwise, it will introduce systematic bias. However, fog / visibility evolution is typically sensitive to time scales (diurnal variation, boundary layer processes, persistence / dissipation phases), making the multi-timescale sample expansion method in this patent unsuitable for expanding fog time series samples. Summary of the Invention
[0009] The technical problem to be solved by this invention is: the low accuracy of time series forecasts in fog forecasting due to the scarcity of high-quality time series samples.
[0010] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0011] A method for time-series fog prediction in scenarios with small sample sizes includes:
[0012] The long-term time series samples, including real-time observations and numerical forecasts, are expanded by splitting them into multiple independent single-time-point samples based on time series decoupling.
[0013] Based on single-time-point samples, a multi-dimensional feature set integrating physical mechanisms and temporal features is constructed to obtain forecast fusion features and compensate for the loss of time series information.
[0014] The data augmentation strategy, which combines stratified sampling, oversampling, and downsampling, is used to balance the sample distribution of forecast fusion features;
[0015] The data-enhanced forecast fusion features are used as input to the forecast model to obtain the initial visibility for hourly forecasts;
[0016] Perform time-series consistency repair on the initial visibility and obtain the repaired visibility.
[0017] In this embodiment, the long-term time series samples, including real-time observations and numerical forecasts, are decomposed into multiple independent single-time-point samples based on time-series decoupling, including:
[0018] The original long-time series sample containing "24-hour input - 15-hour output" was split into 15 independent single-time point samples; each single-time point sample uses the cumulative observation data and numerical forecast data of the 24 hours before 15:00 on the same day as input features.
[0019] In this embodiment, a multi-dimensional feature set integrating physical mechanisms and temporal features is constructed to obtain forecast fusion features, including:
[0020] Based on the meteorological and physical principles of fog formation, the physical mechanism is constructed by identifying vertical structure and near-surface features, including inversion intensity features, wind shear features, and humidity features, as physical mechanism characteristics.
[0021] The time features include the month, hourly sine / cosine encoding, solar altitude angle, and time difference;
[0022] Real-time observation data at 15:00 on the same day was used as the observation anchor point feature, and the observation statistical features of the forecast day were also included.
[0023] By splicing together observation anchor point features, numerical forecast features, physical mechanism features, temporal features, and observational statistical features, forecast fusion features are obtained.
[0024] In this embodiment, the observation anchor point features are shared among multiple single-time-point samples obtained by splitting the same original long-time-series sample.
[0025] In this embodiment, a data augmentation strategy combining hierarchical sampling, oversampling, and downsampling is used to balance the sample distribution of forecast fusion features, including:
[0026] Stratified sampling: fine-grained binning is performed according to visibility threshold, with dominant visibility binning at intervals of 1000~2000 meters;
[0027] Oversampling: Low visibility scarce samples are replicated and expanded, with visibility samples ≤1000 meters expanded by 3 to 5 times and visibility samples 1000 to 6000 meters expanded by 1 to 3 times; and noise is injected into the replicated and expanded samples to simulate the natural disturbance of real meteorological data.
[0028] Downsampling: High visibility samples above 10,000 meters are randomly discarded to balance the overall sample distribution.
[0029] In this embodiment, a custom scoring function is used to optimize the training process of the forecast model; wherein, the custom scoring function is:
[0030] When visibility is less than 1000 meters, the absolute error is used as the core evaluation criterion. Attention is paid to whether the visibility forecast crosses the key threshold, which meets the high-precision forecasting needs of low visibility scenarios.
[0031] When visibility is ≥1000 meters, relative error is used as the core evaluation criterion to balance the prediction accuracy of different visibility ranges.
[0032] In this embodiment, obtaining the repaired visibility includes: obtaining the visibility after timing consistency repair by solving the following optimization problem:
[0033] ;
[0034] In the formula, For the first Initial visibility forecast for the hour, For the first time sequence consistency repair Hourly forecast visibility For the first The confidence weight of the hourly forecast is used to control the degree to which the time-series consistency repair result follows the original forecast at the corresponding time. For the first The smoothing strength of hourly forecasts is used to measure and suppress the magnitude of changes in forecasts between adjacent hours. The minimum objective.
[0035] In this embodiment, the first The smoothing strength of the hourly forecast is obtained using the following formula:
[0036] ;
[0037] In the formula, For the first Hourly predicted smoothing strength For the first The baseline value of smoothness intensity per hour, For the first The fog risk index function for hours For the first Hourly time adjustment factor.
[0038] In this embodiment, the first Hourly smoothness intensity base value It can be obtained through the following formula:
[0039] ;
[0040] ;
[0041] No. Hourly fog risk index function It can be obtained through the following public announcement:
[0042] ;
[0043] No. Hourly time adjustment factor It can be obtained through the following formula:
[0044] ;
[0045] In the formula, For the first The fog potential index for hours is used to characterize the fog potential index at the time of the fog. The strength of the physical conditions that cause fog to form in an hour. This is an indicator function that returns 1 if the condition is true, and 0 otherwise. For the first Hourly solar altitude angle For the first hourly relative humidity, For the first 1000 hPa and ground temperature inversion per hour, For the first Hourly ground wind speed, This is the time difference between the forecast time and the initial reporting time.
[0046] The present invention also provides a system for applying the above-described fog time-series forecasting method for small sample scenarios, comprising:
[0047] The time series decoupling module expands the long-term time series samples based on real-time observation and numerical forecast by splitting them into multiple independent single-time-point samples.
[0048] The feature engineering module constructs a multi-dimensional feature set that integrates physical mechanisms and temporal features based on single-time-point samples, obtains forecast fusion features, and makes up for the loss of time series information;
[0049] The data augmentation module uses a data augmentation strategy that combines hierarchical sampling, oversampling, and downsampling to balance the sample distribution of forecast fusion features;
[0050] The forecast module uses the data-enhanced forecast fusion features as input to the forecast model to obtain the initial visibility for hourly forecasts;
[0051] The timing repair module performs timing consistency repair on the initial visibility and obtains the repaired visibility.
[0052] Compared with the prior art, the beneficial effects of the present invention are:
[0053] To address the technical challenges of high-quality time-series samples, uneven data distribution, complex physical mechanisms, and easy loss of time-series information in fog forecasting, this paper proposes a small-sample fog time-series forecasting technique based on LightGBM. This technique transforms the long-term time-series forecasting task into a single-point regression problem through a time-series decoupling strategy. It balances the sample distribution by combining hierarchical sampling and Gaussian noise injection data augmentation methods, constructs a multi-dimensional feature set integrating physical mechanisms and temporal features, optimizes the model training process using a custom scoring function, and corrects the temporal consistency of the initial forecast results. The output sequence is prone to jitter, discontinuity, and unreasonable jumps, thus reducing model usability and physical plausibility. Experimental results show that, with only 729 original valid time-series samples, this technique achieves accurate hourly visibility forecasts from 22:00 to 12:00 the following morning through sample expansion and feature engineering optimization. It exhibits excellent forecasting performance, especially in low-visibility fog scenarios. Furthermore, the model is lightweight, highly interpretable, and has promising engineering application prospects.
[0054] Temporal decoupling and sample augmentation techniques: To address the challenge of small samples, long-term temporal tasks are transformed into single-point regression problems, increasing the sample size by 15 times. At the same time, temporal awareness is reconstructed through explicit temporal coding, solving the problem of lost temporal information.
[0055] Fine-grained data augmentation strategy: Layered sampling based on visibility binning combined with Gaussian noise injection not only balances the sample distribution, but also improves the model's generalization ability by simulating meteorological disturbances, avoiding overfitting caused by simple replication.
[0056] Physical sensing feature engineering: By integrating features with clear meteorological significance, such as inversion intensity, wind shear, and solar altitude angle, the forecast model predictions are based on both statistical data patterns and the physical mechanisms of fog formation, thus improving the reliability and interpretability of the prediction results.
[0057] Business-oriented model optimization: Combining custom scoring functions with early stopping mechanisms ensures that model training aligns with actual business needs, enabling high-precision forecasting in critical low-visibility scenarios.
[0058] By performing joint optimization based on temporal consistency repair on the predicted result sequence, unreasonable hourly jumps can be significantly reduced while maintaining the ability to fit the original prediction, thereby improving the continuity, stability and physical rationality of the output sequence.
[0059] In time sequence consistency repair, smoothing intensity suppresses hourly jitter and anomalous jumps. By increasing the adjacent difference penalty, the difference in output sequence values is significantly reduced, thus minimizing "sawtooth" fluctuations.
[0060] Fog Potential Index: This index links smoothing intensity to fog physical conditions. When fog maintenance conditions are stronger, the base value of smoothing intensity is larger, suppressing unreasonable jumps. When conditions are not met, smoothing is automatically weakened to avoid over-smoothing.
[0061] Fog Risk Index: When relative humidity is high, temperature inversion is significant, and wind is low, the smoothing intensity is enhanced. Constraints are only strengthened when risk conditions are met, avoiding unnecessary smoothing of effective signals under ordinary weather conditions, thus balancing stability and sensitivity.
[0062] Lead time adjustment factor: to compensate for the accumulation of uncertainty caused by the increase in lead time. It gradually enhances smoothness as the forecast lead time increases, so that there are fewer jitters and jumps in the sequence in the middle and later stages (such as the 10th to 15th hours), thereby reducing the instability of long lead time output.
[0063] Lightweight deployment advantages: The forecast model is built on LightGBM, which has low computing power requirements, supports millisecond-level CPU inference, and can be directly integrated into existing airport meteorological terminals or embedded devices to achieve edge computing deployment without additional hardware investment.
[0064] High interpretability: Through feature importance analysis, the key influencing factors of fog formation can be clearly presented to forecasters, enhancing human-machine trust and assisting forecasters in decision-making;
[0065] Multi-scenario adaptability: In addition to airport visibility forecasting, this technical solution can be transferred to other visibility-sensitive fields such as highway transportation and port shipping, and has broad application and promotion value;
[0066] Significant operational benefits: It can provide accurate early warnings of foggy weather, effectively reducing losses such as flight delays and traffic congestion, and providing strong technical support for public safety. Attached Figure Description
[0067] Figure 1 This is a flowchart of a fog time-series forecasting method for small sample scenarios according to an embodiment of the present invention.
[0068] Figure 2 This is a schematic diagram of timing decoupling in an embodiment of the present invention.
[0069] Figure 3 This is a schematic diagram illustrating the verification of forecast results in an embodiment of the present invention.
[0070] Figure 4 This is a block diagram of a fog time-series forecasting system for small sample scenarios according to an embodiment of the present invention. Detailed Implementation
[0071] To facilitate understanding of the technical solution of the present invention by those skilled in the art, the technical solution of the present invention will now be further described in conjunction with the accompanying drawings.
[0072] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0073] Please see Figure 1 As shown, this invention provides a method for time-series fog prediction in small-sample scenarios, including:
[0074] S10 expands the sample by splitting the long-term time series samples, which include real-time observations and numerical forecasts, into multiple independent single-time-point samples through time series decoupling.
[0075] In one embodiment of the present invention, the input data includes three categories:
[0076] Real-time observation data: Hourly observation data of the station for the 24 hours before 15:00 on the same day, covering indicators such as dominant visibility, runway visual range, temperature, relative humidity, 10-meter wind speed, and high and low cloud cover.
[0077] Numerical forecast data: ECMWF (European Centre for Medium-Range Weather Forecasts) 3-hourly numerical forecast data reported starting at 20:00 the previous day, including temperature, relative humidity, wind speed at multiple pressure levels such as 1000hPa, 950hPa, 925hPa, 850hPa, and 700hPa, as well as sea level pressure and 3-hour precipitation.
[0078] Auxiliary information data: site geographical location information and time information (hour, month, solar altitude angle, etc.).
[0079] In this embodiment, data preprocessing is performed before temporal decoupling of the samples. To address the issue of inconsistent temporal resolution, linear interpolation is used to complete the ECMWF data hourly, filling in the temporal resolution gaps. Quality control is applied to the raw data, removing outliers and samples with excessively high missing values to ensure data reliability. Finally, after preprocessing, a dataset with a uniform structure and standardized format is formed, laying the foundation for subsequent feature engineering and model training.
[0080] Please see Figure 2 As shown, in this embodiment, to address the problem of small sample scarcity, the present invention proposes a time-series decoupling strategy: the original long time-series sample containing "24-hour input - 15-hour output" is split into 15 independent single-time-point samples. Each single-time-point sample uses the cumulative observation data and numerical forecast data of the 24 hours before 15:00 on the current day as input features, and the visibility at the corresponding future time as the prediction target. For example, if the input is the actual observation and numerical forecast data of the 24 hours before 15:00 on the current day, such as when the forecast is carried out on the 14th, the input data is the ground observation and European numerical forecast data of 24 time periods from 16:00 on the 13th to 15:00 on the 14th, and the prediction target is the hourly ground visibility of 15 hours from 22:00 on the 14th to 12:00 on the 15th.
[0081] This strategy expanded the number of samples from the original 729 valid time-series samples to 10,968 single-time-point samples, which greatly alleviated the problem of insufficient model training caused by small samples, and provided a more sufficient sample foundation for subsequent data augmentation.
[0082] S20: Based on single-time-point samples, construct a multi-dimensional feature set that integrates physical mechanisms and temporal features to obtain forecast fusion features and compensate for the loss of time series information.
[0083] In one embodiment of the present invention, in order to compensate for the loss of temporal continuity information after temporal decoupling, and to incorporate the physical mechanism of fog formation, a 35-dimensional feature set containing three types of features (real-time observation data, numerical forecast data, and auxiliary information data) is constructed.
[0084] In this embodiment, obtaining forecast fusion features includes:
[0085] S21, based on the meteorological and physical principles of fog formation, constructs vertical structure and near-surface layer characteristics, including inversion intensity characteristics, wind shear characteristics, and humidity characteristics, as physical mechanism features.
[0086] In this embodiment, the inversion intensity characteristics are: the temperature difference between each pressure layer at 1000hPa, 950hPa, and 925hPa and the ground is calculated to capture the triggering effect of the inversion layer on the formation of dense fog.
[0087] Wind shear characteristics: Calculate the wind speed difference between each pressure layer (1000hPa, 950hPa, 925hPa) and the ground to reflect the influence of vertical wind field changes on fog dissipation.
[0088] Humidity characteristics: including relative humidity at 2 meters above the ground, relative humidity at 1000hPa, 850hPa, and 700hPa pressure layers, highlighting water vapor conditions as the core element for fog formation.
[0089] S22, the time features include the month, hour sine / cosine encoding, solar altitude angle, and time difference.
[0090] In this embodiment, time features are used to reconstruct time-aware capabilities using an explicit time coding strategy:
[0091] Periodic encoding: Hours and months are mapped to sine / cosine spaces (hour_sin, hour_cos, month_sin, month_cos) respectively, preserving the continuous periodicity of the time dimension.
[0092] Solar altitude angle: Calculates the solar altitude angle at the forecast time, accurately correlates it with the solar radiation intensity, and reflects the periodic physical processes of fog dissipation at sunrise and fog formation at night.
[0093] Time difference characteristics: Calculate the time difference between the forecast time and 15:00 on the reporting date to quantify the temporal correlation between the observation data and the forecast time.
[0094] S23 introduces the real-time observation data at 15:00 on the same day as the observation anchor point feature, and also incorporates the observation statistical features of the forecast day.
[0095] In this embodiment, real-time observation data at 15:00 on the same day (dominant visibility, relative humidity, etc.) is introduced as anchor point features to correct systematic biases in numerical forecasts. Simultaneously, observational statistical features such as the lowest and highest visibility on the forecast day are incorporated to enrich the dimensions of the model's input information.
[0096] S24 combines observation anchor point features, numerical forecast features, physical mechanism features, temporal features, and observational statistical features to obtain forecast fusion features.
[0097] In this embodiment, the original long-term time-series sample "past 24 hours input - future 15 hours output" is decoupled into 15 single-time-point samples. For the first... A single time point sample ( =1……15), definition:
[0098] Prediction target: Forecast date Visibility at any given moment.
[0099] Forecast fusion features: These are obtained by splicing together observation anchor point features, numerical forecast features, physical mechanism features, temporal features, and observational statistical features, and are denoted as:
[0100] ;
[0101] in, For the first Hourly forecast fusion characteristics; The characteristics of the observation anchor point at 15:00; Numerical forecast characteristics for the predicted time; The physical mechanism characteristics at the predicted time are calculated from the multi-layer elements of numerical forecasting at the corresponding prediction time; The time characteristics of the predicted time are calculated from the forecast time, including the month, hourly sine and cosine codes, solar altitude angle, and time difference (the number of hours from the forecast time of 15:00). These characteristics are decoupled one-to-one with the samples: each forecast time has its own time characteristics. Observational statistical features. Extracted from the "past 24-hour observation window" (e.g., the lowest / highest visibility of the day, mean, and range of change), these can also serve as shared statistical background features for the 15 single-time-point samples split from the original long-term time series sample, compensating for the loss of continuous temporal information after decoupling.
[0102] The observation anchor point features are shared across multiple single-time-point samples obtained by splitting the same original long-time series sample. That is, the observation anchor point features are identical for all times. This is used to provide the current true state to correct systematic biases in numerical forecasts and reduce drift caused by relying solely on numerical weather forecasts.
[0103] S30 uses a data augmentation strategy that combines stratified sampling, oversampling, and downsampling to balance the sample distribution of forecast fusion features.
[0104] In one embodiment of the present invention, to address the scarcity of low-visibility samples caused by the "long-tail distribution" of visibility data, a three-dimensional data augmentation strategy combining stratified sampling, oversampling, and downsampling is designed:
[0105] Stratified sampling: Fine-grained binning is performed according to visibility thresholds, with dominant visibility binning at intervals of 1000~2000 meters to ensure that samples in each interval are properly characterized during the sampling process.
[0106] Oversampling: Low-visibility, scarce samples are replicated and expanded, with visibility samples ≤1000 meters expanded by 3-5 times and visibility samples 1000-6000 meters expanded by 1-3 times. To avoid overfitting, only 29 continuous features such as temperature, humidity, and wind speed are injected with small Gaussian noise, with the noise amplitude controlled within 1% of the within-class standard deviation, simulating the natural disturbances in real meteorological data.
[0107] Downsampling: High visibility samples above 10,000 meters are randomly discarded, leaving 2,000 samples to balance the overall sample distribution.
[0108] After data augmentation, the total number of samples was further expanded to 15,212, and the proportion of samples in each visibility range tended to be balanced. Among them, the proportion of low visibility samples increased from 2.4% to 6.9%, which effectively improved the model's ability to learn from foggy scenes.
[0109] S40 uses the data-enhanced forecast fusion features as input to the forecast model to obtain the initial visibility for hourly forecasts.
[0110] In one embodiment of the present invention, the forecast model selected is the LightGBM (Light Gradient Boosting Machine) model.
[0111] To address the segmented scoring requirements of fog forecasting, a custom VIS scoring function is designed as both the model evaluation metric and the loss function. The custom scoring function is as follows:
[0112] When visibility is less than 1000 meters, it is a sensitive area of concern for safety management entities such as transportation and aviation. Therefore, absolute error is used as the core evaluation criterion to focus on whether the visibility forecast crosses the key threshold, which meets the high-precision forecasting needs of low visibility (below 1000m) scenarios.
[0113] When visibility is ≥1000 meters, relative error is used as the core evaluation criterion to balance the prediction accuracy across different visibility ranges. See Table 1 for details.
[0114] Table 1. Custom scoring function scoring rules
[0115]
[0116] In one embodiment of the present invention, the training strategy for the prediction model is optimized:
[0117] Parameter tuning: GridSearchCV combined with 5-fold cross-validation was used to traverse and search the core parameters. The optimal parameter combination was: num_leaves=63, max_depth=-1, learning_rate=0.02, feature_fraction=0.9, bagging_fraction=0.8.
[0118] GridSearchCV is a tool used for automatic parameter tuning in machine learning. GridSearch stands for network search, and CV stands for cross-validation. num_leaves is the number of leaf nodes, max_depth is the maximum depth of the decision tree, learning_rate is the learning rate, feature_fraction is the feature sampling rate, and bagging_fraction is the sample sampling rate.
[0119] Dataset partitioning: Stratified sampling was performed based on the visibility binning results, with 80% of the samples used as the training set (12,170 samples) and 20% used as the test set (3,042 samples) to ensure the consistency of the distribution between the training and test sets.
[0120] Early stopping mechanism: Setting stopping_rounds=100 means that if the prediction model shows no improvement in the validation set score after 100 consecutive iterations, training will immediately stop, effectively preventing overfitting in small sample sizes. Here, stopping_rounds represents the number of early stopping rounds.
[0121] In this embodiment, the prediction model reached its optimal state after 1375 training iterations, with an average score of 8.48 per time step on the validation set. The key performance indicators are as follows:
[0122] Root mean square error (RMSE): 1356.8184; Mean absolute error (MAE): 818.7533.
[0123] S50 performs time-series consistency repair on the initial visibility and obtains the repaired visibility.
[0124] In one embodiment of the present invention, after using "temporal decoupling" to independently predict multiple future steps (e.g., 15 hours) hour by hour, there is a lack of temporal constraints between the prediction moments. The output sequence is prone to jitter, discontinuity, and unreasonable jumps, thereby reducing the model's usability and physical rationality. Therefore, it is necessary to repair the temporal consistency. Without changing the single-point prediction model, the prediction results for the next 15 hours are jointly corrected to obtain a continuous and consistent output sequence.
[0125] In this embodiment, the visibility after timing consistency repair is obtained by solving the following optimization problem:
[0126] ;
[0127] In the formula, For the first Initial visibility forecast for the hour, For the first time sequence consistency repair Hourly forecast visibility For the first The confidence weight of the hourly forecast is used to control the degree to which the time-series consistency repair result follows the original forecast at the corresponding time. For the first The smoothing strength of hourly forecasts is used to measure and suppress the magnitude of changes in forecasts between adjacent hours. The minimum objective.
[0128] in, Based on historical verification error statistics, the following was obtained: In the formula, To verify the first set Variance of hourly lead time prediction error; For numerically stable terms, we take 0.00001.
[0129] In this embodiment, the first The smoothing strength of the hourly forecast is obtained using the following formula:
[0130] ;
[0131] In the formula, For the first Hourly predicted smoothing strength For the first The base value of smoothness intensity per hour, For the first The hourly fog risk index function For the first Hourly time adjustment factor.
[0132] No. Hourly smoothness intensity base value It can be obtained through the following formula:
[0133] ;
[0134] ;
[0135] No. Hourly fog risk index function It can be obtained through the following formula:
[0136] ;
[0137] No. Hourly time adjustment factor It can be obtained through the following formula:
[0138] ;
[0139] In the formula, For the first The fog potential index for hours is used to characterize the fog potential index at the time of the fog. The strength of the physical conditions that cause fog to form in an hour. This is an indicator function that returns 1 if the condition is true, and 0 otherwise. For the first Hourly solar altitude angle For the first hourly relative humidity, For the first 1000 hPa and ground temperature inversion per hour, For the first Hourly ground wind speed, This is the time difference between the forecast time and the initial reporting time.
[0140] By jointly optimizing the initial prediction result sequence, we can significantly reduce unreasonable hourly jumps while maintaining the ability to fit the original prediction, thereby improving the continuity, stability, and physical rationality of the output sequence.
[0141] Please see Figure 3 As shown, the forecast results after time series repair are verified:
[0142] The radiation fog case on January 3, 2021: The actual visibility dropped to a minimum of 900 meters, but the forecast model output a warning result of 493 meters in advance, accurately identifying the risk of heavy fog and providing sufficient preparation time for flight scheduling.
[0143] The fog-free case on January 4, 2021: The visibility the previous day dropped to a minimum of 900 meters, and the daytime visibility remained around 2000 meters. However, the forecast model still predicted no heavy fog, and the visibility trend was basically consistent with the actual situation.
[0144] The critical fog case on January 24, 2021: The actual visibility was 1200 meters (close to the fog standard), and the forecast model predicted a value of 600 meters. This was a reasonable defensive forecast and complied with the aviation business's safety principle of "better to be empty than to miss".
[0145] Fog-free case on March 11, 2021: The visibility the previous day dropped to a minimum of 50 meters, then quickly recovered. The forecast model gave a trend forecast of decreasing visibility but no dense fog.
[0146] Critical fog case on March 11, 2021: The actual visibility was good, and the forecast model predicted that the visibility would remain between 1000 and 2000, which was basically consistent with the actual situation.
[0147] The fog-free case on November 14, 2022: Low visibility was observed, but the event was short-lived. As a result, the forecast model predicted lower visibility than the actual visibility, but it was still well above the standard for dense fog.
[0148] Case of persistent low visibility on January 3, 2024: Affected by the low visibility weather of the previous day, the forecast model tends to predict lower visibility. Although there is a certain bias, it reflects the effective use of historical observation information by the forecast model.
[0149] The fog-free case on February 4, 2024: The forecast accurately predicted that no heavy fog would occur, and the visibility trend was basically consistent with the actual situation, indicating a good forecast effect.
[0150] Overall, the forecast model has a strong ability to capture the formation of fog at night and its dissipation at sunrise. After time series repair, the forecast results show very accurate prediction of the time evolution trend.
[0151] Please see Figure 4 As shown, the present invention also provides a system for applying the above-described fog time-series forecasting method for small sample scenarios, comprising:
[0152] The time series decoupling module expands the long-term time series samples based on real-time observation and numerical forecast by splitting them into multiple independent single-time-point samples.
[0153] The feature engineering module constructs a multi-dimensional feature set that integrates physical mechanisms and temporal features based on single-time-point samples, obtains forecast fusion features, and makes up for the loss of time series information.
[0154] The data augmentation module uses a data augmentation strategy that combines stratified sampling, oversampling, and downsampling to balance the sample distribution of forecast fusion features.
[0155] The forecast module uses the data-enhanced forecast fusion features as input to the forecast model to obtain the initial visibility for hourly forecasts.
[0156] The timing repair module performs timing consistency repair on the initial visibility and obtains the repaired visibility.
[0157] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.
[0158] The above embodiments are merely examples of implementation methods of the invention. The scope of protection of the present invention is not limited to the above embodiments. For those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention.
Claims
1. A method for time-series fog forecasting for small sample scenarios, characterized in that, include: The long-term time series sample, which includes real-time observations and numerical forecasts, is expanded by splitting it into multiple independent single-time-point samples based on time-series decoupling. This includes splitting the original long-term time series sample containing "24-hour input - 15-hour output" into 15 independent single-time-point samples. Each single-time-point sample uses the cumulative observation data and numerical forecast data of the 24 hours before 15:00 on the current day as input features. Based on single-time-point samples, a multi-dimensional feature set integrating physical mechanisms and temporal features is constructed to obtain forecast fusion features, compensating for the loss of time-series information. These forecast fusion features include: Based on the meteorological and physical principles of fog formation, the physical mechanism is constructed by identifying vertical structure and near-surface features, including inversion intensity features, wind shear features, and humidity features, as physical mechanism characteristics. The time features include the month, hourly sine / cosine encoding, solar altitude angle, and time difference; Real-time observation data at 15:00 on the same day was used as the observation anchor point feature, and the observation statistical features of the forecast day were also included. By splicing together observation anchor point features, numerical forecast features, physical mechanism features, temporal features, and observational statistical features, forecast fusion features are obtained. The data augmentation strategy, which combines stratified sampling, oversampling, and downsampling, is used to balance the sample distribution of forecast fusion features; The data-enhanced forecast fusion features are used as input to the forecast model to obtain the initial visibility for hourly forecasts; Perform time-series consistency repair on the initial visibility and obtain the repaired visibility.
2. The fog time-series forecasting method for small sample scenarios according to claim 1, characterized in that, The characteristics of the observation anchor points are shared among multiple single-time-point samples obtained by splitting the same original long-time-series sample.
3. The fog time-series forecasting method for small sample scenarios according to claim 1, characterized in that, A data augmentation strategy combining stratified sampling, oversampling, and downsampling is used to balance the sample distribution of forecast fusion features, including: Stratified sampling: fine-grained binning is performed according to visibility threshold, with dominant visibility binning at intervals of 1000~2000 meters; Oversampling: Low visibility scarce samples are replicated and expanded, with visibility samples ≤1000 meters expanded by 3 to 5 times and visibility samples 1000 to 6000 meters expanded by 1 to 3 times; and noise is injected into the replicated and expanded samples to simulate the natural disturbance of real meteorological data. Downsampling: High visibility samples above 10,000 meters are randomly discarded to balance the overall sample distribution.
4. The fog time-series forecasting method for small sample scenarios according to claim 1, characterized in that, A custom scoring function is used to optimize the training process of the forecast model; the custom scoring function is as follows: When visibility is less than 1000 meters, the absolute error is used as the core evaluation criterion. Attention is paid to whether the visibility forecast crosses the key threshold, which meets the high-precision forecasting needs of low visibility scenarios. When visibility is ≥1000 meters, relative error is used as the core evaluation criterion to balance the prediction accuracy of different visibility ranges.
5. The fog time-series forecasting method for small sample scenarios according to claim 1, characterized in that, Obtaining visibility after the repair includes: obtaining visibility after timing consistency repair by solving the following optimization problem: ; In the formula, For the first Initial visibility forecast for the hour, For the first time sequence consistency repair Hourly forecast visibility For the first The confidence weight of the hourly forecast is used to control the degree to which the time-series consistency repair result follows the original forecast at the corresponding time. For the first The smoothing strength of hourly forecasts is used to measure and suppress the magnitude of changes in forecasts between adjacent hours. The minimum objective.
6. The fog time-series forecasting method for small sample scenarios according to claim 5, characterized in that, No. The smoothing strength of the hourly forecast is obtained using the following formula: ; In the formula, For the first Hourly predicted smoothing strength For the first The baseline value of smoothness intensity per hour, For the first The fog risk index function for hours For the first Hourly time adjustment factor.
7. The fog time-series forecasting method for small sample scenarios according to claim 6, characterized in that, No. Hourly smoothness intensity base value It can be obtained through the following formula: ; ; No. Hourly fog risk index function It can be obtained through the following formula: ; No. Hourly time adjustment factor It can be obtained through the following formula: ; In the formula, For the first The fog potential index for hours is used to characterize the fog potential index at the time of... The strength of the physical conditions that cause fog to form in an hour. This is an indicator function that returns 1 if the condition is true, and 0 otherwise. For the first h Hourly solar altitude angle For the first hourly relative humidity, For the first 1000 hPa and ground temperature inversion per hour, For the first Hourly ground wind speed, This is the time difference between the forecast time and the initial reporting time.
8. A system for applying the fog time-series forecasting method for small sample scenarios according to any one of claims 1-7, characterized in that, include: The time series decoupling module expands the long-term time series samples based on real-time observation and numerical forecast by splitting them into multiple independent single-time-point samples. The feature engineering module constructs a multi-dimensional feature set that integrates physical mechanisms and temporal features based on single-time-point samples, obtains forecast fusion features, and makes up for the loss of time series information; The data augmentation module uses a data augmentation strategy that combines hierarchical sampling, oversampling, and downsampling to balance the sample distribution of forecast fusion features; The forecast module uses the data-enhanced forecast fusion features as input to the forecast model to obtain the initial visibility for hourly forecasts; The timing repair module performs timing consistency repair on the initial visibility and obtains the repaired visibility.