Sea fog visibility grade forecasting method based on visibility meter and mesoscale mode
By combining visibility meters and mesoscale models, a sea fog visibility level forecasting model was constructed, which solved the problems of insufficient accuracy and stability in existing sea fog forecasts, and achieved high-precision medium- and long-term forecasts, applicable to coastal transportation and port areas.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AEROSPACE NEWSKY TECHNOLOGY CO LTD
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-12
AI Technical Summary
Existing sea fog forecasting methods rely on numerical models, statistical models, and machine learning. However, due to insufficient and complex data, forecast accuracy and stability are difficult to guarantee. In particular, given the complex formation mechanism of sea fog and its strong spatiotemporal variability, it is difficult to achieve high-precision medium- and long-term forecasts.
By combining visibility meter data and mesoscale models, a visibility forecasting model is constructed through feature extraction, machine learning, and model fusion. The optimal feature variables are selected by fusing visibility meter observation data with data from mesoscale numerical models, and a stacked classifier model is constructed to achieve high-precision sea fog visibility level forecasting.
It achieves seamless high-precision forecasting in the medium and long term, improves forecast accuracy and stability, adapts to different seasons and forecast duration requirements, and is applicable to coastal highways, airports and ports, providing high-precision early warning support.
Smart Images

Figure CN122020555A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visibility forecasting technology, and in particular to a method for forecasting sea fog visibility levels based on a visibility meter and a mesoscale model. Background Technology
[0002] Sea fog refers to a weather phenomenon in the ocean-atmosphere boundary layer where the horizontal visibility is reduced to less than 1 km due to the presence of numerous water droplets or ice crystals. It is a major marine meteorological disaster. Sea fog not only affects maritime navigation and port operations, but also disrupts or even closes coastal highway transportation. Statistics show that 60%–70% of maritime collision accidents are related to sea fog, causing economic losses comparable to those caused by typhoons. Therefore, the impact of low-visibility sea fog events on coastal production safety and economic development is significant, making early prediction extremely urgent. Current sea fog forecasting mainly relies on the following three methods and their combinations:
[0003] (1) Numerical model diagnostic method: Based on physical quantities such as atmospheric liquid water content output by numerical models, the affected area of sea fog is qualitatively judged. The challenge of this method lies in the complexity of sea fog forecasting itself: its formation mechanism is affected by both local climate and weather background, and has strong spatiotemporal variability. At the same time, the parameterization scheme of numerical models for key processes such as turbulent activity and cloud microphysics in the atmospheric boundary layer has significant uncertainties. Coupled with the errors in input data such as initial field and sea surface temperature, sea fog forecasting has become a recognized problem in the field of numerical weather prediction.
[0004] (2) Statistical model method: Qualitative forecasting is carried out by screening effective forecasting factors and establishing a statistical model. This method relies heavily on massive observation data to calibrate and optimize forecasting thresholds. At present, there is a severe lack of atmospheric element observation information above the lower sea surface, and the performance of statistical models is easily affected by local characteristics and seasonal changes, thus limiting their generalization ability.
[0005] (3) Machine learning method: This method uses historical observation data to train a machine learning model to achieve nowcasting of visibility. This method also relies heavily on high-quality, long-term observation data. Its key limitation is that if the input precursor factors fail to fully capture the changing trend of low visibility, the model forecast results will have a huge bias, and the forecast stability will be challenged. Summary of the Invention
[0006] To address the aforementioned problems and technical requirements, this application proposes a sea fog visibility level forecasting method based on a visibility meter and a mesoscale model. The technical solution of this application is as follows: A method for forecasting sea fog visibility levels based on a visibility meter and a mesoscale model, the method comprising: The raw observation data acquired by the visibility meter in historical scenes are preprocessed to obtain visibility level observation data with the same time step as the output data of the mesoscale numerical model. Feature extraction was performed on the historical model dataset of meteorological elements at multiple altitudes in the local area where the visibility meter is located, output by the mesoscale numerical model in historical scenarios. This yielded model forecast data of various meteorological elements and their derivative elements related to visibility at each time point in the local area where the visibility meter is located. The model forecast data of multiple feature variables that are most strongly associated with visibility were then selected. After obtaining a feature sample set with model forecast data of each feature variable as input and visibility level observation data at the same time stamp as sample labels, a visibility forecast model is trained using the feature sample set based on a machine learning algorithm. The model forecast data of each characteristic variable is extracted from the multi-height meteorological elements of the area to be forecasted in the real-time operation scenario of the mesoscale numerical model. The data is then input into the visibility forecast model trained to obtain the sea fog visibility level forecast results for the area to be forecasted.
[0007] The further technical solution involves using derived elements, including model diagnostic atmospheric visibility and other derived elements, to obtain model forecast data of various meteorological elements and their derived elements related to visibility in the local area where the visibility meter is located. For each of the meteorological elements related to visibility and other derived elements other than model diagnostic atmospheric visibility, extract the element value at multiple different typical altitudes in the local area where the visibility meter is located, as well as the average value of the element at different grid points in the local area where the visibility meter is located. The atmospheric extinction coefficient was calculated based on condensate information from historical model datasets of meteorological elements at multiple altitudes within the local area where the visibility meter was located. And obtain model diagnostic atmospheric visibility .
[0008] A further technical solution involves calculating the atmospheric extinction coefficient based on condensate information from historical model datasets of meteorological elements at multiple altitudes within the local area where the visibility meter is located. Including calculations according to the following formula:
[0009] in, Indicates cloud water particle concentration. Indicates cloud ice particle concentration. Indicates the concentration of rainwater particles. Indicates the concentration of snow water particles and:
[0010] in, Indicates the mixing ratio of clouds and water. This indicates the mixing ratio of cloud ice. Indicates the mixing ratio of rainwater. Indicates the mixing ratio of snow and water; Indicates the density of moist air and , Indicates atmospheric pressure. This represents the specific gas constant of dry air. Indicates the water vapor mixing ratio. It represents the absolute temperature of the air.
[0011] The further technical solution involves extracting visibility-related meteorological elements, including temperature, relative humidity, dew point temperature, temperature-dew point difference, sea level pressure, wind direction, and wind speed. Other derived elements besides those for model diagnostic atmospheric visibility include relative eddyivity, divergence, temperature advection, and vertical temperature gradient. The solution also includes extracting the element values for each element at multiple typical altitudes within the local area where the visibility meter is located. For any meteorological element among temperature, relative humidity, dew point temperature, temperature-dew point difference, sea level pressure, wind direction, and wind speed, extract the element value of the meteorological element at the typical ground altitude in the low-altitude region of the local area where the visibility meter is located, and extract the element value of the meteorological element at multiple different typical isobaric surface altitudes in the high-altitude region of the local area where the visibility meter is located. For any of the derived elements among relative vorticity, divergence, temperature advection, and vertical temperature gradient, extract the element values of the derived element at multiple different typical isobaric surface heights in the upper-air region of the local area where the visibility meter is located.
[0012] A further technical solution involves selecting model forecast data that most strongly correlates with visibility, including: The model forecast data and visibility level observation data of candidate features are aligned in time to obtain an initial sample set with model forecast data of multiple candidate features as input and visibility level observation data of the same timestamp as sample labels. The candidate features include a variety of meteorological elements and their derivative elements that are extracted and related to visibility. Samples with visibility level observation data below the visibility level threshold in the initial sample set are classified as low visibility subsets, and samples with visibility level observation data reaching the visibility level threshold are classified as high visibility subsets. The low visibility subsets are oversampled, and the high visibility subsets are undersampled. The low-visibility subset and high-visibility subset, which have been sampled separately, are merged to obtain an extended sample set. Based on the extended sample set, model forecast data of multiple feature variables that are most strongly correlated with visibility are selected to obtain a feature sample set.
[0013] Its further technical solution is to select model forecast data based on an expanded sample set that are most strongly associated with visibility, including for each candidate feature; Using the model forecast data of the candidate features as the horizontal axis and the corresponding visibility level observation data as the vertical axis, a scatter plot of each sample in the extended sample set with respect to the current candidate features is constructed. Under the constraint that the total number of grids does not exceed B(n), the scatter plot is divided into grids according to different grid division methods, and the mutual information values of the model forecast data of candidate features and the visibility level observation data are calculated for each grid division method, where B(n) is a function of the number of samples n in the extended sample set. When the maximum value of mutual information under different grid partitioning methods reaches the contribution threshold, the current candidate feature is selected as the feature variable.
[0014] A further technical solution involves obtaining visibility level observation data with a time step consistent with the output data of the mesoscale numerical model, including: After Butterworth low-pass filtering is applied to the raw observation data acquired by the visibility meter in historical scenarios, the minimum value of the filtered raw observation data within each time step of the mesoscale numerical model output data is extracted and mapped to the corresponding visibility level observation data according to the visibility impact level classification standard.
[0015] A further technical solution involves performing Butterworth low-pass filtering on the raw observation data acquired by the visibility meter in historical scenarios, including: A bilinearly varied discretized Butterworth low-pass filter is used to perform forward filtering on the raw observation data acquired by the visibility meter in historical scenes to obtain a forward intermediate sequence. After reversing the forward intermediate sequence, a discretized Butterworth low-pass filter with bilinear transformation is used for backward filtering to obtain the backward intermediate sequence; the backward intermediate sequence is then reversed to obtain the zero-phase filtered output.
[0016] Its further technical solution is to use a feature sample set to train a visibility forecast model based on a machine learning algorithm, including: Multiple base learner models were constructed based on various machine learning algorithms, and a surrogate model was constructed using the Bayesian method. The hyperparameters of each base learner model were trained using a feature sample set. After obtaining multiple base learner models through pre-training, the meta-model is initialized as a logistic regression model. A stacked classifier is built using the output of each base learner model as the input of the meta-model. The stacked classifier with the best prediction accuracy is trained using the feature sample set to obtain the visibility prediction model.
[0017] A further technical solution is that when using the trained visibility forecast model to forecast the visibility level of sea fog, the effective forecast duration of the visibility forecast model dynamically changes with the time length of the model forecast data of the input feature variables.
[0018] The beneficial technical effects of this application are: This application discloses a sea fog visibility level forecasting method based on a visibility meter and a mesoscale model. This method utilizes machine learning to fuse high-frequency, high-resolution, single-point measurement data from a visibility meter with continuous forecast data from a mesoscale numerical model, constructing a high-precision, medium- to long-term seamless visibility forecasting model capable of predicting sea fog visibility levels. This method not only achieves medium- to long-term seamless forecasting but also significantly improves the accuracy and effectiveness of forecasting and early warning by combining the "precision" of visibility meter observations with the "foresight" of mesoscale numerical models. It can be widely applied to coastal highways, airports, ports, and other fields, providing high-precision and refined forecasting and early warning support for addressing low visibility disasters caused by sea fog. Furthermore, in its model architecture design, this application does not preset a fixed upper limit for forecast duration; its effective forecast time range (i.e., forecast lead time) depends entirely on the duration of the input mesoscale numerical forecast model data, thus possessing strong flexibility and adaptability to meet application scenarios with different forecast duration requirements.
[0019] This application utilizes the maximum mutual information coefficient algorithm to quantify the nonlinear relationship between visibility and various atmospheric features, enabling the autonomous selection and filtering of a subset of optimal feature variables suitable for model training. Furthermore, in model selection, it integrates multiple machine learning algorithms to construct a base learner model for visibility forecasting, employing a stacking strategy for intelligent ensemble integration. This allows for effective training using small sample data, combining the advantages of each model to compensate for the prediction bias of a single model. Compared to WRF models, statistical diagnostics, and single machine learning models, this model demonstrates superior performance in accuracy, stability, and generalization ability, significantly improving forecast accuracy. Simultaneously, by combining feature variable selection strategies and high-frequency iterative training, the weight parameters of each model can be dynamically adjusted. This allows for application in small sample scenarios and enhances the model's adaptability to low visibility forecasts in different seasons, achieving better forecast results.
[0020] This method also supports multi-visibility meter station network forecasting or gridded forecasting by introducing observation data from nearby visibility meter stations. It has good scalability and adaptability and can be widely used in areas affected by low visibility, such as ports, airports, and transportation. Attached Figure Description
[0021] Figure 1This is a flowchart illustrating a sea fog visibility level forecasting method in one embodiment of this application.
[0022] Figure 2 This is a flowchart illustrating a sea fog visibility level forecasting method in another embodiment of this application.
[0023] Figure 3 This is a comparison chart of the sea fog visibility level forecast and the actual result in an example of this application. Detailed Implementation
[0024] The specific embodiments of this application will be further described below with reference to the accompanying drawings.
[0025] This application discloses a method for forecasting sea fog visibility levels based on a visibility meter and a mesoscale model. Please refer to [reference needed]. Figure 1 The flowchart shown illustrates the following steps in this sea fog visibility level forecasting method: Step 110: Preprocess the raw observation data acquired by the visibility meter in historical scenes to obtain visibility level observation data with the same time step as the output data of the mesoscale numerical model.
[0026] Visibility meters operate based on the principle of optical forward scattering, enabling them to acquire time series of high-frequency raw observation data. The data acquisition frequency is typically on the order of minutes, and the measurement range is 10-80,000 meters. However, they are essentially point-based observations, and the raw observation data acquired is only the visibility information of a single point at the location of the visibility meter.
[0027] This step involves preprocessing the raw observation data, including quality control and downsampling. Quality control aims to improve the data quality of the raw observation data, while downsampling facilitates subsequent spatiotemporal alignment with the output data of the low-frequency mesoscale numerical model (The Weather Research and Forecasting Model, WRF). In one embodiment: First, the raw visibility data acquired by the visibility meter in historical scenarios is processed using a Butterworth low-pass filter to remove high-frequency disturbances. Specifically: Let the original sampling frequency of the raw observation data output by the visibility meter be . The frequency of output data from the mesoscale numerical model is The cutoff frequency of the Butterworth low-pass filter used is then... Normalized cutoff angular frequency . Choose a value slightly greater than 1 to establish a transition band, for example, a typical value is .
[0028] In another embodiment, a bilinearly varying discretized Butterworth low-pass filter is used to low-pass filter the original observation data. To eliminate phase shift, forward-backward filtering is employed. Specifically: first, a bilinearly varying discretized Butterworth low-pass filter is used to forward-filter the original observation data acquired by the visibility meter in historical scenes, obtaining a forward intermediate sequence. Then, the forward intermediate sequence is inverted, and a bilinearly varying discretized Butterworth low-pass filter is used for backward filtering, obtaining a backward intermediate sequence. Finally, the backward intermediate sequence is inverted again to obtain a zero-phase filtered output.
[0029] Then, according to the time step of the mesoscale numerical model output data, the filtered raw observation data within each time step is extracted, which is the minimum value of the zero-phase filter output obtained above. This minimum value is then mapped to the corresponding visibility level observation data according to the visibility impact level classification standard. For mapping, the visibility impact level classification in the "Shipiao Meteorological Conditions Level" published by the China Meteorological Administration can be referenced as shown in the table below:
[0030] A typical approach involves a visibility meter collecting data minute by minute and a mesoscale numerical model outputting data every 10 minutes. The zero-phase filter output is then downsampled by 1 / 10 to extract the minimum visibility within each 10 points and mapped to the corresponding visibility level observation data. The resulting time series of visibility level observation data effectively characterizes the lowest visibility level occurring in each corresponding time period, and its time interval is consistent with the step size of the output data from the mesoscale numerical weather prediction model.
[0031] Step 120: Extract features from the historical model dataset of meteorological elements at multiple altitudes in the local area where the visibility meter is located, output by the mesoscale numerical model in historical scenarios, to obtain model forecast data of various meteorological elements and their derivative elements related to visibility at each time point in the local area where the visibility meter is located.
[0032] Based on atmospheric physics, this step first extracts model forecast data of various meteorological elements and their derivative elements related to visibility according to experience, including: (1) Extract model forecast data of various meteorological elements and their derived elements related to visibility within the local area where the visibility meter is located. The derived elements are elements calculated by converting meteorological elements directly output from the mesoscale numerical model. In one embodiment, the derived elements calculated from the meteorological elements include model diagnostic atmospheric visibility and other derived elements. The feature extraction operation includes: (1) For meteorological elements related to visibility and other derived elements other than model diagnostic atmospheric visibility, the model forecast data for each element are extracted in the following two aspects: (a) Extract the feature values for each feature at multiple typical heights within the local area where the visibility meter is located. The meaning of the typical heights differs for different features. Specifically: In one embodiment, the meteorological elements related to visibility include temperature, relative humidity, dew point temperature, temperature-dew point difference, sea level pressure, wind direction, and wind speed. For any one of these meteorological elements, the element value of that element at a typical ground altitude in the low-altitude region within the local area where the visibility meter is located is extracted, as well as the element value of that element at multiple different typical isobaric surface altitudes in the upper-altitude region within the local area where the visibility meter is located. In one embodiment, the typical ground altitude of the low-altitude region corresponding to temperature, relative humidity, dew point temperature, temperature-dew point difference, and sea level pressure is 2 meters above the ground, while the typical ground altitude of the low-altitude region corresponding to wind direction and wind speed is 10 meters above the ground. The multiple different typical isobaric surface altitudes corresponding to each meteorological element in the upper-altitude region include element values at isobaric surface altitudes of 1000 hPa, 925 hPa, 850 hPa, 700 hPa, and 500 hPa.
[0033] In one embodiment, other derived elements calculated from the above meteorological elements, besides model diagnostic atmospheric visibility, include: relative vorticity. divergence Temperature advection Vertical temperature gradient For any of the aforementioned derived elements, extract the element value of that element at multiple different typical isobaric surface heights in the upper-air region of the local area where the visibility meter is located. Since the derived elements are calculated from meteorological elements, in practical applications, it is usually necessary to extract the element values of meteorological elements at multiple different typical isobaric surface heights, and then calculate the element value of the derived element at the corresponding typical isobaric surface height according to the corresponding calculation formula. The specific calculation formula for each of the above derived elements can be found in existing classic formulas, and will not be elaborated here.
[0034] (b) Extract the average value of each feature at different grid points in the local area where the visibility meter is located.
[0035] In this step, the local area where the visibility meter is located can be customized, and different local area ranges can be used when extracting feature values at typical ground altitudes in the low-altitude region and when extracting feature values at typical ground altitudes in the high-altitude region. For example, in one instance, feature values of various meteorological elements at typical ground altitudes in the low-altitude region within a local area centered on the latitude and longitude of the visibility meter and with a radius R of 2° can be extracted. Feature values of various meteorological elements at typical ground altitudes in the high-altitude region within a local area centered on the latitude and longitude of the visibility meter and with a radius R of 5° can be extracted.
[0036] (2) For model diagnostic atmospheric visibility calculated from meteorological elements, the methods for extracting the model diagnostic atmospheric visibility include: calculating the atmospheric extinction coefficient based on the condensate information in the historical model dataset of meteorological elements at multiple altitudes in the local area where the visibility meter is located. Then, the model was used to diagnose atmospheric visibility. The atmospheric extinction coefficient is among them. The calculation formula is:
[0037] in, Indicates cloud water particle concentration. Indicates cloud ice particle concentration. Indicates the concentration of rainwater particles. Indicates the concentration of snow water particles and:
[0038] in, Indicates the mixing ratio of clouds and water. This indicates the mixing ratio of cloud ice. Indicates the mixing ratio of rainwater. These represent the mixing ratios of snow and water, and the units for all four mixing ratios are g / kg. Indicates the density of moist air and The unit is kg / m 3 . This represents atmospheric pressure, measured in Pa. The specific gas constant for dry air is typically taken as 287.05 J / kg. K. This indicates the water vapor mixing ratio, expressed in g / kg. This indicates the absolute temperature of the air, measured in Kelvin (K).
[0039] Step 130: Select model forecast data that are most strongly associated with multiple feature variables of visibility.
[0040] Considering that visibility is a phenomenon resulting from the spatiotemporal variations of multiple atmospheric variables, involving largely nonlinear relationships, and that the inducing factors often differ across scenarios (e.g., visibility varies across seasons), this step, based on the meteorological and derived elements extracted in step 120, further filters to form an optimal subset of feature variables, including: (1) The various meteorological elements and their derivative elements related to visibility obtained in step 120 are used as candidate features. The model forecast data of the candidate features and the visibility level observation data are aligned in time to obtain an initial sample set with the model forecast data of multiple candidate features as input and the visibility level observation data of the same timestamp as sample labels.
[0041] In another embodiment, the initial sample set includes samples extracted from raw observation data based on a single visibility meter, or samples extracted from raw observation data based on multiple visibility meters within a predetermined range.
[0042] (2) The initial sample set is sampled to obtain an extended sample set, thereby expanding the sample. Considering the uneven distribution of visibility data, a composite sampling strategy is adopted in one embodiment to optimize the data distribution of the extended sample set, including: Samples with visibility level observation data below the visibility level threshold in the initial sample set are classified as low visibility subsets, and samples with visibility level observation data reaching the visibility level threshold are classified as high visibility subsets. The visibility level threshold can be set by the user.
[0043] The characteristics of sea fog data result in a sparse sample size in the low-visibility subset, while the high-visibility subset contains the majority of samples. Oversampling is performed on the sparse low-visibility subset to enhance the information of the low-visibility target data. Undersampling is then performed on the majority high-visibility subset to achieve sparse high-visibility data information.
[0044] Finally, the low-visibility subset and high-visibility subset, which have been sampled separately, are merged to obtain an expanded sample set with a relatively balanced class distribution.
[0045] (3) Based on the expanded sample set, the model forecast data of multiple feature variables that are most strongly correlated with visibility are selected to obtain the feature sample set. In one embodiment, the maximum mutual information coefficient (MIC) calculation method for nonlinear relationships is introduced to select the elements with the best correlation as feature variables, including for each candidate feature: Using the model forecast data of the candidate feature as the horizontal axis and the corresponding visibility level observation data as the vertical axis, a scatter plot of each sample in the extended sample set with respect to the current candidate feature is constructed. Under the constraint that the total number of grids does not exceed B(n), the scatter plot is divided into different grid partitioning methods, and the mutual information value between the model forecast data of the candidate feature and the visibility level observation data is calculated for each grid partitioning method. The calculated mutual information value for each grid partitioning method ranges from [0,1], and can quantify the nonlinear relationship between the model forecast data of the current candidate feature and the visibility level observation data: the closer the calculated mutual information value is to 1, the stronger the deterministic correlation between the model forecast data of the current candidate feature and the visibility level observation data; the closer the calculated mutual information value is to 0, the more statistically independent the model forecast data of the current candidate feature and the visibility level observation data are.
[0046] When the maximum mutual information value under different grid partitioning methods reaches a preset contribution threshold, the current candidate feature is selected as a feature variable; otherwise, it is not selected. Here, B(n) is a function of the number of samples n in the expanded sample set, which is usually taken as n. 0.6 .
[0047] For each candidate feature, the maximum mutual information value between it and the visibility level observation data Y is calculated using the method described above. This allows for the selection of feature variables. The model forecast data of each feature variable for each sample in the extended sample set are retained, while the model forecast data of other candidate features are deleted. This results in a feature sample set with the model forecast data of each feature variable as input and the visibility level observation data at the same timestamp as the sample labels.
[0048] Step 140: After obtaining a feature sample set with the model forecast data of each feature variable as input and the visibility level observation data of the same timestamp as sample labels, the visibility forecast model is trained using the feature sample set based on a machine learning algorithm.
[0049] To enable effective training using small sample data and to leverage the strengths of various models to compensate for the prediction bias of a single model, in one embodiment, multiple base learner models are first constructed based on different machine learning algorithms, and a surrogate model is constructed using Bayesian methods. The hyperparameters of each base learner model are then used for training on a feature sample set. In one embodiment, the base learner models constructed using multiple machine learning algorithms include: decision tree, K-nearest neighbor algorithm, support vector machine, random forest, AdaBoost, GBRT, and lightGBM models. The training process for each base learner model is as follows: (1) Initialize the base learner model by selecting the minimum or initial value of each parameter in the key hyperparameter matrix of the model as the initial parameter configuration of the base learner model.
[0050] (2) Establishing a proxy model based on Gaussian processes It is used to fit and evaluate the performance of hyperparameters and base learner models. The hyperparameters of the base learner model.
[0051] (3) Select the next set of hyperparameters to be evaluated by maximizing the acquisition function.
[0052] (4) Use the forecast accuracy to evaluate the performance of the current hyperparameters.
[0053] (5) Update the proxy model, add the results of the new evaluation to the dataset, and refit the Gaussian process.
[0054] (6) Repeat (3) to (5) until the preset number of iterations or convergence condition is reached.
[0055] (7) Return the optimal combination of hyperparameters and output the optimal base learner model, and save it in joblib format.
[0056] After obtaining multiple base learner models through pre-training, a stacking strategy is employed to achieve intelligent ensemble. The output of each base learner model is used as input to a meta-model to build a StackingClassifier. The meta-model is initialized as a logistic regression model, used for subsequent fusion and decision-making regarding the outputs of the base learner models. The visibility forecasting model is obtained by training the stacked classifier with the optimal forecast accuracy using a feature sample set.
[0057] The stacked classifier employs a five-fold cross-validation method. It uses a base learner model trained on the predicted outputs and corresponding sample labels of samples in the feature sample set. The outputs of the base learner models are then used as input features to train the final meta-model. The stacked classifier is trained using the feature sample set to obtain the final prediction model.
[0058] This step integrates multiple machine learning algorithms to build a base learner model and employs a stacking strategy to achieve intelligent ensemble training. It leverages small sample data for effective training and combines the strengths of each model to compensate for the prediction bias of a single model. Compared to WRF mode, statistical diagnostics, and single machine learning models, this model demonstrates superior performance in accuracy, stability, and generalization ability, significantly improving forecast accuracy.
[0059] Step 150: After training the visibility forecast model, it can be used to forecast the visibility level of sea fog. Specifically, this involves extracting model forecast data for each characteristic variable from the multi-height meteorological elements of the area to be forecasted, as output by the mesoscale numerical model in real-time operation. The extraction method is the same as steps 120 and 130 in the training phase, and will not be repeated here. Then, the extracted model forecast data for each characteristic variable is input into the trained visibility forecast model to obtain the sea fog visibility level forecast result for the area to be forecasted.
[0060] The visibility forecast model trained in this application does not have a fixed upper limit for forecast duration. The effective forecast duration depends entirely on the duration of the input mesoscale numerical forecast model data. In other words, when using the trained visibility forecast model to forecast the visibility level of sea fog, the effective forecast duration of the visibility forecast model will dynamically change with the duration of the model forecast data of the input feature variables, thus possessing strong flexibility and adaptability, and can meet application scenarios with different forecast duration requirements.
[0061] Please also refer to... Figure 2 As the observation data from the visibility instrument and the model data from the mesoscale numerical model accumulate, the visibility forecast model will be iteratively updated according to the method of this application. The forecast accuracy will be used as the evaluation index. When the forecast accuracy of the new visibility forecast model exceeds that of the old visibility forecast model, the new visibility forecast model will be activated; otherwise, the old visibility forecast model will continue to be used.
[0062] In one example, the visibility level of sea fog predicted using the visibility forecasting model trained in this application is as follows: Figure 3 The red dashed line in the image shows the actual results of the sea fog visibility level. Figure 3 As shown by the solid green line in the figure, actual measurements show that the visibility forecasting model trained using this application achieves a forecasting accuracy of 87% in this example, exceeding the performance of existing sea fog forecasts.
[0063] The above descriptions are merely preferred embodiments of this application, and this application is not limited to the above embodiments. It is understood that other improvements and variations that can be directly derived or conceived by those skilled in the art without departing from the spirit and concept of this application should be considered to be included within the protection scope of this application.
Claims
1. A method for forecasting sea fog visibility levels based on a visibility meter and a mesoscale model, characterized in that, The sea fog visibility level forecasting method includes: The raw observation data acquired by the visibility meter in historical scenes are preprocessed to obtain visibility level observation data with the same time step as the output data of the mesoscale numerical model. Feature extraction was performed on the historical model dataset of meteorological elements at multiple altitudes in the local area where the visibility meter is located, output by the mesoscale numerical model in historical scenarios. This yielded model forecast data of various meteorological elements and their derivative elements related to visibility at each time point in the local area where the visibility meter is located. The model forecast data of multiple feature variables that are most strongly associated with visibility were then selected. After obtaining a feature sample set with model forecast data of each feature variable as input and visibility level observation data at the same time stamp as sample labels, a visibility forecast model is trained using the feature sample set based on a machine learning algorithm. The model forecast data of each characteristic variable is extracted from the multi-height meteorological elements of the area to be forecasted in the real-time operation scenario of the mesoscale numerical model. The data is then input into the visibility forecast model trained to obtain the sea fog visibility level forecast results for the area to be forecasted.
2. The sea fog visibility level forecasting method according to claim 1, characterized in that, Derivative elements include model diagnostic atmospheric visibility and other derived elements. Model forecast data for various meteorological elements and their derivative elements related to visibility in the local area where the visibility meter is located include: For each of the meteorological elements related to visibility and other derived elements other than model diagnostic atmospheric visibility, extract the element value of the element at multiple different typical altitudes in the local area where the visibility meter is located, as well as the average value of the element value of the element at different grid points in the local area where the visibility meter is located. The atmospheric extinction coefficient was calculated based on condensate information from historical model datasets of meteorological elements at multiple altitudes within the local area where the visibility meter was located. And obtain model diagnostic atmospheric visibility .
3. The sea fog visibility level forecasting method according to claim 2, characterized in that, The atmospheric extinction coefficient was calculated based on condensate information from historical model datasets of meteorological elements at multiple altitudes within the local area where the visibility meter was located. Including calculations according to the following formula: in, Indicates cloud water particle concentration. Indicates cloud ice particle concentration. Indicates the concentration of rainwater particles. Indicates the concentration of snow water particles and: in, Indicates the mixing ratio of clouds and water. This indicates the mixing ratio of cloud ice. Indicates the mixing ratio of rainwater. Indicates the mixing ratio of snow and water; Indicates the density of moist air and , Indicates atmospheric pressure. This represents the specific gas constant of dry air. Indicates the water vapor mixing ratio. It represents the absolute temperature of the air.
4. The sea fog visibility level forecasting method according to claim 2, characterized in that, Meteorological elements related to visibility include temperature, relative humidity, dew point temperature, temperature-dew point difference, sea level pressure, wind direction, and wind speed. Other derived elements besides model diagnostic atmospheric visibility include relative eddy current, divergence, temperature advection, and vertical temperature gradient. The values of each element at multiple typical altitudes within the local area where the visibility meter is located are extracted, including: For any meteorological element among temperature, relative humidity, dew point temperature, temperature-dew point difference, sea level pressure, wind direction, and wind speed, extract the element value of the meteorological element at the typical ground altitude in the low-altitude region of the local area where the visibility meter is located, and extract the element value of the meteorological element at multiple different typical isobaric surface altitudes in the high-altitude region of the local area where the visibility meter is located. For any one of the derived elements among relative vorticity, divergence, temperature advection, and vertical temperature gradient, extract the element values of the derived element at multiple different typical isobaric surface heights in the upper-air region of the local area where the visibility meter is located.
5. The sea fog visibility level forecasting method according to claim 1, characterized in that, The model forecast data that filters out the multiple feature variables most strongly correlated with visibility also includes: The model forecast data and visibility level observation data of candidate features are aligned in time to obtain an initial sample set with model forecast data of multiple candidate features as input and visibility level observation data of the same timestamp as sample labels. The candidate features include a variety of meteorological elements and their derivative elements that are extracted and related to visibility. Samples with visibility level observation data below the visibility level threshold in the initial sample set are classified as low visibility subsets, and samples with visibility level observation data reaching the visibility level threshold are classified as high visibility subsets. The low visibility subsets are oversampled, and the high visibility subsets are undersampled. The low-visibility subset and high-visibility subset, which have been sampled separately, are merged to obtain an extended sample set. Based on the extended sample set, model forecast data of multiple feature variables that are most strongly correlated with visibility are selected to obtain a feature sample set.
6. The sea fog visibility level forecasting method according to claim 5, characterized in that, Model forecast data based on an expanded sample set, which selects multiple feature variables that are most strongly associated with visibility, includes data for each candidate feature. Using the pattern forecast data of the candidate features as the horizontal axis and the corresponding visibility level observation data as the vertical axis, a scatter plot of each sample in the extended sample set with respect to the current candidate features is constructed. Under the constraint that the total number of grids does not exceed B(n), the scatter plot is divided into grids according to different grid division methods, and the mutual information values of the model prediction data and visibility level observation data of the candidate features are calculated for each grid division method, where B(n) is a function of the number of samples n in the extended sample set. When the maximum value of mutual information under different grid partitioning methods reaches the contribution threshold, the current candidate feature is selected as the feature variable.
7. The sea fog visibility level forecasting method according to claim 1, characterized in that, Visibility level observations with time steps consistent with the output data of the mesoscale numerical model include: After Butterworth low-pass filtering is applied to the raw observation data acquired by the visibility meter in historical scenarios, the minimum value of the filtered raw observation data within each time step of the mesoscale numerical model output data is extracted and mapped to the corresponding visibility level observation data according to the visibility impact level classification standard.
8. The sea fog visibility level forecasting method according to claim 7, characterized in that, The Butterworth low-pass filtering process for raw visibility data acquired by the visibility meter in historical scenes includes: A bilinearly varied discretized Butterworth low-pass filter is used to perform forward filtering on the raw observation data acquired by the visibility meter in historical scenes to obtain a forward intermediate sequence. After reversing the forward intermediate sequence, a discretized Butterworth low-pass filter with bilinear transformation is used for backward filtering to obtain the backward intermediate sequence; the backward intermediate sequence is then reversed to obtain the zero-phase filtered output.
9. The sea fog visibility level forecasting method according to claim 1, characterized in that, The process of training a visibility forecast model using the feature sample set based on a machine learning algorithm includes: Multiple base learner models were constructed based on various machine learning algorithms, and a surrogate model was constructed using the Bayesian method. The hyperparameters of each base learner model were trained using a feature sample set. After obtaining multiple base learner models through pre-training, the meta-model is initialized as a logistic regression model. A stacked classifier is built using the output of each base learner model as the input of the meta-model. The stacked classifier with the best prediction accuracy is trained using the feature sample set to obtain the visibility prediction model.
10. The sea fog visibility level forecasting method according to claim 1, characterized in that, When using the visibility forecast model obtained through training to forecast the visibility level of sea fog, the effective forecast duration of the visibility forecast model dynamically changes with the time length of the model forecast data of the input feature variables.