Visibility prediction method based on machine learning
Through the visibility prediction method based on machine learning, and the visibility prediction model is constructed and optimized using historical meteorological data, the problem of inaccurate visibility prediction in traditional methods is solved, and more accurate and reliable visibility prediction is achieved.
Patent Information
- Application Number
- CN202510298621.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The traditional visibility detection method has a large workload and the prediction results are inaccurate when the atmosphere is uneven, making it difficult to effectively predict the impact of precipitation on visibility.
Using machine learning-based visibility prediction method, a visibility prediction model is constructed by acquiring and preprocessing historical meteorological data, a time series feature and meteorological data feature set is extracted, and a visibility prediction model is optimized using loss function and gradient descent to train a visibility prediction model.
Accurate prediction of visibility is achieved, the model prediction results are consistent with the actual observation data, and can effectively capture the nonlinear relationship between precipitation and visibility, improving the accuracy and reliability of visibility prediction.
Smart Images

Figure CN120217045A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visibility prediction, and particularly to a visibility prediction method based on machine learning. Background Art
[0002] Visibility is defined as the maximum distance at which a person with normal vision can identify an object from its background. Visibility has wide applications in fields such as meteorology, aviation, and navigation, and is closely related to the lives of residents. Good visibility is conducive to the smooth travel of residents and the improvement of production efficiency, while poor visibility will reduce the visible distance, increase potential safety hazards in urban transportation, and pose a serious threat to the personal and property safety of residents.
[0003] Precipitation has a certain impact on visibility. The traditional detection method is to remotely measure the distribution of particulate matter in the atmosphere through lidar technology and analyze the echo signal to estimate visibility. However, this method has a large workload, and when the atmosphere is uneven, the atmospheric scattering coefficient will be affected, resulting in inaccurate prediction results. Summary of the Invention
[0004] In view of the above-mentioned prior art, the present invention aims to provide a visibility prediction method based on machine learning, mainly to solve the technical problems existing in the above-mentioned background art.
[0005] To achieve the above object, the technical solution of the embodiment of the present invention is realized as follows:
[0006] A visibility prediction method based on machine learning includes the following steps:
[0007] Obtain historical meteorological data, perform preprocessing on the historical meteorological data, and then perform feature selection to obtain time series features and a meteorological data feature set;
[0008] Construct a visibility prediction model, input the time series features and the meteorological data feature set into the visibility prediction model for training, and use a loss function and gradient descent for optimization to obtain a trained visibility prediction model;
[0009] Perform visibility prediction based on the trained visibility prediction model to obtain a prediction result.
[0010] Optionally, the preprocessing includes outlier detection processing, missing value filling, and normalization processing of the historical meteorological data;
[0011] The historical meteorological data includes precipitation, relative humidity, temperature, wind speed, and visibility data.
[0012] Optionally, the feature selection includes using a rolling window to extract precipitation, visibility, relative humidity, air temperature, and wind speed data at different time intervals, analyzing the relationships between the precipitation, visibility, relative humidity, air temperature, and wind speed data using the Pearson correlation coefficient, determining the optimal time interval, and obtaining the time series features and the meteorological data feature set.
[0013] Optionally, the step of using a rolling window to extract precipitation, visibility, relative humidity, air temperature, and wind speed data at different time intervals is specifically as follows:
[0014] Define the window size as 1 hour, use the rolling window to move at different time intervals to filter out precipitation data with a precipitation of 15 mm or more in 1 hour, and within each selected window, extract the cumulative precipitation, visibility, relative humidity, air temperature, and wind speed data at different time intervals;
[0015] The different time intervals include: 1 minute, 5 minutes, 10 minutes, 15 minutes, 20 minutes, and 30 minutes.
[0016] Optionally, the visibility prediction model is a non-linear power function model, and the specific expression of the model is:
[0017] V = a × R b
[0018] where V is the visibility, R is the cumulative precipitation, and a and b are the parameters of the model.
[0019] Optionally, the optimization using the loss function and gradient descent includes:
[0020] Iteratively train the visibility prediction model by means of the loss function and gradient descent, and use the least squares method for parameter estimation during the iterative training process until the loss function reaches the optimum;
[0021] The loss function is the mean square error function;
[0022] The specific expression of the least squares method is:
[0023] V = a × R b + ∈
[0024] where V is the observed value, R is the independent variable, a and b are the model parameters, and ∈ is the error term.
[0025] Optionally, the parameters of the visibility prediction model are a = 3952.193 and b = -0.517.
[0026] Optionally, it further includes: evaluating the reliability of the visibility prediction model by calculating the coefficient of determination, and the coefficient of determination r 2 = 0.58;
[0027] The expression of the coefficient of determination is as follows:
[0028]
[0029] Among them, SS res is the sum of squared residuals, that is, the sum of the squares of the differences between the actual observed values and the model predicted values; SS tot is the total sum of squares, that is, the sum of the squares of the differences between the actual observed values and the average of the actual observed values;
[0030] The steps for obtaining the coefficient of determination are as follows:
[0031] (1) Calculate the average value, and calculate the average value of the minimum visibility within each of the said time intervals;
[0032] (2) Calculate the total sum of squares:
[0033]
[0034] Among them, V i is the i-th observed value, and n is the total number of observed values, is the average value of the minimum visibility within each of the said time intervals;
[0035] (3) Calculate the predicted value, and use the model to calculate the predicted visibility V* corresponding to each r;
[0036] (4) Calculate the sum of squared residuals SS res :
[0037]
[0038] Among them, V * is the i-th value predicted by the model;
[0039] (5) Calculate the loss function MSE:
[0040]
[0041] (6) Apply Stochastic Gradient Descent (SGD) to update the model parameters:
[0042]
[0043] Among them, a t and b t respectively represent the values of parameters a and b at the t-th iteration, η is the learning rate, and f i is the loss function for the i-th data point;
[0044] (7) Parameter update, update parameters a and b through the least squares method to minimize the loss function;
[0045] (8) Iteration. Repeat steps (2)-(6) until the loss function converges or a predetermined number of iterations is reached;
[0046] (9) Calculate r 2 , substitute SSres and SStot into the above r 2 formula to calculate r 2 = 0.58.
[0047] The beneficial effects of the present invention are as follows: By obtaining historical meteorological data, preprocessing the historical meteorological data, and performing feature selection to obtain time series features and a meteorological data feature set, and feeding the time series features and the meteorological data feature set into a visibility prediction model for training to learn the relationship between visibility and other meteorological factors, and optimizing through a loss function and gradient descent to obtain a trained visibility prediction model; finally, using the trained visibility prediction model to predict actual data, it is found that the prediction results of the model are in good agreement with the actual observed data, indicating that the model can effectively predict visibility. Description of the Drawings
[0048] Figure 1 is a schematic diagram of a visibility prediction method based on machine learning in an embodiment of the present invention;
[0049] Figure 2 is a fitting diagram of a visibility prediction method based on machine learning in an embodiment of the present invention. Detailed Embodiments
[0050] The technical solution of the present invention will be further elaborated in detail below in conjunction with the drawings in the specification and specific embodiments. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. In the following description, the expression "some embodiments" is described, which describes a subset of all possible embodiments. However, it should be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0051] In the following description, a large number of specific details are given to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some technical features well known to those skilled in the art are not described.
[0052] It should be understood that the present invention can be implemented in different forms and should not be construed as limited to the embodiments presented herein. On the contrary, providing these embodiments will make the disclosure thorough and complete, and will fully convey the scope of the present invention to those skilled in the art. And the purpose of the terms used herein is only to describe specific embodiments and is not a limitation of the present invention. When used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the terms "comprising" and / or "including", when used in this specification, determine the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups. When used herein, the term "and / or" includes any and all combinations of the related listed items.
[0053] It should be further noted that when an element is referred to as "fixed to" another element, it can be directly on the other element or there can also be a middle element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be a middle element at the same time. The terms "vertical", "horizontal", "inner", "outer", "left", "right" and similar expressions used herein are only for the purpose of illustration and do not represent the only implementation.
[0054] To thoroughly understand the present invention, detailed structures will be presented in the following description to illustrate the technical solutions proposed by the present invention. The optional embodiments of the present invention are described in detail below. However, in addition to these detailed descriptions, the present invention can also have other implementations.
[0055] Embodiment
[0056] Please refer to the attached Figure 1 , this application provides a visibility prediction method based on machine learning, including the following steps:
[0057] Obtain historical meteorological data, perform preprocessing on the historical meteorological data and then perform feature selection to obtain time series features and a meteorological data feature set; wherein, the historical meteorological data includes precipitation, relative humidity, temperature, wind speed, and visibility data; the preprocessing includes performing outlier detection processing, missing value filling, and normalization processing on the historical meteorological data;
[0058] Specifically, historical meteorological data can be obtained from meteorological stations. Specifically, the minute-level precipitation, relative humidity, temperature, wind speed, and visibility data of Haikou National Automatic Station from 2021 to 2022 are obtained. Through historical meteorological data in different dimensions, the influence of different factors on visibility can be analyzed. The obtained historical meteorological data is preprocessed, including using statistical methods to identify and process outliers to reduce the interference of noise. Then, missing values are filled using the neighbor interpolation method. Finally, data standardization is performed, and the Z-score standardization method is used to standardize the data to ensure that all features are on the same dimension. Then, feature selection is carried out to obtain the time series features from 2021 to 2022 and their corresponding meteorological data feature sets. The meteorological data feature sets include precipitation feature vectors, relative humidity feature vectors, temperature feature vectors, wind speed feature vectors, and visibility feature vectors.
[0059] Construct a visibility prediction model. Input the time series features and the meteorological data feature sets into the visibility prediction model for training, and use the loss function and gradient descent for optimization to obtain a trained visibility prediction model.
[0060] Specifically, input the obtained time series features and meteorological data feature sets into the visibility prediction model. The model can learn the relationship between visibility and other meteorological factors. Use the mean squared error as the loss function, and optimize the parameters of the model through the non-linear least squares method and stochastic gradient descent. Specifically, calculate the loss value of the model using the mean squared error, find the model parameters through the non-linear least squares method to minimize the sum of the squared errors between the model prediction values and the actual observed values, and use stochastic gradient descent to update the model parameters. In each training iteration, calculate the loss, perform backpropagation, and update the model parameters to obtain a trained visibility prediction model.
[0061] Based on the trained visibility prediction model, perform visibility prediction to obtain the prediction result.
[0062] As an optional implementation manner, the feature selection includes: using a rolling window to extract precipitation, visibility, relative humidity, temperature, and wind speed data at different time intervals from the standardized precipitation, relative humidity, temperature, wind speed, and visibility data, analyzing the relationship between precipitation, visibility, relative humidity, temperature, and wind speed using the Pearson correlation coefficient, determining the optimal time interval, and obtaining the time series features and the meteorological data feature sets.
[0063] Specifically, define the window size as 1 hour. Use a rolling window to filter out 1-hour precipitation greater than or equal to 15 mm, ensuring the continuity of events within each window in terms of time. Adopt the rolling window method. Within each selected 1-hour window, move the time intervals (1 minute, 5 minutes, 10 minutes, 15 minutes, 20 minutes, 30 minutes) in sequence to filter out heavy precipitation events that meet the conditions. By using rolling windows with different time intervals within each 1-hour window, the precipitation at different time scales can be analyzed. Then, within each selected window, extract the cumulative precipitation as well as meteorological elements such as visibility, relative humidity, air temperature, and wind speed at different time intervals (1 minute, 5 minutes, 10 minutes, 15 minutes, 20 minutes, 30 minutes). Use the Pearson correlation coefficient to analyze the relationships between different meteorological elements. By analyzing the relationships between the characteristics of each meteorological element at different time intervals and visibility, determine which time interval characteristics are the most important for model prediction. For each time interval, calculate statistical characteristics, such as the mean, maximum, minimum, and standard deviation, and analyze the relationship between precipitation and visibility at different time intervals. It is found that the characteristics of the 10-minute time interval contribute the most to the visibility prediction of the visibility prediction model. Therefore, finally, a time series feature and meteorological data feature set with a 10-minute interval is determined.
[0064] Exemplarily, within a 1-hour window, divide the time into multiple 10-minute intervals. For example: from 0 minutes to 10 minutes is one interval, from 10 minutes to 20 minutes is the second interval, and so on until 60 minutes. Then, within each 10-minute interval, calculate the cumulative precipitation as well as the relative humidity, air temperature, and wind speed during this time period. Then, arrange the data characteristics obtained from each 10-minute interval in chronological order to form a series of data points. For example: considering 1-hour data, there will be 6 data points. In this embodiment, the meteorological data from 2021 to 2022 is used. Therefore, the meteorological data from 2021 to 2022 is taken at 10-minute intervals to obtain a time series feature and meteorological data feature set.
[0065] As an optional implementation manner, the visibility prediction model is a non-linear power function model, and the specific expression of the model is:
[0066] V = a × R b
[0067] where V is visibility, R is cumulative precipitation, and a and b are parameters of the model;
[0068] Optimize using the loss function and gradient descent, including: iteratively train the visibility prediction model by means of the loss function and gradient descent, and use the least squares method for parameter estimation during the iterative training process until the loss function reaches the optimal value;
[0069] The loss function is the Mean Squared Error (MSE), and its formula is:
[0070]
[0071] where V i is the true value of the i-th observation, V* is the predicted value, and n is the total number of observations; this function measures the average of the squares of the differences between the predicted value and the true value.
[0072] Update the model parameters using Stochastic Gradient Descent (SGD):
[0073]
[0074] where a t and b t represent the values of parameters a and b at the t-th iteration respectively, η is the learning rate, and f i is the loss function for the i-th data point. This formula means that in each iteration, we calculate the gradients t and b t under the current parameters a and and then update the parameters in the opposite direction of the gradient to reduce the value of the loss function. In this way, the model parameters are adjusted in the direction of reducing the loss in each iteration.
[0075] The specific expression using the least squares method is:
[0076] V = a × R b + ∈
[0077] where V is the observation, R is the independent variable, a and b are the model parameters, and ∈ is the error term;
[0078] Specifically: Using the mean squared error as the loss function can measure the difference between the visibility prediction value of the visibility prediction model and the actual observation value, and minimize the loss function through gradient descent. In each iteration, calculate the gradient of the loss function with respect to the visibility prediction model parameters and update the parameters in the opposite direction of the gradient to gradually reduce the value of the loss function. Then, perform parameter estimation using the least squares method to find a set of parameters a and b that minimize the mean squared error between the visibility prediction model prediction value and the actual observation value until the loss function reaches the optimum;
[0079] The parameters of the visibility prediction model are a = 3952.193 and b = -0.517;
[0080] As an optional implementation, it further includes: evaluating the reliability of the visibility prediction model by calculating the coefficient of determination, where the coefficient of determination r 2 = 0.58;
[0081] Formula for the coefficient of determination:
[0082]
[0083] where SSres is the sum of squared residuals, i.e., the sum of the squares of the differences between the actual observed values and the model predicted values; SStot is the total sum of squares, i.e., the sum of the squares of the differences between the actual observed values and the average of the actual observed values;
[0084] The specific steps for obtaining the coefficient of determination are as follows:
[0085] (1) Calculate the average: Calculate the average of the dependent variable, the 10-minute minimum visibility V
[0086] (2) Calculate the total sum of squares SStot:
[0087]
[0088] where V i is the i-th observed value and n is the total number of observed values;
[0089] (3) Calculate the predicted values: Use the model to calculate the predicted visibility V* corresponding to each r;
[0090] (4) Calculate the sum of squared residuals SSres:
[0091]
[0092] where V * is the i-th value predicted by the model.
[0093] (5) Calculate the loss function MSE:
[0094]
[0095] (6) Apply Stochastic Gradient Descent (SGD) to update the model parameters:
[0096]
[0097] where, a t and b t respectively represent the values of parameters a and b at the t-th iteration, η is the learning rate, and f i is the loss function for the i-th data point;
[0098] (7) Update the parameters, update parameters a and b by the least squares method to minimize the loss function;
[0099] (8) Iterate, repeat steps (2)-(6) until the loss function converges or reaches a predetermined number of iterations;
[0100] (9) Calculate r 2 : Substitute SSres and SStot into the above r 2 formula for calculation, r 2 = 0.58;
[0101] It should be noted that the coefficient of determination r 2 is an important statistic for measuring model fitting. It represents the fitting degree of the model to the data set. The value range of the coefficient of determination is between 0 and 1. The closer the value is to 1, the better the fitting degree of the model to the data set;
[0102] Please refer to the attached Figure 2 , in the present invention, by determining the parameters a = 3952.193, b = -0.517, the coefficient of determination R 2 = 0.58, indicating that the visibility prediction model of the present invention has a medium degree of explanatory ability; please refer to Figure 2 , the prediction results of the model are in good agreement with the actual observed data, indicating that the model can effectively capture the non-linear relationship between precipitation and visibility, and the model can effectively predict visibility, thereby providing a basis for traffic safety management and meteorological warning.
[0103] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the said claims.
Claims
1. A visibility prediction method based on machine learning, characterized in that: The following steps are involved: Obtain historical meteorological data, perform feature selection after preprocessing the historical meteorological data to obtain time series features and meteorological data feature sets; Constructing a visibility prediction model, inputting the time series features and the meteorological data feature set into the visibility prediction model for training, and optimizing using a loss function and gradient descent to obtain a trained visibility prediction model; Visibility prediction is performed based on the trained visibility prediction model to obtain a prediction result.
2. The visibility prediction method based on machine learning according to claim 1, characterized in that: The preprocessing includes abnormal value detection, missing value filling and standardization of historical meteorological data; The historical meteorological data include precipitation, relative humidity, temperature, wind speed, and visibility data.
3. The visibility prediction method based on machine learning according to claim 1, characterized in that: The feature selection includes using a rolling window to extract precipitation, visibility, relative humidity, temperature, and wind speed data at different time intervals, using the Pearson correlation coefficient to analyze the relationship between precipitation, visibility, relative humidity, temperature, and wind speed data, determining the optimal time interval, and obtaining time series features and a meteorological data feature set.
4. The visibility prediction method based on machine learning according to claim 1, characterized in that: The rolling window is used to extract precipitation, visibility, relative humidity, temperature, and wind speed data at different time intervals, specifically: Define the window size as 1 hour, use the rolling window to move different time intervals to filter out precipitation data greater than or equal to 15 mm in 1 hour, and extract the cumulative precipitation, visibility, relative humidity, temperature, and wind speed data at different time intervals in each selected window; The different time intervals include: 1 minute, 5 minutes, 10 minutes, 15 minutes, 20 minutes, and 30 minutes.
5. The visibility prediction method based on machine learning according to claim 1, characterized in that: The visibility prediction model is a nonlinear power function model, and the specific expression of the model is: V=a×R b Among them, V is visibility, R is accumulated precipitation, and a and b are parameters of the model.
6. The visibility prediction method based on machine learning according to claim 1, characterized in that: The optimization using loss function and gradient descent includes: Iteratively training the visibility prediction model by means of a loss function and gradient descent, and performing parameter estimation by using a least squares method during the iterative training process until the loss function reaches an optimum; The loss function is a mean square error function; The specific expression of the least squares method is: V=a×R b +∈ Among them, V is the observation, R is the independent variable, a and b are model parameters, and ∈ is the error term.
7. The visibility prediction method based on machine learning according to claim 6, characterized in that: The parameters of the visibility prediction model are a=3952.193, b=-0.
517.
8. The visibility prediction method based on machine learning according to claim 7, characterized in that: Also includes: The reliability of the visibility prediction model was evaluated by calculating the coefficient of determination. 2 =0.58; The expression of the determination coefficient is: Among them, SS res is the residual sum of squares, that is, the sum of squares of the differences between the actual observed values and the model predicted values; SS tot is the total sum of squares, that is, the sum of the squares of the differences between the actual observations and the mean of the actual observations; The steps for obtaining the coefficient of determination are: (1) Calculate the average value, that is, the average value of the minimum visibility within each of the time intervals; (2) Calculate the total sum of squares: Among them, V i is the ith observation, n is the total number of observations, is the average value of the minimum visibility during each of the said time intervals; (3) Calculate the predicted value and use the model to calculate the predicted visibility V* corresponding to each r; (4) Calculate the residual sum of squares SS res : Among them, V * is the i-th value predicted by the model; (5) Calculate the loss function MSE: (6) Apply stochastic gradient descent (SGD) to update model parameters: Among them, a t and b t They represent the values of parameters a and b at the tth iteration, η is the learning rate, and f i is the loss function for the i-th data point; (7) parameter updating, updating parameters a and b by the least squares method to minimize the loss function; (8) Iteration: repeat steps (2) to (6) until the loss function converges or reaches a predetermined number of iterations; (9) Calculate r 2 , substitute SSres and SStot into the above r 2 The formula is used to calculate r 2 =0.58.
Citation Information
Cited By
Visibility prediction method
CN121071410A
Multi-model fusion visibility forecasting method
CN121117970A