Visibility grade high-precision inversion method based on satellite data and TabNet model

Through the visibility level inversion method based on satellite data and TabNet model, the uncertainty and parameter dependence problems of visibility level inversion in the prior art are solved, and higher accuracy and interpretability are achieved.

CN120146104APending Publication Date: 2025-06-13NANJING UNIV OF INFORMATION SCI & TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510379774.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the inversion of visibility levels, the prior art has problems such as strong data accuracy and parameter dependence and great uncertainty in the face of complex scenarios and regional differences, especially in high humidity environments.

Method used

The high-precision inversion method of visibility level based on satellite data and TabNet model is adopted to obtain satellite data and ground measured visibility data to perform spatiotemporal matching and quality control, and the performance of the TabNet model is improved through hyperparameter optimization.

Benefits of technology

The accuracy of visibility level inversion and the interpretability of the model are improved, and the inversion effect under complex scenarios and regional differences are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146104A_ABST
    Figure CN120146104A_ABST
Patent Text Reader

Abstract

The invention discloses a visibility grade high-precision inversion method based on satellite data and a TabNet model, and relates to the field of meteorology, and the method comprises the steps: S1, obtaining satellite data and ground actual measurement visibility data of a target region in a to-be-measured time period; s2, performing space-time matching on the satellite data and the ground actually-measured visibility data; s3, performing quality control on the satellite data and the ground actual measurement visibility data, and performing grade division on the visibility data; s4, dividing all data into a training set and a test set, and inputting the training set and the test set into a TabNet visibility grade inversion model for training; and S5, predicting the visibility grade of the target area by using the trained visibility grade inversion model, and outputting an inversion result. According to the high-precision inversion method adopting the steps, model hyper-parameter optimization is carried out on the basis of a common TabNet model, so that the model fitting effect is better, and the development of visibility grade inversion is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of meteorology, and in particular to a method for high-precision inversion of visibility levels based on satellite data and the TabNet model. Background Art

[0002] With the increasingly severe global environmental problems, the research on atmospheric visibility has gradually attracted wide attention. Visibility, as a direct manifestation of atmospheric transparency, is comprehensively affected by factors such as atmospheric aerosols, particulate matter, and relative humidity. In areas with relatively serious air pollution, the reduction of visibility not only reflects the deterioration of air quality, but also has an important impact on traffic safety, climate research, and air pollution control. Especially in the transportation field, low visibility will cause the driver's line of sight to be blocked, thus significantly increasing the probability of traffic accidents, which poses higher requirements for the safety management of transportation modes such as highways, airport operations, and maritime shipping. Therefore, studying the variation law and spatial distribution of visibility has important practical significance for optimizing environmental governance and improving traffic safety.

[0003] The visibility level is an important indicator for evaluating visibility, and is usually used to quantitatively describe the range of ground visibility and its impact on traffic, aviation, and the environment. It divides visibility values into different levels, providing an intuitive and convenient way to measure and express the quality of atmospheric transparency or visibility range. At present, the inversion of visibility levels is mainly completed through two methods - namely, physical methods and statistical methods (including machine learning).

[0004] In physical methods, mainly based on meteorological parameters such as aerosol optical depth (AOD) and relative humidity, a visibility prediction model is established through multiple linear regression or empirical formulas. Such methods are simple to calculate, but have strong dependence on data accuracy and parameters, and show great uncertainty when facing complex scenarios and regional differences. For example, the aerosol hygroscopic effect and non-uniform distribution in a high-humidity environment often lead to model inaccuracy.

[0005] Compared with physical methods, statistical inversion methods (machine learning) are a highly competitive alternative. Machine learning can make full use of a large amount of existing high-quality observational data, and it can fit any complex non-linear function to ensure its applicability after training. However, in the face of a high-dimensional and non-linear task such as visibility level inversion, machine learning has strong dependence on feature engineering and is difficult to capture complex non-linear relationships. At the same time, the capacity of machine learning models is limited and it is difficult to process high-dimensional data. Summary of the Invention

[0006] The purpose of the present invention is to provide a high-precision inversion method for visibility levels based on satellite data and the TabNet model. In view of the characteristics of visibility, the model hyperparameters are optimized on the basis of the ordinary TabNet model, so that the fitting effect of the model is better, thereby promoting the development of visibility level inversion.

[0007] To achieve the above object, the present invention provides a high-precision inversion method for visibility levels based on satellite data and the TabNet model, and the steps are as follows:

[0008] S1. Obtain multi-spectral remote sensing satellite data of a geostationary meteorological satellite in the target area during the period to be measured and ground-measured visibility data of a ground measurement station;

[0009] S2. Perform spatio-temporal matching on the satellite data and the ground-measured visibility data. The matching rule is to select the satellite data with a difference in longitude and latitude from the ground measurement station within 0.04° at the same time and calculate the average value, which represents the satellite data of all spectral bands at this moment;

[0010] S3. Perform quality control on the satellite data and the ground-measured visibility data, and classify the visibility data into levels;

[0011] S4. Divide all the data into a training set and a test set, input them into the TabNet visibility level inversion model for training, and improve the model performance through hyperparameter optimization;

[0012] S5. Use the trained visibility level inversion model to predict the visibility level of the target area and output the inversion result.

[0013] Preferably, in S1, the multi-spectral remote sensing satellite data of the geostationary meteorological satellite includes 14 spectral bands, the spatial resolution is 4KM, and the downloaded time period is the whole hour of each day, with a total of 24 satellite data per day.

[0014] Preferably, in S1, the measured visibility data of the ground measurement station comes from the hourly ground-measured visibility data of the ground measurement stations of the China Meteorological Administration during the period to be measured.

[0015] Preferably, in S3, the process of quality control includes:

[0016] Eliminate the data in which all channels of the satellite data are all 0;

[0017] Eliminate the missing data existing in any channel of the satellite data;

[0018] Eliminate the measured visibility data of the ground measurement stations with an altitude higher than 500m;

[0019] Eliminate the missing data existing in the ground-measured visibility data;

[0020] Remove the data with outliers from the ground-measured visibility data.

[0021] Preferably, in S3, the standard for classifying visibility data is: classify the visibility data into 4 levels according to (0, 200m], (200m, 500m], (500m, 1000m], (1000m, 10000m].

[0022] Preferably, in S4, the ratio of the training set to the test set is 8:2. The input features of the TabNet visibility level inversion model are 18 in total, and the output is the visibility level.

[0023] Preferably, in S4, Optuna is used for automatic hyperparameter tuning. The tuning range includes: setting the initial range of the learning rate of the AdamW optimizer to [0.00001, 0.01], the width n_a of the decision prediction layer to [64, 256], the attention embedding width n_d of each mask to [64, 256], the number of decision steps n_steps in the structure to [3, 10], the number of independent Gated Attention layers n_independent to [1, 5], the sparse regularization coefficient lambda_sparse to [0.000001, 0.01], the weight decay coefficient weight_decay to [0.000001, 0.001], the batch size batch_size to [256, 1024], and setting the early stopping rounds to 50 during the training process.

[0024] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects: Compared with the TabNet model before hyperparameter optimization, the inversion reliability has been optimized, making it more accurate and converging faster. At the same time, compared with other deep learning models, the TabNet model has interpretability, can output the global importance and local importance of features, and improves the inversion effect of visibility levels.

[0025] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0027] Figure 1Schematic flow chart of an embodiment of a high-precision visibility level inversion method based on satellite data and TabNet model of the present invention;

[0028] Figure 2 Prediction error graph of the TabNet visibility level inversion model of the embodiment of the present invention;

[0029] Figure 3 Station accuracy graph of the embodiment of the present invention;

[0030] Figure 4 Thermal map of day-night difference accuracy of the embodiment of the present invention.

[0031] Figure 5 Line graph comparing seasonal accuracies of the embodiment of the present invention.

[0032] Figure 6 Comparison graph of true and predicted visibility levels from December 28, 2023 to December 30, 2023 of the embodiment of the present invention. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0035] Embodiment

[0036] As Figure 1 shown, taking the visibility level inversion in Anhui region from January 1, 2021 to December 31, 2023 as an example, a high-precision visibility level inversion method based on satellite data and TabNet model of the present invention is adopted, and the steps are as follows:

[0037] S1. Download and obtain multi-spectral remote sensing satellite data of geostationary meteorological satellites in Anhui region from January 1, 2021 to December 31, 2023 and ground measured visibility data of ground measurement stations.

[0038] The satellite data used is the full-disk L1 data of Fengyun-4A geostationary meteorological satellite, which comes from the Advanced Geostationary Radiation Imager on Fengyun-4A. It includes 14 spectral bands with a spatial resolution of 4 km. The downloaded time period is the whole hour of each day, that is, the whole-hour data from 00:00 to 23:00, with a total of 24 satellite data per day.

[0039] The measured visibility data of the ground measurement stations comes from the hourly ground-measured visibility data of 1378 ground measurement stations in Anhui region of the China Meteorological Administration from 2021 to 2023.

[0040] S2. Perform spatio-temporal matching on the satellite data and the ground-measured visibility data. The matching rule is to select the satellite data with longitude and latitude differences within 0.04° from those of the ground measurement stations and calculate the average value, representing the satellite data of all spectral bands at this moment.

[0041] S3. Conduct quality control on the satellite data and the ground-measured visibility data, and classify the visibility data into grades. The process of quality control includes:

[0042] Eliminate the data where all channels in the satellite data are 0.

[0043] Eliminate the missing data in any channel of the satellite data.

[0044] Eliminate the measured visibility data of the ground measurement stations with an altitude higher than 500 m.

[0045] Eliminate the missing data in the ground-measured visibility data.

[0046] Eliminate the data with outliers in the ground-measured visibility data.

[0047] Then, classify the visibility data into 4 grades according to the interval standards of (0, 200 m], (200 m, 500 m], (500 m, 1000 m], and (1000 m, 10000 m].

[0048] S4. Divide all the data into a training set and a test set according to the ratio of 8:2, input them into the TabNet visibility grade inversion model for training, and improve the model performance through hyperparameter optimization. The input features of the TabNet visibility grade inversion model are 18 in total, and the output is the visibility grade. The input features include satellite multi-spectral band data, site longitude and latitude, altitude, and satellite channel data time.

[0049] In S4, Optuna is used for automatic hyperparameter tuning. The tuning range includes: setting the initial range of the learning rate of the AdamW optimizer to [0.00001, 0.01], the width n_a of the decision prediction layer to [64, 256], the attention embedding width n_d of each mask to [64, 256], the number of decision steps n_steps in the structure to [3, 10], the number of independent GatedAttention layers n_independent to [1, 5], the sparse regularization coefficient lambda_sparse to [0.000001, 0.01], the weight decay coefficient weight_decay to [0.000001, 0.001], and the batch size batch_size to [256, 1024]. During training, the early stopping rounds are set to 50 to avoid overfitting. Finally, the optimal hyperparameter combination is determined through the tuning results of Optuna for model training and validation.

[0050] S5. Use the trained visibility level inversion model to predict the visibility level of the target area and output the inversion result.

[0051] After training, compare the results of the TabNet visibility level inversion model (hereinafter referred to as the TabNet model) with the random forest model. From the evaluation model, the steps used are S502: Compare with the random forest model after training to evaluate the model. The statistics used are accuracy Accuracy, F1 score F1-Score, and prediction error (PredictionError). Among them, Accuracy and F1-Score are defined as follows:

[0052]

[0053] The specific calculation formula for the recall rate Recall is as follows:

[0054]

[0055] The comparison results are shown in the following table, demonstrating the classification performance of the TabNet model and the random forest model.

[0056] Table 1 - Classification Performance Comparison Table

[0057]

[0058] As can be seen from Table 1, TabNet is significantly superior to Random Forest in terms of overall accuracy (0.618 vs. 0.576) and F1 scores at levels 1 and 2 (0.50 vs. 0.43; 0.54 vs. 0.46), indicating that it has a stronger classification ability for complex categories. At the same time, the advantages of the TabNet model are more prominent in most key indicators, and the overall classification effect is better.

[0059] Figure 2 Shows the comparison between the actual and predicted levels of the TabNet model for different categories. There are two bars for each category in the figure, representing the actual value and the predicted value respectively. As can be seen from the figure, the actual and predicted values for each category are very close in quantity, indicating that the TabNet model has high prediction accuracy. It shows that the model performs relatively evenly across categories and has small prediction errors.

[0060] Figure 3 Shows the accuracy map at the station level, displaying the accuracy distribution of each station in the entire region. The horizontal axis in the figure represents longitude (Lon), the vertical axis represents latitude (Lat), and each point represents the accuracy of a station. The color bar represents the accuracy rate, from light color (low accuracy) to dark color (high accuracy). As can be seen from the figure, there are differences in accuracy values among different stations. The darker regions indicate higher accuracy, while the lighter regions indicate lower accuracy. The accuracy values are relatively uniform in most parts of the figure, indicating that the TabNet model performs relatively consistently across the entire region.

[0061] Figure 4 Shows the accuracy distribution for different time periods (day and night) and different hours. The horizontal axis in the figure represents day and night, the vertical axis represents hours (from 0 to 23), and the number in each cell represents the accuracy rate for the corresponding time period and hour. The color bar represents the magnitude of the accuracy rate, from light color (low accuracy) to dark color (high accuracy). As can be seen from the figure, the accuracy rate during the day is generally higher than that at night. The accuracy rate during the day ranges from 0.62 to 0.70, with the highest accuracy rate of 0.70 at 9 o'clock and the lowest accuracy rates of 0.62 at 4 o'clock and 5 o'clock. The accuracy rate at night ranges from 0.58 to 0.65, with the highest accuracy rate of 0.65 at 0 o'clock and the lowest accuracy rates of 0.58 and 0.59 at 3 o'clock and 4 o'clock respectively. Overall, the TabNet model performs relatively stably at each time period.

[0062] Figure 5It shows the accuracy comparison of the model in different seasons. The horizontal axis represents the seasons, including spring, summer, autumn, and winter, and the vertical axis represents the accuracy. It can be seen from the figure that the model has the highest accuracy in winter, followed by autumn and spring, and the lowest accuracy in summer. Generally speaking, there are certain fluctuations in the accuracy of the model among different seasons, but the overall accuracy remains around 0.62, indicating that the TabNet model performs relatively stably in each season.

[0063] In addition, the results of high-precision inversion of the visibility level in Anhui region from December 28 to December 30, 2023 are as Figure 6 shown. By comparing the data of the three days, the similarities and differences in the spatial distribution between the predicted values and the actual values can be observed. During these three days, the distribution trends of the predicted values and the actual values are basically the same, indicating that the prediction model has a certain degree of accuracy in capturing the spatial changes of the visibility level. These comparison charts further show that the TabNet model can better reflect the distribution of the actual visibility level in most cases.

[0064] The TabNet model captures the detailed changes in the characteristics of water vapor, aerosols, etc. in the atmosphere through deep learning of the multi-spectral data of Fengyun 4A satellite. This achievement can provide a scientific reference for traffic management and environmental governance, and contribute to the coordinated optimization of traffic safety and air quality.

[0065] For the remaining technical features in the above embodiments, those skilled in the art can flexibly select them according to the actual situation to meet different specific actual needs. However, it is obvious to those of ordinary skill in the art that these specific details do not have to be adopted to implement the present invention. In other instances, in order to avoid confusing the present invention, well-known components, structures, or parts are not specifically described, and they are all within the scope of the technical solutions claimed in the claims of the present invention.

[0066] Modifications and changes made by those skilled in the art that do not depart from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention. In the above description, in order to provide a thorough understanding of the present invention, a large number of specific details are elaborated. However, it is obvious to those of ordinary skill in the art that these specific details do not have to be adopted to implement the present invention. In other instances, in order to avoid confusing the present invention, well-known technologies, such as specific construction details, operating conditions, and other technical conditions, are not specifically described.

[0067] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A high-precision inversion method for visibility level based on satellite data and TabNet model, characterized by: Here are the steps: S1. Obtain multispectral remote sensing satellite data of geostationary meteorological satellites and ground-measured visibility data of ground measurement stations in the target area during the measurement period; S2. Perform temporal and spatial matching between satellite data and ground-measured visibility data. The matching rule is to select satellite data with latitude and longitude that are within 0.04° of the ground measurement station and calculate the average value, which represents the satellite data of all spectral bands at this moment. S3. Perform quality control on satellite data and ground-measured visibility data, and classify visibility data; S4, all data are divided into training set and test set, input into TabNet visibility level inversion model for training, and the model performance is improved through hyperparameter optimization; S5. Use the trained visibility level inversion model to predict the visibility level of the target area and output the inversion result.

2. The high-precision inversion method for visibility level based on satellite data and TabNet model according to claim 1, characterized in that: In S1, the multispectral remote sensing satellite data of the geostationary meteorological satellite includes 14 spectral bands, with a spatial resolution of 4KM. The download time period is the entire hour of each day, with a total of 24 satellite data per day.

3. The high-precision inversion method for visibility level based on satellite data and TabNet model according to claim 1, characterized in that: In S1, the measured visibility data of the ground measurement station comes from the hourly ground measurement data of visibility during the measurement period of the ground measurement station of the China Meteorological Administration.

4. The high-precision inversion method for visibility level based on satellite data and TabNet model according to claim 1, characterized in that: In S3, the quality control process includes: Eliminate the data in which all channels in the satellite data are all 0; Eliminate the missing data in any channel of satellite data; Eliminate the measured visibility data of ground measurement stations with an altitude higher than 500m; Eliminate the missing data in the ground-measured visibility data; Eliminate outliers in the ground-measured visibility data.

5. The high-precision inversion method for visibility level based on satellite data and TabNet model according to claim 1, characterized in that: In S3, the standard for classifying the visibility data is: the visibility data is divided into 4 levels according to (0, 200m], (200m, 500m], (500m, 1000m], (1000m, 10000m].

6. The high-precision inversion method for visibility level based on satellite data and TabNet model according to claim 1, characterized in that: In S4, the ratio of training set to test set is 8:

2. The input features of TabNet visibility level inversion model are 18 in total, and the output is visibility level.

7. The high-precision inversion method for visibility level based on satellite data and TabNet model according to claim 1, characterized in that: In S4, Optuna is used to automatically tune hyperparameters. The tuning range includes: setting the initial range of the learning rate of the AdamW optimizer to [0.00001, 0.01], the width of the decision prediction layer n_a to [64, 256], the attention embedding width of each mask n_d to [64, 256], the number of decision steps in the structure n_steps to [3, 10], the number of independent GatedAttention layers n_independent to [1, 5], the sparse regularization coefficient lambda_sparse to [0.000001, 0.01], the weight decay coefficient weight_decay to [0.000001, 0.001], the batch size batch_size to [256, 1024], and setting the number of early stopping rounds to 50 during training.