HY-2A satellite ocean water vapor inversion method based on machine learning

By combining HY-2A satellite data with ERA5 reanalysis data and adopting a variety of machine learning algorithms and Bayesian optimization techniques, the problem of low water vapor inversion accuracy of the HY-2A satellite in ocean environments was solved, high-precision, real-time ocean water vapor inversion was achieved, and the interpretability of the model was enhanced, making it suitable for multiple meteorological monitoring fields.

CN120632292APending Publication Date: 2025-09-12TONGJI UNIV

Patent Information

Application Number
CN202510707628.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The existing HY-2A satellite scanning microwave radiometer has low ocean water vapor inversion accuracy due to the complex nonlinear radiation transmission process in the ocean environment, and the machine learning method is computationally limited within a limited time and has poor interpretability.

Method used

Combining the HY-2A satellite scanning microwave radiometer data with ERA5 reanalysis data, a variety of machine learning algorithms (such as random forest, gradient boosting tree, stacking model, etc.) and Bayesian optimization techniques are used to construct a high-precision ocean water vapor inversion model through data preprocessing, feature extraction and standardization, and the SHAP method is introduced for interpretation.

Benefits of technology

It achieves high-precision, real-time ocean water vapor inversion, improves the robustness and interpretability of the model, and is significantly better than traditional linear regression methods. It is suitable for numerical weather forecasting, extreme weather forecasting and marine climate monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632292A_ABST
    Figure CN120632292A_ABST
Patent Text Reader

Abstract

The invention provides an HY-2A ocean water vapor inversion method based on machine learning, and the method comprises the steps: constructing a high-quality HY-2A scanning microwave radiometer data set through the data preprocessing, matching, feature extraction and normalization processing of HY-2A satellite scanning microwave radiometer data and ERA5 reanalysis data. A plurality of machine learning models are used for training and testing, Bayesian optimization is used for realizing hyper-parameter automatic adjustment, an optimal solution in finite time is obtained, an SHAP method is used for explaining the models, contribution of each characteristic variable is clear, interpretability of the used machine learning models is improved, and finally high-precision inversion of the water vapor over the sea is realized. The method has efficient calculation performance, can realize rapid processing of large-scale marine meteorological data, and improves the interpretability of a machine learning model. The method has important application value in numerical weather forecast, ocean extreme weather forecast, climate monitoring and other geophysical and meteorological fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of remote sensing satellite data processing and artificial intelligence, and in particular to a HY-2A satellite ocean water vapor inversion method based on machine learning. Background Art

[0002] Atmospheric water vapor is a key component of the atmosphere, significantly impacting the hydrological cycle, global energy balance, and climate change. Approximately 80% of this water vapor originates from the ocean, making efficient and accurate monitoring of ocean water vapor crucial for understanding the ocean-atmosphere system.

[0003] In recent years, many technologies have been applied to ocean water vapor retrieval, including radiosondes, global navigation satellite system meteorology, and satellite remote sensing. However, the complex ocean environment leads to high deployment costs for radiosondes and global navigation satellite systems, making large-scale ocean water vapor retrieval difficult.

[0004] With the rapid development of remote sensing satellites and the sensors they carry, satellite remote sensing technology has become the primary method for inverting ocean water vapor. The HY-2A satellite is my country's first ocean environment dynamics satellite, and its onboard scanning microwave radiometer has accumulated a large amount of raw water vapor data.

[0005] The operational technology of the traditional HY-2A satellite scanning microwave radiometer mainly relies on a simple linear regression algorithm. However, in the ocean environment, due to the complex and highly nonlinear radiation transmission process, this method often cannot fully utilize multi-channel brightness temperature data and various environmental parameters, resulting in poor inversion accuracy.

[0006] In recent years, machine learning has demonstrated remarkable performance in big data processing and pattern recognition, particularly in capturing complex nonlinear relationships between input and target variables. For example, in their paper "Study of cold sky calibration and geophysical parameters retrieval for HY-2Asatellite scanning microwave radiometer," W. Zhou et al. employed artificial neural networks, decision trees, and their integrated algorithms to fully exploit the implicit information in multi-source data. They also combined satellite observations (such as HY-2A multi-channel brightness temperature data) with reanalysis data (such as ERA5 total column water vapor data) to achieve in-depth analysis and high-precision retrieval of the spatiotemporal distribution of ocean water vapor. However, machine learning techniques suffer from computational limitations within time constraints and poor interpretability. Summary of the Invention

[0007] To address the high cost of deploying radiosonde balloons and global navigation satellite systems at sea, making it difficult to achieve high-precision inversion of large-scale ocean water vapor, this paper utilizes satellite remote sensing technology to construct a high-quality HY-2A scanning microwave radiometer dataset and combines it with machine learning methods to achieve high-precision inversion of large-scale ocean water vapor. To address the problem that existing HY-2A satellite scanning microwave radiometer operational technology, due to its simple model, fails to take into account the complex ocean environment, the information of the satellite's own instruments, and the difficulty in fully simulating the actual radiation transmission process, resulting in low accuracy in large-scale ocean water vapor inversion, this paper combines comprehensive machine learning methods (decision trees and neural networks, and their ensemble learning) with different ensemble learning methods (homogeneous learner ensembles and heterogeneous learner ensembles). This method considers the complex ocean environment and the satellite's own instrument information as machine learning input features, and leverages the ability of machine learning to effectively capture nonlinear relationships, achieving high-precision inversion of large-scale ocean water vapor using comprehensive machine learning methods. Because machine learning methods often face challenges in determining optimal hyperparameters within a limited timeframe and suffer from poor interpretability, this paper utilizes Bayesian optimization techniques to adjust the hyperparameters of machine learning models, optimizing the model's prediction accuracy within a limited timeframe. Furthermore, this paper introduces SHAP (SHapley Additive exPlanations), a game-theoretic approach that quantitatively evaluates the contribution of various features to water vapor inversion. This approach achieves high-precision ocean water vapor inversion by combining satellite remote sensing data (such as HY-2A multi-channel brightness temperature data) with reanalysis data (such as ERA5 total column water vapor data), combined with rigorous data preprocessing, feature extraction, data standardization, and multiple machine learning algorithms.

[0008] In order to achieve the above object, the technical solution of the present invention is as follows: Step S1: HY-2A satellite SMR and ERA5 data acquisition and preprocessing; The purpose of this step is to obtain the required inverted water vapor observation brightness temperature data from the scanning microwave radiometer of the HY-2A satellite, and to obtain the corresponding environmental data through the ERA5 reanalysis data.

[0009] First, Scanning Microwave Radiometer data will be downloaded from the National Oceanic Administration's Satellite Ocean Application Center platform to obtain corresponding L1B satellite observation data. ERA5 data will be obtained through the European Climate Data Service (ECMWF), which provides high-resolution global reanalysis data.

[0010] The data is then pre-processed, including cleaning and noise removal, to ensure data quality. Data fusion technology is also used to match and synchronize data from two different sources, ensuring consistency and reliability in time and space, providing accurate input data for subsequent analysis.

[0011] Step S2: sample data set construction and feature extraction; Based on the preprocessed data, a sample dataset is constructed for machine learning model training.

[0012] First, based on prior satellite observations and the actual atmospheric radiation transmission process, various meteorological and oceanographic features (such as temperature, humidity, and wind speed) are extracted. These features serve as input to the model. These feature extraction techniques employ conventional methods such as data dimensionality reduction and normalization to reduce computational complexity and improve feature discrimination.

[0013] Then, by combining satellite remote sensing data with numerical meteorological model data, a data set containing multidimensional information was constructed, providing high-quality and comprehensive input data for the next step of model training.

[0014] Step S3: Machine learning model construction and training; In this step, a series of supervised learning machine learning models were constructed to invert ocean water vapor. The core idea of ​​these models is to predict water vapor changes by mapping features related to water vapor changes to the target variable.

[0015] Specifically, we employed ensemble learning algorithms (such as random forests, gradient boosting trees, and stacked models) and deep learning algorithms (such as neural networks) for model training. Through cross-validation and model optimization, we selected the optimal parameters and model architecture to maximize prediction accuracy. During training, we continuously optimized feature selection and model training to ensure robustness and accuracy in various scenarios.

[0016] Step S4: Model interpretation and result comparison.

[0017] After the model training is completed, the model results are analyzed and interpreted.

[0018] By quantitatively evaluating the training results, comparing the errors between the predicted results and the actual observed data, the model performance is evaluated using error analysis indicators (such as mean square error, mean absolute error, etc.).

[0019] At the same time, the model's prediction results are explained by combining SHAP interpretability with physical principles (such as the water vapor-temperature relationship in atmospheric physics) to ensure the scientificity and rationality of the model.

[0020] In order to verify the effectiveness of the model, a comparative analysis will be conducted with existing water vapor prediction methods (such as traditional methods based on physical models) to verify the advantages of this technical solution over traditional methods.

[0021] Furthermore, the step S1 includes: Step S11: Acquisition of HY-2A satellite SMR data and data preprocessing; The data collected by the HY-2A satellite SMR include the following instrument information: The data collected by the present invention include position and time information, observation brightness temperature information of each frequency band (e.g. 6.6, 10.7, 18.7, 23.8, 37 GHz) under different polarization modes (vertical and horizontal polarization), Earth incidence angle, Earth azimuth angle, solar zenith angle and solar flash angle Data acquisition requirements: a. The data source is the National Oceanic Administration Satellite Ocean Application Center (NSOAS).

[0022] Data preprocessing process: During the data preprocessing stage, land data were first removed using land-sea ratio markers, and data within 200 km of the coastline were excluded to prevent land contamination. Next, data with brightness temperatures between 3K and 350K were screened for strict quality control, and data with sun flare angles less than 25° were removed to eliminate solar flare interference. In addition, brightness temperature data for the 18.7–37 GHz channel were screened, retaining only records with a positive difference between vertical and horizontal polarization to ensure the validity of oceanographic retrieval. 18V GHz data with brightness temperatures greater than 240K were further removed to exclude precipitation effects. Finally, the spatial standard deviation distribution of the 23.8 GHz and 37 GHz channels was calculated, and reasonable thresholds were set to exclude tail data, effectively avoiding data contamination caused by abnormally high spatial standard deviations.

[0023] Data storage format: The data is stored in HDF5 format. The storage content includes longitude and latitude, time, brightness temperature of polarization observations in each frequency band, Earth incidence angle, Earth azimuth, solar zenith angle, and solar flare angle. All data retain four significant digits.

[0024] Step S12: Acquisition of ERA5 reanalysis data and data preprocessing; The ERA5 reanalysis data contains a variety of environmental information such as ocean surface brightness temperature, ocean surface wind speed, total water vapor column content, liquid water in clouds, and sea ice. Data acquisition requirements: a. Temporal resolution of 1 hour; b. Horizontal resolution of 0.25°; c. Data source: European Centre for Medium-Range Weather Forecasts (ECMWF).

[0025] Data preprocessing process: The present invention first uses sea ice markers to eliminate the influence of sea ice, then excludes data with ocean surface brightness temperature less than 271.15K to eliminate additional sea ice interference, and simultaneously excludes data with ocean surface wind speeds less than 4m / s and greater than 20m / s, thereby effectively controlling the influence of sea surface roughness and foam.

[0026] Data storage format: a. Stored in a standard gridded format; b. Contains latitude and longitude, time, ocean surface wind speed, ocean surface brightness temperature, cloud liquid water, total columnar water vapor content, and sea ice marker information; c. Retains four significant digits.

[0027] Furthermore, the step S2 includes: Step S21: spatiotemporal matching of HY-2A SMR data and ERA5 data; The preprocessed HY-2A-SMR and ERA5 data were matched based on time, latitude, and longitude information. Bilinear interpolation was used spatially and linear interpolation was used temporally to synchronize the two data types. This resulted in a high-quality matching dataset with a spatial resolution of 0.25° and a temporal resolution of 1 hour, laying a solid foundation for subsequent data analysis and model building.

[0028] Step S22: feature extraction; For the spatiotemporally matched dataset, we leveraged prior physical information to select ocean environmental and instrumental information with high correlations with water vapor characteristics as features for feature extraction. The extracted features include 18 basic variables: longitude and latitude information, polarimetric brightness temperature data at various frequency bands (e.g., 6.6 GHz, 10.7 GHz, 18.7 GHz, 23.8 GHz, and 37 GHz), sea surface temperature, wind speed, and cloud water content, as shown in Table 1. These data reflect the complex ocean information observed by satellite remote sensing and the satellite's own instrumental information, providing solid support for subsequent model training for ocean water vapor retrieval.

[0029] Table 1 18 machine learning input features of the present invention Step S23: data standardization; In this paper, to ensure consistency in the numerical range and units of each extracted feature and effectively eliminate bias caused by differences in data scale, the present invention uses normalization to standardize all features. Specifically, the Z-Score normalization method is used, that is, the mean of each feature is subtracted and divided by its standard deviation, so that the converted data follows a normal distribution with a mean of 0 and a standard deviation of 1.

[0030] This normalization process not only eliminates inconsistencies between physical quantities caused by differences in units and scales, but also makes features comparable within the same numerical range, thereby improving the convergence speed and stability of machine learning models during training. Normalized data effectively reduces the problems of poor accuracy and model instability caused by data scale mismatch, thereby improving the overall model's predictive accuracy and robustness.

[0031] Step S24: data set division and downsampling; After data normalization, downsampling is performed to reduce data density and computational complexity, improve computational efficiency, and ensure sample representativeness. Data sparsification is typically performed using a fixed sampling ratio (e.g., 1:10), with the sample dataset divided according to pre-set proportions (e.g., 60% for training, 20% for testing, and 20% for validation). This step ensures a balanced data distribution and provides sufficient, high-quality samples for subsequent machine learning model training, hyperparameter tuning, and validation, effectively improving the model's generalization and predictive accuracy.

[0032] Furthermore, step S3 includes: Step S31: machine learning model construction; To fully exploit the complex nonlinear relationships between the various extracted features and the water vapor data above the ocean, this paper uses multiple classic machine learning algorithms to build a prediction model. Each model has unique advantages in different data dimensions, and the advantages are complemented through ensemble learning. Specifically, it includes: Multilayer Perceptron (MLP): It uses a feedforward neural network architecture and is designed with at least two hidden layers. It implements nonlinear mapping through activation functions, effectively capturing subtle physical information in the input features and is suitable for deep learning of high-dimensional data features.

[0033] Random Forest (RF): uses the Bagging strategy to construct multiple decision trees, reduces the overfitting risk of a single tree through random sampling and feature subset selection, and uses a majority voting mechanism to improve the robustness of the overall prediction.

[0034] XGBoost: Based on gradient boosting decision trees, it achieves efficient modeling of complex nonlinear relationships by gradually optimizing the loss function. Its built-in regularization mechanism effectively suppresses model overfitting and is particularly suitable for high-dimensional and sparse data scenarios.

[0035] Stacking ensemble model based on heterogeneous learners: adopts a weighted fusion strategy of multi-model prediction results and uses meta-models (such as linear regression or other simple regression models) to integrate the advantages of each base model to further improve the accuracy and stability of the overall prediction.

[0036] This multi-model construction strategy not only fully utilizes the advantages of each algorithm in feature extraction and pattern recognition, but also integrates the prediction results of different models through an integrated approach to achieve high-precision inversion of the temporal and spatial distribution of ocean water vapor, providing comprehensive and reliable data support for subsequent applications.

[0037] Step S32: model hyperparameter tuning; After the standardized sample dataset is constructed, the present invention divides the dataset into 60% as a training set, 20% as a test set, and 20% as a validation set to ensure that the data samples are representative and balanced. During the model training process, the validation set and Bayesian optimization methods are used to automatically tune the key hyperparameters of each model (such as the learning rate and number of hidden layer neurons in MLP, the number and depth of trees in RF, and the step size and regularization coefficient in XGBoost). Bayesian optimization constructs a proxy model to efficiently search in a high-dimensional hyperparameter space, which not only shortens the parameter tuning time but also enables the realization of the globally optimal parameter combination.

[0038] To further improve the generalization ability of the model, the present invention also introduces a K-fold cross-validation strategy during stacking to fully evaluate the performance of the model under different data partitioning to prevent the model performance from being affected by uneven data partitioning or overfitting.

[0039] Step S33: model training and performance evaluation; Each of the aforementioned models was trained using the training set. Early stopping and regularization methods were incorporated during training to ensure stable model convergence and effectively prevent overfitting. After training, the model's predictive performance on the test set was comprehensively evaluated using a series of metrics (such as the coefficient of determination (R²) and root mean square error (RMSE).

[0040] Furthermore, the step S4 includes: Step S41: introduction of a posteriori interpretation method; To reveal the internal decision-making mechanisms of machine learning models in ocean water vapor inversion, this paper introduces the SHAP (SHapley Additive exPlanations) method for posterior interpretation. Based on game theory, this method rationally allocates the marginal contribution of each feature to the predicted output, enabling a feature-by-feature explanation of complex nonlinear models. This provides a transparent and quantitative scientific basis for the model's internal mechanisms, thereby enhancing model interpretability.

[0041] Step S42: quantitative analysis of feature contributions; Using SHAP values, this paper provides a detailed analysis of model prediction results. Specifically, it quantifies the contribution of derived features such as latitude and longitude, brightness temperature in various frequency bands (e.g., 6.6 GHz, 10.7 GHz, 18.7 GHz, 23.8 GHz, and 37 GHz), sea surface temperature, wind speed, cloud water content, and various angular differences to water vapor inversion. Intuitive visualizations show the influence weights and action paths of each feature, ensuring objectivity and accuracy in data interpretation while also providing strong data support for subsequent model optimization and parameter tuning.

[0042] Step S43: comparing model performance and determining the best model; The traditional commercial spaceborne microwave radiometer water vapor inversion uses a simple linear regression formula as follows: in, represents the coefficient, (i=1,2,3,…,9) corresponds to the brightness temperature data of nine channels: 6.6V, 6.6H, 10.7V, 10.7H, 18.7V, 18.7H, 23.8V, 37V and 37H. For the four channels (i=1,2,3,4) at 6.6 GHz and 10.7 GHz, there are = ; For the remaining channels (i=5,6,7,8,9), there are = .

[0043] Based on a posteriori interpretation, this paper systematically compares the performance of various machine learning models on a test set. Using multiple performance metrics, such as the coefficient of determination (R²), root mean square error (RMSE), and feature contribution balance, combined with SHAP interpretation results, the model's fitting accuracy, prediction stability, and prediction accuracy are comprehensively evaluated. Ultimately, through comparative analysis, the optimal model is identified, outperforming traditional linear regression water vapor retrieval methods in overall performance, providing an optimal solution for practical applications.

[0044] Step S44: outputting results and promoting their application; After selecting the optimal model, the water vapor inversion results are output and transmitted via standard interfaces to practical application platforms such as numerical weather forecasting, extreme weather forecasting, and ocean climate monitoring. To further ensure system performance, this invention incorporates a regular evaluation and feedback mechanism. By monitoring the quality of the output results and performing performance review, model parameters and interpretation strategies are continuously optimized to ensure that the water vapor retrieval system maintains high accuracy and robustness during actual operation, thereby providing stable and reliable data support and technical assurance for related fields.

[0045] In summary, the present invention provides a machine learning-based ocean water vapor inversion method based on the HY-2A satellite. This method achieves high-precision inversion of the spatiotemporal distribution of ocean water vapor by combining satellite remote sensing data (such as HY-2A multi-channel brightness temperature data) with ERA5 reanalysis data. This method includes rigorous data preprocessing and quality control measures, such as using land, sea, and sea ice markers to eliminate non-target data, filtering brightness temperature and wind speed ranges to eliminate interference factors, and high-precision storage of observational data in the HDF5 format. Subsequently, through spatiotemporal matching, feature extraction, and data standardization, a multidimensional sample dataset is constructed, including latitude and longitude, brightness temperature in various frequency bands, sea surface temperature, wind speed, and cloud water content. Derived features are then constructed to fully reflect the physical properties of ocean water vapor. Next, the present invention utilizes multiple machine learning algorithms, including multilayer perceptrons, random forests, XGBoost, and stacking ensemble models, combined with Bayesian optimization and cross-validation strategies, to achieve efficient model training and hyperparameter tuning, resulting in prediction results that significantly outperform traditional linear regression methods in both fitting accuracy and robustness. Finally, the SHAP method is used to perform a posteriori interpretation of the prediction results of each model, clarifying the contribution of each feature to the search results. The search results of the optimal model are then exported to application platforms such as numerical weather forecasting and extreme weather monitoring through a standard interface. Overall, this invention not only overcomes the limitations of traditional methods in dealing with complex nonlinear radiation transfer processes, fully leveraging the advantages of machine learning in pattern recognition and data fusion, but also enhances the interpretability of machine learning models, providing reliable and accurate data support and technical guarantees for fields such as climate change monitoring, weather forecasting, and extreme weather event warnings.

[0046] Beneficial effects Compared with the prior art, the present invention has the following advantages: First, the present invention significantly improves data quality by effectively integrating satellite remote sensing data (such as HY-2A multi-channel brightness temperature data) with ERA5 reanalysis data, and adopts strict data preprocessing (including sea and land segmentation, sea ice marking, brightness temperature and wind speed screening, etc.), providing a solid data foundation for subsequent spatiotemporal retrieval of ocean water vapor, thereby achieving high-precision, real-time water vapor retrieval.

[0047] Secondly, the present invention utilizes a variety of advanced machine learning algorithms, including multi-layer perceptrons, random forests, XGBoost, and a stacking ensemble model based on heterogeneous learners. Through Bayesian optimization, hyperparameters are globally tuned to ensure the model's accuracy and stability in complex nonlinear radiation transmission processes. Compared to traditional linear regression methods, the present invention's multi-machine learning model stacking strategy not only fully captures the complex nonlinear relationships between features but also excels in reducing prediction bias and improving robustness.

[0048] In addition, the present invention introduces the SHAP posterior interpretation method to quantitatively analyze the prediction results of each model, and clearly reveals the specific contributions of longitude and latitude, brightness temperature in each frequency band, sea surface temperature, wind speed, cloud water content and their derived characteristics in water vapor retrieval. The results are as follows Figure 4 (a), (b), (c), and the contribution of each base learner in the stacked model meta-learner is also explained a posteriori, as shown in Figure 4 (d) This transparent explanation mechanism not only enhances the interpretability of the model but also provides a reliable basis for further model optimization and parameter tuning.

[0049] Finally, the present invention adopts a standardized data storage and data transmission interface to output the optimized water vapor inversion results to application platforms such as numerical weather forecasting, extreme weather warning, and marine environment assessment, ensuring that the system has high precision and high robustness in practical applications, and has broad application prospects and significant social and economic benefits.

[0050] In summary, the present invention has achieved breakthroughs in improving water vapor inversion accuracy, enhancing model generalization capabilities, and optimizing data quality. It has fully overcome the limitations of traditional technologies in complex marine environments, demonstrating significant technological progress and broad practical application value.

[0051] The technical solution of this invention not only constructs a high-quality HY-2A SMR dataset but also selects a more comprehensive machine learning model to overcome the shortcomings of traditional linear regression in processing complex nonlinear radiation transmission processes. This method fully leverages the advantages of machine learning algorithms in feature mining and pattern recognition, achieving high-precision inversion of ocean water vapor and filling the gap in the high-precision water vapor dataset of the HY-2A satellite scanning microwave radiometer. Furthermore, SHAP is used to provide post-hoc interpretation of machine learning model predictions, increasing the transparency of the machine learning model. This solution has broad application prospects in fields such as numerical weather forecasting, extreme weather forecasting, climate change monitoring, and marine environmental assessment. It also provides reliable data support and technical assurance for related fields, and has significant scientific significance and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 Schematic diagram of the process flow of the method of the present invention; Figure 2 A schematic diagram of constructing a machine learning model for the method of the present invention; Figure 3 Kernel density scatter plots of the water vapor inversion results over the ocean by the HY-2A SMR in an embodiment of the present invention: (a) MLP model, (b) RF model, (c) XGBoost model, (d) stacking model, and (e) traditional model; Figure 4 SHAP values ​​of all input features for the embodiments of the present invention: (a) MLP model, (b) RF model, (c) XGBoost model, and (d) stacked meta-model. DETAILED DESCRIPTION

[0053] The technical solution provided by this application will be further described below in conjunction with specific embodiments and accompanying drawings. The advantages and features of this application will become more apparent with reference to the following description.

[0054] refer to Figure 1 In a preferred embodiment, the present invention uses a HY-2A satellite ocean water vapor inversion method based on machine learning, specifically comprising: Step S1: HY-2A satellite SMR and ERA5 data acquisition and preprocessing; Specifically, step S1 includes: Step S11: Acquisition of 2014 HY-2A satellite SMR data and data preprocessing; The data from the HY-2A satellite SMR include the following instrument information: The data collected by the present invention include position and time information, observation brightness temperature information in various frequency bands (e.g., 6.6, 10.7, 18.7, 23.8, and 37 GHz) under different polarization modes (vertical and horizontal polarization), Earth incidence angle, Earth azimuth angle, solar zenith angle, and solar flare angle.

[0055] Data acquisition requirements: a. The data source is the National Oceanic Administration Satellite Ocean Application Center (NSOAS).

[0056] Data preprocessing process: During the data preprocessing stage, land data were first removed using land-sea ratio markers, and data within 200 km of the coastline were excluded to prevent land contamination. Next, data with brightness temperatures between 3K and 350K were screened for strict quality control, and data with sun flare angles less than 25° were removed to eliminate solar flare interference. In addition, brightness temperature data for the 18.7–37 GHz channel were screened, retaining only records with a positive difference between vertical and horizontal polarization to ensure the validity of oceanographic retrieval. 18V GHz data with brightness temperatures greater than 240K were further removed to exclude precipitation effects. Finally, the spatial standard deviation distribution of the 23.8 GHz and 37 GHz channels was calculated, and reasonable thresholds were set to exclude tail data, effectively avoiding data contamination caused by abnormally high spatial standard deviations.

[0057] Data storage format: The data is stored in HDF5 format. The storage content includes longitude and latitude, time, brightness temperature of polarization observations in each frequency band, Earth incidence angle, Earth azimuth, solar zenith angle, and solar flare angle. All data retain four significant digits.

[0058] Step S12: Obtain 2014 ERA5 reanalysis data and data preprocessing; The ERA5 reanalysis data contains a variety of environmental information such as ocean surface brightness temperature, ocean surface wind speed, total water vapor column content, liquid water in clouds, and sea ice. Data acquisition requirements: a. Temporal resolution of 1 hour; b. Horizontal resolution of 0.25°; c. Data source: European Centre for Medium-Range Weather Forecasts (ECMWF).

[0059] Data preprocessing process: The present invention first uses sea ice markers to eliminate the influence of sea ice, then excludes data with ocean surface brightness temperature less than 271.15K to eliminate additional sea ice interference, and simultaneously excludes data with ocean surface wind speeds less than 4m / s and greater than 20m / s, thereby effectively controlling the influence of sea surface roughness and foam.

[0060] Data storage format: a. Stored in a standard gridded format; b. Contains latitude and longitude, time, ocean surface wind speed, ocean surface brightness temperature, cloud liquid water, total columnar water vapor content, and sea ice marker information; c. Retains four significant digits.

[0061] Step S2: sample data set construction and feature extraction; Specifically, step S2 includes: Step S21: spatiotemporal matching of HY-2A SMR data and ERA5 data; The preprocessed HY-2A-SMR and ERA5 data were matched based on time, latitude, and longitude information. Bilinear interpolation was used spatially and linear interpolation was used temporally to synchronize the two data types. This resulted in a high-quality matching dataset with a spatial resolution of 0.25° and a temporal resolution of 1 hour, laying a solid foundation for subsequent data analysis and model building.

[0062] Step S22: feature extraction; For the spatiotemporally matched dataset, we leveraged prior physical information to select ocean environmental and instrumental information with a high correlation with water vapor characteristics as features for feature extraction. These extracted features include 18 basic variables, including longitude and latitude, polarization-based brightness temperature data for various frequency bands (e.g., 6.6 GHz, 10.7 GHz, 18.7 GHz, 23.8 GHz, and 37 GHz), sea surface temperature, wind speed, and cloud water content. These data reflect the complex ocean information observed by satellite remote sensing and the satellite's own instrumental information, providing solid support for subsequent model training for ocean water vapor retrieval.

[0063] Step S23: data standardization; In this paper, to ensure the consistency of the numerical range and units of each extracted feature and effectively eliminate the deviation caused by data scale differences, the present invention uses normalization processing to standardize all features. Specifically, the Z-Score normalization method is used, that is, the mean of each feature is subtracted and divided by its standard deviation, so that the converted data follows a normal distribution with a mean of 0 and a standard deviation of 1.

[0064] This normalization process not only eliminates inconsistencies between physical quantities caused by differences in units and scales, but also makes features comparable within the same numerical range, thereby improving the convergence speed and stability of machine learning models during training. Normalized data effectively reduces the problems of poor accuracy and model instability caused by data scale mismatch, thereby improving the overall model's predictive accuracy and robustness.

[0065] Step S24: data set division and downsampling; After data normalization, downsampling is performed to reduce data density and computational complexity, improve computational efficiency, and ensure sample representativeness. Data sparsification is typically performed using a fixed sampling ratio (e.g., 1:10), with the sample dataset divided according to pre-set proportions (e.g., 60% for training, 20% for testing, and 20% for validation). This step ensures a balanced data distribution and provides sufficient, high-quality samples for subsequent machine learning model training, hyperparameter tuning, and validation, effectively improving the model's generalization and predictive accuracy.

[0066] Step S3: Machine learning model construction and training Specifically, step S3 includes: Step S31: Machine learning model construction; (e.g. Figure 2 ) To fully exploit the complex nonlinear relationships between the various extracted features and the water vapor data above the ocean, this paper uses multiple classic machine learning algorithms to build a prediction model. Each model has unique advantages in different data dimensions, and the advantages are complemented through ensemble learning. Specifically, it includes: Multilayer Perceptron (MLP): It uses a feedforward neural network architecture and is designed with at least two hidden layers. It implements nonlinear mapping through activation functions, effectively capturing subtle physical information in the input features and is suitable for deep learning of high-dimensional data features.

[0067] Random Forest (RF): uses the Bagging strategy to construct multiple decision trees, reduces the overfitting risk of a single tree through random sampling and feature subset selection, and uses a majority voting mechanism to improve the robustness of the overall prediction.

[0068] XGBoost: Based on gradient boosting decision trees, it achieves efficient modeling of complex nonlinear relationships by gradually optimizing the loss function. Its built-in regularization mechanism effectively suppresses model overfitting and is particularly suitable for high-dimensional and sparse data scenarios.

[0069] Stacking ensemble model based on heterogeneous learners: adopts a weighted fusion strategy of multi-model prediction results and uses meta-models (such as linear regression or other simple regression models) to integrate the advantages of each base model to further improve the accuracy and stability of the overall prediction.

[0070] This multi-model construction strategy not only fully utilizes the advantages of each algorithm in feature extraction and pattern recognition, but also integrates the prediction results of different models through an integrated approach to achieve high-precision inversion of the temporal and spatial distribution of ocean water vapor, providing comprehensive and reliable data support for subsequent applications.

[0071] Step S32: model hyperparameter tuning; After the standardized sample data set is constructed, the present invention divides the data set into 60% as a training set, 20% as a test set, and 20% as a validation set to ensure that the data samples are representative and balanced. During the model training process, the validation set and Bayesian optimization method are used to automatically tune the key hyperparameters of each model (such as the learning rate and the number of hidden layer neurons in MLP, the number and depth of trees in RF, the step size and regularization coefficient in XGBoost, etc.). Bayesian optimization constructs a proxy model to conduct efficient search in a high-dimensional hyperparameter space, which not only shortens the parameter tuning time, but also can achieve the global optimal parameter combination. In order to further improve the generalization ability of the model, the present invention also introduces a K-fold cross-validation strategy when stacking to fully evaluate the performance of the model under different data partitions to prevent the model performance from being affected by uneven data partitioning or overfitting.

[0072] Step S33: model training and performance evaluation; Each of the aforementioned models was trained using the training set. Early stopping and regularization methods were incorporated during training to ensure stable model convergence and effectively prevent overfitting. After training, the model's predictive performance on the test set was comprehensively evaluated using a series of metrics (such as the coefficient of determination (R²) and root mean square error (RMSE).

[0073] Step S4: Model interpretation and result comparison Specifically, step S4 includes: Step S41: introduction of a posteriori interpretation method; To reveal the internal decision-making mechanisms of machine learning models in ocean water vapor inversion, this paper introduces the SHAP (SHapley Additive exPlanations) method for posterior interpretation. Based on game theory, this method rationally allocates the marginal contribution of each feature to the predicted output, enabling a feature-by-feature explanation of complex nonlinear models. This provides a transparent and quantitative scientific basis for the model's internal mechanisms, thereby enhancing model interpretability.

[0074] Step S42: quantitative analysis of feature contributions; Using SHAP values, this paper provides a detailed analysis of model prediction results. Specifically, it quantifies the contribution of derived features such as latitude and longitude, brightness temperature in various frequency bands (e.g., 6.6 GHz, 10.7 GHz, 18.7 GHz, 23.8 GHz, and 37 GHz), sea surface temperature, wind speed, cloud water content, and various angular differences to water vapor inversion. Intuitive visualizations show the influence weights and action paths of each feature, ensuring objectivity and accuracy in data interpretation while also providing strong data support for subsequent model optimization and parameter tuning.

[0075] Step S43: comparing model performance and determining the best model; The traditional commercial spaceborne microwave radiometer water vapor inversion uses a simple linear regression formula as follows: in, represents the coefficient, (i=1,2,3,…,9) corresponds to the brightness temperature data of nine channels: 6.6V, 6.6H, 10.7V, 10.7H, 18.7V, 18.7H, 23.8V, 37V and 37H. For the four channels (i=1,2,3,4) at 6.6 GHz and 10.7 GHz, there are = ; For the remaining channels (i=5,6,7,8,9), there are = .

[0076] Based on a posteriori interpretation, this paper systematically compares the performance of various machine learning models on a test set. Using multiple performance metrics, such as the coefficient of determination (R²), root mean square error (RMSE), and feature contribution balance, combined with SHAP interpretation results, the model's fitting accuracy, prediction stability, and prediction accuracy are comprehensively evaluated. Ultimately, through comparative analysis, the optimal model is identified, outperforming traditional linear regression water vapor retrieval methods in overall performance, providing an optimal solution for practical applications.

[0077] Step S44: outputting results and promoting their application; After selecting the optimal model, the water vapor inversion results are output and transmitted via standard interfaces to practical application platforms such as numerical weather forecasting, extreme weather forecasting, and ocean climate monitoring. To further ensure system performance, this invention incorporates a regular evaluation and feedback mechanism. By monitoring the quality of the output results and performing performance review, model parameters and interpretation strategies are continuously optimized to ensure that the water vapor retrieval system maintains high accuracy and robustness during actual operation, thereby providing stable and reliable data support and technical assurance for related fields.

[0078] The multi-machine learning model stacking strategy of the present invention can not only fully capture the complex nonlinear relationship between each feature, but also perform better in reducing prediction bias and improving robustness. The inversion results and density scatter plots are shown in Figure 2. Figure 3 As shown. Among them, Figure 3 (a) Figure 3 (b) Figure 3 (c) Figure 3 (d) The inversion results and density scatter plots of the MLP model, RF model, XGBoost model, and stacking model, respectively. Figure 3 (e) is the inversion result and density scatter plot of the traditional model. It can be seen that the density scatter plot of the multi-machine learning model stacking strategy of the present invention has smaller prediction deviation and robustness.

[0079] Figure 4 The performance evaluation of these algorithms on the specified test set is shown. Figure 4 (a) MLP model, Figure 4 (b) RF model, Figure 4 (c) SHAP values ​​of XGBoost model for all base learners; Figure 4 (d) SHAP values ​​of the stacked metamodel. The color bar indicates the contribution from high to low.

[0080] The root mean square error (RMSE) values ​​obtained were 2.23 mm for the LR model, 1.55 mm for the MLP, 1.60 mm for the RF, 1.41 mm for the XGBoost, and 1.40 mm for the stacking method. Notably, all machine learning methods achieved RMSE reductions of over 28% compared to the baseline operational model. Furthermore, the PWV estimates generated by the machine learning methods had fewer outliers and a better fit to the regression line than those from the operational model.

[0081] Among all machine learning algorithms, stacked ensemble learning performed the best, leveraging the strengths of heterogeneous learners to reduce both bias and variance. The relative performance of ensemble methods can be attributed to their theoretical foundations: Random Forest (RF, bagging) primarily reduces variance, while XGBoost (boosting) focuses on reducing bias. Consistent with this, XGBoost (RMSE = 1.41 mm) outperformed Random Forest (RMSE = 1.60 mm), but slightly lagged behind the stacked method. XGBoost predictions were closer to the regression line than Random Forest, further supporting the effectiveness of boosting in reducing bias. MLP also outperformed Random Forest, demonstrating the superior feature extraction capabilities of neural networks on this dataset.

[0082] In addition, the present invention introduces the SHAP posterior interpretation method to quantitatively analyze the prediction results of each model, and clearly reveals the specific contributions of longitude and latitude, brightness temperature in each frequency band, sea surface temperature, wind speed, cloud water content and their derived characteristics in water vapor retrieval. The results are as follows Figure 4 (a), (b), (c), and the contribution of each base learner in the stacked model meta-learner is also explained a posteriori, as shown in Figure 4 (d) This transparent explanation mechanism not only enhances the interpretability of the model but also provides a reliable basis for further model optimization and parameter tuning.

[0083] Finally, the present invention adopts a standardized data storage and data transmission interface to output the optimized water vapor inversion results to application platforms such as numerical weather forecasting, extreme weather warning, and marine environment assessment, ensuring that the system has high precision and high robustness in practical applications, and has broad application prospects and significant social and economic benefits.

[0084] In summary, the present invention has achieved breakthroughs in improving water vapor inversion accuracy, enhancing model generalization capabilities, and optimizing data quality. It has fully overcome the limitations of traditional technologies in complex marine environments, demonstrating significant technological progress and broad practical application value.

[0085] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.

Claims

1. A HY-2A satellite ocean water vapor inversion method based on machine learning, characterized in that: By constructing a large-scale spaceborne microwave radiometer dataset and using machine learning methods to achieve high-precision inversion of ocean water vapor from the HY-2A satellite, the following steps are involved: Step S1: HY-2A satellite SMR and ERA5 data acquisition and preprocessing; Step S2: sample data set construction and feature extraction; Step S3: Machine learning model construction and training; Step S4: Model interpretation and result comparison.

2. A HY-2A satellite ocean water vapor inversion method based on machine learning as claimed in claim 1, characterized in that The step S1 comprises: Step S11: Acquisition of HY-2A satellite SMR data and data preprocessing; The HY-2A satellite SMR data includes the following instrument information: a. Position and time information; b. Brightness temperature information observed in different polarization modes in each frequency band; c. Earth incidence angle; d. Earth azimuth angle; e. Solar zenith angle; f. Solar flare angle; Data acquisition requirements: a. The data source is the Satellite Ocean Application Center of the State Oceanic Administration; Data preprocessing process: a. Use land-sea ratio markers to remove land data and exclude data within 200 km of the coastline to prevent land contamination; b. Screen data with brightness temperatures between 3K and 350K for quality control; c. Eliminate data with sun flare angles less than 25° to eliminate the influence of sun flare; d. Screen data with a positive difference between the vertical and horizontal polarization brightness temperatures of 18.7–37 GHz to exclude invalid oceanographic retrievals; e. Eliminate data with 18 V GHz brightness temperatures greater than 240 K to eliminate the influence of precipitation; f. Calculate the spatial standard deviation distribution of the 23.8 GHz and 37 GHz channels and set a threshold to exclude data at the end tails, as their spatial standard deviations are very high and anomalous, and therefore may be contaminated; Data storage format: a. HDF5 format; b. Contains latitude and longitude, time, polarization observation brightness temperature of each frequency band, Earth incidence angle, Earth azimuth, solar zenith angle, and solar flare angle; c. Retains 4 significant digits; Step S12: Acquisition of ERA5 reanalysis data and data preprocessing; The ERA5 reanalysis data includes the following environmental information: a. Ocean surface brightness temperature; b. Ocean surface wind speed; c. Total column water vapor content; d. Cloud liquid water; e. Sea ice; Data acquisition requirements: a. Temporal resolution of 1 hour; b. Horizontal resolution of 0.25°; c. Data source: European Centre for Medium-Range Weather Forecasts; Data preprocessing process: a. Use sea ice markers to remove sea ice effects; b. Exclude data with ocean surface brightness temperatures less than 271.15K to remove additional sea ice effects; c. Exclude data with ocean surface wind speeds less than 4m / s and greater than 20m / s to control for the effects of sea surface roughness and foam; Data storage format: a. Stored in a standard gridded format; b. Contains latitude and longitude, time, ocean surface wind speed, ocean surface brightness temperature, cloud liquid water, total columnar water vapor content, and sea ice marker information; c. Retains four significant digits.

3. A HY-2A satellite ocean water vapor inversion method based on machine learning as claimed in claim 1, characterized in that: The step S2 comprises: Step S21: spatiotemporal matching of HY-2A-SMR data and ERA5 data; The preprocessed HY-2ASMR and ERA5 data were matched based on time, latitude, and longitude information. Bilinear interpolation was used in space, and linear interpolation in time to synchronize the two types of data in the spatiotemporal dimensions. This resulted in a high-quality matching dataset with a spatial resolution of 0.25° and a temporal resolution of 1 hour, laying a solid foundation for subsequent data analysis and model building. Step S22: feature extraction; For the spatiotemporally matched dataset, we leverage prior physical information to select ocean environmental and instrumental information with a high correlation with water vapor characteristics as features for extraction. This data reflects the complex ocean information observed by satellite remote sensing and the satellite's own instrumental information, providing solid support for subsequent model training for ocean water vapor inversion. Step S23: data standardization; Normalization is used to standardize all features. Specifically, the Z-Score normalization method is used, that is, the mean of each feature is subtracted and divided by its standard deviation, so that the transformed data obeys a normal distribution with a mean of 0 and a standard deviation of 1. Step S24: data set division and downsampling; After completing data standardization, the data is downsampled to reduce data density and computational complexity, improve computational efficiency, and ensure the representativeness of the sample.

4. A HY-2A satellite ocean water vapor inversion method based on machine learning as claimed in claim 1, characterized in that The step S3 comprises: Step S31: machine learning model construction; We use a variety of classic machine learning algorithms to build prediction models and achieve complementary advantages through ensemble learning. Specifically, we include: Multilayer Perceptron (MLP): This uses a feedforward neural network architecture with at least two hidden layers. It implements nonlinear mapping through activation functions, effectively capturing subtle physical information in input features and is suitable for deep learning of high-dimensional data features. Random Forest (RF): uses the bagging strategy to construct multiple decision trees, reduces the overfitting risk of individual trees through random sampling and feature subset selection, and uses a majority voting mechanism to improve the robustness of the overall prediction; XGBoost: Based on gradient boosting decision trees, it achieves efficient modeling of complex nonlinear relationships by gradually optimizing the loss function. Its built-in regularization mechanism effectively suppresses model overfitting and is particularly suitable for high-dimensional and sparse data scenarios. Stacking ensemble model based on heterogeneous learners: This model adopts a weighted fusion strategy for multi-model prediction results and uses a meta-model to integrate the advantages of each base model to further improve the accuracy and stability of the overall prediction. Step S32: model hyperparameter tuning; After the standardized sample dataset was constructed, 60% of the dataset was used as a training set, 20% as a test set, and 20% as a validation set to ensure that the data samples were representative and balanced. During model training, validation sets and Bayesian optimization methods are used to automatically tune key hyperparameters for each model. Bayesian optimization constructs surrogate models to efficiently search within a high-dimensional hyperparameter space, shortening parameter tuning time and achieving the globally optimal parameter combination. A K-fold cross-validation strategy is also introduced during stacking to evaluate the model's performance under different data partitions, preventing uneven data partitioning or overfitting from impacting model performance. Step S33: model training and performance evaluation; The above models are trained using the training set, and the early stopping strategy and regularization method are combined in the training process to ensure that the model can converge stably and effectively prevent overfitting. After the training is completed, the prediction performance of the model on the test set is comprehensively evaluated through evaluation indicators.

5. A HY-2A satellite ocean water vapor inversion method based on machine learning as claimed in claim 1, characterized in that: The step S4 comprises: Step S41: introduction of a posteriori interpretation method; To reveal the internal decision-making mechanisms of various machine learning models in the ocean water vapor inversion process, the SHAP method was introduced for posterior interpretation. Based on game theory, this method rationally allocates the marginal contribution of each feature in the predicted output to achieve a factor-by-factor explanation of complex nonlinear models, providing a transparent and quantitative scientific basis for the model's internal mechanisms, thereby enhancing model interpretability. Step S42: quantitative analysis of feature contributions; SHAP values ​​are used to analyze the model prediction results in detail. Specifically, the contribution of derived features such as latitude and longitude, brightness temperature in each frequency band, sea surface temperature, wind speed, cloud water content, and various angle differences in water vapor inversion is quantitatively analyzed. Intuitive visualization charts are used to display the influence weight and action path of each feature, which not only ensures the objectivity and accuracy of data interpretation, but also provides strong data support for subsequent model optimization and parameter tuning. Step S43: comparing model performance and determining the best model; The traditional commercial spaceborne microwave radiometer water vapor inversion uses linear regression, and the formula is as follows: in, represents the coefficient, (i=1,2,3,…,9) corresponds to the brightness temperature data of nine channels: 6.6V, 6.6H, 10.7V, 10.7H, 18.7V, 18.7H, 23.8V, 37V and 37H; for the four channels of 6.6 GHz and 10.7 GHz (i=1,2,3,4), there are = ; For the remaining channels (i=5,6,7,8,9), there are = ; Based on a posteriori interpretations, the performance of various machine learning models on the test set was systematically compared. The fitting accuracy, prediction stability, and prediction accuracy of each model were evaluated using the coefficient of determination (R²), root mean square error (RMSE), and feature contribution balance performance indicators, combined with SHAP interpretation results. Finally, through comparative analysis, the optimal model was identified that outperformed traditional linear regression water vapor retrieval methods in terms of overall performance, providing the optimal solution for practical applications. Step S44: outputting results and promoting their application; The water vapor inversion results after screening by the best model will be output and transmitted to practical application platforms such as numerical weather forecasting, extreme weather forecasting and marine climate monitoring through standard interfaces; in order to ensure system performance, a regular evaluation and feedback mechanism is introduced, and through quality monitoring and performance backtracking of output results, model parameters and interpretation strategies are continuously optimized to ensure that the water vapor retrieval system maintains high precision and high robustness in actual operation, thereby providing stable and reliable data support and technical guarantees for related fields.

Citation Information

Patent Citations

  • High-precision sea surface temperature inversion method based on machine learning

    CN113408742A

  • PM2.5 concentration inversion and constraint fusion method based on multi-model integration

    CN118551675A

  • Method for Predicting Benchmark Value of Unit Equipment Based on XGBoost Algorithm and System thereof

    US20230213895A1

Cited By

  • Scanning microwave radiometer precipitation inversion method based on weight attention mechanism

    CN121305393A

  • Scanning microwave radiometer precipitation retrieval method based on weight attention mechanism

    CN121305393B

  • Pollutant tracing method and system based on intelligent fingerprint database matching

    CN121350782A

  • Ocean-land transition zone paleotopography inference method based on geochemical data and machine learning

    CN121457643A

  • Coastal sea temperature deep learning forecasting method and system based on multi-source data

    CN121502180A