Spaceborne GNSS-R Multi-System Multi-Polarization Sea Surface Wind Speed Estimation Method and System
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-08-14
AI Technical Summary
现有研究多基于单一GNSS系统(如GPS)或单极化接收(LHCP或RHCP),导致信号样本维度受限,未能充分利用多星座(C/G/R/E)与多极化(L/H/V)信息的互补特征,从而限制了模型的综合判别能力与适应性
第一、本发明通过引入高质量特征提取、ERA5环境因子匹配插值、质量控制筛选及多源数据融合,结合Optuna优化的机器学习模型,实现对不同风速尤其是高风速条件下的精确反演,提升了星载GNSS-R的海面风速估计能力与可靠性。本发明提供的Optuna优化过程通过构建非高斯超参数分布模型与概率比值采样准则,能够在高维非凸空间中高效逼近最优解。同时结合权重归一化、核密度估计与修剪机制,实现GNSS-R海面风速机器学习模型的自适应精调,从而在不同轨道几何与极化条件下获得更高的反演精度与模型泛化性能。
Smart Images

Figure CN121476637B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of spaceborne GNSS-R marine environmental monitoring technology, and particularly relates to a method and system for estimating sea surface wind speed using spaceborne GNSS-R multi-system multi-polarization. Background Technology
[0002] Sea surface wind speed is a key physical parameter in meteorological and oceanographic research, and is of great significance for typhoon monitoring, climate change research, marine engineering safety assessment, and analysis of air-sea interaction processes. Current methods for obtaining sea surface wind speed mainly include buoy measurements and active remote sensing methods such as scatterometers and synthetic aperture radar (SAR). However, buoy observations have limitations such as sparse spatial distribution, high equipment maintenance costs, and difficulty in long-term continuous observation. While active remote sensing methods can provide high spatial resolution data, their availability and timeliness under large-scale rapid revisits and extreme weather conditions are insufficient due to limitations in transmission power, incident angle, and polarization configuration, making it difficult to meet the continuous monitoring needs of global sea surface wind fields.
[0003] In recent years, Global Navigation Satellite System Reflectometry (GNSS-R), as an emerging passive microwave remote sensing technology, has become an important direction for sea surface wind speed inversion research due to its advantages such as all-weather, all-day coverage, global coverage, and low cost. GNSS-R receives the reflected echoes of navigation satellite signals on the sea surface and can invert sea surface state information using characteristic parameters such as reflected signal power and delay-Doppler distribution (DDM).
[0004] Based on the above analysis, the problems and shortcomings of the existing technology are as follows: (1) Limitations of single system and single polarization mode. Existing studies are mostly based on a single GNSS system (such as GPS) or single polarization receiver (LHCP or RHCP), which results in limited signal sample dimensionality and fails to fully utilize the complementary features of multi-constellation (C / G / R / E) and multi-polarization (L / H / V) information, thus limiting the model's comprehensive discrimination ability and adaptability.
[0005] (2) There is a systematic bias in the estimation of high wind speed. Traditional empirical models or single machine learning models perform well under low and medium wind speed conditions, but they generally underestimate the wind speed under high wind speed or extreme weather conditions (such as typhoons and strong convection). This is mainly due to the low proportion of high wind speed samples in the training data. The uneven distribution of samples leads to insufficient learning of the model in this range, resulting in a decrease in generalization performance.
[0006] (3) The multi-source information fusion strategy is not perfect. In theory, multiple GNSS systems and multiple polarization modes can provide richer scattering geometry and signal response information. However, existing studies mostly use simple weighted averaging or splicing features for fusion, ignoring the nonlinear relationship and spatiotemporal correlation of signal features between different systems. This results in limited fusion effect and fails to fully leverage the potential advantages of multi-source information.
[0007] (4) Insufficient model optimization and hyperparameter search mechanisms. Some studies use fixed hyperparameters or manual adjustment based on experience, failing to find the optimal parameter combination through systematic optimization (such as Bayesian optimization and cross-validation), thus affecting the generalization accuracy and stability of the model in different regions and sea conditions.
[0008] In summary, the main defects and causes of existing technologies include: The current technology suffers from several shortcomings. Firstly, the limited observation dimensions and insufficient information utilization. Secondly, the unbalanced distribution of training samples leads to inadequate model fitting capabilities in high-wind-speed ranges. Thirdly, the imperfect information fusion and model optimization strategies negatively impact model performance and robustness. Therefore, to address the issues of insufficient accuracy and adaptability to high-wind-speed conditions in sea surface wind speed estimation under multi-GNSS system and multi-polarization observation conditions, it is urgently necessary to propose a spaceborne GNSS-R sea surface wind speed estimation method and system that can comprehensively utilize multi-system and multi-polarization observation information while considering sample distribution balance and model optimization mechanisms. This would overcome existing technological bottlenecks and improve the accuracy and reliability of sea surface wind speed estimation.
[0009] To overcome the problems of insufficient accuracy and poor adaptability of traditional single-system, single-polarization model, the present invention discloses a spaceborne GNSS-R multi-system multi-polarization sea surface wind speed estimation method and system, particularly a spaceborne GNSS-R sea surface wind speed estimation method and system based on the fusion of multiple GNSS systems and multi-polarization models. This method and system can achieve balanced learning and feature optimization fusion of high wind speed samples under multi-source observation conditions, thereby significantly improving the accuracy and stability of wind speed estimation. Summary of the Invention
[0010] This invention is implemented as follows: A spaceborne GNSS-R multi-system multi-polarization sea surface wind speed estimation method includes the following steps: S1, extract the feature value data of spaceborne GNSS-R and the environmental factor data of ERA5; S2 performs quality control on the datasets of multiple GNSS systems and multiple polarization modes of the spaceborne GNSS-R constellation, and performs grid matching interpolation with ERA5 data; S3, based on multiple GNSS systems and multiple polarization modes, divides the training set and test set, and constructs a sea surface wind speed estimation model based on the Optuna optimized machine learning model; S4, increase the proportion of high wind speed dataset in the model, and determine the optimal proportion of high wind speed and the optimal machine learning model for different GNSS systems and different polarization modes. S5 performs multi-GNSS system fusion and multi-polarization mode fusion respectively, and compares it with pure data fusion to determine the best sea surface wind speed estimation model, and outputs sea surface wind speed estimates based on orbit points.
[0011] In step S1, the characteristic data of the spaceborne GNSS-R includes: Spatial positioning parameters, including the latitude and longitude of the mirror point and the incident angle of the mirror point, are used to characterize the reflection geometry. Time and observation information, including DDM generation time, number of channels, satellite number, and surface type; Parameters related to the energy distribution of the DDM include the leading edge slope at the mirror point, the normalized scattering cross section, the raw peak power count, the peak power signal-to-noise ratio, the Doppler column where the peak is located, and the time delay row. The fourth-order moment characteristic parameters of the DDM include kurtosis and skewness, which characterize the statistical properties of the power distribution; the reflection coefficient, which represents the reflection intensity at the mirror point; and the quality label, which provides a marker for the reliability of DDM observations.
[0012] In step S1, the environmental factor data of ERA5 includes: total precipitation, sea surface temperature, significant wave height, 10-meter sea surface wind speed, UTC time, and latitude and longitude; wherein, the 10-meter sea surface wind speed is synthesized from the 10-meter u wind component and v wind component.
[0013] In step S2, the quality control uses Ddm_quality_flag to filter data and remove low-quality data; wherein, Ddm_quality_flag contains multiple binary bits, which are used to mark the position of low-quality data as 1 and record quality problems in different dimensions; The grid matching interpolation includes: spatially using bilinear interpolation to map regular grid ERA5 data to irregular GNSS-R observation point locations, and temporally using linear interpolation to achieve the correspondence between observation time and ERA5 reanalysis time.
[0014] In step S3, the multi-GNSS system includes BDS, GPS, GLONASS and Galileo, and the multi-polarization mode includes LHCP, RHCP, H and V; Based on the correspondence between the multi-GNSS system and the polarization mode configurations of different satellites in the TM-1 constellation, the observation data is divided into 20 data configurations. For each type of data, training and test sets were configured, and CatBoost, LightGBM and XGBoost models optimized based on Optuna were built respectively, for a total of 60 sea surface wind speed estimation models.
[0015] In step S3, the machine learning model adaptively adjusts the parameters of the gradient boosting tree model through the Optuna hyperparameter optimization framework to minimize the weighted mean square error based on cross-validation. The input sample set is: ; In the formula, For the sample set, containing One sample, The total number of samples; For the first The GNSS-R signal feature vector of each sample includes reflected signal intensity, scattering geometry parameters, satellite elevation angle, and polarization information; For the first The actual sea surface wind speed of each sample; These are sample weights, used to balance the importance of different samples. ; Machine learning models are defined as follows: ; In the formula, For the model to the first The predicted value for each sample, This is the model mapping function, used to map input features to output predicted values; For model training parameters, The set of hyperparameters to be optimized includes learning rate, maximum depth, number of leaf nodes, regularization parameter, and sampling ratio; The Optuna optimization process uses the cross-validation weighted mean square error as the objective function, defined as: ; In the formula, The weighted mean squared error for cross-validation is used to measure the model's performance given hyperparameters. The overall verification error is as follows; The number of folds for cross-validation, i.e., dividing the sample into folds. A set of non-overlapping verification subsets; For the first The set of validation samples for the fold; In the first The trained model is then combined using hyperparameters. Below, on the sample The predicted value; For sample index, ; Optuna's optimization objective is: ; In the formula, The optimal combination of hyperparameters is the one that minimizes the cross-validation error. ; The hyperparameter search space is the set of parameters that Optuna samples and evaluates. To verify the square root form of the error, and to ensure that the objective function is consistent with the RMSE dimension, making the optimization results more intuitive; This is a parameter combination operator that minimizes the objective function.
[0016] Furthermore, the optimization process includes: (1) Hyperparameter space setting: Define the parameter range according to the model type; where: learning rate controls the step size of each iteration, The maximum depth of a single decision tree determines the model's complexity and generalization ability. The number of base learners, used to balance bias and variance. ; (2) Trial generation and training: Optuna generates a set of hyperparameters based on a random sampling strategy according to historical trial results. And use the parameters to train the model; (3) Cross-validation and loss calculation: K-fold cross-validation is used to calculate the validation error of each fold. And take the average to get ; (4) Pruning mechanism: Optuna automatically terminates poorly performing experiments based on intermediate evaluation results to save computational resources; (5) Determination of optimal parameters and model training: After a preset number of trials T, the optimal hyperparameters are determined. The final model was obtained by retraining based on all the training data. ; By optimizing the process, the obtained machine learning model can achieve high-precision sea surface wind speed estimation from GNSS-R data, and the output results are as follows: ; In the formula, This is the estimated sea surface wind speed obtained through inversion.
[0017] In step S4, increasing the proportion of the high-wind-speed dataset in the model includes: For observation data from different GNSS systems and under different polarization modes, the proportion of high wind speed sample data was increased sequentially from 1 to 10 times the original sample size. The training and test sets under each proportion condition were modeled and validated using Optuna-optimized machine learning models, resulting in a total of 600 models. Among them, the high wind speed samples are those with sea surface wind speeds greater than or equal to 10 m / s.
[0018] In step S5, the optimal sea surface wind speed estimation model is determined based on the model's performance metrics on the test set. The performance metrics include correlation coefficient R, root mean square error RMSE, mean error ME, and mean absolute percentage error MAPE. The multi-GNSS system fusion includes merging the prediction results of the optimal models of BDS, GPS, GLONASS and Galileo systems, and the multi-polarization mode fusion includes merging the prediction results of the optimal models of LHCP, RHCP, H and V polarization modes. The pure data fusion refers to all data fusion models that do not distinguish between GNSS systems and polarization modes.
[0019] Another objective of this invention is to provide a spaceborne GNSS-R multi-system multi-polarization sea surface wind speed estimation system, which is used to implement the method described above, including: The data acquisition module is used to extract the feature values of the spaceborne GNSS-R and the environmental factor data of ERA5; The data quality control and matching module is used to perform quality control on the datasets of multiple GNSS systems and multiple polarization modes of the spaceborne GNSS-R constellation, and to perform grid matching interpolation with ERA5 data; The model training and optimization module is used to divide the training set and test set based on multiple GNSS systems and multiple polarization modes, and to build a sea surface wind speed estimation model based on the Optuna optimized machine learning model. The high-wind-speed sample weighting module is used to increase the proportion of high-wind-speed datasets in the model, and to determine the optimal proportion of high-wind-speed data and the optimal machine learning model for different GNSS systems and different polarization modes. The fusion and inversion module is used to perform multi-GNSS system fusion, multi-polarization mode fusion and compare with pure data fusion to determine the best sea surface wind speed estimation model and output sea surface wind speed estimates based on orbit points.
[0020] Combining all the above technical solutions, the beneficial effects of this invention are as follows: First, this invention improves the sea surface wind speed estimation capability and reliability of spaceborne GNSS-R by introducing high-quality feature extraction, ERA5 environmental factor matching interpolation, quality control screening, and multi-source data fusion, combined with an Optuna-optimized machine learning model. This achieves accurate inversion under different wind speeds, especially high wind speeds. The Optuna optimization process provided by this invention constructs a non-Gaussian hyperparameter distribution model and a probability ratio sampling criterion, enabling efficient approximation of the optimal solution in a high-dimensional non-convex space. Simultaneously, by combining weight normalization, kernel density estimation, and pruning mechanisms, adaptive fine-tuning of the GNSS-R sea surface wind speed machine learning model is achieved, resulting in higher inversion accuracy and model generalization performance under different orbital geometries and polarizations.
[0021] Secondly, the advantages of multi-source data fusion significantly improve inversion accuracy: By fusing observation data from multiple GNSS systems such as BDS, GPS, GLONASS, and Galileo, as well as multiple polarization modes such as LHCP, RHCP, H, and V, it overcomes the information limitations of traditional single-system, single-polarization modes. It fully utilizes the complementary nature of different constellation signal characteristics and polarization observations to construct a fused dataset containing 20 data configurations, significantly increasing the diversity of effective observation samples. Combined with Optuna-optimized machine learning models, it achieves deep mining of the characteristics of multi-source heterogeneous data. Compared with existing methods, the correlation coefficient R is significantly improved across the entire wind speed range, especially under high wind speed conditions, and the root mean square error (RMSE) is significantly reduced, effectively solving the industry problem of wind speed underestimation under extreme weather conditions.
[0022] Data quality control and environmental factor matching techniques enhance model robustness: A Ddm_quality_flag screening mechanism with multi-dimensional quality indicators is introduced to selectively remove low-quality data such as noise drift, direct signal interference, and geometric parameter anomalies, ensuring the reliability of input features; a spatiotemporal dual-dimensional interpolation method is used to achieve accurate matching between GNSS-R orbit points and ERA5 reanalysis data, controlling the spatiotemporal matching error of 10m wind field environmental factors such as sea surface temperature, significant wave height, and precipitation to within 30 minutes / 10km, providing a high-precision multi-source collaborative data foundation for model training, and significantly improving the model's anti-interference ability and generalization performance under complex sea conditions.
[0023] A dynamic optimization strategy for high-wind-speed samples overcomes the bottleneck of data imbalance: To address the distribution imbalance caused by insufficient high-wind-speed sample ratio (usually <15%) in traditional model training, a 1-10x dynamic oversampling strategy is proposed. By constructing 600 sets of comparative model matrices, the optimal high-wind-speed ratio for each GNSS system and polarization mode is accurately determined, effectively improving the inversion bias under extreme scenarios such as typhoons and severe convection, and filling the technical gap in high-precision monitoring of high-wind-speed segments.
[0024] A multi-strategy fusion framework enables adaptive optimization of technical solutions: a three-layer comparative verification system is established, encompassing multi-GNSS system fusion, multi-polarization mode fusion, and pure data fusion. The optimal model combination is dynamically selected using performance index matrices (R, RMSE, ME, MAPE). Compared to a single data configuration scheme, the multi-strategy fusion model demonstrates improved consistency in inversion accuracy across marine areas, particularly in complex sea states. This forms an adaptive optimization mechanism covering different observation conditions, providing a reusable technical paradigm for operational marine environmental monitoring.
[0025] Third, the GNSS-R sea surface wind speed inversion method based on Optuna optimization proposed in this invention can significantly improve the accuracy and stability of spaceborne GNSS-R wind speed inversion, especially under medium-to-high wind speeds and extreme sea state conditions, outperforming traditional empirical models and single machine learning algorithms. This method can be widely applied in marine meteorological monitoring, typhoon warning, maritime shipping, and engineering safety assessment, demonstrating good engineering feasibility and promising prospects for widespread adoption. By fusing data with existing satellite constellations (such as TM-1 and CYGNSS), it can achieve large-scale, low-cost, all-weather wind speed product production, resulting in significant economic and social benefits. Furthermore, this technical solution possesses strong algorithmic versatility and scalability, and can be ported to other GNSS-R inversion tasks such as sea surface wave height and sea ice monitoring, further enhancing the overall value of the GNSS-R application system.
[0026] Currently, there is a lack of efficient and automated optimization modeling methods for multi-system, multi-polarization observation data of spaceborne GNSS-R, both domestically and internationally. Traditional methods often rely on manually tuned machine learning models, which suffer from problems such as strong subjectivity in parameter selection, low model convergence efficiency, and poor adaptability to complex nonlinear features. This invention introduces the Optuna adaptive hyperparameter optimization algorithm into the field of GNSS-R wind speed inversion for the first time, realizing fully automated search and optimal configuration of model structure and parameters, and constructing a wind speed inversion model that combines high accuracy and high robustness.
[0027] For a long time, GNSS-R sea surface wind speed inversion has suffered from problems such as scarce high-wind-speed samples, insufficient model generalization ability, and low efficiency in feature fusion between different GNSS systems. Traditional empirical models or single-algorithm learning frameworks are unable to handle the complex relationships between multiple systems, multiple polarizations, and multiple scales of features simultaneously, resulting in limited wind speed inversion accuracy, especially significant errors under extreme wind conditions. This invention introduces the Optuna optimization mechanism to achieve dynamic tuning and adaptive fusion of multiple models such as CatBoost, LightGBM, and XGBoost, effectively solving the problems of model underfitting and sample imbalance, thus maintaining excellent inversion stability and accuracy under high wind speed conditions.
[0028] Traditional GNSS-R research often considers model performance primarily dependent on feature selection and physical modeling accuracy, neglecting the importance of hyperparameter space search and fusion optimization, leading to a technical bias of "fixed parameters, local optimization." This invention breaks through this traditional approach, proposing a new paradigm of "multi-system multi-polarization fusion + Optuna global optimization," fully exploring the potential of multi-source GNSS-R observation features and significantly improving the model's adaptability and generalization ability under different wind speed ranges and sea area conditions. This method provides new theoretical support and technical pathways for intelligent and automated GNSS-R inversion, possessing significant innovation and widespread application value. Attached Figure Description
[0029] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure; Figure 1 This is a flowchart of the spaceborne GNSS-R multi-system multi-polarization sea surface wind speed estimation method provided in this embodiment of the invention; Figure 2 This invention provides a comparison chart of the RMSE of sea surface wind speed test sets for 20 data types × 10 wind speed weights × 3 machine learning models for orbit point inversion in the Yellow and Bohai Seas of China, along with its optimal solution. Figure 3 This is a scatter plot of sea surface wind speed in the Yellow and Bohai Seas of China, provided by an embodiment of the present invention, which is a fusion of multiple GNSS systems, multiple polarization modes, and pure data in the Yellow and Bohai Seas region of China. Detailed Implementation
[0030] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0031] The innovation of this invention lies in: (1) This invention proposes an intelligent parameter search mechanism based on the Optuna optimization framework. This method combines the Optuna optimization algorithm with machine learning regression models (such as CatBoost, LightGBM and XGBoost) to automatically search for the optimal hyperparameter combination, effectively avoiding the uncertainty and inefficiency of manual parameter tuning, and significantly improving the generalization performance of the model under different sea conditions and multi-latitude regions.
[0032] (2) This invention introduces a multi-system, multi-polarization GNSS-R observation data fusion strategy. By comprehensively utilizing multiple GNSS systems such as GPS, Galileo, BDS, and GLONASS, as well as the signal characteristics of left and right circular polarization (RHCP / LHCP), a more comprehensive expression of sea surface scattering information is achieved, improving the response sensitivity and stability to complex wind and wave fields.
[0033] (3) This invention designs an input feature construction scheme for noise resistance and feature enhancement. During the model training stage, multi-dimensional input features such as normalized scattering intensity (NBRCS), delayed-Doppler moment (DDM Moments), signal-to-noise ratio (SNR), incident angle, satellite elevation angle, and observation geometric parameters are introduced and dynamically filtered in combination with feature importance ranking mechanism, thereby effectively suppressing noise interference and enhancing the physical interpretability of the model.
[0034] (4) This invention constructs a comprehensive optimization objective function based on cross-validation and multi-index constraints. In the Optuna objective function, this invention simultaneously considers performance indicators such as root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²), and achieves global optimization of model performance through weight combination, so that the final inversion result achieves the best balance between accuracy and stability.
[0035] S1, extract the feature value data of spaceborne GNSS-R and the environmental factor data of ERA5; S2 performs quality control on the datasets of multiple GNSS systems and multiple polarization modes of the spaceborne GNSS-R constellation, and performs grid matching interpolation with ERA5 data; The multiple GNSS systems are BDS, GPS, GLONASS, and Galileo, and the multiple polarization modes are LHCP, RHCP, H, and V. S3, based on multiple GNSS systems and multiple polarization modes, divides the training set and test set, and constructs a sea surface wind speed estimation model based on the Optuna optimized machine learning model; Different machine learning models include CatBoost, LightGBM, and XGBoost; S4, increase the proportion of high wind speed dataset in the model, and determine the optimal proportion of high wind speed and the optimal machine learning model for different GNSS systems and different polarization modes. S5 performs multi-GNSS system fusion and multi-polarization mode fusion respectively, and compares it with pure data fusion to determine the best sea surface wind speed estimation model, and outputs sea surface wind speed estimates based on orbit points.
[0036] In step S1, the feature value data extracted from the Tianmu-1 (TM-1) constellation's onboard GNSS-R includes: Spatial positioning parameters include: mirror point latitude (Sp_lat), longitude (Sp_lon), and mirror point incident angle (Sp_inc_angle), which are used to characterize the reflection geometry. Time and observation information includes: DDM generation time (Ddm_time_utc), number of channels (TM_TD), satellite number (TM_PRN), and surface type (Sp_surface_type). Parameters related to the energy distribution of the DDM include: the leading edge slope (Ddm_sp_les) at the mirror point, the normalized scattering cross section (Ddm_sp_nbrcs), the raw peak power count (Ddm_peak_raw), the peak power signal-to-noise ratio (Ddm_peak_snr), the Doppler column (Ddm_peak_column) where the peak is located, and the time delay row (Ddm_peak_row). The fourth-order moment characteristic parameters of the DDM as a whole include kurtosis and skewness, which are used to characterize the statistical properties of the power distribution.
[0037] In addition, the reflection coefficient (Ddm_sp_reflectivity) is used to represent the reflection intensity at the mirror point, while the quality flag (Ddm_quality_flag) provides a marker for the reliability of DDM observations; In step S1, the environmental factor data of ERA5 includes: total precipitation (TP), sea surface temperature (SST), significant wave height (SWH), wind speed (WS) synthesized from the 10m u wind component and the 10m v wind component, as well as UTC time and latitude and longitude.
[0038] In step S2, the `Ddm_quality_flag` is introduced for data filtering to ensure the reliability of the input data. This flag contains multiple binary bits to record quality issues in different dimensions: if the overall quality is poor, the noise floor is unstable, there is direct signal interference, the geometric estimation of mirror points is unreliable, satellite EIRP information is missing, the scattering matrix has a negative value, or the effective region is invalid, the corresponding flag position is 1, indicating that the data should be removed. Through this quality control mechanism, the impact of noise drift, direct interference, and geometric uncertainty on the accuracy of soil moisture inversion can be effectively avoided, ensuring the robustness and reliability of the feature input.
[0039] The interpolation method for grid matching of multi-GNSS system and multi-polarization mode datasets and ERA5 data in spaceborne GNSS-R constellations is as follows: At the spatial level, bilinear interpolation is used to map regular grid ERA5 data to irregular GNSS-R observation point locations; at the temporal level, linear interpolation is used to achieve the correspondence between observation time and ERA5 reanalysis time.
[0040] In step S3, the pure data configuration of the TM-1 constellation is as follows: all orbital point data can be classified into 1 type of data; The TM-1 constellation has the following multi-GNSS system configuration: it can receive data from BDS, GPS, GLONASS, and Galileo satellite systems, which can be classified into four types of data. The multi-polarization configuration of the TM-1 constellation is as follows: TM01-TM10 all use left-hand circular circular polarization (LHCP), with observation channels C1-C8, corresponding to a defined polarization marker of 0, and experimental scheme number S1. TM11 has multiple polarization modes: channels C1, C3, C5, and C7 use horizontal polarization (H), marker 2, and experimental scheme number S2; channels C2, C4, C6, and C8 use vertical polarization (V), marker 4, and experimental scheme number S3. In contrast, TM12-TM22 are configured with dual polarization: channels C1, C3, C5, and C7 use LHCP, marker 0 (scheme S4); channels C2, C4, C6, and C8 use right-hand circular circular polarization (RHCP), marker 1 (scheme S5). Among them, the polarization modes of GPS, GLONASS, and Galileo correspond to the five types mentioned above, while GLONASS data is all LHCP data with polarization modes of TM01-TM10. Therefore, the experimental scheme does not distinguish between multiple polarizations and can be divided into 15 types of data. Data from different GNSS systems includes the following four categories: BDS, GPS, GLONASS, and Galileo; According to Table 2, the data for different polarization modes can be identified as including the following 15 categories: BDS-S1, GPS-S1, Galileo-S1, BDS-S2, GPS-S2, Galileo-S2, BDS-S3, GPS-S3, Galileo-S3, BDS-S4, GPS-S4, Galileo-S4, BDS-S5, GPS-S5, Galileo-S5; There is also a category where all data is fused together: ALL, which is a pure data fusion model; In summary, there are a total of 20 categories of data; For example, training three algorithms (CatBoost, LightGBM, and XGBoost) on BDS-S1 data can yield the best training model, such as LightGBM. To address the issue of insufficient high-wind-speed data leading to model distortion, we perform high-wind-speed oversampling. This involves multiplying the sea surface wind speeds greater than 10 meters per second in the BDS-S1 dataset by a factor of 1, 2, up to 10, thus expanding the high-wind-speed dataset. Then, we train the LightGBM algorithm again, finding that LightGBM is the optimal algorithm for BDS-S1 data. The highest accuracy in sea surface wind speed retrieval is achieved when the optimal high-wind-speed weight is 10 times.
[0041] Based on the above 20 types of data, 75% of the data was used as the training set and 25% as the test set. Different machine learning models (CatBoost, LightGBM, XGBoost) based on Optuna optimization were constructed, resulting in a total of 60 models.
[0042] In step S4, the proportion of high wind speed dataset in the model is increased, specifically as follows: For observation data from different GNSS systems and under different polarization modes, during the construction of machine learning models, the proportion of high-wind-speed sample data was increased, with the proportion of high-wind-speed sample data increasing sequentially from 1 to 10 times the original sample size. The training and test sets under each proportion condition were modeled and validated using Optuna-optimized machine learning models (CatBoost, LightGBM, XGBoost), resulting in a total of 600 models.
[0043] In step S5, the sea surface wind speed estimation models based on different machine learning models (CatBoost, LightGBM, XGBoost, etc.) optimized by Optuna include: The Optuna hyperparameter optimization framework is used to adaptively adjust the parameters of gradient boosting tree models to minimize the weighted mean square error (WMSE) based on cross-validation and improve the accuracy of GNSS-R sea surface wind (WS) retrieval. The specific implementation is as follows: The input sample set is: ; In the formula, For the sample set, containing One sample, The total number of samples; For the first The GNSS-R signal feature vector of each sample includes reflected signal intensity, scattering geometry parameters, satellite elevation angle, and polarization information; For the first The actual sea surface wind speed of each sample; These are sample weights, used to balance the importance of different samples. It is determined based on the signal-to-noise ratio or measurement uncertainty.
[0044] Machine learning models are defined as follows: ; In the formula, For the model to the first The predicted value for each sample, This is the model mapping function, used to map input features to output predicted values; For model training parameters, The set of hyperparameters to be optimized includes learning rate, maximum depth, number of leaf nodes, regularization parameters reg_alpha and reg_lambda, and sampling ratio subsample, etc. The Optuna optimization process uses the cross-validation weighted mean square error as the objective function, defined as follows: ; In the formula, The weighted mean squared error for cross-validation is used to measure the model's performance given hyperparameters. The overall verification error is as follows; The number of folds for cross-validation, i.e., dividing the sample into folds. A set of non-overlapping verification subsets; For the first The set of validation samples for the fold; In the first The trained model is then combined using hyperparameters. Below, on the sample The predicted value; For sample index, ; Optuna's optimization objective is: ; In the formula, The optimal combination of hyperparameters is the one that minimizes the cross-validation error. ; The hyperparameter search space is the set of parameters that Optuna samples and evaluates. To verify the square root form of the error, and to ensure that the objective function is consistent with the RMSE dimension, making the optimization results more intuitive; This is a parameter combination operator that minimizes the objective function.
[0045] The optimization process includes: (1) Hyperparameter space setting: Define the parameter range according to the model type; where: learning rate controls the step size of each iteration, The maximum depth of a single decision tree determines the model's complexity and generalization ability. The number of base learners, used to balance bias and variance. ; (2) Trial generation and training: Optuna generates a set of hyperparameters based on a random sampling strategy according to historical trial results. And use the parameters to train the model; (3) Cross-validation and loss calculation: K-fold cross-validation is used to calculate the validation error of each fold. And take the average to get ; (4) Pruning mechanism: Optuna automatically terminates poorly performing experiments based on intermediate evaluation results to save computational resources; (5) Determination of optimal parameters and model training: After a preset number of trials T, the optimal hyperparameters are determined. The final model was obtained by retraining based on all the training data. ; By optimizing the process, the obtained machine learning model can achieve high-precision sea surface wind speed estimation from GNSS-R data, and the output results are as follows: ; In the formula, This is the estimated sea surface wind speed obtained through inversion.
[0046] By introducing Optuna's adaptive hyperparameter optimization mechanism, the generalization performance and inversion accuracy of the model can be significantly improved while maintaining the stability of the model structure. This reduces the uncertainty caused by manual parameter tuning and enables unified and efficient optimization of models such as CatBoost, LightGBM, and XGBoost. It is suitable for sea surface wind speed inversion under multiple GNSS systems, multiple polarization modes, and multiple orbit observation conditions.
[0047] Furthermore, the orbital fusion of the TM-1 constellation with sea surface wind speed products is generated, specifically as follows: Based on modeling of a single GNSS system and a single polarization mode, the optimal sea surface wind speed estimation model is determined based on the model's performance indicators on the test set (including correlation coefficient R, root mean square error RMSE, mean error ME, and mean absolute percentage error MAPE). The fusion of multi-GNSS system data and multi-polarization mode data are then carried out respectively. The results of multi-GNSS fusion, multi-polarization mode fusion, and pure data fusion are compared and analyzed. The determined optimal estimation model is used to extrapolate the orbit points point by point to obtain the sea surface wind speed estimate based on the orbit points, and the spatial distribution of sea surface wind speed in the study area is output.
[0048] Another objective of this invention is to provide a spaceborne GNSS-R multi-system multi-polarization sea surface wind speed estimation system. This system is used to implement the aforementioned spaceborne GNSS-R multi-system multi-polarization sea surface wind speed estimation method. The system includes: The data acquisition module is used to extract the feature values of the spaceborne GNSS-R and the environmental factor data of ERA5; The data quality control and matching module is used to perform quality control on the datasets of multiple GNSS systems and multiple polarization modes of the spaceborne GNSS-R constellation, and to perform grid matching interpolation with ERA5 data; The model training and optimization module is used to divide the training set and test set based on multiple GNSS systems and multiple polarization modes, and to build a sea surface wind speed estimation model based on the Optuna optimized machine learning model. The high-wind-speed sample weighting module is used to increase the proportion of high-wind-speed datasets in the model, and to determine the optimal proportion of high-wind-speed data and the optimal machine learning model for different GNSS systems and different polarization modes. The fusion and inversion module is used to perform multi-GNSS system fusion, multi-polarization mode fusion and compare with pure data fusion to determine the best sea surface wind speed estimation model and output sea surface wind speed estimates based on orbit points.
[0049] Example 1: This embodiment of the invention uses the Yellow and Bohai Seas of China (34°-41°N, 118°-127°E) as the study area. TM-1 satellite orbital observation data from February 3rd to March 31st, 2024, are selected and combined with ERA5 reanalysis data to verify the effectiveness of the proposed sea surface wind speed estimation method based on the fusion of multiple GNSS systems and multi-polarization models. Figure 1 As shown, the spaceborne GNSS-R multi-system multi-polarization sea surface wind speed estimation method provided in this embodiment of the invention includes the following steps: Eigenvalue data and ERA5 environmental factor data extraction; 1.1 Extraction of GNSS-R feature value data from spaceborne systems; The GNSS-R payload carried by the TM-1 satellite uses an omnidirectional antenna to receive direct and reflected signals from four major GNSS systems: BDS, GPS, GLONASS, and Galileo, generating a DDM with a spatial resolution of 6km × 1km (along the track × through the track). The following characteristic parameters are extracted: Spatial positioning parameters: The geographic coordinates (Sp_lat, Sp_lon) and incident angle (Sp_inc_angle) of the specular point are obtained through the onboard GPS positioning module and used to characterize the reflection geometry. Time and observation information: DDM generation time (Ddm_time_utc), number of channels (TM_TD), satellite number (TM_PRN), surface type (Sp_surface_type=ocean); DDM energy distribution parameters: front slope (Ddm_sp_les), normalized scattering cross section (Ddm_sp_nbrcs), peak power raw count (Ddm_peak_raw), peak power signal-to-noise ratio (Ddm_peak_snr); peak position (Ddm_peak_column, Ddm_peak_row). The fourth-order moment parameters of the DDM statistical characteristics are: kurtosis (Ddm_kurtosis, reflecting the sharpness of the energy distribution) and skewness (Ddm_skewness, reflecting the symmetry of the energy distribution). Other parameters: reflection coefficient (Ddm_sp_reflectivity), quality flag (Ddm_quality_flag, all binary bits are 0 to indicate reliable data quality, no noise interference, direct signal or geometric error).
[0050] 1.2 ERA5 Environmental Factor Data Extraction and Matching; ERA5 reanalysis data (temporal resolution: 1 hour, spatial resolution: 0.25° × 0.25°) were obtained from the European Centre for Medium-Range Weather Forecasts (ECMWF), and the following environmental factors were extracted: Sea surface temperature (SST), significant wave height (SWH), 10m wind speed (WS, as the true value); total precipitation (TP), timestamp (ERA5_time), and latitude and longitude grid (Lat_grid, Lon_grid).
[0051] Data quality control and grid matching interpolation 2.1 Quality control screening; Based on the Ddm_quality_flag provided by the TM-1 payload, the raw data is filtered, and the binary encoding rules are shown in Table 1.
[0052] Table 1. Quality flags used for data filtering in TM-1 Ddm_quality_flag
[0053] 2.2 Multi-source data grid matching; ERA5 grid data was matched to GNSS-R observation points after quality control through spatiotemporal interpolation; Spatial interpolation: The four-point values of the ERA5 grid points are weighted and averaged using the bilinear interpolation method to obtain the ERA5 data at the GNSS-R specular reflection point.
[0054] Because ERA5 data uses a regular spatiotemporal grid, while TM-1 observations use irregular spatiotemporal points, there is a difference in spatiotemporal resolution. To achieve matching between ERA5 wind speed data and GNSS-R observation points, this invention employs spatial bilinear interpolation and temporal linear interpolation. Spatially, bilinear interpolation maps the regular grid ERA5 data to the irregular GNSS-R observation point locations; temporally, linear interpolation is used to correspond the observation time with the ERA5 reanalysis time. (Target point...) For example, its spatial interpolation process can be expressed as: ; in, , , , The four nearest neighbor grid points around the target point. This represents the wind speed value at the corresponding grid point; Through spatial interpolation, a corresponding wind speed time series can be obtained for each GNSS-R observation location; Time interpolation: Linear interpolation is used to achieve the correspondence between GNSS-R observation time and ERA5 reanalysis time; Subsequently, time interpolation methods were used to further match the sequence to the precise observation times of the GNSS-R: ; in, , For GNSS-R observation time, and The spatial interpolated wind speed for adjacent ERA5 times.
[0055] Model building and optimization; 3.1 Multi-source data configuration and dataset partitioning; Data classification: Based on the combination of GNSS system (BDS / GPS / GLONASS / Galileo) and polarization mode (LHCP / RHCP / H / V), the observation data of the 22 satellites of the TM-1 constellation are divided into 20 categories. The classification scheme is based on Table 2.
[0056] Table 2 Polarization of TM-1 constellation satellites (LHCP: left-hand circular polarization, RHCP: right-hand circular polarization, H: horizontal polarization, V: vertical polarization)
[0057] 3.2 Optuna Optimization and Model Training Parameter search space: The Optuna hyperparameter adaptive optimization algorithm improves the robustness and fitting accuracy of wind speed inversion. The core of Optuna's optimization lies in establishing a hyperparameter probability model. ,in, This is a set of historical experimental results. The objective is to minimize the weighted cross-validation loss function. ; In the formula, It is the cross-validation fold number; It is the first Fold the validation set; This is a Gaussian prior regularization term used to constrain the smoothness of the hyperparameter distribution; It is the penalty coefficient, usually 1. .
[0058] The final optimal solution satisfies: ; Optuna in every trial At that time, the historical hyperparameter-target value pair Divided into two groups: ; in, To be based on quantiles (A performance threshold, typically 0.2, is used.) Two nonparametric distribution estimates are constructed: ; The TPE sampling rule selects new hyperparameters by maximizing the following ratio: ; when The larger the ratio, the more likely the hyperparameter is to produce a lower loss value. Distribution and It is usually obtained by kernel density estimation (KDE): ; ; in, The normalization constant is Let be the covariance matrix for each group.
[0059] To accelerate the search process, Optuna introduces a pruning criterion in each trial: if in the first trial... During round iteration, the current intermediate index satisfy: ; in, This is the historical median performance. If the threshold is dynamic, the experiment will be terminated early.
[0060] Let the total number of trials be The overall optimization process can then be formalized as follows: ; in, The penalty coefficient is... For indicator functions, This represents the number of iterations in the trial.
[0061] To further suppress the influence of outliers, this invention introduces a generalized loss function based on power-law weighting: ; in, Control the weighted nonlinear intensity; when It degenerates into ordinary weighted RMSE. This form exhibits stronger robustness in noise-sensitive GNSS-R wind speed inversion.
[0062] After multiple rounds of searching, the optimal parameter combination was obtained. And update the model parameters based on all training data. : ; The final sea surface wind speed estimation model is as follows: ; Its uncertainty can be calculated using model integration or quantile loss: ; The Optuna optimization process, by constructing a non-Gaussian hyperparameter distribution model and a probability ratio sampling criterion, can efficiently approximate the optimal solution in a high-dimensional non-convex space. Simultaneously, by combining weight normalization, kernel density estimation, and pruning mechanisms, it achieves adaptive fine-tuning of the GNSS-R sea surface wind speed machine learning model, thereby obtaining higher inversion accuracy and model generalization performance under different orbital geometries and polarization conditions.
[0063] Taking the LightGBM model as an example, five hyperparameters were set, including learning rate [0.01, 0.3], tree depth [3, 12], and number of leaf nodes [10, 200]. Optuna was used for 50 iterations of optimization, with the objective function being the root mean square error (RMSE) of the test set. After 50 iterations, Optuna automatically converged to the optimal hyperparameter combination. With this parameter configuration, the model's weighted RMSE on the validation set decreased by approximately 18.7%, and the coefficient of determination... Improved by 0.09.
[0064] For CatBoost and XGBoost models, CatBoost adds an L2 regularization parameter. Boundary division number XGBoost adds regularization terms , and minimum splitting loss For example, after 60 iterations of XGBoost, the optimal combination of hyperparameters is obtained. Its RMSE on the test set is 1.82m / s, which is about 21% lower than that of the manually tuned model.
[0065] Ultimately based on optimal hyperparameters Retrain all model parameters The sea surface wind speed inversion function is obtained as follows: ; It also outputs the predicted wind speed distribution for each orbital point. Independent validation results on GNSS-R multi-system, multi-polarization observation datasets show that, using the Optuna optimization method described in this invention, the correlation coefficient R of the model prediction is improved by an average of 0.06-0.11, and the root mean square error (RMSE) is reduced by 15-25%, demonstrating significant advantages in accuracy and stability.
[0066] Model matrix construction: For 20 types of data × 3 algorithms (CatBoost / LightGBM / XGBoost), a total of 60 training tasks were generated, and parallel training was carried out using a GPU cluster, with a single model training time of ≤30 minutes.
[0067] Data from different GNSS systems includes the following four categories: BDS, GPS, GLONASS, and Galileo; According to Table 2, the data for different polarization modes can be identified as including the following 15 categories: BDS-S1, GPS-S1, Galileo-S1, BDS-S2, GPS-S2, Galileo-S2, BDS-S3, GPS-S3, Galileo-S3, BDS-S4, GPS-S4, Galileo-S4, BDS-S5, GPS-S5, Galileo-S5; There is also a category where all data is merged into one: ALL, which is a pure data fusion model; in total, there are 20 categories of data.
[0068] For example, training three algorithms (CatBoost, LightGBM, and XGBoost) on BDS-S1 data reveals that the optimal training model is LightGBM. To address the issue of insufficient high-wind-speed data leading to model distortion, this invention employs high-wind-speed oversampling. This involves multiplying the sea surface wind speeds greater than 10 meters per second in the BDS-S1 dataset by a factor of 1, 2, up to 10, thus expanding the high-wind-speed dataset before training the LightGBM algorithm. The optimal algorithm for BDS-S1 data is LightGBM, and the highest accuracy in sea surface wind speed retrieval is achieved when the optimal high-wind-speed weight is 10. Similarly, experiments are conducted on the other 20 data types to determine the optimal algorithm and optimal high-wind-speed weights.
[0069] Optimize the proportion of high-wind-speed samples; 4.1 Definition and Oversampling of High Wind Speed Samples; High wind speed samples are defined as those with WS ≥ 10 m / s (critical range for extreme weather). The original data contains a small proportion of high wind speed samples, resulting in an imbalanced distribution. The dataset distribution is adjusted by oversampling (replicating high wind speed samples) at a rate of 1-10 times.
[0070] 4.2 Comparison of model performance under different proportions; 600 models were trained and tested, with 20 data categories × 10 oversampling ratios × 3 models, and RMSE as the core indicator.
[0071] The optimal oversampling ratios for each data type were ultimately determined as follows: BDS-CatBoost (10x), GPS-CatBoost (10x), GLONASS-CatBoost (0x), Galileo-LightGBM (10x), BDS / LHCP-LightGBM (10x), GPS / LHCP-XGBoost (10x), Galileo / LHCP-XGBoost (10x), BDS / H-XGBoost (5x), GPS / H-LightGBM (8x), Galileo / H- CatBoost (4x), BDS / V-XGBoost (1x), GPS / V-LightGBM (9x), Galileo / V-LightGBM (10x), BDS / LHCP-CatBoost (10x), GPS / LHCP-CatBoost (9x), Galileo / LHCP-XGBoost (10x), BDS / RHCP-XGBoost (10x), GPS / RHCP-CatBoost (10x), Galileo / RHCP-XGBoost (10x).
[0072] Multi-source data fusion and optimal model determination; 5.1 Multi-source data fusion; Multi-GNSS system fusion: The prediction results of the optimal models of the four systems, BDS, GPS, GLONASS and Galileo, are fused together.
[0073] Data from different GNSS systems includes the following four categories: BDS, GPS, GLONASS, and Galileo. After optimizing the algorithm and selecting the optimal high wind speed, the optimal models for these four types of data are BDS-CatBoost (10x), GPS-CatBoost (10x), GLONASS-CatBoost (0x), and Galileo-LightGBM (10x). The inversion results of these four optimal models are then combined to obtain the multi-GNSS system fusion results.
[0074] Multipolarity mode fusion: The prediction results of the optimal models of the four polarization modes LHCP, RHCP, H and V of the four systems are fused.
[0075] Table 2 identifies 15 data categories with different polarization modes: BDS-S1, GPS-S1, Galileo-S1, BDS-S2, GPS-S2, Galileo-S2, BDS-S3, GPS-S3, Galileo-S3, BDS-S4, GPS-S4, Galileo-S4, BDS-S5, GPS-S5, and Galileo-S5. After optimizing the algorithm and selecting the optimal high wind speed, the optimal models for these 15 data categories are determined to be BDS / LHCP-LightGBM (10x), GPS / LHCP-XGBoost (10x), Galileo / LHCP-XGBoost (10x), BDS / H-XGBoost (5x), GPS / H-LightGBM (8x), Galileo / H-CatBoost (4x), BDS / V-XGBoost (1x), GPS / V-LightGBM (9x), and Galileo-S5. The inversion results of these 15 optimal models—lileo / V-LightGBM (10x), BDS / LHCP-CatBoost (10x), GPS / LHCP-CatBoost (9x), Galileo / LHCP-XGBoost (10x), BDS / RHCP-XGBoost (10x), GPS / RHCP-CatBoost (10x), and Galileo / RHCP-XGBoost (10x)—are combined to obtain the multi-polarization mode fusion result.
[0076] 5.2 Comparison of results from multi-strategy fusion; Performance comparison of test sets for multi-GNSS fusion, multi-polarization fusion, and pure data fusion (without distinguishing between system and polarization); Multi-GNSS fusion: Full wind speed R=0.87, RMSE=1.710m / s; Multipolar mode fusion: R=0.88 at all wind speeds, RMSE=1.685m / s; Pure data fusion: full wind speed R=0.88, RMSE=1.688m / s; In conclusion, the sea surface wind speed inversion accuracy is highest when the multipolar model fusion is used.
[0077] 5.3 Results of sea surface wind speed inversion at orbital points; Based on three fusion models, the orbital points of the TM-1 satellite in the Yellow and Bohai Seas are displayed point-by-point, generating spatial distribution and scatter plots of sea surface wind speed within the inversion period. The results show that the scatter plot density of the inverted values and the true ERA5 values is concentrated near the 45° line, indicating the reliability of the sea surface wind speed inversion, especially with the multi-polarization model fusion inversion yielding the best results. This invention significantly improves the estimation accuracy and stability in high-wind-speed scenarios, reducing the systematic underestimation problem of high wind speeds; multi-constellation and multi-polarization joint modeling enhances spatial stability and observation redundancy, enabling reliable sea surface wind speed products to be output even when constellation coverage is uneven or partially failed.
[0078] Example 2, the spaceborne GNSS-R multi-system multi-polarization sea surface wind speed estimation system provided in this embodiment of the invention includes: The data acquisition module is used to extract the feature values of the spaceborne GNSS-R and the environmental factor data of ERA5; The data quality control and matching module is used to perform quality control on the datasets of multiple GNSS systems and multiple polarization modes of the spaceborne GNSS-R constellation, and to perform grid matching interpolation with ERA5 data; The model training and optimization module is used to divide the training set and test set based on multiple GNSS systems and multiple polarization modes, and to build a sea surface wind speed estimation model based on the Optuna optimized machine learning model. The high-wind-speed sample weighting module is used to increase the proportion of high-wind-speed datasets in the model, and to determine the optimal proportion of high-wind-speed data and the optimal machine learning model for different GNSS systems and different polarization modes. The fusion and inversion module is used to perform multi-GNSS system fusion, multi-polarization mode fusion and compare with pure data fusion to determine the best sea surface wind speed estimation model and output sea surface wind speed estimates based on orbit points.
[0079] To further demonstrate the positive effects of the above embodiments, the present invention conducts the following experiments based on the above technical solutions.
[0080] like Figure 2 As shown, this figure presents the RMSE comparison results of a test set of 600 sea surface wind speed estimation models constructed using 20 data types, 10 high-wind-speed weights, and 3 machine learning models in the Yellow and Bohai Sea regions of China. The optimal solution for each data configuration is also identified. This figure visually demonstrates that the high-wind-speed sample optimization strategy and Optuna model optimization described in this invention can systematically find the model configuration with the highest inversion accuracy for different GNSS systems and polarization modes.
[0081] like Figure 3 As shown, the scatter density distribution of sea surface wind speed inversion results under three strategies—multi-GNSS system fusion, multi-polarization model fusion, and pure data fusion—is compared with the true values of ERA5 reanalysis data. It can be seen that the scatter points of the inversion results from multi-polarization model fusion are most densely distributed on both sides of the 1:1 baseline representing perfect prediction. This corroborates the quantitative conclusion that "multi-polarization model fusion achieves the highest accuracy in sea surface wind speed inversion," directly demonstrating the effectiveness and superiority of the multi-source fusion strategy proposed in this invention.
[0082] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention and within the spirit and principles of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for estimating sea surface wind speed using a spaceborne GNSS-R multi-system multi-polarization system, characterized in that, The method includes the following steps: S1, extract the feature value data of spaceborne GNSS-R and the environmental factor data of ERA5; S2 performs quality control on the datasets of multiple GNSS systems and multiple polarization modes of the spaceborne GNSS-R constellation, and performs grid matching interpolation with ERA5 data; S3, based on multiple GNSS systems and multiple polarization modes, divides the training set and test set, and constructs a sea surface wind speed estimation model based on the Optuna optimized machine learning model; S4, increase the proportion of high wind speed dataset in the model, and determine the optimal proportion of high wind speed and the optimal machine learning model for different GNSS systems and different polarization modes. S5 performs multi-GNSS system fusion and multi-polarization mode fusion respectively, and compares them with pure data fusion to determine the best sea surface wind speed estimation model and outputs sea surface wind speed estimates based on orbit points; wherein, the best sea surface wind speed estimation model is determined based on the performance indicators of the model on the test set, and the performance indicators include correlation coefficient R, root mean square error RMSE, mean error ME and mean absolute percentage error MAPE. The multi-GNSS system fusion includes merging the prediction results of the optimal models of BDS, GPS, GLONASS and Galileo systems, and the multi-polarization mode fusion includes merging the prediction results of the optimal models of LHCP, RHCP, H and V polarization modes. The pure data fusion refers to all data fusion models that do not distinguish between GNSS systems and polarization modes.
2. The method for estimating sea surface wind speed using a spaceborne GNSS-R multi-system multi-polarization system according to claim 1, characterized in that, In step S1, the characteristic data of the spaceborne GNSS-R includes: Spatial positioning parameters, including the latitude and longitude of the mirror point and the incident angle of the mirror point, are used to characterize the reflection geometry. Time and observation information, including DDM generation time, number of channels, satellite number, and surface type; Parameters related to the energy distribution of the DDM include the leading edge slope at the mirror point, the normalized scattering cross section, the raw peak power count, the peak power signal-to-noise ratio, the Doppler column where the peak is located, and the time delay row. The fourth-order moment characteristic parameters of the DDM include kurtosis and skewness, which characterize the statistical properties of the power distribution; the reflection coefficient, which represents the reflection intensity at the mirror point; and the quality label, which provides a marker for the reliability of DDM observations.
3. The method for estimating sea surface wind speed using a spaceborne GNSS-R multi-system multi-polarization system according to claim 1, characterized in that, In step S1, the environmental factor data of ERA5 includes: total precipitation, sea surface temperature, significant wave height, 10-meter sea surface wind speed, UTC time, and latitude and longitude; wherein, the 10-meter sea surface wind speed is synthesized from the 10-meter u wind component and v wind component.
4. The method for estimating sea surface wind speed using a spaceborne GNSS-R multi-system multi-polarization system according to claim 1, characterized in that, In step S2, the quality control uses Ddm_quality_flag to filter data and remove low-quality data; wherein, Ddm_quality_flag contains multiple binary bits, which are used to mark the position of low-quality data as 1 and record quality problems in different dimensions; The grid matching interpolation includes: spatially using bilinear interpolation to map regular grid ERA5 data to irregular GNSS-R observation point locations, and temporally using linear interpolation to achieve the correspondence between the observation time and the ERA5 at the analysis time.
5. The method for estimating sea surface wind speed using a spaceborne GNSS-R multi-system multi-polarization system according to claim 1, characterized in that, In step S3, the multi-GNSS system includes BDS, GPS, GLONASS and Galileo, and the multi-polarization mode includes LHCP, RHCP, H and V; Based on the correspondence between the multi-GNSS system and the polarization mode configurations of different satellites in the TM-1 constellation, the observation data is divided into 20 data configurations. For each type of data, training and test sets were configured, and CatBoost, LightGBM and XGBoost models based on Optuna optimization were built respectively, for a total of 60 sea surface wind speed estimation models.
6. The method for estimating sea surface wind speed using a spaceborne GNSS-R multi-system multi-polarization system according to claim 1, characterized in that, In step S3, the machine learning model adaptively adjusts the parameters of the gradient boosting tree model through the Optuna hyperparameter optimization framework to minimize the weighted mean square error based on cross-validation. The input sample set is: ; In the formula, For the sample set, containing One sample, The total number of samples; For the first The GNSS-R signal feature vector of each sample includes reflected signal intensity, scattering geometry parameters, satellite elevation angle, and polarization information; For the first The actual sea surface wind speed of each sample; Sample weights are used to balance the importance of different samples. ; Machine learning models are defined as follows: ; In the formula, For the model to the first The predicted value for each sample, This is the model mapping function, used to map input features to output predicted values; For model training parameters, The set of hyperparameters to be optimized includes learning rate, maximum depth, number of leaf nodes, regularization parameter, and sampling ratio; The Optuna optimization process uses the cross-validation weighted mean square error as the objective function, defined as: ; In the formula, The weighted mean squared error for cross-validation is used to measure the model's performance given hyperparameters. Overall verification error; The number of folds for cross-validation, i.e., dividing the sample into folds. A set of non-overlapping verification subsets; For the first The set of validation samples for the fold; In the first The trained model is then combined using hyperparameters. Below, on the sample The predicted value; For sample index, ; Optuna's optimization objective is: ; In the formula, The optimal combination of hyperparameters is the one that minimizes the cross-validation error. ; The hyperparameter search space is the set of parameters that Optuna samples and evaluates. To verify the square root form of the error, and to ensure that the objective function is consistent with the RMSE dimension, making the optimization results more intuitive; This is a parameter combination operator that minimizes the objective function.
7. The spaceborne GNSS-R multi-system multi-polarization sea surface wind speed estimation method according to claim 6, characterized in that, The optimization process includes: (1) Hyperparameter space setting: Define the parameter range according to the model type; where: learning rate controls the step size of each iteration, The maximum depth of a single decision tree determines the model's complexity and generalization ability. The number of base learners, used to balance bias and variance. ; (2) Trial generation and training: Optuna generates a set of hyperparameters based on a random sampling strategy according to historical trial results. And use the parameters to train the model; (3) Cross-validation and loss calculation: K-fold cross-validation is used to calculate the validation error of each fold. And take the average to get ; (4) Pruning mechanism: Optuna automatically terminates poorly performing experiments based on intermediate evaluation results to save computational resources; (5) Determination of optimal parameters and model training: After a preset number of trials T, the optimal hyperparameters are determined. And based on all the training data, the optimal model training parameters are obtained by retraining. Thus, the final model is obtained. ; By optimizing the process, the obtained machine learning model can achieve high-precision sea surface wind speed estimation from GNSS-R data, and the output results are as follows: ; In the formula, The sea surface wind speed estimate obtained from the inversion. These are the optimal training parameters for the model.
8. The method for estimating sea surface wind speed using a spaceborne GNSS-R multi-system multi-polarization system according to claim 1, characterized in that, In step S4, increasing the proportion of the high-wind-speed dataset in the model includes: For observation data from different GNSS systems and under different polarization modes, the proportion of high wind speed sample data was increased sequentially from 1 to 10 times the original sample size. The training and test sets under each proportion condition were modeled and validated using Optuna-optimized machine learning models, resulting in a total of 600 models. Among them, the high wind speed samples are those with sea surface wind speeds greater than or equal to 10 m / s.
9. A spaceborne GNSS-R multi-system multi-polarization sea surface wind speed estimation system, characterized in that, The system is used to implement the method as described in any one of claims 1-8, comprising: The data acquisition module is used to extract the feature values of the spaceborne GNSS-R and the environmental factor data of ERA5; The data quality control and matching module is used to perform quality control on the datasets of multiple GNSS systems and multiple polarization modes of the spaceborne GNSS-R constellation, and to perform grid matching interpolation with ERA5 data; The model training and optimization module is used to divide the training set and test set based on multiple GNSS systems and multiple polarization modes, and to build a sea surface wind speed estimation model based on the Optuna optimized machine learning model. The high-wind-speed sample weighting module is used to increase the proportion of high-wind-speed datasets in the model, and to determine the optimal proportion of high-wind-speed data and the optimal machine learning model for different GNSS systems and different polarization modes. The fusion and inversion module is used to perform multi-GNSS system fusion, multi-polarization mode fusion and compare with pure data fusion to determine the optimal sea surface wind speed estimation model and output sea surface wind speed estimates based on orbit points; wherein, the optimal sea surface wind speed estimation model is determined based on the performance indicators of the model on the test set, and the performance indicators include correlation coefficient R, root mean square error RMSE, mean error ME and mean absolute percentage error MAPE. The multi-GNSS system fusion includes merging the prediction results of the optimal models of BDS, GPS, GLONASS and Galileo systems, and the multi-polarization mode fusion includes merging the prediction results of the optimal models of LHCP, RHCP, H and V polarization modes. The pure data fusion refers to all data fusion models that do not distinguish between GNSS systems and polarization modes.
Citation Information
Patent Citations
Satellite-borne GNSS-R sea surface wind speed inversion method based on multilayer iterative weighted model
CN119045004A
Global lake wind speed estimation method based on satellite-borne BDS / GPS / GLONASS / Galileo reflected signals of China Tangyu No.1
CN119355787A