Residual multi-modal GNSS-R height measurement method based on attention mechanism feature enhancement

By constructing a feature-enhanced residual multimodal model based on an attention mechanism, the problems of complex error models and low information utilization in sea surface measurement using deep learning methods are solved, achieving higher inversion accuracy and stronger model generalization ability.

CN122631033APending Publication Date: 2026-08-25HARBIN INST OF TECH AT WEIHAI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511452377.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing deep learning methods for sea surface altimetry suffer from problems such as complex error models, low information utilization, and noise pollution affecting inversion accuracy. Traditional methods also struggle to effectively utilize multimodal GNSS-R data.

Method used

A new FERM model for sea surface height inversion is constructed by combining a feature-enhanced residual network (FER) and a fully connected neural network (FCNN) and utilizing channel and spatial attention modules to enhance feature extraction, improve the utilization of DDM information and feature attention weights.

Benefits of technology

It significantly improved the accuracy of sea surface height inversion, reduced the mean absolute error and root mean square error, increased the Pearson correlation coefficient, and improved the model's generalization ability and inversion accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122631033A_ABST
    Figure CN122631033A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of satellite altimetry, in particular to a GNSS-R altimetry method based on attention mechanism feature enhancement residual multi-modal, which can significantly improve the altimetry precision. First, a new attention mechanism feature enhancement residual multi-modal FERM model is constructed, and then the FERM model is used to realize altimetry. The FERM model is composed of a feature enhancement residual FER network based on an attention mechanism and a fully connected neural network FCNN. The FER uses a channel attention module and a spatial attention module to distinguish and enhance the DDM channel and spatial effective features extracted by the residual structure, so as to enhance the features from two aspects of improving the DDM information utilization rate and giving the feature attention weight, and then improve the SSH inversion precision of the FCNN from the input quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of satellite altimetry technology, specifically to an attention-based feature-enhanced residual multimodal GNSS-R altimetry method that can significantly improve altimetry accuracy. Background Technology

[0002] Sea level is a core parameter in oceanography and meteorology, reflecting not only the physical state of the ocean but also being closely related to key issues such as the global climate system, ocean circulation, and sea level rise. Spaceborne GNSS-R sea level measurement technology retrieves sea level by receiving and analyzing navigation satellite signals received directly from and reflected from the sea surface. This technology boasts advantages such as abundant signal sources, high spatiotemporal resolution, and the ability to perform measurements 24 / 7 and in all weather conditions. However, traditional methods suffer from drawbacks such as complex error models and significant information loss when compressing a two-dimensional DDM into a one-dimensional model.

[0003] Deep learning methods approximate unknown predicted values ​​in a data-driven manner, fully utilizing various information related to the inversion of unknowns while avoiding the use of highly complex geometric and error models. In recent years, deep learning methods have achieved some success in sea surface altimetry. However, the accuracy of current deep learning-sea surface altimetry fusion algorithms still has significant room for improvement. This is partly due to the high orbits and low signal transmission power of GNSS satellites, which often contaminates the effective DDM signal with background noise, affecting inversion accuracy. Furthermore, the deep learning fusion algorithm requires further research. Summary of the Invention

[0004] This invention addresses the shortcomings and deficiencies of existing technologies by proposing an attention-based feature-enhanced residual multimodal GNSS-R altimeter method that can significantly improve altimeter accuracy.

[0005] This invention achieves its purpose through the following measures: A novel attention-based feature enhancement residual multimodal GNSS-R altimeter method is proposed. The method is characterized by: first, constructing a novel attention-based feature enhancement residual multimodal (FERM) model; and second, utilizing the FERM model for altimeter measurement. The FERM model is constructed by combining a feature enhancement residual (FER) network based on an attention mechanism and a fully connected neural network (FCNN). The FER network utilizes channel attention modules and spatial attention modules to distinguish and enhance the effective features of the DDM channels and space extracted from the residual structure, thereby improving the utilization rate of DDM information and assigning attention weights to the features. This improves the inversion accuracy of SSH by FCNN from the perspective of input quality.

[0006] The FERM model of this invention has two inputs. In the first input, based on the FER architecture, features are extracted from different DDM combinations through several residual blocks with 3×3 convolutional kernels. After the feature extraction process, the two-dimensional DDM data is converted into one-dimensional numerical data, and the heterogeneous data is converted into data with the same structure. The second part is used to input auxiliary parameters. The extracted features are fused with the auxiliary parameters in the additional fully connected layer, and then the fused data is used in FCNN to predict SSH. The input to the convolutional layer of the FERM model is X 1, that is, different DDM combinations of CYGNSS, with auxiliary parameter input as follows: X 2. The model output can be represented as: (1), In the formula, This indicates the output of FERM. L and N These represent the hidden layers determined by grid search and the number of neurons in each layer, respectively. W i and b i They represent the first i The weights and biases of each hidden layer This indicates the output of the FER module. This represents the activation function for each hidden layer.

[0007] In this invention, the FER network used to extract effective features from the DDM in the FERM model is a novel feature-enhancing residual network block suitable for DDM data structures. First, the residual block contains two 3×3 convolutional layers, each followed by a batch normalization layer and a ReLU activation function, as well as a feature enhancement layer. Second, by skipping these two convolutional processes in the cross-layer data path, the input is directly appended to the final ReLU activation function. To concatenate the outputs of the two convolutional layers, they must have the same shape as the input. If the number of channels needs to be changed, an additional 1×1 convolution is added to transform the input into the desired form before the summation operation. Furthermore, the DDM space size of L1b is only 17×11. Pooling layers affect the depth retention of convolutional layers, and further reduction of the space size will lead to the loss of a large amount of valuable information; therefore, pooling layers are discarded. Finally, since the 17×11 DDM data structure is relatively simple, the original number of channels is too large to increase the computational cost of the model. Therefore, it is necessary to reduce the number of channels. An attention-based feature enhancement module is added after the second convolutional layer to enhance the ability of convolutional feature extraction and characterize the importance of effective information.

[0008] In this invention, the feature enhancement layer of the FER network has two sub-blocks: a channel attention module and a spatial attention module. Since convolutional operations extract features by mixing cross-channel and spatial information, an attention-based feature enhancement module is used to apply effective features along both the channel and spatial axes. The channel and spatial attention modules are applied sequentially, allowing each feature map to learn different features along both the channel and spatial axes, thereby enhancing features by improving the utilization of DDM information and assigning attention weights to features. Specifically, the channel attention module generates a channel attention feature map by utilizing the relationships between channels in the feature map. Each channel of the feature is treated as a feature extractor, and the spatial dimension of the input feature map is compressed using average and max pooling methods to condense channel feature information. After pooling, the results are input into a fully connected network and merged to obtain the final extracted inter-channel features, generating the output channel feature map. The model output is represented as follows: (2), In the formula, Represents the output channel feature map, and represents the activation function. and This indicates two hidden layers. This represents the average pooling channel feature map. The model represents the max-pooled channel feature map. The spatial attention module generates a spatial attention feature map by utilizing the importance of spatial information within the feature map channels. Each spatial feature is treated as a feature extractor. The average and max-pooling methods are used to compress the channel dimension of the input feature map to condense the spatial feature information. After pooling, a convolutional network is used to distinguish the importance of the spatial information in the input feature map and generate corresponding feature labels as in-channel features to generate the output spatial feature map. The model output is represented as follows: (3), where, This represents the output spatial feature map. This represents the activation function. This indicates a convolutional layer with a 3×3 kernel. This represents the spatial feature map of average pooling. This represents the feature map of the max pooling space.

[0009] In this invention, the feature enhancement module structure based on the attention mechanism is arranged in the order of channel attention module first, followed by spatial attention module, and generates the final output feature-enhanced feature map. The model output is represented as follows: (4), In the formula, This represents the output feature enhancement feature map. This indicates a spatial attention module. This indicates the channel attention module. This represents the input feature map.

[0010] Before performing height measurement, this invention normalizes the DDM to enhance the convergence of the neural network and improve the generalization ability of the model. The DDM is standardized using the following formula: (5), In the formula, This represents the normalized DDM. Represents any DDM in the dataset. This represents the maximum pixel value of the corresponding DDM in the dataset. The auxiliary parameters also need to be normalized. The normalization formula for the auxiliary parameters is as follows: (6), where, This represents the auxiliary parameters after the normalization operation. Represents any auxiliary parameter in the dataset. This represents the maximum value of the corresponding auxiliary parameter in the dataset.

[0011] This invention uses MAE, RMSE, and PCC to verify the reliability of the method. Smaller MAE and RMSE values ​​indicate a good fit between the predicted and actual values. The formulas are as follows: (9), (10) (11).

[0012] This invention addresses the limitations of current deep learning-based sea surface height measurement methods, such as low information utilization and susceptibility to marginal feature misleading. It proposes a novel attention-based Feature Enhancement Residual Multimodal (FERM) model for retrieval of sea surface height (SSH). FERM is constructed by combining an attention-based Feature Enhancement ResNet (FER) network and a Fully Connected Neural Network (FCNN). The FER network utilizes channel attention modules and spatial attention modules to distinguish and enhance the effective channel and spatial features extracted from the residual structure, thereby improving feature utilization and assigning attention weights to features. This enhances the features and improves the SSH retrieval accuracy of FCNN by improving input quality. This invention uses a DTU18 validation model to evaluate the accuracy of the SSH retrieval model. The validation model is based on the DTU18 global mean sea level (Danmarks Tekniske Universitet 18 Mean Sea). The model consists of the Surface (DTU18MSS) model and the TPXO8 global tidal model. Results show that the mean absolute error (MAE) of the original residual model and the FERM model are 3.84 and 2.56, respectively, and the root mean square error (RMSE) is 5.46 and 3.74, respectively. MAE and RMSE are reduced by approximately 33.33% and 31.50%, respectively. The Pearson correlation coefficient (PCC) is 0.98 and 0.99, respectively, an improvement of approximately 1.02%. This invention further implements the HALF re-tracking algorithm to invert SSH, and the results are compared with those calculated by the FERM method. Based on MAE, RMSE, and PCC, the FERM method is more accurate than the HALF re-tracking method, reducing MAE and RMSE by approximately 62.41% and 50.20%, respectively, and improving PCC by approximately 4.21%. In summary, this invention has significant advantages compared to existing technologies. Attached Figure Description

[0013] Figure 1This is a schematic diagram of the network structure of the FERM model in this invention.

[0014] Figure 2 This is a comparison diagram of basic network structures, in which... Figure 2 (a) is a diagram of the conventional convolutional network block structure; (b) is a diagram of the conventional residual convolutional network block structure; and (c) is a diagram of the feature-enhanced residual network block structure in this invention.

[0015] Figure 3 This is a schematic diagram of the attention sub-block in this invention.

[0016] Figure 4 This is a schematic diagram of the feature enhancement module in this invention.

[0017] Figure 5 This is a data distribution diagram of the training set, validation set, and test set in an embodiment of the present invention.

[0018] Figure 6 This is a channel density distribution diagram of the RM model in an embodiment of the present invention, wherein... Figure 6 (a) is a density distribution map of one channel; (b) is a density distribution map of two channels; (c) is a density distribution map of three channels; and (d) is a density distribution map of four channels.

[0019] Figure 7 This is a channel density distribution diagram of the FERM model in an embodiment of the present invention; wherein Figure 7 (a) is a density distribution map of one channel; (b) is a density distribution map of two channels; (c) is a density distribution map of three channels; and (d) is a density distribution map of four channels.

[0020] Figure 8 These are the error distribution histograms and probability density functions of the inversion results under different input conditions for the two models in this embodiment of the invention.

[0021] Figure 9 This is an error distribution diagram of the HALF re-tracking method and the FERM model in the embodiments of the present invention. Detailed Implementation

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0023] Based on the low reflection signal power, high noise, and binary characteristics of DDM (Distributed Image Module), the architecture of the novel FERM model proposed in this invention mainly consists of FER (Functional Rendering) and FCNN (Functional CNN-based Array Module). FER is used to identify effective features from different combinations of DDMs, and FCNN is used for SSH (Single Image Module) inversion, such as... Figure 1As shown, a new FERM model is constructed. The new FERM model has two main inputs. In the first input, based on the FER architecture, features are extracted from different DDM combinations through several residual blocks with 3×3 convolutional kernels. After feature extraction, the 2D DDM data is converted into 1D numerical data. Heterogeneous data is transformed into data with the same structure. The second part is used to input auxiliary parameters. The extracted features are fused with the auxiliary parameters in an additional fully connected layer. Then, the fused data is used in FCNN for SSH prediction.

[0024] The input to the convolutional layer of the FERM model of this invention is X 1, that is, different DDM combinations of CYGNSS, with auxiliary parameter input as follows: X 2. The model output can be represented as: (1), In the formula, This indicates the output of FERM. L and N These represent the hidden layers determined by grid search and the number of neurons in each layer, respectively. W i and b i They represent the first i The weights and biases of each hidden layer This indicates the output of the FER module. This represents the activation function for each hidden layer.

[0025] This invention also designs a novel feature-enhanced residual network block suitable for DDM data structures, which improves feature representation capabilities from two aspects: improving the utilization rate of DDM information and feature attention weights.

[0026] A Convolutional Neural Network (CNN) is a deep learning model specifically designed for processing grid-structured data (such as images and videos). A typical CNN structure mainly consists of an input layer, convolutional layers, activation function layers, pooling layers, fully connected layers, and an output layer. The convolutional layer is the core component of a CNN, extracting local features from the input data and generating feature maps through convolution operations. The pooling layer reduces the spatial size of the feature maps through pooling operations, reducing interference information while retaining key information, thus accelerating model training and enhancing the model's robustness.

[0027] To improve the performance of CNNs, three key network factors were primarily investigated: depth, width, and cardinality. From the LeNet to the ResNet architecture, networks have become increasingly deeper to achieve richer architectures. ResNet stacks the same residual block topology and skipped connections to build an extremely deep architecture. Zagoruyko and Komodakis proposed increasing the network depth based on the ResNet architecture. In the CIFAR benchmark, the width can exceed that of an extremely wide ResNet with 1001 layers. Xception and ResNeXt proposed methods to increase the network cardinality. Compared to CNNs, ResNet exhibits significantly stronger performance.

[0028] Due to the low power, high noise, and small image size of reflected signals in DDM (Digital Depth Model), the original ResNet residual structure is difficult to adapt to feature extraction from DDM data. Considering these factors, the attention mechanism not only identifies effective features but also further characterizes the effectiveness of those features. Therefore, this invention designs a novel feature enhancement residual network block based on the attention mechanism, suitable for DDM data structures. Figure 2 It demonstrates the differences between the basic structures of CNN, regular convolution, and FER.

[0029] First, the residual block contains two 3×3 convolutional layers, each followed by a batch normalization layer and a ReLU activation function, as well as a feature enhancement layer. Second, by skipping these two convolutional processes in the cross-layer data path, the input is directly appended before the final ReLU activation function. For the outputs of the two convolutional layers to be concatenated, they must have the same shape as the input. If the number of channels needs to be changed, an additional 1×1 convolution is added to transform the input into the desired form before the summation operation. Furthermore, the DDM spatial size of L1 b is only 17×11; pooling layers affect the depth retention of convolutional layers, and further reduction in spatial size would lead to the loss of a large amount of valuable information, therefore, pooling layers are discarded. Finally, since the 17×11 DDM data structure is relatively simple, the original number of channels is too large to increase the computational cost of the model, so the number of channels needs to be reduced. The attention-based feature enhancement module designed in this invention is added after the second convolutional layer to enhance the ability of convolutional feature extraction and characterize the importance of effective information.

[0030] The feature enhancement layer in the FER structure of this invention has two sub-blocks: a channel attention module and a spatial attention module. Since convolutional operations extract features by mixing cross-channel and spatial information, this invention employs an attention-based feature enhancement module to effectively utilize features along the two main directions: the channel and spatial axes. To achieve this, the invention applies the channel and spatial attention modules sequentially, allowing each feature map to learn different features along the channel and spatial axes respectively, thereby enhancing features by improving the utilization of DDM information and assigning attention weights to the features.

[0031] The channel attention module generates channel attention feature maps by utilizing the relationships between channels in the feature maps, treating each channel of a feature as a feature extractor. For example... Figure 3 As shown in the green section, to achieve this, average and max pooling methods are used to compress the spatial dimension of the input feature map to condense channel feature information. After pooling, the results are input into a fully connected network and then merged to obtain the final extracted inter-channel features, generating the output channel feature map. The model output can be represented as: (2), In the formula, This represents the output channel feature map. This represents the activation function. and This indicates two hidden layers. This represents the average pooling channel feature map. This represents the feature map of the max-pooling channel.

[0032] The spatial attention module generates a spatial attention feature map by utilizing the importance of spatial information within the feature map channels, treating each spatial feature as a feature extractor. For example... Figure 3 As shown in purple, to achieve this, average and max pooling methods are used to compress the channel dimension of the input feature map to condense spatial feature information. After pooling, convolutional networks are used to distinguish the importance of spatial information in the input feature map and generate corresponding feature labels as in-channel features to generate the output spatial feature map. The model output can be represented as: (3), where, This represents the output spatial feature map. This represents the activation function. This indicates a convolutional layer with a 3×3 kernel. This represents the spatial feature map of average pooling. This represents the feature map of the max pooling space.

[0033] The feature enhancement module structure based on the attention mechanism consists of the two parts mentioned above, with channel attention and spatial attention complementing each other, such as... Figure 4As shown, the attention modules are arranged in the order of channel attention first, followed by spatial attention, and the final output feature-enhanced feature map is generated. The model output can be represented as: (4), where, This represents the output feature enhancement feature map. This indicates a spatial attention module. This indicates the channel attention module. This represents the input feature map. Example

[0034] CYGNSS is the first satellite system specifically designed for meteorological observation using GNSS reflected signals. It consists of eight small satellites, each simultaneously recording reflected signals from four GNSS satellites at 0.5-second intervals. Its high temporal resolution provides over 200 million DDM images annually. This massive amount of data places a heavy burden on manual processing, and large-scale batch processing is one of the strengths of deep learning algorithms. The CYGNSS dataset can be divided into two forms: 11×17 two-dimensional image data (DDM) and corresponding one-dimensional auxiliary parameter data. The DDM mainly includes raw (Raw Counts) DDM, analog power (Analog) DDM, bistatic radar cross section (BRCS) DDM, and effective scattering area (EFF) generated by the Delay Doppler Mapper Instrument (DDMI). Raw DDM without any processing is a direct measure of dispersed signal power and is commonly used for SSH inversion; Analog DDM is mainly used to evaluate parameters such as signal strength and signal-to-noise ratio; EFF DDM is significant for parameters such as sea surface wind speed and soil moisture; and because BRCS DDM can more significantly represent sea surface wind speed information, it has a stronger correlation with sea surface roughness. In this invention, four dimensions—Raw DDM, BRCS DDM, EFF DDM, and Analog DDM—are used as the original input data. The main auxiliary parameters used in this example mainly include one-dimensional auxiliary parameters from the CYGNSS L1b data product and the main errors in signal path propagation, including ionospheric delay, tropospheric delay, and ERA5, U10, and V10 wind speeds.

[0035] Because the CYGNSS dataset contains two different dimensions of data, it is difficult to utilize them simultaneously using traditional methods. Furthermore, the low reflectance signal power and high noise characteristics of the DDM (Digital Image Dataset) pose significant challenges for researchers when using their own experience to construct feature extraction methods to extract effective information. Therefore, this example employs a feature-enhanced multimodal deep learning approach to address these difficulties. It uses FER (Feature Enhancement) to adaptively extract effective features from the complete DDM while incorporating two attention mechanisms to enhance the focus on effective information during the adaptive extraction process. The two-dimensional image data of the DDM is then converted into one-dimensional numerical data. The extracted features and auxiliary parameters are then fused and used as input data for FCNN to invert the SSH (Synchronous Image Dataset).

[0036] However, there are inconsistencies in wind speed measurements between different satellites, and the conclusions indicate that the wind speed quality obtained by the CY04 satellite is slightly better than that of other satellites. Therefore, this example selects and analyzes CY04 satellite measurement data from January 1, 2020 to March 31, 2021.

[0037] To enhance the convergence of the neural network and improve the model's generalization ability, data preprocessing is required: the DDM is normalized using the following formula: (5), where, This represents the normalized DDM. Represents any DDM in the dataset. This represents the maximum pixel value of the corresponding DDM in the dataset. For the reasons mentioned above, the auxiliary parameters also need to be normalized. The normalization formula for the auxiliary parameters is as follows: (6), where, This represents the auxiliary parameters after the normalization operation. Represents any auxiliary parameter in the dataset. This represents the maximum value of the corresponding auxiliary parameter in the dataset.

[0038] When constructing a deep learning SSH inversion model, due to its data-driven nature, data quality control is particularly important for subsequent dataset use and model training. To ensure the quality of the dataset, the quality control of the CYGNSS dataset is carried out according to the following standards: (1) All Nan samples are discarded; (2) DDM observations should be non-negative; (3) All land data should be discarded; (4) The data quality control flags provided by the L1 level dataset of CYGNSS are all 0; (5) The uncertainty of BRCS is less than 1; (6) The correction gain of DDM is greater than 3; (7) The signal-to-noise ratio (SNR) of DDM is greater than 5; (9) The elevation angle is greater than 60°.

[0039] After quality control and data filtering, daily data were downsampled to alleviate computational burden and overfitting issues. A total of approximately 1×10⁻⁶ data was obtained. 6 There are 100 data samples, with SSH ranging from [−100, +80]m.

[0040] When constructing auxiliary parameter data, additional auxiliary parameters are added to participate in sea surface height inversion: (1) Ionospheric delay: Ionospheric delay will cause a distance error of several meters along the direct and reflected signal paths. Ionospheric delay is estimated using the Global Ionospheric Map (GIM) of the International GNSS Service (IGS). The total ionospheric delay can be calculated using the following formula: (7), where, Indicates the total ionospheric delay. This represents the ionospheric delay from the transmitter to the point of reflection on the mirror. This represents the ionospheric delay from the receiver to the point of reflection on the mirror. This represents the ionospheric delay from the receiver to the transmitter.

[0041] (2) Tropospheric Delay: The UNB3m model is used to estimate the tropospheric delay. Since the influence of the tropospheric delay is mainly concentrated within 10 km above sea level, tropospheric correction is only applicable to the reflected signal path below the altitude of the CYGNSS spacecraft. The total tropospheric delay can be expressed by the following formula: (8), where, Indicates the total tropospheric delay. This represents the tropospheric delay from the transmitter to the point of reflection on the mirror. This represents the tropospheric delay from the receiver to the specular reflection point.

[0042] (3) The ERA5 dataset is the fifth-generation global climate reanalysis dataset released by the European Centre for Medium-Range Weather Forecasts (ECMWF), and it is one of the latest and most widely used climate reanalysis products. Since its spatiotemporal resolution is 1 hour and 0.25 degrees, its spatiotemporal resolution is used for sampling when matching it with the CYGNSS dataset.

[0043] CYGNSS satellite data is spatiotemporally continuous, while the DTU18 mean sea level model is 1' grid data with latitude and longitude. Therefore, this example spatially matches the CYGNSS dataset with the DTU18 mean sea level model (latitude and longitude differing by less than 0.5'). Tidal variations are then estimated and applied to the DTU18 MSS. The CYGNSS dataset and the DTU18 validation model constitute the original dataset.

[0044] The original dataset was divided into training, validation, and test sets according to a certain ratio. These three datasets do not overlap and are used to verify that the model has a certain generalization ability over a time scale. Data from the first 250 days of 2020 was used as the training set, approximately 7 × 10⁻⁶. 6 One sample was used for model training and development. The data from the next 90 days or so was divided into a validation set, approximately 1.7 × 10⁻⁶. 6 A sample of approximately 1.9 × 10⁻⁶ samples was used for hyperparameter optimization and preliminary model performance evaluation. Data from January to March 2021, totaling 91 days, was selected as the test set. 6 There are [number] samples. Test data is not used in model building to prevent data leakage; this test data is only used to check the final accuracy and generalization performance of the model. The time distribution of the data used in the dataset construction is as follows: Figure 5 As shown.

[0045] This example uses MAE, RMSE, and PCC to test the reliability of the method. Smaller MAE and RMSE values ​​indicate a good fit between the predicted and actual values. The closer the PCC is to 1, the better the correlation between the inversion results and the DTU18 validation model. The corresponding formulas are as follows: (9), (10) (11).

[0046] To evaluate the accuracy of SSH inversion, it is necessary to compare and verify it with actual measured SSH. This example uses a validation model to evaluate the accuracy of the SSH inversion model. The DTU18 validation model consists of the DTU18 MSS model and the TPXO8 global tidal model. The SSH in the validation model can be represented as: (12), In the formula, This indicates the sea level height of the validation model. This represents the static sea level height of the DTU18 model. This indicates the tidal correction for the TPXO8 global tidal model.

[0047] This example compares and analyzes the inversion performance of the RM and FERM models to determine the effectiveness of model improvements. RM and FERM represent multimodal deep learning methods consisting of ResNet and FER feature extraction, respectively, and FCNN. Models RM and FERM are trained on the training set, their hyperparameters are optimized on the validation set, and finally, their performance is evaluated on the test set. All models are implemented in PyTorch.

[0048] This invention summarizes the number of input and output channels for each layer of the two models used. In the feature extraction stage, RM and FERM employ the same channel design scheme. X The value of is one to four, representing the number of input channels in the original DDM. During the SSH inversion stage, all models use the same FCNN design structure with five hidden layers, each with 200 neurons. First, this invention conducts a detailed study on the data sensitivity of the two models. This invention provides a statistical explanation of the relationship between the inversion results of the two models and the data input channels under full DDM input conditions. Previous studies mostly used Raw DDM for sea surface altimetry inversion, therefore this type of DDM was used as input data for single-channel applications. However, there is no experience to refer to for subsequent channel additions, so a full combination study was conducted. For two-channel input, Raw DDM was retained as one input channel while being randomly combined with three other types. For three-channel input, Raw DDM was also retained as one input channel while two other three types were randomly selected and combined.

[0049] It can be seen that the two models have similar impacts on the input DDM data. Under two-channel input conditions, both models perform best with the combination of Raw DDM and BRCS DDM, and the FERM model outperforms the RM model in all three combinations. Under three-channel input conditions, both models perform best with the combination of Raw DDM, BRCS DDM, and EFF DDM, and the FERM model shows a greater performance improvement rate than the RM model across all three combinations. Therefore, in subsequent studies, this invention adopted the combination with the best results in the table as the combination method for further research, that is, using Raw DDM and BRCS DDM as input data for two-channel input, and using Raw DDM, BRCS DDM, and EFFDDM as input data for three-channel input.

[0050] Figure 6 The scattering density map represents the inversion results of the RM model and the SSH DTU18 validation model, and a detailed study of the RM model was conducted from both spatial and channel perspectives. The results show that for input data with one to four 2D image channels, the MAE of the RM model are 3.92, 3.87, 3.85, and 3.84, respectively; the RMSE is 5.50, 5.44, 5.41, and 5.46, respectively; and the PCC is 0.98 for all four. The improvement rates for MAE and RMSE are 1.78% and 0.72%, respectively. These results indicate that the original RM model has low sensitivity to the number of input DDM channels, and the improvement rate is small as the number of data channels increases.

[0051] Figure 7The scattering density maps of the inversion results compared to the SSH DTU18 validation model are shown, and the same study was conducted on the new FERM model. The results show that for input data with one to four 2D image channels, the MAE of the FERM model are 3.36, 2.80, 2.71, and 2.56, respectively; the RMSE is 4.77, 4.07, 3.88, and 3.74, respectively; and the PCC is 0.98, 0.98, 0.99, and 0.99, respectively. With increasing channel number, the MAE of the FERM model increases by 23.60%, the RMSE by 21.59%, and the PCC further improves to 0.99. These results indicate that the improved FERM model is highly sensitive to the number of input DDM channels; with increasing data channel number, both evaluation metrics improve by more than 20%.

[0052] In summary, the FERM model inversion results outperform the RM model in overall performance and show stronger correlation with the DTU18 validation model. This indicates that both the spatial and channel modules effectively enhance the RM model's ability to extract inter-channel information and mine deep spatial information in DDM, thereby improving the accuracy of the SSH inversion results.

[0053] Figure 8 The left figure shows the statistical histograms of the error distribution of the two models under different conditions compared to the DTU18 validation model for SSH inversion. The statistical results show that under the single-channel input condition of the DDM, the errors of the two models between -5 and 5 m are 77.82% and 72.27%, respectively, and the inversion errors between -10 and 10 m are 95.45% and 93.19%, respectively. Under the four-channel input condition of the DDM, the errors of the two models between -5 and 5 m are 87.44% and 73.15%, respectively, and the inversion errors between -10 and 10 m are 97.56% and 93.60%, respectively. These results indicate that under both different conditions, the FERM model shows varying degrees of improvement in error distribution compared to the RM model.

[0054] Figure 8The right figure illustrates the SSH probability density function (PDF) of the two models under different conditions and the DTU18 validation model. The SSH results retrieved by the FERM model are largely consistent with those of the DTU18 validation model. However, under the DDM four-channel input condition, compared to the DTU18 validation model, the FERM model retrieves a much higher probability of SSH occurring between -36 to -25 m, -10 to -3 m, and 14 to 17 m, while the probability of SSH occurring between -53 to -40 m, 29 to 36 m, and 45 to 54 m is significantly lower. This is mainly due to the imbalanced distribution of data with various SSH values ​​in the training dataset.

[0055] This invention provides a statistical description of the performance metrics of these two models. Under different input channel numbers, the FERM model achieved better inversion results. Under single-channel conditions, the MAE of the RM model and the FERM model were 3.92 and 3.36, respectively, and the RMSE of the RM model and the FERM model were 5.50 and 4.77, respectively. The improvement rates of the two accuracy evaluation metrics were 14.28% and 13.27%, respectively, demonstrating the effectiveness of the spatial module in enhancing the model's spatial feature extraction capability. Under the DDM four-channel input condition, the MAE of the RM model and the FERM model were 3.84 and 2.56, respectively, and the RMSE of the RM model and the FERM model were 5.46 and 3.74, respectively. The PCC of the RM model and the FERM model were 0.98 and 0.99, respectively. The improvement rates of the three accuracy evaluation metrics were 33.33%, 31.50%, and 1.02%, respectively. Excluding the spatial module's enhancement of the model's spatial feature extraction capability, the improvement rates were 19.05% and 18.23%, respectively, demonstrating the effectiveness of the channel module in enhancing the model's channel feature extraction capability. This means that the proposed FERM model based on the FER residual structure effectively distinguishes and enhances the channel and spatial features of the DDM. Simultaneously, the FERM inversion results improve synchronously with the increase in the number of input channels. The feature enhancement module effectively and comprehensively improves the RM's ability to extract data features from the DDM, thus achieving excellent SSH inversion results.

[0056] However, the primary purpose of the CYGNSS satellite is not ocean altimetry, and it has not been optimized for ocean altimetry. Furthermore, due to the limited resolution of the DDM time delay, narrow receiver bandwidth, and weak peak antenna gain, the inversion error remains relatively large.

[0057] To demonstrate that the FERM model outperforms the traditional spaceborne GNSS-R SSH inversion method, the HALF method was implemented, and the inversion accuracy of the two methods was evaluated: The traditional SSH inversion method uses the HALF retracking method to calculate the time delay between the reflected signal and the direct signal. This method selects the position of 70% of the maximum correlation power of the leading edge of the correlated waveform as the retracking point, based on previous research experience. Furthermore, an error model was established to correct various errors present in the time delay measurement.

[0058] Figure 9 The error distribution statistics of the global SSH obtained by the FERM model and the HALF re-tracking method are shown, indicating that the FERM model provides more reliable inversion results. The FERM model and the HALF re-tracking method have 87.45% and 42.92% of their inversion errors between −5 and 5 m, respectively, and 97.57% and 75.24% of their inversion errors between −10 and 10 m, respectively.

[0059] FERM outperforms traditional re-tracking algorithms in terms of accuracy metrics MAE, RMSE, and PCC. The FERM model significantly improves the accuracy of SSH inversion. MAE and RMSE are reduced by 62.41% and 50.20%, respectively, while PCC is improved by 4.21%. This is mainly because traditional re-tracking methods compress single Raw DDM information into a scalar, which inherently loses a significant amount of information. Furthermore, due to complex sea surface conditions, this scalar cannot accurately describe SSH information. The attention-based FERM model proposed in this invention not only improves the utilization rate of input DDM information but also enhances effective features and suppresses interfering features by assigning attention weights to features. Therefore, better inversion accuracy can be achieved.

Claims

1. A multimodal GNSS-R altimeter measurement method based on attention mechanism feature enhancement residuals, characterized in that, First, a novel attention-based feature-enhanced residual multimodal (FERM) model is constructed. Second, the FERM model is used to perform height measurement. The FERM model is constructed by combining a feature-enhanced residual network (FER) based on an attention mechanism and a fully connected neural network (FCNN). The FER model uses channel attention modules and spatial attention modules to distinguish and enhance the effective features of the DDM channel and spatial data extracted from the residual structure, thereby improving the utilization rate of DDM information and enhancing the features by assigning attention weights to the features. This improves the inversion accuracy of FCNN for SSH from the perspective of input quality.

2. The attention-based feature-enhanced residual multimodal GNSS-R altimeter measurement method according to claim 1, characterized in that, The FERM model has two inputs. In the first input, based on the FER architecture, features are extracted from different DDM combinations through several residual blocks with 3×3 convolutional kernels. After feature extraction, the 2D DDM data is converted into 1D numerical data, transforming heterogeneous data into data with the same structure. The second input is for auxiliary parameters. The extracted features are fused with auxiliary parameters in an additional fully connected layer, and then the fused data is used in FCNN for SSH prediction. The input to the convolutional layer of the FERM model is... X 1, that is, different DDM combinations of CYGNSS, with auxiliary parameter input as follows: X 2. The model output is represented as: (1), In the formula, This indicates the output of FERM. L and N These represent the hidden layers determined by grid search and the number of neurons in each layer, respectively. W i and b i They represent the first i The weights and biases of each hidden layer This indicates the output of the FER module. This represents the activation function for each hidden layer.

3. The attention-mechanism-based feature-enhanced residual multimodal GNSS-R altimeter measurement method according to claim 2, characterized in that, The FER network used to extract effective features from the DDM in the FERM model is a novel feature-enhancing residual network block suitable for the DDM data structure. First, the residual block contains two 3×3 convolutional layers, each followed by a batch normalization layer and a ReLU activation function, as well as a feature enhancement layer. Second, by skipping these two convolutional processes in the cross-layer data path, the input is directly appended to the final ReLU activation function. To concatenate the outputs of the two convolutional layers, they must have the same shape as the input. If the number of channels needs to be changed, an additional 1×1 convolution is added to transform the input into the desired form before the summation operation. Furthermore, the DDM spatial size of L1 b is only 17×11. Pooling layers affect the depth of the convolutional layers, and further reduction of the spatial size will lead to the loss of a large amount of valuable information, so pooling layers are discarded. Finally, since the 17×11 DDM data structure is relatively simple, the original number of channels is too large to increase the computational cost of the model, so the number of channels needs to be reduced. An attention-based feature enhancement module is added after the second convolutional layer to enhance the ability of convolutional feature extraction and characterize the importance of effective information.

4. The attention-mechanism-based feature-enhanced residual multimodal GNSS-R altimeter measurement method according to claim 3, characterized in that, The feature enhancement layer in the FER network has two sub-blocks: the channel attention module and the spatial attention module. Since convolutional operations extract features by mixing cross-channel and spatial information, an attention-based feature enhancement module is used to apply effective features along both the channel and spatial axes. The channel and spatial attention modules are applied sequentially, allowing each feature map to learn different features along both the channel and spatial axes, thereby enhancing features by improving the utilization of DDM information and assigning attention weights to features. Specifically, the channel attention module generates channel attention feature maps by utilizing the relationships between channels in the feature maps. Each channel is treated as a feature extractor, and average and max pooling methods are used to compress the spatial dimension of the input feature maps to condense channel feature information. After pooling, the results are input into a fully connected network and merged to generate the final extracted inter-channel features, producing the output channel feature map. The model output is represented as follows: (2), In the formula, This represents the output channel feature map. This represents the activation function. and This indicates two hidden layers. This represents the average pooling channel feature map. The model represents the max-pooled channel feature map. The spatial attention module generates a spatial attention feature map by utilizing the importance of spatial information within the feature map channels. Each spatial feature is treated as a feature extractor. The average and max-pooling methods are used to compress the channel dimension of the input feature map to condense the spatial feature information. After pooling, a convolutional network is used to distinguish the importance of the spatial information in the input feature map and generate corresponding feature labels as in-channel features to generate the output spatial feature map. The model output is represented as follows: (3), where, This represents the output spatial feature map. This represents the activation function. This indicates a convolutional layer with a 3×3 kernel. This represents the spatial feature map of average pooling. This represents the feature map of the max pooling space.

5. The attention-mechanism-based feature-enhanced residual multimodal GNSS-R altimeter measurement method according to claim 4, characterized in that, The feature enhancement module structure based on the attention mechanism is arranged in the order of channel attention module first and then spatial attention module, and generates the final output feature-enhanced feature map. The model output is represented as follows: (4), In the formula, This represents the output feature enhancement feature map. This indicates a spatial attention module. This indicates the channel attention module. This represents the input feature map.

6. The attention-based feature-enhanced residual multimodal GNSS-R altimeter measurement method according to claim 1, characterized in that, Before performing the height measurement, in order to enhance the convergence of the neural network and improve the generalization ability of the model, the DDM was normalized. The DDM was standardized using the following formula: (5), In the formula, This represents the normalized DDM. Represents any DDM in the dataset. This represents the maximum pixel value of the corresponding DDM in the dataset. The auxiliary parameters also need to be normalized. The normalization formula for the auxiliary parameters is as follows: (6), In the formula, This represents the auxiliary parameters after the normalization operation. Represents any auxiliary parameter in the dataset. This represents the maximum value of the corresponding auxiliary parameter in the dataset.

7. The attention-mechanism-based feature-enhanced residual multimodal GNSS-R altimeter measurement method according to claim 6, characterized in that, The reliability of this method is tested using MAE, RMSE, and PCC. Smaller MAE and RMSE values ​​indicate a good fit between the predicted and actual values. The formulas are as follows: (9), (10), (11)。