Submarine topography super-resolution reconstruction method based on window displacement multi-source fusion

Through the window displacement multi-source fusion seabed topography super-resolution reconstruction method, the residual U-Net attention neural network model is used to fuse multi-source data and position encoding, which solves the problems of lack of geological interpretability and accuracy attenuation in the reconstruction results in the existing technology, and realizes high-precision seabed topography super-resolution reconstruction.

CN120746833APending Publication Date: 2025-10-03NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510867423.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing super-resolution reconstruction methods for seabed topography fail to effectively integrate physical quantities such as gravity anomalies and vertical deviations, resulting in a lack of geological interpretability in the reconstruction results, accuracy verification that is divorced from actual needs, and significant accuracy degradation when the resolution is increased, especially the insufficient ability to reconstruct medium-frequency terrain features.

Method used

A seabed topography super-resolution reconstruction method based on window displacement multi-source fusion is adopted. By constructing a residual U-Net attention neural network model, combining multi-source data such as gravity anomalies and vertical deviations with position encoding, the water depth prediction value of high-resolution grid points is output.

Benefits of technology

It significantly improves the reconstruction accuracy, ensures the accuracy stability when the resolution is increased, enhances the reconstruction ability of medium-frequency terrain features, provides geological interpretability, and breaks through the limitations of traditional image super-resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120746833A_ABST
    Figure CN120746833A_ABST
Patent Text Reader

Abstract

The invention discloses a submarine topography super-resolution reconstruction method based on window displacement multi-source fusion, and aims to solve the problem of precision attenuation caused by absence of multi-beam truth value evaluation and resolution improvement of an existing neural network submarine topography reconstruction method. The method comprises the following steps of: obtaining gravity anomaly, gravity vertical gradient, vertical line deviation component, submarine topography background and ship survey water depth multi-source data of a target sea area; constructing a residual U-Net attention neural network model, and performing training by taking multi-source data extracted by a 11 * 11 window and position codes as input; moving the trained model sampling window in a staggered manner according to a target resolution step length; extracting data points in the window during movement and injecting position codes; and outputting a high-resolution water depth predicted value based on the position code and the window data. The reconstruction precision is improved by fusing physical quantities, the multi-resolution robustness is guaranteed by a window displacement mechanism, and the medium-frequency feature reconstruction capability is enhanced by a residual U-Net attention model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of seabed topography surveying and mapping, and in particular relates to a seabed topography super-resolution reconstruction method based on window displacement multi-source fusion. Background Art

[0002] The high-resolution construction of seabed topography is of great value to underwater navigation and marine science research. Currently, it mainly relies on two methods: satellite gravity inversion and ship-borne acoustic detection. The satellite gravity inversion method has a wide coverage but low resolution, and can only achieve kilometer-level terrain reconstruction; ship-borne multi-beam detection has high accuracy but insufficient coverage, and has only completed the mapping of a small part of the world's sea areas. Existing technologies have constructed seabed topography products such as ETOPO and GEBCO by fusing multi-source data, but there are still many sea areas that lack measured data support. In recent years, the seabed topography super-resolution method based on neural networks has made progress, but there are fundamental defects:

[0003] First, existing methods treat digital seabed topography models as pure image processing, relying solely on low-resolution DBM grid data for super-resolution reconstruction. Physical quantities directly related to topographic structure, such as gravity anomalies and vertical deviations, are not incorporated, resulting in a lack of geological interpretability in the reconstruction results.

[0004] Second, the model evaluation uses high-resolution DBM data as a reference standard, rather than actual shipborne multi-beam sounding data, which makes the accuracy verification out of touch with actual application requirements.

[0005] Third, there is a significant loss of accuracy when the resolution is increased, especially the ability to reconstruct medium-frequency terrain features such as seamount edges and fault zones.

[0006] Furthermore, existing technologies fail to effectively integrate high-precision gravity data derived from the new SWOT satellites, limiting their ability to capture subtle terrain features. These shortcomings render current methods unable to meet the urgent need for high-precision, multi-resolution seafloor topography data in scenarios such as deep-sea exploration and geological disaster warning. Summary of the Invention

[0007] In order to solve the above technical problems, the present invention proposes a seabed topography super-resolution reconstruction method based on window displacement multi-source fusion to solve the problems existing in the above-mentioned prior art.

[0008] In a first aspect, to achieve the above-mentioned objectives, the present invention provides a method for super-resolution reconstruction of seabed topography based on window displacement multi-source fusion, comprising the following steps:

[0009] Acquire multi-source data of the target sea area, including gravity anomaly data, gravity vertical gradient data, vertical deviation component data, seabed topography background data, and ship-measured water depth data;

[0010] A residual U-Net attention neural network model was constructed, using multi-source data extracted from an 11×11 window and position encoding as input, and ship-measured water depth data as labels for model training.

[0011] The trained model sampling window is shifted according to the target resolution step size;

[0012] During the dislocation movement, all data points in the window are extracted and injected into the position code;

[0013] Based on the location coding and multi-source data within the window, the water depth prediction value of the high-resolution grid point is output.

[0014] In a second aspect, the present invention further provides a seabed topography super-resolution reconstruction system based on window displacement multi-source fusion, which is used to implement a seabed topography super-resolution reconstruction method based on window displacement multi-source fusion, and the system comprises:

[0015] Multi-source data acquisition module, used to obtain gravity anomaly data, gravity vertical gradient data, vertical deviation component data, seabed topography background data and ship-measured water depth data of the target sea area;

[0016] The neural network training module is used to build a residual U-Net attention neural network model, using multi-source data extracted from an 11×11 window and position encoding as input and ship-measured water depth data as labels for model training;

[0017] The window shift control module is used to shift the trained model sampling window according to the target resolution step size;

[0018] The position code injection module is used to extract all data points in the window and inject position codes during the window displacement movement process;

[0019] The super-resolution prediction module is used to output the water depth prediction value of high-resolution grid points based on position coding and multi-source data within the window.

[0020] In a third aspect, the present invention further provides a computer terminal device, comprising:

[0021] one or more processors;

[0022] a memory, coupled to the processor, for storing one or more programs;

[0023] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the method as described in any one of the first aspects.

[0024] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method described in any one of the first aspects are implemented.

[0025] Compared with the prior art, the present invention has the following advantages and technical effects:

[0026] The present invention provides a method for super-resolution reconstruction of seabed topography based on window-shift multi-source fusion. This method overcomes the problem that existing seabed topography super-resolution methods are out of physics. By integrating physical quantities such as gravity anomalies and vertical deviations with multi-beam true value evaluation, the reconstruction accuracy is significantly improved. A window-shift mechanism and position encoding are used to resolve data resolution differences and achieve accuracy stability when the resolution is improved. The residual U-Net attention model combines multi-head attention and jump connections to specifically enhance the reconstruction capability of medium-frequency terrain features and effectively suppress extreme terrain prediction errors. Dynamic weighted fusion of multi-source data provides geological interpretability, breaking through the limitations of traditional image super-resolution. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are provided to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0028] Figure 1 This is a distribution diagram of multi-beam sounding water depth data in the experimental sea area of ​​an embodiment of the present invention;

[0029] Figure 2 This is a flowchart illustrating a window shift super-resolution reconstruction solution according to an embodiment of the present invention;

[0030] Figure 3 This is a structural diagram of the residual U-Net attention model according to an embodiment of the present invention;

[0031] Figure 4 This is the reconstruction effect diagram of each seabed topography super-resolution model in an embodiment of the present invention (30" resolution);

[0032] Figure 5 This is a heat map of the prediction effect errors of each model in different sea areas according to the embodiment of the present invention (30” resolution);

[0033] Figure 6 This is a scatter point fitting diagram of the prediction effect errors of each model in different sea areas according to the embodiment of the present invention (30” resolution);

[0034] Figure 7 This is the rendering of the super-resolution reconstruction of each model of the seabed terrain in an embodiment of the present invention (15" resolution);

[0035] Figure 8This is the absolute error distribution diagram of the prediction effect of each model in different sea areas according to the embodiment of the present invention (15” resolution);

[0036] Figure 9 This is a scatter plot of the prediction error of each model in different sea areas according to the embodiment of the present invention (15” resolution);

[0037] Figure 10 This is a statistical diagram of the error interval distribution of the model results in different sea areas according to the embodiment of the present invention (30” resolution);

[0038] Figure 11 This is a statistical diagram of the error interval distribution of the model results in different sea areas according to the embodiment of the present invention (15” resolution);

[0039] Figure 12 The power spectrum density diagram (30” resolution) of the model results in different sea areas according to the embodiment of the present invention is shown in FIG.

[0040] Figure 13 This is a power spectrum density diagram (15” resolution) of the model results in different sea areas according to an embodiment of the present invention. DETAILED DESCRIPTION

[0041] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0042] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0043] High-resolution reconstruction of seafloor topography holds significant economic and scientific importance for humanity, providing crucial data support for underwater navigation and marine science research. Current neural network-based seafloor super-resolution methods generally rely on simple image super-resolution of the Digital Bathymetry Model (DBM). These methods fail to effectively incorporate physical quantities related to seafloor structure and lack objective and effective accuracy assessment based on multi-beam sounding data. To address this issue, this paper proposes a window-shifted seafloor super-resolution scheme based on multi-source fusion. A residual U-Net attention seafloor super-resolution model, adapted to this scheme and featuring powerful spatial feature extraction capabilities, is constructed. Experiments were conducted on 30- and 15-inch super-resolution reconstructions in two key sea areas with high spatial structural heterogeneity: the Xisha Islands and the Hawaii-Emperor Seamount Chain. Results demonstrate the effectiveness of both the window-shifted seafloor super-resolution scheme and the residual U-Net attention seafloor super-resolution model, achieving superior reconstruction accuracy at 15- and 30-inch resolutions in both areas compared to GEBCO and a basic neural network model. Notably, the super-resolution reconstruction accuracy based on this scheme does not significantly decrease with increasing resolution, and it is particularly effective for reconstructing medium-frequency terrain features. This invention introduces multi-beam detection as an objective criterion for evaluating reconstruction accuracy, providing a new approach for the study of local high-resolution seafloor topography reconstruction based on multi-source data fusion.

[0044] Super-resolution reconstruction of seafloor topography is a long-standing challenge in deep-sea exploration and mapping, with immense scientific and economic value. Accurate seafloor topography data helps geologists study plate tectonics, understand the dynamic history of ocean basins and seafloor geology, and effectively predict natural disasters such as seamount eruptions and tsunamis. Furthermore, as the lower boundary of the ocean, the structure of the seafloor influences ocean circulation, indirectly influencing Earth's climate and, consequently, the living environment for humans and other organisms. Shipborne acoustic sounding is currently the most accurate method for detecting seafloor topography. To date, only 25% of the world's oceans have been mapped to a resolution of 400 meters using shipborne acoustic measurements (Yu et al., 2024). For other areas not yet covered by shipborne acoustic measurements, depths are often determined using gravity-based depth inversion.

[0045] Since the 1990s, a series of studies have been conducted on inverting seafloor topography using ocean gravity data calculated from satellite sea surface height field data (Sandwell et al., 1997; Anderson et al., 1998). A representative example is the work of Smith and Sandwell et al., who used satellite altimetry and gravity anomaly data to develop a 15-km resolution S&S regression model for the region between 72°N and 72°S (Smith et al., 1994). These studies, combined with classical mathematical physics theory (Parker, 1973), exploited the strong correlation between the gravity field and seafloor geological structure to construct theoretical models for inverting seafloor topography using gravity. With the continuous advancement of satellite exploration technology, the quality of satellite altimetry data has continued to improve. The increasingly accurate ocean gravity field provides important data support for seafloor topography research. Scholars in the field have conducted numerous studies on seafloor topography inversion based on satellite altimetry (Luo Jia et al., 2002; Fan Diao, 2018; Sui Xiaohong et al., 2017; Hu Minzhang et al., 2012; Hu Minzhang et al., 2013; Ouyang Mingda et al., 2016; Li Qianqian et al., 2016). In 2024, the Sandwell team used the latest SWOT satellite altimetry data to calculate more refined and less noisy global ocean gravity data, and confirmed that this gravity field data can more accurately reflect the seafloor topography and geomorphology characteristics of the corresponding region (Yu et al., 2024).

[0046] Different methods of seabed topography detection have their own advantages and disadvantages. Satellite altimetry gravity theory inversion method has a wide coverage area but low accuracy, while shipborne bathymetry has high accuracy but small coverage area (Smith et al., 1993; et al., 2019), scholars have integrated a variety of seabed topography detection methods (direct ship survey, gravity inversion, etc.) and constructed seabed topography DBM products using the idea of ​​multi-source data fusion (Hu Minzhang et al., 2015; et al., 2019), with the ETOPO series (C. Amante et al., 2009), DTU series, S&S series, STO_IEU series (Fan et al., 2022), and GEBCO series (GEBCO group, 2023) being widely known. Although these data products are constantly being updated with new multibeam bathymetric data and improved in accuracy and resolution, 75% of the ocean area is still not covered by ship-based bathymetry. At current efficiency levels, achieving global multibeam coverage within ten years is unlikely (Yu et al., 2024).

[0047] In order to construct seabed topography more efficiently and accurately, scholars have begun to use artificial intelligence methods to conduct data-driven seabed topography construction research (Jena et al., 2013; Annan et al., 2022; Sun et al., 2022; Wan et al., 2023; Zhou et al., 2023; Zhou et al., 2024; Harper et al., 2024; Ge et al., 2025). These methods use deep learning algorithms to achieve breakthroughs in the accuracy of inverting seabed topography and the effectiveness of fusing multi-source gravity data, but they have never considered how to use limited data to construct seabed topography with higher resolution. In recent years, there have been several studies dedicated to DBM data super-resolution, such as ESRGAN (Sonogashira et al., 2020), which reconstructs high-resolution bathymetric data for the waters surrounding Japan; the DEM-SRNet global model (Zhang et al., 2022), which uses transfer learning to transfer land patterns to the seafloor; and the STFET model (Cai et al., 2023), which combines a deformable convolutional neural network (Deformable Conv Neural Network) with a Transformer. However, all of these super-resolution methods essentially perform simple image super-resolution on DBM grid data, using high-resolution DBM data as a reference for measuring reconstruction results, rather than actual multi-beam sounding data for comparison. Consequently, the model accuracy is insufficiently convincing, leaving significant room for improvement before practical application. Irisawa used the graph convolutional network PU-GCN to upsample the point cloud data of ship-surveyed water depth (Irisawa et al., 2024). This is the first method to use a neural network method to perform super-resolution reconstruction of ship-surveyed water depth data, and achieved good super-resolution reconstruction results. However, this method is only a preliminary exploration of super-resolution based on ship-surveyed data. The confidence level of the points generated by the model based on the training data is unknown and lacks reliability.

[0048] From the current research status in the above fields, it can be seen that in the task of super-resolution reconstruction of submarine topography based on artificial intelligence methods, most studies are limited to super-resolution research on visual images, ignoring the actual physical characteristics of the submarine topography. At the same time, few studies have introduced reliable multi-beam sounding water depth data to reasonably evaluate the effectiveness of the methods. At present, no relevant research has adopted the idea of ​​multi-source fusion to add gravity-related physical quantities in the super-resolution process to further improve the accuracy and interpretability of submarine topography super-resolution reconstruction. To address the above problems, this paper proposes a window-shift multi-source fusion submarine topography super-resolution reconstruction scheme and constructs a residual U-Net attention multi-source data fusion super-resolution model adapted to this scheme. This model introduces the latest gravity-related data (gravity anomaly, vertical deviation, and vertical gravity gradient) calculated based on SWOT satellite altimetry data (Yu et al., 2024) as external information into the 2x (30” resolution) and 4x (15” resolution) super-resolution reconstruction tasks. When evaluating the model effect calculation indicators, independent and objective high-resolution multi-beam sounding data that did not participate in the model training is used as the ground truth reference to ensure the practical value of the method. At the same time, in order to preliminarily test the adaptability of the method to different terrains, two regions with completely different terrain characteristics, China's Xisha area and the Hawaii-Emperor Seamounts, were selected to evaluate the model effect. The super-resolution reconstruction effect of the model was compared with the high-resolution GEBCO accuracy and other basic neural network models to comprehensively evaluate the effectiveness of the scheme.

[0049] Example 1

[0050] like Figure 1 As shown, this embodiment provides a method for super-resolution reconstruction of seabed terrain based on window displacement multi-source fusion, including:

[0051] Acquire multi-source data of the target sea area, including gravity anomaly data, gravity vertical gradient data, vertical deviation component data, seabed topography background data, and ship-measured water depth data;

[0052] A residual U-Net attention neural network model was constructed, using multi-source data extracted from an 11×11 window and position encoding as input, and ship-measured water depth data as labels for model training.

[0053] The trained model sampling window is shifted according to the target resolution step size;

[0054] During the dislocation movement, all data points in the window are extracted and injected into the position code;

[0055] Based on the location coding and multi-source data within the window, the water depth prediction value of the high-resolution grid point is output.

[0056] As an implementation method in this embodiment, the multi-source data includes:

[0057] Gravity anomaly data, gravity vertical gradient data, vertical deviation north component and east component data calculated based on SWOT satellite altimetry data;

[0058] Global open seafloor topography grid data products;

[0059] Raw water depth data from shipborne multi-beam soundings;

[0060] The ship-measured bathymetric data were filtered using the 3σ rule to eliminate extreme values ​​and interpolated to a grid aligned with the target resolution.

[0061] Specifically, the data involved is described below:

[0062] 1. Seafloor topography grid data products:

[0063] The present invention uses GEBCO2024 seafloor topography grid data as the background data for the model. GEBCO is a project jointly supported by the International Hydrographic Organization (IHO) and the Intergovernmental Oceanographic Commission (IOC) of UNESCO. It aims to provide global seafloor topography data and continuously update it to improve the accuracy of global seafloor topography mapping. Its resolution is 15 arc seconds, making it the latest and highest-resolution public global seafloor topography dataset in this field. In order to give full play to the deep learning model's ability to correct and optimize the errors of existing topographic products and obtain water depth data that is closer to multi-beam sounding values, high-quality product data is required as the basic background. GEBCO2024 obviously meets the needs of the present invention. The data can be obtained at the following link: https: / / www.gebco.net / data-products / gridded-bathymetry-data / .

[0064] 2. Ocean gravity anomaly data:

[0065] This paper incorporates the latest gravity anomaly data calculated from SWOT satellite altimetry data, published by a team at the Scripps Institution of Oceanography (SIO) at the University of California. This data, also gridded with a 1-arc-minute resolution, has been shown to more precisely reflect the correlation between gravity and small-scale features of the seafloor topography. Using this data as a feature can theoretically further improve the model's ability to capture subtle topographic features, thereby enhancing the accuracy of super-resolution reconstructions. The data can be accessed at the following link: https: / / topex.ucsd.edu / pub / archive / grav_SWOT / .

[0066] 3. Ocean gravity vertical gradient data:

[0067] This paper incorporates the latest vertical gravity gradient data calculated from SWOT satellite altimetry data, published by the Scripps Institution of Oceanography (SIO) at the University of California. This vertical gravity gradient data is gridded with a resolution of 1 arc minute. The vertical gravity gradient can better reflect the high-frequency characteristics of terrain. Including it as part of the data can further improve the model's performance in predicting steep terrain and increase physical interpretability. The data can be accessed at the following link: https: / / topex.ucsd.edu / pub / archive / grav_SWOT / .

[0068] 4. Ocean gravity vertical deviation north component and east component data:

[0069] The present invention introduces two types of vertical deviation north component and east component data. The first type is the latest vertical deviation north-south component and east-west component data calculated based on SWOT satellite altimetry data disclosed by the Scripps Institution of Oceanography (SIO) at the University of California. Both data are grid data with a resolution of 1 arc minute. The vertical deviation is the angle between the actual gravity direction in the earth's gravity field and the normal direction of the reference ellipsoid. It is an important reference standard in the conversion of elevation systems. The vertical deviation can reflect the change in the density of underground materials and reveal the tectonic activity of the earth's crust. The introduction of this physical quantity can provide more information for the model to capture more specific and subtle geological features. The data can be obtained at the following link https: / / topex.ucsd.edu / pub / archive / grav_SWOT / .

[0070] 5. Shipborne multi-beam water depth data:

[0071] The multi-beam bathymetric data used in this invention comes from two sources. One source is dense multi-beam bathymetric data from the Xisha Islands, China, provided by the Guangzhou Marine Geological Survey, a partner organization. This data corresponds to the Xisha Islands region of China used in the experiment, with a sampling interval of 300 meters. To align with other grid data points during model training and evaluation, this invention downsamples the data to a 15-arc-second grid, then uses data corresponding to a 1-arc-minute resolution for model training. High-resolution data points corresponding to 30-arc-second and 15-arc-second resolutions are only used during the performance evaluation process. The other source is multi-beam bathymetric data from the Emperor Seamount Chain, collected by the National Centers for Environmental Information (NCEI) under the National Oceanic and Atmospheric Administration (NOAA). This invention uses the data from these two sea areas and interpolates them to a 15-arc-second resolution grid. Similarly, only data corresponding to a 1-arc-minute resolution are used as labels for model training. High-resolution data points corresponding to 30-arc-second and 15-arc-second resolutions are used only for model performance evaluation. All multi-beam bathymetric data are pre-processed using the 3σ rule to eliminate extreme values. Figure 1The distribution of multibeam sounding bathymetric data for the two aforementioned sea areas is shown in Table 1. The basic parameters of the data for all sea areas are shown. The water depths of the Xisha Islands (16°–18°N, 110.5°–112.5°E) range from near-surface depth (0 m) to -2604.28 m, indicating the presence of coral platforms adjacent to deep-sea basins (standard deviation = 486.07 m) (Ma et al., 2011). The Emperor Seamounts (30°–50°N, 165°–175°E) are located within a flat plain (mean depth -7710.41 m, standard deviation 1155.26 m) interspersed with intermediate-depth seamounts (-279.18 m) (Watts et al., 1978). As shown in Table 1, statistical analysis reveals that the topographic features of these two sea areas differ significantly, and they also exhibit strong spatial heterogeneity. This further validates the model's ability to extract complex features from adjacent areas. NCEI data can be accessed at the following link: https: / / www.ncei.noaa.gov / products / seafloor-mapping.

[0072] Table 1

[0073]

[0074]

[0075] As an implementation method of this embodiment, the process of generating the position code includes:

[0076] Calculate the projection coordinates of the high-resolution prediction points on the low-resolution grid;

[0077] Rounding the projection coordinates as the window center coordinates;

[0078] The relative offset between the predicted point and the window center is divided by half the window length to generate a two-dimensional position encoding vector.

[0079] Specifically, the process includes:

[0080] To balance the resolution and accuracy of seabed topography reconstruction, the present invention proposes a window-shifted multi-source fusion terrain super-resolution reconstruction scheme for the first time. Limited by the fact that the SWOT gravity-derived data (gravity anomalies, vertical gravity gradients, and north and east components of vertical deviation) currently available for multi-source fusion reconstruction of seabed topography are all 1' resolution, the present invention can only ensure that a reliable terrain reconstruction model can be trained on 1' resolution data points under normal circumstances. The present invention uses all 1' resolution data for model training and verification, and uses high-resolution data points that were not involved in the training process as test points for the present invention. After the model is trained, its sampling window is shifted (for 30" super-resolution reconstruction tasks, the shift is 30" in step size, and for 15" super-resolution reconstruction tasks, the shift is 15" in step size). All available data points within the window coverage are used for sampling, and the relative offset between the window and the target sampling point is recorded as the model's position code, so that the model can more accurately locate its position during the window-shifted super-resolution reconstruction process. During the super-resolution prediction process, the present invention faces a common challenge: due to differences in grid resolution, the data points extracted by the window may not be accurately aligned with the predefined grid points, resulting in incomplete or misaligned input data. For example, when extracting a high-resolution window from low-resolution data (such as a 1' grid), these points may slightly offset the grid position, resulting in fewer valid data points than expected. To address this problem, the present invention adopts two strategies. First, through data filling technology (using the mean to fill the boundary area), it ensures that each window always maintains the full size (for example, 11×11 points) to avoid missing data. Secondly, the present invention introduces a position encoding mechanism to capture the relative offset of the point and the center of the window, and inputs the position encoding as part of the feature into the model during the process of dislocation displacement super-resolution reconstruction of the sliding sampling window, so that the model can accurately perceive the relative position of the target point during the window movement in the super-resolution stage. The definition of position encoding in the present invention is shown in the formula.

[0081]

[0082] in, It represents the projection position (x_pos, y_pos) of each high-resolution prediction point (i, j) on the low-resolution grid. center =round(x pos ),j center =round(y pos ), which represents the window center coordinates (i_center, j_center) obtained using the round function.

[0083] This approach not only improves the spatial robustness of the predictions, but also ensures that the model can produce accurate super-resolution results even in the presence of imperfect data alignment. Figure 2 The process of the window shift super-resolution reconstruction scheme is intuitively demonstrated.

[0084] To achieve detailed seafloor topography reconstruction, we designed a baseline model based on a multilayer perceptron (MLP) (Rumelhart et al., 1986). This model takes as input flattened multi-source window features, including gravity anomalies, gravity gradients, vertical deviation north and east components, and GEBCO background field data. To enhance the model's spatial capture for super-resolution reconstruction, a positional encoding is added, with an input dimension of 11×11×5+2. The model architecture consists of three fully connected layers with 256, 128, and 64 hidden units, respectively. Each layer is followed by a ReLU activation function and a dropout layer (Nitish et al., 2014) with a dropout rate of 0.3 to enhance generalization and prevent overfitting. The final layer outputs a single water depth value for regression prediction. The MLP design is simple and straightforward, effectively capturing the nonlinear relationship between input features and water depth, making it a suitable performance benchmark for complex models. However, its flattening of window features ignores the spatial structure of the data, potentially limiting its ability to model local spatial patterns. Nevertheless, the model is computationally efficient and parameter adjustment is intuitive, providing a reliable reference for the development of subsequent complex models.

[0085] To fully exploit the spatial characteristics of multi-source data, a seafloor topography model was designed based on the Convolutional Neural Network (CNN) architecture (Lecun et al., 1998). The model accepts a five-channel input of 11×11×5, with each channel corresponding to a data source (gravity anomaly, gravity gradient, vertical deviation north component, vertical deviation east component, and GEBCO background field data). The model consists of two layers of convolutional operations: the first layer uses 16 3x3 convolutional kernels, and the second layer uses 32 3x3 convolutional kernels. Each layer is followed by a Reluctant Unit (ReLU) activation function to introduce nonlinearity. Subsequently, adaptive global average pooling is used to compress the feature map into a fixed-size vector, reducing computational complexity and enhancing the model's adaptability to varying input sizes. Finally, the feature vector is mapped to the water depth prediction value through a two-layer fully connected network (32 units and 1 unit). A dropout layer (with a dropout rate of 0.3) is added to mitigate overfitting. Convolutional neural networks effectively capture local spatial patterns within a window and are more suitable for processing geophysical data with spatial correlation than multilayer perceptrons. It has fewer parameters and higher computational efficiency.

[0086] To adaptively adjust the importance weights of multi-source data fusion, this paper designs a basic attention model, incorporating the attention mechanism (Bahdanau et al., 2015) to enhance the model's ability to model diverse data sources. The model first constructs independent encoders for each data source (gravity anomaly, gravity gradient, vertical deviation north component, vertical deviation east component, and GEBCO background field data). Each encoder consists of two fully connected networks (with 128 and 64 hidden units) that map the flattened window features into a high-dimensional feature space. Subsequently, the feature vectors and positional encodings of the five data sources are concatenated, and the weights of each data source are calculated using an attention module. The attention module consists of two fully connected networks and uses the softmax function to generate normalized weights for weighted fusion of the features from each data source. Furthermore, the model introduces residual connections, which preserve the original feature information through linear projection and improve training stability. Finally, the weighted features are mapped to water depth predictions via a two-layer fully connected network. The core advantage of the basic attention model lies in its ability to dynamically adjust the contribution of each data source, providing an interpretable weight distribution and thus revealing the impact of different geophysical data on water depth prediction. This model is particularly suitable for multi-source data fusion scenarios, but under the independent encoder design, the interaction relationship between data sources may be ignored.

[0087] As an implementation method in this embodiment, the construction process of the residual U-Net attention neural network model includes:

[0088] Extract features through the encoder: perform convolution operations using three layers of residual convolution blocks, perform batch normalization and ReLU activation on each layer, and downsample through max pooling;

[0089] Apply multi-head attention mechanism through bottleneck layer;

[0090] The decoder performs upsampling, connects the skip link, and outputs the water depth value through 1×1 convolution.

[0091] As an implementation method in this embodiment, the execution process of the multi-head attention mechanism includes:

[0092] Modeling global dependencies on the feature maps output by the encoder;

[0093] Dynamically weighted fusion of multi-source data features.

[0094] Specifically, the residual U-Net attention neural network model includes:

[0095] In order to capture the spatial features of the input data more efficiently and accurately, the present invention constructs a residual U-Net attention model with position encoding for the task of super-resolution reconstruction of seabed terrain. Position encoding is additionally added in the process of inputting five types of multi-source data. The model uses a three-layer residual convolution block (output channel is [64, 128, 256]) for feature extraction, and each layer contains two convolutions, batch normalization and ReLU activation. Downsampling is performed through the maximum pooling layer, and the model applies a multi-head attention mechanism (8 heads) (Vaswani et al., 2018) in the bottleneck layer to enhance the dependency modeling between features. The decoder gradually recovers the features through upsampling and jump connections, and finally outputs the water depth prediction value through 1×1 convolution. This model uses U-Net for multi-scale feature extraction, adds residual connections to ensure training stability, and also has the feature fusion capability brought by the multi-head attention mechanism. It is specially designed to capture the complex spatial features of the input data. The detailed structure of the residual U-Net attention model is as follows Figure 3 As shown in Table 2, the basic information of all neural network models involved in the present invention is shown in Table 2.

[0096] Table 2

[0097]

[0098]

[0099] As an implementation method in this embodiment, the model training process includes:

[0100] Set the root mean square error as the loss function;

[0101] Add entropy regularization term to constrain the attention weight distribution;

[0102] The Adam optimizer is used to perform parameter updates, and a linear learning rate warm-up strategy and validation loss early stopping mechanism are applied.

[0103] Specifically, all models use the mean square error (MSE) loss function, and the loss function definition formula is as follows.

[0104]

[0105] Where N is the number of valid points, Pred is the predicted water depth output by the model, and Truth is the true multi-beam ship bathymetric data for the corresponding area. All neural network model training processes described in this invention use the Adam optimizer for parameter updates. The epoch count is set to 300, the initial learning rate is set to 0.0001, and a linear warm-up strategy is employed: the learning rate is gradually increased from 0.00001 to the target learning rate over the first 10 epochs to avoid initial gradient oscillation. An early stopping mechanism is introduced during training; training is terminated if the validation loss does not decrease for 50 consecutive epochs. For the attention model, an additional entropy regularization term -λH(w) (λ = 0.01) is added to constrain the distribution of attention weights and prevent over-concentration on a single data source. The learning rate is dynamically adjusted based on the validation loss using the ReduceLROnPlateau strategy, with a decay factor of 0.5 and a minimum learning rate cap of 0.000001. The batch size is uniformly set to 64, and the maximum number of training epochs is 300. All experiments were conducted on an NVIDIA Tesla A100 GPU platform.

[0106] The task of constructing seafloor topography differs from conventional visual imagery in two key ways. First, seafloor topography data has high variance and a wide range of values. During the feature extraction phase of 2D-to-2D topography reconstruction, small values ​​adjacent to large-valued areas are easily overlooked by the smoothing effect during normalization, resulting in a model that lacks the ability to capture detailed information. Second, the boundaries between topographic features in seafloor topography data are not sharp. What appear to be two different-colored areas are actually gradients, with fluctuations caused by gradient changes. 2D-to-2D topography reconstruction struggles to directly capture these subtle variations. For these two reasons, the present invention employs a field-to-point strategy during data preprocessing that achieves high precision for each point and accounts for spatial correlation. For each target prediction point, the present invention uses all features within an 11×11 grid around the target prediction point to construct the water depth at that point. Building on the GEBCO background topography data, the present invention introduces four additional features: gravity anomaly, gravity gradient, vertical deviation north component, and vertical deviation east component, to enhance model performance and interpretability. Therefore, the present invention organizes the data samples into a 5-channel 11×11 tensor mapping to the water depth at a single point. When extracting data points from an 11×11 window around each point, there is a potential problem of window crossing at the boundary points. To address this problem, the present invention adopts a global mean filling strategy, using the global mean of the channel to fill in the empty values ​​at the boundary window.

[0107] This method uses multibeam bathymetric data interpolated onto a high-resolution grid as a reference for ground truth. To objectively evaluate the model's effectiveness, a masking operation is performed to select only those points covered by multibeam bathymetric values ​​as valid points. All 1-minute resolution multibeam sounding data from these valid points is used as the training data label, while the remaining high-resolution multibeam sounding data serves as the benchmark for evaluating the effectiveness of high-resolution reconstruction. Table 3 shows the number of valid points at different resolutions. It should be noted that to ensure objective evaluation of the model's effectiveness, the 30-arc-second and 15-arc-second data points used in the subsequent super-resolution reconstruction verification process exclude all 1-arc-minute resolution points used in training.

[0108] Table 3

[0109]

[0110] Before model training, all data were z-score normalized to ensure stable training of the neural network. The formula for data normalization is shown below.

[0111]

[0112] Where X can be a matrix of any data channel, μ represents the mean of such data, and σ represents the standard deviation of such data.

[0113] The evaluation indicators of the present invention for data include root mean square error (RMSE) and mean absolute error (MAE), both of which are indicators used to measure the accuracy of the model. The formula for the root mean square error is as follows.

[0114]

[0115] The formula for mean absolute error is as follows:

[0116]

[0117] This paper uses both of these metrics to measure model accuracy, providing a more comprehensive picture of the model's accuracy. RMSE can amplify the impact of large errors during calculations, and since seabed topography construction itself involves large numerical problems, the RMSE metric provides a more sensitive reflection of model accuracy.

[0118] Verification of the effectiveness of the residual U-Net attention model based on the window moving multi-source fusion seabed terrain super-resolution reconstruction scheme:

[0119] The purpose of this part of the experiment is to test the effectiveness of the residual U-Net attention model under the window moving multi-source fusion seabed topography super-resolution reconstruction scheme, and analyze the experimental results from the perspective of accuracy. The present invention combines the gravity anomaly, vertical gravity gradient, vertical deviation north component, vertical deviation east component and GEBCO data corresponding to the two processed sea areas calculated based on SWOT sea surface height data into a data tensor input model of N×5×11×11 specifications (where N is the number of valid points), and performs model training and super-resolution reconstruction (30" resolution and 15" resolution) for different sea areas respectively. The specific process is as follows: First, during the training process, the present invention uses all 1' resolution gravity-related data, the aligned 1' resolution GEBCO data and 1' resolution multi-beam sounding data labels as training sets, and uses different neural network structures to train basic models in different sea areas. Then, the data sampling window of the model is shifted to achieve seabed topography reconstruction at high-resolution points (for the 30" super-resolution reconstruction task, the shift step is 30", and for the 15" super-resolution reconstruction task, the shift step is 15). During the super-resolution reconstruction process of the shifted displacement model, all data points within the 11×11 sampling window will be extracted as input data for the prediction model. At this time, the position encoding that records the relative position of the prediction point and the data point in the sampling window will also be added as part of the input to further enhance the spatial perception ability of the model during the high-resolution reconstruction process. Finally, the 30" and 15" super-resolution reconstruction data output by the model are compared with the pre-prepared high-resolution multi-beam sounding data and GEBCO data that were not involved in the training and interpolated to the grid points to evaluate the reconstruction accuracy.

[0120] 1. 30” resolution super-resolution reconstruction of seafloor topography:

[0121] The intuitive effect of the model on the super-resolution reconstruction of the seabed topography at 30” resolution is as follows Figure 4 As shown in the figure, the terrain reconstruction results of the residual U-Net attention model in both sea areas are closest to those of the multi-beam sounding data. The reconstruction of some edge areas has clearer boundaries than the GEBCO model and other neural network models. This initially reflects the effectiveness of the model.

[0122] In order to more intuitively reflect the comparison effect between models, the error heat map of the absolute error distribution of the model is drawn for the experiments of different models in different regions, such as Figure 4 As shown in , and the scatter fitting diagram that can more intuitively reflect the numerical fitting situation, such as Figure 5 As shown in the figure. The darker the color of the difference distribution graph, the smaller the difference between the model output and the multi-beam detection. The more concentrated the points of the scatter fitting graph are and the closer they are to the straight line, the stronger the model's fitting effect on the true value. Figure 6As shown in the figure, the absolute error maps of the super-resolution reconstruction using the residual U-Net attention model are closest to black (0) in both the Xisha Islands and the Emperor Seamount Chain. The scatter plot also shows that this model's points are relatively compact, with no obvious clustering of large deviation points, demonstrating the best true value fit.

[0123] Table 4 shows the output values ​​of different models and various statistical indicators of GEBCO on the test set for this region. Analysis of the data shows that both attention models significantly underperform the background field in terms of the model accuracy metrics MAE and RMSE, with RMSE decreasing by 62.1% and 66.6% respectively compared to the background GEBCO. Both models also outperform the classic multilayer perceptron and convolutional neural network models in terms of accuracy.

[0124] Table 4

[0125]

[0126]

[0127] Table 4 summarizes the quantitative evaluation results of various deep learning models for seafloor depth prediction at 30” resolution. The analysis compares the models with the ground truth multibeam soundings and the GEBCO terrain dataset using key metrics including mean absolute error (MAE), root mean square error (RMSE), mean, standard deviation, maximum, and minimum values.

[0128] Overall, the Residual U-Net Attention model performed well in both test areas, with significantly lower MAE and RMSE values ​​than other models and GEBCO. In the Xisha Islands, the model achieved a MAE of 23.90m and an RMSE of 63.18m, significantly outperforming GEBCO's MAE (85.82m) and RMSE (131.29m). In the Hawaiian-Emperor Seamount Chain, the model achieved a MAE of 52.62m and an RMSE of 141.75m, still significantly outperforming GEBCO's MAE and RMSE, with reductions of 10.5% and 55.5%, respectively. This demonstrates that the model significantly improves depth prediction accuracy in complex terrain. The RMSE, a metric more sensitive to large errors, showed a significantly lower reduction than that of GEBCO, demonstrating that the 30-inch high-resolution seafloor topography reconstructed by the Residual U-Net Attention model is more stable in areas prone to large errors.

[0129] In contrast, the simpler multi-layer perceptron (MLP), convolutional neural network (CNN), and basic attention models all performed poorly in super-resolution reconstruction of the two ocean areas. This indicates that these models lack the robustness to extract spatial correlations in the data for this type of dislocated and shifted super-resolution reconstruction task. The basic neural network methods used as baselines all exhibited inferior accuracy to the GEBCO data in super-resolution reconstruction of 30-inch resolution.

[0130] 2. 15” resolution super-resolution reconstruction of seabed topography:

[0131] The experimental results of the model for super-resolution reconstruction of seabed topography with a resolution of 15” are as follows: Figure 7 As shown in the figure, different sub-figures show the visual effects of different models on terrain reconstruction in different areas. It can be seen that for the Xisha Islands, only the attention model and the residual U-Net attention model achieve visual results comparable to the ground truth from multibeam soundings, while the GEBCO data clearly exhibits some noise. In the Hawaiian-Emperor Seamount Chain, the colors of the residual U-Net attention and GEBCO data are closer to the ground truth. This preliminary demonstration of the effectiveness of the residual U-Net attention model in the 15" resolution task.

[0132] In order to intuitively reflect the error comparison between models, the present invention draws an error heat map reflecting the absolute error distribution of the model for experiments with different models in different regions, such as Figure 8 As shown in the error heatmap, the absolute error distribution of the residual U-Net attention model in the 15" super-resolution task is almost completely black, indicating the distribution with the lowest absolute error value. The less effective attention model and GEBCO both have light-colored patches indicating larger errors. It can be initially noted from the figure that the GEBCO data has some large errors (manifested as sporadic bright spots), which affect its conclusions based on the root mean square error metric.

[0133] At the same time, the present invention also draws a scatter point fitting diagram that can intuitively show the numerical fitting situation, such as Figure 9As shown in the figure. Darker colors in the difference distribution plot indicate smaller discrepancies between the model output and multibeam detection, while closer concentrations of points in the scatter plot indicate stronger model fit. The scatter plots reveal that both the Residual U-Net Attention and GEBCO control groups achieve relatively good ground truth fit in both sea areas. In the Xisha Islands, the Residual U-Net Attention data point distribution is significantly closer to a straight line and more compact, while the GEBCO distribution is more dispersed, indicating a higher average error level. In the Emperor Seamount Chain region, while the Residual U-Net Attention and GEBCO methods generally offer similar point-to-line fits, GEBCO exhibits a denser cluster of outlier errors, corresponding to the scattered bright spots found in the error heatmap above, indicating the presence of large errors.

[0134] Table 5

[0135]

[0136]

[0137] Table 5 shows the quantitative evaluation results of various deep learning models for seafloor depth prediction at 15” resolution. The analysis compares the models with the ground truth multibeam soundings and the GEBCO terrain dataset using key metrics including mean absolute error (MAE), root mean square error (RMSE), mean, standard deviation, maximum, and minimum values.

[0138] In the Xisha Islands, China, the Residual U-Net Attention model performed exceptionally well, achieving the lowest MAE of 21.85m and RMSE of 82.42m. This represents a 75.2% and 43.3% decrease compared to GEBCO's MAE (88.17m) and RMSE (145.45m), respectively. This demonstrates the significant effectiveness of the Residual U-Net Attention model on the 15" resolution super-resolution task. The accuracy of other neural network models on the 15" resolution super-resolution reconstruction task was significantly weaker than that of the Residual U-Net Attention model.

[0139] In the Hawaiian-Emperor Seamounts region, model performance varied significantly. The Residual U-Net Attention model again performed exceptionally well, achieving a MAE of 44.94m and an RMSE of 126.34m. These performances were significantly lower than those of the GEBCO (58.77m) and RMSE (318.32m) in this region, representing decreases of 23.5% and 60.3%, respectively. Other neural network models with a single architecture still performed poorly on the 15" resolution super-resolution reconstruction task, exhibiting larger errors and performing worse than the GEBCO model.

[0140] Analysis in different sea areas shows that the residual U-Net attention model consistently outperforms other models in super-resolution reconstruction of seabed topography at 15” resolution, with the lowest MAE and RMSE indicators. It is also worth noting that all experimental data in the two sea areas show that the accuracy of the residual U-Net attention model in super-resolution reconstruction at 15” resolution has not been significantly attenuated due to the increase in resolution. The window staggered super-resolution reconstruction scheme based on the residual U-Net attention model can improve the reconstruction resolution of the target sea area while ensuring that its accuracy relative to the true label remains reliable. This is due to its powerful spatial feature extraction capabilities and flexible multi-source data integration capabilities based on the attention mechanism.

[0141] 3. Stability evaluation of super-resolution reconstruction performance of seafloor topography with moving windows at multiple resolutions:

[0142] The above experimental results verify the excellent performance of the residual U-Net attention model under the window displacement-based seabed topography super-resolution reconstruction scheme at different resolutions. The present invention will combine the analysis of the spatial domain and the frequency domain to further illustrate the stability of the reconstruction performance of this scheme under multiple resolutions. Table 6 shows the comparison of the residual U-Net attention performance at different resolutions in two sea areas. It is found that as the target resolution of the super-resolution task increases, the reconstruction accuracy performance not only does not significantly attenuate, but has better performance in the MAE indicator. This shows that the window displacement seabed topography super-resolution scheme has the robustness of multi-resolution reconstruction. Under this scheme, the super-resolution reconstruction model trained using 1' resolution data has fully grasped the terrain pattern of the corresponding sea area. Only in the Xisha Sea area of ​​China does the root mean square error index increase with the increase of resolution.

[0143] Table 6

[0144]

[0145] To further illustrate the data phenomenon in the table, Figure 10 and Figure 11The statistics of the proportion of super-resolution error points of each model in different ranges at 30" and 15" resolutions are shown respectively. In the Xisha Sea area of ​​China, the proportion of data points with errors greater than 300 meters increases with the increase in resolution (the ratio of error points greater than 300 meters for the Residual U-Net Attention model increases from 0.8% to 0.9%). After increasing the resolution, the increase in the proportion of large errors causes the RMSE indicator, which is sensitive to large errors, to increase significantly. In the Hawaiian-Emperor Seamount Chain area, the proportion of data points with errors greater than 300 meters decreases with the increase in resolution (the ratio of error points greater than 300 meters for the Residual U-Net Attention model decreases from 0.2% to 0.1%). The RMSE change trend in this sea area is consistent with the MAE decline. Among all model error distributions, the Residual U-Net Attention model has a much smaller proportion of large errors than other models. The effective reduction of large errors is also the fundamental reason why it has the best RMSE indicator effect. It can also be found that in the Emperor Seamount Chain region, the errors of all models in the range of 10-20m are over 90%. This shows that although all models have effectively grasped the terrain pattern information of this sea area, the key to the large differences in the performance of each model lies in the model's ability to reduce large errors. The stronger spatial feature extraction capability of the residual U-Net attention model plays an important role here.

[0146] In order to more comprehensively evaluate and illustrate the model effects, the present invention performs Fast Fourier Transform (FFT) on the predicted values ​​and multi-beam true values ​​output by different models in different sea areas, and analyzes the performance of the model from a frequency domain perspective. Figure 12 and Figure 13 The power spectrum density diagrams of each model in different sea areas in the 30" and 15" resolution reconstruction tasks respectively reflect the power density under different frequency terrain features. For the Xisha Sea area of ​​China, the 15" resolution power spectrum density diagram shows more detailed information (more peaks) in the medium frequency area compared to the 30" resolution power spectrum density diagram, and the power density value in the high frequency area is significantly increased (the power exceeds 10 7dB), indicating that the increased resolution in this area shows a significant increase in mid- and high-frequency detail. In China's Xisha Sea area, the power spectral density curves of the model output values ​​of the residual U-Net attention model have the highest overlap with the power spectral density curves of the true values, regardless of whether the resolution is 15" or 30". This demonstrates the excellent performance of the residual U-Net attention model in the frequency domain. It can also be noted that the power spectral density curves of all model output values ​​and the true values ​​have a high degree of overlap in the low-frequency region (0-50) and the high-frequency region (300-350). They all show a significant difference in performance in the medium-frequency region, indicating that the window-shift seafloor topography super-resolution scheme combined with the residual U-Net attention model has a targeted effect on the high-resolution reconstruction of mid-frequency topographic features.

[0147] For the Emperor Seamount Chain region, high-frequency and low-frequency terrain features have higher power. The 15” resolution power spectrum density map also shows more detailed information in the medium-frequency region (several peaks in the 500 to 2500 frequency range) compared to the 30” resolution power spectrum density map. Combined with the statistics of error distribution mentioned above, it can be seen that the error rate within 20m for all models here is more than 90%. The key point to distinguish the performance of model accuracy indicators lies in large-value errors. When this situation is converted to the frequency domain space, it can be seen that the true value and predicted power spectrum density curves of almost all models are highly overlapped, because a small number of scattered large-value errors will not appear in the power spectrum density curve. The above power spectrum density map shows that the model will not fail due to the increase in resolution in different terrain feature frequency bands, which shows that the sliding window seabed topography high-resolution reconstruction scheme is robust to different resolution reconstruction performance in the target area.

[0148] In summary, this paper addresses the challenges of super-resolution reconstruction of submarine topography by proposing a neural network reconstruction scheme with window offset and shifting. A residual U-Net attention model is constructed to achieve super-resolution reconstruction at 2x (30") and 4x (15") resolution. A multi-source fusion model is first trained using all 1'×1' resolution data from a specified area. The trained model is then offset and shifted to use all data within an 11×11 grid centered on the high-resolution point (SWOT-derived gravity anomalies, vertical gravity gradients, east and north components of vertical deviation, and GEBCO data) as input to reconstruct water depth at the high-resolution point. This scheme is tested in two areas, the Xisha Islands and the Hawaiian-Emperor Seamount Chain, using multibeam ship-based data as a benchmark for error metrics. Experimental results demonstrate that this method achieves excellent super-resolution reconstruction results in both test areas, significantly outperforming single-structure neural networks (MLP, CNN, and basic attention models) and outperforming the GEBCO dataset, which is currently commonly used in the field. In the 30” resolution terrain reconstruction experiment, the performance of the residual U-Net in the test area decreased by 51.9% (China's Xisha area) and 55.5% (Hawaii-Emperor Seamount Chain area) compared with the RMSE indicators of GEBCO; in the 15” resolution terrain reconstruction experiment, the performance of the residual U-Net in the test area decreased by 43.3% (China's Xisha area) and 60.3% (Hawaii-Emperor Seamount Chain area) compared with GEBCO. The experimental results at different resolutions further show that this method has strong applicability in 2x and 4x resolution super-resolution tasks. Detailed error statistics and frequency domain analysis further demonstrate that the accuracy of the reconstructed terrain based on the window moving seabed topography super-resolution reconstruction scheme will not be significantly attenuated due to the increase in resolution, and it has multi-resolution robustness, and this scheme has a targeted effect on the reconstruction of medium-frequency terrain features. Unlike the simple basic neural network used in the experimental comparison group, the residual U-Net attention model combines the spatial feature extraction capability of the convolutional layer and the multi-source data integration capability of the attention mechanism's adaptive weights. It has powerful spatial feature extraction and neighborhood generalization capabilities. These characteristics make it suitable for the super-resolution terrain construction scheme of the dislocated window movement. At the same time, the present invention adds a dynamic input layer and dynamic encoding, so that the model can stably extract features within the window range during the dislocated movement and more accurately locate the positional relationship between the data sampling window and the target point.

[0149] Based on this, an embodiment of the present invention provides a method for super-resolution reconstruction of seabed topography based on window displacement and multi-source fusion. The present invention overcomes the problem that existing seabed topography super-resolution methods are out of touch with physical laws, and significantly improves reconstruction accuracy by fusing physical quantities such as gravity anomalies and vertical deviations with multi-beam true value evaluation. A window displacement mechanism and position encoding are used to resolve data resolution differences and achieve accuracy stability when the resolution is improved. The residual U-Net attention model combines multi-head attention and jump connections to specifically enhance the reconstruction capability of medium-frequency terrain features and effectively suppress extreme terrain prediction errors. Dynamic weighted fusion of multi-source data provides geological interpretability, breaking through the limitations of traditional image super-resolution.

[0150] Example 2

[0151] Based on the same general inventive concept, the present invention also provides a seabed topography super-resolution reconstruction system based on window displacement multi-source fusion. The seabed topography super-resolution reconstruction system based on window displacement multi-source fusion provided by the present invention is described below. The seabed topography super-resolution reconstruction system based on window displacement multi-source fusion described below can be referenced to the seabed topography super-resolution reconstruction method based on window displacement multi-source fusion described above. The system includes:

[0152] Multi-source data acquisition module, used to obtain gravity anomaly data, gravity vertical gradient data, vertical deviation component data, seabed topography background data and ship-measured water depth data of the target sea area;

[0153] The neural network training module is used to build a residual U-Net attention neural network model, using multi-source data extracted from an 11×11 window and position encoding as input and ship-measured water depth data as labels for model training;

[0154] The window shift control module is used to shift the trained model sampling window according to the target resolution step size;

[0155] The position code injection module is used to extract all data points in the window and inject position codes during the window displacement movement process;

[0156] The super-resolution prediction module is used to output the water depth prediction value of high-resolution grid points based on position coding and multi-source data within the window.

[0157] As an implementation method of this embodiment, the multi-source data acquisition module includes:

[0158] Satellite gravity data unit, used to obtain gravity anomaly, gravity vertical gradient, vertical deviation north component and east component data calculated based on SWOT satellite altimetry data;

[0159] Terrain background unit, used to load global public seabed terrain grid data products;

[0160] The ship survey data unit is used to collect the original water depth data of shipborne multi-beam detection, eliminate extreme values ​​​​by the 3σ rule, and interpolate to the target resolution grid.

[0161] As an implementation method of this embodiment, the position code injection module includes:

[0162] Projection coordinate calculation unit, used to calculate the projection coordinates of high-resolution prediction points on the low-resolution grid;

[0163] a window center positioning unit, configured to round the projection coordinates as the window center coordinates;

[0164] The offset normalization unit is used to divide the relative offset between the prediction point and the window center by half the window length to generate a two-dimensional position encoding vector.

[0165] As an implementation in this embodiment, the neural network training module includes:

[0166] The feature extraction unit uses three layers of residual convolution blocks to perform convolution operations, with batch normalization and ReLU activation in each layer, and downsampling through maximum pooling;

[0167] Attention fusion unit, which applies multi-head attention mechanism at the bottleneck layer;

[0168] The terrain reconstruction unit is connected to the skip link through upsampling operation and outputs the water depth value through 1×1 convolution.

[0169] As an implementation in this embodiment, the attention fusion unit includes:

[0170] A global dependency modeling unit, used to establish global dependencies on the feature maps output by the encoder;

[0171] Dynamic weighting unit, used to fuse multi-source data features and assign weights.

[0172] As an implementation in this embodiment, the neural network training module further includes:

[0173] Loss function configuration unit, set the root mean square error as the loss function;

[0174] Regularization constraint unit, adding entropy regularization term to constrain the attention weight distribution;

[0175] Optimize the control unit, use the Adam optimizer to perform parameter updates, configure the linear learning rate warm-up strategy and the verification loss early stopping mechanism.

[0176] It should be understood that the seabed topography super-resolution reconstruction system based on window displacement multi-source fusion provided by the embodiment of the present invention has all the advantages of the seabed topography super-resolution reconstruction method based on window displacement multi-source fusion provided by the above embodiments.

[0177] In this embodiment, a computer terminal device is further provided, including:

[0178] one or more processors;

[0179] a memory, coupled to the processor, for storing one or more programs;

[0180] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the above-mentioned seabed topography super-resolution reconstruction method based on window displacement multi-source fusion.

[0181] In this embodiment, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned seabed topography super-resolution reconstruction method based on window displacement multi-source fusion are implemented.

[0182] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for super-resolution reconstruction of seabed topography based on window displacement multi-source fusion, characterized in that: The following steps are involved: Acquire multi-source data of the target sea area, including gravity anomaly data, gravity vertical gradient data, vertical deviation component data, seabed topography background data, and ship-measured water depth data; A residual U-Net attention neural network model was constructed, using multi-source data extracted from an 11×11 window and position encoding as input, and ship-measured water depth data as labels for model training. The trained model sampling window is shifted according to the target resolution step size; During the dislocation movement, all data points in the window are extracted and injected into the position code; Based on the location coding and multi-source data within the window, the water depth prediction value of the high-resolution grid point is output.

2. The method according to claim 1, characterized in that The multi-source data includes: Gravity anomaly data, gravity vertical gradient data, vertical deviation north component and east component data calculated based on SWOT satellite altimetry data; Global open seafloor topography grid data products; Raw water depth data from shipborne multi-beam soundings; The ship-measured bathymetric data were filtered using the 3σ rule to eliminate extreme values ​​and interpolated to a grid aligned with the target resolution.

3. The method according to claim 1, characterized in that The generation process of the position code includes: Calculate the projection coordinates of the high-resolution prediction points on the low-resolution grid; Rounding the projection coordinates as the window center coordinates; The relative offset between the predicted point and the window center is divided by half the window length to generate a two-dimensional position encoding vector.

4. The method according to claim 1, wherein The construction process of the residual U-Net attention neural network model includes: Extract features through the encoder: perform convolution operations using three layers of residual convolution blocks, perform batch normalization and ReLU activation on each layer, and downsample through max pooling; Apply multi-head attention mechanism through bottleneck layer; The decoder performs upsampling, connects the skip link, and outputs the water depth value through 1×1 convolution.

5. The method according to claim 4, characterized in that The execution process of the multi-head attention mechanism includes: Modeling global dependencies on the feature maps output by the encoder; Dynamically weighted fusion of multi-source data features.

6. The method according to claim 1, characterized in that The model training process includes: Set the root mean square error as the loss function; Add entropy regularization term to constrain the attention weight distribution; The Adam optimizer is used to perform parameter updates, and a linear learning rate warm-up strategy and validation loss early stopping mechanism are applied.

7. A seabed topography super-resolution reconstruction system based on window displacement multi-source fusion, characterized in that: The system comprises: Multi-source data acquisition module, used to obtain gravity anomaly data, gravity vertical gradient data, vertical deviation component data, seabed topography background data and ship-measured water depth data of the target sea area; The neural network training module is used to build a residual U-Net attention neural network model, using multi-source data extracted from an 11×11 window and position encoding as input and ship-measured water depth data as labels for model training; The window shift control module is used to shift the trained model sampling window according to the target resolution step size; The position code injection module is used to extract all data points in the window and inject position codes during the window displacement movement process; The super-resolution prediction module is used to output the water depth prediction value of high-resolution grid points based on position coding and multi-source data within the window.

8. The system according to claim 7, characterized in that The multi-source data acquisition module includes: Satellite gravity data unit, used to obtain gravity anomaly, gravity vertical gradient, vertical deviation north component and east component data calculated based on SWOT satellite altimetry data; Terrain background unit, used to load global public seabed terrain grid data products; The ship survey data unit is used to collect the original water depth data of shipborne multi-beam detection, eliminate extreme values ​​​​by the 3σ rule, and interpolate to the target resolution grid.

9. A computer terminal device, characterized in that: include: one or more processors; a memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Gravity anomalous body inversion method based on PMU-Net deep learning network

    CN120910810A

  • Gravity anomaly body inversion method based on PMU-net deep learning network

    CN120910810B

  • Large-scene water depth model reconstruction method based on sparse single-beam sounding data constraint

    CN121190698A

  • Water depth mapping method and system based on multi-source data fusion analysis

    CN122156351A