A water depth inversion system and method for turbid nearshore waters
Patent Information
- Application Number
- CN202611009408.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-29
AI Technical Summary
[0017]本发明的目的在于提供一种面向浑浊近岸水域的水深反演系统及方法,用于解决现有技术中卫星水深反演技术在复杂近岸环境中存在的空间非平稳性表达不足、深度区间差异建模不足、分区与回归割裂以及缺少预测不确定性输出等问题
[0042]与现有技术相比,根据本发明的一种面向浑浊近岸水域的水深反演系统及方法,能够提高整体水深反演精度并显著缓解空间非平稳性问题。
Smart Images

Figure CN122835336A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of marine surveying and mapping technology, and in particular to a water depth inversion system and method for turbid nearshore waters. Background Technology
[0002] Nearshore shallow water depth information is crucial foundational data for waterway maintenance, port operations, coastal engineering, sediment transport analysis, marine environmental monitoring, and disaster risk assessment. Current high-precision water depth measurements typically rely on single-beam / multi-beam sonar bathymetry, shipborne surveying, or airborne lidar bathymetry. While these active measurement methods offer high accuracy, they require specialized equipment and fieldwork, and suffer from drawbacks such as high cost, long turnaround times, significant susceptibility to sea conditions and weather, and difficulty in providing high-frequency coverage of large nearshore areas.
[0003] To reduce measurement costs and improve spatial coverage, satellite remote sensing water depth retrieval technology has gradually become an important supplementary means of acquiring nearshore water depth. Existing satellite water depth retrieval methods mainly include the following categories:
[0004] Empirical or semi-empirical models: These models are based on the statistical relationship between reflectivity in the blue, green, and red visible light bands and water depth, or establish water depth inversion formulas through logarithmic band ratios. These methods are simple in structure, computationally efficient, and effective in clear, relatively uniform shallow water areas.
[0005] Semi-analytical physical models: Based on the theory of water radiative transfer, these models introduce prior parameters such as water composition, sediment reflectivity, suspended matter, chlorophyll, and atmospheric correction to model the optical processes of a water column. While the physical meaning of this type of method is relatively clear, obtaining the parameters is difficult, and its engineering application is costly.
[0006] Traditional machine learning models, such as random forests, support vector regression, and gradient boosting trees, learn the nonlinear mapping relationship between remote sensing reflectance and water depth by constructing features such as spectral indices, band ratios, and spatial descriptors.
[0007] Deep learning models, such as convolutional neural networks, U-Net, Transformer, or attention mechanism models, automatically learn spectral-spatial features through local image patches or multi-scale feature extraction, improving the ability to retrieve water depth in complex nearshore environments.
[0008] In some research or engineering processes, K-means clustering, optical partitioning, or artificial region division are first used to divide the water area into different optical sub-regions, and then water depth regression models are established for each sub-region. This "partitioning first, regression later" approach can alleviate the heterogeneity problem caused by different water conditions to some extent.
[0009] However, existing technologies still have the following shortcomings in turbid nearshore waters, especially in areas where harbor basins, channels, tidal flats, and nearshore engineering facilities are interspersed:
[0010] First, the global unified mapping capability is insufficient. Existing empirical models, machine learning models, or ordinary CNN models usually assume that there is a relatively stable "reflectivity-water depth" relationship in the same image scene. However, in turbid water bodies, suspended sediment, bottom sediment type, hydrodynamic conditions, and human disturbances will cause the same water depth to correspond to different spectral responses, making it difficult for a single global model to adapt.
[0011] Second, ordinary convolutional networks focus on local features and are insufficient in modeling cross-regional dependencies. CNNs are good at extracting the spatial-spectral texture of local image patches, but their receptive field is limited, making it difficult to capture the correlation between regions that are far apart but have similar optical conditions, and also making it difficult to fully characterize the regional differences between harbor basins, channels and nearshore shoals.
[0012] Third, traditional clustering partitioning is disconnected from regression models. Existing optical partitioning or K-means clustering is usually treated as an independent step before modeling. The clustering results cannot be updated in reverse according to the water depth prediction error, nor can they form a closed-loop optimization with attention weights, deep feature learning, and uncertainty estimation.
[0013] Fourth, there is a lack of adaptive representation for depth ranges. Shallow water is strongly affected by bottom reflection, while medium and deep water are more significantly affected by water column absorption and scattering. The spectral attenuation mechanisms differ across different depth ranges. If existing models use a uniform feature representation, they are prone to range-specific biases in shallow, transitional, and deep water sections.
[0014] Fifth, there is a lack of pixel-level prediction confidence output. Most water depth inversion models only output a single water depth value and cannot simultaneously provide prediction uncertainty or confidence maps. As a result, when applied in high-risk areas such as waterways and harbor basins, users find it difficult to determine which areas have reliable results and which areas require careful interpretation or supplementary field measurements.
[0015] In addition to the problems mentioned above, existing optical partitioning or clustering-based depth inversion methods typically employ a two-stage processing flow: first, the study area is partitioned based on human experience, K-means clustering, or other rules; then, depth inversion models are established separately for each sub-region. While this approach can mitigate the impact of regional optical differences to some extent, the partitioning or clustering process is independent of the subsequent depth prediction process. It cannot update the partitioning results based on the final prediction error, and it is difficult to achieve unified optimization with attention weights, deep feature learning, and uncertainty estimation.
[0016] The information disclosed in this background section is intended only to enhance the understanding of the overall background of the invention and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0017] The purpose of this invention is to provide a water depth inversion system and method for turbid nearshore waters, which solves the problems of insufficient spatial nonstationarity expression, insufficient modeling of depth interval differences, fragmentation of partitioning and regression, and lack of prediction uncertainty output in the existing satellite water depth inversion technology in complex nearshore environments.
[0018] To achieve the above objectives, the present invention provides a water depth inversion system for turbid nearshore waters, comprising:
[0019] The data acquisition and preprocessing unit is used to acquire and process multispectral remote sensing images covering the nearshore waters of the target and measured water depth data of the same area to form a sample set including multiple target sample points; and to crop a local image block of a fixed window size from the multispectral remote sensing image with the pixel corresponding to each target sample point in the sample set as the center.
[0020] The feature extraction unit is used to input the local image block into the backbone module of the convolutional neural network to extract multi-scale spectral spatial features;
[0021] The coordinate region encoding unit is used to normalize the latitude and longitude coordinates of each target sample point, input them into the multilayer perceptron and generate a coordinate embedding vector, and fuse the coordinate embedding vector with the multi-scale spectral spatial features to obtain the region fusion feature.
[0022] The region-aware attention modeling unit is used to input the region fusion features into a two-layer heterogeneous multi-head attention module to establish long-distance dependencies between different spatial locations and different feature dimensions, and introduce residual connections and normalization operations after each attention layer to output region-aware features; wherein, the first-layer attention module is used to capture diverse region associations, and the second-layer attention module is used to integrate and abstract the feature relationships output by the first layer.
[0023] A deep-aware clustering aggregation unit is used to set up several learnable deep prototypes, calculate modality weights based on the similarity between the region-aware features of the sample and different deep prototypes, and perform weighted aggregation of the deep prototypes based on the modality weights to obtain deep-aware features; wherein, each deep prototype represents a potential deep-related feature modality.
[0024] The water depth prediction and uncertainty estimation unit is used to input the depth sensing features into two parallel output heads; wherein, the two parallel output heads include a water depth prediction output head and an uncertainty output head; the water depth prediction output head is used to predict the water depth, and the uncertainty output head outputs the prediction variance based on the heteroscedastic Gaussian assumption.
[0025] The multispectral remote sensing image is a Sentinel-2 Level-2A surface reflectance product; the processing of the multispectral remote sensing image covering the nearshore waters of the target includes: quality screening, image synthesis, and quality control of measured water depth data.
[0026] The coordinate region encoding unit fuses the coordinate embedding vector with the multi-scale spectral spatial features by concatenating the coordinate embedding vector with the multi-scale spectral spatial features along the feature dimension; the coordinate embedding vector is used to characterize the spatial prior information of the geographical sub-region where the sample is located.
[0027] The first layer attention module uses a first number of attention heads for global dependency modeling, and the second layer attention module uses a second number of attention heads for feature integration and abstraction; the first number of attention heads is greater than the second number of attention heads.
[0028] The modal weights are obtained by using the similarity between region-aware features and each depth prototype, and then transforming them using a soft assignment function.
[0029] The coordinate region encoding unit, the region-aware attention modeling unit, the depth-aware clustering aggregation unit, and the water depth prediction and uncertainty estimation unit all participate in end-to-end backpropagation training.
[0030] The training loss function under the heteroscedastic Gaussian assumption is as follows:
[0031]
[0032] in, To measure the actual water depth, To predict water depth, To predict variance.
[0033] The water depth prediction and uncertainty estimation unit is further configured to: after the model training is completed, input the effective water body pixels into the system pixel by pixel, and output a spatially continuous water depth prediction map and a corresponding uncertainty map, wherein the uncertainty map represents the water depth prediction confidence at each pixel location.
[0034] In one embodiment of the present invention, a method for water depth inversion in turbid nearshore waters includes:
[0035] S1. Acquire and process multispectral remote sensing images covering the nearshore waters of the target and measured water depth data of the same area to form a sample set including multiple target sample points; take the pixel corresponding to each target sample point in the sample set as the center and crop a local image block of fixed window size in the multispectral remote sensing image.
[0036] S2. Input the local image block into the backbone module of the convolutional neural network to extract multi-scale spectral spatial features;
[0037] S3. After normalizing the latitude and longitude coordinates of each target sample point, input them into a multilayer perceptron and generate a coordinate embedding vector. Then, fuse the coordinate embedding vector with the multi-scale spectral spatial features to obtain the regional fusion features.
[0038] S4. Input the region fusion features into two heterogeneous multi-head attention modules to establish long-distance dependencies between different spatial locations and different feature dimensions, and introduce residual connections and normalization operations after each attention layer to output region-aware features; wherein, the first attention module is used to capture diverse region associations, and the second attention module is used to integrate and abstract the feature relationships output by the first layer.
[0039] S5. Several learnable deep prototypes are set through the deep perception clustering module. Modal weights are calculated based on the similarity between the region perception features of the sample and different deep prototypes. The deep prototypes are then weighted and aggregated according to the modal weights to obtain deep perception features. Each deep prototype represents a potential deep-related feature modality.
[0040] S6. Input the depth-sensing features into two parallel output heads; wherein, the two parallel output heads include: a water depth prediction output head and an uncertainty output head; the water depth prediction output head predicts the water depth, and the uncertainty output head outputs the prediction variance based on the heteroscedastic Gaussian assumption.
[0041] In one embodiment of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the water depth inversion method for turbid nearshore waters as described above.
[0042] Compared with existing technologies, the water depth inversion system and method for turbid nearshore waters according to the present invention can improve the overall water depth inversion accuracy and significantly alleviate the problem of spatial nonstationarity. Attached Figure Description
[0043] Figure 1 This is a structural diagram of a water depth inversion system for turbid nearshore waters according to an embodiment of the present invention;
[0044] Figure 2This is a flowchart of a water depth inversion method for turbid nearshore waters according to an embodiment of the present invention. Detailed Implementation
[0045] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings, but it should be understood that the scope of protection of the present invention is not limited to the specific embodiments.
[0046] Unless otherwise expressly stated, throughout the specification and claims, the term "comprising" or its variations such as "including" or "comprises" shall be understood to include the stated elements or components without excluding other elements or other components.
[0047] like Figure 1 As shown, a water depth inversion system for turbid nearshore waters according to a preferred embodiment of the present invention includes the following units:
[0048] The data acquisition and preprocessing unit 100 is used to acquire and process multispectral remote sensing images covering the nearshore waters of the target and measured water depth data of the same area to form a sample set including multiple target sample points.
[0049] Specifically, visible light bands with strong water penetration and high spatial resolution from multispectral remote sensing images are prioritized. During preprocessing, the remote sensing images undergo quality screening and image compositing, and the measured water depth data are subjected to coordinate unification and quality control to form a sample set containing image features, latitude and longitude coordinates, and measured water depth labels.
[0050] After forming the sample set, a local image patch of fixed window size is cropped from the multispectral remote sensing image, centered on the pixel corresponding to each target sample point in the sample set, and used as the spectral-spatial input for subsequent feature extraction units. This local image patch contains both the band reflectance information of the target pixel itself and the water optical continuity and spatial texture information of neighboring pixels. Through this input method, the model no longer relies solely on the spectral values of a single pixel, but can combine the spatial context information of the surrounding neighborhood, thereby improving the stability of water depth inversion in complex nearshore waters.
[0051] The feature extraction unit 200 is used to input the local image patch into the backbone module of a convolutional neural network to extract multi-scale spectral spatial features. Specifically, the local image patch is input into the backbone module of the convolutional neural network, which consists of a two-dimensional convolutional layer, a normalization layer, and a nonlinear activation function, capable of progressively converting low-level band responses into high-level feature representations. This unit is mainly used to capture the local water body optical response within the target pixel and its neighborhood, providing basic features for subsequent region-aware modeling and depth-aware aggregation.
[0052] The coordinate region encoding unit 300 is used to normalize the latitude and longitude coordinates of each target sample point, input them into a multilayer perceptron (MLP), and generate a coordinate embedding vector. This coordinate embedding vector is then fused with the multi-scale spectral spatial features to obtain a region fusion feature. Specifically, to express the spatial differences between different geographic sub-regions, the latitude and longitude coordinates of each sample are normalized and input into the MLP to generate a coordinate embedding vector. This coordinate embedding vector represents the spatial prior information of the sample's region, such as the optical differences between different areas like harbor basins, channels, nearshore shoals, and open water.
[0053] Subsequently, the convolutional features are fused with the coordinate embedding vector to obtain the region fusion features. In this way, the model can utilize geographical location constraints on spatial non-stationary phenomena such as "same spectrum, different water depth" or "same water depth, different spectrum", thereby improving the model's adaptability in complex nearshore waters.
[0054] The region-aware attention modeling unit 400 is used to input the region fusion features into a two-layer heterogeneous multi-head attention module to establish long-distance dependencies between different spatial locations and different feature dimensions. Residual connections and normalization operations are introduced after each attention layer to output region-aware features. Specifically, the fused region fusion features are input into a two-layer heterogeneous multi-head attention module to establish long-distance dependencies between different spatial locations and different feature dimensions. The first-layer attention module captures diverse region relationships, and the second-layer attention module integrates and abstracts the feature relationships output by the first layer. After each attention layer, residual connections and normalization operations are introduced to preserve both the attention-enhanced features and the original fused features, avoiding information degradation caused by deep transformations. After this module, the model can obtain region-aware features to express cross-regional dependencies and spatially non-stationary optical structures.
[0055] A depth-aware clustering aggregation unit 500 is used to set several learnable depth prototypes. Modal weights are calculated based on the similarity between the region-aware features of a sample and different depth prototypes. The depth prototypes are then weighted and aggregated according to these modal weights to obtain depth-aware features. Each depth prototype represents a potential depth-related feature modality. Specifically, a depth-aware clustering aggregation unit is introduced based on the region-aware features. This unit sets several learnable depth prototypes, each representing a potential depth-related feature modality, such as very shallow water, shallow-to-medium water, medium-to-deep water, and deep water modalities. Based on the similarity between the region-aware features of a sample and different depth prototypes, the weights for belonging to different depth modalities are automatically calculated, and feature aggregation is performed based on these weights to obtain depth-aware features. Unlike the traditional two-stage method of "clustering first, then regression," in this invention, depth prototypes, modal weights, and subsequent water depth prediction errors all participate in end-to-end training, enabling the clustering results to directly serve the final water depth prediction target. Through this mechanism, the model can form differentiated feature expressions for different water depth ranges, reducing the prediction error in the transition zone between shallow, medium and deep water.
[0056] The water depth prediction and uncertainty estimation unit 600 is used to input the depth-sensing features into two parallel output heads. These two parallel output heads include a water depth prediction output head and an uncertainty output head. The water depth prediction output head is used to predict water depth, and the uncertainty output head outputs the prediction variance based on the heteroscedastic Gaussian assumption. Specifically, the depth-sensing features are input into two parallel output heads: one outputs the predicted water depth, and the other outputs the predicted uncertainty. The water depth prediction output head is used to generate continuous water depth values, and the uncertainty output head is used to estimate the prediction reliability of the model under the current sample conditions. This embodiment of the invention employs the heteroscedastic Gaussian assumption, enabling joint optimization of the predicted water depth error and the prediction uncertainty. When a certain area is significantly affected by factors such as suspended sediment, water turbidity, changes in bottom sediment, interference from port structures, or sparse samples, the model can provide a higher uncertainty indication, thereby assisting users in judging the reliability of the prediction results.
[0057] In a preferred embodiment of the present invention, the multispectral remote sensing image is a Sentinel-2 Level-2A surface reflectance product. Processing the multispectral remote sensing image covering the target nearshore waters includes: quality screening, image compositing, and quality control of measured water depth data. Specifically, the remote sensing image uses the blue B2, green B3, and red B4 bands, which have strong water penetration and high spatial resolution. Quality screening includes screening for clouds, cloud shadows, and anomalous pixels. Image compositing involves median compositing of multiple images within the same season or a window adjacent to the measured time to reduce the impact of instantaneous cloud shadows, waves, and short-term sea state changes. Quality control of the measured water depth data includes coordinate unification, outlier removal, duplicate point merging, and effective pixel matching. Through the above processing, a high-quality sample set containing image features, latitude and longitude coordinates, and measured water depth labels is formed.
[0058] Preferably, the coordinate region encoding unit fuses the coordinate embedding vector with the multi-scale spectral spatial features by: concatenating the coordinate embedding vector and the multi-scale spectral spatial features along the feature dimension; the coordinate embedding vector is used to represent the spatial prior information of the geographical sub-region where the sample is located. Specifically, the multilayer perceptron maps the normalized latitude and longitude coordinates into a coordinate embedding vector that matches the dimension of the multi-scale spectral spatial features, allowing the two to be directly concatenated along the feature dimension. The coordinate embedding vector is used to represent the spatial prior information of the geographical sub-region where the sample is located, enabling the model to utilize geographical location constraints on spatial non-stationary phenomena such as "same spectrum, different water depth" or "same water depth, different spectrum".
[0059] Preferably, the first-layer attention module uses a first number of attention heads for global dependency modeling, and the second-layer attention module uses a second number of attention heads for feature integration and abstraction; the first number of attention heads is greater than the second number of attention heads. Specifically, the two-layer heterogeneous multi-head attention modules adopt a hierarchical attention structure of "multi-head global dependency modeling - few-head feature integration and abstraction". The first-layer attention module uses a larger number of attention heads (e.g., 8 heads) to capture diverse regional relationships and establish global dependencies across regions and feature dimensions; the second-layer attention module uses a smaller number of attention heads (e.g., 4 heads) to integrate and abstract the feature relationships output by the first layer, focusing on key features while controlling computational complexity. The residual connection adds the output of each attention layer to the input of that layer and then performs layer normalization, so that the attention-enhanced features and the original region fusion features are preserved together, avoiding information degradation caused by deep transformations.
[0060] Preferably, the modality weights are obtained by calculating the similarity between region-aware features and each depth prototype, and then transforming it using a soft allocation function. Specifically, the depth-aware clustering aggregation unit sets K learnable depth prototypes, each representing a potential depth-related feature modality. The modality weights are obtained by calculating the similarity between the region-aware features of the sample and each depth prototype, and then transforming it using a soft allocation function (such as the softmax function). The weighted aggregation is a weighted sum of the modality weights over each depth prototype to obtain the depth-aware features. The soft allocation function ensures that each modality weight is non-negative and that the sum is 1, making the feature aggregation process interpretable.
[0061] Preferably, the coordinate region encoding unit, region-aware attention modeling unit, depth-aware clustering aggregation unit, and depth prediction and uncertainty estimation unit jointly participate in end-to-end backpropagation training. Specifically, during model training, units such as local feature extraction, coordinate region encoding, region-aware attention, depth-aware clustering, depth prediction, and uncertainty estimation jointly participate in end-to-end optimization. Coordinate region encoding, attention modeling, depth prototype aggregation, and depth regression jointly participate in backpropagation, enabling the model to automatically adjust region features and depth modalities according to the final depth prediction target. Through this mechanism, unlike the traditional two-stage method of "partitioning first, then regression," the region feature representation and depth modality allocation in this invention are not performed independently, but form a closed-loop optimization with the depth prediction error, avoiding the cumulative propagation of partitioning errors.
[0062] Preferably, the training loss function under the heteroscedastic Gaussian assumption is:
[0063]
[0064] in, To measure the actual water depth, To predict water depth, To predict variance, this invention employs a heteroscedastic Gaussian negative log-likelihood loss function, enabling joint optimization of depth prediction error and prediction uncertainty. The first term of this loss function... This is a regularization term to prevent the model from infinitely amplifying the prediction variance, thus reducing the second term; the second term As a weighted regression term, it automatically reduces the weight of the regression loss on samples with large prediction variance, thereby improving the model's robustness to samples with high uncertainty.
[0065] Preferably, the water depth prediction and uncertainty estimation unit is further configured to, after model training is complete, input effective water body pixels pixel by pixel into the system and output a spatially continuous water depth prediction map and a corresponding uncertainty map, wherein the uncertainty map represents the water depth prediction confidence at each pixel location. Specifically, after training is complete, effective water body pixels in the target image are input pixel by pixel into the model, and a spatially continuous water depth prediction map and a corresponding uncertainty map are output. The water depth prediction map is used to display the nearshore underwater topographic pattern, and the uncertainty map is used to highlight areas requiring special interpretation or caution, such as harbor basins, channels, and sparse boundary sample areas.
[0066] This embodiment uses the turbid waters near the coast of Qinhuangdao as a specific research area to provide specific parameters and experimental results of the present invention in practical application.
[0067] 1. Study Area and Data
[0068] The study area is the nearshore turbid waters of Qinhuangdao, which includes various typical nearshore scenarios such as harbor basins, channels, nearshore shorelines and open waters. The water turbidity, sediment conditions, human engineering activities and hydrodynamic environment vary significantly, making it suitable for verifying the applicability of this invention in water depth inversion in complex nearshore waters.
[0069] The multispectral remote sensing imagery used the Sentinel-2 Level-2A surface reflectance product, selecting the blue B2, green B3, and red B4 bands. Cloud, cloud shadows, and anomalous pixels were screened from the remote sensing images. Median composites were then performed on multiple images within the same season or adjacent time windows to mitigate the impact of transient cloud shadows, waves, and short-term sea state changes. A total of 2401 measured water depth samples were collected, ranging from approximately 1.02 m to 20.24 m, covering harbor basins, channels, and nearshore shorelines. Quality control measures, including coordinate unification, outlier removal, duplicate point merging, and effective pixel matching, were implemented for the measured water depth data, resulting in a sample set containing image features, latitude and longitude coordinates, and measured water depth labels.
[0070] 2. Construction of local image patches
[0071] For each target water depth sample point, a local image patch of a fixed window size is cropped from the Sentinel-2 multispectral image, centered on that pixel, and used as the spectral-spatial input of the model. This local image patch contains both the band reflectance information of the target pixel itself and the water optical continuity and spatial texture information of neighboring pixels.
[0072] 3. Network Structure and Training
[0073] The region-aware deep clustering convolutional neural network used in this embodiment includes the following modules:
[0074] (1) SDBCNN spectral-spatial feature extraction module: It consists of two-dimensional convolutional layers, batch normalization layers and ReLU activation function, which gradually convert the multispectral input of local image blocks into high-level feature representation.
[0075] (2) Coordinate region encoding module: The normalized latitude and longitude coordinates are mapped into coordinate embedding vectors that match the convolution feature dimensions through a multilayer perceptron (MLP), and the coordinate embedding vectors are concatenated with the convolution features in the feature dimension to obtain the region fusion feature.
[0076] (3) Two-layer heterogeneous multi-head attention module: The first attention module is used for global dependency modeling to capture diverse regional relationships; the second attention module is used for feature integration and abstraction. Residual connections and layer normalization operations are introduced after each attention layer to output region-aware features.
[0077] (4) Depth-aware clustering aggregation module: K learnable depth prototypes are set, each prototype representing a potential depth-related feature modality. Modality weights are obtained by calculating the similarity between the region-aware features of the sample and each depth prototype, and then transforming it using the softmax soft allocation function. The depth-aware features are obtained by weighted summation of each depth prototype using the modality weights.
[0078] (5) Depth Prediction and Uncertainty Estimation Module: Depth-sensing features are input into two parallel output heads. The depth prediction head outputs continuous depth values, while the uncertainty output head outputs the prediction variance based on the heteroscedastic Gaussian assumption. The training loss function is:
[0079]
[0080] in, To measure the actual water depth, To predict water depth for the model, The variance predicted by the model.
[0081] During model training, the sample set is divided into training, validation, and test sets. Strategies such as optimizers, learning rate scheduling, and gradient pruning are employed to improve training stability. Units including local feature extraction, coordinate region encoding, region-aware attention, depth-aware clustering, depth prediction, and uncertainty estimation work together in end-to-end optimization.
[0082] 4. Experimental Results
[0083] (1) Overall performance comparison
[0084] Table 1 shows the overall performance comparison between the present invention and the main control model.
[0085]
[0086] As shown in Table 1, the optimized RA-SDBCNN of this invention achieves R²=0.845, RMSE=1.686 m, and MAE=0.951 m on the test set, which is significantly better than the Base SDBCNN's R²=0.714, RMSE=2.287 m, and MAE=1.649 m, demonstrating a substantial improvement in overall fitting ability and error control. With the addition of coordinate region encoding alone, the model's R² increased from 0.714 to 0.822, and the RMSE decreased from 2.287 m to 1.803 m, indicating that explicit spatial priors can effectively help the model distinguish the differences in optical response among different functional sub-regions. Further introducing region-aware attention and depth-aware clustering on top of coordinate region encoding further improves the complete model's R² to 0.845 and reduces the RMSE to 1.686 m, demonstrating that cross-region dependency modeling and deep modality representation can jointly enhance model performance.
[0087] (2) Error performance in different depth ranges
[0088] Table 2 shows the error performance of the present invention in different water depth ranges.
[0089]
[0090] As shown in Table 2, compared with Base SDBCNN, the present invention reduces RMSE and MAE in different depth ranges such as 0-5 m, 5-10 m, 10-15 m and greater than 15 m, and shows more stable error control capability, especially in shallow to medium water and deep water areas.
[0091] (3) Confidence assessment of prediction
[0092] The uncertainty of the model output is strongly correlated with the absolute prediction error. In the experiment, the Pearson correlation coefficient was 0.812 and the Spearman correlation coefficient was 0.785, which can indicate potential low-confidence locations in high-risk areas such as port basins and waterways.
[0093] (4) Spatial generalization ability
[0094] Under the spatial isolation evaluation protocol, the present invention still maintains an RMSE of approximately 1.77 m, indicating that the framework is more adaptable to spatially separated heteroproton regions than ordinary random partitioning tests.
[0095] 5. Full-scene reasoning output
[0096] After training, the effective water pixels from the target Sentinel-2 image are input pixel by pixel into the trained model, which outputs a spatially continuous water depth prediction map and a corresponding uncertainty map. The water depth prediction map is used to display the underwater topography of the nearshore area of Qinhuangdao, while the uncertainty map is used to highlight areas that require special interpretation or should be used with caution, such as harbor basins, channels, and sparse boundary samples.
[0097] Correspondingly, this embodiment also provides a method for water depth inversion in turbid nearshore waters, such as... Figure 2 As shown, it includes the following steps:
[0098] S1. Acquire and process multispectral remote sensing images covering the nearshore waters of the target and measured water depth data of the same area to form a sample set including multiple target sample points; take the pixel corresponding to each target sample point in the sample set as the center and crop a local image block of fixed window size in the multispectral remote sensing image.
[0099] S2. Input the local image block into the backbone module of the convolutional neural network to extract multi-scale spectral spatial features;
[0100] S3. After normalizing the latitude and longitude coordinates of each target sample point, input them into a multilayer perceptron and generate a coordinate embedding vector. Then, fuse the coordinate embedding vector with the multi-scale spectral spatial features to obtain the regional fusion features.
[0101] S4. Input the region fusion features into two heterogeneous multi-head attention modules to establish long-distance dependencies between different spatial locations and different feature dimensions, and introduce residual connections and normalization operations after each attention layer to output region-aware features; wherein, the first attention module is used to capture diverse region associations, and the second attention module is used to integrate and abstract the feature relationships output by the first layer.
[0102] S5. Several learnable deep prototypes are set through the deep perception clustering module. Modal weights are calculated based on the similarity between the region perception features of the sample and different deep prototypes. The deep prototypes are then weighted and aggregated according to the modal weights to obtain deep perception features. Each deep prototype represents a potential deep-related feature modality.
[0103] S6. Input the depth-sensing features into two parallel output heads; wherein, the two parallel output heads include: a water depth prediction output head and an uncertainty output head; the water depth prediction output head predicts the water depth, and the uncertainty output head outputs the prediction variance based on the heteroscedastic Gaussian assumption.
[0104] Preferably, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the water depth inversion method for turbid nearshore waters as described above.
[0105] The foregoing description of specific exemplary embodiments of the invention is for illustrative and explanatory purposes. These descriptions are not intended to limit the invention to the precise forms disclosed, and it will be apparent that many changes and variations can be made in accordance with the foregoing teachings. The exemplary embodiments were chosen and described in order to explain the specific principles of the invention and its practical application, thereby enabling those skilled in the art to implement and utilize various different exemplary embodiments of the invention, as well as various different choices and variations. The scope of the invention is intended to be defined by the claims and their equivalents.
Claims
1. A water depth inversion system for turbid nearshore waters, characterized in that, include: The data acquisition and preprocessing unit is used to acquire and process multispectral remote sensing images covering the nearshore waters of the target and measured water depth data of the same area to form a sample set including multiple target sample points; and to crop a local image block of a fixed window size from the multispectral remote sensing image with the pixel corresponding to each target sample point in the sample set as the center. The feature extraction unit is used to input the local image block into the backbone module of the convolutional neural network to extract multi-scale spectral spatial features; The coordinate region encoding unit is used to normalize the latitude and longitude coordinates of each target sample point, input them into the multilayer perceptron and generate a coordinate embedding vector, and fuse the coordinate embedding vector with the multi-scale spectral spatial features to obtain the region fusion feature. The region-aware attention modeling unit is used to input the region fusion features into a two-layer heterogeneous multi-head attention module to establish long-distance dependencies between different spatial locations and different feature dimensions, and introduce residual connections and normalization operations after each attention layer to output region-aware features; wherein, the first-layer attention module is used to capture diverse region associations, and the second-layer attention module is used to integrate and abstract the feature relationships output by the first layer. A deep-aware clustering aggregation unit is used to set up several learnable deep prototypes, calculate modality weights based on the similarity between the region-aware features of the sample and different deep prototypes, and perform weighted aggregation of the deep prototypes based on the modality weights to obtain deep-aware features; wherein, each deep prototype represents a potential deep-related feature modality. The water depth prediction and uncertainty estimation unit is used to input the depth sensing features into two parallel output heads; wherein, the two parallel output heads include a water depth prediction output head and an uncertainty output head; the water depth prediction output head is used to predict the water depth, and the uncertainty output head outputs the prediction variance based on the heteroscedastic Gaussian assumption.
2. The system according to claim 1, characterized in that, The multispectral remote sensing image is a Sentinel-2 Level-2A surface reflectance product; the processing of multispectral remote sensing images covering the nearshore waters of the target includes: quality screening, image synthesis, and quality control of measured water depth data.
3. The system according to claim 1, characterized in that, The coordinate region encoding unit fuses the coordinate embedding vector with the multi-scale spectral spatial features by: concatenating the coordinate embedding vector with the multi-scale spectral spatial features along the feature dimension; the coordinate embedding vector is used to characterize the spatial prior information of the geographical sub-region where the sample is located.
4. The system according to claim 1, characterized in that, The first layer attention module uses a first number of attention heads for global dependency modeling, and the second layer attention module uses a second number of attention heads for feature integration and abstraction; the first number of attention heads is greater than the second number of attention heads.
5. The system according to claim 1, characterized in that, The modal weights are obtained by using the similarity between region-aware features and each depth prototype, and then transforming them using a soft assignment function.
6. The system according to claim 1, characterized in that, The coordinate region encoding unit, region-aware attention modeling unit, depth-aware clustering aggregation unit, and water depth prediction and uncertainty estimation unit jointly participate in end-to-end backpropagation training.
7. The system according to claim 1, characterized in that, The training loss function under the heteroscedastic Gaussian assumption is: in, To measure the actual water depth, To predict water depth, To predict variance.
8. The system according to claim 1, characterized in that, The water depth prediction and uncertainty estimation unit is also used to: after the model training is completed, input the effective water body pixels into the system pixel by pixel, and output a spatially continuous water depth prediction map and a corresponding uncertainty map, wherein the uncertainty map represents the water depth prediction confidence at each pixel location.
9. A method for water depth inversion in turbid nearshore waters, characterized in that, include: S1. Acquire and process multispectral remote sensing images covering the nearshore waters of the target and measured water depth data of the same area to form a sample set including multiple target sample points; take the pixel corresponding to each target sample point in the sample set as the center and crop a local image block of fixed window size in the multispectral remote sensing image. S2. Input the local image block into the backbone module of the convolutional neural network to extract multi-scale spectral spatial features; S3. After normalizing the latitude and longitude coordinates of each target sample point, input them into a multilayer perceptron and generate a coordinate embedding vector. Then, fuse the coordinate embedding vector with the multi-scale spectral spatial features to obtain the regional fusion features. S4. Input the region fusion features into two heterogeneous multi-head attention modules to establish long-distance dependencies between different spatial locations and different feature dimensions, and introduce residual connections and normalization operations after each attention layer to output region-aware features; wherein, the first attention module is used to capture diverse region associations, and the second attention module is used to integrate and abstract the feature relationships output by the first layer. S5. Several learnable deep prototypes are set through the deep perception clustering module. Modal weights are calculated based on the similarity between the region perception features of the sample and different deep prototypes. The deep prototypes are then weighted and aggregated according to the modal weights to obtain deep perception features. Each deep prototype represents a potential deep-related feature modality. S6. Input the depth-sensing features into two parallel output heads; wherein, the two parallel output heads include: a water depth prediction output head and an uncertainty output head; the water depth prediction output head predicts the water depth, and the uncertainty output head outputs the prediction variance based on the heteroscedastic Gaussian assumption.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the water depth inversion method for turbid nearshore waters as described in claim 9.