Synthesis method and device of sonar image, electronic equipment and storage medium

By decomposing the cross-modal image conversion task into a multi-stage processing, and combining multi-scale feature extraction of optical images with acoustic property fusion, high-quality sonar images are generated. This solves the problems of blurred edges and severe artifacts in existing sonar images, and improves the realism and usability of sonar images.

CN121505087BActive Publication Date: 2026-04-17启元实验室
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
启元实验室
Filing Date
2026-01-07
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing cross-modal image conversion technologies suffer from problems such as blurred edges, severe artifacts, and poor image quality when generating sonar images. Furthermore, sonar data acquisition is costly and the amount of data is limited.

Method used

The cross-modal image conversion task is decomposed into three stages: viewpoint conversion, shadow generation, and style conversion. Through a multi-stage end-to-end training strategy, multi-scale feature extraction and acoustic property fusion are performed using optical images to generate sonar shadow feature maps, which are then converted into realistic sonar style images through progressive denoising.

Benefits of technology

The generated sonar images are highly consistent with real sonar data in terms of texture, granularity, and contrast, which improves the realism and usability of the images, overcomes the learning difficulty and uncertainty of direct end-to-end conversion, and ensures semantic consistency and structural integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505087B_ABST
    Figure CN121505087B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, electronic device, and storage medium for synthesizing sonar images, relating to the field of computer technology. The sonar image synthesis method includes: converting a basic optical image into an overhead view image and a depth map according to a preset viewpoint conversion model; determining the weights of different features in the overhead view image and depth map based on their contribution to sonar imaging, and performing multi-scale feature extraction on the overhead view image and depth map based on these weights to determine a corresponding hierarchical fusion map; determining the acoustic wave reflection intensity map and sonar shadow region of the underwater object surface based on the overhead view image, depth map, predefined acoustic material library, and hierarchical fusion map, and fusing the acoustic wave reflection intensity map and sonar shadow region to determine a sonar shadow feature map; and converting the sonar shadow feature map into a realistic sonar-style image based on a preset progressive denoising method. This application improves the conversion effect from optical images to sonar images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and for example to a method, apparatus, electronic device, and storage medium for synthesizing sonar images. Background Technology

[0002] Underwater environmental perception is a core technology in fields such as marine engineering, underwater robot navigation, and marine resource exploration. Optical imaging and sonar imaging are two mainstream sensing methods for underwater environmental perception, using light waves and sound waves as carriers, respectively, to achieve environmental detection. Sonar systems utilize the long-distance propagation characteristics of sound waves in water, enabling detection at the hundred-meter level in low-visibility environments, but they suffer from low resolution, significant noise interference, and a lack of color information. Furthermore, underwater sonar data acquisition requires specialized equipment, which is expensive and yields limited data. Optical images, on the other hand, offer advantages such as high resolution, rich texture and color information, and underwater optical image acquisition is less costly compared to sonar images.

[0003] Currently, cross-modal image conversion technology is developing rapidly, aiming to generate sonar images using prior knowledge of optical images to address the problems of scarce sonar data and insufficient realism in simulations. Among related technologies, cross-modal image conversion methods are mainly divided into physical model-based simulation methods and deep learning-based style transfer methods. The former is limited by complex marine environmental parameters and computational costs, making real-time simulation difficult. While the latter uses adversarial generative networks to learn mapping relationships, the fundamental differences between the two modalities in imaging principles generally result in problems such as blurred edges and severe artifacts in the conversion results. Therefore, the current methods for cross-modal image conversion and subsequent synthesis of sonar images are relatively difficult, leading to poor image quality in the final sonar images. Summary of the Invention

[0004] This application aims to provide a method, apparatus, electronic device, and storage medium for synthesizing sonar images.

[0005] According to one aspect of this application, a method for synthesizing sonar images is proposed, comprising: acquiring a basic optical image of an underwater object captured by an optical camera, and converting the basic optical image into an overhead view image and a depth map according to a preset viewpoint conversion model; determining the weights of different features based on their contribution to sonar imaging in the overhead view image and depth map, and performing multi-scale feature extraction on the overhead view image and depth map based on the weights to determine the corresponding hierarchical fusion map; determining the acoustic wave reflection intensity map and sonar shadow region on the surface of the underwater object based on the overhead view image, depth map, predefined acoustic material library, and hierarchical fusion map, and fusing and determining a sonar shadow feature map based on the acoustic wave reflection intensity map and sonar shadow region; and converting the sonar shadow feature map into a realistic sonar style image based on a preset progressive denoising method.

[0006] According to one aspect of this application, a sonar image synthesis apparatus is provided, comprising:

[0007] The image conversion model is used to acquire basic optical images of underwater objects captured by an optical camera, and convert the basic optical images into overhead view images and depth maps according to a preset view conversion model.

[0008] The hierarchical image determination module is used to determine the weights of different features based on their contribution to sonar imaging in the overhead view image and depth map, and to perform multi-scale feature extraction on the overhead view image and depth map based on the weights in order to determine the corresponding hierarchical fusion map.

[0009] The image fusion module is used to determine the acoustic wave reflection intensity map and sonar shadow area of ​​the underwater object surface based on the overhead view image, depth map, predefined acoustic material library and hierarchical fusion map, and to fuse and determine the sonar shadow feature map based on the acoustic wave reflection intensity map and sonar shadow area.

[0010] The denoising module is used to convert sonar shadow feature maps into realistic sonar-style images based on a preset stepwise denoising method.

[0011] According to one aspect of this application, an electronic device is provided, comprising: a processor; and a memory storing a computer program that, when executed by the processor, causes the processor to perform the method described above.

[0012] According to one aspect of this application, a non-transitory computer-readable medium is proposed, on which readable instructions are stored, which, when executed by a processor, cause the processor to perform the method described above.

[0013] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this application.

[0014] Beneficial effects:

[0015] The embodiments provided in this application decompose the complex cross-modal conversion task into three targeted stages: viewpoint conversion, shadow generation, and style conversion, effectively reducing the learning difficulty and uncertainty of direct end-to-end conversion. This step-by-step approximation strategy can inject and retain key features at each level, thereby fundamentally overcoming the problems of severe loss of image details and texture distortion caused by the overly general conversion process, ensuring the semantic consistency and structural integrity of the generated image. This application also deeply integrates acoustic properties (such as reflection intensity) with physical material parameters, and calculates and generates sound wave reflection intensity maps and sonar shadow regions, so that the synthesized sonar image strictly follows the basic physical laws of underwater acoustic imaging, avoiding the problem of results not matching actual sonar detection data due to ignoring physical mechanisms. By extracting multi-scale features from the overhead viewpoint and depth information and adaptively weighting them according to their contribution to the final sonar imaging, priority can be given to and the feature levels with the greatest impact on the imaging results can be strengthened. Combined with shadow generation based on a physical model, the finally fused sonar shadow feature map provides an intermediate representation with rich information and clear physical meaning for subsequent style conversion. In the final style transfer stage, based on the preset stepwise denoising method, the noise and artifacts introduced during the generation process can be effectively suppressed, and the initial sonar feature map can be smoothly converted into a final image that is highly consistent with the real sonar data in terms of texture, graininess and contrast, which further improves the realism and usability of the final sonar image. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings, without exceeding the scope of protection claimed by this application.

[0017] Figure 1 This is a schematic diagram of the overall process of sonar image synthesis provided in the embodiments of this application;

[0018] Figure 2 A flowchart illustrating the sonar image synthesis method provided in this application embodiment;

[0019] Figure 3 A block diagram of a sonar image synthesis apparatus provided in the embodiments of this application;

[0020] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0021] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that this application will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.

[0022] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0023] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0024] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0025] It should be understood that although the terms first, second, third, etc., may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Therefore, the first component discussed below may be referred to as the second component without departing from the teachings of this application. As used herein, the term "and / or" includes all combinations of any one and more of the associated listed items.

[0026] Figure 1 This is a schematic diagram illustrating the overall process of sonar image synthesis provided in an embodiment of this application. Figure 1As shown, this application can pre-set a multi-stage end-to-end training strategy for training the viewpoint transformation model, sonar shadow generation model, and sonar image stylization model. The multi-stage end-to-end training strategy constructs a unified training framework, chaining the aforementioned models to achieve a complete multi-stage optical image to sonar image conversion process. The multi-stage forward propagation mechanism allows data to flow sequentially between modules, with the output of each stage serving as both input to the next-level module and as an independent supervised object, achieving dual constraints of staged and overall goals. The specific process is as follows: In the first stage, the normal viewpoint optical image is input into the viewpoint transformation model. After feature extraction, depth estimation, and viewpoint transformation, the output is a top-down viewpoint optical image, which serves as input to the sonar shadow generation model. In the second stage, the top-down viewpoint optical image is input into the sonar shadow generation model, combining with sonar physical imaging to generate sonar shadow features. In the third stage, the sonar shadow features are input into the sonar image stylization model, and after passing through a diffusion model framework, a sonar stylization image is output. Through this multi-stage flow mechanism, the data gradually approaches the target output during the transmission process, while each intermediate result is supervised, preventing error accumulation between modules.

[0027] The multi-stage supervised loss function sets a specific loss function for each module, and the total loss function is formed by weighted summation, which updates the parameters of the entire network through backpropagation. The specific losses of the viewpoint transition model include depth estimation loss and viewpoint alignment loss. The depth estimation loss is calculated using mean squared error, and the formula is as follows:

[0028]

[0029] in, Used to characterize the depth estimation loss Used to characterize the predicted depth value Used to represent the true depth value For the number of pixels, This is the current pixel.

[0030] The view alignment loss uses cross-entropy loss, and the corresponding formula is:

[0031]

[0032] in, pixel coordinates ( Mapped to ( The probability of ), therefore the total loss of the viewpoint switching model. for:

[0033]

[0034] in, , Let be the weight coefficient, and satisfy... + =1.

[0035] The training of the sonar shadow generation model focuses on the accurate modeling of physical properties, including the prediction of sound wave reflection intensity and the calculation of shadow regions. During training, physical constraints based on acoustic principles are introduced to ensure that the predicted reflection intensity distribution conforms to the acoustic properties of different materials, and that the generated shadow regions accurately reflect geometric occlusion relationships. The model's loss function integrates physical modeling and feature learning, as shown in the formula:

[0036]

[0037] in, Used to characterize the sonar shadow region. and These represent the predicted reflection intensity and shadow distribution, respectively. and These are the true values ​​corresponding to reflection intensity and shadow distribution, respectively; Losses due to material classification; , , and These are the weighting parameters used to balance the various loss terms; The matching loss used to characterize the predicted reflection intensity and the corresponding material's true acoustic properties is to constrain the generated reflection intensity to conform to the inherent acoustic reflection law of the material, ensuring that the sonar reflection characteristics of different materials are true and accurate.

[0038] The sonar image stylization model employs a diffusion model training process. By learning the data distribution of real sonar images, the model gains the ability to generate high-quality sonar images from noise and conditional features. The loss function combines the denoising objective with style consistency, as shown in the formula:

[0039]

[0040] and These are the predicted noise and the actual noise, respectively. To combat the losses, This results in a loss of consistency in equipment characteristics. , and These are the weight parameters.

[0041] The final total system loss function for:

[0042]

[0043] in, , and The stage loss weights and satisfy During training, the Adam (Adaptive Moment Estimation) optimizer is used to optimize the total loss function. The gradient is then passed from the sonar image style model to the viewpoint transformation model layer by layer through the backpropagation algorithm to achieve parameter updates at each stage.

[0044] Optical images from a normal viewing angle can be passed layer by layer to the viewpoint transformation model, the sonar shadow generation model, and the sonar image style model, outputting the final sonar image, i.e., a real sonar style image.

[0045] The models described above can all contain multiple sub-modules to perform different operations. The viewpoint transformation model includes sub-modules for multi-scale feature extraction, depth estimation, and viewpoint transformation, and the viewpoint transformation stage can be evaluated using overhead view quality and depth accuracy. The sonar shadow generation model can include sub-modules for adaptive multi-scale feature extraction (combined with depth maps), reflectance intensity prediction, shadow region localization, and feature fusion (fusion of reflectance intensity maps and shadow maps), and the shadow generation stage can be evaluated using reflectance intensity accuracy and shadow region accuracy. The sonar image style model can include sub-modules for diffusion models (including forward noise addition, reverse noise reduction, and cross-scale fusion) and device characteristic simulation (implemented using parametric noise modulation), and the style transfer stage can be evaluated using image quality, texture consistency, and physical realism.

[0046] Feedback and optimization can also be carried out based on the above-mentioned phase evaluation results.

[0047] For specific implementation details, please refer to the following examples.

[0048] Figure 2 A flowchart illustrating a sonar image synthesis method provided in an embodiment of this application. Figure 2 As shown, the method includes steps S20, S21, S22 and S23.

[0049] In step S20, a basic optical image of an underwater object is acquired by an optical camera, and the basic optical image is converted into an overhead view image and a depth map according to a preset view conversion model.

[0050] In this application, an optical camera is used to photograph target objects in an underwater environment, i.e., underwater objects. Images directly captured by the optical camera can serve as the base optical image, which is an optical image from a predefined normal viewing angle. The normal viewing angle can be the side / eye-level viewpoint of the optical camera. A viewpoint transformation model can be pre-set; the base optical image is input into this model, and the top-down view image and its corresponding depth map are directly output.

[0051] In step S21, the weights of different features are determined based on their contribution to sonar imaging in the overhead view image and depth map, and multi-scale feature extraction is performed on the overhead view image and depth map based on the weights to determine the corresponding hierarchical fusion map.

[0052] In this application, the contribution of different hierarchical features to sonar imaging and the correspondence between contribution and weight allocation can be preset in advance, so as to set the weight of different features accordingly.

[0053] A sonar shadow generation model can be pre-set to convert the overhead view image into an intermediate representation characterizing key features of sonar imaging, namely a sonar shadow feature map. The sonar shadow generation model can consist of two parts. The first part can extract hierarchical features and their corresponding hierarchical fusion maps from the input overhead view image and depth map.

[0054] In step S22, the acoustic wave reflection intensity map and sonar shadow area of ​​the underwater object surface are determined based on the overhead view image, depth map, predefined acoustic material library and hierarchical fusion map, and the sonar shadow feature map is determined by fusing the acoustic wave reflection intensity map and sonar shadow area.

[0055] In this application, the process of "determining hierarchical feature representation and weight allocation" in step S21 provides high-quality features after screening for the "prediction process of sonar reflection intensity distribution and shadow area". The low-scale and high-scale feature information extracted from the hierarchical fusion map are used as hierarchical features. The weight allocation further strengthens the features useful for prediction and weakens irrelevant information. The final output "high-weight key features" are used as input for the prediction of reflection intensity distribution and shadow area.

[0056] The second part of the sonar shadow generation model can output a sound wave reflection intensity map and a sonar shadow region by utilizing information from the overhead view image, the corresponding information from the depth map, the corresponding information from the predefined acoustic material library, and the hierarchical features in the hierarchical fusion map. The sound wave reflection intensity map is the information on the intensity of sound wave reflection generated by the surface of an underwater object, estimated by the sonar shadow generation model. The sonar shadow region is the shadow area formed by the obstruction of sound wave propagation.

[0057] The sonar shadow model obtains a sonar shadow feature map through adaptive fusion of the acoustic reflection intensity map and the sonar shadow region.

[0058] In step S23, the sonar shadow feature map is converted into a real sonar style image based on a preset stepwise denoising method.

[0059] In this application, a sonar image style model can be pre-constructed. This model is based on a diffusion model architecture and transforms the sonar shadow features in the sonar shadow feature map into a realistic sonar style image through a stepwise denoising process. The sonar shadow features can include visual features such as object outlines, strong reflective areas, shadow distribution, and basic geometric structures, all of which can be extracted from the sonar shadow feature map.

[0060] In some implementations, the object outline is derived from the texture edge features extracted from the overhead view RGB image (i.e., the overhead view image); the strong reflection area is directly derived from the predicted sound wave reflection intensity map. The area with a pixel value higher than a set threshold in the sound wave reflection intensity map is the area with strong sound wave reflection energy, which is the strong reflection area; the shadow distribution is directly derived from the predicted sonar shadow area. The area marked as "shadow area" (e.g., pixel value of 0) in the sonar shadow area is the shadow range formed by the obstruction of sound wave propagation, which constitutes the shadow distribution feature; the basic geometric structure is derived from the depth map and the scene structure features under the overhead view, and is obtained by the model through the analysis of depth information and features.

[0061] This application decomposes the complex cross-modal conversion task into three targeted stages: viewpoint conversion, shadow generation, and style conversion, effectively reducing the learning difficulty and uncertainty of direct end-to-end conversion. This step-by-step approximation strategy can inject and retain key features at each level, fundamentally overcoming the problems of severe loss of image details and texture distortion caused by the overly general conversion process, thus ensuring the semantic consistency and structural integrity of the generated image. This application also deeply integrates acoustic properties (such as reflection intensity) with physical material parameters, generating sound wave reflection intensity maps and sonar shadow regions by calculation. This ensures that the synthesized sonar image strictly follows the basic physical laws of underwater acoustic imaging, avoiding the problem of results not matching actual sonar detection data due to ignoring physical mechanisms. By extracting multi-scale features from the overhead viewpoint and depth information and adaptively weighting them according to their contribution to the final sonar imaging, it can prioritize and strengthen the feature levels that have the greatest impact on the imaging results. Combined with shadow generation based on a physical model, the final fused sonar shadow feature map provides an information-rich and physically meaningful intermediate representation for subsequent style conversion. In the final style transfer stage, based on the preset stepwise denoising method, the noise and artifacts introduced during the generation process can be effectively suppressed, and the initial sonar feature map can be smoothly converted into a final image that is highly consistent with the real sonar data in terms of texture, graininess and contrast, which further improves the realism and usability of the final sonar image.

[0062] According to some embodiments, the viewpoint transformation model includes a hierarchical convolutional multi-scale feature extraction network, a depth estimation module, and a viewpoint transformation module. Specifically, the process of converting a basic optical image into an overhead view image and a depth map involves: acquiring a basic optical image of an underwater object captured by an optical camera; extracting features from the basic optical image using the multi-scale feature extraction network to determine multi-scale feature information from low to high order; optimizing the multi-scale feature information and fusing the optimized multi-scale feature information according to a preset ratio to generate fused feature information; obtaining a depth map corresponding to the fused feature information through the depth estimation module; and constructing a coordinate transformation matrix based on the viewpoint transformation module, the depth map, and the camera intrinsic parameters of the optical camera, and mapping pixels in the basic optical image to corresponding positions in the overhead view based on an end-to-end differentiable transformation method and the depth map to generate the overhead view image.

[0063] In this application, a basic optical image captured by an optical camera is first obtained. Viewpoint transformation can be achieved using three modules: a multi-scale feature extraction network, depth estimation, and viewpoint transformation. The multi-scale feature extraction network employs a hierarchical convolutional structure, using convolutional kernels of different sizes to extract features from the input normal-view optical image. Small-sized convolutional kernels are used to extract detailed features, while large-sized kernels capture global features. Simultaneously, a ReLU activation function is used for non-linear transformation to enhance the network's feature representation capability. This yields multi-scale feature information from low to high order. Low-order features retain detailed information such as edges and textures, while high-order features capture semantic information such as object categories and scene layout. Finally, features are fused through feature concatenation and weighted summation.

[0064] The depth estimation module employs an encoder-decoder architecture, using fused multi-scale features as input. The encoder consists of multiple convolutional layers that progressively compress the feature space through convolution operations, reducing the size of the feature map while increasing the number of feature channels to extract richer feature information. The decoder, correspondingly, restores the feature dimensions through operations such as deconvolution. The number of deconvolutional layers is the same as the number of convolutional layers in the encoder, ultimately outputting a depth map corresponding to the input image size. During training, geometric prior knowledge is introduced to improve the accuracy of depth prediction, and geometric constraints are used to optimize the depth estimation loss function.

[0065] The view transformation module constructs a coordinate transformation matrix from the normal viewpoint to the overhead viewpoint based on the acquired depth map and camera intrinsic parameters. Let the camera intrinsic parameter matrix be K:

[0066]

[0067] in, and For camera focal length, and The master point coordinates. The depth value of a pixel in the depth map. The coordinates of this pixel in a normal viewing angle are Then its corresponding three-dimensional spatial coordinates can be expressed as:

[0068]

[0069] By setting the camera parameters for the overhead view, a coordinate transformation matrix is ​​constructed. , three-dimensional spatial coordinates Switch to coordinates from an overhead view ,Right now:

[0070]

[0071] Then, through the internal parameter matrix from an overhead perspective Projecting it onto the image plane yields the pixel coordinates from an overhead view. ,in, Automatic acquisition based on camera calibration and device specifications stems from hardware characteristics. Utilizing an end-to-end differentiable transformation method, the coordinate transformation is made differentiable throughout the training process, facilitating parameter optimization via backpropagation. Pixels from the normal viewpoint are mapped to their corresponding positions in the overhead viewpoint based on their depth information, achieving viewpoint transformation. For dynamic scenes, optical flow estimation is introduced to obtain object motion information, calculating the displacement vectors of objects in adjacent frames. The mapping position of the moving object is dynamically adjusted based on this displacement vector to ensure the stability of the transformation, that is:

[0072]

[0073] in, This is used to represent the mapped position information of moving objects in a dynamic scene from an overhead viewpoint. Based on the above process, an overhead view image is obtained.

[0074] This application utilizes a hierarchical convolutional structure, enabling the network to extract comprehensive feature information from basic optical images, ranging from low-level texture to high-level semantics. Subsequently, multi-scale features are optimized and fused according to a preset ratio. This process effectively enhances contextual information beneficial for underwater scene understanding and suppresses interference from the water medium in the optical image (such as scattering and noise), generating complete and robust fused features, providing a high-quality data foundation for subsequent depth estimation and viewpoint transformation. The depth estimation module uses the fused features obtained in the previous step as input and directly outputs a dense depth map. This design ensures pixel-level alignment of depth information with the original image, providing accurate and indispensable 3D spatial geometric constraints for subsequent viewpoint transformation and acoustic modeling, fundamentally guaranteeing the physical rationality and spatial accuracy of the synthesized sonar image. The viewpoint transformation module innovatively combines camera intrinsics, the depth map, and the coordinate transformation matrix, employing an end-to-end differentiable transformation method to accurately map pixels in the basic optical image to the overhead view. This method not only ensures the geometric correctness of the perspective transformation process and eliminates deformation caused by perspective differences, but its "differentiable" characteristic also enables the entire transformation model to be jointly optimized end-to-end through gradient backpropagation, which significantly improves the collaborative efficiency of each module and the overall accuracy of the final output.

[0075] According to some embodiments, the viewpoint transformation model includes a hybrid network structure combining sliding window feature aggregation and an attention mechanism, a depth estimation module, and a viewpoint transformation module. Under this viewpoint transformation model structure, the process of converting a base optical image into an overhead time-lapse image and a depth map specifically involves: acquiring a base optical image of an underwater object captured by an optical camera; extracting multi-scale features from the base optical image based on windowing of different receptive fields of the hybrid network structure, and fusing the multi-scale features according to dynamic weights determined by adaptive calculation of the attention mechanism to determine fused feature information; generating a depth map corresponding to the pixels of the base optical image through the depth estimation module, wherein the depth map is a dense depth map corresponding to the pixels of the base optical image; and converting the base optical image into an overhead view image through the viewpoint transformation module.

[0076] In this application, a basic optical image captured by an optical camera is first acquired and processed. Feature extraction employs a hybrid network structure combining sliding window feature aggregation and an attention mechanism to extract multi-scale features from the input normal-view optical image. For the input optical image, a sliding window feature aggregation operation is first performed. Based on the scale characteristics of different targets in the image, various sliding windows of different sizes are set, such as 3×3, 5×5, and 7×7. The window's sliding stride is set to 1 or 2 to ensure dense sliding across the image, thereby comprehensively capturing local detail features under different receptive fields. Smaller windows can focus on subtle textures and edge information in the image, while larger windows can acquire features from a wider range of local regions. Through this multi-window setup, multi-scale extraction of local details in the image is achieved.

[0077] After acquiring local features from different windows, a self-attention mechanism is introduced to model global semantic relationships. This mechanism mines long-distance dependencies between features by calculating the association weights between each feature point and all other feature points. Then, low-order detail features and high-order semantic features are fused using dynamic weights to form a feature map with spatial precision and semantic integrity. The dynamic weights are determined based on feature importance. A lightweight convolutional network learns the features, outputting a weight matrix of the same size as the feature map. Low-order and high-order features are multiplied by their corresponding weights and then summed to obtain the fused feature map, providing feature support for subsequent depth estimation. The fusion process can be represented as follows:

[0078]

[0079] in, The feature map used to represent the fusion of high-order detail features and high-order semantic features is called fused feature information. For low-level detail features, For higher-order semantic features, and These are the dynamic weights corresponding to low-order detail features and high-order semantic features, respectively, and satisfy the following conditions: .

[0080] Depth estimation, based on extracted visual features, generates a dense depth map corresponding to the pixels of the input image through implicit geometric inference. Its core lies in learning a mapping function. This method can associate two-dimensional image features F with three-dimensional spatial depth information D. The mapping function is learned through a deep neural network. The network takes the fused feature map as input, passes it through multiple convolutional and fully connected layers for nonlinear transformation, and outputs the corresponding depth value.

[0081] During inference, a global-local joint optimization strategy is employed. Globally, depth anchors are determined by understanding the overall scene, utilizing semantic information and global features from the entire image. For example, a global pooling layer is used to obtain overall scene features, which are then passed through a small network to predict multiple depth anchors. These anchor points reflect the depth range of different areas in the scene.

[0082] Local enhancement of texture and edges finely adjusts depth. For each pixel, the depth is locally corrected based on surrounding local features such as texture gradients and edge information. The initially predicted depth value is adjusted by calculating the feature differences between the pixel and its neighboring pixels.

[0083] Global and local regions interact through an attention mechanism. Global information provides contextual constraints for local predictions. Attention weights are calculated between each local region and the global depth anchor point. The predicted depth value of the local region is then constrained by the global anchor point based on these weights.

[0084]

[0085] in, The value used to represent the local depth prediction after global-local attention constraints is the actual constrained depth value after combining local predictions and global anchor point corrections. This represents the average depth value of the local area. The original depth prediction value for the local region is output by the depth prediction network based on local features. The attention weights for local regions and global depth anchors are calculated using an attention mechanism. This is the depth value of the global depth anchor point.

[0086] In addition, local details refine the global estimate by feeding back information from areas in the local features that deviate significantly from the global anchor point to the global optimization process, thereby adjusting the value of the global anchor point.

[0087] The final output depth map is obtained through adaptive weight fusion. For each pixel, the globally optimized depth value is fused. and the depth value after local optimization , fusion formula, where The output is derived from global features through a deeply optimized network. The above is obtained by deep optimization of local features.

[0088] Dense Depth Map .

[0089] in, The weights are adaptive and determined by the feature confidence of each pixel.

[0090] The conversion process for the perspective conversion module can be referenced from the above process.

[0091] This application utilizes sliding windows with different receptive fields to effectively capture full-scale visual features, from local texture to global structure. Subsequently, an attention mechanism is used to assign appropriate dynamic weights to features at different scales, prioritizing and enhancing the feature information most critical for subsequent sonar imaging. This design overcomes the limitations of traditional single-scale convolutional networks in feature extraction, providing a rich and robust feature foundation for generating accurate and clear top-down view and depth maps. The depth estimation module can directly generate a dense depth map corresponding one-to-one with the input image pixels based on the fused rich feature information. This step is crucial because it transforms two-dimensional visual information into three-dimensional spatial geometric information, providing indispensable depth data support for subsequent physically-based acoustic modeling (such as shadow calculation and reflection intensity simulation), fundamentally ensuring the geometric accuracy of the entire sonar synthesis process. The view transformation module uses the fused features and dense depth map obtained in the preceding steps to convert the side-view basic optical image into a standard sonar top-down view image through precise geometric projection and image transformation. This conversion corrects for geometric distortions and structural discrepancies caused by differences in perspective, ensuring that the generated intermediate image maintains the same spatial layout as the real sonar image, thus laying a correct perspective foundation for simulating acoustic phenomena in subsequent stages.

[0092] According to some embodiments, in the process of determining the hierarchical fusion map, the contribution of different features in the overhead view image and depth map to sonar imaging can be obtained; based on the contribution, a preset attention mechanism and a transformation function, the weights are determined; the weights, the overhead view image and the depth map are input into a preset sonar shadow generation model for multi-scale feature extraction, so as to output the hierarchical fusion map corresponding to the extracted hierarchical features.

[0093] In this application, the overhead view image and depth map are input into a pre-defined sonar shadow generation model. This model takes the overhead view optical image and depth map output by the view conversion module as input, and first maps the two types of data to a unified feature space through a specific embedding layer. The image is segmented into a sequence of image blocks through an overlapping image block embedding operation. Each image block has a size of P×P pixels, a stride of S, and an overlap of PS, i.e.:

[0094]

[0095] in, This is used to characterize the feature representation of an RGB image from an overhead view, after the overlapping image patch embedding operation is performed and combined with the position encoding matrix, and then mapped to a unified feature space. It is a holistic representation, representing PositionalEncoding (of RGB), which is RGB-based position embedding or RGB position encoding; This is an overhead view RGB image within an overhead view image; right Perform overlapping image patch embedding operation to segment the image into a sequence of image patches with a step size of S and an overlap of PS, and map them to the feature space. The source is the overhead view RGB image output by the view conversion module. This is a positional encoding matrix for RGB image blocks, used to preserve spatial location information; its source is an encoding matrix generated based on the spatial location of the image blocks.

[0096] The depth map is processed using the same embedding method, resulting in:

[0097]

[0098] in, This is used to represent the feature representation of the depth map from the overhead view after the overlapping image patch embedding operation and the corresponding position encoding matrix, which is then mapped to a unified feature space. As a whole, it represents the location embedding based on the depth map; This is a depth map; To The overlapping image patch embedding operation is performed, and the source is the overhead view depth map output by the view conversion module. This is the corresponding location encoding matrix, used to preserve spatial location information. It is derived from the encoding matrix generated based on the spatial location of the depth image patch.

[0099] To fully utilize the complementarity of RGB and depth information, a cross-attention mechanism is employed to achieve intermodal feature interaction. The fusion process is as follows:

[0100]

[0101] The network employs a multi-stage pyramid structure, with each stage containing a varying number of Transformer blocks to process feature maps at different resolutions. Each Transformer block utilizes a spatially reduced attention mechanism, employing adaptive sampling to reduce computational complexity while preserving key information. An adaptive weight allocation mechanism dynamically adjusts the importance of features at different scales through a combination of channel attention and spatial attention. Channel attention is calculated as follows:

[0102]

[0103] in, represents the channel attention weights for the i-th feature map, used to dynamically adjust the importance of features from different channels. The MLP (Multi-layer Perceptron) transforms the pooled features, using the sigmoid activation function to map the output values ​​to the 0-1 range. For the i-th feature map Perform global average pooling.

[0104] The contribution of different pre-defined features to sonar imaging is obtained. A pre-defined attention mechanism is used to evaluate the feature correlation between the overhead view image and the depth map based on the contribution. Weights are then generated through normalization using a transformation function. In this application, the transformation function can be softmax.

[0105] The spatial attention weights (or simply weights) of the i-th feature map are: The hierarchical features after weight modulation are as follows:

[0106]

[0107] in, This is an element-wise multiplication operation.

[0108] Multi-scale feature fusion is achieved through a feature pyramid network structure, employing a top-down path and lateral connections to ensure the effective integration of information at different scales.

[0109]

[0110] in, The pyramid features of the i-th stage in the feature pyramid network are used for multi-scale fusion. For the pyramid features of stage i+1 Upsampling; For weighted features Perform lateral connection.

[0111] The final fused feature map, i.e., the hierarchical fused map for:

[0112]

[0113] in, This is a characteristic of the pyramids in Phase 1. This is a characteristic of the second phase of the pyramids, and so on; For downsampling, This is a channel-dimensional splicing operation.

[0114] This application first evaluates the "contribution" of different features to the final sonar imaging, combining physical priors with a data-driven attention mechanism to generate dynamic weights, prioritizing and retaining the most critical information for sonar image generation. Through a pre-defined attention mechanism and transformation function, the "contribution" is quantified into precise weights. This method adaptively strengthens important feature channels and weakens redundant or noisy channels, ensuring that the "hierarchical fusion map" received by the subsequent generation model is a purified and enhanced high-quality feature representation, improving the quality of the input data from the source and laying the foundation for generating high-fidelity sonar images. The weighted multi-source data is then input into the sonar shadow generation model for multi-scale feature extraction. The final output "hierarchical fusion map" is no longer a simple accumulation of the original data. It is a feature pyramid that deeply integrates different scales and sources, and is weighted by importance, promoting accurate modeling of shadows and reflections in subsequent steps. Furthermore, in complex underwater environments, the quality and reliability of optical images and depth information vary with the scene. This application, through dynamic weight allocation, can adaptively adjust the fusion ratio of the two data sources.

[0115] According to some embodiments, the process of fusing and determining the sonar shadow feature map can specifically include: extracting the incident angle information corresponding to the hierarchical fusion map for geometric calculation from the depth map; obtaining the camera intrinsic parameters from the overhead view and extracting the image texture features corresponding to the hierarchical fusion map from the overhead view image; extracting the material acoustic properties corresponding to the hierarchical fusion map from the acoustic material library; determining the sound wave reflection intensity map based on the incident angle information, camera intrinsic parameters, image texture features, and material acoustic properties; constructing scene geometry based on the depth map to determine the occlusion relationship on the sound wave propagation path based on the scene geometry, thereby determining the sonar shadow region under the occlusion relationship; and fusing the sonar shadow region and the sound wave reflection intensity map according to weights to generate the sonar shadow feature map.

[0116] In this application, the incident angle information and image texture features corresponding to the hierarchical fusion map are first extracted from the depth map and the overhead view image, and the material acoustic properties corresponding to the hierarchical fusion map are extracted from the acoustic material library.

[0117] Based on the depth map, the normal vector and sound wave incident angle of each pixel in the scene are calculated. Assuming the sound wave is incident along a direction perpendicular to the overhead plane, a pixel in the layered fusion map is used. angle of incidence It can be calculated using the gradient information of the depth map, that is:

[0118]

[0119] in, This is the gradient vector of the depth map at this pixel. Let be the unit vector representing the incident direction of the sound wave. The prediction of reflection intensity also incorporates the influence of the material's acoustic properties. The model first classifies different materials in the scene using a material recognition network and outputs a material probability distribution. :

[0120]

[0121] in, This is a material classification network that outputs a probability distribution containing common materials. Based on the identified material type, the model assigns a corresponding acoustic reflection coefficient to each material, and the base reflection intensity of different materials... The weighted fusion yielded:

[0122]

[0123] in, Let be the probability distribution of the i-th material. Let be the basic reflectance coefficient of the i-th material.

[0124] To ensure that the prediction results conform to the laws of acoustic physics, the model introduces physical prior constraints to correct for the reflection intensity. The final prediction of the reflection intensity R combines the geometric angle of incidence and material properties.

[0125]

[0126] in, Use the Sigmoid activation function; and These are the parameters of the convolutional layer. A hierarchical fusion diagram; Angle attenuation coefficient related to the material; To avoid extremely small constants with zero values, thus ensuring that the intensity value of the strong reflection region is significantly higher than that of other regions.

[0127] The sound wave reflection intensity map is obtained based on the above process.

[0128] Occlusion relationship reasoning based on scene geometry first constructs a 3D point cloud using a depth map. ,in This represents the depth value. For each pixel, calculate all points along the sound wave propagation path; if such points exist... satisfy:

[0129]

[0130]

[0131]

[0132] If the pixel is in shadow, it is determined that the pixel is in shadow. To achieve end-to-end learning, the model transforms the occlusion detection process into a convolution operation, outputting a shadow probability map through a 1×1 convolutional layer. and the following constraints are applied:

[0133]

[0134] in, For point cloud computing-based occlusion labels (1 indicates occlusion, 0 indicates no occlusion); and These are the parameters of the convolutional layer; The input feature map, i.e., the hierarchical fusion map, is the output of the preceding feature extraction module.

[0135] Final sonar shadow feature map The fusion weights are obtained through adaptive fusion of the reflection intensity map and the shadow map. Predicted by a separate convolutional branch:

[0136]

[0137]

[0138] The fusion weights have a range of values. By learning to automatically balance the contribution of reflection intensity and shadow information, The image shows the reflection intensity, derived from the calculated sonar reflection intensity. Model training employs a hybrid loss function, combining pixel-level regression and perceptual loss to ensure that the generated feature maps structurally and effectively capture sonar characteristics while maintaining end-to-end differentiability.

[0139] This application combines incident angle information from depth maps, geometric projection relationships from camera intrinsic parameters, surface texture features from optical images, and acoustic properties from a material library. It calculates acoustic wave reflection intensity maps using a multi-parameter coupled physical model, overcoming the limitations of traditional methods that rely on a single data source. This ensures that the generated reflection intensity map not only conforms to the physical laws of acoustic reflection (such as its relation to incident angle and material) but also reflects the scattering effect of sound waves on the micro-geometry implied by the visual texture of the object's surface, significantly improving the realism and accuracy of the reflection information. Based on the 3D scene geometry constructed from the depth map, the sonar shadow region is determined by calculating the occlusion relationships along the sound wave propagation path. This process strictly follows the imaging mechanism of sonar in real underwater environments, ensuring that the generated shadows are consistent with the real sonar images in shape, position, and extent, fundamentally solving the distortion problems such as unreasonable shadow shapes and incorrect positions caused by data-driven methods. Finally, the acoustic wave reflection intensity map and the sonar shadow region are synthesized into a unified sonar shadow feature map through weighted fusion. This feature map simultaneously encodes the scene's reflection properties and geometric occlusion information, forming an intermediate representation that is highly instructive for subsequent style transfer, and providing the most critical data support for generating the final sonar image that combines visual realism and physical consistency.

[0140] According to some embodiments, the stepwise denoising method includes a noise addition method and an iterative denoising method. In the process of converting a sonar style image, specifically: multi-level features are extracted from the sonar shadow feature map based on a preset noise prediction network, wherein the multi-level features correspond to features corresponding to coarse-grained structures and fine-grained textures; in the forward diffusion stage, Gaussian noise is added to the sonar shadow feature map according to the noise addition method to determine the current noise state; in the reverse diffusion stage, based on the multi-level features, the current noise state, and the iterative denoising method, the sonar shadow feature map is iteratively predicted and noise is eliminated to determine the real sonar style image.

[0141] In this application, the sonar image style model is based on a diffusion model architecture. It transforms the sonar shadow feature map into a realistic sonar-style image through a progressive denoising process, and incorporates device characteristic simulation to enhance realism. The core of the diffusion model lies in constructing a Markov chain process, progressively adding Gaussian noise to the sonar shadow feature map through a multi-step forward diffusion process, and then generating the image through a reverse denoising process. The noise prediction network adopts the U-net architecture, acquiring multi-level features of the sonar image from coarse-grained structure to fine-grained texture through a multi-level cross-scale feature fusion mechanism. Based on the above settings, the sonar shadow feature map is progressively denoised to obtain a realistic sonar-style image.

[0142] In some implementations, the noisy image is assumed to be... Time embedding Using the conditional feature c as input, and following the noise addition method, the temporal embedding is implemented through sinusoidal positional encoding to predict the added noise:

[0143]

[0144] in, Gaussian noise predicted by the noise prediction network and added to the sonar image; The image at time t contains noise, derived from the image after noise was gradually added during the forward diffusion process; t is the time step of the diffusion process. For sinusoidal position encoding embedding at time step t, conditional features The cross-attention mechanism is used to input information into the network at various scales, ensuring that the generation process makes full use of the information from sonar shadow features. This is a noise prediction network.

[0145] To effectively preserve object contours and geometric structure information in sonar shadow features during generation, the model employs a multi-level cross-scale feature fusion mechanism. Each layer of the encoder fuses with corresponding scale features of the sonar shadow feature map through a feature alignment module, while the decoder further enhances the transmission of structural information through skip connections and attention mechanisms. Equipment characteristic modeling simulates the imaging characteristics of different sonar devices parametrically, dynamically adding noise and artifacts based on device characteristics, ensuring that artifacts match the object structure and conform to physical origins. Noise simulation is achieved through Gaussian distribution sampling, and the artifact part generates striped textures matching the beam angle based on the gradient direction of the object contour edges, ensuring that the artifact distribution is consistent with the physical origins of the object structure.

[0146] The inverse denoising process generates images through iterative prediction and noise removal, employing a deterministic sampling strategy to improve generation efficiency. Given sonar shadow features... and equipment parameters From pure noise Stepwise noise reduction:

[0147]

[0148] in, Characterizes the sonar image state from time t to time t-1 during the reverse denoising process; , The cumulative fidelity coefficient for the diffusion process originates from a pre-defined noise scheduling table, such as linear or cosine scheduling. The noise predicted by the noise prediction network; It is an adjustable randomness parameter; Control the degree of determinism in sampling; .

[0149] The model adopts Figure 1 The end-to-end training method combines adversarial loss, perceptual loss, and structural similarity loss in its loss function, ensuring that the generated images closely resemble real sonar images in style and detail while preserving the scene structure.

[0150] This application extracts multi-level features, ranging from coarse-grained structure to fine-grained texture, through a noise prediction network, enabling the model to simultaneously understand the macroscopic contours and microscopic details of the image during denoising. This multi-scale feature guidance mechanism ensures the directionality of generation, effectively preventing structural distortion or texture blurring during iteration, and guaranteeing the high integrity of the final image in terms of structure and detail. In the forward stage, Gaussian noise is added to randomize the data, and then in the reverse stage, noise is removed progressively and iteratively based on multi-level features. This process is essentially a robust approximation from a random distribution to the distribution of the target sonar image. It not only generates images with extremely high visual quality and noise patterns highly consistent with real sonar data, but also avoids the monotony of the generated results by introducing randomness, thus improving the diversity of the synthesized data. The progressive denoising strategy adopted in this application decomposes the complex style transfer task into a series of simple denoising sub-tasks, greatly reducing the learning difficulty of the model. Each step only needs to predict and remove a small amount of noise, thereby stably and controllably reconstructing detailed sonar textures.

[0151] According to some embodiments, during the forward diffusion phase, the noise intensity and distribution pattern of the forward diffusion phase are adjusted according to the frequency band characteristics and beamwidth of the sonar device to add Gaussian noise.

[0152] In this application, device characteristic simulation is achieved through parametric noise modulation. The device characteristic simulation through parametric noise modulation mainly involves the forward diffusion process, dynamically adjusting the noise intensity and distribution pattern according to parameters such as the frequency band characteristics and beamwidth of the sonar device to inject device-specific noise textures and artifacts; the adjustment is maintained during the reverse process to ensure the consistency of the generated image in terms of realism.

[0153] This application adaptively adjusts the intensity and distribution pattern of noise during the forward diffusion process based on the frequency band characteristics and beamwidth of a specific sonar device. This allows the noise addition process to simulate the inherent imaging characteristics differences caused by the physical differences of different sonar devices. By incorporating device parameters as conditions into the forward process, this application guides the model to not only reconstruct the general appearance of the sonar image during reverse denoising, but also to learn to reconstruct the image style under specific device parameters. This injection of physical priors greatly constrains the generation process, enabling the model to learn the data distribution of the target sonar image more accurately and effectively avoiding the generation of unrealistic or device-irrelevant invalid images.

[0154] According to other embodiments, Figure 1The specific process of feedback and optimization based on the multi-stage evaluation results includes:

[0155] By constructing a phased comprehensive evaluation framework, each stage is quantitatively evaluated and dynamically adjusted, and the overall system performance is continuously improved through a feedback optimization mechanism. The phased comprehensive evaluation framework establishes specific evaluation standards and quantitative indicators for each processing module in the system, ensuring the output quality of each stage through a multi-level evaluation system. The evaluation indicators for the viewpoint transformation stage mainly revolve around two core dimensions: the quality of the overhead view generation and the accuracy of depth information. Image quality indicators include peak signal-to-noise ratio and structural similarity index, while depth information accuracy is evaluated using mean absolute error.

[0156] When the evaluation system detects distortion or depth estimation bias in the viewpoint transition, it first analyzes the specific type of bias and then adjusts the viewpoint transformation parameters accordingly, including the selection of camera intrinsic matrix, transformation matrix, and interpolation method. For the depth estimation network, the system dynamically adjusts the network weights and parameters of the geometric inference-related layers. Simultaneously, the system identifies the specific scene types that cause the bias and adds training samples for those specific scene types.

[0157] The evaluation of the sonar shadow generation stage focuses on three key dimensions: reflection intensity and shadow area prediction accuracy, physical property fidelity, and indicators including shadow integrity (SI) and echo intensity consistency (EIC). I represents the percentage of the intersection between the effective shadow region and the real shadow region in the generated image, i.e.:

[0158]

[0159] in, The shadow regions in the image are derived from the shadow map output by the sonar shadow generation model; These are real shadow areas, derived from shadow maps annotated in real-world scenes.

[0160] EIC is measured by the Barthold distance of the echo intensity distribution, using the following formula:

[0161]

[0162] Where p is the intensity distribution probability and k is the intensity level number, which is derived from the number of levels divided according to the range of echo intensity values. To generate the probability distribution of image echo intensity level k; This represents the probability distribution of the echo intensity level k in the true image.

[0163] Based on the calculated index results, if deviations exist, the physical model parameters are dynamically calibrated. By analyzing the deviation patterns between the predicted results and the actual data, the material-related parameters and angle attenuation coefficients in the acoustic reflection model are automatically adjusted. The evaluation of the sonar image stylization stage employs a multi-dimensional comprehensive quality assessment system, integrating traditional image quality indicators, perceptual quality assessment, and sonar-specific texture consistency indicators. Traditional image quality indicators include peak signal-to-noise ratio, structural similarity index, and mean square error, which assess the similarity between the generated image and the real sonar image at the pixel level. Perceptual quality assessment uses a deep learning-based perceptual loss function, comparing the semantic similarity between the generated image and the real image through high-level features extracted by a pre-trained network.

[0164] When the generated image is detected to have deficiencies in certain quality dimensions, the system analyzes the specific problem type and takes corresponding optimization measures. If the perceptual quality assessment shows that semantic information is lost, the system will increase the influence weight of conditional features; if the texture consistency assessment shows that the texture pattern does not conform to sonar features, the system will adjust the parameter settings of the device characteristic simulation; if traditional quality indicators show that pixel-level errors are large, the system will optimize the number of diffusion sampling steps and strategies.

[0165] System performance evaluation not only focuses on output quality but also simultaneously monitors operational performance metrics such as processing speed, memory usage, and computing resource utilization, providing quantitative data for hardware configuration optimization and algorithm efficiency improvement. Processing speed monitoring includes the individual processing time of each stage, the overall workflow, and the end-to-end latency of each stage. By establishing performance benchmarks and real-time monitoring, performance bottlenecks and anomalies can be identified promptly. Memory usage monitoring covers peak memory usage, memory allocation patterns, and memory leak detection, ensuring system stability during long-term operation.

[0166] Computational resource utilization monitoring includes CPU (Central Processing Unit) utilization, GPU (Graphics Processing Unit) utilization, and memory bandwidth utilization. Through analysis of resource usage patterns, it identifies inefficient resource allocation. The system establishes a correlation analysis model between performance and quality metrics, maximizing performance through algorithm optimization and parameter tuning while ensuring output quality. When excessive processing time is detected at a certain stage, the system analyzes the specific causes of the bottleneck, and possible optimization measures include model pruning, quantization optimization, parallel computing strategy adjustments, and improvements to the caching mechanism.

[0167] The feedback tuning system also includes an adaptive load balancing mechanism that dynamically adjusts the computational allocation at each stage based on actual processing needs and hardware resource availability.

[0168] The following describes an apparatus embodiment of this application, which can be used to perform the method embodiment of this application. For details not disclosed in the apparatus embodiment of this application, please refer to the method embodiment of this application.

[0169] Figure 3 This is a block diagram of a sonar image synthesis apparatus provided in an embodiment of this application. Figure 3 As shown, the sonar image synthesis device 300 includes an image conversion model 301, a hierarchical image determination module 302, an image fusion module 303, and a noise reduction module 304.

[0170] Image conversion model 301 is used to acquire basic optical images of underwater objects captured by an optical camera, and convert the basic optical images into overhead view images and depth maps according to a preset view conversion model.

[0171] The hierarchical image determination module 302 is used to determine the weights of different features based on the contribution of different features in the overhead view image and depth map to sonar imaging, and to perform multi-scale feature extraction on the overhead view image and depth map based on the weights to determine the corresponding hierarchical fusion map.

[0172] The image fusion module 303 is used to determine the acoustic wave reflection intensity map and sonar shadow area of ​​the surface of the underwater object based on the overhead view image, depth map, predefined acoustic material library and hierarchical fusion map, and to fuse and determine the sonar shadow feature map based on the acoustic wave reflection intensity map and sonar shadow area.

[0173] The denoising module 304 is used to convert the sonar shadow feature map into a real sonar style image based on a preset stepwise denoising method.

[0174] Optionally, the viewpoint transformation model includes a multi-scale feature extraction network with a hierarchical convolutional structure, a depth estimation module, and a viewpoint transformation module; the image transformation model 301 is specifically used for:

[0175] Acquire basic optical images of underwater objects captured by an optical camera;

[0176] Based on a multi-scale feature extraction network, features are extracted from the basic optical image to determine multi-scale feature information from low to high order.

[0177] Multi-scale feature information is optimized, and the optimized multi-scale feature information is fused according to a preset ratio to generate fused feature information;

[0178] Based on the depth estimation module, a depth map corresponding to the fused feature information is output;

[0179] Based on the viewpoint transformation module, depth map, and camera intrinsics of the optical camera, a coordinate transformation matrix is ​​constructed. Then, based on an end-to-end differentiable transformation method and depth map, the pixels in the basic optical image are mapped to the corresponding positions in the overhead view to generate an overhead view image.

[0180] Optionally, the viewpoint transformation model includes a hybrid network structure combining sliding window feature aggregation and attention mechanisms, a depth estimation module, and a viewpoint transformation module; the image transformation model 301 is specifically used for:

[0181] Acquire basic optical images of underwater objects captured by an optical camera;

[0182] Multi-scale features are extracted from the basic optical image by windowing different receptive fields of the hybrid network structure, and the multi-scale features are fused according to the dynamic weights determined by the adaptive calculation of the attention mechanism to determine the fused feature information.

[0183] Based on the depth estimation module, a depth map corresponding to the pixels of the base optical image is generated, wherein the depth map is a dense depth map corresponding to the pixels of the base optical image;

[0184] The basic optical image is converted into an overhead view image using the perspective conversion module.

[0185] Optionally, the layer image determination module 302 is specifically used for:

[0186] To obtain the contribution of different features in the overhead view image and depth map to sonar imaging;

[0187] Weights are determined based on contribution, a pre-defined attention mechanism, and a transformation function;

[0188] The weights, overhead view image, and depth map are input into a pre-defined sonar shadow generation model for multi-scale feature extraction, and the resulting hierarchical fusion map is output.

[0189] Optionally, the image fusion module 303 is specifically used for:

[0190] Extract the incident angle information corresponding to the hierarchical fusion map for geometric calculations from the depth map;

[0191] Obtain camera intrinsic parameters from an overhead view and extract image texture features corresponding to the hierarchical fusion map from the overhead view image;

[0192] Extract the material acoustic properties corresponding to the hierarchical fusion map from the acoustic material library;

[0193] The sound wave reflection intensity map is determined based on the incident angle information, camera intrinsic parameters, image texture features, and material acoustic properties.

[0194] Scene geometry is constructed based on depth maps to determine occlusion relationships on the sound wave propagation path, thereby identifying sonar shadow areas under occlusion relationships.

[0195] Based on the weights, the sonar shadow region and the acoustic wave reflection intensity map are fused to generate a sonar shadow feature map.

[0196] Optionally, the stepwise denoising method includes noise augmentation and iterative denoising; the denoising module 304 is specifically used for:

[0197] Multi-level features are extracted from the sonar shadow feature map based on a pre-set noise prediction network. The multi-level features correspond to the features corresponding to coarse-grained structure and fine-grained texture.

[0198] During the forward diffusion phase, Gaussian noise is added to the sonar shadow feature map according to the noise increase method to determine the current noise state;

[0199] In the backdiffusion stage, based on multi-level features, the current noise state, and the iterative denoising method, the sonar shadow feature map is iteratively predicted and noise is eliminated to determine the true sonar style image.

[0200] Optionally, during the forward diffusion phase, the denoising module 304 adds Gaussian noise to the sonar shadow feature map according to the noise increase method to determine the current noise state. Specifically, this is used for:

[0201] During the forward diffusion phase, the noise intensity and distribution pattern are adjusted according to the frequency band characteristics and beamwidth of the sonar equipment to add Gaussian noise.

[0202] The device performs functions similar to those described above; other functions are described in the preceding descriptions and will not be repeated here.

[0203] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 4 As shown, the electronic device 400 of this embodiment may include a memory 401 and a processor 402.

[0204] The memory 401 stores a computer program, which, when executed by the processor 402, causes the processor 402 to perform the method described in the above embodiment.

[0205] The processor 402 and the memory 401 are connected, for example, via a bus.

[0206] Optionally, the electronic device 400 may also include a transceiver. It should be noted that in practical applications, the transceiver is not limited to one, and the structure of the electronic device 400 does not constitute a limitation on the embodiments of this application.

[0207] Processor 402 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 402 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0208] A bus can include a pathway for transmitting information between the aforementioned components. The bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one thick line is used in the diagram, but this does not imply that there is only one bus or one type of bus.

[0209] The memory 401 can be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or it can be EEPROM (Electrically Erasable Programmable Read Only Memory), CD. ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed discs, laser discs, optical discs, digital universal discs, Blu-ray discs, etc.), disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0210] The memory 401 is used to store application code that executes the solution of this application, and its execution is controlled by the processor 402. The processor 402 is used to execute the application code stored in the memory 401 to implement the content shown in the foregoing method embodiments.

[0211] Electronic devices include, but are not limited to: mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Servers can also be included. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0212] The electronic device in this embodiment can be used to execute the method of any of the above embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.

[0213] This application also provides a non-transitory computer-readable storage medium storing computer-readable instructions thereon, which, when executed by a processor, cause the processor to perform the method as described in the above embodiments.

[0214] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a non-transitory computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0215] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. Furthermore, any changes or modifications made by those skilled in the art based on the ideas of this application, and on the specific implementation methods and application scope of this application, are all within the scope of protection of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method of synthesizing sonar images, characterized by, include: Acquire basic optical images of underwater objects captured by an optical camera, and convert the basic optical images into overhead view images and depth maps according to a preset viewpoint conversion model; Based on the contribution of different features in the overhead view image and the depth map to sonar imaging, the weights of the different features are determined, and multi-scale feature extraction is performed on the overhead view image and the depth map based on the weights to determine the corresponding hierarchical fusion map. Based on the overhead view image, the depth map, the predefined acoustic material library, and the hierarchical fusion map, the acoustic wave reflection intensity map and the sonar shadow region on the surface of the underwater object are determined, and a sonar shadow feature map is determined by fusing the acoustic wave reflection intensity map and the sonar shadow region. Based on a preset stepwise denoising method, the sonar shadow feature map is converted into a real sonar style image; The step of determining the acoustic wave reflection intensity map and sonar shadow region of the underwater object surface based on the overhead view image, the depth map, the predefined acoustic material library, and the hierarchical fusion map, and then fusing and determining the sonar shadow feature map based on the acoustic wave reflection intensity map and the sonar shadow region, includes: Extract the incident angle information corresponding to the hierarchical fusion map used for geometric calculations from the depth map; Obtain camera intrinsic parameters from an overhead view, and extract image texture features corresponding to the hierarchical fusion map from the overhead view image; Extract the material acoustic properties corresponding to the hierarchical fusion map from the acoustic material library; The sound wave reflection intensity map is determined based on the incident angle information, the camera intrinsic parameters, the image texture features, and the material acoustic properties; Based on the depth map, scene geometry is constructed to determine the occlusion relationship on the sound wave propagation path, thereby determining the sonar shadow region under the occlusion relationship; Based on the weights, the sonar shadow region and the acoustic wave reflection intensity map are fused to generate the sonar shadow feature map.

2. The method of claim 1, wherein, The viewpoint transformation model includes a multi-scale feature extraction network with a hierarchical convolutional structure, a depth estimation module, and a viewpoint transformation module. The step of acquiring a basic optical image of an underwater object captured by an optical camera, and converting the basic optical image into an overhead view image and a depth map according to a preset viewpoint transformation model, includes: Acquire basic optical images of underwater objects captured by an optical camera; Based on the multi-scale feature extraction network, feature extraction is performed on the basic optical image to determine multi-scale feature information from low to high order; The multi-scale feature information is optimized, and the optimized multi-scale feature information is fused according to a preset ratio to generate fused feature information. The depth estimation module obtains the depth map corresponding to the fused feature information. Based on the viewpoint transformation module, the depth map, and the camera intrinsic parameters of the optical camera, a coordinate transformation matrix is ​​constructed. Then, based on the end-to-end differentiable transformation method and the depth map, the pixels in the basic optical image are mapped to the corresponding positions in the overhead view to generate the overhead view image.

3. The method of claim 1, wherein, The viewpoint transformation model includes a hybrid network structure combining sliding window feature aggregation and attention mechanism, a depth estimation module, and a viewpoint transformation module. The step of acquiring a basic optical image of an underwater object captured by an optical camera, and converting the basic optical image into an overhead view image and a depth map according to a preset viewpoint transformation model, includes: Acquire basic optical images of underwater objects captured by an optical camera; Multi-scale features in the basic optical image are extracted based on the windowing of different receptive fields of the hybrid network structure, and the multi-scale features are fused according to the dynamic weights determined by the adaptive calculation of the attention mechanism to determine the fused feature information. The depth estimation module generates a depth map corresponding to the pixels of the base optical image, wherein the depth map is a dense depth map corresponding to the pixels of the base optical image. The perspective transformation module converts the basic optical image into the overhead view image.

4. The method of claim 1, wherein, The step involves determining the weights of different features in the overhead view image and the depth map based on their contribution to sonar imaging, and then performing multi-scale feature extraction on the overhead view image and the depth map based on these weights to determine the corresponding hierarchical fusion map, including: The contribution of different features in the overhead view image and the depth map to sonar imaging is obtained; The weights are determined based on the contribution level, the preset attention mechanism, and the transformation function. The weights, the overhead view image, and the depth map are input into a preset sonar shadow generation model for multi-scale feature extraction, so as to output a hierarchical fusion map corresponding to the extracted hierarchical features.

5. The method according to any one of claims 1 to 4, characterized in that, The progressive denoising method includes noise augmentation and iterative denoising. The stepwise denoising method based on a preset stepwise denoising approach, which converts the sonar shadow feature map into a realistic sonar-style image, includes: Multi-level features are extracted from the sonar shadow feature map based on a preset noise prediction network, wherein the multi-level features correspond to features corresponding to coarse-grained structure and fine-grained texture. During the forward diffusion phase, Gaussian noise is added to the sonar shadow feature map according to the noise increase method to determine the current noise state; During the back-diffusion phase, based on the multi-level features, the current noise state, and the iterative denoising method, the sonar shadow feature map is iteratively predicted and noise is eliminated to determine the true sonar style image.

6. The method according to claim 5, characterized in that, During the forward diffusion phase, Gaussian noise is added to the sonar shadow feature map according to the noise increase method to determine the current noise state, including: During the forward diffusion phase, the noise intensity and distribution pattern of the forward diffusion phase are adjusted according to the frequency band characteristics and beamwidth of the sonar device to add Gaussian noise.

7. A sonar image synthesis apparatus, characterized in that, include: An image conversion model is used to acquire a basic optical image of an underwater object captured by an optical camera, and to convert the basic optical image into an overhead view image and a depth map according to a preset view conversion model. The hierarchical image determination module is used to determine the weights of the different features based on their contribution to sonar imaging in the overhead view image and the depth map, and to perform multi-scale feature extraction on the overhead view image and the depth map based on the weights to determine the corresponding hierarchical fusion map. The image fusion module is used to determine the acoustic wave reflection intensity map and sonar shadow region of the underwater object surface based on the overhead view image, the depth map, the predefined acoustic material library and the hierarchical fusion map, and to fuse and determine the sonar shadow feature map based on the acoustic wave reflection intensity map and the sonar shadow region. The denoising module is used to convert the sonar shadow feature map into a real sonar style image based on a preset stepwise denoising method. Specifically, the image fusion module is used for: Extract the incident angle information corresponding to the hierarchical fusion map used for geometric calculations from the depth map; Obtain camera intrinsic parameters from an overhead view, and extract image texture features corresponding to the hierarchical fusion map from the overhead view image; Extract the material acoustic properties corresponding to the hierarchical fusion map from the acoustic material library; The sound wave reflection intensity map is determined based on the incident angle information, the camera intrinsic parameters, the image texture features, and the material acoustic properties; Based on the depth map, scene geometry is constructed to determine the occlusion relationship on the sound wave propagation path, thereby determining the sonar shadow region under the occlusion relationship; Based on the weights, the sonar shadow region and the acoustic wave reflection intensity map are fused to generate the sonar shadow feature map.

8. An electronic device, characterized in that, include: processor; A memory storing a computer program that, when executed by the processor, causes the processor to perform the method as described in any one of claims 1-6.

9. A non-transitory computer-readable storage medium, characterized in that, It stores computer-readable instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Underwater image processing and target identification method and device, storage medium and electronic equipment

    CN117079117A

  • Automatic enhancement method for sonar image

    CN117853359A