Underwater sonar and optical image feature fusion method and system

By adopting a feature fusion network with a cross-modal attention mechanism in the fusion of underwater sonar and optical image features, the problem that existing methods are not effective in complex underwater environments is solved, and higher target perception accuracy and environmental adaptability are achieved.

CN120671083APending Publication Date: 2025-09-19HARBIN ENG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510794350.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing underwater sonar and optical image feature fusion methods do not work well in complex underwater environments. They find it difficult to fully utilize the advantages of the two modalities, resulting in insufficient target perception accuracy and robustness.

Method used

A cross-modal attention mechanism is used to interact sonar data with optical image features. By constructing a feature fusion network, including a multi-scale deformable attention mechanism and a fully connected layer, efficient fusion of sonar and optical features is achieved.

Benefits of technology

It improves the accuracy and environmental adaptability of underwater target perception and enhances the robustness of target recognition, especially in scenes with blurred target boundaries and low contrast.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671083A_ABST
    Figure CN120671083A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater sonar and optical image feature fusion method and system, and belongs to the technical field of deep learning. According to the method, the problem that the feature fusion effect of an existing method is poor due to the fact that feature complementarity mining and robustness are insufficient is solved. According to the method, the sonar data and the optical image features are interacted by optimizing the feature level fusion strategy and adopting the cross-modal attention mechanism, the complementarity of the cross-modal features is enhanced, and the feature fusion effect is improved. The long-distance detection capability of sonar data is combined with the high-resolution detail information of the optical image, so that the accuracy of target sensing is improved. Aiming at complex working conditions of underwater illumination change, turbid environment and noise interference, through sonar and optical feature level fusion of the method, adaptability of the model to different underwater environments can be improved, and robustness of target identification is ensured. The method can be applied to feature fusion of underwater sonar and optical images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning technology, and specifically relates to a method and system for fusing underwater sonar and optical image features. Background Art

[0002] In the field of underwater target perception, sonar and optical imagery are two common sensing methods. Sonar can penetrate turbid water and provide target information at long distances, but its imaging resolution is low and it is susceptible to noise and multipath effects. In contrast, optical imagery has higher resolution and can provide rich, detailed information, but its detection range is short due to limitations in underwater lighting conditions and visibility. Therefore, a single sensing method cannot achieve high-precision target detection and recognition in complex underwater environments. In recent years, underwater perception methods that fuse multimodal data have received widespread attention. Among them, how to efficiently fuse the features of sonar and optical imagery to improve the accuracy and robustness of target perception has become a new research hotspot.

[0003] Currently, methods for fusing underwater sonar and optical images primarily include pixel-level, feature-level, and decision-level fusion strategies. However, pixel-level fusion requires strict alignment of data from different modalities, making it difficult to adapt to complex underwater environments. Decision-level fusion, on the other hand, only post-processes the results of independent recognition of each modality and fails to fully utilize cross-modal information. Feature-level fusion achieves a balance between information utilization and computational complexity, but existing methods still lack cross-modal feature alignment, feature complementarity mining, and robustness. Therefore, the feature fusion performance of existing methods remains poor. An efficient underwater sonar and optical image feature fusion method is urgently needed to fully leverage the advantages of both modalities and improve underwater target perception and environmental adaptability. Summary of the Invention

[0004] The purpose of the present invention is to solve the problem of poor feature fusion effect of existing methods and to propose a method and system for fusing underwater sonar and optical image features.

[0005] The technical solution adopted by the present invention to solve the above technical problems is: According to one aspect of the present invention, a method for fusing underwater sonar and optical image features comprises the following steps: Step 1: The underwater robot is equipped with an optical camera and a sonar system, which are used to synchronously acquire optical images and acoustic data; Step 2: After preprocessing the optical image, extract the optical features in the preprocessed image ; After preprocessing the acoustic data, extract the acoustic features from the preprocessed acoustic data ; Step 3: Construct a feature fusion network for fusing optical features and acoustic features, wherein the feature fusion network includes a first CALayer module, a second CALayer module, a third CALayer module, and a first fully connected layer; Step 4: Extracting optical features Harmonic characteristics After processing, the processed optical features, processed acoustic features, and optical features Harmonic characteristics As the input of the feature fusion network, the feature fusion result is output through the feature fusion network.

[0006] Furthermore, after the optical image is preprocessed, the optical features in the preprocessed image are extracted. ; Specifically: Enhance the optical image and then use a deep neural network to extract optical features from the enhanced image .

[0007] Furthermore, after the acoustic data is preprocessed, the acoustic features in the preprocessed acoustic data are extracted. Specifically: The acoustic data is processed with noise reduction and signal enhancement in sequence, and then the Transformer model is used to extract acoustic features from the processed acoustic data. .

[0008] Furthermore, the extracted optical features Harmonic characteristics Processing, specifically: Optical characteristics Perform position encoding to obtain position encoding results , and then give the extracted optical features Add position encoding , and obtain the processed optical features :

[0009] Acoustic characteristics Perform position encoding to obtain position encoding results , and then give the extracted acoustic features Add position encoding , get the processed acoustic features :

[0010] Furthermore, the working process of the feature fusion network is as follows: Step 1: ,Will 、 and As the input of the first CALayer module, get the output of the first CALayer module ; Step 2: ,Will 、 and As the input of the second CALayer module, get the output of the second CALayer module ; Step 3: ,Will 、 and As the input of the third CALayer module, get the output of the third CALayer module ; Step 4: Optical features and Perform splicing and use the splicing result as the input of the first fully connected layer, and output the fusion feature through the first fully connected layer :

[0011] in, Represents feature concatenation operation; Linear represents a fully connected layer.

[0012] Furthermore, the first CALayer module includes a multi-scale deformable attention mechanism and a fully connected layer, and the working process of the first CALayer module is: Will 、 and As the input of the multi-scale deformable attention mechanism, the output of the multi-scale deformable attention mechanism is used as the input of the fully connected layer, and the output of the fully connected layer is used as the output of the first CALayer module :

[0013] in, Represents the first CALayer module.

[0014] According to another aspect of the present invention, an underwater sonar and optical image feature fusion system includes an optical image acquisition module, an acoustic data acquisition module, an optical feature extraction module, an acoustic feature extraction module, an optical feature processing module, an acoustic feature processing module, and a feature fusion module; wherein: The optical image acquisition module and the acoustic data acquisition module are used to acquire optical images and acoustic data respectively; The optical feature extraction module is used to extract the features of the optical image, and the optical feature processing module is used to process the features of the optical image; The acoustic feature extraction module is used to extract features of acoustic data, and the acoustic feature processing module is used to process the features of acoustic data; The feature fusion module is used to fuse the processed optical features and the processed acoustic features to obtain a feature fusion result.

[0015] Furthermore, the optical feature extraction module is used to extract features of the optical image, and the optical feature processing module is used to process the features of the optical image; the specific process is: Step S1: Enhance the optical image and then use a deep neural network to extract optical features from the enhanced image. ; Step S2: Optical characteristics Perform position encoding to obtain position encoding results , and then give the extracted optical features Add position encoding , and obtain the processed optical features : .

[0016] Furthermore, the acoustic feature extraction module is used to extract features of acoustic data, and the acoustic feature processing module is used to process the features of acoustic data; the specific process is: Step A1: Perform noise reduction and signal enhancement on the acoustic data in sequence, and then use the Transformer model to extract acoustic features from the processed acoustic data. ; Step A2: Acoustic features Perform position encoding to obtain position encoding results , and then give the extracted acoustic features Add position encoding , get the processed acoustic features : .

[0017] Furthermore, the working process of the feature fusion module is as follows: Step 1: ,Will 、 and As the input of the first CALayer module, get the output of the first CALayer module ; The working process of the first CALayer module is: Will 、 and As the input of the multi-scale deformable attention mechanism, the output of the multi-scale deformable attention mechanism is used as the input of the fully connected layer, and the output of the fully connected layer is used as the output of the first CALayer module :

[0018] in, Represents the first CALayer module; Step 2: ,Will 、 and As the input of the second CALayer module, get the output of the second CALayer module ; Step 3: ,Will 、 and As the input of the third CALayer module, get the output of the third CALayer module ; Step 4: Optical features and Perform splicing and use the splicing result as the input of the first fully connected layer, and output the fusion feature through the first fully connected layer :

[0019] in, Represents feature concatenation operation; Linear represents a fully connected layer.

[0020] The beneficial effects of the present invention are: The present invention optimizes the feature-level fusion strategy and adopts a cross-modal attention mechanism to interact sonar data with optical image features, thereby enhancing cross-modal feature complementarity and improving feature fusion effects. The long-range detection capability of sonar data is combined with the high-resolution detail information of optical images, thereby improving the accuracy of target perception. For complex working conditions such as underwater lighting changes, turbid environments, and noise interference, the sonar and optical feature-level fusion of the method of the present invention can improve the adaptability of the model to different underwater environments and ensure the robustness of target recognition. At the same time, the method of the present invention can achieve target detection accuracy that is superior to existing methods in different underwater scenarios, especially in scenarios with blurred target boundaries and low contrast, significantly improving the perception capabilities of underwater unmanned equipment and providing reliable technical support for underwater detection, monitoring, and operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a schematic diagram of the CALayer module based on the multi-scale deformable attention mechanism; Figure 2 It is a feature fusion flowchart based on the alignment of acoustic features and optical features in the optical feature space. DETAILED DESCRIPTION

[0022] Specific implementation method 1: Combination Figure 2 This embodiment describes a method for fusing underwater sonar and optical image features, which specifically includes the following steps: Step 1: Depth control of the underwater robot is achieved through precise path planning. The underwater robot is equipped with a high-resolution multispectral optical camera and a broadband multibeam sonar system. These systems simultaneously acquire high-definition optical images and high-precision acoustic data, ensuring precise temporal and spatial correspondence between the two types of data using spatiotemporal synchronization technology. Step 2: After preprocessing the optical image, extract the optical features in the preprocessed image ; After preprocessing the acoustic data, extract the acoustic features from the preprocessed acoustic data The extracted optical and acoustic features not only retain the key information of the original data, but also provide standardized input for subsequent data fusion and application; Step 3: Construct a feature fusion network for fusing optical features and acoustic features, wherein the feature fusion network includes a first CALayer module, a second CALayer module, a third CALayer module, and a first fully connected layer; Step 4: Extracted optical features Harmonic characteristics After processing, the processed optical features, processed acoustic features, and optical features Harmonic characteristics As the input of the feature fusion network, the feature fusion result is output through the feature fusion network.

[0023] To address the problems in underwater target recognition and detection scenarios, where sonar has a long detection distance but obtains less feature information, while optical cameras obtain more precise feature information but have a smaller detection range, this paper proposes a new feature fusion method. This method enables underwater robots to comprehensively utilize the respective advantages of sound and light when performing target recognition and detection, achieving preliminary screening of distant targets and accurate recognition of close-range targets.

[0024] Specific embodiment 2: This embodiment differs from the specific embodiment 1 in that after preprocessing the optical image, the optical features in the preprocessed image are extracted. Specifically: Enhance the optical image and then use a deep neural network to extract optical features from the enhanced image .

[0025] Other steps and parameters are the same as those in the first embodiment.

[0026] The present invention can eliminate blur and color distortion caused by water bodies through enhancement processing, thereby ensuring the accuracy of optical feature extraction.

[0027] Specific embodiment three: This embodiment is different from specific embodiment one or two in that after preprocessing the acoustic data, the acoustic features in the preprocessed acoustic data are extracted. Specifically: The acoustic data is processed with noise reduction and signal enhancement in sequence, and then the Transformer model is used to extract acoustic features from the processed acoustic data. .

[0028] Other steps and parameters are the same as those in the first or second embodiment.

[0029] By processing acoustic data, environmental interference can be removed, data quality can be improved, and the accuracy of acoustic feature extraction can be ensured.

[0030] Specific embodiment 4: This embodiment is different from any one of the specific embodiments 1 to 3 in that the optical features to be extracted are Harmonic characteristics Processing, specifically: Optical characteristics Perform position encoding to obtain position encoding results , and then give the extracted optical features Add position encoding , and obtain the processed optical features :

[0031] Acoustic characteristics Perform position encoding to obtain position encoding results , and then give the extracted acoustic features Add position encoding , get the processed acoustic features :

[0032] The other steps and parameters are the same as those in the first to third embodiments.

[0033] Specific embodiment 5: This embodiment differs from specific embodiments 1 to 4 in that the working process of the feature fusion network is as follows: Step 1: ,Will 、 and As the input of the first CALayer module, get the output of the first CALayer module ; Step 2: ,Will 、 and As the input of the second CALayer module, get the output of the second CALayer module ; Step 3: ,Will 、 and As the input of the third CALayer module, get the output of the third CALayer module ; Step 4: Optical features and Perform splicing and use the splicing result as the input of the first fully connected layer, and output the fusion feature through the first fully connected layer :

[0034] in, Represents feature concatenation operation; Linear represents a fully connected layer.

[0035] The other steps and parameters are the same as those in the first to fourth embodiments.

[0036] Specific implementation method six: combination Figure 1This embodiment differs from any one of the first to fifth embodiments in that the first CALayer module includes a multi-scale deformable self-attention mechanism and a fully connected layer, and the working process of the first CALayer module is as follows: Will 、 and As the input of the multi-scale deformable attention mechanism, the output of the multi-scale deformable attention mechanism is used as the input of the fully connected layer, and the output of the fully connected layer is used as the output of the first CALayer module :

[0037] in, Represents the first CALayer module.

[0038] The other steps and parameters are the same as those in the first to fifth embodiments.

[0039] The CALayer module of the present invention can align sonar features to optical features in the optical feature space. The working process of each CALayer module in the present invention is the same, and each CALayer module retains the valid information of the previous alignment through residual connections. Through a step-by-step alignment method, the acoustic features aligned with the optical features are finally obtained.

[0040] Specific embodiment seven: This embodiment describes an underwater sonar and optical image feature fusion system, which includes an optical image acquisition module, an acoustic data acquisition module, an optical feature extraction module, an acoustic feature extraction module, an optical feature processing module, an acoustic feature processing module, and a feature fusion module; wherein: The optical image acquisition module and the acoustic data acquisition module are used to acquire optical images and acoustic data respectively; The optical feature extraction module is used to extract the features of the optical image, and the optical feature processing module is used to process the features of the optical image; The acoustic feature extraction module is used to extract features of acoustic data, and the acoustic feature processing module is used to process the features of acoustic data; The feature fusion module is used to fuse the processed optical features and the processed acoustic features to obtain a feature fusion result.

[0041] Specific embodiment eight: This embodiment differs from specific embodiment seven in that the optical feature extraction module is used to extract the features of the optical image, and the optical feature processing module is used to process the features of the optical image; the specific process is as follows: Step S1: Enhance the optical image and then use a deep neural network to extract optical features from the enhanced image. ; Step S2: Optical characteristics Perform position encoding to obtain position encoding results , and then give the extracted optical features Add position encoding , and obtain the processed optical features : .

[0042] Other steps and parameters are the same as those in the seventh embodiment.

[0043] Specific embodiment 9: This embodiment differs from specific embodiment 7 or 8 in that the acoustic feature extraction module is used to extract the features of acoustic data, and the acoustic feature processing module is used to process the features of acoustic data; the specific process is as follows: Step A1: Perform noise reduction and signal enhancement on the acoustic data in sequence, and then use the Transformer model to extract acoustic features from the processed acoustic data. ; Step A2: Acoustic features Perform position encoding to obtain position encoding results , and then give the extracted acoustic features Add position encoding , get the processed acoustic features : .

[0044] Other steps and parameters are the same as those in the seventh or eighth embodiment.

[0045] Specific embodiment 10: This embodiment differs from any one of specific embodiments 7 to 9 in that the working process of the feature fusion module is as follows: Step 1: ,Will 、 and As the input of the first CALayer module, get the output of the first CALayer module ; The working process of the first CALayer module is: Will 、 and As the input of the multi-scale deformable attention mechanism, the output of the multi-scale deformable attention mechanism is used as the input of the fully connected layer, and the output of the fully connected layer is used as the output of the first CALayer module :

[0046] in, Represents the first CALayer module; Step 2: ,Will 、 and As the input of the second CALayer module, get the output of the second CALayer module ; Step 3: ,Will 、 and As the input of the third CALayer module, get the output of the third CALayer module ; Step 4: Optical features and Perform splicing and use the splicing result as the input of the first fully connected layer, and output the fusion feature through the first fully connected layer :

[0047] in, Represents feature concatenation operation; Linear represents a fully connected layer.

[0048] The other steps and parameters are the same as those in any one of the seventh to ninth embodiments.

[0049] The above examples are merely illustrative of the calculation model and process of the present invention and are not intended to limit the embodiments of the present invention. Persons skilled in the art will readily appreciate that other variations or modifications based on the above description are possible. This list of embodiments is not exhaustive; however, any obvious variations or modifications derived from the technical solution of the present invention remain within the scope of protection of the present invention.

Claims

1. A method for fusing underwater sonar and optical image features, characterized in that: The method specifically comprises the following steps: Step 1: The underwater robot is equipped with an optical camera and a sonar system, which are used to synchronously acquire optical images and acoustic data; Step 2: After preprocessing the optical image, extract the optical features in the preprocessed image ; After preprocessing the acoustic data, extract the acoustic features from the preprocessed acoustic data ; Step 3: Construct a feature fusion network for fusing optical features and acoustic features, wherein the feature fusion network includes a first CALayer module, a second CALayer module, a third CALayer module, and a first fully connected layer; Step 4: Extracting optical features Harmonic characteristics After processing, the processed optical features, processed acoustic features, and optical features Harmonic characteristics As the input of the feature fusion network, the feature fusion result is output through the feature fusion network.

2. The method for fusing underwater sonar and optical image features according to claim 1, wherein: After the optical image is preprocessed, the optical features in the preprocessed image are extracted. ; Specifically: Enhance the optical image and then use a deep neural network to extract optical features from the enhanced image .

3. The method for fusing underwater sonar and optical image features according to claim 1, wherein: After the acoustic data is preprocessed, the acoustic features in the preprocessed acoustic data are extracted. ; Specifically: The acoustic data is processed with noise reduction and signal enhancement in sequence, and then the Transformer model is used to extract acoustic features from the processed acoustic data. .

4. The method for fusing underwater sonar and optical image features according to claim 1, wherein: The extracted optical features Harmonic characteristics Processing, specifically: Optical characteristics Perform position encoding to obtain position encoding results , and then give the extracted optical features Add position encoding , and obtain the processed optical features : Acoustic characteristics Perform position encoding to obtain position encoding results , and then give the extracted acoustic features Add position encoding , get the processed acoustic features : 。 5. The method for fusing underwater sonar and optical image features according to claim 4, characterized in that: The working process of the feature fusion network is as follows: Step 1: ,Will 、 and As the input of the first CALayer module, get the output of the first CALayer module ; Step 2: ,Will 、 and As the input of the second CALayer module, get the output of the second CALayer module ; Step 3: ,Will 、 and As the input of the third CALayer module, get the output of the third CALayer module ; Step 4: Optical features and Perform splicing and use the splicing result as the input of the first fully connected layer, and output the fusion feature through the first fully connected layer : in, Represents feature concatenation operation; Linear represents a fully connected layer.

6. The method for fusing underwater sonar and optical image features according to claim 5, characterized in that: The first CALayer module includes a multi-scale deformable attention mechanism and a fully connected layer, and the working process of the first CALayer module is: Will 、 and As the input of the multi-scale deformable attention mechanism, the output of the multi-scale deformable attention mechanism is used as the input of the fully connected layer, and the output of the fully connected layer is used as the output of the first CALayer module : in, Represents the first CALayer module.

7. An underwater sonar and optical image feature fusion system, comprising an optical image acquisition module, an acoustic data acquisition module, an optical feature extraction module, an acoustic feature extraction module, an optical feature processing module, an acoustic feature processing module, and a feature fusion module; wherein: The optical image acquisition module and the acoustic data acquisition module are used to acquire optical images and acoustic data respectively; The optical feature extraction module is used to extract the features of the optical image, and the optical feature processing module is used to process the features of the optical image; The acoustic feature extraction module is used to extract features of acoustic data, and the acoustic feature processing module is used to process the features of acoustic data; The feature fusion module is used to fuse the processed optical features and the processed acoustic features to obtain a feature fusion result.

8. The underwater sonar and optical image feature fusion system according to claim 7, characterized in that: The optical feature extraction module is used to extract the features of the optical image, and the optical feature processing module is used to process the features of the optical image; the specific process is: Step S1: Enhance the optical image and then use a deep neural network to extract optical features from the enhanced image. ; Step S2: Optical characteristics Perform position encoding to obtain position encoding results , and then give the extracted optical features Add position encoding , and obtain the processed optical features : 。 9. The underwater sonar and optical image feature fusion system according to claim 8, characterized in that: The acoustic feature extraction module is used to extract the features of the acoustic data, and the acoustic feature processing module is used to process the features of the acoustic data; the specific process is as follows: Step A1: Perform noise reduction and signal enhancement on the acoustic data in sequence, and then use the Transformer model to extract acoustic features from the processed acoustic data. ; Step A2: Acoustic features Perform position encoding to obtain position encoding results , and then give the extracted acoustic features Add position encoding , get the processed acoustic features : 。 10. The underwater sonar and optical image feature fusion system according to claim 9, characterized in that: The working process of the feature fusion module is as follows: Step 1: ,Will 、 and As the input of the first CALayer module, get the output of the first CALayer module ; The working process of the first CALayer module is: Will 、 and As the input of the multi-scale deformable attention mechanism, the output of the multi-scale deformable attention mechanism is used as the input of the fully connected layer, and the output of the fully connected layer is used as the output of the first CALayer module : in, Represents the first CALayer module; Step 2: ,Will 、 and As the input of the second CALayer module, get the output of the second CALayer module ; Step 3: ,Will 、 and As the input of the third CALayer module, get the output of the third CALayer module ; Step 4: Optical features and Perform splicing and use the splicing result as the input of the first fully connected layer, and output the fusion feature through the first fully connected layer : in, Represents feature concatenation operation; Linear represents a fully connected layer.

Citation Information

Cited By

  • Underwater sonar and optical camera multi-sensor combined image calibration method

    CN121937542A