Submarine cable pipeline detection method based on underwater acousto-optic fusion
By adopting underwater acousto-optical fusion technology in the detection of submarine submarine cable pipelines, combining convolution-Transformer network and multimodal decoder, the problem of insufficient detection accuracy and robustness is solved, and a more efficient and reliable detection effect is achieved.
Patent Information
- Application Number
- CN202510578489.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
At this stage, the detection of submarine submarine cable pipelines has problems with insufficient detection accuracy and robustness, mainly due to the influence of underwater environmental factors on single modal data.
Through the detection method based on underwater acoustic and optical image data, the underwater sensor equipment is used to simultaneously collect acoustic images and optical image data, perform preprocessing and feature extraction, and feature fusion is performed using parallel double-branched convolution-Transformer network, and detection and classification is performed through multimodal decoder structure and detection regression head/classification head.
This method improves detection accuracy and robustness by fusion of data from acoustic and optical modes.
Smart Images

Figure CN120107772A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to underwater detection related fields, and in particular to a submarine cable pipeline detection method based on underwater sound and light fusion. Background Art
[0002] Underwater detection tasks, especially the detection of submarine cables and pipelines, are important technical requirements in the fields of marine engineering, underwater resource exploration, environmental monitoring, etc. However, due to the particularity of the underwater environment, such as insufficient light, turbid water, biological interference and other factors, underwater detection tasks face many challenges. In the detection of submarine cables and pipelines, traditional methods mainly rely on single acoustic or optical sensors. Acoustic sensors can penetrate turbid water and capture acoustic information on the seabed, but their accuracy in target recognition is limited; while optical sensors can provide high-resolution image data, but their performance will be greatly reduced in the case of insufficient light or turbid water. Both acoustic and optical data are severely affected by underwater environmental factors, making it difficult for single modal data to provide sufficient information for accurate detection.
[0003] Among the current related technologies, submarine cable pipeline detection has technical problems such as insufficient detection accuracy and robustness. Summary of the invention
[0004] The present application provides a submarine cable pipeline detection method based on underwater acoustic-optical fusion, uses underwater sensor equipment to simultaneously collect acoustic image and optical image data, preprocesses the collected data, extracts high-quality sonar rectangular image data and standard optical image data, and uses a parallel dual-branch convolution-Transformer network to extract features of the two modal data. Through attention perception and weighted fusion mechanism, the acoustic modal embedding feature set and the optical modal embedding feature set are fused to obtain acoustic-optical multimodal fusion-level features, and the fused features are decoded using a multimodal decoder structure to obtain acoustic-optical fusion pixel-level features. Detection and classification are performed through detection regression head and classification head and other technical means. By fusing the data of the two modalities of acoustic image and optical image, the complementary advantages of the two modalities in representing submarine cable pipelines are fully utilized, thereby achieving the technical effect of improving detection accuracy and enhancing detection robustness.
[0005] The present application provides a submarine cable pipeline detection method based on underwater acoustic-optical fusion, comprising: acquiring underwater sonar image data and underwater optical image data containing a target submarine cable pipeline through underwater sensor equipment; preprocessing and extracting the underwater sonar image data and underwater optical image data respectively to obtain sonar rectangular image data and standard optical image data; using a parallel double-branch convolution-Transformer network to extract features of the sonar rectangular image data and the standard optical image data respectively to obtain an acoustic modal embedding feature set and an optical modal embedding feature set; performing attention perception and weighted fusion on the acoustic modal embedding feature set and the optical modal embedding feature set to obtain acoustic-optical multimodal fusion level features; decoding the acoustic-optical multimodal fusion level features through a multimodal decoder structure to obtain acoustic-optical fusion pixel-level features; using a detection regression head and a classification head to detect and classify the acoustic-optical fusion pixel-level features to obtain a submarine cable pipeline detection result.
[0006] In a possible implementation, the sonar rectangular image data and the standard underwater optical image data are obtained, and the following processing is performed: a rectangular image conversion formula is constructed: ,in, They represent the input underwater sonar image data and the preprocessed sonar rectangular image data respectively. are the height and width of the image, respectively. Represents the preprocessing reconstruction algorithm set, which includes angle conversion , traverse the projection matrix ; Based on the rectangular image conversion formula, the underwater sonar image data is preprocessed by window conversion to obtain the sonar rectangular image data; and a lightweight underwater image enhancement algorithm is constructed: , ,in, represents the preprocessed standard optical image data, is an activation function for deep learning. and It is a convolution layer with a kernel size of 7×7 and a maximum pooling layer. Represents input underwater optical image data expanded along the RGB three-color space; the underwater optical image data is preprocessed based on the lightweight underwater image enhancement algorithm to obtain the standard optical image data.
[0007] In a possible implementation, the acoustic modality embedding feature set and the optical modality embedding feature set are obtained, and the following processing is performed: an acoustic code extraction formula is constructed through the acoustic encoder in the parallel dual-branch convolution-Transformer network as follows: ,in, represents the embedding vector extracted by the convolution-Transformer structure, is the sonar rectangular image data, is the corresponding result of the embedding vector after the acoustic encoder, Represents the convolutional-Transformer used for inference, the encoder contains fully connected layers , represents the activation function for deep learning, is a scalar parameter used to scale the output of softmax; based on the acoustic coding extraction formula, the sonar rectangular image data is encoded and extracted to obtain the acoustic modality embedding feature set; through the optical encoder in the parallel dual-branch convolution-Transformer network, the optical coding extraction formula is constructed as follows: ,in, , , Respectively represent the first and The feature representation matrix on the image blocks up to the Nth image block is The corresponding result generated after the optical encoder is , and , represents a weight operator for balancing multiple image blocks; encoding and extracting the standard optical image data based on the optical encoding extraction formula to obtain the optical modality embedding feature set.
[0008] In a possible implementation, the acoustic-optical multimodal fusion level feature is obtained by performing the following processing: performing attention perception on the acoustic modality embedding feature set and the optical modality embedding feature set based on a dynamic attention fusion network to obtain a multimodal feature attention map; and constructing a feature alignment formula: ,in represents the KL divergence loss function used, and its corresponding inputs are and , which is the first feature regions and optical feature embedding feature regions, N is the total number of feature regions; based on the feature alignment formula, feature alignment operation is performed on the acoustic mode embedding feature set and the optical mode embedding feature set; and a multimodal feature fusion formula is constructed: ,in, is the optical modality embedding feature set, is the acoustic modal embedding feature set, The multimodal feature attention map is used as a weight map, and the acoustic modal embedding feature set and the optical modal embedding feature set after feature alignment are fused based on the feature fusion formula to obtain the acoustic-optical multimodal fusion level feature.
[0009] In a possible implementation, the multimodal feature attention map is obtained, and the following processing is performed: According to the dynamic attention fusion network, the attention fusion network formula is constructed as follows: ,in, represents the attention map corresponding to the acoustic modality embedding feature set, represents the attention map corresponding to the optical modality embedding feature set, represents the maximum pooling layer in deep learning, represents the average pooling layer in deep learning, Indicates that the convolution kernel size is convolutional layer; performing attention perception on the acoustic modality embedding feature set and the optical modality embedding feature set based on the attention fusion network formula to obtain an acoustic embedding feature attention map and an optical embedding feature attention map; performing superposition fusion and channel confusion on the acoustic embedding feature attention map and the optical embedding feature attention map to obtain the multimodal feature attention map.
[0010] In a possible implementation, the multimodal feature attention map is obtained, and the following processing is performed: a superposition fusion formula is constructed: ,in, represents the consistent attention map of superposition fusion, represents the fusion weight of the acoustic embedding feature attention map, Represents the fusion weight of the optical embedding feature attention map, which is dynamically updated during the learning process; based on the superposition fusion formula, the acoustic embedding feature attention map and the optical embedding feature attention map are superimposed and fused to generate a consistent attention map; construct a channel confusion formula: ,in, Represents multiple groups of convolutional layers with 7 convolution kernels. is the activation function used in deep learning. ; Based on the channel confusion formula, channel confusion is performed on the consistent attention map to obtain the multimodal feature attention map.
[0011] In a possible implementation, the acquisition of the acousto-optic fusion pixel-level features performs the following processing: using the multimodal decoder structure to introduce differential convolution to obtain a feature decoding process, and the feature decoding process is specifically: ,in, is a conventional Vanilla convolution operation. Represents multiple differential convolution operations, including center differential and angle differential convolution operations; constructs a reparameterization process: ,in, Representing convolution kernels corresponding to multiple convolution operations; based on the feature decoding process and the re-parameterization process, the acoustic-optical multimodal fusion level feature is feature decoded to obtain the acoustic-optical fusion pixel-level feature.
[0012] In a possible implementation, the submarine cable pipeline detection result is obtained, and the following processing is performed: a prediction box regression function is constructed according to the detection regression head: ,in, Represents the center point b of the predicted box and the center point of the real box The Euclidean distance of represents the length of the diagonal line of the minimum circumscribed moment between two center points, Represents the intersection-over-union ratio of the predicted box and the true box, Representation parameters The corresponding impact factor is represents a parameter for measuring the consistency of aspect ratio; performing anchor frame detection on the acoustic-optical fusion pixel-level feature based on the prediction frame regression function to obtain a target feature prediction frame; constructing a target classification process according to the classification head, and the target classification process is specifically as follows: ,in, Indicates The probability that an image belongs to each category, is the fully connected layer, is a classification function; based on the target classification process, the target feature prediction frame is classified and determined to obtain the submarine cable pipeline detection result.
[0013] The submarine cable pipeline detection method based on underwater acoustic-optical fusion proposed in the present application is to first collect underwater sonar image data and underwater optical image data containing the target submarine cable pipeline through underwater sensor equipment, then pre-process and extract the underwater sonar image data and underwater optical image data respectively to obtain sonar rectangular image data and standard optical image data, and then use a parallel double-branch convolution-Transformer network to extract features of the sonar rectangular image data and the standard optical image data respectively to obtain an acoustic modal embedding feature set and an optical modal embedding feature set, and then perform attention perception and weighted fusion on the acoustic modal embedding feature set and the optical modal embedding feature set to obtain acoustic-optical multimodal fusion-level features, and then decode the acoustic-optical multimodal fusion-level features through a multimodal decoder structure to obtain acoustic-optical fusion pixel-level features, and finally use a detection regression head and a classification head to detect and classify the acoustic-optical fusion pixel-level features to obtain the submarine cable pipeline detection results, thereby achieving the technical effect of improving detection accuracy and enhancing detection robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solution of the embodiment of the present invention, the accompanying drawings of the embodiment of the present invention will be briefly introduced below. A flow chart is used in the present application to illustrate the operations performed by the method according to the embodiment of the present application. It should be understood that the preceding or following operations are not necessarily performed accurately in order. On the contrary, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or more operations can be removed from these processes.
[0015] Figure 1 A schematic flow chart of a method for detecting submarine cables and pipelines based on underwater acoustic and optical fusion provided in an embodiment of the present application.
[0016] Figure 2 A schematic diagram of the process of preprocessing and extracting underwater sonar image data and underwater optical image data in a submarine cable pipeline detection method based on underwater acoustic and optical fusion provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.
[0018] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0019] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. The terms "including" and "having" and any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by technicians in the technical field to which this application belongs. The terms used herein are for the purpose of describing the embodiments of the present application only.
[0020] The present application embodiment provides a submarine cable pipeline detection method based on underwater sound and light fusion, such as Figure 1 As shown, the method includes: Step S100, acquiring underwater sonar image data and underwater optical image data containing the target submarine cable pipeline through underwater sensor equipment.
[0021] Specifically, the underwater sensor device refers to a sensor device that can work in an underwater environment and is used to collect sonar and optical image data. The underwater sensor device integrating sonar and optical sensors is deployed on the seabed or nearby waters, and the device is started to collect sonar image data and optical image data at the same time.
[0022] Step S200, pre-processing and extracting the underwater sonar image data and the underwater optical image data respectively to obtain sonar rectangular image data and standard optical image data.
[0023] Specifically, the collected data is preprocessed using an image processing algorithm to improve the image quality. The underwater sonar image data is subjected to denoising, enhancement and other processing to obtain clear sonar rectangular image data, which is preprocessed sonar image data and has clear rectangular shapes and features. The underwater optical image data is subjected to correction, enhancement and other processing to obtain standard optical image data, which is preprocessed optical image data and meets certain standards and specifications.
[0024] like Figure 2 As shown, in a possible implementation, the sonar rectangular image data and the standard underwater optical image data are obtained, and step S200 further includes step S210, constructing a rectangular image conversion formula: ,in, They represent the input underwater sonar image data and the preprocessed sonar rectangular image data respectively. are the height and width of the image, respectively. Represents the preprocessing reconstruction algorithm set, which includes angle conversion , traverse the projection matrix Specifically, a rectangular image conversion formula is constructed using a mathematical formula and an image processing algorithm, wherein the rectangular image conversion formula is used to convert the input forward-looking sonar fan-shaped image data into a rectangular shape, and the preprocessing reconstruction algorithm is an image processing algorithm for preprocessing and reconstructing an image.
[0025] Step S220, performing window conversion preprocessing on the underwater sonar image data based on the rectangular image conversion formula to obtain the sonar rectangular image data. Specifically, the window conversion preprocessing is performed using image processing software or an algorithm. The rectangular image conversion formula constructed in step S210 is applied to the input underwater sonar image data, and the image is preprocessed by window conversion through steps such as angle conversion and traversing the projection matrix in the algorithm, and the preprocessed sonar rectangular image data is output.
[0026] Step S230, constructing a lightweight underwater image enhancement algorithm: , ,in, represents the preprocessed standard optical image data, is an activation function for deep learning. and It is a convolution layer with a kernel size of 7×7 and a maximum pooling layer. Represents the input underwater optical image data expanded along the RGB three-color space. Specifically, a lightweight underwater image enhancement algorithm is constructed using deep learning algorithms and image processing technology. The lightweight underwater image enhancement algorithm is a lightweight image processing algorithm used to improve the quality of underwater optical images. The convolution layer with a convolution kernel size of 7×7 and the maximum pooling layer are used to extract image features and perform dimensionality reduction. The input underwater optical image data is expanded along the RGB three-color space to enable the algorithm to process color images and reduce the color cast caused by the refraction and scattering of light underwater.
[0027] Step S240, preprocessing the underwater optical image data based on the lightweight underwater image enhancement algorithm to obtain the standard optical image data. Specifically, the lightweight underwater image enhancement algorithm is preprocessed using image processing software or algorithm. The lightweight underwater image enhancement algorithm constructed in step S230 is applied to the input underwater optical image data, and the image is preprocessed through steps such as the convolution layer, the maximum pooling layer and the activation function in the algorithm, and the preprocessed standard optical image data is output. Underwater images are often affected by factors such as illumination and turbidity, resulting in poor image quality. This implementation method preprocesses sonar images and optical images by constructing a rectangular image conversion formula and a lightweight underwater image enhancement algorithm, thereby improving image quality, making them more suitable for subsequent feature extraction and detection classification, and improving the detection accuracy and efficiency of submarine cables and pipelines.
[0028] Step S300, using a parallel dual-branch convolution-Transformer network to perform feature extraction on the sonar rectangular image data and the standard optical image data respectively, to obtain an acoustic modality embedding feature set and an optical modality embedding feature set.
[0029] Specifically, a parallel dual-branch convolution-Transformer network is used for feature extraction. The parallel dual-branch network refers to a network structure with two parallel branches, each branch processing different types of image data. Convolutional neural network (CNN) is a deep learning network structure used for image feature extraction; Transformer network is a deep learning network based on self-attention mechanism, which can process sequence data and extract global features. Sonar rectangular image data and standard optical image data are input into the parallel dual-branch network respectively, and the local features of the image are extracted using the convolutional neural network (CNN), and the global features and contextual information of the image are extracted using the Transformer network. The extracted features are combined into acoustic modality embedding feature sets and optical modality embedding feature sets.
[0030] In a possible implementation, the step of obtaining the acoustic modality embedding feature set and the optical modality embedding feature set, step S300 further includes step S310, constructing an acoustic coding extraction formula through the acoustic encoder in the parallel dual-branch convolution-Transformer network as follows: ,in, represents the embedding vector extracted by the convolution-Transformer structure, is the sonar rectangular image data, is the corresponding result of the embedding vector after the acoustic encoder, Represents the convolutional-Transformer used for inference, the encoder contains fully connected layers , represents the activation function for deep learning, is a scalar parameter used to scale the output of softmax. Specifically, the acoustic encoder contains a fully connected layer and an activation function in deep learning to increase the nonlinear expression ability of the network. Used to stabilize the training process.
[0031] Step S320, encoding and extracting the sonar rectangular image data based on the acoustic coding extraction formula to obtain the acoustic modal embedding feature set. Specifically, the sonar rectangular image data is used as input and sent to the acoustic encoder, and image features are extracted through convolution operations and Transformer structures according to the acoustic coding extraction formula, and the acoustic modal embedding feature set is output as input for subsequent steps.
[0032] Step S330, constructing an optical code extraction formula through the optical encoder in the parallel dual-branch convolution-Transformer network as follows: ,in, , , Respectively represent the first and The feature representation matrix on the image blocks up to the Nth image block is The corresponding result generated after the optical encoder is , and , Represents a weight operator, which is used to balance multiple image blocks. Specifically, the optical encoder infers the information of image feature changes in different area blocks, divides the input image (i.e., standard optical image data) into multiple image blocks, and extracts features from each image block separately. The weight operator is used to balance the features of multiple image blocks to ensure the comprehensiveness and accuracy of feature extraction.
[0033] Step S340, encode and extract the standard optical image data based on the optical coding extraction formula to obtain the optical modality embedding feature set. Specifically, the standard optical image data is used as input and sent to the optical encoder, the image is divided into multiple image blocks according to the optical coding extraction formula, and feature extraction is performed on each image block respectively, and the features of each image block are balanced using a weight operator to obtain the optical modality embedding feature set. This implementation method can process sonar images and optical images simultaneously through a parallel dual-branch structure, thereby improving detection efficiency. The convolution-Transformer structure combines the advantages of CNN and Transformer, which can not only process local features but also capture global dependencies, which enables the network to adapt to seabed environments and submarine cable pipeline morphologies of different complexities, thereby improving detection accuracy and robustness.
[0034] Step S400, performing attention perception and weighted fusion on the acoustic modality embedding feature set and the optical modality embedding feature set to obtain acoustic-optical multimodal fusion level features.
[0035] Specifically, an attention mechanism is used for feature fusion, which can simulate human attention to information and improve the accuracy of feature extraction and fusion. Attention is perceived on the acoustic modality embedding feature set and the optical modality embedding feature set, the weight of each feature is calculated, and the features are weighted fused according to the weight to obtain the acoustic-optical multimodal fusion level features.
[0036] In a possible implementation, the step of obtaining the acoustic-optical multimodal fusion-level features, step S400 further includes step S410, performing attention perception on the acoustic modal embedding feature set and the optical modal embedding feature set based on a dynamic attention fusion network, and obtaining a multimodal feature attention map. Specifically, the dynamic attention fusion network is a network structure that can dynamically assign different weights according to the importance of input features, and is used to help the model pay attention to more valuable information in different modalities. Based on the dynamic attention fusion network, the acoustic modal embedding feature set and the optical modal embedding feature set are subjected to attention perception to generate a multimodal feature attention map, which reflects the importance of different feature areas in different modalities.
[0037] Step S420, constructing a feature alignment formula: ,in represents the KL divergence loss function used, and its corresponding inputs are and , that is, the acoustic feature embedding feature regions and optical feature embedding feature regions, N is the total number of feature regions. Specifically, a feature alignment formula is constructed, and the KL divergence loss function is used to minimize the difference between acoustic feature embedding and optical feature embedding, so that they are aligned in space and channels, ensuring that features of different modalities can be compared in the same feature space.
[0038] Step S430, performing feature alignment operation on the acoustic modal embedding feature set and the optical modal embedding feature set based on the feature alignment formula. Specifically, performing feature alignment operation on the acoustic modal embedding feature set and the optical modal embedding feature set based on the feature alignment formula, by optimizing the KL divergence loss function, the difference between the acoustic feature embedding and the optical feature embedding is minimized, and comparison and fusion can be performed in the same feature space.
[0039] Step S440, constructing a multimodal feature fusion formula: ,in, is the optical modality embedding feature set, is the acoustic modal embedding feature set, Specifically, the multimodal feature fusion formula is used to fuse features from different modalities according to certain rules or weights to generate a comprehensive feature representation, and the fusion is performed based on the weight map.
[0040] Step S450, using the multimodal feature attention map as a weight map, based on the feature fusion formula, the acoustic modal embedding feature set and the optical modal embedding feature set after feature alignment are fused to obtain the acoustic-optical multimodal fusion-level feature. Specifically, using the multimodal feature attention map as a weight map, based on the multimodal feature fusion formula, the acoustic modal embedding feature set and the optical modal embedding feature set after feature alignment are fused to generate a comprehensive acoustic-optical multimodal fusion-level feature, which contains both acoustic modal information and optical modal information, and the information is weighted according to their importance. This implementation method enables the model to capture more details and features by fusing information from different modalities, thereby improving the accuracy of submarine cable pipeline detection.
[0041] In a possible implementation, the step of obtaining a multimodal feature attention map, step S410 further includes step S411, constructing an attention fusion network formula according to the dynamic attention fusion network as follows: ,in, represents the attention map corresponding to the acoustic modality embedding feature set, represents the attention map corresponding to the optical modality embedding feature set, represents the maximum pooling layer in deep learning, represents the average pooling layer in deep learning, Indicates that the convolution kernel size is Specifically, the attention fusion network formula captures the correlation and importance between the two modal features by calculating the attention map of the acoustic modality embedding feature set and the optical modality embedding feature set. The maximum pooling layer and the average pooling layer are used to extract the global information of features from different angles, and the convolution layer is used to adjust the feature dimension and fuse the features from the maximum pooling layer and the average pooling layer.
[0042] Step S412, based on the attention fusion network formula, the acoustic modality embedding feature set and the optical modality embedding feature set are subjected to attention perception to obtain an acoustic embedding feature attention map and an optical embedding feature attention map. Specifically, the acoustic modality embedding feature set and the optical modality embedding feature set are respectively input into the attention fusion network formula, global information is extracted through the maximum pooling layer and the average pooling layer, the convolution layer is used to fuse the features from the two pooling layers, and an acoustic embedding feature attention map and an optical embedding feature attention map are generated.
[0043] Step S413, superimposes, fuses and channels the acoustic embedded feature attention map and the optical embedded feature attention map to obtain the multimodal feature attention map. Specifically, the acoustic embedded feature attention map and the optical embedded feature attention map are superimposed in the channel dimension, and the channel confusion operation is performed on the superimposed feature map, that is, the channel order of the feature map is disrupted to further enhance the interaction and fusion effect of the features, and finally, a multimodal feature attention map is obtained, which reflects the correlation and importance between the acoustic modal and optical modal features. This implementation method can capture the correlation and importance between the two modal features by calculating the attention map between the acoustic modal and optical modal features, thereby effectively utilizing multimodal information to improve detection accuracy.
[0044] In a possible implementation, the step of obtaining the multimodal feature attention map further includes step S4131, constructing a superposition fusion formula: ,in, represents the consistent attention map of superposition fusion, represents the fusion weight of the acoustic embedding feature attention map, Represents the fusion weight of the optical embedding feature attention map, which is dynamically updated during the learning process. Specifically, the fusion weight is dynamically updated during the learning process to balance the contribution of acoustic modality and optical modality features in the superposition fusion process.
[0045] Step S4132, based on the superposition fusion formula, the acoustic embedding feature attention map and the optical embedding feature attention map are superimposed and fused to generate a consistent attention map. Specifically, the acoustic embedding feature attention map and the optical embedding feature attention map are used as inputs, substituted into the superposition fusion formula, and the consistent attention map after superposition fusion is calculated according to the dynamically updated fusion weight.
[0046] Step S4133, construct a channel confusion formula: ,in, Represents multiple groups of convolutional layers with 7 convolution kernels. is the activation function used in deep learning. Specifically, the channel confusion formula is used to enhance the interaction and fusion of features, further explore more valuable multimodal information in the consistent attention map, and avoid noise interference caused by modal fusion. Through multiple sets of convolution-activation operations, the channels of the consistent attention map are shuffled and recombined. After each convolution operation, the activation function is used to increase nonlinearity. In the formula, Indicates channel splicing.
[0047] Step S4134, channel confusion is performed on the consistent attention map based on the channel confusion formula to obtain the multimodal feature attention map. Specifically, the consistent attention map is used as input, substituted into the channel confusion formula, and the channels of the input features are shuffled and recombined through multiple sets of convolution-activation operations until the final multimodal feature attention map is obtained. This implementation method obtains a more accurate and comprehensive multimodal feature attention map through superposition fusion and channel confusion, thereby improving the accuracy and robustness of submarine cable pipeline detection.
[0048] Step S500: decoding the acoustic-optical multimodal fusion level feature through a multimodal decoder structure to obtain an acoustic-optical fusion pixel-level feature.
[0049] Specifically, a multimodal decoder structure is used for decoding, and the multimodal decoder structure refers to a network structure that can process multiple modal features and perform decoding. The acoustic-optical multimodal fusion level features are input into the multimodal decoder structure, and the decoder is used to decode the features to obtain the acoustic-optical fusion pixel-level features, that is, the features that can reflect the information of each pixel in the image.
[0050] In a possible implementation, the step of obtaining the acousto-optic fusion pixel-level features, step S500 further includes step S510, using the multimodal decoder structure to introduce differential convolution to obtain a feature decoding process, wherein the feature decoding process is specifically as follows: ,in, is a conventional Vanilla convolution operation. Represents multi-difference convolution operation, including center difference and angle difference convolution operations. Specifically, the conventional Vanilla convolution operation is a standard convolution operation used to extract local features in feature maps. Differential convolution is a special convolution operation that captures finer spatial structure information by calculating the differences between adjacent pixels (or pixels in a specific direction) in the input feature map. Multi-difference convolution operations include center difference and angle difference convolution operations. The center difference convolution calculates the difference between the center pixel and its neighboring pixels, and the angle difference convolution focuses on capturing the features of the edges or corners of the image. Combining the conventional Vanilla convolution operation with the multi-difference convolution operation forms a feature decoding process that can more comprehensively capture and process multi-scale and multi-directional information in the image.
[0051] Step S520, constructing a re-parameterization process: ,in, Represents the convolution kernels corresponding to various convolution operations. Specifically, by constructing a reparameterization process, various convolution operations (including conventional convolution and differential convolution) are integrated into a unified framework. The reparameterization technology allows the use of multiple convolution kernels for feature extraction during training, and these convolution kernels are merged into one during inference to reduce the amount of computation and memory consumption, thereby simplifying the model structure without sacrificing performance. The convolution kernels corresponding to the various convolution operations include kernels for conventional convolution and kernels for differential convolution.
[0052] Step S530, based on the feature decoding process and the reparameterization process, the acoustic-optical multimodal fusion level feature is feature decoded to obtain the acoustic-optical fusion pixel-level feature. Specifically, in combination with the feature decoding process in step S510 and the reparameterization process in step S520, the acoustic-optical multimodal fusion level feature is decoded to obtain the acoustic-optical fusion pixel-level feature. This implementation method can capture the spatial structure information in the image more finely by introducing differential convolution, thereby enhancing the feature extraction capability. The application of reparameterization technology allows the use of multiple convolution kernels for feature extraction during training, and simplifies the model structure during inference, reduces the amount of calculation and memory consumption, and optimizes the calculation efficiency.
[0053] Step S600: Use a detection regression head and a classification head to detect and classify the acoustic-optical fusion pixel-level features to obtain a submarine cable pipeline detection result.
[0054] Specifically, detection and classification are performed using a detection regression head and a classification head. The detection regression head refers to the network part used for position regression, which can determine the position of the target in the image. The classification head refers to the network part used for classification, which can determine the type or state of the target. The acoustic-optical fusion pixel-level features are input into the detection regression head and the classification head, and the detection regression head is used to perform position regression on the features to determine the position of the submarine cable pipeline. The classification head is used to classify the features to determine the type or state of the submarine cable pipeline, and the detection results are output, including information such as the position, type and state of the submarine cable pipeline. The embodiment of the present application utilizes underwater sensor equipment to simultaneously collect acoustic image and optical image data, preprocesses the collected data, extracts high-quality sonar rectangular image data and standard optical image data, and uses a parallel dual-branch convolution-Transformer network to extract features of the two modal data. Through attention perception and weighted fusion mechanism, the acoustic modal embedding feature set and the optical modal embedding feature set are fused to obtain acoustic and optical multimodal fusion-level features, and the fused features are decoded using a multimodal decoder structure to obtain acoustic and optical fusion pixel-level features. Detection and classification are performed through detection regression head and classification head, and other technical means. By fusing the data of the two modalities of acoustic image and optical image, the complementary advantages of the two modalities in representing submarine cable pipelines are fully utilized, thereby achieving the technical effect of improving detection accuracy and enhancing detection robustness.
[0055] In a possible implementation, the step S600 of obtaining the submarine cable pipeline detection result further includes a step S610 of constructing a prediction box regression function according to the detection regression head: ,in, Represents the center point b of the predicted box and the center point of the real box The Euclidean distance of represents the length of the diagonal line of the minimum circumscribed moment between two center points, Represents the intersection-over-union ratio of the predicted box and the true box, Representation parameters The corresponding impact factor is Represents a parameter that measures the consistency of aspect ratio. Specifically, the prediction box regression function is used to calculate the difference between the prediction box (i.e., the target position box predicted by the model) and the true box (i.e., the actual target position box) and minimize this difference. The Euclidean distance between the center point of the prediction box and the center point of the true box reflects the deviation in position. The length of the minimum circumscribed moment diagonal measures the deviation in scale between the prediction box and the true box. The ratio of the intersection area of the prediction box and the true box to the union area reflects the degree of shape matching between the two.
[0056] Step S620, based on the prediction frame regression function, anchor frame detection is performed on the acoustic-optical fusion pixel-level features to obtain the target feature prediction frame. Specifically, anchor frame detection is performed on the acoustic-optical fusion pixel-level features using the constructed prediction frame regression function. Anchor frames are a series of pre-set rectangular frames of different scales and aspect ratios, which are used as the starting point for target detection. By calculating the difference between each anchor frame and the true frame, and applying the prediction frame regression function for optimization, the target feature prediction frame is finally obtained.
[0057] Step S630: construct a target classification process according to the classification header, and the target classification process is specifically as follows: ,in, Indicates The probability that an image belongs to each category, is the fully connected layer, is a classification function. Specifically, the target classification process is used to classify the target feature prediction box. The process includes a fully connected layer for extracting features and a classification function for mapping features to category probabilities. The target feature prediction box is classified by calculating the probability that each image belongs to each category.
[0058] Step S640, classify and determine the target feature prediction frame based on the target classification process to obtain the submarine cable pipeline detection result. Specifically, the target feature prediction frame is classified and determined using the constructed target classification process, and the features of the target feature prediction frame are input into the fully connected layer. After the features are extracted, the probabilities of each category are calculated by the classification function. Then, according to the set threshold or maximum probability principle, the category to which the target feature prediction frame belongs is determined, thereby obtaining the detection result of the submarine cable pipeline. This implementation method optimizes the position and scale of the prediction frame by constructing a prediction frame regression function, so that it more accurately approaches the real frame; by constructing a target classification process, the prediction frame is classified and determined to determine the category to which it belongs. This combination method not only ensures the accuracy of detection, but also improves the robustness of detection.
[0059] The above specific implementation manner does not constitute a limitation to the protection scope of the present application. It should be understood by those skilled in the art that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application. In some cases, the actions or steps recorded in the present application can be performed in an order different from that in the embodiment and can still achieve the desired results. In addition, the process depicted in the accompanying drawings does not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A submarine cable pipeline detection method based on underwater sound and light fusion, characterized in that: The method comprises: Acquire underwater sonar image data and underwater optical image data including the target submarine cable pipeline through underwater sensor equipment; Preprocessing and extracting the underwater sonar image data and the underwater optical image data respectively to obtain sonar rectangular image data and standard optical image data; Using a parallel dual-branch convolution-Transformer network to extract features from the sonar rectangular image data and the standard optical image data, respectively, to obtain an acoustic modal embedding feature set and an optical modal embedding feature set; Performing attention perception and weighted fusion on the acoustic modality embedding feature set and the optical modality embedding feature set to obtain acoustic-optical multimodal fusion-level features; Decoding the acoustic-optical multimodal fusion level features through a multimodal decoder structure to obtain acoustic-optical fusion pixel-level features; The detection regression head and the classification head are used to detect and classify the acoustic-optical fusion pixel-level features to obtain the submarine cable pipeline detection results.
2. The method for detecting submarine cables and pipelines based on underwater sound and light fusion according to claim 1, characterized in that: The obtaining of sonar rectangular image data and standard underwater optical image data comprises: Construct the rectangular image transformation formula: ,in, They represent the input underwater sonar image data and the preprocessed sonar rectangular image data respectively. are the height and width of the image, respectively. Represents the preprocessing reconstruction algorithm set, which includes angle conversion , traverse the projection matrix ; Performing window conversion preprocessing on the underwater sonar image data based on the rectangular image conversion formula to obtain the sonar rectangular image data; Building a lightweight underwater image enhancement algorithm: , ,in, represents the preprocessed standard optical image data, is an activation function for deep learning. and It is a convolution layer with a kernel size of 7×7 and a maximum pooling layer. represents the input underwater optical image data expanded along the RGB three-color space; The underwater optical image data is preprocessed based on the lightweight underwater image enhancement algorithm to obtain the standard optical image data.
3. The method for detecting submarine cables and pipelines based on underwater sound and light fusion according to claim 2, characterized in that: The step of obtaining an acoustic mode embedding feature set and an optical mode embedding feature set comprises: Through the acoustic encoder in the parallel dual-branch convolution-Transformer network, the acoustic coding extraction formula is constructed as follows: ,in, represents the embedding vector extracted by the convolution-Transformer structure, is the sonar rectangular image data, is the corresponding result of the embedding vector after the acoustic encoder, Represents the convolutional-Transformer used for inference, the encoder contains fully connected layers , represents the activation function for deep learning, is a scalar parameter used to scale the output of softmax; Based on the acoustic coding extraction formula, encoding and extracting the sonar rectangular image data to obtain the acoustic modal embedding feature set; Through the optical encoder in the parallel dual-branch convolution-Transformer network, the optical encoding extraction formula is constructed as follows: ,in, , , Respectively represent the first and The feature representation matrix on the image blocks up to the Nth image block is The corresponding result generated after the optical encoder is , and , Represents a weight operator, which is used to balance multiple image blocks; The standard optical image data is encoded and extracted based on the optical encoding extraction formula to obtain the optical modality embedding feature set.
4. The method for detecting submarine cables and pipelines based on underwater sound and light fusion according to claim 3 is characterized in that: The step of obtaining the acoustic-optical multimodal fusion level features comprises: Performing attention perception on the acoustic modality embedding feature set and the optical modality embedding feature set based on a dynamic attention fusion network to obtain a multimodal feature attention map; Construct feature alignment formula: ,in represents the KL divergence loss function used, and its corresponding inputs are and , which is the first feature regions and optical feature embedding feature regions, N is the total number of feature regions; Performing a feature alignment operation on the acoustic mode embedding feature set and the optical mode embedding feature set based on the feature alignment formula; Construct multimodal feature fusion formula: ,in, is the optical modality embedding feature set, is the acoustic modal embedding feature set, is the weight graph; The multimodal feature attention map is used as a weight map, and the acoustic modal embedding feature set and the optical modal embedding feature set after feature alignment are fused based on the feature fusion formula to obtain the acoustic-optical multimodal fusion level feature.
5. The method for detecting submarine cables and pipelines based on underwater sound and light fusion according to claim 4, characterized in that: The obtaining of a multimodal feature attention map comprises: According to the dynamic attention fusion network, the attention fusion network formula is constructed as follows: , ,in, represents the attention map corresponding to the acoustic modality embedding feature set, represents the attention map corresponding to the optical modality embedding feature set, represents the maximum pooling layer in deep learning, represents the average pooling layer in deep learning, Indicates that the convolution kernel size is The convolutional layer; Based on the attention fusion network formula, the acoustic modality embedding feature set and the optical modality embedding feature set are subjected to attention perception to obtain an acoustic embedding feature attention map and an optical embedding feature attention map; The acoustic embedding feature attention map and the optical embedding feature attention map are superimposed, fused and channel-confused to obtain the multimodal feature attention map.
6. The method for detecting submarine cables and pipelines based on underwater sound and light fusion according to claim 5, characterized in that: The obtaining of the multimodal feature attention map comprises: Construct the superposition fusion formula: ,in, represents the consistent attention map of superposition fusion, represents the fusion weight of the acoustic embedding feature attention map, represents the fusion weight of the optical embedding feature attention map, which is dynamically updated during the learning process; Based on the superposition fusion formula, the acoustic embedding feature attention map and the optical embedding feature attention map are superimposed and fused to generate a consistent attention map; Construct channel confusion formula: ,in, Represents multiple groups of convolutional layers with 7 convolution kernels. is the activation function used in deep learning. ; Channel confusion is performed on the consistent attention map based on the channel confusion formula to obtain the multimodal feature attention map.
7. The method for detecting submarine cables and pipelines based on underwater sound and light fusion according to claim 6, characterized in that: The step of obtaining the acousto-optic fusion pixel-level features comprises: The multimodal decoder structure is used to introduce differential convolution to obtain a feature decoding process, which is specifically as follows: ,in, is a conventional Vanilla convolution operation. Represents multiple differential convolution operations, including central differential and angular differential convolution operations; Build a re-parameterized process: ,in, Represents the convolution kernels corresponding to various convolution operations; Based on the feature decoding process and the re-parameterization process, the acoustic-optical multimodal fusion level feature is feature decoded to obtain the acoustic-optical fusion pixel-level feature.
8. The method for detecting submarine cables and pipelines based on underwater sound and light fusion according to claim 7, characterized in that: The obtaining of the submarine cable pipeline detection result comprises: According to the detection regression head, a prediction box regression function is constructed: ,in, Represents the center point b of the predicted box and the center point of the real box The Euclidean distance of represents the length of the diagonal line of the minimum circumscribed moment between two center points, Represents the intersection-over-union ratio of the predicted box and the true box, Representation parameters The corresponding impact factor is It represents the parameter to measure the consistency of aspect ratio; Perform anchor frame detection on the acoustic-optical fusion pixel-level feature based on the prediction frame regression function to obtain a target feature prediction frame; According to the classification head, a target classification process is constructed, and the target classification process is specifically as follows: ,in, Indicates The probability that an image belongs to each category, is the fully connected layer, is the classification function; The target feature prediction frame is classified and determined based on the target classification process to obtain the submarine cable pipeline detection result.
Citation Information
Patent Citations
Area pre-detection-based underwater suspended sonar target identification method
CN115810144A
RGB-D cross-modal interactive fusion mechanical arm grabbing detection method based on Transform-CNN hybrid architecture
CN116912608A
Detection method and device based on inherent attribute characteristics of underwater target
CN117809168A
Dual-mode underwater dam crack detection method based on acousto-optic image fusion
CN118154993A
Automatic substrate glass surface defect detection method and system based on machine vision
CN119006469A
Cited By
Multi-modal underwater target detection method and system based on layered feature alignment
CN121477213A
Multimodal underwater target detection method and system based on hierarchical feature alignment
CN121477213B
Underwater image target detection method and system based on multi-modal fusion
CN121482380A
A Method and System for Underwater Image Target Detection Based on Multimodal Fusion
CN121482380B
Deep sea polymetallic nodule image segmentation method based on multi-modal data fusion
CN121937472A