Underwater pipeline fault detection method and device

Through multimodal fusion technology, visual, acoustic and laser data are used to train detection models, the problem of limited detection capabilities of traditional underwater pipeline detection technology is solved, and the fault detection effect with high accuracy and robustness is achieved.

CN120212445AActive Publication Date: 2025-06-27DONGHAI LAB

Patent Information

Application Number
CN202510685069.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-27
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Traditional underwater pipeline detection technology relies on a single sensor, has limited detection capabilities, cannot fully reflect the health status of the pipeline, and is prone to missing fault conditions.

Method used

The multimodal fusion method is adopted to train and detect the model through visual images, acoustic images and laser point cloud data, and use the first feature extraction module, multimodal fusion module, second feature extraction module and fault classifier to realize multimodal fusion and fault recognition of data.

Benefits of technology

It improves the accuracy and robustness of underwater pipeline fault detection, can accurately identify different types of pipeline faults in a variety of underwater environments, and enhances detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120212445A_ABST
    Figure CN120212445A_ABST
Patent Text Reader

Abstract

The invention provides an underwater pipeline fault detection method and device. The underwater pipeline fault detection method provided by the invention comprises the following steps: firstly, for each position in a plurality of positions selected from an underwater pipeline in advance, acquiring a group of data of the position to obtain a plurality of groups of data; training a detection model by using multiple groups of data; and finally, inputting a group of to-be-detected data to be detected into the trained detection model to obtain a detection result. According to the underwater pipeline fault detection method and device provided by the invention, information of multi-mode sensors such as a camera, a sonar and a laser radar is fused, learnable feature prompts are added into embedded representation of each mode, a model is guided to better utilize information of each mode, a multi-level structure based on a residual low-rank adapter is adopted, and the fault detection accuracy is improved. The multi-modal information is injected into the adapter through the cross attention mechanism, so that the pre-trained detection model can better utilize the information from the multi-modal sensor, and more accurate underwater pipeline fault detection is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of fault detection, and particularly to an underwater pipeline fault detection method and device. Background Art

[0002] Underwater pipelines are important infrastructure for energy transportation and are widely used in the transportation of energy such as oil, natural gas, and chemicals. In the underwater environment, pipelines often face complex geographical conditions and harsh environmental factors, such as water flow scouring, high pressure at deep water, and corrosion caused by seawater. These factors increase the difficulty and risk of pipeline detection.

[0003] Traditional underwater pipeline detection technologies mainly rely on single sensor devices, such as underwater robots, remotely operated vehicles (ROVs), sonar detectors, etc. However, the detection capabilities of single sensor devices are limited and cannot comprehensively reflect the health status of pipelines, easily missing pipeline fault conditions. Summary of the Invention

[0004] In view of this, this application provides an underwater pipeline fault detection method and device to achieve accurate and efficient underwater pipeline detection.

[0005] Specifically, this application is implemented through the following technical solutions:

[0006] The first aspect of this application provides an underwater pipeline fault detection method, and the method includes:

[0007] For each of a plurality of positions pre-selected from an underwater pipeline, obtain a set of data at that position, obtaining multiple sets of data; a set of data includes visual images, acoustic images, and laser point clouds;

[0008] Train a detection model using the multiple sets of data; the detection model includes a first feature extraction module, a multi-modal fusion module, a second feature extraction module, and a fault classifier; the first feature extraction module is used to extract the initial features of each modality, and splice the learnable prompt features corresponding to the modality on the initial features of each modality to form spliced visual features, spliced acoustic features, and spliced point cloud features; the multi-modal fusion module is used to fuse the spliced features to obtain the final fused features; the second feature extraction module includes multiple encoding modules, each encoding module includes a plurality of cascaded encoding layers, a low-rank adapter connected to the output end of the second-to-last encoding layer of the multiple encoding layers, and a weighted fusion layer connected to the output ends of the last encoding layer of the multiple encoding layers and the output end of the low-rank adapter at the same time; the first encoding layer of the multiple encoding layers of the first encoding module is used to receive the visual spliced features, and the input end of the low-rank adapter of the first encoding module is connected to the output end of the multi-modal fusion module; the low-rank adapters of two adjacent encoding modules are connected, and the weighted fusion layer of the previous encoding module among the two adjacent encoding modules is connected to the first encoding layer of the latter encoding module;

[0009] Input a set of data to be detected into the trained detection model to obtain a detection result.

[0010] The second aspect of this application provides an underwater pipeline fault detection device, the device includes an acquisition module, a training module, and a detection module; wherein, the acquisition module is used to obtain a set of data for each of multiple positions pre-selected from the underwater pipeline to obtain multiple sets of data; a set of data includes visual images, acoustic images, and laser point clouds;

[0011] The training module is used to train a detection model using the multiple sets of data; the detection model includes a first feature extraction module, a multi-modal fusion module, a second feature extraction module, and a fault classifier; the first feature extraction module is used to extract the initial features of each modality, and splice the learnable prompt features corresponding to the modality on the initial features of each modality to form spliced visual features, spliced acoustic features, and spliced point cloud features; the multi-modal fusion module is used to fuse the spliced features to obtain the final fused features; the second feature extraction module includes multiple encoding modules, each encoding module includes a plurality of cascaded encoding layers, a low-rank adapter connected to the output end of the second-to-last encoding layer of the multiple encoding layers, and a weighted fusion layer connected to the output ends of the last encoding layer of the multiple encoding layers and the output end of the low-rank adapter at the same time; the first encoding layer of the multiple encoding layers of the first encoding module is used to receive the visual spliced features, and the input end of the low-rank adapter of the first encoding module is connected to the output end of the multi-modal fusion module; the low-rank adapters of two adjacent encoding modules are connected, and the weighted fusion layer of the previous encoding module among the two adjacent encoding modules is connected to the first encoding layer of the latter encoding module;

[0012] The detection module is used to input a set of data to be detected into the trained detection model to obtain a detection result.

[0013] For the underwater pipeline fault detection method and device provided in this application, the first feature extraction module respectively extracts features from visual images, acoustic images and laser point clouds, and splices learnable prompt features on the basis of initial features, so that data of different modalities have better expression ability before fusion. The multi-modal fusion module adopts a learnable fusion method to make the information of visual, acoustic and point cloud data complementary, effectively enhancing the extraction ability of key fault features and overcoming the limitations that may be caused by the environment in a single modality. Further, the second feature adopts a hierarchical coding structure based on a low-rank adapter, allowing information to be gradually optimized in coding layers of different depths, making the semantics of the features more complete and enhancing the feature expression ability. In this way, the three work together, enabling the detection model to accurately identify different types of pipeline faults under various underwater environments (such as insufficient light, noise interference, etc.), and improving the detection accuracy.

[0014] In addition, the parameters of the coding layer and the weighting layer are fixed to ensure the stability of feature extraction, while the parameters of the low-rank adapter and the multi-modal fusion module are learnable, enabling the model to perform adaptive optimization in different environments, improving the generalization ability, enhancing the detection accuracy and robustness, and reducing the training parameters at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flowchart of the first embodiment of the underwater pipeline fault detection method provided in this application;

[0016] Figure 2 It is a schematic structural diagram of the detection model provided by an exemplary embodiment of this application;

[0017] Figure 3 It is a schematic structural diagram of the first feature extraction module shown by an exemplary embodiment of this application;

[0018] Figure 4 It is a schematic internal structure diagram of the low-rank adapter shown by an exemplary embodiment of this application;

[0019] Figure 5 It is a schematic diagram of the first embodiment of the underwater pipeline fault detection device provided in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.

[0021] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0022] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0023] Specific embodiments are given below to introduce the technical solutions of this application in detail.

[0024] Figure 1 This is a flowchart of the first embodiment of the underwater pipeline fault detection method provided for this application. Figure 2 This is a structural diagram of the detection model shown in an exemplary embodiment of this application. Please refer to Figure 1 and Figure 2 simultaneously. The method provided in this embodiment may include:

[0025] S101. For each of a plurality of positions pre-selected from an underwater pipeline, obtain a set of data at this position, obtaining multiple sets of data; a set of data includes visual images, acoustic images, and laser point clouds.

[0026] Specifically, the underwater pipeline fault detection method and device provided in this application can perform fault detection on underwater pipelines for transporting natural energy. For example, it can perform detection of transportation safety in underwater transportation pipelines for transporting natural energy such as oil or natural gas.

[0027] It should be noted that underwater pipelines for transporting natural energy usually extend for several kilometers. The underwater pipelines at different locations may have significant differences in the types and severities of faults at various parts of the underwater pipeline due to factors such as environmental pressure, corrosion degree, mechanical stress, or geological activities. Therefore, during the detection process, multiple locations of the underwater pipeline can be selected for separate detections. In this way, it can ensure effective coverage of high-risk areas at the underwater pipeline and avoid missing overall hidden dangers inside the underwater pipeline due to local detections.

[0028] It should be noted that the multiple locations pre-selected from the underwater pipeline are selected according to actual needs, and this application does not limit them. For example, in one embodiment, when the underwater pipeline transports natural energy such as oil or natural gas, leaks are likely to occur at the weld interface positions of the underwater pipeline. Therefore, the weld interfaces at various parts of the underwater pipeline can be selected as the multiple locations for fault detection.

[0029] Furthermore, after selecting multiple locations where fault detection needs to be performed on the underwater pipeline, for each location among the multiple locations, obtain the visual image, acoustic image, and laser point cloud at that location as a set of data, and obtain multiple sets of data corresponding to the multiple locations.

[0030] It should be noted that in specific implementation, appropriate data acquisition tools can be selected according to actual needs to collect the visual image, acoustic image, and laser point cloud of each location. This application does not limit them. For example, in one embodiment, a pressure-resistant underwater optical camera is used to collect the visual image of the underwater pipeline, an ultrasonic sensor is used to collect the acoustic image of the underwater pipeline, and a lidar sensor is used to collect the laser point cloud of the underwater pipeline.

[0031] Optionally, in a possible implementation manner, after collecting multiple sets of data, each set of data among the multiple sets of data can also be preprocessed to remove the noise therein.

[0032] It should be noted that the visual image, acoustic image, and laser point cloud in each set of data are all collected by sensor devices. The data collected by sensors underwater are often easily affected by various environmental interferences, resulting in a large amount of noise in the collected original data. Therefore, the collected original data can be preprocessed to improve the quality of the visual image, acoustic image, and laser point cloud and ensure the accuracy of subsequent detections.

[0033] Optionally, in a possible implementation manner, each set of data can be preprocessed by the following method:

[0034] For the visual image, determine the denoising algorithm and the parameters of the denoising algorithm according to the characteristics of the visual image, and perform denoising processing on the visual image based on the determined denoising algorithm and the parameters of the denoising algorithm;

[0035] For acoustic images, common features are determined based on the features of visual images and acoustic images, and the acoustic images are denoised based on the common features.

[0036] For point cloud data, isolated points and discrete points are removed based on semantic similarity and semantic coherence.

[0037] Specifically, when denoising visual images, the denoising algorithm and the parameters of the denoising algorithm can be determined according to features such as the local contrast of the visual image (the local contrast of the visual image is represented by LC) and the noise standard deviation (the noise standard deviation of the visual image is represented by σ).

[0038] For example, in a possible implementation, when the image features are low contrast (the local contrast LC of the low-contrast image feature < 0.3) and high noise (the noise standard σ of the high-noise image feature > 15), the non-local mean denoising algorithm is used to denoise the visual image, and the parameters used can be: search window 21×21, similarity window 7×7, attenuation coefficient h = 0.6σ. Further, when the image features are high contrast and low noise, the adaptive median filter denoising algorithm is used to denoise the visual features, and the parameters used can be the window extended to 7×7 and the threshold is 255.

[0039] Further, when denoising acoustic images, the acoustic images can be denoised by determining common features based on the features of visual images and acoustic images.

[0040] It should be noted that the common features of visual images and acoustic images refer to the feature expressions that jointly represent the true physical characteristics of underwater pipelines in visual images and acoustic images. For example, the common features in visual images and acoustic images include structural features, abnormal features, and geometric features; among them, structural features include pipeline axes, girth welds, and support structure positions, etc.; abnormal features include that the corrosion area of the underwater pipeline appears as color patches visually and shows enhanced scattering acoustically; geometric features include three-dimensional morphological features such as pipe diameter changes and depression deformations. Then, the acoustic images are denoised based on the obtained common features.

[0041] For example, in a possible implementation, the pipeline axis can be extracted from a visual image through edge detection and Hough transform. In an acoustic image, the reflection trajectory of the pipeline can be extracted through reflection pattern and echo analysis, and this trajectory represents the pipeline axis. Then, the noise type is determined according to the characteristics of the acoustic image (the noise types include Gaussian noise, scattering noise, speckle noise, etc.). Next, the pipeline axis parameters in the visual image (the pipeline axis parameters include direction, position, etc.) and the pipeline axis parameters in the acoustic image are registered to ensure their spatial consistency. Further, the acoustic image is denoised according to the pipeline axis parameters. When specifically denoising, for example, first, a mask region containing the pipeline structure (i.e., the strong reflection region near the axis) is constructed with the pipeline axis extracted from the acoustic image as the center to protect the pipeline region. Then, a noise denoising method such as non-local means filtering or guided filtering is used to selectively denoise the noise region outside the mask region.

[0042] Further, when denoising the point cloud data, the point cloud data can be classified through semantic segmentation, and the underwater pipeline body and other objects (such as water bodies, attachments, etc.) are marked. Then, using the neighborhood information of each point cloud in the point cloud data, the semantic consistency of each point cloud is calculated. Among them, if the semantic similarity of a certain point cloud with other point clouds in its neighborhood is lower than a preset threshold, it is marked as an isolated point. Immediately afterwards, a clustering algorithm is used to cluster the points of the same category, and small clusters and scattered discrete points are removed.

[0043] It should be noted that in this embodiment, different methods are used to denoise the visual image, acoustic image, and point cloud data. The denoising algorithm and parameters of the visual image are adaptively determined according to its characteristics, which can more accurately remove noise while retaining effective information. The denoising of the acoustic image is based on the common characteristics with the visual image, making the cross-modal denoising process more robust. The denoising of the point cloud data considers semantic information, avoiding misdeleting important points by simple statistical filtering methods and improving the quality of the point cloud data. The denoising methods of the visual, acoustic, and point cloud data respectively adapt to their characteristics, which can enhance the adaptability in different environments. In addition, the denoising of the acoustic image not only depends on its own characteristics but also combines the visual image characteristics to determine the common characteristics, making the denoising more stable and having physical consistency.

[0044] In summary, this solution adopts a highly targeted denoising strategy for visual, acoustic, and point cloud data, improving the denoising effect, data consistency, computational efficiency, and system robustness, providing high-quality data support for subsequent processing.

[0045] S102. Train a detection model using the multiple groups of data.

[0046] Specifically, refer to Figure 2, the detection model includes a first feature extraction module, a multimodal fusion module, a second feature extraction module, and a fault classifier.

[0047] Among them, the first feature extraction module is used to extract features from visual images, acoustic images, and laser point clouds respectively, obtain initial visual features, initial acoustic features, and initial point cloud features, and splice the learnable prompt features corresponding to each modality (denote the learnable prompt features as tokens) on the initial features of each modality to form spliced visual features, spliced acoustic features, and spliced point cloud features.

[0048] Optionally, Figure 3 The following is a schematic structural diagram of the first feature extraction module shown in an exemplary embodiment of the present application. Please refer to Figure 3 , in a possible implementation, the first feature extraction module includes a visual image encoder, an acoustic image encoder, and a laser point cloud encoder; the visual image encoder is used to extract features from visual images; the acoustic image encoder is used to extract features from acoustic images; the laser point cloud encoder is used to extract features from laser point clouds;

[0049] Among them, the network parameters of the visual image encoder and the laser point cloud encoder are fixed, and the network parameters of the acoustic image encoder are learnable.

[0050] Specifically, the detection model takes the visual image as the backbone, and uses a large visual model (CLIP model, based on the Transformer architecture) that processes visual image information as the backbone network for fault detection. The visual image encoder adopts the convolutional part in the backbone network of the large visual model. By extracting features from visual image information through the visual image encoder, initial visual features can be obtained .

[0051] Furthermore, in a possible implementation, the acoustic image encoder uses a multi-layer residual-based convolutional network to extract features from acoustic images to obtain initial acoustic features . In addition, the laser point cloud encoder uses the VoxelNet network to extract 3D voxel features from the laser point cloud data. After that, the obtained 3D voxel features are flattened along the Z-axis to generate a bird's-eye view, and initial point cloud features can be obtained .

[0052] It should be noted that, among them, represents the number of inputs of each sample during training, represents the length of the feature map, represents the width of the feature map, represents the number of channels of the features. For example, the initial visual feature , then at this time, it represents that the number of visual image samples for a single input is 8, the length of the output initial visual feature map is 20, the width of the output initial visual feature map is 15, and the number of channels of the output initial visual features is 2048.

[0053] In this embodiment, the network parameters of the visual image encoder and the lidar point cloud encoder are fixed, and the network parameters of the acoustic image encoder are learnable.

[0054] Specifically, the network parameters refer to the weights and biases of the encoder. The fact that the network parameters of the visual image encoder and the lidar point cloud encoder are fixed means that the weights and biases inside the visual image encoder and the lidar point cloud encoder will not be updated during the training process of the detection model, and their existing feature extraction capabilities are retained; the fact that the network parameters of the acoustic image encoder are learnable means that its weights and biases will be continuously updated during the training process to adapt to the requirements of the current task and optimize its own feature extraction capabilities.

[0055] Further, please continue to refer to Figure 4 , after the first feature extraction module extracts the initial visual features, initial acoustic features, and initial point cloud features from the visual image, acoustic image, and point cloud data respectively, then flatten the dimensions of the initial visual features, initial acoustic features, and initial point cloud features into a one-dimensional length dimension to obtain the embedding tokens corresponding to each modality (each modality includes vision, acoustics, and point cloud). After that, concatenate the prompt tokens corresponding to each modality on the embedding tokens of each modality, and then the concatenated visual features, concatenated acoustic features, and concatenated point cloud features are formed.

[0056] It should be noted that the prompt tokens corresponding to each modality can indicate the detection model to recognize the specific modality corresponding to the input when processing the input. After that, concatenating the embedding token of each modality with the prompt token can obtain a new token, which contains all the information of the modality prompt and modality embedding at the same time. In this way, the detection model can effectively fuse the information of different modalities later, help the detection model understand the relationship between modalities, and then can perform more effective multi-modal reasoning and processing.

[0057] Referring to the previous description, it can be understood that the concatenated visual features, concatenated acoustic features, and concatenated point cloud features respectively represent the new tokens corresponding to the image modality, acoustic modality, and point cloud modality; among them, the concatenated visual features are denoted as , the concatenated acoustic features are denoted as , the concatenated point cloud features are denoted as , represents the length of the new token, that is, the dimension of the concatenated features.

[0058] It should be noted that the modality prompt tokens corresponding to each modality are all learnable parameters. Through continuous training of the modality prompt tokens, the values of the tokens can be adjusted during the training process to better help the model understand and fuse information from different modalities, thereby enhancing the multi-modal processing ability of the detection model. For example, the prompt tokens of the visual modality will be gradually adjusted during the training process to help the visual features fuse better with the features of other modalities during splicing; the prompt tokens of the acoustic modality will be gradually adjusted during the training process to help the model better identify and understand audio data; the prompt tokens of the point cloud modality will be continuously adjusted during the training process to learn how to identify the best point cloud data.

[0059] Furthermore, please continue to refer to Figure 2 , the detection model further includes a multi-modal fusion module, which is used to fuse the spliced features to obtain the final fused feature, that is, to fuse the spliced visual feature, the spliced acoustic feature and the spliced point cloud feature to obtain the final fused feature.

[0060] Specifically, the multi-modal fusion module is specifically used to fuse the spliced visual feature, the spliced acoustic feature and the spliced point cloud feature to obtain an initial fused feature, perform a non-linear transformation on the initial fused feature to obtain an enhanced fused feature, and perform a weighted process on the spliced visual feature and the enhanced fused feature to obtain the final fused feature.

[0061] It can be understood that the multi-modal fusion module may include a fusion layer, a feed-forward network and a fully connected layer. The fusion layer is specifically used to fuse the spliced visual feature, the spliced acoustic feature and the spliced point cloud feature to obtain an initial fused feature; the feed-forward network is used to perform a non-linear transformation on the initial fused feature to obtain an enhanced fused feature; the fully connected layer is used to perform a weighted process on the spliced visual feature and the enhanced fused feature to obtain the final fused feature.

[0062] The following introduces the specific implementation principle inside the multi-modal fusion module:

[0063] Specifically, for example, the spliced features can be directly fused based on a weighted method or a splicing method to obtain an initial fused feature. In addition, in a possible implementation manner, when fusing the spliced features, the spliced visual feature and the spliced acoustic feature can be first fused based on a first cross-attention module, using the spliced visual feature as the query, the spliced acoustic feature as the key and value, to obtain a first fused feature. Further, based on a second cross-attention module, using the first fused feature as the query, the spliced point cloud feature as the key and value, the first fused feature and the spliced point cloud feature are fused to obtain the initial fused feature.

[0064] Optionally, in a possible implementation, fusing the spliced visual features, the spliced acoustic features, and the spliced point cloud features to obtain an initial fusion feature includes:

[0065] Step 1: For each type of modal data, calculate the confidence of the modal data according to the feature information of the modal data.

[0066] In this step, for visual image data, the confidence can be calculated based on feature information such as the clarity, contrast, and noise level of the image. For example, the Laplacian variance can be used to calculate the image blurriness as the confidence; for another example, a CNN model can be used to learn and extract the image quality features in the visual image features and predict its confidence.

[0067] For acoustic image data, the quality of the acoustic image can be evaluated based on the signal-to-noise ratio or spectral energy distribution of the acoustic image data, and its confidence can be calculated. For example, after calculating the short-time Fourier transform, the proportion of the high-energy region can be calculated to determine the signal quality, and the confidence can be calculated based on the signal quality.

[0068] For point cloud image data, the confidence of the point cloud modal data can be calculated based on the sparsity, point density, reconstruction integrity, etc. of the point cloud. For example, the density histogram of the point cloud can be calculated. If the density is low, it indicates that the point cloud data is relatively sparse and the confidence is low.

[0069] Step 2: Normalize the confidence of various modal data, and determine the normalized result as the weighted weight corresponding to the modal data.

[0070] In this step, the confidence of the image modal data, the acoustic modal data, and the point cloud modal data are normalized respectively.

[0071] It should be noted that the Softmax method can be used to normalize the confidence of the three modal data. Among them, assuming that the confidence of the image modal data, the acoustic modal data, and the point cloud modal data are respectively , , , then the following formula is used for normalization:

[0072] ;

[0073] ;

[0074] ;

[0075] Among them, is the weighted weight corresponding to the image modal data; is the weighted weight corresponding to the acoustic modal data; is the weighted weight corresponding to the point cloud modal data.

[0076] Step 3: Fuse the spliced visual feature, the spliced acoustic feature, and the spliced point cloud feature according to the weighted weights corresponding to the respective modal data to obtain an initial fusion feature.

[0077] In this step, according to the following formula, the spliced visual feature, the spliced acoustic feature, and the spliced point cloud feature are fused by weighted summation to obtain an initial fusion feature:

[0078] ;

[0079] where is the initial fusion feature; is the spliced visual feature; is the spliced acoustic feature; is the spliced point cloud feature.

[0080] It should be noted that in this embodiment, when obtaining the initial fusion feature based on the spliced visual feature, the spliced acoustic feature, and the spliced power supply feature, an adaptive weighting mechanism is introduced, and the weighted weights are dynamically adjusted according to the confidence levels of different modal data, which can improve the fusion effect of multi-modal data, improve the accuracy of the initial fusion feature, enable the detection model to have stronger adaptability to changes in the quality of different modal data, and thus improve the robustness and accuracy of underwater pipeline fault detection.

[0081] Furthermore, in a possible implementation manner, a non-linear transformation is performed on the initial fusion feature to obtain an enhanced fusion feature, including:

[0082] Based on a feedforward network, a non-linear transformation is performed on the initial fusion feature using the first formula to obtain an enhanced fusion feature; the first formula is:

[0083] ;

[0084] where the is the initial fusion feature;

[0085] the is the enhanced fusion feature;

[0086] the and the are learnable parameter matrices, the and the are learnable bias terms, and the is a non-linear activation function.

[0087] Further, in a possible implementation, the spliced visual feature and the enhanced fusion feature are weighted to obtain a final fusion feature, including:

[0088] Based on the fully connected layer, the spliced visual feature and the enhanced fusion feature are weighted using the second formula to obtain the final fusion feature; the second formula is:

[0089] ;

[0090] wherein, the is the final fusion feature;

[0091] the is the spliced visual feature;

[0092] the is the enhanced fusion feature;

[0093] the is a learnable parameter vector,

[0094] It should be noted that visual, acoustic, and point cloud data each provide information in different dimensions. After fusing the three, the complementarity of different modality features can be achieved. Further, through non-linear transformation, deeper feature relationships can be learned, enhancing the expressive power of the features, making the fused features more abstract and rich. Finally, by further fusing the visual features and the non-linearly transformed features, the visual features cannot be completely replaced by other modality features during the final fusion, the original visual information can be maintained, ensuring that the visual information still dominates, and the visual perception will not be weakened. Additionally, extra information such as depth and material can be supplemented through the features of other modalities, improving the overall perception ability. Moreover, through the learnable parameters, the network is allowed to automatically adjust the weights of the features, which can not only improve the information interaction ability but also make the multi-modal contribution degrees different in different scenarios, enhancing the generalization ability.

[0095] Further, please refer to Figure 2, the detection model further includes a second feature extraction module, which includes a plurality of encoding modules. Each encoding module includes a plurality of cascaded encoding layers, a low-rank adapter connected to the output end of the penultimate encoding layer of the plurality of encoding layers, and a weighted fusion layer connected to the output ends of the last encoding layer of the plurality of encoding layers and the output end of the low-rank adapter at the same time; the first encoding layer of the plurality of encoding layers of the first encoding module is used to receive the visual splicing features, and the input end of the low-rank adapter of the first encoding module is connected to the output end of the multi-modal fusion module; the low-rank adapters of two adjacent encoding modules are connected, and the weighted fusion layer of the previous encoding module among two adjacent encoding modules is connected to the first encoding layer of the latter encoding module; the parameters of the encoding layer and the weighted layer in each encoding module are fixed, and the parameters of the low-rank adapter are learnable.

[0096] It should be noted that the specific number of encoding modules is set according to actual needs, and in this embodiment, it is not limited. In addition, the encoding layer in the encoding module is a Transformer encoding layer. That is, the model part involved in the logic line from the visual image to the fault classifier can adopt an existing pre-trained large visual model. Further, for the multiple Transformer encoding layers in this model part, they are divided into multiple parts, and each part forms an encoding module. Further, for this encoding module, the last layer of the transformer encoding layer is taken for adaptation processing to form a multi-level adaptation processing, that is, the output of the penultimate Transformer encoding layer is input to the low-rank adapter at the same time.

[0097] Specifically, Figure 4 is a schematic internal structure diagram of the low-rank adapter shown in an exemplary embodiment of the present application. Please refer to Figure 4 , in a possible implementation, the low-rank adapter includes a cross-attention module, a dimensionality reduction module, a non-linear processing module, and a dimensionality increase module;

[0098] Among them, the cross-attention module is used to use the main line features from the encoding layer as queries, and use the side line features from the multi-modal fusion module or the previous low-rank adapter as keys and values to fuse the main line features and the side line features to obtain fused features;

[0099] The dimensionality reduction module is used to perform dimensionality reduction processing on the fused features to obtain dimensionality-reduced fused features;

[0100] The non-linear processing module is used to perform non-linear processing on the dimensionality-reduced fused features to obtain processed features;

[0101] The dimensionality increase module is used to process the processed features to obtain dimensionality-increased fused features; among them, the dimension of the dimensionality-increased fused features is the same as the dimension of the fused features.

[0102] Specifically, the cross-attention module in the low-rank adapter receives the main-line features from the encoding layer ( Figure 2 the transformer encoder in it), and the side-line features from the low-rank adapter in the multi-modal fusion module or the previous encoding module (for the sake of convenience of description, the features output by the low-rank adapter in the multi-modal fusion module or the previous encoding module are denoted as side-line features). Through the attention mechanism, they are fused (wherein, the main-line features are used as the query Query, and the side-line features are used as the key Key and value Value) to obtain the fused features.

[0103] Referring to the previous description, the main-line features are the features directly transmitted in the backbone network of the detection model, and the side-line features come from the low-rank adapter, which are the auxiliary information in the detection model and provide additional feature information for the main-line features in a cross-modal or cross-level manner. In this embodiment, the features from the transformer encoding layer are the main-line features, and the features from the low-rank adapter in the multi-modal fusion module or the previous encoding module are the side-line features.

[0104] Further, the dimension of the fused features is reduced by the dimension reduction module in the low-rank adapter to reduce the complexity of the fused features. For example, a single-layer fully connected layer can be used as the dimension reduction module to reduce the feature dimension C of the fused features to Csmall. Then, the non-linear processing module in the low-rank adapter performs non-linear processing on the dimension-reduced features to obtain the processed features, and the processed features have strong non-linear expression ability.

[0105] It should be noted that the non-linear processing module in the low-rank adapter can be selected according to actual needs, and in this application, it is not limited. For example, in one embodiment, the ReLU activation function can be used for non-linear processing.

[0106] Further, a dimension increase process is performed by the dimension increase module in the low-rank adapter. For example, a fully connected layer can be used to restore the feature dimension of the processed module from Csmall to the dimension C to obtain the dimension-increased fused features, so that it can be ensured that the output dimension-increased fused features are compatible with the backbone network during data transmission.

[0107] Finally, referring to Figure 2, for an encoding module, the upsampled features output by the low-rank adapter are fused with the output of the last Transformer encoding layer in the encoding module where it is located through a weighted fusion layer, and the fused result is input into the first encoding layer of the next encoding module; at the same time, the upsampled features of the low-rank adapter are also input into the low-rank adapter inside the next encoding module for further processing. In other words, for the low-rank adapter, its input involves the main features transmitted from the encoding layer and the side features transmitted from the multi-modal fusion module or the previous low-rank adapter.

[0108] Referring to the previous description, in this embodiment, by setting the second feature extraction module as a multi-level structure based on the residual low-rank adapter and injecting multi-modal information into the low-rank adapter through the cross-attention mechanism, the pre-trained vision large model can better utilize multi-modal information, effectively adapt to the underwater pipeline detection task, improve the detection accuracy and robustness, and reduce the training parameters at the same time.

[0109] It should be noted that the parameters of the encoding layer and the weighting layer in each encoding module are fixed, which means that the weights of the encoding layer and the weighting layer are not updated during the training stage of the detection model. These parameters have the general representation ability in multiple scenarios through pre-training before the detection model is trained. Even without weight update, the stable feature extraction ability can be achieved; while the parameters of the low-rank adapter are learnable, which means that in the encoding module, the parameters of the low-rank adapter will be gradually updated during the training process of the detection model. Thus, on the premise of maintaining the stability of the vision backbone network, by adjusting the parameters of the low-rank adapter, the adaptive adjustment for specific tasks can be realized.

[0110] Optionally, in a possible implementation, in different encoding modules, low-rank adapters with variable capacities (i.e., the ranks of the low-rank adapters in different Transformer encoding modules are different) are adopted to adapt to the information requirements at different levels. In this way, the flexibility of feature transformation can be effectively improved.

[0111] Furthermore, in a possible implementation, the rank of the low-rank adapter in each encoding module is determined based on the features input to the low-rank adapter.

[0112] Specifically, the rank of the low-rank adapter is the dimension obtained after the dimensionality reduction module in the low-rank adapter reduces the dimensionality of the fused features. It should be noted that the rank of the low-rank adapter will affect the efficiency and performance of the low-rank adapter. In this embodiment, selecting low-rank adapters with different ranks for different features can ensure the balance between the efficiency and performance of the low-rank adapter to the greatest extent and adapt to multiple scenarios.

[0113] Optionally, in a possible implementation, the method for determining the rank of the low-rank adapter of each encoding module includes:

[0114] Step 1: Process the features input to the low-rank adapter using a multi-layer perceptron to obtain an initial predicted value.

[0115] In this step, the features input to the low-rank adapter are input into the multi-layer perceptron. The multi-layer perceptron extracts implicit complexity information (the complexity information includes feature entropy, activation sparsity, cross-modal conflict intensity, etc.) from the input features. Then, the obtained complexity information is mapped to an initial predicted value of the rank.

[0116] Step 2: Activate the initial predicted value using an activation function to obtain the rank of the low-rank adapter.

[0117] In this step, the initial predicted value of the rank obtained in Step 1 is input into the selected activation function (the selected activation function can be the Sigmoid function). The activation function performs activation processing on the initial predicted value, adjusts the range of the initial predicted value, ensures that the value of the rank is within a reasonable range, and finally obtains the rank of the low-rank adapter.

[0118] It should be noted that in this embodiment, each encoding module receives feature information at different levels, and the expression capabilities and information densities of these features between levels are different. To adapt to different information requirements, the rank of the low-rank adapter is dynamically determined according to the input features of each encoding module. Through this dynamic adjustment mechanism, a larger rank can be used in higher-level encoding modules to effectively capture global features, while a smaller rank is used in lower-level encoding modules to avoid excessive computational overhead. In this way, not only can the expression ability of the features be improved, but also the computational overhead can be taken into account at the same time.

[0119] S103: Input a set of data to be detected into the trained detection model to obtain a detection result.

[0120] Specifically, after the training of the detection model in the previous steps, in this step, a set of data to be detected can be input into the trained detection model after denoising processing, and the detection model outputs the result.

[0121] Specifically, the first feature extraction module in the detection model first extracts the initial visual features of the visual image, the initial acoustic features of the acoustic image, and the initial point cloud features of the point cloud data in the data to be detected; then, the learned prompt features (tokens) corresponding to each modality are concatenated onto the initial features of each modality to form concatenated visual features, concatenated acoustic features, and concatenated point cloud features; after that, the multi-modal fusion module first performs weighted fusion on the concatenated visual features, concatenated acoustic features, and concatenated point cloud features to obtain initial fusion features, then performs non-linear transformation on the initial fusion features to obtain enhanced fusion features, and finally, performs weighted processing on the concatenated visual features and the enhanced fusion features to obtain final fusion features; then, the final fusion features are input into the low-rank adapter in the first encoding module, and the concatenated visual features are input into the first layer of the transform encoding layer in this encoding module, and are processed layer by layer through the internal transform encoding layer. After that, the output of the second-to-last transform encoding layer in this encoding module is input into the low-rank adapter, and the low-rank adapter outputs dimension-increased fusion features based on the features transmitted from the multi-modal fusion modality and the features transmitted from the second-to-last transform encoding layer. The dimension-increased fusion features and the features output by the last transform encoding layer are input into the weighted fusion layer, and the weighted fusion layer performs weighted fusion to obtain the result.

[0122] Further, this result is used as input features and is input into the first layer of the transform encoding layer in the next encoding module again. At the same time, the dimension-increased fusion features output by the low-rank adapter in the previous encoding module are also input into the low-rank adapter in this next encoding module;...; after cyclic processing by multiple encoding modules, the output of the weighted fusion layer of the last encoding module is input into the fault classifier, and the fault classifier outputs the final detection result.

[0123] It should be noted that the fault classifier determines the fault type according to the input content by setting an empirical database; among them, the empirical database can use an existing database or be set by the operator according to experience. In this application, it is not limited.

[0124] In this embodiment, by fusing the information of multi-modal sensors such as cameras, sonars, and lidars, and adding learnable feature prompts to the embedded representations of each modality, the model can be guided to better utilize the information of each modality; and a multi-level structure based on a residual low-rank adapter is adopted, and multi-modal information is injected into the low-rank adapter through a cross-attention mechanism. In this way, the pre-trained detection model can better utilize the information from multi-modal sensors, effectively adapt to the underwater pipeline detection task, and while improving the detection accuracy and robustness, reduce the training parameters.

[0125] The underwater pipeline fault detection method provided in this embodiment. The first feature extraction module extracts features from visual images, acoustic images, and laser point clouds respectively, and splices learnable prompt features on the basis of the initial features, so that data of different modalities have better expression ability before fusion. The multimodal fusion module adopts a learnable fusion method to make the information of visual, acoustic, and point cloud data complementary, effectively enhancing the extraction ability of key fault features and overcoming the limitations caused by the possible influence of a single modality on the environment. Further, the second feature adopts a hierarchical coding structure based on a low-rank adapter, allowing information to be gradually optimized in coding layers of different depths, making the semantics of the features more complete and enhancing the feature expression ability. In this way, the three work together, enabling the detection model to accurately identify different types of pipeline faults in various underwater environments (such as insufficient light, noise interference, etc.) and improving the detection accuracy.

[0126] In addition, the parameters of the coding layer and the weighting layer are fixed to ensure the stability of feature extraction, while the parameters of the low-rank adapter and the multimodal fusion module are learnable, enabling the model to perform adaptive optimization in different environments, improving the generalization ability, enhancing the detection accuracy and robustness, and reducing the training parameters at the same time.

[0127] In summary, the method provided in this embodiment, through multimodal data fusion and hierarchical coding structure, solves the difficulties of multimodal fusion and the challenges of environmental interference faced by traditional underwater pipeline fault detection, and realizes efficient, accurate, and robust underwater pipeline fault detection.

[0128] Corresponding to the foregoing embodiment of an underwater pipeline fault detection method, the present application also provides an embodiment of an underwater pipeline fault detection device.

[0129] Figure 5 It is a schematic diagram of the first embodiment of the underwater pipeline fault detection device provided by the present application. Please refer to Figure 5 , the device provided in this embodiment includes an acquisition module 510, a training module 520, and a detection module 530; wherein, the acquisition module 510 is used to obtain a set of data at each of multiple positions pre-selected from an underwater pipeline, obtaining multiple sets of data; a set of data includes a visual image, an acoustic image, and a laser point cloud.

[0130] The training module 520 is used to train a detection model by using the multiple sets of data; the detection model includes a first feature extraction module, a multi-modal fusion module, a second feature extraction module, and a fault classifier; the first feature extraction module is used to extract initial features of each modality, and splice the learnable prompt features corresponding to the modality on the initial features of each modality to form spliced visual features, spliced acoustic features, and spliced point cloud features; the multi-modal fusion module is used to fuse the spliced features to obtain final fused features; the second feature extraction module includes a plurality of encoding modules, and each encoding module includes a plurality of cascaded encoding layers, a low-rank adapter connected to the output end of the second-to-last encoding layer of the plurality of encoding layers, and a weighted fusion layer connected to the output ends of the last encoding layer of the plurality of encoding layers and the output end of the low-rank adapter; the first encoding layer of the plurality of encoding layers of the first encoding module is used to receive the visual spliced features, and the input end of the low-rank adapter of the first encoding module is connected to the output end of the multi-modal fusion module; the low-rank adapters of two adjacent encoding modules are connected, and the weighted fusion layer of the previous encoding module of two adjacent encoding modules is connected to the first encoding layer of the subsequent encoding module;

[0131] The detection module 530 is used to input a set of data to be detected into the trained detection model to obtain a detection result.

[0132] The device in this embodiment can be used to execute Figure 1 the steps of the method embodiment shown. The specific implementation principle and process are similar and will not be elaborated here.

[0133] For the implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method for details, which will not be elaborated here.

[0134] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative work.

[0135] The above are only the preferred embodiments of this application, and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included within the scope of protection of this application.

Claims

1. An underwater pipeline fault detection method, characterized in that, The method includes: For each of a plurality of positions pre-selected from an underwater pipeline, a set of data at this position is acquired to obtain multiple sets of data; a set of data includes a visual image, an acoustic image, and a laser point cloud; Using the multiple sets of data to train a detection model; the detection model includes a first feature extraction module, a multi-modal fusion module, a second feature extraction module, and a fault classifier; the first feature extraction module is used to extract initial features of each modality, and splice learnable prompt features corresponding to this modality on the initial features of each modality to form spliced visual features, spliced acoustic features, and spliced point cloud features; the multi-modal fusion module is used to fuse the spliced features to obtain a final fused feature; the second feature extraction module includes multiple encoding modules, and each encoding module includes a plurality of cascaded encoding layers, a low-rank adapter connected to the output end of the penultimate encoding layer of the plurality of encoding layers, and a weighted fusion layer connected to the output ends of the last encoding layer of the plurality of encoding layers and the output end of the low-rank adapter at the same time; the first encoding layer of the plurality of encoding layers of the first encoding module is used to receive the visual spliced features, and the input end of the low-rank adapter of the first encoding module is connected to the output end of the multi-modal fusion module; the low-rank adapters of two adjacent encoding modules are connected, and the weighted fusion layer of the previous encoding module among two adjacent encoding modules is connected to the first encoding layer of the latter encoding module; Input a set of data to be detected into the trained detection model to obtain a detection result.

2. The method according to claim 1, characterized in that, The multi-modal fusion module is specifically used to fuse the spliced visual features, the spliced acoustic features, and the spliced point cloud features to obtain an initial fused feature, perform a non-linear transformation on the initial fused feature to obtain an enhanced fused feature, and perform a weighted process on the spliced visual feature and the enhanced fused feature to obtain a final fused feature.

3. The method according to claim 1, characterized in that The first feature extraction module includes a visual image encoder, an acoustic image encoder, and a laser point cloud encoder; the visual image encoder is used to extract features from the visual image; the acoustic image encoder is used to extract features from the acoustic image; the laser point cloud encoder is used to extract features from the laser point cloud; Among them, the network parameters of the visual image encoder and the laser point cloud encoder are fixed and unchanged, and the network parameters of the acoustic image encoder are learnable.

4. The method according to claim 2, wherein The performing a non-linear transformation on the initial fused feature to obtain an enhanced fused feature includes: Based on a feed-forward network, performing a non-linear transformation on the initial fused feature using a first formula to obtain an enhanced fused feature; the first formula is: ; Among them, the is the initial fusion feature; The said is the said enhanced fusion feature; The and the are learnable parameter matrices, the and the are learnable bias terms, and the is a non-linear activation function.

5. The method according to claim 2 or 4, characterized in that, The performing a weighted process on the spliced visual feature and the enhanced fused feature to obtain a final fused feature includes: Based on a fully-connected layer, performing a weighted process on the spliced visual feature and the enhanced fused feature using a second formula to obtain a final fused feature; the second formula is: ; Among them, the is the final fusion feature; The said is the said spliced visual feature; The said is the said enhanced fusion feature; The said is a learnable parameter vector, .

6. The method according to claim 1, characterized in that, The low-rank adapter includes a cross-attention module, a dimensionality reduction module, a non-linear processing module, and a dimensionality increase module; Among them, the cross-attention module is used to take the main-line features from the encoding layer as queries, and take the sub-line features from the multi-modal fusion module or the previous low-rank adapter as keys and values, and fuse the main-line features and the sub-line features to obtain fused features; The dimensionality reduction module is used to perform dimensionality reduction processing on the fused features to obtain dimensionally reduced fused features; The non-linear processing module is used to perform non-linear processing on the dimensionally reduced fused features to obtain processed features; The dimensionality increase module is used to process the processed features to obtain dimensionally increased fused features; among them, the dimension of the dimensionally increased fused features is the same as the dimension of the fused features.

7. The method according to claim 1, wherein After obtaining a set of data at the position, the method further includes: For visual images, determine a denoising algorithm and the parameters of the denoising algorithm according to the features of the visual image, and perform denoising processing on the visual image based on the determined denoising algorithm and the parameters of the denoising algorithm; For acoustic images, determine common features based on the features of visual images and acoustic images, and perform denoising processing on acoustic images based on the common features; For point cloud data, remove isolated points and discrete points based on semantic similarity and semantic coherence.

8. The method according to claim 2, characterized in that, The fusing the spliced visual features, the spliced acoustic features and the spliced point cloud features to obtain an initial fused feature includes: For each type of modal data, calculate the confidence of the modal data according to the feature information of the modal data; Perform normalization processing on the confidences of various modal data, and determine the normalization processing result as the weighted weight corresponding to the modal data; Fuse the spliced visual features, the spliced acoustic features and the spliced point cloud features according to the weighted weights corresponding to the modal data to obtain an initial fused feature.

9. The method according to claim 6, wherein The ranks of the low-rank adapters of different encoding modules are different; the rank of the low-rank adapter of each encoding module is determined based on the features input to the low-rank adapter; The method for determining the rank of the low-rank adapter of each encoding module includes: Use a multi-layer perceptron to process the features input to the low-rank adapter to obtain an initial predicted value; Use an activation function to perform activation processing on the initial predicted value to obtain the rank of the low-rank adapter.

10. An underwater pipeline fault detection device, characterized in that, The device includes an acquisition module, a training module, and a detection module; among them, The acquisition module is used to obtain a set of data at each of a plurality of positions pre-selected from an underwater pipeline, to obtain multiple sets of data; a set of data includes visual images, acoustic images, and laser point clouds; The training module is used to train a detection model using the multiple sets of data; the detection model includes a first feature extraction module, a multi-modal fusion module, a second feature extraction module, and a fault classifier; the first feature extraction module is used to extract initial features of each modality and concatenate the learnable prompt features corresponding to each modality on the initial features of each modality to form a concatenated visual feature, a concatenated acoustic feature, and a concatenated point cloud feature; the multi-modal fusion module is used to fuse the concatenated features to obtain a final fused feature; the second feature extraction module includes multiple encoding modules, and each encoding module includes a plurality of encoding layers connected in cascade, a low-rank adapter connected to the output end of the second-to-last encoding layer of the plurality of encoding layers, and a weighted fusion layer connected to the output ends of the last encoding layer of the plurality of encoding layers and the output end of the low-rank adapter at the same time; the first encoding layer of the plurality of encoding layers of the first encoding module is used to receive the visual concatenated feature, and the input end of the low-rank adapter of the first encoding module is connected to the output end of the multi-modal fusion module; the low-rank adapters of two adjacent encoding modules are connected, and the weighted fusion layer of the previous encoding module among two adjacent encoding modules is connected to the first encoding layer of the subsequent encoding module. The detection module is used to input a set of data to be detected into the trained detection model to obtain a detection result.

Citation Information

Patent Citations

  • Flotation froth image segmentation method and device based on multi-modal data fusion

    CN116258719A

  • Transform-based spatio-temporal context target tracking method and system

    CN117315293A

  • Three-dimensional target detection method based on multi-modal fusion and deformable attention

    CN117975436A

  • Acousto-optic fusion submarine pipe cable state identification method

    CN118447378A

  • Image segmentation method, medical image segmentation system and computer terminal

    CN118735949A

Cited By

  • Audio processing method and device, in-vehicle infotainment device and medium

    CN121483233A

  • Generated image detection method and device, electronic equipment and readable storage medium

    CN121527532A