Multi-mode propeller fault diagnosis method and diagnosis system
By constructing a quality-aware gating network, the acoustic and visual signal quality is quantified in real time, and the interaction of deep features is dynamically controlled. This solves the problem of insufficient robustness in underwater propeller fault diagnosis and achieves high-precision and adaptive diagnosis in complex environments.
Patent Information
- Application Number
- CN202511516098.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-10-23
AI Technical Summary
In complex and harsh underwater environments, existing multimodal diagnostic technologies rely on signal characteristics, but the quality of the original signal is unknown, resulting in insufficient robustness and difficulty in achieving accurate and reliable propeller fault diagnosis.
By constructing a quality-aware gating network, the quality of acoustic and visual signals is quantified in real time, the interaction intensity of deep features is dynamically controlled, high-quality modal features are enhanced, and low-quality modal features are corrected, thereby achieving high robustness and accuracy in fault diagnosis.
It maintains stable diagnostic performance in harsh environments such as noise, low light, and occlusion, achieving higher diagnostic accuracy and adaptability, avoiding interference from low-quality information, and making the training process more stable.
Smart Images

Figure CN121008566A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of underwater vehicle fault diagnosis, in particular to a multi-modal propeller fault diagnosis method and system. BACKGROUND
[0002] The propeller of an autonomous underwater vehicle (AUV) is a core propulsion and power component, playing a vital role in the fields of ocean exploration, resource exploration, underwater engineering, national defense security, etc. However, when operating in a real underwater environment with dim light, turbid water, and complex background noise, the propeller is prone to failure due to impact, fatigue, cavitation, or entanglement with foreign objects. Once these faults occur, they can cause a series of serious consequences, such as reduced propulsion efficiency and shortened endurance, damaged components or system failure, reduced control performance and increased safety risks, and interrupted tasks and economic losses.
[0003] Therefore, developing a technology that can automatically, real-time, and accurately detect and diagnose various fault states of underwater propellers has great practical significance and application value for ensuring the safety of AUV equipment, improving mission success rate, and reducing the total life cycle operating cost.
[0004] Currently, there are several technical paths for fault diagnosis of underwater propellers, mainly including methods based on operating parameter analysis, methods based on vibration analysis, and methods based on external perception such as acoustics and vision.
[0005] The method based on operating parameter analysis indirectly infers the load state of the propeller by monitoring the current, voltage, speed, torque, and other electrical or control parameters of the propulsion motor. When the propeller is entangled or severely damaged, causing abnormal increase in load, these parameters will show identifiable patterns. The advantages of this method are fast response, low cost, and no need to install additional sensors on the wet end. The disadvantages are that it is an indirect measurement, which is prone to misjudgment of normal working condition changes such as AUV rapid acceleration and encountering strong ocean currents as faults, and it is difficult to distinguish the specific types of faults: for example, entanglement and broken propeller may produce similar load characteristics.
[0006] The method based on vibration analysis is a classic means of mechanical equipment fault diagnosis. It captures abnormal vibration signals caused by problems such as mass imbalance, blade damage, or bearing wear by installing acceleration sensors on the hull or shafting near the propeller. Its advantages are that it is extremely sensitive to mechanical faults and the technology is mature. The disadvantages are that in underwater applications, vibration signals are severely attenuated when transmitted in water, and are easily confused with other structural vibrations of the ship or AUV itself, resulting in low signal-to-noise ratio, and the installation and maintenance of sensors are relatively difficult.
[0007] External perception-based diagnostic methods use non-contact sensors to directly observe the propeller and its surrounding environment, with acoustics and vision as the main means, which is also a highly promising development direction.
[0008] Acoustic diagnostic methods involve using hydrophones to collect acoustic signals radiated by the propeller as it rotates underwater. When the propeller has cracks, is unbalanced, or experiences cavitation, its acoustic characteristics, such as spectral lines at specific frequencies and the overall noise level, will exhibit identifiable changes. Its advantages include the ability to penetrate turbid water and sensitivity to non-surface faults such as internal damage or hydrodynamic anomalies. Its disadvantages include susceptibility to severe interference from the AUV's self-noise from the motor and electronic controls, as well as from marine environmental noise from other vessels and marine life, and difficulty in accurately locating the physical location of the fault.
[0009] Visual diagnostic methods utilize underwater cameras to directly capture images or videos of the propeller, which are then analyzed using computer vision algorithms. Deep learning-based target detection and segmentation techniques have been proven effective in identifying blade breakage, defects, and entanglement with foreign objects such as fishing nets and ropes. The advantages of this method are its intuitiveness and accuracy; once a clear image of the fault is captured, the type and location of the fault can be quickly identified. The disadvantage is that its performance is entirely limited by underwater lighting and visibility; its effectiveness decreases sharply or even fails completely in dark or murky waters.
[0010] As can be seen from the above, diagnostic methods based on different physical quantities such as operating parameters, vibration, acoustics, and vision each have their own advantages and disadvantages. For example, electrical and vibration signals respond quickly to internal mechanical faults and load changes, while acoustics and vision can provide more direct evidence about the acoustic signature and physical form of the fault. Therefore, multimodal diagnostic strategies that integrate information from multiple sensors have become an inevitable trend and research hotspot for overcoming the bottleneck of single information sources and improving the overall performance of the system.
[0011] Among numerous modal combinations, the fusion of acoustics and vision is particularly noteworthy and promising in the specific scenario of underwater propeller diagnostics. This is because acoustic signals are highly sensitive to dynamic anomalies within the equipment, such as abnormal noises caused by cracks, and to fluid dynamic changes, such as cavitation, while visual information can most directly capture external physical damage, such as blade breakage or foreign object entanglement. The effective combination of the two can create a comprehensive, three-dimensional monitoring system, from "internal auscultation" to "external visual inspection," providing a wealth of information for accurate and reliable fault diagnosis.
[0012] Therefore, current cutting-edge research generally employs deep learning models, especially fusion architectures based on attention mechanisms, to process audiovisual data. These methods, through complex network designs, attempt to enable models to autonomously learn the deep correlations between acoustic and visual features, and intelligently weight and fuse information accordingly, aiming to achieve superior diagnostic performance compared to single-modal approaches.
[0013] However, when applying these advanced fusion strategies to real, harsh underwater environments, a common and deep-seated inherent flaw exists: the fusion process of existing models relies almost entirely on the features extracted from the signal, while the real-time quality of the original signal carrying these features is "unknowable." This "quality blind spot" is precisely the key bottleneck restricting the robustness of current multimodal diagnostic technologies in practical applications. Summary of the Invention
[0014] The purpose of this invention is to overcome the above-mentioned defects in the existing technology and propose a multi-mode propeller fault diagnosis method and system. It uses the characteristics of high-quality modes to dominate, enhance and correct the characteristics of low-quality modes. Starting from the source quality of the signal, it actively and in real time controls the intensity of deep feature interaction, and can maintain high robustness and accuracy in propeller fault diagnosis even in complex and harsh environments.
[0015] The technical solution of this invention is: a multi-mode propeller fault diagnosis method, comprising the following steps: S1. Multimodal data acquisition and preprocessing: S2. Real-time quantification of the quality of acoustic signal data and visual signal data obtained in step S1; S3. Construct a quality-aware gating network model; S4. Train the quality-aware gating network model and use the trained quality-aware gating network model to output the propeller fault type.
[0016] In step S1, hydrophones deployed near the thruster are used to collect real-time or offline acoustic signals during propeller operation, and the collected acoustic signal data is preprocessed: The acquired raw acoustic waveform data is segmented and converted into a two-dimensional time-spectrum graph using short-time Fourier transform. The time-spectrum graph is then resized and normalized. The underwater camera deployed near the thruster synchronously acquires real-time or offline image sequences of the propeller's working area, and preprocesses the acquired visual signal data: The size of the acquired image frames is uniformly adjusted, the images are converted into tensor format, and then standardized.
[0017] In step S2, the acoustic signal data and visual signal data are evaluated for quality to obtain an acoustic quality score. and visual quality score ; The specific implementation process for quality assessment of acoustic signal data is as follows: S2.1.A.1. Divide the complete acoustic waveform signal into M non-overlapping short time frames; S2.1.A.2. Within each frame, the noise level is estimated using the minimum statistical method: each frame is further divided into several sub-segments, and the lowest energy value among all sub-segments is taken as the noise energy estimate for the current frame. Signal energy It is then calculated by subtracting the noise energy estimate from the total energy of the frame; S2.1.A.3. By averaging the signal-to-noise ratio (SNR) of all valid frames in decibels, an acoustic quality score that stably reflects the SNR level of the entire audio segment is obtained. The calculation formula is as follows: , in, Indicates the total number of frames; This represents the signal energy of the m-th frame; This represents the noise energy estimate for the m-th frame; The specific implementation process for quality assessment of visual data is as follows: S2.1.B.1, Evaluate the color measurement index UICM, sharpness measurement index UISM, and contrast measurement index UIConM of the image; S2.1.B.2. Perform a linear weighted sum of the three indicators UICM, UISM, and UIConM to obtain the final visual quality score. The calculation formula is as follows: in, , , All of these are preset weighting coefficients. , , .
[0018] The obtained acoustic quality score and visual quality score Standardization processes are performed, including offline and online statistical phases. The offline statistical phase includes the following steps: S2.2.1. The acoustic waveforms and visual images collected under different operating conditions are combined into a dataset. For each sample in this dataset, the original acoustic quality score of each sample is calculated. and visual quality score ; S2.2.2 Collect the acoustic quality scores of all samples. Value, calculate its global mean and global standard deviation ; Calculate the visual quality score for all samples. global mean of values and global standard deviation ; The above four statistical measures , , , Stored as preset parameters; The steps involved in the online statistics phase include: S2.2.3. Utilize the parameters stored in the offline statistical phase and process them using standardized formulas: , , in, This represents the standardized acoustic quality score. The standardized visual quality score is represented by the two parameters mentioned above, which together form a two-dimensional quality feature vector. .
[0019] The specific implementation process of step S3 is as follows: S3.1 Determine the data for input to the quality-aware gating network: The input data includes acoustic time-spectrum data, visual image sequence data, and quality feature vector data; The input acoustic time-spectrum data is (Tensor [B, 1, F, T]), where 1 represents the number of channels, B represents the batch, F represents the frequency of the spectrum, and T represents the time dimension of the spectrum. The input visual image sequence data is Where 3 represents the red, green, and blue color channels of the image. H represents the length of the visual sequence, and W represents the image height. The input quality feature vector data is ; S3.2 Extraction of acoustic and visual features; S3.3 Generation of dynamic gating weights; S3.4, Gated Cross-Attention Fusion; S3.5, Final Feature Generation and Multi-Task Prediction.
[0020] The extraction of acoustic features includes the following steps: S3.2. A.1, Local Time-Frequency Pattern Extraction: Extracting the input acoustic time-frequency pattern data... The input is a CNN network that performs a sliding scan across a two-dimensional time-frequency plane through several convolutional and pooling layers to capture local, fundamental acoustic patterns and output a sequence of feature maps. ,in Indicates the number of feature channels. This represents the frequency after dimensionality reduction. This represents the time dimension after dimensionality reduction; S3.2. A.2, Serialization and Dimension Reshaping: Feature Map Sequence Merging or flattening along the frequency dimension to form a time step of The feature dimension of each time step is Feature sequences ,in , representing the dimension of the feature vector at each time step; S3.2.A.3, Global Temporal Dependency Modeling: Modeling feature sequences... Input the Transformer encoder, capture the long-distance dependency between any two time steps in the sequence through the multi-head self-attention mechanism inside the Transformer encoder, and output the sequence features; S3.2.A.4 Feature Aggregation and Output: The sequence features output by the Transformer encoder are aggregated into a single, fixed-dimensional acoustic feature vector through a time-dimensional average pooling operation. ,in This represents the hidden layer dimension of the Transformer; The extraction of visual features includes the following steps: S3.2.B.1 Single-frame feature extraction: Extracting the input image sequence Split by frame, and divide each frame image The data is independently fed into the backbone network of the MobileViT model to generate a high-level semantic feature vector for each frame of the image. S3.2.B.2, Sequence Feature Aggregation: Collecting all features in the sequence. The feature vectors generated from the frame image form a feature vector sequence. Where D_frame represents the feature dimension of a single frame output by the MobileViT backbone network; S3.2.B.3, Temporal Pooling and Output: The feature vector sequence obtained in the above steps is subjected to average pooling in the temporal dimension to obtain a single, fixed-dimensional aggregated visual feature vector. ; Dynamic gating weight generation includes the following steps: S3.3.1 Input and Preprocessing: The quality-aware gating network receives a two-dimensional quality feature vector as the input layer; S3.3.2 Nonlinear Feature Mapping: The two-dimensional quality feature vector of the input layer is fed into a hidden layer containing 64 neurons, and the modified linear unit ReLU is used as the activation function; S3.3.3, Weight logits generation: The output of the hidden layer is fed into the output layer with only one neuron to generate an unnormalized weight value logit; S3.3.4 Normalized Output: The logit values of the weights are processed by the Sigmoid activation function. The Sigmoid function maps the logit values to the interval (0, 1), thus obtaining the dynamic gating weights. The output of this step is the dynamic gating weight data. .
[0021] In step S3.4, the acoustic feature vector, visual feature vector, and dynamic gating weights are received. Through parallel visual feature enhancement paths and acoustic feature enhancement paths, the acoustic feature vector after intelligent fusion enhancement is output. and visual feature vectors ; The specific processing steps for the visual feature enhancement path are as follows: S3.4.A.1, using visual feature vectors As a query, using acoustic feature vectors Using these as keys and values, the acoustic features that the visual features should focus on are calculated, and an attention-weighted acoustic information vector is generated. The calculation formula is as follows: , , in, From visual feature vectors , and From acoustic feature vectors ; Indicates the first The query projection matrix of each attention head; Indicates the first The key projection matrix of each attention head; Indicates the first The value projection matrix of each attention head; The output projection matrix represents the visual enhancement path; Indicates the number of heads of attention; Indicates the dimension of each head; Indicates a splicing operation; S3.4.A.2. Multiply the attention-weighted acoustic information vector with the dynamic gating weights; add this weighted result to the original visual feature vector through a residual connection, and then perform layer normalization to obtain the enhanced visual feature vector. The calculation formula is as follows: , in Indicates the layer normalization processing function; The specific processing steps for the acoustic feature enhancement path are as follows: S3.4.B.1, using acoustic feature vectors As a query, using visual feature vectors Using these as keys and values, the visual features that the acoustic features should focus on are calculated, and an attention-weighted visual information vector is generated. The calculation formula is as follows: , , in, From acoustic feature vectors , and From visual feature vectors ; Indicates the first The query projection matrix of each attention head; Indicates the first The key projection matrix of each attention head; Indicates the first The value projection matrix of each attention head; The output projection matrix represents the acoustic enhancement path; Indicates the number of heads of attention; Indicates the dimension of each head; Indicates a splicing operation; S3.4.B.2, Attention-weighted visual information vectors and Multiplication; combined with the original acoustic feature vector via residual concatenation. The features are added together and then normalized to obtain the enhanced acoustic feature vector. The calculation formula is as follows: .
[0022] The specific implementation process of step S3.5 is as follows: S3.5.1 Final Feature Generation: The enhanced acoustic feature vector obtained after gated cross-attention fusion and visual feature vectors Utilizing dynamic gating weights again We perform a weighted summation to obtain the fusion features ultimately used for the main task decision. The calculation formula is as follows: , S3.5.2 Multi-task prediction: A multi-task learning strategy is adopted, in which different feature vectors are input into three parallel classification prediction heads. The three prediction heads are the main classification prediction head, the acoustic auxiliary prediction head, and the visual auxiliary prediction head. The data processing process of each prediction head is as follows: The input to the main classification prediction head is the fused features. First, the feature is transformed and non-linearly mapped through several fully connected linear layers. Each linear layer is followed by a ReLU activation function. The last linear layer maps the feature dimension to a dimension equal to the number of fault categories and directly outputs the original prediction score of the fault category. The input to the acoustic-assisted prediction head is an unfused acoustic feature vector. The system processes data through a linear layer, a ReLU activation function, and a Dropout layer, and outputs a raw prediction score for the fault category based on pure acoustic features. The input to the visual-assisted prediction head is an unfused visual feature vector. The system processes data through a linear layer, a ReLU activation function, and a Dropout layer, outputting a raw prediction score for the fault category based on pure visual features.
[0023] In step S4, during the training of the quality-aware gated network model, a multi-task learning strategy including main loss and auxiliary loss is adopted, and its total loss is... The calculation formula is: , in, This represents the main classification loss, which incorporates the fused features. The prediction result is obtained by inputting the main prediction head, and then the cross-entropy is calculated with the true label. Represents the acoustic-assisted classification loss, which converts the unfused acoustic feature vectors... The prediction result is obtained by inputting the acoustic-assisted prediction head, and then the cross-entropy is calculated with the real label to obtain the result. The visual-assisted classification loss represents the loss of the unfused visual feature vectors. The prediction result is obtained by inputting the visual aid prediction head, and then the cross-entropy is calculated with the real label to obtain the result. This represents the pre-defined hyperparameters.
[0024] This application also discloses a fault diagnosis system for implementing the above-mentioned multimodal propeller fault diagnosis method, comprising: The data acquisition module is used to collect acoustic signals and visual image sequences during propeller operation and to preprocess the collected data. The quality assessment module is used to perform real-time quality quantification assessment of preprocessed acoustic and visual data, and generate standardized quality feature vectors. The feature extraction module is used to extract depth features from the preprocessed acoustic time-spectrum map and visual image sequence, and output acoustic feature vectors and visual feature vectors. The gated fusion module is used to generate dynamic gate weights based on the quality feature vector and to achieve intelligent fusion of acoustic and visual features through a gated cross-attention mechanism. The classifier module is used to output the propeller fault type prediction result by passing the fused feature vector through the main classification prediction head and the auxiliary prediction head.
[0025] The beneficial effects of this invention are: (1) High robustness: This application can sense and quantify the quality of input data in real time. When the quality of any modality data decreases, it can adaptively reduce its weight in the fusion process, effectively avoiding interference from low quality or erroneous information. Therefore, it can still maintain stable diagnostic performance in harsh environments such as noise, low light, and occlusion. (2) High precision: Through the gated cross-attention mechanism proposed in this application, the information of high-quality modalities can be used more fully to supplement and correct low-quality modalities, realizing deeper and more efficient modal complementarity, thereby obtaining more discriminative fusion features and higher diagnostic precision; (3) Strong adaptability: This application does not require separate model design for different environmental noise or image quality levels. Its quality perception and dynamic fusion mechanism enable it to automatically adapt to various complex and changing working environments. (4) Stable and efficient training: The introduction of auxiliary loss ensures that each single-modal branch can learn meaningful features, avoids excessive dependence on a certain modality during model training, and makes the training process of the whole model more stable and the convergence effect better. Attached Figure Description
[0026] Figure 1 This is a flowchart of the method described in this application; Figure 2 The diagnostic results are output using the method described in this application; Figure 3 Acoustic signals collected by hydrophones under various fault conditions; Figure 4 These are underwater optical images captured by an underwater camera; Figure 5 It is a propeller fault diagnosis confusion matrix obtained from experiments in clean water. Figure 6 The results are experimental comparisons of the robustness of the method described in this application with existing propeller fault diagnosis methods in clean water bodies. Figure 7 The results are experimental comparisons of the robustness of the method described in this application with existing propeller fault diagnosis methods in slightly turbid water. Figure 8 The results are experimental comparisons of the robustness of the method described in this application with existing propeller fault diagnosis methods in heavily turbid water bodies. Figure 9 The results are experimental comparisons of the robustness of the method described in this application with existing propeller fault diagnosis methods in noisy water bodies. Figure 10 The results are experimental comparisons of the robustness of the method described in this application with existing propeller fault diagnosis methods in low-light water bodies. Figure 11 The results are experimental comparisons of the robustness of the method described in this application with existing propeller fault diagnosis methods in complex degraded water bodies. Detailed Implementation
[0027] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0028] Specific details are set forth in the following description to provide a full understanding of the invention. However, the invention can be practiced in many ways other than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0029] This application proposes a multimodal propeller fault diagnosis method, which includes the following specific steps, and the flowchart of the method is shown below. Figure 1 As shown.
[0030] The first step is multimodal data acquisition and preprocessing.
[0031] In this embodiment, hydrophones deployed near the propeller are used to collect acoustic data: real-time or offline acoustic signals during propeller operation are acquired using hydrophones. The acoustic signals are typically waveform data containing timestamps and can be in .wav or .npy format.
[0032] The acquired acoustic data is preprocessed. The raw acoustic waveform data is segmented and converted into a two-dimensional time-spectrum graph using a short-time Fourier transform (STFT). This time-spectrum graph is then resized and normalized to meet the input requirements of the subsequent acoustic feature extraction model.
[0033] At the same time, visual data is collected using underwater cameras deployed near the propeller: real-time or offline image sequences of the propeller's working area are collected simultaneously using underwater cameras, and can be in .jpg or .png format.
[0034] The acquired visual data is preprocessed. The acquired image frames are resized uniformly. In this embodiment, the image size is uniformly resized to 224x224 pixels. Then, the images are converted to tensor format and standardized to meet the input requirements for subsequent visual feature extraction.
[0035] The second step is to quantify the quality of the acoustic and visual data obtained in the first step in real time.
[0036] One of the innovations of this application is that it can perform real-time, quantitative quality assessment of input multimodal data, including acoustic and visual data, without relying on any manual annotation, providing a key decision-making basis for subsequent dynamic fusion.
[0037] First, conduct a quality assessment of the acoustic data.
[0038] In the process of quality assessment of acoustic data, the segmental signal-to-noise ratio (SegSNR) algorithm is used to calculate each segment of the original acoustic waveform before preprocessing to evaluate the clarity of the signal under normal background noise, and finally outputs a scalar form of acoustic quality score. The specific calculation steps of the segmented signal-to-noise ratio algorithm are as follows.
[0039] First, the complete acoustic waveform is divided into The short frames are non-overlapping, and in this embodiment, the duration of each short frame is 20 milliseconds.
[0040] Subsequently, within each frame, the noise level is estimated using the minimum statistics method.
[0041] Each frame is further divided into several smaller segments. In this embodiment, each frame is divided into 4 segments, and the lowest energy value among all segments is taken as the noise energy estimate of the current frame. Signal energy It is then calculated by subtracting the noise energy estimate from the total energy of the frame.
[0042] Finally, by averaging the signal-to-noise ratio (SNR) of all valid frames in decibels, an acoustic quality score that stably reflects the SNR level of the entire audio segment is obtained. The calculation formula is as follows: , in, Indicates the total number of frames. This represents the signal energy of the m-th frame; This represents the noise energy estimate for the m-th frame.
[0043] Second, conduct a quality assessment of the visual data.
[0044] In the process of visual data quality assessment, an underwater image quality evaluation algorithm is used to calculate the quality of each frame of the image before preprocessing. This algorithm comprehensively evaluates the image's color, sharpness, and contrast, and outputs a scalar visual quality score by linearly weighting these three key perceptual indicators. The specific calculation process is as follows.
[0045] The Underwater Image Colorfulness Measure (UICM) is a metric used to quantify the color saturation and color cast of an image. In practice, the image is first converted from the RGB color space to the CIELab color space, and then the statistical variances of the red-green and yellow-blue channels are calculated to obtain the combined UICM value.
[0046] The Underwater Image Sharpness Measure (UISM) is a metric used to evaluate the overall sharpness and detail richness of an image. In implementation, the Sobel operator is used to calculate the edge intensity maps of the R, G, and B channels of the image, and then a weighted average is performed using the information entropy of each channel as weights to obtain the UISM value.
[0047] The Underwater Image Contrast Measure (UIConM) is a metric used to measure the overall contrast of an image. In practice, it first extracts the image's luminance channels, then calculates the UIConM value by analyzing the asymmetric distribution of its luminance and local contrast.
[0048] Finally, the three metrics UICM, UISM, and UIConM are linearly weighted and summed to obtain the final visual quality score. The calculation formula is as follows: , in, , , All are preset weighting coefficients. In this embodiment, , , .
[0049] Third, standardize the quality score.
[0050] The acoustic quality fraction obtained in the above steps and visual quality score Z-score standardization is performed using the mean and standard deviation obtained pre-calculated from a large amount of data under different operating conditions. The purpose of this step is to eliminate differences in units and numerical ranges between different quality assessment algorithms, enabling quality scores to be compared and input into subsequent networks under a unified standard. This step includes an offline statistical phase and an online computation phase.
[0051] (a) Offline statistics stage: Prepare a representative large-scale dataset containing acoustic waveforms and visual images collected under various operating conditions, such as different lighting, water turbidity, and background noise levels.
[0052] For each sample in this dataset, calculate the original acoustic quality score of each sample using the formula described above. and visual quality score .
[0053] Collect acoustic quality scores for all samples Value, calculate its global mean and global standard deviation .
[0054] Similarly, calculate the visual quality score for all samples. global mean of values and global standard deviation These four statistics ( , , , ) are stored as preset parameters.
[0055] (II) Online Calculation Stage: The parameters stored during the offline statistical phase are processed using the Z-score standardization formula: , , in, This represents the standardized acoustic quality score. This represents the standardized visual quality score. The two parameters above constitute a two-dimensional quality feature vector. This vector will be used as the input to the subsequent gating network.
[0056] The third step is to construct a quality-aware gating network based on quality perception.
[0057] The quality-aware gating network receives the processed audio-visual data streams and real-time quality feature vectors from previous steps. Internally, through the collaborative work of a series of sophisticated modules, such as feature extraction, dynamic gating weight generation, and gating cross-attention fusion, it ultimately outputs accurate fault diagnosis results. Its detailed structure and data processing flow are as follows.
[0058] First, determine the data to be input to the quality-aware gating network: The data to be input to the quality-aware gating network includes acoustic time-spectrum data, visual image sequence data, and quality feature vector data.
[0059] The input acoustic time-spectrum data is as follows: Where 1 represents the number of channels, B represents the batch, F represents the frequency of the spectrum, and T represents the time dimension of the spectrum.
[0060] The above data originates from the processing of the original acoustic waveform. First, a 1-second original acoustic waveform is loaded and its length is normalized. Next, the one-dimensional waveform is converted into a two-dimensional time-spectrum graph using a short-time Fourier transform (STFT), and the amplitude spectrum is further converted to a decibel (dB) scale to compress the dynamic range. Finally, the processed decibel spectrum is converted into a tensor conforming to the input format of a convolutional neural network.
[0061] The input visual image sequence data is Where 3 represents the red, green, and blue color channels of the image. H represents the length of the visual sequence, H represents the image height, and W represents the image width.
[0062] The above data originates from the processing of original image frames. First, a preset number of images, such as 5 frames, are uniformly sampled from the video to form a sequence. Then, each frame in the sequence is sequentially resized, converted from image format to tensor, and its values are normalized. The single-frame tensors processed through the above steps are finally stacked in the time dimension to form visual sequence data. In this embodiment, the image size is adjusted to 224*224 pixels.
[0063] The input quality feature vector data is .
[0064] Second, the extraction of acoustic and visual features.
[0065] (a) Acoustic feature extraction.
[0066] The acoustic feature extractor in this embodiment is built on the existing mature Convolutional Neural Network (CNN) and Transformer architecture, and its purpose is to extract deep time-frequency features from the acoustic time-spectrum graph. The specific processing flow for acoustic feature extraction using this acoustic feature extractor is as follows.
[0067] 1. Local time-frequency pattern extraction: Extracting the input acoustic time-frequency spectrum data First, a convolutional feature extraction module is fed in. This module contains three convolutional blocks, each consisting of a convolutional layer (3×3 convolutional kernel), a batch normalization layer, a ReLU activation function, and a max-pooling layer (2×2 pooling window) connected sequentially, with feature channels of 32, 64, and 128 respectively. This module performs a sliding scan on the two-dimensional time-frequency plane through multiple convolutional and pooling layers to capture local, fundamental acoustic patterns, such as harmonic structures and transient impulses. This process converts the original spectrogram into a sequence of feature maps. ,in Indicates the number of feature channels. This represents the frequency after dimensionality reduction. This represents the time dimension after dimensionality reduction.
[0068] 2. Serialization and Dimension Reshaping: The feature map sequence output by the CNN network is merged or flattened along the frequency dimension to form a sequence with a time step of [missing information]. The feature dimension of each time step is Feature sequences ,in , where represents the dimension of the feature vector at each time step.
[0069] 3. Global Temporal Dependency Modeling: The above feature sequence is input into a standard Transformer encoder. Through the multi-head self-attention mechanism inside the Transformer encoder, the model can capture the long-distance dependency between any two time steps in the sequence and output the sequence features, thereby understanding the evolution of the acoustic signal over the entire time axis when a fault occurs.
[0070] 4. Feature Aggregation and Output: The sequence features output by the Transformer encoder are aggregated into a single, fixed-dimensional acoustic feature vector through a time-dimensional average pooling operation. , where D_model represents the hidden layer dimension of the Transformer; this vector is the feature representation of the core information of the current acoustic signal.
[0071] (ii) Visual feature extraction.
[0072] The visual feature extractor in this embodiment employs an existing, lightweight pre-trained visual transformer model, MobileViT, designed to efficiently extract visual features representing the physical morphology of a propeller from image sequences. The specific processing flow for visual feature extraction using the MobileViT model is as follows.
[0073] 1. Single-frame feature extraction: Extracting features from the input image sequence. Split by frame. Divide each frame image... The data is independently fed into the backbone network of the MobileViT model. Inside the model, by combining the local information processing capability of convolution and the global information capture capability of Transformer, a high-level semantic feature vector is generated for each frame of image.
[0074] 2. Sequence Feature Aggregation: Collects all features in the sequence. The feature vectors generated from the frame image form a feature vector sequence. Where D_frame represents the feature dimension of a single frame output by the MobileViT backbone network; 3. Temporal Pooling and Output: Perform average pooling on the feature vector sequence obtained in the above steps to obtain a single, fixed-dimensional aggregated visual feature vector. This vector integrates information from the entire image sequence, representing the overall visual state of the propeller during that time period.
[0075] Third, the generation of dynamic gating weights.
[0076] The two-dimensional quality feature vector output from the second step The input is fed into a quality-aware gating network based on a standard multilayer perceptron architecture to generate dynamic gating weights. .
[0077] The specific processing flow of this network is as follows.
[0078] 1. Input and preprocessing.
[0079] The quality-aware gating network receives two-dimensional quality feature vectors as input layers.
[0080] 2. Nonlinear feature mapping.
[0081] The two-dimensional quality feature vector from the input layer is fed into a hidden layer containing 64 neurons, with the Modified Linear Unit (ReLU) used as the activation function. This step aims to learn a deep nonlinear relationship between the quality score and the modal reliability it represents.
[0082] 3. Generation of weight logits.
[0083] The output of the hidden layer is fed into the output layer, which has only one neuron, generating an unnormalized logit value.
[0084] 4. Normalized output.
[0085] The logit values obtained from the above steps are processed using the Sigmoid activation function. The Sigmoid function maps the logit values to the (0, 1) interval, thereby ensuring dynamic gating of the weights. It is a scalar value between 0 and 1. This weight intuitively represents the confidence level of the acoustic modality relative to the visual modality at the current moment. The output after this step is the dynamic gating weight data. .
[0086] Fourth, gating cross-attention fusion.
[0087] This step is the core of the dynamic fusion achieved in this application. By gating and weighting the standard cross-attention mechanism, dynamic and asymmetric interaction of information between modalities is realized. This module receives the acoustic feature vector output from the above steps. Visual feature vectors and dynamic gating weights Through the following two parallel information enhancement paths, the final output is an acoustic feature vector enhanced by intelligent fusion. and visual feature vectors .
[0088] In the visual feature enhancement path, high-quality acoustic information is used to enhance or supplement the expression of visual features. The specific implementation process is described below.
[0089] 1. Attention Calculation: Perform a multi-head cross-attention operation.
[0090] In this operation, visual feature vectors As a query, using acoustic feature vectors This step uses keys and values. It calculates which parts of the acoustic features the visual features should "focus on" and generates an attention-weighted acoustic information vector. The calculation formula is as follows: , , in, From visual feature vectors , and From acoustic feature vectors ; Indicates the first The query projection matrix of each attention head; Indicates the first The key projection matrix of each attention head; Indicates the first The value projection matrix of each attention head; The output projection matrix represents the visual enhancement path; This indicates the number of attention heads; in this embodiment, its value is 8. This represents the dimension of each head; in this embodiment, its value is 32. This indicates a splicing operation.
[0091] 2. Gating weighting and updating.
[0092] The attention-weighted acoustic information vector output from the previous step is multiplied by the dynamic gating weights. Then, this weighted result is added to the original visual feature vector via a residual connection, and after layer normalization, the enhanced visual feature vector is obtained. The calculation formula is as follows: , in The representation layer normalization processing function.
[0093] Through this step, when the acoustic modality has high reliability, i.e. When the value approaches 1, visual features will absorb more information from acoustics for enhancement; conversely, they will absorb less, thus avoiding interference from low-quality acoustic information.
[0094] In the acoustic feature enhancement path, high-quality visual information is used to correct or supplement acoustic features, and this path is executed in parallel with the visual enhancement path. The specific implementation process is described below.
[0095] Ⅰ. Attention Calculation: Perform the multi-head cross-attention operation once more.
[0096] acoustic feature vectors As a query, using visual feature vectors This step uses key and value pairs. It calculates which parts of the visual features the acoustic features should focus on and generates an attention-weighted visual information vector. The calculation formula is as follows: , , in, From acoustic feature vectors ; and From visual feature vectors ; Indicates the first The query projection matrix of each attention head; Indicates the first The key projection matrix of each attention head; Indicates the first The value projection matrix of each attention head; The output projection matrix represents the acoustic enhancement path; This indicates the number of attention heads; in this embodiment, its value is 8. This represents the dimension of each head; in this embodiment, its value is 32. This indicates a splicing operation.
[0097] II. Gating weighting and updating.
[0098] The attention-weighted visual information vector obtained in the previous step is combined with... Multiply. Then, similarly, combine with the original acoustic feature vector via residual concatenation. The features are added together and then normalized to obtain the enhanced acoustic feature vector. The calculation formula is as follows: .
[0099] This asymmetric gating weight ensures that when the visual modality has high credibility, i.e. Approaching 0 When the value approaches 1, the acoustic features can absorb more information from the vision for correction.
[0100] Through the aforementioned quality-score-driven, asymmetric gated cross-attention mechanism, intelligent and bidirectional enhancement of intermodal information is achieved. This allows high-quality modalities to dominate the information interaction process while effectively suppressing potential "noise" interference from low-quality modalities. The final output is an enhanced feature vector that has undergone deep fusion and optimization. and .
[0101] Fifth, final feature generation and multi-task prediction, which specifically includes the following steps.
[0102] (a) Final feature generation.
[0103] The enhanced acoustic feature vector obtained after gated cross-attention fusion and visual feature vectors Utilizing dynamic gating weights again We perform a weighted summation to obtain the fusion features ultimately used for the main task decision. The calculation formula is as follows: .
[0104] (ii) Multi-task prediction.
[0105] To enhance the stability and effectiveness of model training and ensure that each single-modal feature extractor receives adequate supervision and training, this application employs a multi-task learning strategy. This strategy inputs different feature vectors into three parallel classification prediction heads, each of which is a standard feedforward neural network module. This feedforward neural network module typically consists of multiple linear layers, activation functions, and dropout layers, and its structure is designed to map high-dimensional features to the final fault category probability.
[0106] The three prediction heads are the main classification prediction head, the acoustic auxiliary prediction head, and the visual auxiliary prediction head. The data processing procedure for each prediction head is as follows.
[0107] The input to the main classification prediction head is the fused features. The main classification prediction head is a multilayer perceptron (MLP) classifier. The main classification prediction head first processes the input features through one or more fully connected linear layers for dimensionality transformation and non-linear mapping; in this embodiment, it maps from 256 dimensions to 128 dimensions. Each linear layer is typically followed by a ReLU activation function to increase the model's non-linear expressive power. To prevent overfitting, a Dropout layer is added to the network for regularization, for example, randomly deactivating some neurons with a probability of 0.1. The last linear layer maps the feature dimensions to a dimension equal to the number of fault categories. In this embodiment, the fault categories are divided into seven categories: healthy, minor injury, severe injury, 3mm rope entanglement, propeller loss, fishing net entanglement, and plastic bag entanglement. Finally, without passing through an activation function, the raw prediction score for each fault category is directly output. When calculating the loss or making the final prediction, this score is converted into a probability distribution for each category using a Softmax function.
[0108] The input to the acoustic-assisted prediction head is an unfused acoustic feature vector. The internal structure of the acoustic-assisted prediction head is similar to that of the main classification prediction head; it is also an independent multilayer perceptron classifier. It receives acoustic feature vectors. The algorithm processes data through a series of linear layers, ReLU activation functions, and Dropout layers, ultimately outputting the raw score corresponding to the number of fault categories. In this embodiment, the Dropout layer of the acoustic-assisted prediction head typically has a slightly higher deactivation rate, such as 0.2, to enhance the regularization effect. Finally, the raw prediction score of the fault category based on pure acoustic features is output.
[0109] The input to the visual-assisted prediction head is an unfused visual feature vector. The visual-assisted prediction head is completely equivalent to the acoustic-assisted prediction head; it is also an independent multilayer perceptron classifier that receives visual feature vectors. The same computational process is performed, namely linear layers, ReLU activation, and Dropout layers. Finally, the raw prediction score of the fault category based on pure visual features is output.
[0110] The fourth step is model training and output. The specific implementation process is described below.
[0111] First, the loss function is calculated.
[0112] During the training phase, a multi-task learning strategy incorporating both main and auxiliary losses is employed, with the total loss... The calculation formula is: , in, This represents the main classification loss, which incorporates the fused features. The prediction result is obtained by inputting the main prediction head, and then the cross-entropy is calculated with the true label. Represents the acoustic-assisted classification loss, which converts the unfused acoustic feature vectors... The prediction result is obtained by inputting the acoustic-assisted prediction head, and then the cross-entropy is calculated with the real label to obtain the result. The visual-assisted classification loss represents the loss of the unfused visual feature vectors. The prediction result is obtained by inputting the visual aid prediction head, and then the cross-entropy is calculated with the real label. All three losses mentioned above are calculated using the standard cross-entropy loss function.
[0113] This represents a pre-defined hyperparameter, set to 0.4 in this embodiment, used to balance the weights of the auxiliary loss. This design allows for direct supervision of the single-modal feature extraction path, ensuring the stability and effectiveness of model training.
[0114] Second, the output of propeller fault diagnosis categories.
[0115] During the inference phase, the output of the master classification prediction head based on fused features is used as the final and most reliable fault diagnosis category. The final output of this method is as follows: Figure 2 As shown.
[0116] This application also discloses an intelligent fault diagnosis system capable of implementing the above method, the system comprising a data acquisition module, a quality assessment module, a feature extraction module, a gating fusion module, and a classifier module.
[0117] The data acquisition module is used to collect acoustic signals and visual image sequences during propeller operation and to preprocess the collected data.
[0118] The quality assessment module is used to perform real-time quality quantification assessment of preprocessed acoustic and visual data, generating standardized quality feature vectors.
[0119] The feature extraction module is used to extract depth features from the preprocessed acoustic time-spectrum map and visual image sequence, and output acoustic feature vectors and visual feature vectors.
[0120] The gating fusion module is used to generate dynamic gating weights based on the quality feature vector and to achieve intelligent fusion of acoustic and visual features through a gating cross-attention mechanism.
[0121] The classifier module is used to output the propeller fault type prediction result by passing the fused feature vector through the main classification prediction head and the auxiliary prediction head.
[0122] During the experiments conducted for this application, acoustic signals under various fault conditions were collected using a hydrophone, such as... Figure 3 As shown. Underwater optical images under various malfunctions captured by an underwater camera, such as... Figure 4 As shown, the propeller fault diagnosis confusion matrix was obtained during experiments in clean water. This method achieves an accuracy of 98.1% in diagnosing propeller fault types. Figure 5 As shown. Under different environmental conditions, the robustness of the method proposed in this application compared with existing propeller fault diagnosis methods is as follows: Figure 6 to Figure 11 As shown in the figure. Through comparative experiments, it can be seen that under different environments, this application can perceive and quantify the quality of input data in real time. When the quality of any modality data deteriorates, it can adaptively reduce its weight in the fusion process, effectively avoiding interference from low-quality or erroneous information. Therefore, it can still maintain stable diagnostic performance in harsh environments such as noise, low light, and occlusion, and has high robustness.
[0123] The multimodal propeller fault diagnosis method and system provided by this invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from the principles of this invention, and these improvements and modifications also fall within the protection scope of the claims of this invention. The above description of the disclosed embodiments enables those skilled in the art to implement or use this invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this invention. Therefore, this invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-mode propeller fault diagnosis method, characterized in that, Includes the following steps: S1. Multimodal data acquisition and preprocessing: S2. Real-time quantification of the quality of acoustic signal data and visual signal data obtained in step S1; S3. Construct a quality-aware gating network model; S4. Train the quality-aware gating network model and use the trained quality-aware gating network model to output the propeller fault type.
2. The multi-mode propeller fault diagnosis method according to claim 1, characterized in that, In step S1, hydrophones deployed near the thruster are used to collect real-time or offline acoustic signals during propeller operation, and the collected acoustic signal data is preprocessed: The acquired raw acoustic waveform data is segmented and converted into a two-dimensional time-spectrum graph using short-time Fourier transform. The time-spectrum graph is then resized and normalized. The underwater camera deployed near the thruster synchronously acquires real-time or offline image sequences of the propeller's working area, and preprocesses the acquired visual signal data: The size of the acquired image frames is uniformly adjusted, the images are converted into tensor format, and then standardized.
3. The multi-mode propeller fault diagnosis method according to claim 1, characterized in that, In step S2, the acoustic signal data and visual signal data are evaluated for quality to obtain an acoustic quality score. and visual quality score ; The specific implementation process for quality assessment of acoustic signal data is as follows: S2.1.A.
1. Divide the complete acoustic waveform signal into M non-overlapping short time frames; S2.1.A.
2. Within each frame, the noise level is estimated using the minimum statistical method: each frame is further divided into several sub-segments, and the lowest energy value among all sub-segments is taken as the noise energy estimate for the current frame. Signal energy It is then calculated by subtracting the noise energy estimate from the total energy of the frame; S2.1.A.
3. By averaging the signal-to-noise ratio (SNR) of all valid frames in decibels, an acoustic quality score that stably reflects the SNR level of the entire audio segment is obtained. The calculation formula is as follows: , in, Indicates the total number of frames; This represents the signal energy of the m-th frame; This represents the noise energy estimate for the m-th frame; The specific implementation process for quality assessment of visual data is as follows: S2.1.B.1, Evaluate the color measurement index UICM, sharpness measurement index UISM, and contrast measurement index UIConM of the image; S2.1.B.
2. Perform a linear weighted sum of the three indicators UICM, UISM, and UIConM to obtain the final visual quality score. The calculation formula is as follows: in, , , All of these are preset weighting coefficients. , , .
4. The multi-mode propeller fault diagnosis method according to claim 3, characterized in that, The obtained acoustic quality score and visual quality score Standardization processes are performed, including offline and online statistical phases. The offline statistical phase includes the following steps: S2.2.
1. The acoustic waveforms and visual images collected under different operating conditions are combined into a dataset. For each sample in this dataset, the original acoustic quality score of each sample is calculated. and visual quality score ; S2.2.2 Collect the acoustic quality scores of all samples. Value, calculate its global mean and global standard deviation ; Calculate the visual quality score for all samples. global mean of values and global standard deviation ; The above four statistical measures , , , Stored as preset parameters; The steps involved in the online statistics phase include: S2.2.
3. Utilize the parameters stored in the offline statistical phase and process them using standardized formulas: , , in, This represents the standardized acoustic quality score. The standardized visual quality score is represented by the two parameters mentioned above, which together form a two-dimensional quality feature vector. .
5. The multi-mode propeller fault diagnosis method according to claim 1, characterized in that, The specific implementation process of step S3 is as follows: S3.1 Determine the data for input to the quality-aware gating network: The input data includes acoustic time-spectrum data, visual image sequence data, and quality feature vector data; The input acoustic time-spectrum data is (Tensor [B, 1, F, T]), where 1 represents the number of channels, B represents the batch, F represents the frequency of the spectrum, and T represents the time dimension of the spectrum. The input visual image sequence data is Where 3 represents the red, green, and blue color channels of the image. H represents the length of the visual sequence, and W represents the image height. The input quality feature vector data is ; S3.2 Extraction of acoustic and visual features; S3.3 Generation of dynamic gating weights; S3.4, Gated Cross-Attention Fusion; S3.5, Final Feature Generation and Multi-Task Prediction.
6. The multi-mode propeller fault diagnosis method according to claim 5, characterized in that, The extraction of acoustic features includes the following steps: S3.
2. A.1, Local Time-Frequency Pattern Extraction: Extracting the input acoustic time-frequency pattern data... The input is a CNN network that performs a sliding scan across a two-dimensional time-frequency plane through several convolutional and pooling layers to capture local, fundamental acoustic patterns and output a sequence of feature maps. ,in Indicates the number of feature channels. This represents the frequency after dimensionality reduction. This represents the time dimension after dimensionality reduction; S3.
2. A.2, Serialization and Dimension Reshaping: Feature Map Sequence Merging or flattening along the frequency dimension to form a time step of The feature dimension of each time step is Feature sequences ,in , representing the dimension of the feature vector at each time step; S3.2.A.3, Global Temporal Dependency Modeling: Modeling feature sequences... Input the Transformer encoder, capture the long-distance dependency between any two time steps in the sequence through the multi-head self-attention mechanism inside the Transformer encoder, and output the sequence features; S3.2.A.4 Feature Aggregation and Output: The sequence features output by the Transformer encoder are aggregated into a single, fixed-dimensional acoustic feature vector through a time-dimensional average pooling operation. ,in This represents the hidden layer dimension of the Transformer; The extraction of visual features includes the following steps: S3.2.B.1 Single-frame feature extraction: Extracting the input image sequence Split by frame, and divide each frame image The data is independently fed into the backbone network of the MobileViT model to generate a high-level semantic feature vector for each frame of the image. S3.2.B.2, Sequence Feature Aggregation: Collecting all features in the sequence. The feature vectors generated from the frame image form a feature vector sequence. Where D_frame represents the feature dimension of a single frame output by the MobileViT backbone network; S3.2.B.3, Temporal Pooling and Output: The feature vector sequence obtained in the above steps is subjected to average pooling in the temporal dimension to obtain a single, fixed-dimensional aggregated visual feature vector. ; Dynamic gating weight generation includes the following steps: S3.3.1 Input and Preprocessing: The quality-aware gating network receives a two-dimensional quality feature vector as the input layer; S3.3.2 Nonlinear Feature Mapping: The two-dimensional quality feature vector of the input layer is fed into a hidden layer containing 64 neurons, and the modified linear unit ReLU is used as the activation function; S3.3.3, Weight logits generation: The output of the hidden layer is fed into the output layer with only one neuron to generate an unnormalized weight value logit; S3.3.4 Normalized Output: The logit values of the weights are processed by the Sigmoid activation function. The Sigmoid function maps the logit values to the interval (0, 1), thus obtaining the dynamic gating weights. The output of this step is the dynamic gating weight data. .
7. The multi-mode propeller fault diagnosis method according to claim 6, characterized in that, In step S3.4, the acoustic feature vector, visual feature vector, and dynamic gating weights are received. Through parallel visual feature enhancement paths and acoustic feature enhancement paths, the acoustic feature vector after intelligent fusion enhancement is output. and visual feature vectors ; The specific processing steps for the visual feature enhancement path are as follows: S3.4.A.1, using visual feature vectors As a query, using acoustic feature vectors Using these as keys and values, the acoustic features that the visual features should focus on are calculated, and an attention-weighted acoustic information vector is generated. The calculation formula is as follows: , , in, From visual feature vectors , and From acoustic feature vectors ; Indicates the first The query projection matrix of each attention head; Indicates the first The key projection matrix of each attention head; Indicates the first The value projection matrix of each attention head; The output projection matrix represents the visual enhancement path; Indicates the number of heads of attention; Indicates the dimension of each head; Indicates a splicing operation; S3.4.A.
2. Multiply the attention-weighted acoustic information vector with the dynamic gating weights; add this weighted result to the original visual feature vector through a residual connection, and then perform layer normalization to obtain the enhanced visual feature vector. The calculation formula is as follows: , in Indicates the layer normalization processing function; The specific processing steps for the acoustic feature enhancement path are as follows: S3.4.B.1, using acoustic feature vectors As a query, using visual feature vectors Using these as keys and values, the visual features that the acoustic features should focus on are calculated, and an attention-weighted visual information vector is generated. The calculation formula is as follows: , , in, From acoustic feature vectors , and From visual feature vectors ; Indicates the first The query projection matrix of each attention head; Indicates the first The key projection matrix of each attention head; Indicates the first The value projection matrix of each attention head; The output projection matrix represents the acoustic enhancement path; Indicates the number of heads of attention; Indicates the dimension of each head; Indicates a splicing operation; S3.4.B.2, Attention-weighted visual information vectors and Multiplication; combined with the original acoustic feature vector via residual concatenation. The features are added together and then normalized to obtain the enhanced acoustic feature vector. The calculation formula is as follows: 。 8. The multi-mode propeller fault diagnosis method according to claim 7, characterized in that, The specific implementation process of step S3.5 is as follows: S3.5.1 Final Feature Generation: The enhanced acoustic feature vector obtained after gated cross-attention fusion and visual feature vectors Utilizing dynamic gating weights again We perform a weighted summation to obtain the fusion features ultimately used for the main task decision. The calculation formula is as follows: , S3.5.2 Multi-task prediction: A multi-task learning strategy is adopted, in which different feature vectors are input into three parallel classification prediction heads. The three prediction heads are the main classification prediction head, the acoustic auxiliary prediction head, and the visual auxiliary prediction head. The data processing process of each prediction head is as follows: The input to the main classification prediction head is the fused features. First, the feature is transformed and non-linearly mapped through several fully connected linear layers. Each linear layer is followed by a ReLU activation function. The last linear layer maps the feature dimension to a dimension equal to the number of fault categories and directly outputs the original prediction score of the fault category. The input to the acoustic-assisted prediction head is an unfused acoustic feature vector. The system processes data through a linear layer, a ReLU activation function, and a Dropout layer, and outputs a raw prediction score for the fault category based on pure acoustic features. The input to the visual-assisted prediction head is an unfused visual feature vector. The system processes data through a linear layer, a ReLU activation function, and a Dropout layer, outputting a raw prediction score for the fault category based on pure visual features.
9. The multi-mode propeller fault diagnosis method according to claim 1, characterized in that, In step S4, during the training of the quality-aware gated network model, a multi-task learning strategy including main loss and auxiliary loss is adopted, and its total loss is... The calculation formula is: , in, This represents the main classification loss, which incorporates the fused features. The prediction result is obtained by inputting the main prediction head, and then the cross-entropy is calculated with the true label. Represents the acoustic-assisted classification loss, which converts the unfused acoustic feature vectors... The prediction result is obtained by inputting the acoustic-assisted prediction head, and then the cross-entropy is calculated with the real label to obtain the result. The visual-assisted classification loss represents the loss of the unfused visual feature vectors. The prediction result is obtained by inputting the visual aid prediction head, and then the cross-entropy is calculated with the real label to obtain the result. This represents the pre-defined hyperparameters.
10. A fault diagnosis system for implementing the multi-mode propeller fault diagnosis method according to any one of claims 1-9, characterized in that, include: The data acquisition module is used to collect acoustic signals and visual image sequences during propeller operation and to preprocess the collected data. The quality assessment module is used to perform real-time quality quantification assessment of preprocessed acoustic and visual data, and generate standardized quality feature vectors. The feature extraction module is used to extract depth features from the preprocessed acoustic time-spectrum map and visual image sequence, and output acoustic feature vectors and visual feature vectors. The gated fusion module is used to generate dynamic gate weights based on the quality feature vector and to achieve intelligent fusion of acoustic and visual features through a gated cross-attention mechanism. The classifier module is used to output the propeller fault type prediction result by passing the fused feature vector through the main classification prediction head and the auxiliary prediction head.
Citation Information
Patent Citations
Autonomous underwater vehicle (AUV) adaptive fault diagnosis method based on discriminative feature learning method
CN110244689A
AUV propeller multi-source fusion fault diagnosis method and system based on deep learning
CN114544155A
Underwater propeller fault diagnosis method and system
CN116417013A
Underwater robot fault diagnosis method based on multi-channel full convolutional neural network
CN116933173A
Data aggregation method based on multi-modal features
CN120372555A
Cited By
Multi-branch fault detection method and system for electric energy metering assembly line
CN121980247A
A method and system for multi-branch fault detection in power metering pipelines
CN121980247B
Virtual reality navigation rehabilitation training system and method
CN122117239A
Intelligent monitoring method for fault diagnosis and prediction of underwater equipment
CN122221185A