Fault diagnosis method for mechanical system of water jet propulsion device based on time-frequency prototype network

By using a time-frequency prototype network-based method to generate RGB time-frequency maps and combining them with a deep learning model, the problem of lack of interpretability in the fault diagnosis of water jet propulsion devices is solved, and highly reliable fault diagnosis and visual monitoring are achieved.

CN121990136APending Publication Date: 2026-05-08HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN INST OF TECH
Filing Date
2026-01-29
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing deep learning methods lack process interpretability and post-event interpretability in the fault diagnosis of water jet propulsion devices, making it difficult to achieve automated real-time monitoring, and users cannot intuitively see the diagnostic basis of the model.

Method used

A time-frequency prototype network-based approach is adopted to generate RGB time-frequency maps through sensor data acquisition and time-frequency transformation. Combined with a deep feature extractor and a prototype decision layer, a diagnostic model is constructed to achieve visualization and credibility verification of fault diagnosis.

Benefits of technology

The process of fault diagnosis for water jet propulsion devices has been made interpretable and highly reliable, providing intuitive visualization of physical characteristics and logical support based on historical cases, thus avoiding misdiagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121990136A_ABST
    Figure CN121990136A_ABST
Patent Text Reader

Abstract

The invention provides a water jet propulsion device mechanical system fault diagnosis method based on a time-frequency prototype network, and the method comprises the steps: carrying out the data collection and time-frequency transformation through a sensor measurement point, and obtaining a single-channel time-frequency matrix; obtaining a fused RGB time-frequency graph by using an RGB mapping rule, and carrying out image processing on the fused RGB time-frequency graph to obtain a standard input sample; performing staged training on the constructed initial model to obtain a diagnosis model; and performing reasoning diagnosis and verification on the standard input sample by using the diagnosis model to obtain a fault diagnosis result, a credibility basis and a high-resolution evidence. According to the invention, accurate, efficient and interpretable fault diagnosis can be realized, visual monitoring of the fault diagnosis process and case-based reasoning type afterwards attribution of the diagnosis result can be realized, and a high-precision and high-reliability fault diagnosis scheme is provided for intelligent operation and maintenance of ships.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent ship operation and maintenance and fault diagnosis technology, and in particular to a fault diagnosis method for a waterjet propulsion system based on a time-frequency prototype network. Background Technology

[0002] The waterjet propulsion system is a complex power and control system involving multi-physics coupling. Based on its functional structure and signal characteristics, it mainly consists of three parts: a control system, a hydraulic system, and a mechanical system. The control system includes a main control circuit, a signal conditioning circuit, and a power drive circuit, responsible for command processing and signal amplification. The hydraulic system, composed of a hydraulic pump, proportional directional valves, and hydraulic cylinders, is responsible for driving the steering and reversing mechanisms. The mechanical system, the core power transmission unit, mainly consists of an impeller, guide vanes, and a shaft system, and is subjected to long-term fluid vibration and mechanical friction.

[0003] Because ships operate in complex marine environments, the mechanical systems of waterjet propulsion systems are constantly subjected to the coupled effects of fluid excitation, unsteady cavitation loads, and mechanical friction. Common failure modes include impeller blade damage, abnormal blade tip clearance, and bearing wear. If these faults are not detected and addressed in a timely manner, they can lead to a significant decrease in propulsion efficiency and, in severe cases, even catastrophic mechanical failures. Therefore, accurate, efficient, and interpretable fault diagnosis is crucial for ensuring the safe operation of waterjet propulsion systems.

[0004] Currently, traditional fault diagnosis methods mainly rely on signal processing techniques, requiring professional vibration analysts to manually interpret spectral characteristics, which is inefficient and difficult to automate in real time. In recent years, with the development of deep learning technology, convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have been widely used in mechanical fault diagnosis, achieving significant breakthroughs in fault classification accuracy.

[0005] However, existing deep learning diagnostic methods still face severe challenges in practical shipbuilding engineering applications, mainly in the following two aspects: (1) Lack of process interpretability, i.e., serious "black box" problem. Traditional deep neural networks map vibration signals into abstract feature vectors, and maintenance personnel cannot intuitively see which part of the signal the model is based on for judgment. In waterjet propulsion systems, fluid noise often masks weak mechanical fault characteristics. If the model only fits the background noise rather than the fault impact characteristics, it will lead to serious false alarms or missed alarms, and users will not be able to detect it during the diagnostic process. (2) Lack of post-event interpretability and case support. When the model outputs an alarm of "major impeller damage", it often only gives a confidence probability and cannot provide physical evidence. For decision-makers, such a lack of evidence makes it difficult to use as a basis for shutdown maintenance decisions. Summary of the Invention

[0006] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a fault diagnosis method for mechanical systems of waterjet propulsion devices based on time-frequency prototype networks. This method can achieve accurate, efficient, and interpretable fault diagnosis, as well as visualize the fault diagnosis process and perform case-based reasoning-based post-hoc attribution of diagnostic results. This provides a high-precision and high-reliability fault diagnosis solution for intelligent ship maintenance.

[0007] To achieve the above objectives, the present invention provides the following solution: a fault diagnosis method for a waterjet propulsion device mechanical system based on a time-frequency prototype network, comprising: Data is collected and transformed using sensor points set on the mechanical system of the water jet propulsion device to obtain a single-channel time-frequency matrix, and a spatial coordinate system is defined for physical mapping. Using the defined RGB mapping rules, the single-channel time-frequency matrix is ​​transformed into a fused RGB time-frequency image, and then the fused RGB time-frequency image is processed to obtain a standard input sample. A state dataset including five health states is collected at the sensor measurement points. The initial model is then trained in stages using the state dataset to obtain a diagnostic model. The diagnostic model is used to perform inference diagnosis on the standard input sample to obtain fault diagnosis results and credibility basis, and the fault diagnosis results are verified to obtain high-resolution evidence.

[0008] Optionally, data is acquired and time-frequency transformed using sensor points installed on the water jet propulsion system to obtain a single-channel time-frequency matrix, and a spatial coordinate system is defined for physical mapping, including: In a mechanical system based on a water jet propulsion device, sensor measuring points are set on the housing of the first-stage impeller, the housing of the pump bearing, the housing of the second-stage impeller, and the housing of the main shaft bearing. Data is collected at these sensor measuring points to obtain multi-channel raw data. For a mechanical system based on a water jet propulsion device, the direction parallel to the center line of the propulsion pump main shaft and pointing towards the water jet is defined as the X-axis, the direction from the sensor mounting point and passing through the geometric center of the pump body is defined as the Y-axis, and the direction along the tangent of the casing is defined as the Z-axis, thus obtaining a spatial coordinate system for physical mapping. Based on the multi-channel raw data, using continuous wavelet transform based on complex Morlet wavelets, and through scaling and translation operations, the time window is automatically narrowed in the high-frequency band to capture the impact position at the microsecond level, and the time window is automatically widened in the low-frequency band to refine the frequency components, thus obtaining a single-channel time-frequency matrix.

[0009] Optionally, using the defined RGB mapping rules, the single-channel time-frequency matrix is ​​transformed into a fused RGB time-frequency image, and then the fused RGB time-frequency image is subjected to image processing to obtain standard input samples, including: Set the defined Z-axis to the R channel, the Y-axis to the G channel, and the X-axis to the B channel to complete the definition of the RGB mapping rule; Based on the single-channel time-frequency matrix, the R channel is used to capture the rotational centrifugal force and lateral ramming characteristics, the G channel is used to capture the radial impact and the stiffness change of the supporting structure, and the B channel is used to capture the axial thrust fluctuation caused by fluid-structure interaction, so as to obtain a fused RGB time-frequency map. Based on the fused RGB time-frequency image, the size is unified, and for each acquisition time period, a set of multidimensional image tensors containing spatial and time-frequency features is generated to obtain standard input samples.

[0010] Optionally, a state dataset including five health states is collected at the sensor measurement points. This state dataset is then used to train the initial model in stages to obtain a diagnostic model, including: By combining a deep feature extractor and a designed prototype decision layer, an initial model is obtained. The deep feature extractor is used to map the standard input samples into a high-dimensional feature map. Based on the high-dimensional feature map, the prototype decision layer is used to calculate the Euclidean distance between the input feature block and each prototype vector to complete the initial inference. Based on the sensor measurement points, state datasets of five health states, including normal state, blade tip gap too small, blade tip gap too large, impeller major damage, and impeller minor damage, are collected under different working conditions. The state datasets are randomly shuffled and divided into training set and test set in a 7:3 ratio. The design includes a hybrid loss function comprising a classification accuracy loss function, a clustering loss function, and a separation loss function. The initial model is trained in stages using the training set, the test set, and the hybrid loss function to obtain a diagnostic model.

[0011] Optionally, the prototype decision layer is used to preset multiple learnable prototype vectors and assign multiple learnable prototype vectors to each type of fault; wherein each learnable prototype vector corresponds to a potential image patch in the standard input sample.

[0012] Optional, phased training includes: The weights of the deep feature extractor are frozen, the prototype decision layer and the fully connected layer are trained, and the prototype vector is initialized using the Adam optimizer and a preset learning rate to complete the warm-up training phase. The thawed deep feature extractor is jointly trained with the prototype decision layer, and a learning rate decay strategy is introduced until the hybrid loss function converges, thus completing the joint fine-tuning training phase. The learnable prototype vector is replaced with the nearest real image patch in the training set to complete the prototype projection training phase.

[0013] Optionally, the diagnostic model is used to perform inference diagnosis on the standard input sample to obtain fault diagnosis results and credibility basis, and the fault diagnosis results are verified to obtain high-resolution evidence, including: The standard input sample is input into the diagnostic model, and feature mapping, Euclidean distance measurement and distance transformation are performed to obtain the category conditional probability. Then, based on the category conditional probability, the fault diagnosis result and confidence basis are output. Based on the fault diagnosis results and the credibility criteria, pixel-level weighting of the hierarchical gradient is performed using hierarchical activation mapping to generate activation maps for each layer. The activation maps of different resolutions are interpolated, aligned to the input size, and fused and superimposed to obtain a multi-scale fused high-resolution heatmap. Then, the high-resolution heatmap is physically consistent to obtain high-resolution evidence.

[0014] Optionally, the physical consistency verification includes fishbone frequency verification and RGB channel coupling verification.

[0015] This invention discloses the following technical effects by providing a fault diagnosis method for the mechanical system of a waterjet propulsion device based on a time-frequency prototype network: 1. A process-interpretable feature learning strategy based on physical semantic mapping was constructed. Compared to the traditional "black box" model that inputs abstract one-dimensional signals, this method endows the data with clear physical meaning at the input end. Through continuous wavelet transform (CWT) and RGB mapping strategy of orthogonal physical channels, the originally invisible spatial coupling vibration of the water jet propulsion device is transformed into intuitive color textures. This feature imaging method allows maintenance personnel to intuitively perceive the physical location and dynamic characteristics of the fault through image color before diagnosis, realizing the first layer of process interpretation from signal processing to physical feature visualization.

[0016] 2. Achieved logically transparent diagnosis based on prototype matching. Addressing the shortcomings of traditional deep learning, which only outputs probability values ​​and lacks decision-making basis, this method utilizes the prototype matching mechanism of Mech-ProtoNet to establish a reliable diagnostic logic based on historical precedents. The model's decision-making uses the most similar real historical prototype slice retrieved from the feature space as the cause of the fault, providing traceable logical support for fault diagnosis and significantly improving the reliability of the diagnostic results.

[0017] 3. A post-hoc interpretability verification method based on high-resolution visual source tracing is provided. To verify whether the diagnostic model's decisions conform to the physical fault mechanism, this method introduces the LayerCAM algorithm to achieve pixel-level fault feature localization. Compared with the fuzzy localization of traditional gradient interpretation methods, this method can generate detailed multi-scale heatmaps, clearly outlining the weak higher harmonics and transient impact textures in the vibration signal. This allows experts to verify the model's focus from a post-hoc perspective, effectively avoiding misdiagnosis caused by the model fitting background noise.

[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of the method flow provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the multi-channel fusion and imaging framework provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of a prototype network for fault diagnosis of a water jet propulsion device provided in an embodiment of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] Example 1 like Figure 1 As shown, this invention provides a fault diagnosis method for the mechanical system of a waterjet propulsion device based on a time-frequency prototype network, comprising: Step 1: Data acquisition and time-frequency transformation are performed using sensor measurement points set on the mechanical system of the water jet propulsion device to obtain a single-channel time-frequency matrix, and a spatial coordinate system is defined for physical mapping; specifically including: 1.1 A mechanical system based on a water jet propulsion device, wherein sensor measuring points are set on the housing of the first-stage impeller, the housing of the pump bearing, the housing of the second-stage impeller, and the housing of the main shaft bearing, respectively, and data is collected at the sensor measuring points to obtain multi-channel raw data; 1.2 For a mechanical system based on a water jet propulsion device, the direction parallel to the center line of the propulsion pump shaft and pointing towards the water jet is defined as the X-axis, the direction from the sensor mounting point and passing through the geometric center of the pump body is defined as the Y-axis, and the direction along the tangent of the casing is defined as the Z-axis, thus obtaining a spatial coordinate system for physical mapping. 1.3 Based on the aforementioned multi-channel raw data, using continuous wavelet transform based on complex Morlet wavelets, and through scaling and translation operations, the time window is automatically narrowed in the high-frequency band to capture microsecond-level impact positions, and the time window is automatically widened in the low-frequency band to refine the frequency components, thus obtaining a single-channel time-frequency matrix.

[0024] Step 2: Using the defined RGB mapping rules, the single-channel time-frequency matrix is ​​converted into a fused RGB time-frequency image. Then, the fused RGB time-frequency image undergoes image processing to obtain standard input samples. Specifically, this includes: 2.1 Set the defined Z-axis as the R channel, the Y-axis as the G channel, and the X-axis as the B channel to complete the definition of the RGB mapping rule; 2.2 Based on the single-channel time-frequency matrix, the R channel is used to capture the rotational centrifugal force and lateral ramming characteristics, the G channel is used to capture the radial impact and the stiffness change of the supporting structure, and the B channel is used to capture the axial thrust fluctuation caused by fluid-structure interaction, so as to obtain a fused RGB time-frequency map; 2.3 Based on the fused RGB time-frequency image, the size is unified, and for each acquisition time period, a set of multidimensional image tensors containing spatial and time-frequency features is generated to obtain standard input samples.

[0025] Step 3: Collect a state dataset including five health states at the sensor measurement points. Use the state dataset to train the initial model in stages to obtain a diagnostic model; specifically including: 3.1 Combining a deep feature extractor and a designed prototype decision layer, an initial model is obtained. The deep feature extractor is used to map the standard input sample into a high-dimensional feature map. Based on the high-dimensional feature map, the prototype decision layer is used to calculate the Euclidean distance between the input feature block and each prototype vector to complete the initial inference. The prototype decision layer is used to preset multiple learnable prototype vectors and assign multiple learnable prototype vectors to each type of fault. Each learnable prototype vector corresponds to a potential image block in the standard input sample.

[0026] 3.2 Based on the sensor measurement points, collect state datasets of five health states under different working conditions, including normal state, blade tip gap too small, blade tip gap too large, impeller major damage and impeller minor damage. After randomly shuffling the state datasets, divide them into training set and test set in a 7:3 ratio. 3.3 A hybrid loss function is designed, which includes a classification accuracy loss function, a clustering loss function, and a separation loss function. The initial model is trained in stages using the training set, the test set, and the hybrid loss function to obtain a diagnostic model.

[0027] Phased training includes: The weights of the deep feature extractor are frozen, the prototype decision layer and the fully connected layer are trained, and the prototype vector is initialized using the Adam optimizer and a preset learning rate to complete the warm-up training phase. The thawed deep feature extractor is jointly trained with the prototype decision layer, and a learning rate decay strategy is introduced until the hybrid loss function converges, thus completing the joint fine-tuning training phase. The learnable prototype vector is replaced with the nearest real image patch in the training set to complete the prototype projection training phase.

[0028] Step 4: Utilize the diagnostic model to perform inference diagnosis on the standard input sample, obtaining fault diagnosis results and credibility criteria, and verify the fault diagnosis results to obtain high-resolution evidence. Specifically, this includes: 4.1 Input the standard input sample into the diagnostic model, perform feature mapping, Euclidean distance measurement and distance transformation to obtain the category conditional probability, and then output the fault diagnosis result and confidence basis according to the category conditional probability; 4.2 Based on the fault diagnosis results and the credibility criteria, pixel-level weighting of the hierarchical gradient is performed using hierarchical activation mapping to generate activation maps for each layer. Activation maps of different resolutions are interpolated, aligned to the input size, and fused and superimposed to obtain a multi-scale fused high-resolution heatmap. Then, physical consistency verification is performed on the high-resolution heatmap to obtain high-resolution evidence. The physical consistency verification includes basalt frequency verification and RGB channel coupling verification.

[0029] Example 2 This invention proposes a fault diagnosis method for waterjet propulsion devices that integrates multi-channel time-frequency imaging and an interpretable prototype network. The method uses a time-frequency prototype network as its core, converting high-frequency vibration signals into time-frequency spectra and matching them with fault prototypes based on similarity, thus achieving transparency and case-based diagnostic decisions. This method aims to address the serious "black box" problem, lack of physical feature localization capability, and lack of historical case support in existing deep learning models for waterjet propulsion device diagnosis, providing a high-precision and high-reliability fault diagnosis solution for intelligent ship maintenance. The overall process includes three core steps: quality enhancement and time-frequency imaging of multi-channel vibration signals, training and prototype learning of the time-frequency prototype network, and visualization verification and case attribution of diagnostic results. It is applicable to the mechanical system condition monitoring and intelligent diagnosis of various waterjet propulsion devices.

[0030] This invention first collects vibration signals from key components such as the main shaft bearing of a waterjet propulsion device using multi-channel sensors. These signals are then converted into time-frequency spectra using continuous wavelet transform, forming the input for the model. Next, a time-frequency prototype network is trained to learn time-frequency prototypes of typical faults, and a prototype matching mechanism is used to classify and visualize the faults. The specific implementation steps of the method are as follows: Step 1: Multichannel fusion and imaging Data acquisition experiments were conducted on the mechanical system of a ship's power plant. The layout and number of sensors are shown in the figure. Based on the sensor layout of the waterjet propulsion system, the system includes four key measuring points: the first-stage impeller housing, the pump bearing housing, the second-stage impeller housing, and the main shaft bearing housing. To accurately decouple vibration characteristics in different dimensions, this invention establishes a spatial mapping rule based on the equipment's coordinate system, specifically defined as follows: X-axis (axial direction): Parallel to the centerline of the propulsion pump main shaft, pointing in the direction of water jet. This direction signal mainly reflects the axial movement characteristics caused by impeller thrust pulsation and fluid momentum changes.

[0031] Y-axis (radial / vertical): Defined as the direction of the normal to the sensor mounting point, with a straight path passing through the geometric center of the pump body. This direction signal is most sensitive to structural loosening and vertical misalignment in the direction of gravity.

[0032] Z-axis (tangential / horizontal): Perpendicular to the YZ plane, i.e., along the tangent direction of the casing. This direction signal mainly reflects the mass imbalance of the rotating components and the lateral fluid excitation force.

[0033] The triaxial vibration accelerometer was glued to the housings of the first-stage impeller, pump bearing, second-stage impeller, and main shaft bearing. The Y-axis passed through the pump center, while the X and Z axes aligned with the diagram. Vibration acceleration signals were collected. The sensor measurement points and prediction system channels for the mechanical system are shown in Table 1. Data acquisition experiments were conducted at speeds of 400 r / min, 600 r / min, and 800 r / min, maintaining each condition for 5 minutes. The sampling interval was 1 second, the sampling frequency was 51.2 kHz, and the number of sampling points per sample was 8192.

[0034] Table 1 Wiring Table for Measuring Points in Mechanical Systems

[0035] like Figure 2 As shown, although the processed vibration data has numerical integrity, it still exists in the form of a one-dimensional time series, making it difficult to process directly by convolutional neural networks and to intuitively present the non-stationary impact characteristics of the water jet propulsion device under complex fluid-structure interaction. Therefore, this step aims to convert the multi-channel one-dimensional signal into a two-dimensional color time-frequency spectrum rich in physical information. Through the "RGB channel fusion" technique, the spatial vibration coupling relationship is explicitly mapped into the color texture information of the image.

[0036] The specific implementation process includes the following three sub-steps: (1) Signal time-frequency transformation: Fault signals in waterjet propulsion systems often exhibit two characteristics: low-frequency fluid pulsation and high-frequency mechanical impact. While the traditional Fast Fourier Transform (FFT) provides accurate frequency information, it completely loses the time dimension, making it impossible to pinpoint the exact moment the fault occurred. On the other hand, the Short-Time Fourier Transform (STFT), limited by the Heisenberg uncertainty principle, cannot simultaneously accommodate the time resolution of high-frequency signals and the frequency resolution of low-frequency signals due to its fixed-width window function.

[0037] Therefore, this invention uses Continuous Wavelet Transform (CWT) as the core of time-frequency analysis. CWT automatically narrows the time window in the high-frequency band to capture the impact position at the microsecond level through scaling and translation operations, and automatically widens the time window in the low-frequency band to refine the frequency components, which can accurately extract weak mechanical fault features from strong background fluid noise.

[0038] The wavelet basis selection uses the complex Morlet wavelet as the basis function, which is mathematically defined as a complex sine wave modulated by a Gaussian envelope: (1) in, For the center frequency, This is the bandwidth parameter.

[0039] The complex Morlet wavelet was chosen for the following reasons: First, at the physical level, its bilaterally decaying sine wave is highly similar to the mechanical impact response. Typical mechanical faults in waterjet propulsion devices produce a "shock-decay" vibration response under excitation. The real part of the complex Morlet wavelet is essentially a bilaterally decaying sine wave, a form that closely matches the free decaying vibration waveform of a mechanical system after impact. According to the principle of matched filters, the more similar the analysis basis function is to the waveform of the signal under test, the larger the magnitude of the correlation coefficient after transformation, thus maximizing the signal-to-noise ratio of the fault characteristics.

[0040] Secondly, from a mathematical perspective, as an analytic wavelet, its complex modulus can directly extract the instantaneous envelope energy. Compared to real wavelets, the complex Morlet wavelet is an analytic wavelet in complex form. It can extract not only the amplitude information of the signal but also preserve the phase information. In fault diagnosis, we are more concerned with the fluctuation of vibration energy over time, and the coefficients after the complex Morlet wavelet transform... For a complex number, its modulus is... Directly corresponding to the instantaneous envelope energy of the signal, it can naturally eliminate interference fringes caused by phase fluctuations, generating a smooth and clearly textured time-frequency energy map. This characteristic is crucial for generating high-quality RGB time-frequency maps, effectively avoiding feature map blurring caused by vibration phase alignment errors, thus providing the clearest input samples with the clearest texture features for the subsequent Mech-ProtoNet network.

[0041] The transformation process applies to any channel ( vibration signal The transformation is performed to obtain its time-frequency coefficient matrix. ,in This is the scale factor (corresponding to frequency). This is the translation factor (corresponding to time). Calculate the squared modulus of the wavelet coefficients. The obtained time-frequency energy spectrum directly reflects the distribution of vibration energy with time and frequency.

[0042] (2) Spatial grouping of measurement points and RGB mapping: Based on the sensor layout of the waterjet propulsion system, the system comprises four key measuring points: the first-stage impeller housing, the pump bearing housing, the second-stage impeller housing, and the main shaft bearing housing. The RGB channel mapping strategy, based on the previously described physical definition of channel mapping, maps the time-frequency spectra in three orthogonal directions to the three color channels of a color image, constructing an RGB fusion spectrum containing full-space vibration information.

[0043] The Z-axis (horizontal tangential component) is mapped to the R channel to capture rotational centrifugal force and lateral sweeping characteristics; the Y-axis (vertical radial component) is mapped to the G channel to capture radial impact and supporting structure stiffness variations; and the X-axis (longitudinal axial component) is mapped to the B channel to capture axial thrust fluctuations caused by fluid-structure interaction. Through this mapping, the originally independent single-axis vibration signals are fused into a two-dimensional image with a specific color texture. The synthesized RGB image is no longer an ordinary picture; its colors directly reflect the spatial modes of vibration. When a fault occurs, due to mass imbalance (Z-axis surge) accompanied by uneven fluid thrust (X-axis surge), the fused image will exhibit a magenta highlight texture area, thus achieving intuitive visualization of the fault mode in color space, indicating that at that moment and frequency, the equipment experienced strong X-axis and Z-axis coupled vibration.

[0044] (3) Image standardization and sample construction: The four types of RGB time-frequency images generated by fusion (corresponding to four measurement points respectively) were uniformly resized. (For example (pixels) to meet the input requirements of subsequent deep learning models. Finally, for each time period of data collection, the system generates a set of multi-dimensional image tensors containing spatial and temporal features, which serve as standard input samples for the third-step Mech-ProtoNet model. This processing method not only achieves data dimensionality compression but also forces the network to simultaneously focus on the "frequency features" (texture) and "spatial orientation features" (color) of vibrations during the learning process, thereby significantly improving the ability to identify complex fault modes.

[0045] Step 2: Train Mech-ProtoNet Based on the generated RGB image set containing spatial-temporal features, this step aims to construct and train an interpretable deep diagnostic model called Mech-ProtoNet. This model does not directly learn the black-box mapping from input images to fault labels, but instead autonomously learns prototypes representing typical textures of various faults in the feature space, achieving transparent diagnosis based on physical feature similarity. The specific implementation process includes four stages: model architecture construction, dataset configuration, loss function design, and training strategy. The network framework is as follows: Figure 3 As shown.

[0046] (1) Model architecture construction: Mech-ProtoNet consists of two parts: a deep feature extractor and a prototype decision layer.

[0047] The feature extractor uses a pre-trained ResNet-50 or DenseNet-121 as the backbone network. This network is responsible for receiving the input RGB time-frequency image. (size And through multiple convolution and pooling operations, it is mapped into a high-dimensional feature map. This process is equivalent to extracting an abstract "texture fingerprint" from the original time-frequency image.

[0048] The prototype decision-making layer is the core innovative module of this invention. The network is pre-defined. A learnable prototype vector Each fault category is assigned One prototype. Each prototype These are not abstract numbers, but directly correspond to feature maps. one of the or The model calculates the latent image patches of the input image and the relationships between them. Euclidean distance. If a local texture in the input image is very close to the prototype vector of a specific fault category, the activation value at that location increases sharply. This mechanism ensures that the model reasons based on the logic that "this part of the image looks like a typical fault case," rather than an unknowable mathematical fit.

[0049] (2) Dataset Configuration: The dataset has five health states: normal state, blade tip gap too small, blade tip gap too large, impeller major damage, and impeller minor damage. For each health state, 300 sets of data were collected at speeds of 400 r / min, 600 r / min, and 800 r / min. Each set contains 24,576 (collected 3 times) or 32,768 (collected 4 times) sample points. Each operating condition includes four sensor measurement points: the first-stage impeller housing, the pump bearing housing, the second-stage impeller housing, and the main shaft bearing housing, for a total of 12 channels of data, as shown in Table 2 below.

[0050] Table 2 Device Status Dataset

[0051] The dataset was randomly shuffled and divided into a 70% training set and a 30% test set to ensure that the model could generalize under unseen conditions.

[0052] (3) Hybrid Loss Function Design: To force the model to learn representative fault prototypes, Mech-ProtoNet abandons the single cross-entropy loss and adopts a joint optimization strategy. Total Loss Function The definition is as follows: (2) Classification accuracy loss ( Clustering loss: This ensures the model's accuracy in predicting fault categories and penalizes incorrect classifications. ): (3) It mandates that image features belong to a certain category. It must be as close as possible to the prototype of the class. This makes the same type of fault highly cohesive in the feature space.

[0053] Separation loss ( ): (4) It mandates image features It is necessary to keep away from prototypes that are not of that type. For example, the image features of "normal state" must be kept as far away as possible from the prototype of "impeller damage" in order to establish a clear fault boundary in the feature space.

[0054] (4) Training Strategy: A phased training method is adopted to accelerate convergence and ensure the clarity of the prototype's physical semantics. During the warm-up phase, the weights of the backbone network are frozen, and only the prototype layer and the final fully connected layer are trained. The Adam optimizer is used in this phase, with a learning rate set to [value missing]. The goal is to quickly initialize prototype vectors with distinctiveness.

[0055] During the joint fine-tuning phase, the last few layers of the backbone network are unfrozen and jointly trained with the prototype layers. A learning rate decay strategy is introduced, halving the learning rate every 10 epochs until the loss function converges.

[0056] Every 5 epochs, a "projection" operation is performed during the prototype projection phase. This involves transferring the learned abstract prototype vectors... The vector is replaced with the nearest real image patch in the training set. This step ensures that the final prototype is not a virtual mathematical vector, but a real-world slice of a time-frequency graph, thus providing a real-world physical case to support "post-hoc interpretability".

[0057] Step 3: Dual Interpretability Verification and Intelligent Diagnostic Output After training with Mech-ProtoNet, the model acquires the ability to extract features and match prototypes for waterjet propulsion device faults. This step involves validation on a test set and deployment of an online diagnostic process. Unlike traditional deep learning, which only outputs fault category probabilities, this invention aims to open the "black box" of deep learning by backpropagating gradient flow to precisely visualize the time-frequency feature regions that the network focuses on during decision-making. This verifies the physical consistency of the diagnostic logic and provides maintenance personnel with physically based decision support.

[0058] The specific implementation process is as follows: 1. Prototype matching mechanism based on Euclidean distance metric: This step constructs a compact feature metric space and uses the geometric distance between the input features and the prototype vector to quantify the probability of fault category attribution. Let the input time-frequency image be... The backbone network maps it to feature vectors in the latent feature space. ,in These are the network parameters. The set of prototypes learned by the model is defined as follows: ,in The model calculates the feature vector. With the A prototype square Euclidean distance : (5) This distance function directly characterizes the degree of deviation of the current sample from the typical failure mode in the manifold space.

[0059] To transform geometric distance into a probability distribution, this invention performs probabilistic inference based on radial basis functions (RBF) and employs a kernel-based soft allocation strategy. For the target category... The conditional probability that it belongs to this category Defined as the sum of normalized exponents of similarity among all prototypes in this category: (6) in is the scaling factor. This formula shows that classification decisions strictly depend on the neighborhood structure of feature points in the metric space; the closer the distance, the higher the posterior probability, thus providing a rigorous numerical interpretation.

[0060] Suppose the model output classifies the current sample as "major impeller damage." The underlying quantitative interpretation is as follows: the feature vector of the current input sample, in the feature space, has a distance of only 0.15 to a certain prototype under the "major impeller damage" category, while its closest distance to the prototype in the "normal state" category is as high as 12.8. Based on this significant distance difference, the system determines that a fault has occurred. Besides classification, this distance metric also has a "rejection" function. If the distance between the current sample and all known prototypes exceeds a set safety threshold... The system will not force a classification, but will instead output "unknown anomaly". This is extremely valuable in practical engineering, as it can effectively prevent the model from randomly classifying unseen fault modes.

[0061] 2. High-resolution multi-scale visual source tracing based on LayerCAM: To achieve refined visual tracing of the decision-making process of Mech-ProtoNet, this invention abandons the traditional global average pooling class interpretation method (such as CAM, Grad-CAM, Grad-CAM++) and instead adopts the LayerCAM (LayerClass Activation Mapping) algorithm.

[0062] Traditional Gradient Weighted Class Activation Mapping (Grad-CAM) and its improved version (Grad-CAM++) primarily rely on the gradient information of the last convolutional layer of the backbone network. They calculate the weight of each feature map by performing global average pooling on the backpropagation gradient. While this "averaging" operation can locate the approximate region of the target, it inevitably destroys the spatial structural information of the feature maps, resulting in the generated interpretive heatmaps often appearing as blurry clumps. For the vibration time-frequency map of a waterjet propulsion device, key fault features often manifest as extremely narrow frequency bands or transient impact textures. The low-resolution heatmaps generated by traditional methods cannot accurately delineate the contours of these subtle features, failing to meet the needs of precise diagnosis.

[0063] LayerCAM proposes a pixel-level gradient-based weighted strategy to generate high-quality, class-specific, and pixel-level consistent class activation maps. Its core innovations encompass two dimensions: Preservation of spatial heterogeneity weights: LayerCAM no longer averages gradients, but directly uses the positive gradients backpropagated to the convolutional layers as weights. This means that for each spatial location on the feature map... Each feature has its own independent weight, which directly reflects the contribution of the feature at that location to the final classification. Even in some shallower convolutional layers, LayerCAM can use this pixel-level weight to separate subtle texture features from complex background noise.

[0064] Multi-scale hierarchical feature combination: Different convolutional layers in a deep neural network contain semantic information at different levels. Shallow convolutional layers preserve high-resolution spatial details (edges, textures), while deep convolutional layers contain abstract semantic category information. LayerCAM can simultaneously extract feature maps of different depths in the backbone network to generate activation maps, and combine the fine contours of the shallow layers with the semantic localization of the deep layers through a specific fusion mechanism.

[0065] To address the applicability of this invention to the fault diagnosis of water jet propulsion devices, the input data is a time-frequency spectrum fused from RGB channels. The aforementioned characteristics of LayerCAM enable it to accurately locate frequency components and clearly separate the energy distribution of the fundamental frequency and higher harmonics, rather than blurring them into a single, indistinct mass; furthermore, LayerCAM can precisely define the activation intensity of different RGB channels over a specific time period.

[0066] Given LayerCAM's significant advantages in fine-grained feature localization and multi-scale information fusion, this embodiment integrates it into the interpretation module of the intelligent diagnostic system. The specific implementation process and mathematical expression are described below: (1) Pixel-level weighting of layered gradients: Unlike Grad-CAM++, which only utilizes the last convolutional layer and performs global average pooling on the gradient, LayerCAM applies pixel-level weighting to each specific convolutional layer in the backbone network. Calculate its feature map Each spatial location weight The weights are determined by the pixel-level response of the positive gradient: (7) This formula directly uses the gradient from backpropagation as the spatial weight of the feature map, avoiding the loss of spatial information caused by averaging operations and ensuring sensitivity to minute vibration textures.

[0067] (2) Multi-scale feature fusion mechanism: The model extracts feature maps from different stages of the backbone network (e.g., Conv3_x, Conv4_x, Conv5_x of ResNet) and generates activation maps for each level. : (8) Subsequently, bilinear interpolation is used to align activation maps of different resolutions to the input image size, and then they are fused and overlaid. (9) By fusing high-resolution texture features from shallow layers with advanced semantic features from deep layers, LayerCAM can simultaneously and accurately locate both wideband background noise (captured by deep layers) and extremely narrowband fault harmonics (captured by shallow layers).

[0068] (3) Physical consistency verification based on LayerCAM: The generated high-resolution heatmap is superimposed on the RGB time-frequency map for refined physical verification. First, a fine structure verification is performed to observe whether the highlighted outline of the heatmap accurately delineates the "fish spine" frequency (BPF) band. If LayerCAM can clearly separate the energy lines of the fundamental frequency and harmonics, rather than blurring them into a mess, it proves that the model has truly learned the frequency structure characteristics of the fault. Then, RGB channel coupling verification is performed to check the color channel weights of the heatmap at specific pixels. For example, in the "excessive gap" fault, the heatmap should accurately illuminate the specific time-frequency block at the intersection of the Z channel (tangential) and the X channel (axial) to verify the model's ability to capture multidimensional coupled vibrations.

[0069] Therefore, this invention provides a fault diagnosis method for the mechanical system of a waterjet propulsion device based on a time-frequency prototype network. This method can achieve accurate, efficient and interpretable fault diagnosis, as well as visualize the fault diagnosis process and perform case-based reasoning-based post-hoc attribution of the diagnosis results. This provides a high-precision and high-reliability fault diagnosis solution for intelligent ship operation and maintenance.

[0070] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0071] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for fault diagnosis of a mechanical system of a waterjet propulsion device based on a time-frequency prototype network, characterized in that, include: Data is collected and transformed using sensor points set on the mechanical system of the water jet propulsion device to obtain a single-channel time-frequency matrix, and a spatial coordinate system is defined for physical mapping. Using the defined RGB mapping rules, the single-channel time-frequency matrix is ​​transformed into a fused RGB time-frequency image, and then the fused RGB time-frequency image is processed to obtain a standard input sample. A state dataset including five health states is collected at the sensor measurement points. The initial model is then trained in stages using the state dataset to obtain a diagnostic model. The diagnostic model is used to perform inference diagnosis on the standard input sample to obtain fault diagnosis results and credibility basis, and the fault diagnosis results are verified to obtain high-resolution evidence.

2. The method for fault diagnosis of a waterjet propulsion device mechanical system based on a time-frequency prototype network according to claim 1, characterized in that, Data is acquired and time-frequency transformed using sensor points installed on the water jet propulsion system to obtain a single-channel time-frequency matrix. A spatial coordinate system is then defined for physical mapping, including: In a mechanical system based on a water jet propulsion device, sensor measuring points are set on the housing of the first-stage impeller, the housing of the pump bearing, the housing of the second-stage impeller, and the housing of the main shaft bearing. Data is collected at these sensor measuring points to obtain multi-channel raw data. For a mechanical system based on a water jet propulsion device, the direction parallel to the center line of the propulsion pump main shaft and pointing towards the water jet is defined as the X-axis, the direction from the sensor mounting point and passing through the geometric center of the pump body is defined as the Y-axis, and the direction along the tangent of the casing is defined as the Z-axis, thus obtaining a spatial coordinate system for physical mapping. Based on the multi-channel raw data, using continuous wavelet transform based on complex Morlet wavelets, and through scaling and translation operations, the time window is automatically narrowed in the high-frequency band to capture the impact position at the microsecond level, and the time window is automatically widened in the low-frequency band to refine the frequency components, thus obtaining a single-channel time-frequency matrix.

3. The method for fault diagnosis of a waterjet propulsion device mechanical system based on a time-frequency prototype network according to claim 2, characterized in that, Using the defined RGB mapping rules, the single-channel time-frequency matrix is ​​transformed into a fused RGB time-frequency image. Then, the fused RGB time-frequency image undergoes image processing to obtain standard input samples, including: Set the defined Z-axis to the R channel, the Y-axis to the G channel, and the X-axis to the B channel to complete the definition of the RGB mapping rule; Based on the single-channel time-frequency matrix, the R channel is used to capture the rotational centrifugal force and lateral ramming characteristics, the G channel is used to capture the radial impact and the stiffness change of the supporting structure, and the B channel is used to capture the axial thrust fluctuation caused by fluid-structure interaction, so as to obtain a fused RGB time-frequency map. Based on the fused RGB time-frequency image, the size is unified, and for each acquisition time period, a set of multidimensional image tensors containing spatial and time-frequency features is generated to obtain standard input samples.

4. The method for fault diagnosis of a waterjet propulsion device mechanical system based on a time-frequency prototype network according to claim 3, characterized in that, A state dataset including five health states is collected at the sensor measurement points. This state dataset is then used to train the initial model in stages to obtain a diagnostic model, including: By combining a deep feature extractor and a designed prototype decision layer, an initial model is obtained. The deep feature extractor is used to map the standard input samples into a high-dimensional feature map. Based on the high-dimensional feature map, the prototype decision layer is used to calculate the Euclidean distance between the input feature block and each prototype vector to complete the initial inference. Based on the sensor measurement points, state datasets of five health states, including normal state, blade tip gap too small, blade tip gap too large, impeller major damage, and impeller minor damage, are collected under different working conditions. The state datasets are randomly shuffled and divided into training set and test set in a 7:3 ratio. The design includes a hybrid loss function comprising a classification accuracy loss function, a clustering loss function, and a separation loss function. The initial model is trained in stages using the training set, the test set, and the hybrid loss function to obtain a diagnostic model.

5. The method for fault diagnosis of a waterjet propulsion device mechanical system based on a time-frequency prototype network according to claim 4, characterized in that, The prototype decision layer is used to preset multiple learnable prototype vectors and assign multiple learnable prototype vectors to each type of fault; wherein each learnable prototype vector corresponds to a potential image patch in the standard input sample.

6. The method for fault diagnosis of a waterjet propulsion device mechanical system based on a time-frequency prototype network according to claim 5, characterized in that, Phased training includes: The weights of the deep feature extractor are frozen, the prototype decision layer and the fully connected layer are trained, and the prototype vector is initialized using the Adam optimizer and a preset learning rate to complete the warm-up training phase. The thawed deep feature extractor is jointly trained with the prototype decision layer, and a learning rate decay strategy is introduced until the hybrid loss function converges, thus completing the joint fine-tuning training phase. The learnable prototype vector is replaced with the nearest real image patch in the training set to complete the prototype projection training phase.

7. The method for fault diagnosis of a waterjet propulsion device mechanical system based on a time-frequency prototype network according to claim 6, characterized in that, The diagnostic model is used to perform inference diagnosis on the standard input sample to obtain fault diagnosis results and credibility basis. The fault diagnosis results are then verified to obtain high-resolution evidence, including: The standard input sample is input into the diagnostic model, and feature mapping, Euclidean distance measurement and distance transformation are performed to obtain the category conditional probability. Then, based on the category conditional probability, the fault diagnosis result and confidence basis are output. Based on the fault diagnosis results and the credibility criteria, pixel-level weighting of the hierarchical gradient is performed using hierarchical activation mapping to generate activation maps for each layer. The activation maps of different resolutions are interpolated, aligned to the input size, and fused and superimposed to obtain a multi-scale fused high-resolution heatmap. Then, the high-resolution heatmap is physically consistent to obtain high-resolution evidence.

8. The method for fault diagnosis of a waterjet propulsion device mechanical system based on a time-frequency prototype network according to claim 7, characterized in that, The physical consistency verification includes fishbone-shaped frequency verification and RGB channel coupling verification.