Fault diagnosis method for aero-engine bearing driven by large audio model

Through multi-stage processing and cross-modal attention mechanisms, combined with LoRA technology, the direct generation of interpretable fault labels is solved, and the existing methods fail to utilize acoustic signal similarity and lack of specialized models is achieved, achieving efficient and accurate aero engine bearing fault diagnosis.

CN120336929APending Publication Date: 2025-07-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510496242.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing fault diagnosis methods fail to fully utilize the inherent similarity between vibration signals and acoustic signals, the output needs to be post-processed, and the lack of large-scale models designed specifically for aircraft engine bearing fault diagnosis, limiting the effectiveness of predictive maintenance.

Method used

Through multi-stage processing, the audio input recognition capability of a large-scale language model is established, combined with cross-modal attention mechanism and LoRA technology, interpretable fault labels and diagnostic explanations can be directly generated to achieve knowledge transfer and parameter reduction.

Benefits of technology

It improves the accuracy and interpretability of aircraft engine bearing fault diagnosis, reduces the computing resource requirements, and provides reliable predictive maintenance solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336929A_ABST
    Figure CN120336929A_ABST
Patent Text Reader

Abstract

The invention discloses a bearing fault diagnosis method for an aero-engine driven by an audio large model, and relates to the technical field of bearing fault diagnosis. The method comprises the following steps: establishing the recognition capability of a large-scale language model on audio input through multi-stage processing to obtain a pre-training basic model; constructing a paired data set of aero-engine bearing vibration signals and detailed text description of the aero-engine bearing vibration signals, and performing amplitude normalization and coding on the vibration signals to obtain audio samples; establishing connection between the audio samples and the text modals through a cross-modal attention mechanism; and taking the audio sample and the text mode which are connected together as the input of a pre-training basic model, and directly generating a text fault tag and a diagnosis explanation in an autoregression mode by utilizing the generation capability of the model. According to the invention, a system capable of directly outputting interpretable and operable fault diagnosis results is established by using the general mode recognition capability of a large-scale audio pre-training model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bearing fault diagnosis, and particularly to a method for diagnosing faults of an aero-engine bearing driven by an audio large model. Background Art

[0002] As a core component of modern aviation and aerospace industries, the aero-engine provides power for various aircrafts, from commercial airplanes to space exploration vehicles. Its reliable operation is crucial for aviation safety, mission success, and prevention of catastrophic failures that can lead to significant human and economic losses. In these complex propulsion systems, aero-engine bearings are particularly critical components that operate under extreme conditions such as high speeds, temperature variations, and significant mechanical stresses. The failure of these bearings can lead to severe engine malfunctions, making them a major focus of condition monitoring and maintenance protocols. Therefore, the fault diagnosis of aero-engine bearings has become an important research field, providing methods for detecting, identifying, and predicting potential faults before catastrophic events occur. Early and accurate bearing fault detection can significantly extend the engine life, optimize maintenance schedules, and ensure operational safety.

[0003] Traditional bearing fault diagnosis methods typically rely on a two-step process of feature extraction and subsequent classification. However, such methods are limited by their dependence on domain-specific feature engineering and traditional machine learning classifiers, and may not be able to fully capture the complex dynamic characteristics of aero-engine bearing systems.

[0004] In recent years, data-driven methods, especially deep learning methods, have revolutionized the field of fault diagnosis by directly performing end-to-end learning from raw vibration signals. These methods automatically extract hierarchical features, eliminating the need for manual feature engineering required by traditional methods. Convolutional neural networks (CNNs) have been widely adopted because of their ability to capture local patterns and spatial hierarchies in data. In addition to CNNs, recurrent neural networks, especially long short-term memory (LSTM) networks, have also gained attention in vibration signal analysis. LSTMs are good at capturing temporal dependencies in sequential data and are suitable for modeling the time-varying characteristics of vibration signals. Architectures based on Transformers, which utilize self-attention mechanisms to capture long-range dependencies and context information, have also become powerful tools for vibration signal analysis.

[0005] However, the above data-driven methods usually output logits or confidence scores, which require post-processing to obtain actionable insights, while industrial applications demand efficient and interpretable solutions. Recently, the emergence and development of large language models (LLMs) and large multimodal models (LMMs) have provided transformative advantages in this regard. These models demonstrate excellent capabilities in generating human-like text, understanding complex queries, and providing contextually relevant responses. Despite these advancements, the application of large-scale models in fault diagnosis is still in its infancy. Whether leveraging the text modality of LLMs or extending to the visual modality of LMMs, these methods fail to recognize a fundamental characteristic of bearing vibration signals, namely their inherent similarity to acoustic phenomena. Bearing vibration signals, as mechanical oscillations propagated through a medium, manifest as time-varying waveforms with frequency, amplitude, and temporal patterns, which are very similar to sound waves. This acoustic similarity is manifested in their shared spectral characteristics, such as harmonic peaks and resonance frequencies, which are routinely analyzed in audio processing but not fully utilized in fault diagnosis. Additionally, despite the critical importance of aerospace propulsion systems and the severe consequences of their failures, large-scale models specifically designed for aerospace engine bearing fault diagnosis have not been developed. The lack of such specialized models represents a significant gap in the current research field and limits the effectiveness of predictive maintenance in aerospace applications. Summary of the Invention

[0006] The object of the present invention is to provide a method for diagnosing aerospace engine bearing faults driven by a large audio model, aiming to address the deficiencies of existing fault diagnosis technologies in the following aspects: Firstly, existing deep learning methods usually output logical values or confidence scores, which require post-processing to obtain actionable diagnostic results; Secondly, most methods fail to recognize the inherent similarity between vibration signals and acoustic signals and do not fully utilize the advanced research results in the acoustic field; Finally, the lack of large-scale models specifically designed for aerospace engine bearing fault diagnosis limits the effective application of predictive maintenance in the aerospace field. The present invention aims to establish a system that can directly output interpretable and actionable fault diagnosis results by leveraging the general pattern recognition capabilities of large-scale audio pre-trained models, thereby improving the accuracy, interpretability, and practicality of aerospace engine bearing fault diagnosis.

[0007] Specifically, the present invention provides a method for diagnosing aerospace engine bearing faults driven by a large audio model, including: Establishing the recognition ability of a large language model for audio input through multi-stage processing to obtain a pre-trained basic model, enabling the model to extract meaningful features from complex acoustic signals, where the multi-stage processing includes basic acoustic understanding, supervised fine-tuning, and direct preference optimization; Construct a paired dataset of aero-engine bearing vibration signals and their detailed text descriptions, normalize and encode the vibration signals to obtain audio samples; establish a connection between the audio samples and the text modality through a cross-modal attention mechanism; among them, when encoding the audio, by setting the LoRA layer, modify the processing of the vibration signals so that the encoder can recognize the acoustic characteristics of various bearing faults; Use the connected audio samples and text modality as the input of the pre-trained base model, and utilize the model's generation ability to directly generate text fault labels and diagnostic explanations in an autoregressive manner; among them, establish adaptation points based on the LoRA technology in the self-attention mechanism and the feed-forward network of the pre-trained base model, so as to achieve knowledge transfer and adapt the general audio understanding ability of the model to the specific field of aero-engine bearing vibration signals.

[0008] Furthermore, in the basic acoustic understanding stage, by converting the audio input into a text output and processing various types of audio inputs, the model learns to extract meaningful features from complex acoustic signals.

[0009] Furthermore, the supervised fine-tuning stage further enhances the model's analysis ability for complex audio content; the direct preference optimization stage uses a comparison scoring mechanism to evaluate alternative responses to the same audio input, and the scores indicate their relative quality.

[0010] Furthermore, the LoRA freezes the pre-trained weight matrix and introduces a trainable low-rank adaptation layer to achieve knowledge transfer and adapt the general audio understanding ability to the specific field of aero-engine bearing vibration signals.

[0011] Furthermore, the detailed text description of the aero-engine bearing vibration signal specifically includes: vibration source, noise characteristics, fault type, severity, location, and related operating conditions.

[0012] Furthermore, the pre-trained base model directly generates text fault labels and diagnostic explanations in an autoregressive manner, and its generation process is as follows: ; where represents the token predicted at position t, represents the vocabulary space, represents the probability distribution of the model predicting the next token under the given conditions, represents the encoding of the vibration signal, including all previously generated tokens.

[0013] Furthermore, the training objective of model optimization is expressed as the cross-entropy loss function on the fault label token sequence: ; Among them represents the fault diagnosis data set composed of vibration signal-label pairs (v, y), represents the probability of generating the correct label given the audio embedding and the previous tokens .

[0014] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: The present invention uses the general pattern recognition ability of large-scale audio models for vibration signal analysis, enabling the model to extract meaningful features from complex acoustic signals. The proposed Vibration Signal Alignment (VSA) mechanism successfully bridges the gap between general audio knowledge and domain-specific vibration patterns, enabling the model to adapt its acoustic representation to the characteristics of bearing vibration signals. The Generative Fault Classification (GFC) method utilizes the generative ability of large language models to directly output interpretable fault labels, eliminating the need for post-processing steps in traditional methods. By using the LoRA technique, parameter reduction is achieved, greatly reducing the computational resource requirements and making the technology more practical. Compared with the prior art, the present invention not only improves the accuracy of fault diagnosis but also takes into account computational efficiency and interpretability, providing a reliable technical solution for the predictive maintenance of aeroengines and showing promise for broader industrial applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings described herein are used to provide a further understanding of the embodiments of the present invention and constitute a part of the present invention. The illustrative embodiments and descriptions thereof are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is the overall flowchart of the diagnostic model framework proposed by the present invention; Figure 2 is the schematic diagram of the implementation details of the aeroengine bearing fault diagnosis method driven by the large audio model proposed by the present invention; Figure 3 is the acquisition device for verifying the bearing vibration data used in the present invention, where (a) is the overall structure diagram of the test bench, (b) is the schematic diagram of the installation positions of the acceleration sensors (A1, A2) and the reference coordinate system, and (c) is the physical diagram of the test shaft with three roller bearings (B1, B2, B3). DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] In order to make the technical problems, technical solutions, and beneficial effects to be solved by the present invention clearer, the present invention will be further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0017] AsFigure 1 and Figure 2 As shown in Figure 2 , the present invention provides a method for diagnosing faults in aero-engine bearings driven by an audio large model, comprising the following steps: Step 1: Audio knowledge acquisition: Establish the recognition ability of a large language model for audio input through multi-stage processing to obtain a pre-trained basic model, enabling the model to extract meaningful features from complex acoustic signals. The multi-stage processing includes basic acoustic understanding, supervised fine-tuning, and direct preference optimization.

[0018] Specifically, the vibration data of the aero-engine bearing is collected through an acceleration sensor. Therefore, a segment of vibration signal sample collected can be expressed as: represents the number of samples of a segment of sensor signal. To obtain the basic audio understanding ability, this method adopts a multi-stage processing process: Firstly, in the basic acoustic understanding stage, the audio input (represented by the microphone icon, for example) is converted into a text output (such as "<[en]>Hello there!"). By processing various audio inputs, the model learns to extract meaningful features from complex acoustic signals.

[0019] Secondly, supervised fine-tuning (SFT) is performed to enhance the analysis ability of the model. This stage focuses on higher-level audio understanding tasks, such as answering complex questions about the audio content, such as "How does the speaker feel?" and "What kind of sound is this?", and providing answers such as "Very happy. The sound of rain." Finally, the direct preference optimization (DPO) method is implemented. This technique uses a comparison scoring mechanism, in which alternative responses to the same audio input are evaluated, and the scores indicate their relative quality (for example, answer 1 scores 8.5 compared to answer 2 scores 2.0). The optimization objective formula is: ; where, represents the system to be optimized, is the reference system, and are the better and worse responses respectively, σ is the sigmoid function, β is a hyperparameter controlling the optimization intensity, is the input audio signal data.

[0020] Step 2: Vibration Signal Alignment (VSA): Construct a paired dataset of aero-engine bearing vibration signals and their detailed text descriptions, normalize and encode the vibration signals to obtain audio samples; establish a connection between the audio samples and the text modality through a cross-modal attention mechanism, enabling the model to adapt its acoustic representation to the specific characteristics of the bearing vibration signals; among which, when encoding the audio, the processing of the vibration signals is modified by setting the LoRA layer, enabling the encoder to recognize the acoustic characteristics of various bearing faults.

[0021] The vibration signal alignment mechanism aims to bridge the fundamental gap between general audio knowledge and the specific characteristics of aero-engine bearing vibrations. This is achieved through an alignment process that retains the rich hierarchical representations inherent in large-scale audio models while adapting them to the specialized patterns of mechanical vibrations.

[0022] Proper alignment requires establishing a connection between the vibration signal and its corresponding semantic interpretation. To this end, the present invention constructs a dataset paired with vibration signals and detailed text descriptions that express their diagnostic significance.

[0023] Each text description systematically captures key fault-related information, including the vibration source, noise characteristics, fault type, severity, location, and relevant operating conditions. This pairing strategy enables the model to develop a nuanced understanding of the relationship between acoustic features and their diagnostic implications.

[0024] The alignment process begins with precise signal preparation to ensure optimal encoding. The original vibration signal is amplitude-normalized through a statistical transformation: ; where and represent the mean and standard deviation of the signal respectively, and α is a calibration factor for optimizing the dynamic range of the audio encoder. This normalization ensures consistent amplitude characteristics under different operating conditions and sensor configurations, significantly enhancing the model's robustness to variations in signal acquisition parameters.

[0025] The normalized vibration signal is converted into an audio embedding through an audio encoder - domain adaptation encoder embedded with the LoRA layer: ; where represents the audio encoder, represents the frozen parameters of the pre-trained model, contains the trainable LoRA parameters. This formulation enables selective adaptation of the encoder parameters while retaining its underlying acoustic understanding ability. The encoder projects the temporal vibration patterns into a high-dimensional embedding space , where La represents the sequence length and d is the embedding dimension.

[0026] The cross-modal attention mechanism then establishes a connection between the audio and text modalities and calculates the attention weights as follows: ; where represents the text embedding representation obtained by processing the text input (such as instruction prompts and context information) through the text encoder, represents the audio embedding representation obtained by processing the normalized vibration signal through the audio encoder, W Q and W K are learnable projection matrices, and d k is the dimension of the key vector. This mechanism allows the model to selectively focus on relevant regions of the vibration signal when interpreting or generating diagnostic text, effectively establishing a vibration-to-language mapping that can capture subtle fault indication patterns.

[0027] Step 3: Generative Fault Classification (GFC): Use the concatenated audio samples and text modality as the input to the pre-trained base model. Leverage the generative ability of the model to directly generate text fault labels and diagnostic explanations in an autoregressive manner, eliminating the need for post-processing steps in traditional methods and achieving seamless integration with subsequent analysis capabilities, enabling the model to not only classify faults but also provide interpretable labels and respond to queries regarding fault characteristics. Among them, adaptation points based on the LoRA technique are established in the self-attention mechanism and feed-forward network of the pre-trained base model to achieve knowledge transfer and adapt the general audio understanding ability of the model to the specific domain of aero-engine bearing vibration signals.

[0028] Generative Fault Classification represents a paradigm shift in fault diagnosis, leveraging the inherent generative ability of large language models to directly produce interpretable diagnostic outputs. Different from traditional deep learning methods, which output numerical logits that require post-processing, this method enables the model to generate human-readable fault labels and natural language explanations, significantly enhancing interpretability and operability for maintenance personnel.

[0029] Traditional fault classification methods typically employ multi-class classifiers that output the probability distribution of predefined fault categories, requiring argmax operations and label mapping to determine the final classification. In contrast, this method utilizes the autoregressive nature of the LLM to directly generate text fault labels. This approach eliminates intermediate post-processing steps, simplifies the diagnostic process, and simultaneously provides an output in a format immediately understandable to maintenance technicians. The generation process is as follows: ; where represents the token predicted at position t, represents the vocabulary space, Represents the probability distribution of the model predicting the next token under given conditions, Represents the encoding of the vibration signal, including all previously generated tokens. This autoregressive generation continues until a complete fault diagnosis is generated.

[0030] The training objective for model optimization is formulated as a cross-entropy loss function on the fault label token sequence: ; where represents the fault diagnosis dataset consisting of vibration signal-label pairs (v, y), represents the probability of generating the correct token given the audio embedding and the previous tokens .

[0031] A unique advantage of generative fault classification is its seamless integration with subsequent analysis capabilities. By preserving the generative nature of the underlying LLM, this method can not only classify faults but also provide interpretable labels and respond to queries regarding fault characteristics. This is formulated as a conditional generation task: ; where r represents the generated response, q represents the subsequent query, and y is the initially generated fault label. This process enables context analysis considering the original vibration signal and diagnostic history, facilitating a deeper exploration of fault characteristics and impacts.

[0032] It should be noted that, in order to effectively transfer general audio knowledge to the field of aero-engine bearing fault diagnosis, the present invention introduces a domain adaptation method based on the Low-Rank Adaptation (LoRA) technique when performing the above-mentioned Steps 2 and 3. LoRA freezes the pre-trained weight matrix and introduces a trainable low-rank adaptation layer to achieve knowledge transfer and adapt the general audio understanding ability to the specific domain of aero-engine bearing vibration signals.

[0033] The principle of the above adaptation process is to effectively integrate domain-specific knowledge through low-rank decomposition of weight updates. Specifically: For the pre-trained weight matrix in the model, adaptation is achieved as follows:

[0034] where and are low-rank matrices with rank . This decomposition significantly reduces the number of trainable parameters, from d×k to r×(d + k).

[0035] During the adaptation process, the original pre-trained weight matrix Remain frozen, only the low-rank matrices A and B are updated. The forward pass calculation is modified as follows:

[0036] where x represents the input of the layer and h is the output. For ease of stable training, matrix A is usually initialized with random Gaussian values, while matrix B is initialized with zeros, ensuring that at the start of training, 。

[0037] In practice, a scaling factor α is also introduced to control the magnitude of adaptation:

[0038] Here, α is a hyperparameter that can be set to be equal to the rank r or adjusted individually. This scaling factor helps to stabilize the training, especially when trying different rank values.

[0039] The LoRA technique adaptation in the present invention is applied at multiple key points. Specifically, within the audio encoder, the LoRA layer modifies the processing of vibration signals, enabling the system to recognize the unique acoustic characteristics of various bearing faults. For the language model component, the adaptation points are established in the self-attention mechanism (especially in the query, key, and value projections) and the feed-forward network.

[0040] To better illustrate the technical effects of the present invention, a specific embodiment is used to experimentally verify the present invention.

[0041] Dataset Introduction: This study uses the Aero Bearing Vibration Dataset provided by the Dynamic and Identification Research Group (DIRG) of the Politecnico di Torino for experimental verification. This dataset was collected on a test platform designed for high-speed aero bearings (such as Figure 3 ), and these bearings can operate at speeds up to 35,000 rpm. Figure 3 (a) shows an overall view of the test bench. This test platform consists of a shaft driven by a high-speed spindle, which is supported by two identical rolling bearings, and a third larger bearing is used to apply a radial load. Figure 3 (b) shows a schematic diagram of the installation positions of the acceleration sensors (A1, A2) and the reference coordinate system. This figure shows the specific positions where the two triaxial acceleration sensors A1 and A2 are installed on the support structures of bearings B1 and B2, and marks the reference coordinate system (x, y, z) for data acquisition. Figure 3(c) Physical diagram of the test shaft with three roller bearings (B1, B2, B3), which shows the test shaft removed from the test bench and assembled with three roller bearings (B1 and B3 are test bearings, and B2 is the loading bearing). The bearing at position B1 was systematically modified to incorporate various fault conditions. Vibration signals were recorded by triaxial IEPE accelerometers mounted at two critical locations, namely, on the bracket of the damaged bearing and on the bracket of the bearing applying the load. The dataset contains records from seven different bearing states: one healthy state (0A) and six damaged states (1A - 6A). The dataset was collected at a sampling frequency of 51,200 Hz and 4 dB of Gaussian white noise was added to simulate the high-noise operating environment of the bearings.

[0042] Experimental conditions: In this part, a comprehensive experimental evaluation and comparative study on the performance of the proposed model in aero-engine bearing fault diagnosis were carried out. The experiments were conducted on a workstation equipped with an NVIDIA A10 GPU (24GB VRAM) to meet the memory requirements and training batch requirements of large models. All deep learning models were implemented in the PyTorch framework. During the training process, an optimizer with a learning rate of 1×10^-5, LoRA rank r = 16, and LoRA scaling factor α = 32 were used. The training batch size was set to 32, and gradient accumulation was performed for 16 steps. A linear warm-up schedule of 5% of the total steps was implemented to ensure training stability.

[0043] Experimental content: To verify the effectiveness of the model, we conducted two main experiments. The first experiment was an ablation study, aiming to systematically evaluate the contribution of each core component of the present invention to the overall performance. The present invention designed various component combination schemes, including using only the base model (FM) - the pre-trained base model obtained in the above step 1, the base model with vibration signal alignment (FM + VSA), the base model with generative fault classification (FM + GFC), and the complete framework (FM + VSA + GFC). By comparing the performance metrics of these different configurations on the same test set, the contribution degree of each component and its synergistic effect were quantified. This step-by-step addition of components enabled us to clarify the actual value of each innovative module, especially the role of the vibration signal alignment mechanism in bridging the gap between general audio knowledge and specific domain vibration patterns, and the advantage of generative fault classification in providing directly interpretable outputs. All experiments used the same dataset division, evaluation metrics, and training parameters, ensuring the fairness of the comparison and the reliability of the results.

[0044] The second experiment is a comprehensive comparison with existing methods, the purpose of which is to verify the superiority of the present invention over the existing technologies in the field of aviation bearing fault diagnosis. We selected a variety of representative deep learning architectures as baselines, including the learning architecture provided by the present invention - AeroGPT, a multi-scale attention one-dimensional CNN (MA1DCNN) designed for vibration signal analysis, deep residual networks (ResNet18 and ResNet50) that perform well in various classification tasks, long short-term memory networks (LSTM) that are good at capturing the temporal dependencies of sequence data, and a hybrid architecture (Conv-Transformer) that combines convolutional layers and Transformer self-attention mechanisms. In the experiment, we paid special attention to the balanced performance of the model on non-defective and defective samples, which is crucial for industrial applications because both high-precision fault detection is required to minimize expensive false alarms and high recall rates are required to prevent catastrophic failures. Through a comprehensive evaluation of multiple indicators, such as accuracy, precision, F1 score, AUC, and AP value, we are able to comprehensively compare the performance of each method on the task of aviation bearing fault diagnosis. At the same time, in order to more intuitively demonstrate the ability of each model to distinguish between normal and abnormal samples, we also conducted an anomaly score distribution analysis to evaluate the sensitivity and robustness of the model in identifying subtle abnormal signals.

[0045] Experimental results: Ablation Study Results Table 1 Ablation study results

[0046] The results of the ablation study clearly demonstrate the contributions of each component to the model performance. As shown in Table 1, when using the basic model (FM) alone, the performance is extremely limited, with an accuracy of only 14.87%, a precision of 6.31%, and an F1-score of 4.20%, which is basically equivalent to random guessing, highlighting a significant domain gap between general audio understanding and vibration signal fault diagnosis. When introducing the vibration signal alignment (VSA) component but not using generative fault classification, the performance improves - the accuracy increases to 20.65%, the precision improves to 12.47%, and the F1-score reaches 10.83%. Although this improvement is limited, it indicates that the VSA stage initially establishes a connection between general audio knowledge and vibration-specific patterns, laying the foundation for subsequent fault classification. When directly combining the generative fault classification (GFC) with the basic model, the performance is significantly improved. The FM+GFC configuration achieves an accuracy of 97.21%, a precision of 97.84%, and an F1-score of 97.52%, which are 82.34%, 91.53%, and 93.32% higher than using FM alone, respectively. This huge progress shows that even without explicit domain alignment, parameter-efficient fine-tuning can effectively adapt the basic model to the target task. The complete framework (FM+VSA+GFC) achieves the best performance in all metrics: an accuracy of 98.94%, a precision of 99.16%, and an F1-score of 99.02%. Compared with the FM+GFC configuration, the complete framework further improves the accuracy by 1.73%, the precision by 1.32%, and the F1-score by 1.50%, which is equivalent to reducing the misclassification rate by 62.1%, the precision error by 61.1%, and the F1-score defect by 60.5%. These results confirm the synergistic effect of the VSA stage, which, although providing limited benefits when used alone, can significantly enhance the effect of the subsequent GFC stage, improving the accuracy and stability of fault diagnosis by aligning the representations of the basic model with the vibration signal domain.

[0047] Results of comparison with existing methods

[0048] A comprehensive comparison with existing deep learning methods further validates the excellent performance of the present invention. As shown in Table 2, the proposed method significantly outperforms the comparative methods in all evaluation metrics. In terms of overall accuracy, the present method reaches 98.94%, exceeding the second-best method MA1DCNN (96.21%) by 2.73 percentage points. This improvement highlights the effectiveness of leveraging pre-trained audio knowledge for vibration signal analysis. Similar improvements are also observed in terms of precision and F1-score, with the present method reaching 99.16% and 99.02% respectively, while MA1DCNN is 96.31% and 96.22%. Notably, the present method demonstrates excellent and balanced performance in both non-defect and defect categories, with an accuracy of 99.46% for non-defect categories and 98.42% for defect categories, a difference of only 1.04 percentage points between them. This balance is crucial for industrial applications, which require both high-precision fault detection to minimize costly false alarms and high recall to identify defective components to prevent catastrophic failures. In contrast, traditional deep learning models show significant performance differences between these two types of samples. For example, MA1DCNN reaches an accuracy of 99.15% on non-defect samples but only 93.26% on defect samples, a difference of 5.89 percentage points. This difference is even more significant in other models, with the Conv-Transformer having an accuracy gap of up to 24.95 percentage points between these two categories (96.11% for non-defect vs. 71.16% for defect). The superior performance of the present method can be attributed to several key factors: First, the pre-training of the base model on diverse audio data provides rich acoustic representations that can capture general patterns relevant to vibration analysis; second, the vibration signal alignment stage effectively bridges the domain gap between general audio and specific vibration patterns, as demonstrated by ablation studies; finally, the generative fault classification method allows the model to utilize its inherent language generation ability to produce more informative and contextually relevant fault diagnosis results, not only improving classification accuracy but also enhancing the practical value and interpretability of the diagnostic output. These results together demonstrate the significant advantages of the method of the present invention in the field of aeroengine bearing fault diagnosis, providing a more reliable and practical technical solution for predictive maintenance.

[0049] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0050] The present invention is described with reference to the flowcharts and / or block diagrams of systems, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0051] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0052] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0053] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0054] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. An aero-engine bearing fault diagnosis method driven by an audio large model, characterized in that Including: Establish the recognition ability of a large - scale language model for audio input through multi - stage processing to obtain a pre - trained basic model, enabling the model to extract meaningful features from complex acoustic signals. The multi - stage processing includes basic acoustic understanding, supervised fine - tuning, and direct preference optimization; Construct a paired dataset of aero - engine bearing vibration signals and their detailed text descriptions, perform amplitude normalization and encoding on the vibration signals to obtain audio samples; establish a connection between the audio samples and the text modality through a cross - modal attention mechanism; among them, when encoding the audio, by setting the LoRA layer, modify the processing of the vibration signals so that the encoder can recognize the acoustic features of various bearing faults; Use the connected audio samples and text modality as the input of the pre - trained basic model, and utilize the generation ability of the model to directly generate text fault labels and diagnostic explanations in an autoregressive manner; among them, establish adaptation points based on the LoRA technique in the self - attention mechanism and the feed - forward network of the pre - trained basic model, so as to achieve knowledge transfer and adapt the general audio understanding ability of the model to the specific field of aero - engine bearing vibration signals.

2. The method for diagnosing faults of an aero-engine bearing driven by an audio large model according to claim 1, wherein: In the basic acoustic understanding stage, by converting audio input into text output, and through processing various types of audio input, the model learns to extract meaningful features from complex acoustic signals.

3. The method for diagnosing faults of an aero-engine bearing driven by an audio large model according to claim 1, characterized in that: The supervised fine - tuning stage further enhances the model's analysis ability for complex audio content; the direct preference optimization stage uses a comparison scoring mechanism to evaluate alternative responses to the same audio input, and the scores indicate their relative quality.

4. The method for diagnosing faults of an aero-engine bearing driven by an audio large model according to claim 1, wherein: The LoRA freezes the pre - trained weight matrix and introduces a trainable low - rank adaptation layer to achieve knowledge transfer and adapt the general audio understanding ability to the specific field of aero - engine bearing vibration signals.

5. The method for diagnosing faults of an aero-engine bearing driven by an audio large model according to claim 1, wherein: The detailed text description of the aero - engine bearing vibration signal specifically includes: vibration source, noise characteristics, fault type, severity, location, and relevant operating conditions.

6. The method for diagnosing faults of an aero-engine bearing driven by an audio large model according to claim 1, wherein: The pre - trained basic model directly generates text fault labels and diagnostic explanations in an autoregressive manner, and its generation process is as follows: ; wherein represents the token predicted at position t, represents the vocabulary space, represents the probability distribution for the model to predict the next token given the conditions, represents the encoding of the vibration signal, including all previously generated tokens.

7. The method for diagnosing faults of an aero-engine bearing driven by an audio large model according to claim 6, wherein: The training objective of model optimization is expressed as a cross - entropy loss function on the fault label marking sequence: ; Among them represents a fault diagnosis data set composed of vibration signal-label pairs (v, y), represents the probability of generating the correct label given the audio embedding and the previous tokens as described above.

Citation Information

Cited By

  • Vertical large model-based aeronautical manufacturing field fault diagnosis method and system

    CN121456409A