A Soft Contrastive Learning Fault Diagnosis Method Based on Question-Answering Dialogue of Large Language Models

The data enhancement strategy and layered soft contrast learning generated through the large language model question-and-answer dialogues solve the problem of data enhancement and negative sample comparison mechanism in self-supervised comparison learning, and realizes efficient and intelligent fault diagnosis of rotating mechanical components, which is suitable for equipment such as bearings and gears.

CN120145057BActive Publication Date: 2025-08-05WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510601474.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-05
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The existing self-supervised comparison learning small sample fault diagnosis method has defects in the insufficient adaptability of data enhancement technology and the design of negative sample comparison mechanism, resulting in poor feature extraction effect and classification boundary deviation, which is difficult to meet the diversified needs of rotating mechanical components.

Method used

Using a method based on the Q&A dialogue of the large language model, the adaptive data enhancement strategy is automatically generated through the thinking chain large language model, combined with hierarchical soft contrast learning, and using the instance-level and time-level soft label allocation mechanisms to perform pre-training and fine-tuning of the feature encoder to achieve intelligent fault diagnosis.

Benefits of technology

It improves data quality and model generalization performance, improves the accuracy and adaptability of fault diagnosis, and can achieve efficient and intelligent fault identification under the conditions of few samples. It is suitable for rotating mechanical equipment such as bearings and gears.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145057B_ABST
    Figure CN120145057B_ABST
Patent Text Reader

Abstract

The present invention discloses a soft contrast learning fault diagnosis method based on large language model question and answer dialogue, including: inputting the structural parameters and working conditions information of the target rotating machinery component into the large language model with a chain of thought to obtain the strong and weak data augmentation strategies and parameters adapted to this component; using a hierarchical soft label assignment mechanism to perform instance-level and time-level hierarchical measurement on the unlabeled vibration samples, and constructing an instance-level soft label matrix and a time-level soft label matrix respectively; inputting the strongly augmented samples, weakly augmented samples and their soft labels into the soft contrast learning module, and pre-training the feature encoder through hierarchical soft label feature refinement, cross prediction and context consistency tasks; adding a fault diagnosis classification head module to the pre-trained feature encoder, and fine-tuning with a small amount of labeled vibration samples to obtain a fault diagnosis model; performing fault diagnosis through the fault diagnosis model. The present invention can achieve fault diagnosis of automated data augmentation and hierarchical soft contrast learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault diagnosis of mechanical key components, and particularly relates to a soft contrast learning fault diagnosis method based on question-and-answer dialogue of a large language model. Background Art

[0002] Significant progress has been made in the intelligent fault diagnosis of large rotating machinery components based on deep learning. However, most current studies rely on a large amount of labeled data for supervised training, which is not practical in actual engineering applications. On the one hand, the process of analyzing and labeling a large amount of monitoring signals to determine the health category labels is extremely time-consuming and costly; on the other hand, it is difficult for mechanical systems to operate for a long time under fault conditions, resulting in extremely limited available fault data. This limitation restricts the popularization of deep learning methods in actual application scenarios and leads the research focus to the development of fault diagnosis technologies under few-shot conditions.

[0003] Although labeled fault data is scarce, advanced sensor acquisition systems can provide a large amount of unlabeled monitoring signals. These unlabeled signals are easy to obtain and contain rich fault-related information, which has important research value. Self-Supervised Learning (SSL) and Contrastive Learning (CL) have shown unique advantages in using unlabeled data. Specifically, SSL learns generalized feature representations by designing pre-training tasks (such as predicting masked regions, restoring shuffled sequences, etc.); CL constructs "positive sample pairs" and "negative sample pairs" and uses a contrastive loss function to guide the model to learn the differences between features, encouraging the feature vectors of "positive sample pairs" to be similar and keeping the feature vectors of "negative sample pairs" at a distance, so as to effectively extract discriminative features.

[0004] However, there are still some key problems to be solved in the existing self-supervised contrast learning few-shot fault diagnosis method (Self-Supervised Contrastive Learning for Few-shot Fault Diagnosis, SSCLFD):

[0005] (1) Conventional data augmentation techniques are not suitable for vibration signal features

[0006] Current SSCLFD methods generally rely on general data augmentation techniques such as random masking, cropping, noise addition, etc. However, these techniques are not optimized for the characteristics of mechanical component vibration signals, ignoring the unique patterns of vibration signals in terms of time, frequency characteristics, and sequential dependencies. Blindly adopting these augmentation methods may disrupt the inherent patterns of fault signals, resulting in the loss of time-frequency information and seriously affecting the feature extraction effect and the diagnostic performance of the model. More critically, designing specialized data augmentation strategies for different diagnostic objects usually requires relying on rich domain knowledge, and as the fault equipment or working conditions change, the data augmentation methods need to be frequently adjusted. This method is not only cumbersome but also lacks flexibility and intelligence, making it difficult to meet the diverse needs in practical engineering applications.

[0007] (2)Misjudgment of negative sample pairs and insufficient feature aggregation

[0008] In the process of constructing negative sample pairs in the existing contrastive learning framework, the potential similarity between negative samples in the same batch is not fully considered. Samples of the same type of fault may be wrongly labeled as "negative sample pairs" and subjected to unified "distant operations", causing these "false negative sample pairs" to be separated in the feature space, making the model learn irrelevant or misleading differences, resulting in the problem of insufficient intra-class aggregation ability, further leading to classification boundary deviation and reducing the accuracy of fault identification.

[0009] It can be seen that the specialized design of data augmentation techniques and the optimization of the negative sample contrast mechanism are the key bottlenecks restricting the performance of self-supervised contrastive learning for few-shot fault diagnosis. Therefore, aiming at the characteristics of mechanical component vibration signals, proposing a more intelligent and flexible data augmentation strategy and simultaneously improving the negative sample contrast mechanism to effectively enhance the feature aggregation of similar samples has become the key breakthrough point for improving the performance of the SSCLFD method. Summary of the Invention

[0010] To solve the problems of unreasonable data augmentation strategy generation and contrast mechanism in self-supervised contrastive learning under the condition of few samples of rotating machinery components, the present invention provides a soft contrastive learning fault diagnosis method based on question-and-answer dialogue of a large language model. Based on the extensive knowledge of the large language model LLM and the structural parameters and working condition information of the equipment to be diagnosed, a data augmentation strategy is automatically generated, combined with a hierarchical soft contrastive learning method at the instance level and time level to learn discriminative features from a large number of unlabeled samples, and a feature encoder Encoder is fine-tuned using a small number of samples to achieve fault diagnosis of various rotating machinery components under the conditions of few samples and insufficient expert knowledge.

[0011] In the first aspect of the present invention below, a soft contrastive learning fault diagnosis method based on question-and-answer dialogue of a large language model is provided, and this method includes:

[0012] Obtain a chain-of-thought large language model for fault diagnosis data augmentation strategy recommendations. Input the structural parameters and operating conditions of the target rotating machinery component into the chain-of-thought large language model to obtain strong and weak data augmentation strategies and parameters adapted to this component;

[0013] Obtain unlabeled vibration samples of the target rotating machinery component. Use a hierarchical soft label assignment mechanism to perform instance-level and time-level hierarchical metrics on the unlabeled vibration samples, respectively construct an instance-level soft label matrix and a time-level soft label matrix, and assign soft labels to the samples based on the constructed instance-level soft label matrix and time-level soft label matrix;

[0014] Perform data augmentation on the unlabeled vibration samples based on the strong and weak data augmentation strategies and parameters adapted to this component to obtain strongly augmented samples and weakly augmented samples, and input the strongly augmented samples, weakly augmented samples and their soft labels into the soft contrastive learning module to pre-train the feature encoder through hierarchical soft label feature refinement, cross prediction and context consistency tasks;

[0015] Add a fault diagnosis classification head module to the pre-trained feature encoder and fine-tune it using a small number of labeled vibration samples to obtain a fault diagnosis model;

[0016] Input the vibration samples of the target rotating machinery component into the fault diagnosis model to obtain a fault diagnosis result.

[0017] In some embodiments, obtaining a chain-of-thought large language model for fault diagnosis data augmentation strategy recommendations, inputting the structural parameters and operating conditions of the target rotating machinery component into the chain-of-thought large language model, and obtaining strong and weak data augmentation strategies and parameters adapted to this component includes:

[0018] In the conversational user interface, define the role of the large language model as an expert in the field of mechanical fault diagnosis, and carry out conversations and Q&A with it, completing four key steps step by step: putting forward basic requirements, clarifying precautions, refining implementation strategies, and unifying output formats. Based on the extensive knowledge base of the large language model and through the way of human-computer interaction dialogue prompts, let the general large language model understand the user's needs and make it have the ability of chain-of-thought reasoning;

[0019] Convert the Q&A dialogue based on the conversational user interface into a dialogue history corpus record that the large language model can understand, and this record contains the step-by-step prompt process of the entire data augmentation construction strategy;

[0020] The conversation history corpus is fed into a large language model. Through contextual learning and training, the large language model understands user needs and is able to recommend fault diagnosis data enhancement strategies. This model evolves from a general large language model into a specialized thinking chain large language model for fault diagnosis data enhancement strategies, called the LLM-FDDA expert. This model can gradually acquire structural parameters and operating condition information of the equipment to be diagnosed, calculate fault characteristic frequencies, reason about data enhancement methods, and recommend parameters.

[0021] The structural parameters and working condition information of the target rotating machinery parts are input into the thinking chain large language model. The thinking chain large language model will calculate the fault characteristic frequency according to the part object type and infer the appropriate strong and weak data enhancement strategies and their parameters.

[0022] In some embodiments, the method further comprises:

[0023] Based on the Thinking Chain Large Language Model and the fault diagnosis model, a rotating machinery component fault diagnosis system combining LLM and soft comparative learning is built. New structural parameters and operating condition information of the rotating machinery component to be diagnosed are input, and the system automatically generates new strong and weak data enhancement strategies and parameters, as well as fault diagnosis results, for fault identification of various rotating machinery components.

[0024] When the target rotating machinery component is a bearing, the structural parameters include the bearing's inner diameter, outer diameter, rolling element diameter, number of rolling elements, and contact angle; the operating condition information includes load, rotational frequency, temperature, and sampling frequency;

[0025] When the target rotating mechanical component is a gear, the structural parameters include the number of teeth and pressure angle, and the operating condition information includes load, rotational frequency, temperature, and sampling frequency.

[0026] In some embodiments, a hierarchical soft label assignment mechanism is used to perform instance-level and time-level hierarchical metrics on unlabeled vibration samples, construct an instance-level soft label matrix and a time-level soft label matrix, and assign soft labels to the samples based on the constructed instance-level soft label matrix and time-level soft label matrix, including:

[0027] Based on dynamic time warping (DTW), the distance between unlabeled vibration samples is calculated, and an instance-level distance matrix is constructed. The values on the diagonal of the distance matrix are smoothed and normalized. The normalized distance matrix is then converted into an instance-level similarity matrix by subtracting the distance from 1. Finally, the sigmoid function is used to assign an instance-level soft label matrix to the sample. :

[0028]

[0029] Where, Representation sample aand the sample b The distance between them is calculated by dynamic time warping (DTW); is the temperature parameter; is the weight parameter; is the diagonal element of the identity matrix;

[0030] According to the convolutional network structure of the feature encoder, calculate the time-axis distance after each layer of convolution, construct the time-level weight matrix, and perform a smooth transformation on the time-lag distance of the unlabeled vibration samples at different convolutional layers based on the sigmoid function to obtain the time-level soft label matrix :

[0031]

[0032] In the formula, p and q are the time-point indices; is the absolute difference between time points, representing the time-lag distance; is the temperature parameter.

[0033] In some embodiments, according to the convolutional network structure of the feature encoder, calculate the time-axis distance after each layer of convolution, and construct the time-level weight matrix, including:

[0034] For the feature map output by each layer of convolution, construct a distance matrix of size T×T by calculating the absolute distance dist = |p - q| between time points, which is called the time-level weight matrix; where T represents the time dimension size of the feature map of this layer;

[0035] The temperature parameter , where m is the kernel size of the pooling layer in different convolutional layers, k is the depth where the convolutional layer is located, is a hyperparameter.

[0036] In some embodiments, input the strongly augmented samples, weakly augmented samples, and their soft labels into the soft contrast learning module, and pre-train the feature encoder through hierarchical soft label feature refinement, cross-prediction, and context consistency tasks, including:

[0037] (1) Hierarchical soft label feature refinement task

[0038] Input the strongly augmented samples and weakly augmented samples into the feature encoder composed of n layers of convolutional networks, and calculate the instance-level soft contrast loss and the time-level soft contrast loss of the local features based on the local features obtained by each layer of convolutional network in the feature encoder , and the instance-level soft contrast loss and temporal soft contrast loss Accumulate layer by layer to get the final instance-level soft contrast loss and temporal soft contrast loss ;

[0039] (2) Cross-prediction task

[0040] Strongly enhanced samples output global features after passing through the feature encoder , from the global features Extract the feature representation of the first t time steps and input it into the autoregressive Transformer model to generate the context feature representation ; Using linear converter W k Mapping context feature representations to prediction representations at different time steps; where the linear transformer W k Corresponding to multiple future time steps k ;

[0041] For each future time step k , the predicted representation is compared with the real future time step in the weakly enhanced sample k The inner product calculation is performed on the feature representation of to obtain the similarity between the two. Combined with the Log-Softmax operation, the similarity is normalized and the negative logarithm is taken to obtain the contrast loss matrix;

[0042] By extracting the diagonal elements of the contrast loss matrix, the consistency between the predicted representation and the true representation in time steps, that is, the positive sample similarity, is calculated, and the loss is accumulated, and finally the Noise Contrastive Estimation loss is used. Calculate the cross-prediction loss from strong enhancement samples to weak enhancement samples ;in, B is the batch size; b The first b samples; i For the i time steps; b ’ Index of all samples in the same batch, including the current sample b and other samples, namely, negative sample candidate sets; is the representation of real future samples, is the representation of the predicted sample;

[0043] Similarly, calculate the cross-prediction loss from weakly enhanced samples to strongly enhanced samples , the cross prediction loss from strong enhanced samples to weak enhanced samples Cross-prediction loss from weakly enhanced samples to strongly enhanced samples The sum is used as the final cross - prediction loss ;

[0044] (3)Context consistency task

[0045] The context feature representation of strongly augmented samples and the context feature representation of weakly augmented samples respectively go through a non - linear projection head for dimensionality reduction, and the context consistency loss between the two features of the output is calculated through the Normalized Temperature - scaled Cross Entropy Loss : :

[0046]

[0047] In the formula, the numerator is the exponential form of the similarity of positive sample pairs, and the denominator is the sum of the exponentials of the similarities between the current sample and all other samples. N is the number of samples, represents a positive sample pair, that is, the strongly augmented sample and the weakly augmented sample of a certain sample, sim represents the cosine similarity function, is an indicator function to ensure that itself is not included; is the temperature parameter;

[0048] (4)Total loss

[0049] The hierarchical soft label feature refinement loss , cross - prediction loss and context consistency loss are weighted and summed to obtain the total loss function , and the formula is as follows:

[0050]

[0051] In the formula, , , are weights;

[0052] The feature encoder is pre - trained through the total loss function.

[0053] In some embodiments, the calculation method of the instance - level soft contrast loss is:

[0054]

[0055] In the formula, t is the time - point index, a is the sample index, and L is the loss matrix after softmax normalization of the instance - level similarity matrix and taking the negative logarithm.S L is the lower left triangular part of the instance-level soft label matrix, S R is the upper right triangular part of the instance-level soft label matrix, B is the batch size, and T is the time step;

[0056] Time-level soft contrast loss The calculation method is as follows:

[0057]

[0058] In the formula, b is the sample index in the batch, and are respectively the lower left and upper right triangular parts of the time-level soft label matrix.

[0059] In some of the embodiments, the fault diagnosis classification head module is a single-layer linear layer, which is responsible for mapping the high-dimensional feature representation to the probability distribution of the classification labels; the input dimension of the linear layer is equal to the product of the output channels of the last convolutional layer of the feature encoder and the time step length, and the output dimension is equal to the number of the final classification labels.

[0060] According to a second aspect of the present invention, there is provided a computer device, including: a processor and a memory, the memory stores a program or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the soft contrast learning fault diagnosis method for question-and-answer conversations based on large language models described in any one of the first aspects are implemented.

[0061] According to a third aspect of the present invention, there is provided a readable storage medium, on which a program or instructions are stored, and when the program or instructions are executed by the processor, the steps of the soft contrast learning fault diagnosis method for question-and-answer conversations based on large language models described in any one of the first aspects are implemented.

[0062] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0063] (1) The present invention enables the LLM to gradually learn and understand the user's needs from historical conversation questions and answers, and obtain the thinking chain reasoning ability to decompose complex problems. Using the large language model (LLM) with a thinking chain after training, combined with the structural parameters and working conditions information of rotating machinery, an appropriate vibration signal data enhancement strategy is automatically recommended to achieve intelligent and targeted strong and weak data enhancement, effectively improving the data quality and the generalization performance of the model.

[0064] (2) The present invention designs a self-supervised soft contrast learning strategy to mine potential discriminative feature representations from a large amount of unlabeled data, and combines a small number of labeled samples for fine-tuning to achieve high-precision fault diagnosis. At the same time, by using the time-level and instance-level soft label mechanisms, the similarity relationship between samples is effectively quantified, the possible mislabeling problem in the traditional negative sample construction process is overcome, and the feature clustering effect and boundary recognition ability are enhanced.

[0065] (3) The present invention is not only applicable to bearing-type mechanical equipment, but also can be adapted to other types of rotating mechanical equipment such as gears. The LLM with a chain of thought can infer the fault characteristic frequencies according to the structure and working conditions of the new equipment, and use its huge knowledge reserve to provide appropriate data augmentation methods. Combined with the proposed soft contrast learning strategy, fault diagnosis under few-shot conditions can be achieved, with good generality and scalability.

[0066] (4) The present invention greatly reduces the engineer's dependence on complex fault diagnosis and analysis means. By comparing with other advanced algorithms in real fault diagnosis tests, the superiority of the present invention in few-shot fault diagnosis tasks is verified. More importantly, the generality of the present invention is verified on different intelligent diagnosis model architectures, indicating that it can be adapted to various fault diagnosis model architectures and has a practical promotion prospect.

[0067] In summary, the present invention uses intelligent data augmentation based on LLM and a self-designed hierarchical soft contrast learning framework to provide an efficient, intelligent and easy-to-operate solution by automatically providing data augmentation, feature extraction and fault diagnosis processes. It can not only solve the bottleneck of few-shot fault diagnosis in actual engineering scenarios, but also has the beneficial effects of intelligence, precision, generalization and high efficiency and convenience, providing a feasible attempt direction for rotating machinery fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a schematic diagram of the overall process of a soft contrast learning fault diagnosis method based on large language model question and answer dialogue provided by an embodiment of the present application;

[0069] Figure 2 It is a training and working flow chart of an LLM with a chain of thought provided by an embodiment of the present application;

[0070] Figure 3 It is a schematic diagram of a pseudo-code example of data augmentation executed by an LLM provided by an embodiment of the present application;

[0071] Figure 4 It is a flow chart of an instance-level soft label assignment mechanism provided by an embodiment of the present application;

[0072] Figure 5Flowchart of a time-level soft label assignment mechanism provided by an embodiment of this application;

[0073] Figure 6 Framework diagram of a soft contrast learning provided by an embodiment of this application;

[0074] Figure 7 Schematic diagram of functional modules of a rotating machinery component fault diagnosis system based on LLM and soft contrast learning provided by an embodiment of this application;

[0075] Figure 8 Diagram of a fault diagnosis test device provided by an embodiment of this application; wherein, Figure 8 In (a) is the fault test platform for bearings and gears, Figure 8 In (b) is the bearing fault diagram, Figure 8 In (c) is the gear fault diagram;

[0076] Figure 9 Comparison chart of few-shot fault diagnosis performance under different algorithms provided by an embodiment of this application;

[0077] Figure 10 Schematic diagram of the hardware structure of a computer device provided by an embodiment of this application. Detailed implementation manners

[0078] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0079] Obviously, the accompanying drawings in the following description are only some examples or embodiments of this application. For those of ordinary skill in the art, without creative efforts, this application can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in this application, some design, manufacturing or production changes based on the technical content disclosed in this application are only conventional technical means and should not be understood as the content disclosed in this application being insufficient.

[0080] Reference to "embodiment" in the present application means that the specific features, structures or characteristics described in combination with the embodiment can be included in at least one embodiment of the present application. The phrase appears at various positions in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in the present application can be combined with other embodiments without conflict.

[0081] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the ordinary meaning understood by those with ordinary skills in the technical field to which the present application pertains. The words such as "a", "an", "one", "the" and the like involved in the present application do not indicate a quantity limitation and can represent a singular or plural number. The terms "comprising", "including", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The term "plurality" involved in the present application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third" and the like involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0082] Currently, the existing self-supervised contrastive learning small-sample fault diagnosis method (SSCLFD) has the following problems:

[0083] (1) Insufficient adaptability of data augmentation techniques

[0084] Conventional data augmentation techniques, such as random masking, cropping, etc., fail to fully consider the unique laws of vibration signals in terms of time, frequency characteristics and sequential dependence. This blind augmentation may damage the inherent pattern of vibration signals, resulting in loss of feature information. In addition, the existing methods cannot automatically adjust the augmentation strategy according to the changes in the type of faulty equipment or working conditions, and rely on domain experts to design data augmentation methods specifically, increasing the implementation difficulty of fault diagnosis and restricting the improvement of the intelligent level.

[0085] (2) Defects in the negative sample contrast mechanism

[0086] Existing contrastive learning mechanisms ignore the similarities between samples and apply a uniform "away operation" to all negative pairs. This mechanism can easily mislabel samples of the same type within a batch as "negative pairs," causing them to be separated in the feature space. The model learns irrelevant or misleading feature differences rather than effective discriminative representations of fault characteristics. This flaw results in insufficient intra-class aggregation capabilities and affects the accuracy of classification boundaries.

[0087] The difficulty of solving the above problems is:

[0088] (1) The dependence of data augmentation method design on professional knowledge

[0089] Designing automated and appropriate data augmentation methods tailored to the specific characteristics of vibration signals typically requires the in-depth involvement and extensive expertise of fault diagnosis experts. This dependency not only increases the difficulty of practice but also limits the method's intelligence and practical applicability.

[0090] (2) Differentiation processing strategy without negative sample comparison mechanism

[0091] Existing technologies lack effective metrics for measuring the similarity of negative sample pairs, and no negative sample comparison strategy capable of performing differentiated processing has been developed. This results in the model being unable to accurately cluster similar samples and easily falling into the dilemma of learning misleading features.

[0092] To this end, this application provides a soft contrastive learning fault diagnosis method based on a large language model question-answering dialogue. The method comprises the following steps: First, leveraging the extensive knowledge base of a large language model (LLM), a human-computer interactive dialogue prompt enables the general LLM to understand user needs, equipping it with chain-of-thought reasoning capabilities, and then providing targeted vibration data enhancement strategies and parameters. This reduces the need for specialized knowledge while automatically integrating the geometric parameters and operating condition information of mechanical components to provide appropriate data enhancement strategies. Second, a hierarchical soft contrastive learning strategy is used to hierarchically process unlabeled samples at the instance and temporal levels. By quantifying the similarity of negative sample pairs from the same batch, it effectively prevents similar negative sample pairs from being misclassified, focuses on extracting discriminative features between classes, and improves the feature clustering effect and the accuracy of classification boundaries. Finally, a small number of labeled samples are used to fine-tune the model to achieve fault diagnosis in various scenarios. This method can implement fault diagnosis using automated data enhancement and hierarchical soft contrastive learning even with limited sample size and expert knowledge, demonstrating excellent versatility and scalability.

[0093] The present application provides a soft contrastive learning fault diagnosis method based on a large language model question-answering dialogue, comprising the following steps:

[0094] 1) Input the structural parameters and operating condition information of rotating machinery components into a large language model (LLM) that already understands user needs and has a thought chain, and output a dedicated data enhancement strategy and appropriate parameters that meet the needs of the component;

[0095] 2) Using a hierarchical metric method at the instance level and time level, we calculate the instance-level distance of all training samples. At the same time, we construct a time-level distance matrix and assign soft labels to samples at different levels.

[0096] 3) The strong and weak enhancement samples and their soft labels of all training samples are jointly input into the soft contrastive learning module. Based on the three preconditions of hierarchical soft label feature refinement, cross prediction, and context consistency, the feature encoder is pre-trained to mine discriminative representations between unlabeled samples.

[0097] 4) Add a fault diagnosis classification head module to the pre-trained feature encoder, named LLM-COTAugSCLFD, and fine-tune it using a small number of labeled samples;

[0098] 5) Freeze the model weights of the fine-tuned LLM-COTAugSCLFD and test the classification performance of the model on real samples;

[0099] 6) Based on the above-mentioned method, a rotating machinery component fault diagnosis system combining LLM and soft comparative learning is constructed. When new structural and operating condition information of the rotating machinery component to be diagnosed is input, the system can automatically generate new data enhancement strategies and fault diagnosis results for fault identification of various rotating machinery, with versatility and scalability.

[0100] The LLM, which features thought chaining, gradually acquires contextual thought chaining capabilities by learning from historical question-and-answer (Q&A) dialogue corpus. This process does not require engineers to possess in-depth expertise in fault diagnosis and signal processing. Simply by providing the LLM with the structural parameters and current operating condition information of the component being diagnosed, the LLM can obtain data augmentation strategies and parameters appropriate for the component.

[0101] When the object to be diagnosed is a component of a bearing-type machinery, the structural parameters refer to the inner diameter, outer diameter, rolling element diameter, number of rolling elements and contact angle of the bearing; when the object to be diagnosed is a component of a gear-type machinery, the structural parameters refer to the number of teeth and pressure angle; the operating condition information includes load, rotational frequency, temperature and the sampling frequency of the monitoring signal.

[0102] The dedicated data augmentation strategies include two types: strong augmentation is the multi-layer frequency band coefficient augmentation method in the wavelet domain; weak augmentation is the Fourier frequency domain amplitude scaling augmentation method. The parameters are the hyperparameters required in their respective augmentation methods. The parameters in strong augmentation include the wavelet basis function and the scale coefficient; while the parameter in weak augmentation is the amplitude scaling coefficient.

[0103] The strongly and weakly augmented samples are obtained by two dedicated data augmentation strategies, that is, each sample has two perspectives: strongly augmented samples and weakly augmented samples.

[0104] The instance-level hierarchical metric method is the actual-level similarity matrix between the original vibration data based on dynamic time warping (DTW). The time-level hierarchical metric is to smooth and transform the time lag distance of different convolutional layers of the Encoder based on the Sigmoid function to obtain the time-level soft label matrix, which measures the feature weights extracted by each convolutional layer.

[0105] The soft contrast learning module mainly includes: vibration data augmentation, Encoder feature encoding, hierarchical soft label feature refinement, cross prediction, and context consistency. The pre-training of the feature encoder Encoder is to carry out self-supervised contrast learning based on a large number of unlabeled samples to improve the representation ability of the autoregressive model for sample features and thus improve the performance in downstream classification tasks. The weight parameters of the pre-trained feature encoder Encoder in step 4) are inherited from the weight parameters and state dictionary saved in the last step of step 3). The main advantage of fine-tuning is that it can use the pre-trained weights as a good initialization, which speeds up the convergence rate, and only a small number of labeled samples can be used for efficient training for downstream tasks.

[0106] The fault diagnosis classification head module is a linear layer, which is responsible for mapping the high-dimensional feature representation to the probability distribution of classification labels. The input dimension of the linear layer is equal to the product of the output channel number of the last convolutional layer and the time step length, and the output dimension is equal to the number of final classification categories.

[0107] The fault diagnosis system for rotating machinery components based on LLM and soft contrast learning has functions and advantages such as intelligent data augmentation, self-supervised soft contrast learning, generality and scalability, and convenience in engineering applications. The training process of the LLM with the chain of thought is as follows:

[0108] a. The user defines the role of GPT in the conversational user interface (CUI) and conducts conversations and Q&A, completing four key steps step by step: putting forward basic requirements, clarifying precautions, refining execution strategies, and unifying the output format. This step-by-step chain of conversations decomposes complex problems into multi-step reasoning, and finally has the function of suggesting fault diagnosis data augmentation strategies.

[0109] b. Convert the Q&A dialogue (Q1A1~QnAn) from the previous step into a dialogue history corpus that can be understood by the LLM. This corpus records the step-by-step prompting process of the entire data augmentation construction strategy, enabling the chain-of-thought LLM to gradually complete the acquisition of the structural parameters and operating conditions of the device to be diagnosed, the calculation of fault characteristic frequencies, the inference of data augmentation methods, and parameter suggestions.

[0110] c. Call the application programming interface (API), input the dialogue history corpus into the LLM, and enable the LLM to understand the user's needs from it and possess the ability to perform step-by-step inference on the data augmentation strategy in this specific field of fault diagnosis through few-shot learning.

[0111] c. The general LLM (GPT4) evolves into a specific model dedicated to suggesting data augmentation strategies for fault diagnosis, called the LLM-FDDA expert.

[0112] d. The user provides the structural parameters and operating conditions of the components of the device to be diagnosed or a new device to the LLM-FDDA expert, which will automatically calculate the fault characteristic frequencies according to the component object type and infer appropriate strong and weak data augmentation strategies and their parameters.

[0113] This method leverages the extensive knowledge base of the large language model (LLM). Through the way of dialogue prompting, it enables the general LLM to understand the user's needs, endows it with the ability of chain-of-thought reasoning, and specifically gives the data augmentation strategy for vibration data and the parameters therein. While reducing the need for professional knowledge, it automatically integrates the geometric parameters and operating conditions of mechanical components to provide appropriate data augmentation strategies, which not only ensures the diversity of self-supervised samples but also maintains the time dependence of vibration signals. In addition, the proposed hierarchical soft contrastive learning strategy performs hierarchical processing on unlabeled samples at the instance level and time level. By quantifying the similarity of negative sample pairs in the same batch, it effectively prevents similar negative sample pairs from being wrongly separated, avoids the model learning irrelevant or misleading features, focuses on the extraction of discriminative features between classes, and improves the feature clustering effect and the accuracy of the classification boundary. This method can achieve automated data augmentation and hierarchical soft contrastive learning-based fault diagnosis under the conditions of few samples and insufficient expert knowledge. It is not only applicable to the bearing state recognition of single faults and compound faults but can also be extended to other mechanical equipment such as gears, with good generality and scalability. In summary, the proposed method effectively overcomes the limitations of the existing technology and provides an efficient, intelligent, and highly adaptable solution for rotating machinery fault diagnosis under the background of insufficient expert knowledge and small samples.

[0114] Specifically, as Figure 1 shown, the soft contrastive learning fault diagnosis method based on large language model Q&A dialogue in the embodiment of the present invention specifically includes the following steps:

[0115] 1) The user defines the role of the GPT as an expert in the field of mechanical fault diagnosis in the conversational user interface (CUI) and conducts a dialogue and Q&A with the GPT, such as Figure 2 The specific steps are as follows:

[0116] a. Propose basic requirements: Ask GPT the following question: "When using self-supervised contrastive learning to perform fault diagnosis tasks on rotating machinery vibration signals, a single sample is often converted into a strongly enhanced and a weakly enhanced sample. How can we use knowledge in the specific field of fault diagnosis to enhance the vibration signal? Please provide some appropriate data augmentation methods."

[0117] b. Clarify the precautions: In response to GPT's question, "Please recommend several data enhancement methods that are suitable for rotating machinery vibration signal fault diagnosis, especially those that cannot destroy the timing dependency of the vibration signal. Furthermore, the data enhancement methods provided should preferably be related to traditional data analysis methods in the field of fault diagnosis."

[0118] c. Refine the execution strategy: Ask GPT, "Based on the relevant data enhancement methods for time series comparative learning, please provide me with a strong / weak enhancement method suitable for the vibration signals of rotating machinery bearings. In particular, I hope to utilize fault diagnosis and analysis methods, such as FFT transform, wavelet transform, bandpass filtering, and other data processing methods mentioned above. In addition, during the data enhancement process, it is required to maintain periodic fault characteristics."

[0119] d. Unified output format: Question to GPT: "The bearing information is: pitch diameter 31.00mm, ball diameter 6.25mm, number of balls 9, contact angle 0°, rotational speed 500rpm, sampling frequency 10k. Failure modes include inner race, outer race, and rolling element failures. Please recommend a suitable data augmentation method and the parameters used, and provide the corresponding Python code."

[0120] e. Convert the CUI-based conversation history (Q1A1 to QnAn) into a conversation history corpus that LLM can understand, including a step-by-step guide to building a data augmentation strategy.

[0121] f. Call the application programming interface (API) to input the conversation history corpus into the LLM. Through contextual learning and training, the LLM can understand user needs and gradually master the ability to decompose complex problems into multi-step reasoning. Ultimately, it has the function of recommending fault diagnosis data augmentation strategies. It grows from the general GPT4 to a specific model dedicated to fault diagnosis data augmentation strategy recommendations, called LLM-FDDA Expert.

[0122] g. Provide the structural parameter of components of the device to be diagnosed or new device (inner diameter, outer diameter, diameter of rolling elements, number of rolling elements, contact angle, etc. of bearings) and operating conditions information (load, rotational frequency, temperature, and sampling frequency of monitoring signals) as the "current problem" to the LLM-FDDA expert, such as Figure 3 shown, which will automatically calculate the fault characteristic frequency according to the component object type, and infer the appropriate data augmentation strategy of strong and weak and its parameters.

[0123] 2) Use the soft label assignment mechanism to assign instance-level and time-level soft labels to unlabeled vibration samples. The specific steps are as follows:

[0124] a. Based on the dynamic time warping technology, as Figure 4 shown, calculate the distance in the data space between the unlabeled original vibration data, construct the distance matrix between training samples, smooth the values on the diagonal of the distance matrix and perform normalization of the distance matrix, and finally convert it to an instance-level similarity matrix through 1 - distance. Among them, distance and similarity are complementary relationships. The greater the distance, the smaller the similarity. Then, based on the instance-level similarity matrix, combine the sigmoid function to assign the instance-level soft label matrix . Among them, represents the distance between sample a and sample b , which is calculated by DTW; is the temperature parameter, which controls the smoothness of the soft label. When it is larger, the label distribution is more extreme; is the weight parameter, which controls the influence of the diagonal. When it is larger, the label is more concentrated on the diagonal (self-similarity); is the diagonal element of the identity matrix (representing self-similarity).

[0125] b. According to the convolutional network structure of the feature encoder Encoder, as Figure 5 shown, after each layer of convolution, calculate the time-axis distance matrix of the feature map of this layer. Specifically, for the feature map (shape is [B, C, T], where T is the time dimension) output by each layer of convolution, construct a distance matrix of size T×T by calculating the absolute distance dist = |p - q| between time points, which is also called the time-level weight matrix. Then use the sigmoid function to convert the distance to the time-level soft label matrix , which performs a smooth transformation on the time lag distance through the Sigmoid function, so that the weights of adjacent time points are larger, and the weights of time points far away are smaller. Among them, p and q are the time point indices; is the absolute difference between time points, representing the lag distance; is the temperature parameter, which controls the distribution smoothness of the time weight and is related to the network structure of the Encoder. , m is the kernel size of the pooling layer in different convolutional layers. k is the depth where the convolutional layer is located. is a hyperparameter. It can be found that as the depth of the convolutional layer increases, the similarity between adjacent time steps decreases, and larger soft labels should be assigned to compensate for the semantic differences between adjacent time steps.

[0126] 3) Use a large number of unlabeled vibration samples as training samples to carry out soft contrast learning. As shown in Figure 6 , mine the discriminative representations through the three established upstream pretext tasks (hierarchical soft label feature refinement, cross-prediction, and context consistency). The specific steps are as follows:

[0127] a. Use the soft label assignment method provided in step 2) to precompute or cache the instance-level soft label matrix between all original unlabeled vibration samples offline and the time-level soft label matrix;

[0128] b. Divide all training samples into several batches and execute the soft contrast learning module for each batch separately;

[0129] c. Use the data augmentation method proposed in step 1) to generate strong and weak augmented samples for this batch of training samples respectively. The relationship between the augmented samples inherits from the soft label matrix of the original training samples;

[0130] d. All augmented samples (strong and weak augmented samples) in this batch are feature-encoded by the Encoder (consisting of a three-layer convolutional network, gradually extracting the local and global features of the training samples) to obtain the local and global high-dimensional representations , is the high-dimensional representation of the strong augmented sample, is the high-dimensional representation of the weak augmented sample;

[0131] e. Extract the local representations output by each convolutional network layer of the Encoder, and combine with the soft label matrix to perform hierarchical soft label feature refinement at the time and instance levels respectively, and calculate the hierarchical soft label feature refinement loss of the local features, where is the weighting coefficient, is the instance-level soft contrast loss, is the time-level soft contrast loss.

[0132] All enhanced samples (strong and weak enhanced samples) in this batch are respectively passed through the Encoder, and the local features obtained by each layer of the convolutional network are all subject to instance-level and temporal-level loss evaluation. The final instance-level and temporal-level losses are respectively accumulated by the instance-level and temporal-level losses of each layer of the three-layer convolutional network. This multi-level loss calculation design is to ensure that the model can learn good temporal correlation and instance correlation representations at different feature extraction levels. The specific instance-level and temporal-level loss calculation methods are as follows:

[0133] The calculation method of the instance-level soft contrast loss is . Among them, t is the time point index, a is the sample index, L is the loss matrix after softmax normalization of the instance-level similarity matrix and taking the negative logarithm, S L is the lower left triangular part of the instance-level soft label matrix, S R is the upper right triangular part of the instance-level soft label matrix, B is the batch size, T is the time step; using the upper and lower triangular parts of the soft label matrix S L and S R to weight the softmax loss and calculate the instance-level soft contrast loss. Through the instance-level soft label, the model can learn the similarity relationship between samples.

[0134] The calculation method of the temporal-level soft contrast loss is . Among them, b is the sample index in the batch, and are respectively the lower left and upper right triangular parts of the temporal-level soft label matrix, and the other parameters are the same as . Using the upper and lower triangular parts of the temporal-level soft label matrix to weight the softmax loss and calculate the temporal-level soft contrast loss. Through the time lag relationship, the model can learn the dependence between different time points.

[0135] f. The features of the strong and weak enhanced samples after passing through the feature encoder Encoder are used as global features to perform the cross-prediction task. The cross-prediction interface task is divided into two parts, strong enhanced sample → weak enhanced sample and weak enhanced sample → strong enhanced sample , and the steps of the two are similar. Taking the calculation of as an example, the steps are as follows:

[0136] ① From the global features output by the strong enhanced sample after passing through the Encoder Extract the feature representations of the first t time steps and input them into an autoregressive Transformer model to generate context feature representations . Subsequently, use a set of linear transformers W k (corresponding to multiple future time steps k ) to map to the prediction representations at different time steps.

[0137] ② For each future time step k , calculate the inner product between the representation predicted by the model and the feature representation of the true future time step k in the weakly augmented sample to obtain their similarity. Combine with the Log-Softmax operation to normalize the similarity and take the negative logarithm to obtain the contrastive loss matrix.

[0138] ③ By extracting the diagonal elements, calculate the consistency (positive sample similarity) between the predicted representation and the true representation at the time step, and accumulate the losses. Finally, obtain through the Noise Contrastive Estimation loss . Among them, B is the batch size; b is the b th sample (positive sample) in the current batch; b ’ is the index of all samples in the same batch (including the current sample b and other samples), that is, the negative sample candidate set; is the representation of the true future sample, is the representation of the predicted sample.

[0139] Cross-prediction loss from weakly augmented samples to strongly augmented samples Similarly, it can also be calculated. Take the sum of and as the cross-prediction loss of the training samples . By minimizing , the model can effectively learn the temporal dependence relationship of time series data.

[0140] g. The process of the context consistency pretext task is as follows:

[0141] ① The global feature representations of the strongly and weakly augmented samples generated by the autoregressive Transformer and Can be used as their respective context feature representations, and respectively undergo dimensionality reduction processing through a non-linear projection head, combined with the Normalized Temperature-scaled Cross Entropy Loss (NT-Xent Loss) loss Calculate the context consistency loss between the two features output , by minimizing , improve the similarity between the context feature representations of strong and weak augmented samples, so that strong and weak augmented samples are predicted as the same class of samples as much as possible, and improve the robustness of the features extracted by the network, where The calculation method of is as follows:

[0142]

[0143] In the formula, the numerator is the exponential form of the similarity of positive sample pairs, and the denominator is the sum of the exponential forms of the similarities between the current sample and all other samples (including negative samples), N is the number of samples, represents a positive sample pair, that is, the strong and weak augmented samples of a certain sample, sim represents the cosine similarity function, is an indicator function to ensure that itself is not included; is the temperature parameter, used to adjust the sensitivity of similarity calculation.

[0144] ② By minimizing , improve the context feature representations of strong and weak augmented samples between the similarities, so that strong and weak augmented samples are predicted as the same class of samples as much as possible, and improve the robustness of the features extracted by the network.

[0145] h. Refine the loss of the hierarchical soft label features of the training samples , the cross-prediction loss and the context consistency loss are weighted and summed, and the formula is as follows:

[0146]

[0147] i. Calculate the gradient of the loss with respect to each parameter of the Encoder through the Chain Rule, update the gradient and backpropagate, and use the optimizer Adam to update the parameters to optimize the feature extraction ability of the Encoder. When the predetermined number of iterations is reached, complete the pre-training of the feature encoder Encoder, and save its weight parameters and state dictionary

[0148] 4) Use a small number of labeled vibration samples to fine-tune the model, such as Figure 1As shown in Step 4, the specific steps are as follows:

[0149] a. Load the pre-trained weights, which can provide a good initialization for the training of downstream tasks, accelerate convergence, and reduce the amount of data required for training;

[0150] b. Delete the classification layer weights of the original model and add a fault diagnosis classification head module (linear layer) to complete the training of the fault diagnosis classification task;

[0151] c. Use a small number of labeled samples to update the LLM-COTAugSCLFD weights through supervised fine-tuning training;

[0152] d. Save the weights of the fine-tuned LLM-COTAugSCLFD.

[0153] 5) Use the vibration samples in the test set to test the performance of the fine-tuned model. The specific steps are as follows:

[0154] a. Set the model to evaluation mode, freeze the model weights to turn off gradient calculation, save video memory and accelerate inference;

[0155] b. Initialize the evaluation metrics, including the loss of all batches, the accuracy of all batches, etc.;

[0156] c. Perform forward calculation through the data in the validation set;

[0157] d. Use the cross-entropy loss function to calculate the loss of the classification task, and store the categories predicted by the LLM-COTAugSCLFD model and the true labels;

[0158] e. Return the average loss and accuracy, and print the classification report to evaluate the overall performance of the model on the validation set.

[0159] 6) Based on the method described in the embodiments of this application, establish Figure 7 The rotating machinery component fault diagnosis system based on LLM and soft contrast learning shown in Figure 8 Taking the fault diagnosis test device shown in Figure 8 As an example, (a) in Figure 8 is the fault test platform for bearings and gears, Figure 8 in (b) is the bearing fault diagram,

[0160] a. Select the type of the component to be diagnosed (bearing or gear);

[0161] b. Extract the structural parameters and working conditions information of the target component and give it to the LLM-FDDA expert, who will give targeted data augmentation strategies and suggestions on their important parameters based on this information;

[0162] c. Collect vibration monitoring signals without labels as training samples. First, perform data augmentation, then calculate the soft labels of the original data, and finally conduct self-supervised soft contrastive learning and fine-tuning to obtain the LLM-COTAugSCLFD fault model;

[0163] d. Conduct real-time fault diagnosis of vibration monitoring signals in a real environment and evaluate the performance of the model.

[0164] Among them, the rotating machinery component fault diagnosis system with LLM and soft contrastive learning has the following functions:

[0165] Intelligent data augmentation: The LLM with a chain of thought can automatically recommend appropriate vibration signal data augmentation methods based on the structural parameters and working conditions of bearing-type mechanical equipment, improving data quality.

[0166] Self-supervised soft contrastive learning: Through the designed self-supervised soft contrastive learning strategy, the system can fully explore the potential discriminative feature representations in a large amount of unlabeled data. Combine a small number of labeled samples for fine-tuning to achieve high-precision fault diagnosis.

[0167] Generality and scalability: When the object to be diagnosed switches from bearing-type equipment to gear-type mechanical equipment, the system is still effective. The LLM can automatically combine the structural parameters and working conditions of gear equipment, infer the fault characteristic frequencies, and recommend suitable data augmentation methods.

[0168] Engineering application convenience: The system significantly reduces the need for engineers to master complex fault diagnosis analysis methods and provides an intelligent, general, and efficient fault diagnosis solution.

[0169] 7) To prove the superiority and generality of the method of this application, the model was fine-tuned with 5, 10, and 20 samples respectively, and the classification accuracies under different fault diagnosis algorithms were compared. As Figure 9 shown, the method described in this invention can still achieve an accuracy of about 90% under the condition of very few samples (5). More importantly, as shown in Table 1, on different intelligent diagnosis model architectures (MLP, CNN, LSTM, and Transformer), the strategy proposed in this invention can improve the classification performance of the original fault diagnosis model, verifying the generality of this invention and indicating that it can be adapted to various fault diagnosis model architectures and has a practical promotion prospect.

[0170] Table 1 Performance Table

[0171]

[0172] In addition, combined with Figure 1The soft contrast learning fault diagnosis method based on large language model question and answer dialogue described in the embodiments of the present application can be implemented by a computer device. Figure 10 It is a schematic hardware structure diagram of the computer device in the embodiments of the present application. As Figure 10 shown, the device may include a processor 301 and a memory 302 storing computer program instructions.

[0173] Specifically, the above-mentioned processor 301 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0174] Among them, the memory 302 may include a mass memory for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In appropriate cases, the memory 302 may include removable or non-removable (or fixed) media. In appropriate cases, the memory 302 may be internal or external to the data processing device. In a particular embodiment, the memory 302 is non-volatile memory. In a particular embodiment, the memory 302 includes a read-only memory (ROM) and a random access memory (RAM). In appropriate cases, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory (FLASH), or a combination of two or more of these. In appropriate cases, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended date out dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0175] The memory 302 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 301.

[0176] The processor 301 reads and executes the computer program instructions stored in the memory 302 to implement any one of the soft contrast learning fault diagnosis methods for question-and-answer conversations based on large language models in the above embodiments.

[0177] In some embodiments, the point cloud generation device may further include a communication interface 303 and a bus 300. Among them, as Figure 10 shown, the processor 301, the memory 302, and the communication interface 303 are connected through the bus 300 to complete mutual communication.

[0178] The communication interface 303 is used to implement communication between various modules, devices, units, and / or devices in the embodiments of the present application. The communication interface 303 can also implement data communication with other components, such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations, etc.

[0179] The bus 300 includes hardware, software, or both, and couples components of the point cloud generation device to each other. The bus 300 includes, but is not limited to, at least one of the following: Data Bus, Address Bus, Control Bus, Expansion Bus, Local Bus. By way of example and not limitation, the bus 300 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable bus or a combination of two or more of these. In suitable cases, the bus 300 may include one or more buses. Although embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0180] The computer device can execute the soft contrast learning fault diagnosis method for question-and-answer dialogue based on the large language model in the embodiments of the present application based on the rendering device, so as to implement the combination Figure 1 of the described soft contrast learning fault diagnosis method for question-and-answer dialogue based on the large language model.

[0181] In addition, in combination with the soft contrast learning fault diagnosis method based on large language model question-and-answer dialogue in the above embodiments, an embodiment of the present application can provide a computer-readable storage medium to implement. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the soft contrast learning fault diagnosis methods based on large language model question-and-answer dialogue in the above embodiments is implemented.

[0182] In summary, the present invention calls the API to enable the LLM to gradually learn and understand the user's needs from historical dialogue Q&A, and obtain the ability of thinking chain reasoning for decomposing complex problems. By using the large language model (LLM) with a chain of thought after training, combined with the structural parameters and working conditions information of rotating machinery, it automatically recommends an appropriate vibration signal data enhancement strategy to achieve intelligent and targeted strong and weak data enhancement, effectively improving the data quality and the generalization performance of the model.

[0183] In addition, the present invention designs a self-supervised soft contrast learning strategy to mine potential discriminative feature representations in a large amount of unlabeled data, and combines a small number of labeled samples for Fine-Tuning to achieve high-precision fault diagnosis. At the same time, by using the time-level and instance-level soft label mechanisms, the similarity relationship between samples is effectively quantified, the possible mislabeling problem in the traditional negative sample construction process is overcome, and the feature clustering effect and boundary recognition ability are enhanced.

[0184] It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should be considered as the scope described in this specification. In addition, according to the needs of implementation, each step / component described in the present application can be split into more steps / components, or two or more steps / components or partial operations of steps / components can be combined into new steps / components to achieve the purpose of the present invention.

[0185] Those skilled in the art can easily understand that the above embodiments only represent several implementation manners of the present application, and the description is relatively specific and detailed, but it cannot be understood as a limitation on the scope of the invention patent. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A soft contrastive learning fault diagnosis method based on a large language model question-answering dialogue, characterized in that: The method includes: Obtain a thinking chain large language model for fault diagnosis data enhancement strategy recommendations, input the structural parameters and operating condition information of the target rotating machinery component into the thinking chain large language model, and obtain strong and weak data enhancement strategies and parameters suitable for the component; Obtain unlabeled vibration samples of target rotating machinery parts, perform instance-level and time-level hierarchical measurements on the unlabeled vibration samples using a hierarchical soft label assignment mechanism, construct instance-level and time-level soft label matrices, and assign soft labels to the samples based on the constructed instance-level and time-level soft label matrices. Based on the strong and weak data augmentation strategies and parameters adapted to the component, the unlabeled vibration samples are augmented to obtain strongly and weakly enhanced samples. These samples and their soft labels are then fed into the soft contrastive learning module. The feature encoder is pre-trained through layered soft label feature refinement, cross-prediction, and context consistency tasks. Add a fault diagnosis classification head module to the pre-trained feature encoder and fine-tune it using a small number of labeled vibration samples to obtain a fault diagnosis model; Inputting vibration samples of target rotating machinery parts into a fault diagnosis model to obtain fault diagnosis results; The hierarchical soft label assignment mechanism is used to perform instance-level and time-level hierarchical measurements on unlabeled vibration samples, constructing instance-level soft label matrices and time-level soft label matrices respectively. Soft labels are then assigned to samples based on the constructed instance-level soft label matrices and time-level soft label matrices, including: Based on dynamic time warping (DTW), the distance between unlabeled vibration samples is calculated, and an instance-level distance matrix is constructed. The values on the diagonal of the distance matrix are smoothed and normalized. The normalized distance matrix is then converted into an instance-level similarity matrix by subtracting the distance from 1. Finally, the sigmoid function is used to assign an instance-level soft label matrix to the sample. : Where, Representation sample a and samples b The distance between them is calculated by dynamic time warping DTW; is the first temperature parameter, used to control the smoothness of the soft label; is the weight parameter; are the diagonal elements of the identity matrix; According to the convolutional network structure of the feature encoder, the time axis distance is calculated after each convolution layer, the time-level weight matrix is constructed, and the time lag distance of the unlabeled vibration samples at different convolution layers is smoothly converted based on the sigmoid function to obtain the time-level soft label matrix : Where, p and q Index for time points; is the absolute gap between time points, indicating the time lag distance; is the second temperature parameter, which is used to control the smoothness of the distribution of time weight.

2. The soft contrastive learning fault diagnosis method based on large language model question-answering dialogue according to claim 1 is characterized in that: Obtain a thinking chain large language model for fault diagnosis data enhancement strategy recommendations. Input the structural parameters and operating condition information of the target rotating machinery component into the thinking chain large language model to obtain strong and weak data enhancement strategies and parameters suitable for the component, including: In the conversational user interface, the large language model is defined as an expert in the field of mechanical fault diagnosis. A dialogue and question-and-answer session is conducted with the model, completing four key steps: establishing basic requirements, clarifying precautions, refining execution strategies, and standardizing output formats. Based on the large language model's extensive knowledge base and through interactive dialogue prompts, the general large language model understands user needs and develops the ability to reason through thought chains. Convert question-answering conversations based on conversational user interfaces into conversation history records that can be understood by large language models. This record includes step-by-step instructions for building a data augmentation strategy. The conversation history corpus is fed into a large language model. Through contextual learning and training, the large language model understands user needs and is able to recommend fault diagnosis data enhancement strategies. This model evolves from a general large language model into a specialized thinking chain large language model for fault diagnosis data enhancement strategies, called the LLM-FDDA expert. This model can gradually acquire structural parameters and operating condition information of the equipment to be diagnosed, calculate fault characteristic frequencies, reason about data enhancement methods, and recommend parameters. The structural parameters and working condition information of the target rotating machinery parts are input into the thinking chain large language model. The thinking chain large language model will calculate the fault characteristic frequency according to the part object type and infer the appropriate strong and weak data enhancement strategies and their parameters.

3. The soft contrastive learning fault diagnosis method based on large language model question-answering dialogue according to claim 2 is characterized in that: The method further includes: Based on the Thinking Chain Large Language Model and the fault diagnosis model, a rotating machinery component fault diagnosis system combining LLM and soft comparative learning is built. New structural parameters and operating condition information of the rotating machinery component to be diagnosed are input, and the system automatically generates new strong and weak data enhancement strategies and parameters, as well as fault diagnosis results, for fault identification of various rotating machinery components. When the target rotating machinery component is a bearing, the structural parameters include the bearing's inner diameter, outer diameter, rolling element diameter, number of rolling elements, and contact angle; the operating condition information includes load, rotational frequency, temperature, and sampling frequency; When the target rotating mechanical component is a gear, the structural parameters include the number of teeth and pressure angle, and the operating condition information includes load, rotational frequency, temperature, and sampling frequency.

4. The soft contrastive learning fault diagnosis method based on large language model question-answering dialogue according to claim 1 is characterized in that: According to the convolutional network structure of the feature encoder, the time axis distance is calculated after each convolution layer, and the time-level weight matrix is constructed, including: For the feature map output by each convolution layer, a distance matrix of size T×T is constructed by calculating the absolute distance dist=|pq| between the time points, which is called the time-level weight matrix; where T represents the time dimension size of the feature map of this layer; The second temperature parameter ,in m is the kernel size of the pooling layer in different convolutional layers, k is the depth of the convolutional layer, is a hyperparameter.

5. The soft contrastive learning fault diagnosis method based on large language model question-answering dialogue according to claim 1 is characterized in that: The strong and weak enhancement samples and their soft labels are input into the soft contrastive learning module. The feature encoder is pre-trained through hierarchical soft label feature refinement, cross prediction and context consistency tasks, including: (1) Hierarchical soft label feature refinement task The strong enhancement samples and weak enhancement samples are input into the feature encoder composed of n layers of convolutional networks, based on the local features obtained by each layer of convolutional network in the feature encoder. Compute instance-level soft contrast loss for local features and temporal soft contrast loss , the instance-level soft contrast loss of the n-layer convolutional network and temporal soft contrast loss Accumulate layer by layer to get the final instance-level soft contrast loss and temporal soft contrast loss ; (2) Cross-prediction task Strongly enhanced samples output global features after passing through the feature encoder , from the global features Extract the feature representation of the first t time steps and input it into the autoregressive Transformer model to generate the context feature representation ; Using linear converter W k Mapping context feature representations to prediction representations at different time steps; where the linear transformer W k Corresponding to multiple future time steps k ; For each future time step k , the predicted representation is compared with the real future time step in the weakly enhanced sample k The inner product calculation is performed on the feature representation of to obtain the similarity between the two. Combined with the Log-Softmax operation, the similarity is normalized and the negative logarithm is taken to obtain the contrast loss matrix; By extracting the diagonal elements of the contrast loss matrix, the consistency between the predicted representation and the true representation in time steps, that is, the positive sample similarity, is calculated, and the loss is accumulated, and finally the Noise Contrastive Estimation loss is used. Calculate the cross-prediction loss from strong enhancement samples to weak enhancement samples ;in, B is the batch size; b The first b samples; i For the i time steps; b ’ Index of all samples in the same batch, including the current sample b and other samples, namely, negative sample candidate sets; is the representation of real future samples, is the representation of the predicted sample; Similarly, calculate the cross-prediction loss from weakly enhanced samples to strongly enhanced samples , the cross prediction loss from strong enhancement samples to weak enhancement samples Cross-prediction loss from weakly enhanced samples to strongly enhanced samples The sum is the final cross-prediction loss ; (3) Contextual consistency task Contextual feature representation of strongly enhanced samples Contextual feature representation of weakly enhanced samples Each of them is processed by a nonlinear projection head for dimensionality reduction and the Normalized Temperature-scaled CrossEntropy Loss is used. Calculate the context consistency loss between the two features of the output : Where, the numerator is the exponential form of the similarity of the positive sample pair, and the denominator is the sum of the exponential similarities between the current sample and all other samples. N is the sample size, Represents a positive sample pair, that is, a strong enhancement sample and a weak enhancement sample of a certain sample, sim represents the cosine similarity function, For an indicator function, make sure not to include itself; is the third temperature parameter, used to adjust the sensitivity of similarity calculation; (4) Total losses The layered soft label feature refinement loss , cross prediction loss and context consistency loss Perform weighted summation to obtain the total loss function , the formula is as follows: Where, 、 、 is the weight; The feature encoder is pre-trained using a total loss function.

6. The soft contrastive learning fault diagnosis method based on large language model question-answering dialogue according to claim 5 is characterized in that: Instance-level soft contrast loss The calculation method is: Where t is the time point index, a is the sample index, and L is the loss matrix after performing softmax normalization and taking the negative logarithm of the instance-level similarity matrix. S L is the lower left triangular part of the instance-level soft label matrix, S R is the upper right triangular part of the instance-level soft label matrix, B is the batch size, and T is the time step; Temporal soft contrast loss The calculation method is: Where b is the sample index in the batch, and They are the lower left and upper right triangular parts of the time-level soft label matrix, respectively.

7. The soft contrastive learning fault diagnosis method based on large language model question-answering dialogue according to claim 1 is characterized in that: The fault diagnosis classification head module is a linear layer that is responsible for mapping high-dimensional feature representations to the probability distribution of classification labels. The input dimension of the linear layer is equal to the product of the number of output channels of the last convolutional layer of the feature encoder and the time step length, and the output dimension is equal to the final number of classification labels.

8. A computer device, characterized in that: include: A processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the soft contrast learning fault diagnosis method based on large language model question-answering dialogue as described in any one of claims 1 to 7 are implemented.

9. A readable storage medium, characterized in that: Programs or instructions are stored thereon, and when the programs or instructions are executed by the processor, the steps of the soft contrast learning fault diagnosis method based on large language model question-answering dialogue as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Time sequence equipment fault diagnosis method based on comparison self-supervised learning

    CN116070128A

  • Steel production equipment fault diagnosis system based on large language model

    CN119917965A