Fault Detection Method for Rotating Machinery Based on Projection Distribution Learning of Monitoring Data

Through a method based on monitoring data projection distribution learning, combined with self-attention Transformer and differential attention deep neural network, the problem of large amount of information correlation and calculation in rotating mechanical equipment failure detection is solved, and efficient fault detection is achieved.

CN117992881BActive Publication Date: 2025-07-08CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410178647.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-09
Publication Date
2025-07-08
Estimated Expiration
2044-02-09

AI Technical Summary

Technical Problem

The existing methods are difficult to effectively capture the comprehensive monitoring data information of rotating machinery equipment, and the traditional Transformer model is computationally expensive and difficult to deploy, making it difficult to solve the problems of abnormal sparsity and proximity abnormality resolution.

Method used

Using a method based on monitoring data projection distribution learning, combining the self-attention Transformer branch and the differential attention deep neural network branch, a training sample set of monitoring data projection distribution data is formed through fast Fourier transform and data projection, and fault detection is performed using the VS-Transformer model.

Benefits of technology

It improves the characteristics of rotating mechanical equipment fault data, can better capture the differences and similarities between fault samples, has good robustness and detection performance, and avoids the calculation and storage problems of traditional models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117992881B_ABST
    Figure CN117992881B_ABST
Patent Text Reader

Abstract

The present invention describes a fault detection method for rotating mechanical equipment based on the learning of the projection distribution of monitoring data. Multivariate time series data during equipment operation is collected through vibration monitoring, and the data for each dimension is preprocessed to form training sample data. After dividing the training sample data into time windows, fast Fourier transform, data slice segmentation, and monitoring data projection are performed to form a training sample set of monitoring data projection distribution data. The training sample data and the training sample set of monitoring data projection distribution data are input into the VS-Transformer model. In the VS-Transformer model, a self-attention Transformer branch and a differential attention deep neural network branch are set up to calculate the training sample data and the training sample set of monitoring data projection distribution data respectively, and the difference value between the normal data and the detected data of the rotating mechanical equipment is obtained. A classifier is used to classify the detected data according to the difference value, thereby detecting the fault type of the rotating mechanical equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rotating machinery and equipment fault detection, and particularly relates to a rotating machinery and equipment fault detection method based on the learning of the projection distribution of monitoring data. Background Art

[0002] Rotating machinery and equipment is widely used in industrial production and life, such as aerospace, engines, and various mining machinery. As one of the cores of various industrial machinery, rotating machinery and equipment will have faults such as bearing fatigue spalling, wear, corrosion, and fracture during use, which will affect the normal operation of the equipment and may cause casualties in severe cases. For rolling bearings operating under high-speed conditions, the probability of failure is higher and the harm is greater. Therefore, researching the abnormal detection method of rotating machinery and equipment faults and timely discovering the faults that occur in rotating machinery and equipment has very high practical value.

[0003] With the development and application of deep learning, anomaly detection based on deep learning has become a hot issue in contemporary research fields. Deep learning algorithms have deep hierarchical networks that can automatically learn and extract multi-level features from raw data. In recent years, the attention mechanism has been proposed as a technology to enhance the performance of other models. Its unique information extraction model has greater research potential and has been widely studied and applied by scholars. The field of monitoring data also requires the internal association and global information extraction ability of the attention mechanism.

[0004] Many researchers use self-attention to improve the stability of model training. Self-attention can be appropriately designed into the model, and even self-attention can be used as the basis for constructing a feature extraction module, so that the dependence of the model on external information can be reduced, and the internal connection of features can be more fully extracted. However, it is difficult to capture the comprehensive information of mechanical monitoring data through the correlation calculation of a single sensor. Existing methods lack attention to the information association between samples and can only focus on the internal abnormal time series information of a single abnormal time series. This time series information will be confused by a large number of normal patterns and some similar abnormal information, making it difficult for the model to solve the problems of anomaly sparsity and adjacent anomaly discrimination. At the same time, the Transformer structure often faces the problems of large memory space occupation and huge computational complexity and is difficult to be deployed and applied. Summary of the Invention

[0005] To solve the problems existing in the above-mentioned prior art, the present invention provides a rotating machinery and equipment fault detection method based on the learning of the projection distribution of monitoring data.

[0006] The technical solution of the present invention is as follows:

[0007] A rotating machinery and equipment fault detection method based on the learning of the projection distribution of monitoring data includes the following steps:

[0008] S1, set the working state monitoring data of rotating mechanical equipment operating normally and rotating mechanical equipment operating with multiple different single fault types as training samples;

[0009] S2, through vibration monitoring, collect the multivariate time series data of the training samples during their operation, select the data of each variable dimension in the multivariate time series data for normalization preprocessing, and form training sample data;

[0010] The training sample data includes normal data collected during normal operation of rotating mechanical equipment and detected data collected during multiple different fault operations of rotating mechanical equipment classified by single fault;

[0011] Set labels for the training samples and label them as 0, 1, 2,.., N according to fault classification;

[0012] S3, divide the training sample data into time windows, and perform fast Fourier transform, data slice segmentation and monitoring data projection on the data within each divided time window to form a training sample set of monitoring data projection distribution data;

[0013] S4, take the training sample data and the training sample set of monitoring data projection distribution data as input data, and input them into the VS-Transformer model; set a self-attention Transformer branch and a differential attention deep neural network branch in the VS-Transformer model, and calculate the training sample data and the training sample set of monitoring data projection distribution data respectively to obtain the difference value between the normal data of rotating mechanical equipment operating normally and the detected data of fault operation;

[0014] S5, use a classifier to perform fault classification on the detected data according to the difference value, so as to detect the fault type of the rotating mechanical equipment.

[0015] Further, in step S2, the formula for normalizing the data is:

[0016]

[0017] In the formula,

[0018] is the mathematical expectation of the training sample;

[0019] σ x is the standard deviation;

[0020] x t-1 is the training sample data before normalization;

[0021] x t is the normalized training sample data.

[0022] Further, step S3 includes the following sub-steps:

[0023] S301, dividing the training sample data into time windows of the same length using the operating cycle of the rotating machinery;

[0024] S302, performing a fast Fourier transform on the data within each of the divided time windows to obtain the real part data and the imaginary part data of the transformed data sequence; directly connecting the real part data and the imaginary part data in sequence for re-splicing to form a data sequence training sample data [x ft .

[0025] S303, dividing the data within each of the divided time windows into K data slices; establishing a projection space and dividing the projection space into K corresponding to the data slices; projecting the data slices onto their corresponding projection spaces as monitoring data projections;

[0026] S304, splicing the data sequence training sample data with the monitoring data projection and performing embedding encoding to obtain a monitoring data projection distribution data training sample set.

[0027] Further, the calculation formula for the length of the time window is:

[0028]

[0029] In the formula,

[0030] L w is the length of the time window;

[0031] rpm is the operating cycle of the rotating machinery, unit: revolutions per minute;

[0032] f is the sampling frequency, unit: times per second.

[0033] Further, the relationship between the lengths of different time windows is determined by the following formula:

[0034] L w = concat(O * L w-1 , L - O * L w-1 )

[0035] In the formula,

[0036] O is the data overlap rate;

[0037] L w is the length of the new time window;

[0038] L w-1 is the length of the previous time window;

[0039] concat is an operation to concatenate two time series data.

[0040] Furthermore, the formula for the projection of the monitoring data is:

[0041]

[0042] In the formula,

[0043] i represents the i-th data slice and the corresponding projection space, with the value range [1..K];

[0044] exp is the exponential function;

[0045] represents the segmentation of the i-th data slice of the data after the fast Fourier transform operation;

[0046] represents the projection value of the monitoring data.

[0047] Furthermore, the training sample data of the data sequence and the projection of the monitoring data are concatenated using the concat function; the embedding encoding is a one-hot encoding of classification labels with the parameter set between 256 and 2048.

[0048] Furthermore, in step S4,

[0049] S401 sets the self-attention Transformer branch as an intermediate layer in the VS-Transformer model, and the steps are as follows:

[0050] S401-1. multi-head takes values from 1 to 6;

[0051] S401-2. The number of attention blocks is 1 to 5 layers, and each layer consists of a self-attention module, a regularization layer Normalize layer, a Feed Forward layer, and a softmax layer. Each layer outputs a result and a residual to the next layer;

[0052] S401-3. The input of the first layer is the training sample data of the data sequence or the training sample set of the projection distribution of the monitoring data. It is activated through the softmax function to output an intermediate layer with a length of 64 to 2048, and a cross-entropy optimization target Loss1 is obtained by comparing with the label.

[0053] S402 sets the differential attention depth neural network branch as an intermediate layer in the VS-Transformer model, and the steps are as follows:

[0054] S402-1. The value of multi-head ranges from 1 to 6;

[0055] S402-2. The number of attention blocks is 1 to 5 layers, and the input is the monitoring data projection [x i proj or the training sample set of the monitoring data projection distribution data [x d ;

[0056] S402-3. Set the calculation process as: enter the input data into two Linear branches, perform linear transformations 1 to 5 times respectively, and then one of them performs an exponential operation to obtain two sets of parameters; perform a connection on the two sets of parameters, and perform a differential calculation with the previous result to form the optimization objective Loss2, as shown in the formula:

[0057]

[0058] In the formula,

[0059] q s is the true distribution of the monitoring data;

[0060] p s is the prior distribution of the monitoring data;

[0061] is the expectation of the true distribution of the monitoring data;

[0062] x, y are the sample data and the corresponding labels;

[0063] x proj is the sample projection.

[0064] S403. Based on Loss1 and Loss2, form the final optimization objective, and the final optimization objective is composed of αLoss1+(1-α)Loss2, and 0<α<1.

[0065] Furthermore, through the following formula, initialize the dual-branch attention of the l-th layer of the attention block in step S401-2:

[0066]

[0067]

[0068] In the formula,

[0069] respectively represent the query, key, value and associated sequence of self-attention;

[0070] They respectively represent the parameter matrices of the l-th layer in Q, K, V, and σ;

[0071] It represents self-attention.

[0072] Furthermore, in step S401-3, the output representation is obtained through the following formula:

[0073] x l = softmax(Feed Forword(LNz l )) + z l )#

[0074] In the formula,

[0075] represents the input of the l-th layer encoder;

[0076] represents the hidden output of the l-th layer encoder;

[0077] d mode l is a hyperparameter, and its value ranges from 50 to 5000.

[0078] The difference between the projection distribution for output and the projection of normal data is described as a KL divergence parameter matrix through the following formula:

[0079]

[0080]

[0081] In the formula,

[0082] μ l and ∑ l are the variational parameter matrices of the logistic normal distribution;

[0083] d scale represents the number of distribution parameters for matching the classification dimension.

[0084] Furthermore, in the VS-Transformer model, the learning rate is set to 0.001 and decays exponentially to 0.00000001, and dropout is set to 0.1 - 0.5.

[0085] Furthermore, in step S5, the classifier includes two feedforward multi-layer perceptrons and two activation layers overlapping alternately, and dropout layers are added in front of the two sub-layers respectively to randomly discard some input features; the two activation layers are Relu and softmax activations respectively; the input of the classifier is spliced by the outputs of the self-attention Transformer branch and the differential attention deep neural network branch.

[0086] Furthermore, the fault classification process of the classifier is represented by the following formula:

[0087] CLA(x L ) = softmax(ReLU(x L W c1 + b c1 )W c 2 + b c2 )#

[0088] In the formula,

[0089] are the weights and biases of these two layers respectively;

[0090] N cla is the number of categories, and d ff is the dimension of the classifier input.

[0091] The beneficial effects of the present invention are as follows:

[0092] 1. The fault detection method recorded in the present invention uses a deep learning attention model, designs a dual-branch neural network structure, and adds a differential attention deep neural network branch for extracting the amplitude distribution statistical features in the distributed time-frequency domain. The differential attention deep neural network branch structure can, when extracting features, improve the attention of the Transformer model to the distribution characteristics of the fault data of rotating machinery equipment and magnify the attention to the abnormal data features.

[0093] 2. The fault detection method recorded in the present invention can pay more attention to the time and frequency domain information and the correlation of information distribution between samples, can better capture the differences and similarities between fault samples, has good fault diagnosis performance, and avoids the disadvantages such as limited long-range association ability and parallel computing ability of traditional models.

[0094] 3. The method for abnormal detection of rotating machinery equipment faults based on the monitoring data projection distribution mechanism recorded in the present invention has good robustness and can still maintain good detection performance under the influence of various complex environments and noises. Description of the Drawings

[0095] Figure 1: Types of faults and label schematic diagrams of bearings and gears during the operation process in the exemplary embodiments of the present invention;

[0096] Figure 2 : Schematic diagram of the process flow of the fault detection method for rotating mechanical equipment in the exemplary embodiments of the present invention;

[0097] Figure 3 : Schematic diagrams of the overall structure of the VS-Transformer model and the encoder structure in the exemplary embodiments of the present invention;

[0098] Figure 4 : Schematic diagram of the dataset used to verify the method of the present invention in the exemplary embodiments of the present invention;

[0099] Figure 5 : Schematic diagrams of the comparison results between the VS-Transformer model and other models under each dataset in the exemplary embodiments of the present invention;

[0100] Figure 6 : Schematic diagram of the prediction accuracy results of the VS-Transformer model compared with other models on the dataset under the noisy environment added in the exemplary embodiments of the present invention. Detailed implementation manners

[0101] The present invention will be described in detail below in the form of embodiments in conjunction with the accompanying drawings.

[0102] The present invention describes a fault detection method for rotating mechanical equipment based on the learning of the projection distribution of monitoring data, which specifically includes the following steps.

[0103] S1. According to the types of faults generated by the rotating mechanical equipment during the operation process, through label making, the monitoring data of the working state of the rotating mechanical equipment is set as a plurality of corresponding training samples.

[0104] The training samples include the normal data collected during the normal operation of the rotating mechanical equipment and the detected data collected during the operation of the rotating mechanical equipment with multiple different faults classified by a single fault;

[0105] As a preference:

[0106] a. The rotating mechanical equipment in this embodiment takes bearings and gears as examples, and Figure 1 shows the types of faults and label schematics of bearings and gears during the operation process.

[0107] b. Set labels for the training samples and mark the labels as 0, 1, 2,.., N according to the fault classification.

[0108] S2. Through vibration monitoring, collect the multivariate time series data of the training sample during its operation. Select the data of each variable dimension in the multivariate time series data for normalization preprocessing to obtain the preprocessed data set of the time series, which is defined as the training sample data [x t .

[0109] Through normalization preprocessing, the multivariate time series data can be standardized, and its formula is:

[0110]

[0111] In the formula,

[0112] is the mathematical expectation of the training sample;

[0113] σ x is the standard deviation;

[0114] x t-1 is the training sample data before normalization;

[0115] x t is the training sample data after normalization.

[0116] Preferably:

[0117] a. The training sample data [x t includes: normal data collected during the normal operation of rotating machinery and equipment; detected data collected during the operation of rotating machinery and equipment with a single fault type.

[0118] b. The collection of multivariate time series data is completed by monitoring with triaxial vibration sensors installed on rotating machinery and equipment; the number of triaxial vibration sensors is preferably set to 2 - 5.

[0119] S3. Divide the training sample data [x t into time windows, and perform fast Fourier transform, data slice segmentation, and monitoring data projection on the data in each divided time window to form a monitoring data projection distribution data set, which is defined as the monitoring data projection distribution data training sample set [x d .

[0120] Specifically, it includes the following steps:

[0121] S301. Divide the training sample data [x t into time windows:

[0122] Use the operating cycle of the rotating machinery and equipment to divide time windows of the same length. The calculation formula for the length of the time window is:

[0123]

[0124] In the formula,

[0125] L w is the length of the time window;

[0126] rpm is the operating cycle of the rotating machinery and equipment, unit: revolutions per minute;

[0127] f is the sampling frequency, unit: times per second.

[0128] Preferably:

[0129] a. Since the relationship between a new time window and the previous time window in the time series data is that some data in the new time window may overlap with the previous one, therefore, the relationship between the lengths of different time windows is determined by the following formula:

[0130] L w = concat(O * L w-1 , L - O * L w-1 )

[0131] In the formula,

[0132] O is the data overlap rate;

[0133] L w is the length of the new time window;

[0134] L w-1 is the length of the previous time window;

[0135] concat is an operation to concatenate two time series data.

[0136] b. When the data overlap rate is greater than 0, when inputting to the VS-Transformer model described below, a buffer is needed to store the overlapping part of the data.

[0137] S302. Perform a fast Fourier transform on the data within each divided time window to obtain the real part and the imaginary part of the transformed data sequence;

[0138] Connect the real part data and the imaginary part data directly in sequence and perform re - splicing to form a data sequence training sample data [x w with the same length L ft .

[0139] S303. Divide the data within each divided time window into K data slices, where K is selected as an integer multiplication factor of the length L w value.

[0140] S304. Establish a monitoring data projection distribution data training sample set [xd :

[0141] a. Establish a projection space of length E, divide the projection space into K, and number them sequentially according to the position as [p1, p2, …, p i , …, p k ; The method for establishing the projection space is to allocate a memory space with the same dimension as the sample set and a length of E;

[0142] b. Number the data slices in chronological order Project the i-th data slice onto the projection space p i to obtain the monitoring data projection [x i pro j].

[0143] The monitoring data projection [x i proj formula is:

[0144]

[0145] In the formula,

[0146] i represents the i-th data slice and the corresponding projection space, with a value range of [1..K];

[0147] exp is the exponential function;

[0148] represents the segmentation of the i-th data slice after the fast Fourier transform operation;

[0149] represents the detection data projection value.

[0150] S305, Concatenate the data sequence training sample data [x ft and the monitoring data projection [x i proj using the concat function, and then perform embedding encoding to obtain the concatenated data after information enhancement, that is, the monitoring data projection distribution data training sample set [x d .

[0151] Preferably:

[0152] The embedding encoding is the one-hot encoding of the classification label, and the parameter is set between 256 and 2048.

[0153] S4, Use the training sample data [x t and the monitoring data projection distribution data training sample set [x dAs input data, it is input into the VS-Transformer model for fault detection of rotating machinery for learning the projection distribution;

[0154] In the VS-Transformer model, a self-attention Transformer branch and a differential attention deep neural network branch are set up to calculate the training sample data [x t and the projection distribution data training sample set [x d of the monitoring data respectively, so as to obtain the difference values between the normal data of the rotating machinery and the detected data of various faults.

[0155] Preferably:

[0156] In the VS-Transformer model, the learning rate is set to 0.001 and decays exponentially to 0.00000001, and dropout is set to 0.1 - 0.5.

[0157] Specifically, it includes the following steps:

[0158] S401. The self-attention Transformer branch is an intermediate layer in the VS-Transformer model, and the setting steps are as follows:

[0159] S401-1. The value of multi-head is 1 - 6;

[0160] S401-2. The number of attention blocks is 1 - 5 layers. Each layer consists of a self-attention module, a regularization layer Normalize layer, a Feed Forward layer and a softmax layer. Each layer will output a result and a residual to the next layer;

[0161] The dual-branch attention of the l-th layer of the attention block is initialized through the following formula:

[0162]

[0163]

[0164] In the formula,

[0165] respectively represent the query, key, value and associated sequence of the self-attention;

[0166] respectively represent the parameter matrices of the l-th layer in Q, K, V, σ;

[0167] represents the self-attention.

[0168] S401 - 3. The input of the first layer is the training sample data [x ft of the data sequence after the fast Fourier transform operation or the training sample set of the projection distribution data of the monitoring data [x d . The intermediate layer with a length of 64 - 2048 is activated through the softmax function and compared with the label to obtain a cross - entropy optimization objective Loss1;

[0169] The output representation is carried out through the following formula:

[0170] x l = softmax(Feed Forword(LN(z l )) + z l )#

[0171] In the formula,

[0172] represents the input of the l - th layer encoder;

[0173] represents the hidden output of the l - th layer encoder;

[0174] d model is a hyperparameter, and its value is between 50 - 5000.

[0175] S402. The differential attention deep neural network branch is an intermediate layer in the VS - Transformer model, and the setting steps are as follows:

[0176] S402 - 1. multi - head takes values from 1 to 6;

[0177] S402 - 2. The number of attention blocks is 1 - 5 layers, and the input is the monitoring data projection [x i proj or the training sample set of the projection distribution data of the monitoring data [x d ;

[0178] S402 - 3. Set the calculation process as: the input data enters two Linear branches, and each does 1 - 5 linear transformations. One of them then does an exponential operation to obtain two sets of parameters; connect the two sets of parameters and do a differential calculation with the previous result to form an optimization objective Loss2, as shown in the formula:

[0179]

[0180] In the formula,

[0181] q s is the true distribution of the monitoring data;

[0182] ps is the prior distribution of the monitoring data;

[0183] is the expectation of the true distribution of the monitoring data;

[0184] x and y are the sample data and the corresponding labels;

[0185] x proj is the sample projection.

[0186] S403. Form a final optimization objective based on Loss1 and Loss2. The final optimization objective is composed of αLoss1+(1 - α)Loss2, where 0 < α < 1.

[0187] Appendix Figure 3 shows a schematic diagram of the differential attention depth neural network branch structure, which is used to output the difference between the projection distribution and the projection of normal data, and is described by the following formula as the KL divergence parameter matrix:

[0188]

[0189]

[0190] In the formula,

[0191] μ l and ∑ l are the variational parameter matrices of the logistic normal distribution;

[0192] d scale represents the number of distribution parameters for matching the classification dimension.

[0193] S5. Use a classifier to determine the fault classification in the detected data according to the difference value, so as to detect the fault type of the rotating mechanical equipment.

[0194] Specifically, it includes the following steps:

[0195] S501. The classifier includes two feedforward multi-layer perceptrons and two activation layers overlapping alternately, and a dropout layer is added in front of each of the two sub-layers to randomly discard some input features; the two activation layers are Relu and softmax activations respectively.

[0196] S502. The input of the classifier is spliced by the outputs of the self-attention Transformer branch and the differential attention depth neural network branch.

[0197] S503. The fault classification process is represented by the following formula:

[0198] CLA(x L ) = softmax(ReLUx LW c1 + b c1 )W c2 + b c2 )#

[0199] wherein

[0200] are the weights and biases of these two layers respectively;

[0201] N cla is the number of categories, and d ff is the dimension of the classifier input.

[0202] As shown in Appendix Figure 4 to Appendix Figure 6 the method described in this embodiment is described below through experimental verification.

[0203] The method for fault detection of rotating machinery based on the projection distribution learning of monitoring data described in this exemplary embodiment is verified through existing and commonly used anomaly detection data sets, including but not limited to: CWRU data set, UPB data set, MFPT data set, JUN data set, and SEU data set.

[0204] Among them, the first four data sets are all bearing data sets, and the faulty bearings are generally distributed on the inner ring, rolling elements, and outer ring of the bearings; the SEU data set is a gearbox data set. By changing the load torque and speed of the motor, experimental data under different working conditions are obtained. The specific description is as follows:

[0205] 1. The CWRU data set is provided by the experimental bearing data of Case Western Reserve University based on vibration status. This data set contains data of normal bearings and faulty bearings, where the faulty bearings are implanted with faults through electrical discharge machining, and data are collected at 3 separate locations near the motor using accelerometers;

[0206] 2. The UPB data set consists of an experimental data set based on vibration and motor current data collected by the University of Paderborn in Germany. A total of 26 faulty bearings and 6 healthy bearings were tested, and the faulty bearings were subdivided into 12 artificial damages and 14 natural damages;

[0207] 3. The MFPT data set is prepared from the MFPT set of fault repairs. This data set includes experimental data of 3 normal bearings, experimental data of 17 faulty bearings, and 3 real-world sample data;

[0208] 4. The JNU data set is an experimental bearing data set collected and shared by Jiangnan University. This data set includes a total of 4 normal bearings and 9 faulty bearing data;

[0209] 5. The SEU data set is a gearbox data set collected and shared by Southeast University.

[0210] The core attributes of the above abnormal detection dataset are all vibration data.

[0211] To design experiments for different factors, the dataset was re-partitioned, and the details are as attached Figure 4 shown; the comparison results of the VS-Transformer model with other models under each dataset are as attached Figure 5 shown; attached Figure 6 shows the prediction accuracy results of the VS-Transformer model compared with other models on the dataset under the added noise environment. Through the above comparison, it can be clearly concluded that: the method described in this exemplary embodiment can obtain better detection effects and shows significant beneficial effects in the abnormal detection of rotating machinery equipment.

[0212] A brief description of the VS-Transformer model for rotating machinery equipment fault detection based on projection distribution learning constructed by the deep learning network is as follows:

[0213] This model includes two main parts: an encoder feature extraction network and a classifier feature mapping network. In this embodiment, the encoding in the VS-Transformer model is used to increase the differential learning encoding of the distribution projection. The VS-Transformer model requires two types of data to complete the training work: the normal data of the rotating machinery equipment and the data to be detected. The normal data uses the normal data in the training sample data [x t , and the data to be detected is the target of the fault category to be identified, that is, the data to be detected in the training sample data [x t . The data to be detected is labeled according to the fault type, as shown in the attachment Figure 1 shown.

[0214] A brief description of the abnormal classification judgment criterion of the VS-Transformer model for equipment fault detection is as follows:

[0215] For example, if there is a fault in the rotating machinery equipment, for the vibration caused by the fault, the monitoring data sequence collected by the three-axis vibration sensor will be different from the normal operation of the equipment, and there will be differences in each cycle. After the VS-Transformer model calculates the output value for this data, the result will be different from the calculation result of the normal data. Assuming that the difference value formed by the fault is an abnormal value, the difference values calculated by different faults through the model will be different, thus forming different value ranges outside the normal data. When the VS-Transformer model calculates the output value, the method described above will calculate which fault difference value range the model calculation output is in, so as to judge the fault type.

[0216] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed.

Claims

1. A fault detection method for rotating machinery based on learning the projection distribution of monitoring data, characterized in that, Including the following steps: S1. Set the working condition monitoring data of the rotating mechanical equipment under normal operation and the rotating mechanical equipment operating with multiple different single fault types as training samples; S2. Through vibration monitoring, collect the multivariate time series data of the training samples during their operation, select the data of each variable dimension in the multivariate time series data for normalization preprocessing, and form training sample data; The training sample data includes normal data collected during the normal operation of the rotating mechanical equipment and detected data collected during the operation of the rotating mechanical equipment with multiple different faults classified by single faults; Set labels for the training samples and mark the labels as 0, 1, 2,.., N according to the fault classification; S3. Divide the training sample data into time windows, and perform fast Fourier transform, data slice segmentation, and monitoring data projection on the data in each divided time window to form a training sample set of monitoring data projection distribution data; S4. Use the training sample data and the training sample set of monitoring data projection distribution data as input data and input them into the VS-Transformer model; set a self-attention Transformer branch and a differential attention deep neural network branch in the VS-Transformer model, and calculate the training sample data and the training sample set of monitoring data projection distribution data respectively to obtain the difference value between the normal data of the rotating mechanical equipment under normal operation and the detected data of the fault operation; S5. Use a classifier to classify the detected data according to the difference value, so as to detect the fault type of the rotating mechanical equipment; the input of the classifier is spliced by the outputs of the self-attention Transformer branch and the differential attention deep neural network branch; Set the differential attention deep neural network branch as an intermediate layer in the VS-Transformer model, and the steps are as follows: S4-1. The value range of multi-head is from 1 to 6; S4-2. The number of attention blocks ranges from 1 to 5 layers, and the input is the projection of the monitoring data [x i proj or the training sample set of the projection distribution data of the monitoring data [x d ; The first branch of the differential attention module is sequentially connected to the Norm layer and the Exp layer, and the output parameters μ0 and σ0 are output; The second branch of the differential attention module inputs the calculation results of the self-attention module Q and K, and after connecting the Linear layer and the Softmax layer, it is sequentially connected to the two branches. One of the branches is connected to the Linear layer to output the parameter μ l , and the other branch is sequentially connected to the Linear layer and the Exp layer to output the parameter σ l ; S4-3. Set the calculation process as follows: Input the input data into two branches to obtain two sets of parameters μ0, σ0 and μ l , σ l , concatenate the two sets of parameters once, perform a difference calculation with the previous result to form the optimization objective Loss2, as shown in the following formula: In the formula, q s For monitoring the true distribution of data; p s For the prior distribution of monitoring data; For the expectation of the true distribution of monitoring data; x, y are sample data and corresponding labels; x proj Is the sample projection.

2. The fault detection method for rotating machinery based on the learning of the projection distribution of monitoring data according to claim 1, wherein In step S2, the formula for the normalization preprocessing is: Wherein, is the mathematical expectation of the training samples; σ x is the standard deviation; x t-1 is the training sample data before normalization; x t is the training sample data after normalization.

3. A method for fault detection of rotating mechanical equipment based on learning the projection distribution of monitoring data according to claim 1, characterized in that, Step S3 specifically includes the following sub-steps: S301. Use the operating cycle of the rotating mechanical equipment to divide the training sample data into time windows of the same length; In S302, perform a fast Fourier transform on the data within each of the divided time windows to obtain the real part data and the imaginary part data of the transformed data sequence; directly splice the real part data and the imaginary part data in sequence to form a data sequence training sample data [x ft ; S303. Divide the data in each divided time window into K data slices; establish a projection space and divide the projection space into K corresponding to the data slices; project the data slices to their corresponding projection spaces; the formula for the monitoring data projection is: where \(i\) represents the \(i\)-th data slice and the corresponding projection space, and the value range is \([1, K]\); exp is the exponential function; represents the segmentation of the \(i\)-th data slice of the data after the fast Fourier transform operation; represents the projection value of the monitoring data of S304. Splice the data sequence training sample data and the monitoring data projection, and perform embedding encoding to obtain a training sample set of monitoring data projection distribution data.

4. A method for fault detection of rotating machinery based on learning of projection distribution of monitoring data according to claim 3, characterized in that, The formula for the length of the time window is: In the formula, L w is the length of the time window; rpm is the operating cycle of the rotating mechanical equipment, unit: revolutions per minute; f is the sampling frequency, unit: times per second.

5. A method for fault detection of rotating mechanical equipment based on learning of projection distribution of monitoring data according to claim 3 or 4, characterized in that The training sample data of the data sequence and the projection of the monitoring data are concatenated using the concat function; the embedding encoding is set as the one-hot encoding of the classification labels with the parameter between 256 and 2048.

6. A method for rotating machinery fault detection based on learning of projection distribution of monitoring data according to claim 3, characterized in that In step S4: S401 sets the self-attention Transformer branch as an intermediate layer in the VS-Transformer model, and the steps are as follows: S401-1. The value range of multi-head is from 1 to 6; S401-2. The number of attention blocks ranges from 1 to 5 layers. Each layer consists of a self-attention module, a regularization layer Normalize layer, a Feed Forward layer and a softmax layer. Each layer outputs a result and a residual to the next layer; S401-3. The input of the first layer is the training sample data of the data sequence or the training sample set of the projection distribution data of the monitoring data. It is activated by the softmax function to output an intermediate layer with a length of 64 - 2048, and a cross-entropy optimization target Loss1 is obtained by comparing with the label; S402. Based on Loss1 and Loss2, a final optimization target is formed. The final optimization target is composed of αLoss1+(1-α)Loss2, and 0<α<1.

7. A method for rotating machinery fault detection based on the projection distribution learning of monitoring data according to claim 6, characterized in that The output representation is carried out through the following formula: x l = softmax(Feed Forward(LN(z l )) + z l ) In the formula, Denote the output of the l-th layer encoder; Represents the output of the hidden layer of the l-th layer encoder; d model is a hyperparameter, and its value ranges from 50 to 5000; And the difference between the output projection distribution and the normal data projection is described as the KL divergence parameter matrix through the following formula: In the formula, and is the variational parameter matrix of the logistic normal distribution; d scale Indicates the number of distribution parameters matching the classification dimension.

8. A method for fault detection of rotating mechanical equipment based on learning the projection distribution of monitoring data according to claim 6, characterized in that, In the VS-Transformer model, the learning rate is set to 0.001 and decays exponentially to 0.00000001, and dropout is set to 0.1 - 0.

5.

9. A method for fault detection of rotating mechanical equipment based on learning of projection distribution of monitoring data according to claim 1, 2, 3 or 6, characterized in that In step S5, the fault classification process is represented by the following formula: CLA(x L ) = softmax(ReLU(x L W c1 + b c1 )W c2 + b c2 ) In the formula, They are the weights and biases of these two layers respectively; N cla is the number of categories, and d ff is the dimension of the classifier input.

10. A method for fault detection of rotating machinery based on learning of projection distribution of monitoring data according to claim 9, characterized in that In step S5, the classifier includes two feed-forward multi-layer perceptrons and two activation layers overlapping alternately, and dropout layers are added before the two sub-layers respectively to randomly discard some input features; The two activation layers are Relu and softmax activations respectively.

Citation Information

Patent Citations

  • Human body action recognition method of graph convolutional neural network capable of automatically distinguishing and enhancing spatial and temporal features

    CN111339845A

  • Atrial fibrillation identification method and device based on Transformer

    CN113855037A