Fault Diagnosis Method and Device for Rotating Machinery Based on Sample Amplification

By reconstructing the phase space and amplifying the samples of vibration signals from rotating machinery, fault samples that conform to the true distribution are generated, solving the problem of sample imbalance in fault diagnosis of rotating machinery and improving the accuracy of fault diagnosis.

CN120653926BActive Publication Date: 2026-04-03GUANGDONG OCEAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

The sample imbalance problem exists in the fault diagnosis of rotating machinery equipment, resulting in low accuracy of fault diagnosis, especially insufficient recognition rate for a few types of faults.

Method used

Vibration signals of rotating machinery under different fault states are collected, standardized, and then reconstructed in phase space. The proportion and weight of false nearest neighbors are determined, a new fault sample phase space trajectory matrix is ​​generated, and the samples are merged to form a fault dataset. A trained fault diagnosis model is then used for diagnosis.

Benefits of technology

It improves the accuracy of fault diagnosis for rotating machinery, avoids imbalance between fault samples and normal samples, enhances the ability to express fault features, ensures that newly generated fault samples conform to the true distribution, and improves diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653926B_ABST
    Figure CN120653926B_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for fault diagnosis of rotating machinery based on sample amplification. It reconstructs the phase space of one-dimensional vibration signals, converting them into high-dimensional signals to enhance the expressive power of fault features. For each spatial sample, a false nearest neighbor ratio is determined, and a corresponding weight is assigned based on this ratio. This allows low-stability samples located in edge noise regions to be assigned lower weights, while high-stability samples are assigned higher weights. The higher-weighted spatial samples (reference samples) serve as the benchmark for data amplification. When generating new fault samples based on the benchmark and target neighbor samples, the newly generated fault samples are ensured to be close to the benchmark, preventing them from deviating from the true distribution and thus improving the reliability of the newly generated fault samples. Furthermore, the increased number of fault samples avoids imbalance between fault diagnosis samples and normal samples, improving the accuracy of fault diagnosis for rotating machinery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fault diagnosis technology for rotating machinery, and in particular to a method and apparatus for fault diagnosis of rotating machinery based on sample amplification. Background Technology

[0002] In rotating machinery power systems, the rotating machinery itself serves as the core power output equipment, and its operational status directly affects the safety, stability, and production efficiency of the entire system. This type of equipment typically includes steam turbines, gas turbines, compressors, and generators. These devices often operate under extreme conditions such as high temperature, high pressure, and high speed, while simultaneously enduring the coupled effects of mechanical stress, thermal stress, and media corrosion. Key components of rotating machinery (such as bearings, gears, and rotors) are highly susceptible to typical failures such as wear, cracks, imbalance, and misalignment. Even more critically, in modern industrial production, rotating machinery often requires continuous 24-hour operation, placing it in a state of fatigue accumulation over extended periods, further exacerbating the risk of malfunctions.

[0003] In related technologies, the core challenge in the field of fault diagnosis is the imbalanced sample problem. During fault diagnosis, the sample size for equipment operating normally typically exceeds 99%. This extreme imbalance in data distribution significantly impacts the performance of fault diagnosis models, primarily in the following ways: First, classification algorithms tend to classify most samples as normal, leading to a significant decrease in the recognition rate of minority faults such as bearing spalling and gear tooth breakage. Second, rare fault patterns often lack sufficient training samples, making it difficult for the model to learn effective fault feature representations.

[0004] Therefore, a new method for diagnosing faults in rotating machinery is needed to solve the aforementioned problems in the existing technology. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention provides a method and apparatus for fault diagnosis of rotating machinery based on sample amplification, in order to solve or partially solve the technical problem in the prior art where the imbalance of fault diagnosis samples affects the accuracy of fault diagnosis of rotating machinery.

[0006] A first aspect of the present invention provides a method for fault diagnosis of rotating machinery based on sample amplification, the method comprising:

[0007] Multiple sets of first vibration signals of the rotating machinery under different fault states are collected, and the multiple sets of first vibration signals are standardized to obtain a first standard signal; multiple sets of second vibration signals of the rotating machinery under normal state are collected, and the multiple sets of second vibration signals are standardized to obtain a second standard signal.

[0008] The first standard signal is reconstructed in phase space to obtain a first phase space trajectory matrix; the second standard signal is reconstructed in phase space to obtain a second phase space trajectory matrix.

[0009] Determine the false nearest neighbor ratio of each spatial sample in the first phase space trajectory matrix; determine the weight corresponding to each spatial sample based on the false nearest neighbor ratio of each spatial sample; and determine the benchmark sample and the target nearest neighbor sample based on the weight of each spatial sample.

[0010] A new fault sample phase space trajectory matrix is ​​generated based on the benchmark sample and the target nearest neighbor sample. The new fault sample phase space trajectory matrix and the second phase space trajectory are merged to obtain the fault dataset.

[0011] The fault dataset is used as a pre-built fault diagnosis model for training, and the trained fault diagnosis model is used to diagnose faults in rotating machinery.

[0012] In the above scheme, the first standard signal includes multiple components, and the step of reconstructing the first standard signal into a first phase space trajectory matrix includes:

[0013] Under different delay times, the autocorrelation function value of the first standard signal is determined using the autocorrelation function, resulting in multiple autocorrelation function values;

[0014] Each autocorrelation function value is compared with a preset first threshold. The autocorrelation function value that matches the first preset threshold is determined as the target autocorrelation function value. The delay time corresponding to the target autocorrelation function value is taken as the optimal delay time.

[0015] Determine the optimal embedding dimension of the first phase space trajectory matrix;

[0016] The first standard signal is reconstructed in a high-dimensional phase space based on the optimal embedding dimension and the optimal delay time to obtain the first phase space trajectory matrix.

[0017] In the above scheme, determining the autocorrelation function value of the first standard signal using the autocorrelation function includes:

[0018] Using formula Determine the autocorrelation function value R(τ) of the first standard signal; where,

[0019] N is the total length of the first standard signal, τ is the delay time, t is the sampling point of the first standard signal, and x(t) is the t-th first standard signal. Let σ be the mean of all first standard signals, and let σ be the variance of all first standard signals.

[0020] In the above scheme, the step of reconstructing the first standard signal in a high-dimensional phase space based on the optimal embedding dimension and the optimal delay time to obtain the first phase space trajectory matrix includes:

[0021] According to the formula The first standard signal is reconstructed in a high-dimensional phase space to obtain the first phase space trajectory matrix T;

[0022] M is the total number of spatial samples in the first phase space trajectory matrix, and X is... M Let x(M) be the Mth spatial sample in the first phase space trajectory matrix, x(M) be the value of the Mth sampling point of the first standard signal in the time series, τ′ be the optimal delay time, and d be the value of the first standard signal in the first phase space trajectory matrix. * Let be the optimal embedding dimension.

[0023] In the above scheme, determining the proportion of false nearest neighbors for each spatial sample in the first phase space trajectory matrix includes:

[0024] According to the formula Determine the proportion of false nearest neighbors R(X) for each spatial sample. i );

[0025] The X i The d represents the i-th spatial sample in the first phase space trajectory matrix, where i = 1, 2, ..., M; * For the optimal embedding dimension, the For the i-th spatial sample in the d-dimensional first phase space trajectory matrix, the X is the trajectory matrix in the first phase space of d-dimensional space. i The j-th nearest neighbor sample, where k is the number of nearest neighbors of the i-th spatial sample, and the For the i-th spatial sample in the d+1 dimensional first phase space trajectory matrix, X is the first phase space trajectory matrix in d+1 dimensions. i+1 The j-th nearest neighbor sample, the R th For a preset distance threshold, II() is an indicator function used to indicate when... When the time is right, the output is 1.

[0026] In the above scheme, determining the weight corresponding to each spatial sample based on the proportion of false nearest neighbors of each spatial sample includes:

[0027] According to the formula Determine the i-th spatial sample X i weight w(X) i );in,

[0028] The R(X) i ) represents the i-th spatial sample X i The false nearest neighbor ratio, where θ is a preset second threshold, and the i-th spatial sample is any spatial sample in the first phase space trajectory matrix.

[0029] In the above scheme, determining the benchmark sample based on the weight of each spatial sample includes:

[0030] According to the formula Determine the i-th spatial sample X i The first probability P(X) of the baseline sample i );

[0031] Spatial samples with a first probability greater than a third threshold are defined as the baseline samples; the third threshold is the average of the first probabilities of all spatial samples; wherein...

[0032] The w(X) i ) is the i-th spatial sample X i The weights are M, where M is the total number of spatial samples in the first phase space trajectory matrix.

[0033] In the above scheme, determining the target nearest neighbor sample based on the weight of each spatial sample includes:

[0034] Determine the nearest neighbor set for each spatial sample;

[0035] For each nearest neighbor sample in the nearest neighbor set, according to the formula The second probability P(X) is used to determine each of the nearest neighbor samples as the target nearest neighbor sample. j |X i );

[0036] The nearest neighbor samples whose second probability is greater than the fourth threshold are determined as the target nearest neighbor samples; the fourth threshold is the average of the second probabilities of all nearest neighbor samples; wherein,

[0037] The w(X) j ) is the j-th nearest neighbor sample X j The weights, M being the total number of spatial samples in the first phase space trajectory matrix, and N... k (Xi ) is the set of nearest neighbor samples of the i-th spatial sample, where k is the total number of the nearest neighbor samples.

[0038] In the above scheme, generating a new fault sample phase space trajectory matrix based on the reference sample and the target nearest neighbor sample includes:

[0039] According to formula X new =X i′ +λ×w(X i′ )×(X j′ -X i′ Generate a new fault space sample X at the corresponding position in the first phase space trajectory matrix. new The phase space trajectory matrix of the new fault sample is obtained;

[0040] The X i′ For the i′-th reference sample, the w(X) i′ ) represents the weight of the i′-th reference sample, λ is the interpolation position correction coefficient, and the value of λ ranges from [0,1]. X j′ Let i be the i′-th target nearest neighbor sample.

[0041] A second aspect of the present invention provides a fault diagnosis device for rotating machinery based on sample amplification, characterized in that the device comprises:

[0042] The processing unit is used to collect multiple sets of first vibration signals of the rotating machinery under different fault states, perform standardization processing on the multiple sets of first vibration signals to obtain a first standard signal; and collect multiple sets of second vibration signals of the rotating machinery under normal state, perform standardization processing on the multiple sets of second vibration signals to obtain a second standard signal.

[0043] The reconstruction unit is used to perform phase space reconstruction on the first standard signal to obtain a first phase space trajectory matrix; and to perform phase space reconstruction on the second standard signal to obtain a second phase space trajectory matrix.

[0044] A determining unit is used to determine the false nearest neighbor ratio of each spatial sample in the first phase space trajectory matrix; determine the weight corresponding to each spatial sample based on the false nearest neighbor ratio of each spatial sample; and determine the benchmark sample and the target nearest neighbor sample based on the weight of each spatial sample.

[0045] The generation unit is used to generate a new fault sample phase space trajectory matrix based on the reference sample and the target nearest neighbor sample, and merge the new fault sample phase space trajectory matrix with the second phase space trajectory to obtain a fault dataset;

[0046] The training unit is used to train the fault dataset as a pre-built fault diagnosis model, and to use the trained fault diagnosis model to diagnose faults in rotating machinery.

[0047] This invention provides a method and apparatus for fault diagnosis of rotating machinery based on sample amplification. The method includes: acquiring multiple sets of first vibration signals of the rotating machinery under different fault states; standardizing the multiple sets of first vibration signals to obtain a first standard signal; acquiring multiple sets of second vibration signals of the rotating machinery under normal state; standardizing the multiple sets of second vibration signals to obtain a second standard signal; reconstructing the phase space of the first standard signal to obtain a first phase space trajectory matrix; reconstructing the phase space of the second standard signal to obtain a second phase space trajectory matrix; determining the false nearest neighbor ratio of each spatial sample in the first phase space trajectory matrix; determining the weight corresponding to each spatial sample according to the false nearest neighbor ratio of each spatial sample; determining a reference sample and a target nearest neighbor sample according to the weight of each spatial sample; generating a new fault sample phase space trajectory matrix according to the reference sample and the target nearest neighbor sample; and connecting the new fault sample phase space trajectory matrix with the second phase space trajectory matrix. The inter-trajectories are merged to obtain a fault dataset. This fault dataset is then used to train a pre-built fault diagnosis model, which is then used to diagnose faults in rotating machinery. In this way, by reconstructing the phase space of the first vibration signal, the one-dimensional vibration signal is converted into a high-dimensional vibration signal, enhancing the expressive power of fault features. Furthermore, a false nearest neighbor ratio is determined for each spatial sample, and the corresponding weight is determined based on this ratio. This allows low-stability samples located in edge noise regions to be assigned lower weights, while high-stability samples are assigned higher weights. The higher-weighted spatial samples (reference samples) serve as the benchmark for data augmentation. When generating new fault samples based on the benchmark samples and target neighbor samples, it ensures that the newly generated fault samples are close to the benchmark samples, preventing them from deviating from the true distribution and thus improving the reliability of the newly generated fault samples. As the number of fault samples increases, the imbalance between fault diagnosis samples and normal samples can be avoided, improving the accuracy of fault diagnosis for rotating machinery. Attached Figure Description

[0048] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0049] Figure 1A schematic flowchart of a fault diagnosis method for rotating machinery based on sample amplification according to an embodiment of the present invention is shown.

[0050] Figure 2 A schematic diagram showing the result of predicting a test set using conventional methods according to an embodiment of the present invention is illustrated.

[0051] Figure 3 A schematic diagram of the results of predicting a test set using a sample amplification-based fault diagnosis method for rotating machinery according to an embodiment of the present invention is shown.

[0052] Figure 4 A schematic diagram of a fault diagnosis device for rotating machinery based on sample amplification according to an embodiment of the present invention is shown. Detailed Implementation

[0053] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0054] This invention provides a method for fault diagnosis of rotating machinery based on sample amplification, such as... Figure 1 As shown, the method mainly includes the following steps:

[0055] S110: Collect multiple sets of first vibration signals of the rotating machinery under different fault states, and perform standardization processing on the multiple sets of first vibration signals to obtain a first standard signal; collect multiple sets of second vibration signals of the rotating machinery under normal state, and perform standardization processing on the multiple sets of second vibration signals to obtain a second standard signal.

[0056] In this step, vibration sensors or acceleration sensors installed on rotating machinery can be used to collect multiple sets of first vibration signals of the rotating machinery under different fault conditions and multiple sets of second vibration signals of the rotating machinery under normal conditions.

[0057] The signal lengths of multiple sets of first vibration signals and multiple sets of second vibration signals can be unified using zero-padding or interpolation methods. The unified first vibration signals are then standardized to obtain a first standard signal; similarly, the multiple second vibration signals are standardized to obtain a second standard signal. This ensures that both the first and second standard signals conform to a standard normal distribution with a mean of 0 and a standard deviation of 1.

[0058] The first and second standard signals include multiple ones.

[0059] S111, perform phase space reconstruction on the first standard signal to obtain a first phase space trajectory matrix; perform spatial reconstruction on the second standard signal to obtain a second phase space trajectory matrix.

[0060] Since the first and second standard signals are one-dimensional vibration signals, in order to improve the expression of fault characteristics, this invention needs to reconstruct the phase space of the first standard signal to obtain the first phase space trajectory matrix.

[0061] Similarly, in order to facilitate the fusion of the amplified fault signal with the normal vibration signal, a second standard signal is also needed for spatial reconstruction to obtain the second phase spatial trajectory matrix.

[0062] The method for phase space reconstruction of the first standard signal and the second standard signal in this invention is the same. Here, we will take the phase space reconstruction of the first standard signal as an example for explanation:

[0063] In one implementation, phase space reconstruction is performed on the first standard signal to obtain a first phase space trajectory matrix, including:

[0064] Under different delay times, the autocorrelation function value of the first standard signal is determined by the autocorrelation function, and multiple autocorrelation function values ​​are obtained;

[0065] Each autocorrelation function value is compared with a preset first threshold. The autocorrelation function value that matches the first preset threshold is determined as the target autocorrelation function value. The delay time corresponding to the target autocorrelation function value is taken as the optimal delay time.

[0066] Determine the optimal embedding dimension of the first phase space trajectory matrix;

[0067] The first standard signal is reconstructed in a high-dimensional phase space based on the optimal embedding dimension and the optimal delay time, resulting in the first phase space trajectory matrix.

[0068] In one implementation, determining the autocorrelation function value of the first standard signal using the autocorrelation function includes:

[0069] The autocorrelation function value R(τ) of the first standard signal is determined using formula (1);

[0070]

[0071] In formula (1), N is the total length of the first standard signal, τ is the delay time, t is the sampling point of the first standard signal, and x(t) is the t-th first standard signal. Let σ be the mean of all first standard signals, and σ be the variance of all first standard signals.

[0072] In one implementation, the first standard signal is reconstructed in a high-dimensional phase space according to the optimal embedding dimension and the optimal delay time to obtain a first phase space trajectory matrix, including:

[0073] The first standard signal is reconstructed in the high-dimensional phase space according to formula (2) to obtain the first phase space trajectory matrix T;

[0074]

[0075] In formula (2), M is the total number of spatial samples in the first phase spatial trajectory matrix, and X M Let x(M) be the Mth spatial sample in the first phase spatial trajectory matrix, x(M) be the value of the Mth sampling point of the first standard signal in the time series, and x(M+τ′) be the value of the (M+τ′)th sampling point of the first standard signal in the time series. * -1)×τ′) represents the M+(d)th time series of the first standard signal. * -1)×τ′ values ​​of sampling points, where τ′ is the optimal delay time, d * This represents the optimal embedding dimension.

[0076] Specifically, phase space reconstruction of the first standard signal requires two important parameters: the optimal delay time and the optimal embedding dimension of the matrix.

[0077] The process of determining the optimal delay time is as follows:

[0078] This invention can pre-set multiple delay times τ, and calculate the autocorrelation function value of the first standard signal using formula (1) at each delay time. When it is determined that a certain autocorrelation function value is consistent with the first threshold, the autocorrelation function value is determined as the target autocorrelation function value, and the time delay corresponding to the target autocorrelation function value is determined as the optimal delay time τ′. The first threshold can be set according to specific circumstances, for example, it can be 1 / e.

[0079] The process of determining the optimal embedding dimension is as follows:

[0080] When initially reconstructing the first-phase spatial trajectory matrix, the optimal dimension is unknown. Therefore, it is necessary to construct initial first-phase spatial trajectory matrices of different dimensions for the one-dimensional vibration signal based on the initial embedding dimension. This facilitates subsequent calculations of the initial first-phase spatial trajectory matrices of different dimensions to determine the optimal embedding dimension. The initial embedding dimension is typically 3.

[0081] For example, when the embedding dimension is 3, the corresponding initial first phase space trajectory matrix is:

[0082]

[0083] When the embedding dimension is 4, the corresponding initial first phase space trajectory matrix is:

[0084] By analogy, multiple high-dimensional first phase space trajectory matrices can eventually be obtained.

[0085] For the initial first phase spatial trajectory matrix under different dimensions, determine the false nearest neighbor ratio of each spatial sample in the initial first phase spatial matrix. If the mean of the false nearest neighbor ratio of all spatial samples in the initial first phase spatial matrix is ​​lower than the preset false nearest neighbor ratio threshold (which can be 0.01), it indicates that the dimension is the optimal embedding dimension.

[0086] For example, if the embedding dimension is 5, the mean proportion of false neighbors is 0.02; and the embedding dimension is 6, the mean proportion of false neighbors is 0.005, then the optimal embedding dimension is 6.

[0087] The method for determining the proportion of false nearest neighbors can refer to the method for determining the proportion of false nearest neighbors in step S112, and will not be repeated here.

[0088] S112, determine the false nearest neighbor ratio of each spatial sample in the first phase space trajectory matrix, determine the weight corresponding to each spatial sample according to the false nearest neighbor ratio of each spatial sample, and determine the benchmark sample and the target nearest neighbor sample according to the weight of each spatial sample.

[0089] To ensure that subsequent fault sample amplification more closely reflects the actual fault distribution, this invention requires determining the proportion of false neighbors for each spatial sample in the first phase spatial trajectory matrix. Based on this proportion, the weight of each spatial sample is determined, and then the benchmark sample and target neighbor sample are identified according to their respective weights. Since the benchmark and target neighbor samples are selected based on weights—specifically, samples with higher weights are chosen as the benchmark and target neighbor samples—this improves the reliability of the fault samples. Furthermore, when amplifying fault samples based on these benchmark and target neighbor samples, the amplified fault samples are less likely to deviate from the true distribution.

[0090] In one implementation, determining the proportion of false nearest neighbors for each spatial sample in the first phase space trajectory matrix includes:

[0091] The proportion of false neighbors R(X) for each spatial sample is determined according to formula (3). i ):

[0092]

[0093] In formula (3), X iLet d be the i-th spatial sample in the first phase spatial trajectory matrix, i = 1, 2, ..., M; * For the optimal embedding dimension, Let i be the i-th spatial sample in the d-dimensional first phase space trajectory matrix. X is the trajectory matrix in the first phase space of d-dimensional space. i Let j be the j-th nearest neighbor sample, and k be the number of nearest neighbors of the i-th spatial sample. For the i-th spatial sample in the d+1 dimensional first phase space trajectory matrix, X is the first phase space trajectory matrix in d+1 dimensions. i+1 The j-th nearest neighbor sample, R th The preset distance threshold is used, and II() is an indicator function used to indicate when... When the output is 1, it represents X. j It is X i False nearest neighbors; otherwise, output 0, representing X. j It is X i A true neighbor.

[0094] When selecting the nearest neighbor for the i-th spatial sample, it is necessary to calculate the Euclidean distance between each of the remaining spatial samples and the i-th spatial sample, arrange all Euclidean distances in ascending order, and select the spatial samples corresponding to the first k Euclidean distances as the nearest neighbors of the i-th spatial sample. k can be set according to the actual situation and is not restricted here.

[0095] According to formula (3), the proportion of false neighbors for each spatial sample in the first phase spatial trajectory matrix can be determined. Then, the weight corresponding to each spatial sample can be determined based on the proportion of false neighbors for each spatial sample, including:

[0096] The i-th spatial sample X is determined according to formula (4). i weight w(X) i ):

[0097]

[0098] In formula (4), R(X) i ) represents the i-th spatial sample X i The false nearest neighbor ratio, θ is a preset second threshold, and the i-th spatial sample is any spatial sample in the first phase space trajectory matrix.

[0099] For example, suppose the second threshold is 0.6, if the i-th spatial sample X i If the proportion of false nearest neighbors is greater than 0.6, it indicates that the i-th spatial sample X... i Unreliable, low-stability sample, X iThe weight is reduced to 1 - 0.6 = 0.4. If the i-th spatial sample X i If the proportion of false nearest neighbors is less than or equal to 0.6, then it indicates that the i-th spatial sample X... i Reliable, a highly stable sample, then X i The weight is set directly to 1.

[0100] After the weight of each spatial sample is determined, the benchmark sample and the target nearest neighbor sample need to be determined based on the weight of each spatial sample. This ensures that when new fault samples are generated by interpolation based on the benchmark sample and the target nearest neighbor sample, the newly generated fault samples are more consistent with the actual fault distribution.

[0101] In one implementation, determining the benchmark sample based on the weight of each spatial sample includes:

[0102] The i-th spatial sample X is determined according to formula (5). i The first probability P(X) of the baseline sample i ):

[0103]

[0104] Spatial samples with a first probability greater than a third threshold are defined as baseline samples; the third threshold is the average of the first probabilities of all spatial samples; where,

[0105] w(X i ) represents the i-th spatial sample X i The weights are M, where M is the total number of spatial samples in the first phase spatial trajectory matrix.

[0106] Specifically, after calculating the first probability of each spatial sample, the average of the first probabilities of all spatial samples can be determined, and this average is used as the third threshold. Then, spatial samples with a first probability greater than the third threshold can be determined as benchmark samples.

[0107] Similarly, the target nearest neighbor samples are determined based on the weight of each spatial sample, including:

[0108] Determine the nearest neighbor set for each spatial sample;

[0109] For each nearest neighbor sample in the nearest neighbor set, the second probability P(X) of each nearest neighbor sample being the target nearest neighbor sample is determined according to formula (6). j |X i ):

[0110]

[0111] In formula (6), the nearest neighbor samples whose second probability is greater than the fourth threshold are determined as the target nearest neighbor samples; the fourth threshold is the average of the second probabilities of all nearest neighbor samples; where,

[0112] w(X j Let X be the j-th nearest neighbor sample. j The weights are M, M is the total number of spatial samples in the first phase spatial trajectory matrix, and N is the weight. k (X i Let be the set of nearest neighbor samples of the i-th spatial sample, and k be the total number of nearest neighbor samples.

[0113] This determines the baseline sample and the target nearest neighbor sample. In this embodiment, the baseline sample is denoted as S113. A new fault sample phase space trajectory matrix is ​​generated based on the baseline sample and the target nearest neighbor sample. The new fault sample phase space trajectory matrix and the second phase space trajectory are merged to obtain the fault dataset.

[0114] Then, a new fault sample phase space trajectory matrix can be generated based on the benchmark sample and the target nearest neighbor sample. The new fault sample phase space trajectory matrix and the second phase space trajectory are then merged to obtain the fault dataset.

[0115] In one implementation, generating a new fault sample phase space trajectory matrix based on a reference sample and the target nearest neighbor sample includes:

[0116] According to formula (7), a new fault space sample X is generated at the corresponding position in the first phase space trajectory matrix. new The phase space trajectory matrix of the new fault sample is obtained as follows:

[0117] X new =X i′ +λ×w(X i′ )×(X j′ -X i′ (7)

[0118] In formula (7), X i′ For the i′-th benchmark sample, w(X) i′ X represents the weight of the i′-th reference sample, λ is the interpolation position correction coefficient used to control the interpolation position, and the value of λ ranges from [0,1]. j′ Let i be the i′-th target nearest neighbor sample.

[0119] Weight w(X) i′ This is used to dynamically adjust the interpolation magnitude. The larger the weight, the closer the generated new fault sample is to the baseline sample, thus avoiding deviation from the true fault distribution. j′ -X i′ This represents the difference vector between the target's nearest neighbor samples and the baseline samples, used to determine the generation direction of new faulty samples.

[0120] To detect whether a new fault sample is located at the edge of the distribution or in a noisy region, the method further includes the following after generating the new fault sample:

[0121] The proportion of false neighbors R(X) of the new faulty sample is determined according to formula (8). new ):

[0122]

[0123] In formula (8), X new New fault sample; d * For the optimal embedding dimension, For new fault samples in the d-dimensional first-phase space trajectory matrix, Let j′ be the j′-th nearest neighbor sample of the new fault sample in the d-dimensional first phase space trajectory matrix, and k′ be the number of nearest neighbors of the new fault sample. For a new fault sample in the d+1 dimensional first phase space trajectory matrix, Let R be the j′-th nearest neighbor sample of the new fault sample in the d+1-dimensional first phase space trajectory matrix. th The preset distance threshold is used, and II() is an indicator function used to indicate when... When the time is right, the output is 1.

[0124] Then, the weights of the new faulty samples are updated based on the proportion of their false nearest neighbors, including:

[0125] The weights of the new fault samples are updated according to formula (9) to obtain the weights w of the new fault samples. updated (X new ):

[0126]

[0127] In formula (9), w(X) i ) represents the weight of the i′th benchmark sample, w(X′) j ) represents the weight of the j′-th target's nearest neighbor sample. The weights of the parent samples inherited by the new faulty sample, (1-R(X)) new ()) represents adjusting the corresponding weights based on the reliability of the new faulty samples. Then, the steps for generating new faulty samples are repeated until the number of faulty samples reaches the expected number (e.g., the same as the number of normal samples). This invention achieves this through (1-R(X) newThe weights of new fault samples are adjusted accordingly. In the subsequent process of generating new fault samples, the generation process can be adaptively adjusted according to the weights of the new fault samples (for example, new fault samples with small weights will not be identified as baseline samples), so that the generated samples are more consistent with the real fault distribution and the reliability of the entire dataset is improved.

[0128] After the phase space trajectory matrix of the new fault sample is determined, it is merged with the second phase space trajectory to obtain the fault dataset. Since the number of fault samples and normal samples in the fault dataset is the same, the balance of the samples can be ensured. Therefore, after training the pre-built fault diagnosis model using the fault dataset, even if the subsequent diagnosis is performed on a small number of samples, the fault features can still be identified with high accuracy, thus improving the accuracy of fault diagnosis.

[0129] S114, The fault dataset is used as a pre-built fault diagnosis model for training, and the trained fault diagnosis model is used to diagnose faults in rotating machinery.

[0130] Once the fault dataset is determined, it is used as a pre-built fault diagnosis model for training, and the trained fault diagnosis model is used to diagnose faults in rotating machinery.

[0131] The pre-built fault diagnosis model can be a Gaussian Process Regression (GRP) model, or other models, such as neural networks.

[0132] The most important aspect of Gaussian process regression (GPR) models is selecting a suitable kernel function and setting initial values ​​for hyperparameters to determine the prior model in the form of a probability distribution. The covariance function in a GPR model is the central moment of the random output variable corresponding to two random input points in space. It measures the degree of similarity or correlation between different samples and is a key factor affecting the predictive performance of the GPR model.

[0133] Assuming the fault dataset is D, and each fault space sample in the fault dataset has a corresponding fault label, and assuming the GRP model follows a Gaussian distribution, then the improved weighted covariance function K... w (X a ,X b It can be defined as:

[0134]

[0135] In formula (10), w(X) a Let w(X) be the weight of any sample in the fault dataset. b ) Sample X in the fault dataset b The weight of X;b For X a Nearest neighbor samples; σ f is the signal variance hyperparameter, used to control the overall magnitude of the covariance function. l is the length scale parameter, used to determine the smoothness of feature variations. ||X a -X b || Represents X b and X a The Euclidean distance is used to measure the similarity between two samples; The noise variance is represented by the Kronecker delta function δ, used to account for sensor measurement errors or signal disturbances. ab (1 when a = b, 0 otherwise) Ensure that the diagonal elements of the covariance matrix include noise correction to avoid matrix singularity problems.

[0136] It can be seen that the model suppresses the interference of edge or noisy samples by adjusting the weights of the samples, and uses the probability output of Gaussian process to quantify the prediction results, thereby improving the accuracy of fault diagnosis.

[0137] To further demonstrate the reliability of the accuracy of the fault diagnosis method provided by this invention, a comparative explanation is provided between ordinary fault diagnosis methods and the fault diagnosis method of this invention:

[0138] Vibration signals under different states were collected using an accelerometer. Since the normal state time of the equipment is often much longer than the abnormal state time, the number of abnormal state samples is much smaller than the number of normal samples, leading to a data sample imbalance problem between different states. The experimental data includes 100 normal state samples and 10 fault state samples, each sample with a length of 256. The number of normal state samples remained unchanged at 100, while the number of fault state samples was expanded from 10 to 100 to balance the prediction results of the classifier for normal state samples. Taking the confusion matrix after model convergence as an example, the number of normal and fault state samples in the test set remained the same, and the training and test sets were divided in a 7:3 ratio. Figure 2 As shown, the accuracy rate on the test set was 93.3%.

[0139] Fault samples are generated using the method of this invention, balancing fault-type samples with normal-state samples, and then input into a Gaussian process regression model for classification. For example... Figure 3 As shown, the model can accurately predict fault state samples without misclassification, and the accuracy on the test set can reach 100%.

[0140] This invention enhances the expressive power of fault features by reconstructing the phase space of the first vibration signal, transforming the one-dimensional vibration signal into a high-dimensional vibration signal. Furthermore, it determines the proportion of false nearest neighbors for each spatial sample and then assigns a corresponding weight based on this proportion. This allows low-stability samples located in edge noise regions to be assigned lower weights, while high-stability samples are assigned higher weights. The higher-weighted spatial samples (reference samples) serve as the benchmark for data augmentation. When generating new fault samples based on the benchmark samples and target neighbor samples, it ensures that the newly generated fault samples are close to the benchmark samples, preventing them from deviating from the true distribution and thus improving the reliability of the newly generated fault samples. The increased number of fault samples also avoids imbalance between fault diagnosis samples and normal samples, improving the accuracy of fault diagnosis for rotating machinery.

[0141] Based on the same inventive concept as in the foregoing embodiments, this embodiment also provides a fault diagnosis device for rotating machinery based on sample amplification, such as... Figure 4 As shown, the device includes:

[0142] Processing unit 41 is used to collect multiple sets of first vibration signals of the rotating machinery under different fault states, perform standardization processing on the multiple sets of first vibration signals to obtain a first standard signal; and collect multiple sets of second vibration signals of the rotating machinery under normal state, perform standardization processing on the multiple sets of second vibration signals to obtain a second standard signal.

[0143] The reconstruction unit 42 is used to perform phase space reconstruction on the first standard signal to obtain a first phase space trajectory matrix; and to perform phase space reconstruction on the second standard signal to obtain a second phase space trajectory matrix.

[0144] The determining unit 43 is used to determine the false nearest neighbor ratio of each spatial sample in the first phase space trajectory matrix; determine the weight corresponding to each spatial sample according to the false nearest neighbor ratio of each spatial sample; and determine the benchmark sample and the target nearest neighbor sample according to the weight of each spatial sample.

[0145] The generation unit 44 is used to generate a new fault sample phase space trajectory matrix based on the reference sample and the target nearest neighbor sample, and merge the new fault sample phase space trajectory matrix with the second phase space trajectory to obtain a fault dataset;

[0146] Training unit 45 is used to train the fault dataset as a pre-built fault diagnosis model and use the trained fault diagnosis model to diagnose faults in rotating machinery.

[0147] Since the apparatus described in the embodiments of this invention is used to implement the sample amplification-based fault diagnosis method for rotating machinery according to the embodiments of this invention, those skilled in the art can understand the specific structure and variations of the apparatus based on the method described in the embodiments of this invention, and therefore will not be described in detail here. All apparatuses used in the methods of the embodiments of this invention fall within the scope of protection of this invention.

[0148] Through one or more embodiments of the present invention, the present invention has the following beneficial effects or advantages:

[0149] This invention provides a method and apparatus for fault diagnosis of rotating machinery based on sample augmentation. The method includes: acquiring multiple sets of first vibration signals of the rotating machinery under different fault states; standardizing the multiple sets of first vibration signals to obtain a first standard signal; acquiring multiple sets of second vibration signals of the rotating machinery under normal state; standardizing the multiple sets of second vibration signals to obtain a second standard signal; reconstructing the phase space of the first standard signal to obtain a first phase space trajectory matrix; reconstructing the phase space of the second standard signal to obtain a second phase space trajectory matrix; determining the false nearest neighbor ratio of each spatial sample in the first phase space trajectory matrix; determining the weight corresponding to each spatial sample according to the false nearest neighbor ratio of each spatial sample; determining a reference sample and a target nearest neighbor sample according to the weight of each spatial sample; generating a new fault sample phase space trajectory matrix based on the reference sample and the target nearest neighbor sample; and integrating the new fault sample phase space trajectory matrix with the second phase space... Trajectories are merged to obtain a fault dataset. This fault dataset is then used to train a pre-built fault diagnosis model, which is used to diagnose faults in rotating machinery. By reconstructing the phase space of the first vibration signal, a one-dimensional vibration signal is converted into a high-dimensional vibration signal, enhancing the expressive power of fault features. Furthermore, a false nearest neighbor ratio is determined for each spatial sample, and the corresponding weight is determined based on this ratio. This allows low-stability samples located in edge noise regions to be assigned lower weights, while high-stability samples are assigned higher weights. The higher-weighted spatial samples (reference samples) serve as the benchmark for data augmentation. When generating new fault samples based on the benchmark samples and target neighbor samples, it ensures that the newly generated fault samples are close to the benchmark samples, preventing them from deviating from the true distribution and improving the reliability of the newly generated fault samples. The increased number of fault samples also helps avoid imbalance between fault diagnosis samples and normal samples, improving the accuracy of fault diagnosis for rotating machinery.

[0150] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0151] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for fault diagnosis of rotating machinery based on sample amplification, characterized in that, The method includes: Multiple sets of first vibration signals of the rotating machinery under different fault states are collected, and the multiple sets of first vibration signals are standardized to obtain a first standard signal; multiple sets of second vibration signals of the rotating machinery under normal state are collected, and the multiple sets of second vibration signals are standardized to obtain a second standard signal. The first standard signal is reconstructed in phase space to obtain a first phase space trajectory matrix; the second standard signal is reconstructed in phase space to obtain a second phase space trajectory matrix. Determine the false nearest neighbor ratio of each spatial sample in the first phase space trajectory matrix; determine the weight corresponding to each spatial sample based on the false nearest neighbor ratio of each spatial sample; and determine the benchmark sample and the target nearest neighbor sample based on the weight of each spatial sample. A new fault sample phase space trajectory matrix is ​​generated based on the benchmark sample and the target nearest neighbor sample. The new fault sample phase space trajectory matrix and the second phase space trajectory are merged to obtain the fault dataset. The fault dataset is used as a pre-built fault diagnosis model for training, and the trained fault diagnosis model is used to diagnose faults in rotating machinery.

2. The method as described in claim 1, characterized in that, The first standard signal includes multiple components, and the step of reconstructing the phase space of the first standard signal to obtain the first phase space trajectory matrix includes: Under different delay times, the autocorrelation function value of the first standard signal is determined using the autocorrelation function, resulting in multiple autocorrelation function values; Each autocorrelation function value is compared with a preset first threshold. The autocorrelation function value that matches the first threshold is determined as the target autocorrelation function value. The delay time corresponding to the target autocorrelation function value is taken as the optimal delay time. Determine the optimal embedding dimension of the first phase space trajectory matrix; The first standard signal is reconstructed in a high-dimensional phase space based on the optimal embedding dimension and the optimal delay time to obtain the first phase space trajectory matrix.

3. The method as described in claim 2, characterized in that, The step of determining the autocorrelation function value of the first standard signal using the autocorrelation function includes: Using formula Determine the autocorrelation function value of the first standard signal ;in, N The total length of the first standard signal, the The delay time is t, where t is the sampling point of the first standard signal. For the first t The first standard signal, the The mean of all first standard signals, the Let be the variance of all first standard signals.

4. The method as described in claim 2, characterized in that, The step of reconstructing the first standard signal in a high-dimensional phase space based on the optimal embedding dimension and the optimal delay time to obtain the first phase space trajectory matrix includes: According to the formula The first standard signal is reconstructed in a high-dimensional phase space to obtain the first phase space trajectory matrix T; The M The total number of spatial samples in the first phase space trajectory matrix, the The first phase space trajectory matrix is ​​the first... M A spatial sample, the The first standard signal in the time series M The values ​​of each sampling point, the For the optimal delay time, the Let be the optimal embedding dimension.

5. The method as described in claim 1, characterized in that, Determining the proportion of false nearest neighbors for each spatial sample in the first phase space trajectory matrix includes: According to the formula Determine the proportion of false nearest neighbors for each spatial sample. ; The The first phase space trajectory matrix is ​​the first... i One spatial sample, i =1,2…… M The For the optimal embedding dimension, the for d The first phase space trajectory matrix in the first dimension i A spatial sample, the for d The first phase space trajectory matrix described The j The nearest neighbor samples, the k For the first i The number of nearest neighbors of a spatial sample, the for d+ The first phase space trajectory matrix in 1D i One spatial sample, for d+ The 1D first phase space trajectory matrix described The j The nearest neighbor samples, the The preset distance threshold, the ( ) is an indicator function used to indicate when When the time is right, the output is 1.

6. The method as described in claim 1, characterized in that, The step of determining the weight corresponding to each spatial sample based on the proportion of false nearest neighbors for each spatial sample includes: According to the formula Determine the first i spatial samples weight ;in, The For the first i spatial samples The proportion of false nearest neighbors, the The second threshold is a preset threshold, the first i Each spatial sample is any spatial sample in the first phase spatial trajectory matrix.

7. The method as described in claim 1, characterized in that, The step of determining the benchmark sample based on the weight of each spatial sample includes: According to the formula Determine the first i spatial samples The first probability of the benchmark sample ; Spatial samples with a first probability greater than a third threshold are defined as the baseline samples; the third threshold is the average of the first probabilities of all spatial samples; wherein... The For the first i spatial samples The weights, the M This represents the total number of spatial samples in the first phase space trajectory matrix.

8. The method as described in claim 1, characterized in that, The step of determining the target nearest neighbor sample based on the weight of each spatial sample includes: Determine the nearest neighbor set for each spatial sample; For each nearest neighbor sample in the nearest neighbor set, according to the formula The second probability of determining each of the nearest neighbor samples as the target nearest neighbor sample. ; The nearest neighbor samples whose second probability is greater than the fourth threshold are determined as the target nearest neighbor samples; the fourth threshold is the average of the second probabilities of all nearest neighbor samples; wherein, The For the first j Nearest neighbor samples The weights, the For the first i The set of nearest neighbor samples of a spatial sample, the k For the first i The number of nearest neighbors of a spatial sample.

9. The method as described in claim 1, characterized in that, The step of generating a new fault sample phase space trajectory matrix based on the benchmark sample and the target nearest neighbor sample includes: According to the formula New fault space samples are generated at the corresponding positions in the first phase space trajectory matrix. The phase space trajectory matrix of the new fault sample is obtained; The For the first The reference samples, For the first The weights of the reference samples, the These are the interpolation position correction coefficients. The value range of is [0,1]. For the first Target nearest neighbor samples.

10. A fault diagnosis device for rotating machinery based on sample amplification, characterized in that, The device includes: The processing unit is used to collect multiple sets of first vibration signals of the rotating machinery under different fault states, perform standardization processing on the multiple sets of first vibration signals to obtain a first standard signal; and collect multiple sets of second vibration signals of the rotating machinery under normal state, perform standardization processing on the multiple sets of second vibration signals to obtain a second standard signal. The reconstruction unit is used to perform phase space reconstruction on the first standard signal to obtain a first phase space trajectory matrix; and to perform phase space reconstruction on the second standard signal to obtain a second phase space trajectory matrix. A determining unit is used to determine the false nearest neighbor ratio of each spatial sample in the first phase space trajectory matrix; determine the weight corresponding to each spatial sample based on the false nearest neighbor ratio of each spatial sample; and determine the benchmark sample and the target nearest neighbor sample based on the weight of each spatial sample. The generation unit is used to generate a new fault sample phase space trajectory matrix based on the reference sample and the target nearest neighbor sample, and merge the new fault sample phase space trajectory matrix with the second phase space trajectory to obtain a fault dataset; The training unit is used to train the fault dataset as a pre-built fault diagnosis model, and to use the trained fault diagnosis model to diagnose faults in rotating machinery.

Citation Information

Patent Citations

  • Patch clamp electrophysiological data processing method and system

    CN119632565A

  • Transformer online life prediction method under unbalanced multi-source small sample condition

    CN119807849A