High-reliability fault diagnosis method based on physical-data joint optimization driving and related device

Through the physical-data joint optimization-driven fault diagnosis method, combined with prior physics knowledge and deep learning technology, the problems of high false alarm rate and insufficient diagnosis real-time in the existing technology are solved, achieving higher fault diagnosis accuracy and real-time.

CN120046003APending Publication Date: 2025-05-27UNIV OF ELECTRONICS SCI & TECH OF CHINA

Patent Information

Application Number
CN202510201944.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27

Smart Images

  • Figure CN120046003A_ABST
    Figure CN120046003A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of rotating equipment fault diagnosis, and discloses a high-reliability fault diagnosis method driven by physical-data joint optimization and a related device. The method comprises the following steps: acquiring vibration data of operation of the rotating equipment; inputting the vibration data into a pre-trained deep learning diagnosis model fusing prior physical knowledge and data, and obtaining a fault diagnosis result of the to-be-diagnosed rotating equipment; wherein the deep learning diagnosis model fusing the physical knowledge and the data is composed of a data-driven feature extraction module, a feature extraction module based on the physical knowledge, a feature fusion layer and a fault classification layer; the architecture parameters of the deep learning diagnosis model are obtained by a self-adaptive optimization method based on a constrained Gaussian process model. A high-frequency wireless vibration sensor fused with edge calculation is also integrated and is used for real-time fault diagnosis. According to the method, the constraint of physical knowledge is introduced into the data-driven diagnosis model, so that the diagnosis interpretability and reliability can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of rotating equipment fault diagnosis, and relates to a highly reliable fault diagnosis method and related device driven by physical-data joint optimization. Background Art

[0002] Equipment condition monitoring and fault diagnosis are crucial in the operation and maintenance of modern industrial equipment. A real-time, accurate, and highly reliable fault diagnosis method can timely detect potential problems, reduce downtime, and extend the equipment life. However, due to factors such as sensor signals being susceptible to noise and environmental interference, the diagnostic model's insufficient adaptability to complex working conditions, and data processing delays, existing diagnostic technologies generally suffer from high false alarm rates and insufficient diagnostic real-time performance. A high false alarm rate may lead to normal equipment being misjudged as faulty, triggering unnecessary shutdowns and maintenance, increasing maintenance costs. At the same time, frequent false alarms will also disrupt the production process, reduce efficiency, and delay production plans. On the other hand, with the increase in equipment operating speed and working condition complexity, existing methods are difficult to timely process massive monitoring data and accurately judge faults. Such delays may result in faults not being detected in time, missing the best maintenance window, and further increasing the equipment operation risk.

[0003] Traditional fault diagnosis methods, such as time-frequency domain statistical analysis, Hilbert transform, and Fourier spectrum analysis, usually rely on prior physical models and empirical rules, and predict potential faults by analyzing specific features in equipment operation data. These methods have certain diagnostic capabilities and interpretability under clear fault modes and stable working conditions. However, in actual industrial applications, equipment operating conditions are often complex and variable, and fault modes have a high degree of uncertainty and complexity. In this case, traditional methods, due to their heavy reliance on preset fault mode libraries and physical models, lack the adaptability to newly emerging unknown fault modes, resulting in a decline in their diagnostic accuracy under complex working conditions and being difficult to meet the requirements of modern industry for high reliability. In addition, traditional time-domain and frequency-domain analysis methods mainly capture the basic statistical characteristics of signals, such as mean, frequency, and amplitude, etc., and are difficult to effectively mine the non-linear, non-stationary characteristics in signals and the key information hidden in noise, resulting in lower diagnostic result accuracy. Especially when signal features are fuzzy or masked, false alarms are likely to occur.

[0004] In recent years, fault diagnosis methods based on deep learning have received extensive attention and applications. Deep learning methods can automatically analyze large-scale data of device monitoring signals, extract potential health information from them, and thus improve the accuracy of fault diagnosis. Commonly used deep learning models, such as convolutional neural networks (CNNs), long short-term memory networks (LSTMs), and autoencoders, can automatically discover potential fault information from a large amount of high-frequency vibration signals without pre-defining fault features. However, most of the existing methods are still mainly data-driven, lacking in-depth understanding of device physical principles and fault mechanisms, and having certain deficiencies in interpretability and physical constraints, resulting in the diagnostic results being easily affected by noise and thus having a relatively high false alarm rate. Therefore, how to combine physical knowledge with deep learning models, make full use of existing prior physical knowledge and measured data information to reduce the false alarm rate of diagnosis and improve the real-time performance of diagnosis has become a key technical problem urgently to be solved in the current research in the field of fault diagnosis. Summary of the Invention

[0005] The purpose of the present invention is to overcome the above-mentioned shortcomings of the existing technologies and provide a highly reliable fault diagnosis method and related device driven by joint optimization of physics and data.

[0006] In the first aspect of the present invention, a highly reliable fault diagnosis method driven by joint optimization of physics and data is provided, including:

[0007] Extracting time-frequency domain features of vibration signals based on prior physical knowledge;

[0008] Building a deep learning network that fuses prior physical knowledge and data;

[0009] Optimizing adaptive network parameters based on a constrained Gaussian process model;

[0010] Among them, the purpose of extracting time-frequency domain features of vibration signals based on prior physical knowledge is to convert prior physical knowledge into features that can be read by deep learning. The present invention mainly considers the time-domain and frequency-domain features contained in bearing vibration signals. In terms of time-domain features, the present invention mainly involves three features: skewness, kurtosis, and root mean square error.

[0011] Although the present invention only uses skewness, kurtosis, and root mean square error in bearing vibration signals as a set of statistical health features to represent prior physical knowledge, it should be noted that other statistical features can also be integrated with the subsequent proposed deep network as prior physical knowledge. In terms of frequency-domain features, the original vibration signal is processed using the Hilbert transform to obtain the corresponding envelope vibration signal.

[0012] The envelope spectrum F(ω) is obtained by calculating the envelope signal a(t) using the fast Fourier transform method.

[0013] Through the spectrum F(ω), a specified range [f FCF -f se , f FCF +f se is selected to calculate the fault spectrum energy within the predefined sidebands.

[0014] The purpose of building a deep learning network that fuses prior physical knowledge and data is to integrate prior physical knowledge and deep learning models. Therefore, during the knowledge-based deep network configuration phase, both a data-driven feature extraction module and a knowledge-based feature extraction module are set up.

[0015] The knowledge-based deep network framework includes a knowledge-based feature extraction module and a data-driven feature extraction module composed of multiple convolutional layers, batch normalization layers, activation layers, pooling layers, feature fusion layers, and classification layers.

[0016] In each convolutional layer, the input is convolved through a learnable convolutional kernel to generate a new feature map, which serves as the input passed to the next layer.

[0017] The rectified linear unit (ReLU) can be used as the activation function in the neurons of the intermediate hidden layers. Then, the max pooling layer retains the key features while reducing the dimension.

[0018] It is worth noting that there are no parameters in the pooling layer. The last pooling layer is connected to the fully connected layer. In the proposed deep learning model, three dense layers are used to map the extracted feature vectors to the system health state.

[0019] The data-driven features obtained by the dense layer are obtained through the learning process of the network (involving weight adjustment and non-linear transformation), which provides a useful discriminative representation of the original data. To utilize prior physical knowledge and deep learning, a feature fusion layer is defined to fuse the data-driven features MD and knowledge-based features MK extracted from the first dense layer of the CNN network. Then, the new concatenated feature vector Xmerg will be used as the fused feature map for fault classification.

[0020] In the proposed method, a dropout layer is used, which randomly discards some information in the training data to prevent overfitting of the deep network. The dropout rate is usually set between 0 and 1 to improve the generalization ability of the model by reducing the co-adaptation of the training data. Subsequently, Softmax is used as the activation function for bearing diagnosis.

[0021] The purpose of adaptive network parameter optimization based on the constrained Gaussian process model is to quickly obtain the architecture parameters of a deep network with high generalization ability. The optimization of the learnable parameters in the deep network is obtained using the backpropagation method. At the same time, the deep network architecture parameters are determined using an adaptive constrained Gaussian process.

[0022] The generalization ability of a knowledge-based deep network varies with different architecture parameters and learnable parameters. Therefore, its structure design can be expressed as an optimization problem.

[0023] Let θ and P represent the architecture parameters and learnable parameters of the deep network, respectively. Let σ represent the metric based on the generalization ability, which is used to accurately quantify the generalization ability of the deep network.

[0024] It should be noted that the value range of the metric σ based on the generalization ability is [-1, 1]. Considering the constraints imposed by σ, an optimization method based on the constrained Gaussian process (CGP) is adopted as the optimizer.

[0025] It should be noted that the set θ of virtual observation positions v should be dense enough so that the constraints are satisfied with a high enough probability at any input position in a certain bounded set . Therefore, a target probability p target ∈ [0, 1) is specified to find the set θ v . When θ v satisfies the constraints at all virtual observation positions, the probability that the constraints hold for any θ * ∈ Ω is at least p target .

[0026] Here, the non-negative fixed number υ is used to ensure that p c can be increased by observations with additional noise.

[0027] Within the predefined design space, m random network architectures are generated using Latin hypercube sampling The corresponding deep networks will be trained to obtain the corresponding generalization ability values G KIDN = [G e1 , G e2 ,..., G em T . Based on a set of initial samples a constrained Gaussian process model is constructed using the maximum likelihood function.

[0028] After constructing the initial constrained Gaussian process model, the expected improvement metric is used to quantify the potential contribution of new sample points to the current model.

[0029] Subsequently, the sample point with the maximum expected improvement value is selected to update the constrained Gaussian process model. The update process will be iteratively repeated until the maximum expected improvement value is lower than the critical threshold I c . Here, I c = |G emin / |100 represents 1% of the absolute value of the current best global minimum estimate.

[0030] Using the proposed method for extracting time-frequency domain features of vibration signals based on prior physical knowledge, the corresponding time-frequency domain features are extracted, including skewness, kurtosis, root mean square error, and fault spectrum energy. Subsequently, using the method for constructing a deep learning network that fuses prior physical knowledge and data, a deep learning diagnosis model is constructed. The constructed model is a network with three convolutional layers, three pooling layers, and three fully connected layers.

[0031] Furthermore, 2340 samples are used, with 70% for training and the rest for validation. Adaptive network parameter optimization based on a constrained Gaussian process model is adopted to determine the optimal structural parameters, and the optimal structural parameters are obtained after 9 iterations. The number of filters in the three convolutional layers is 9, 8, and 10 respectively. The filter sizes of the three convolutional layers are 84, 90, and 41 respectively. The pooling window sizes are 4, 4, and 23 respectively. The dropout rate is 0.32, and the optimal number of neurons in the fully connected layers is 5, 29, and 22 respectively. The validation accuracy in the last 5 training cycles of the training curve based on the validation set under the optimal structural parameters is 100%. To verify the advantages of the proposed method compared with other similar methods, the fault diagnosis accuracies of different methods are used. The results show that the proposed method can accurately identify faults.

[0032] In the second aspect of the present invention, a wireless vibration acquisition device integrating edge computing that can run a highly reliable fault diagnosis algorithm driven by physical-data joint optimization is provided, which consists of a vibration acquisition board, an edge computing circuit board, a power supply battery, a main control circuit board, and a liquid crystal display screen.

[0033] Vibration acquisition board: Embedded with a vibration signal acquisition module, an integrated electronic piezoelectric (IEPE) sensor is used to collect high-frequency vibration signals.

[0034] Edge computing circuit board: This module is embedded with an edge computing module, mainly used to run a highly reliable fault diagnosis method driven by physical-data joint optimization to ensure the real-time and reliable assessment of the operating state of the device.

[0035] Power supply battery: Used to supply power to the fault diagnosis device.

[0036] Liquid crystal display screen: An embedded result display module is used to display calculation results and device status information in real time. In addition, the integrated small liquid crystal screen enables the sensor to have an independent display function. Even when wireless communication is restricted or remote monitoring cannot obtain data in a timely manner, users can still monitor the operation of the device in real time through the liquid crystal screen.

[0037] Wireless transmitter: An embedded wireless transmission module. In this module, wireless transmission technology is used for bearing vibration monitoring, and a wireless communication module is integrated to achieve remote monitoring and analysis of vibration data and calculation results.

[0038] In the third aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned high-reliability fault diagnosis method based on physical-data joint optimization driving are realized.

[0039] In the fourth aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned high-reliability fault diagnosis method based on physical-data joint optimization driving are realized.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] The modeling method that integrates prior physical knowledge and data improves the accuracy of the diagnosis result: The present invention fully combines the time-domain and frequency-domain characteristics in the vibration signal, and uses this as the fusion of prior physical knowledge and key health information in the condition monitoring data to improve the accuracy of the diagnosis model. Compared with the prior art, the fault diagnosis accuracy rate of the present invention is increased by more than 99%.

[0042] The adaptive optimization method based on the constrained Gaussian process model realizes the efficiency of the diagnostic model architecture parameter adjustment process: The present invention adopts an adaptive optimization algorithm based on the constrained Gaussian process model, which can quickly search for the optimal architecture parameters of the diagnostic model. Compared with the prior art, the present invention effectively overcomes the deficiencies of the traditional model architecture parameter adjustment process that strongly relies on subjective experience and takes a long time in the adjustment process, and provides an efficient and intelligent solution for the fault diagnosis of complex equipment.

[0043] The high-frequency wireless vibration sensor that integrates edge computing ensures the real-time processing and transmission of complex equipment fault diagnosis information: The present invention integrates vibration signal acquisition, signal processing and fault diagnosis, result display, and wireless transmission. Through edge computing technology, real-time fault diagnosis is directly carried out at the front end of the device, significantly improving the response speed and data processing efficiency, and providing real-time technical support for the accurate diagnosis and effective prevention of early faults. Brief Description of the Drawings

[0044] Figure 1 This is the technical roadmap of the embodiment of the present invention.

[0045] Figure 2 This is the knowledge-based deep network framework diagram of the embodiment of the present invention.

[0046] Figure 3 This is the adaptive network parameter optimization flowchart based on the constrained Gaussian process model of the embodiment of the present invention.

[0047] Figure 4 This is the schematic diagram of the publicly available bearing vibration data signal of Case Western Reserve University in the embodiment of the present invention. Figure 4 (a) is the schematic diagram of the data signal of the inner raceway fault of the bearing. Figure 4 (b) is the schematic diagram of the data signal of the roller element fault. Figure 4 (c) is the schematic diagram of the data signal of the outer raceway fault. Figure 4 (d) is the schematic diagram of the healthy signal.

[0048] Figure 5 This is the optimization iteration process of the network parameters in the embodiment of the present invention.

[0049] Figure 6 This is the training curve based on the validation set under the optimal structural parameters in the embodiment of the present invention.

[0050] Figure 7 This is a highly reliable fault diagnosis device based on the joint optimization of physics and data in the embodiment of the present invention. Detailed implementation manners

[0051] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0052] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0053] The present invention will be further described in detail below with reference to the drawings:

[0054] See Figure 1 , in an embodiment of the present invention, a highly reliable fault diagnosis method based on physical-data joint optimization drive is provided, including the following steps:

[0055] Step 1, the purpose of extracting the time-frequency domain features of the vibration signal based on prior physical knowledge is to convert the prior physical knowledge into features that can be read by deep learning. The present invention mainly considers the time-domain and frequency-domain features contained in the bearing vibration signal. In terms of time-domain features, the present invention mainly involves three features: skewness, kurtosis, and root mean square error. Among them, skewness can be defined by the mathematical formula:

[0056]

[0057] In the formula, μ and are the mean and standard deviation of the vibration signal x(t) respectively, and N is the number of data points. Kurtosis is mainly used to describe the distribution of the vibration signal around its average value and can be expressed as:

[0058]

[0059] The root mean square error can be defined by the mathematical formula:

[0060]

[0061] Although the present invention only uses skewness, kurtosis, and root mean square error in the bearing vibration signal as a set of statistical health features to represent prior physical knowledge, it should be noted that other statistical features can also be used as prior physical knowledge and integrated with the subsequent proposed deep network. In terms of frequency-domain features, the Hilbert transform is used to process the original vibration signal to obtain the corresponding envelope vibration signal. The principle of the Hilbert transform can be expressed as:

[0062]

[0063] In the formula, (p.v.) represents the Cauchy principal value, and H(t) represents the sample of the Hilbert transform signal at time t. By coupling x(t) and H(t), the analytic signal z(t) of x(t) is expressed as:

[0064] z(t) = x(t) + iH(t) = a(t)e iΨ(t)

[0065] In the formula, the instantaneous phase of x(t) Subsequently, the corresponding envelope signal a(t) can be calculated as:

[0066] a(t) = |z(t)| = (x 2 (t) + H 2 (t)) 1 / 2

[0067] Using the fast Fourier transform method to calculate the envelope signal a(t), the envelope spectrum F(ω) is obtained. The process can be expressed as

[0068]

[0069] In the spectrum, the fault characteristic frequencies fFCF of different types of bearing faults can be expressed as

[0070]

[0071] In the formula, fFCF,out, fFCF,in, and fFCF,ball represent the fault characteristic frequencies of the bearing outer race, inner race, and ball respectively, db is the ball diameter, dp is the pitch diameter, Nb is the number of balls, and β is the ball contact angle. Through the spectrum F(ω), a specified range [fFCF - fse, fFCF + fse] is selected to calculate the fault spectrum energy within the predefined sidebands. In the present invention, the fault spectrum energy is defined as:

[0072]

[0073] In the formula, SEout, SEin, and SEball represent the fault spectrum energies of the outer race defect, inner race defect, and ball defect respectively, and fse represents half of the bandpass range.

[0074] Step 2, the purpose of building a deep learning network that fuses prior physical knowledge and data is to integrate prior physical knowledge and a deep learning model. Therefore, in the knowledge-based deep network configuration stage, a data-driven feature extraction module and a knowledge-based feature extraction module are simultaneously set up.

[0075] Such as Figure 2As shown, the knowledge-based deep network framework includes a knowledge-based feature extraction module and a data-driven feature extraction module composed of multiple convolutional layers, batch normalization layers, activation layers, pooling layers, feature fusion layers, and classification layers. Among them, the forward propagation process of the data-driven feature extraction module in the proposed deep network can be expressed as

[0076] f(x(t)) = f L (...f 2 (f 1 (x(t), p (1) ), p 2 )..., p (L) )

[0077] In the formula, p(1), p(2), …, p(L) are learnable parameters, such as weights and biases; L represents the number of stages; f1, f2, …, fL are the corresponding operations in these stages. In each convolutional layer, the input is convolved through a learnable convolutional kernel to generate a new feature map, which is used as the input passed to the next layer. If Xl-1 and Xl represent the input and output of the convolutional layer respectively, then the j-th feature map can be expressed as:

[0078]

[0079] In the formula, * represents the convolution operation, fact(+) represents the activation function, k represents the convolutional kernel, and b represents the bias. The rectified linear unit (ReLU) can be used as the activation function in the neurons of the intermediate hidden layer. Then, the max pooling layer retains the key features while reducing the dimension, which can be expressed as:

[0080]

[0081] In the formula, (i', j') and (i, j) represent the positions of the input feature map Xin and the output feature map Xout respectively, and w and h represent the width and height of the pooling window. It should be noted that there are no parameters in the pooling layer. The last pooling layer is connected to the fully connected layer. In the proposed deep learning model, three dense layers are used to map the extracted feature vectors to the system health state. The dense layer has a matrix-vector multiplication operation, and the result passes through the activation function, which is expressed as:

[0082] M D = f act (X input · k + b)

[0083] In the formula, X inputRepresents the input of the dense layer, MD represents the output feature vector of the dense layer, k represents the weight data, "·" represents the dot product of the input and the corresponding weights, and b is the bias value. The data-driven features obtained from the dense layer are obtained through the learning process of the network (involving weight adjustment and non-linear transformation), which provides a useful discriminative representation of the original data. To utilize prior physical knowledge and deep learning, a feature fusion layer is defined to fuse the data-driven features MD extracted from the first dense layer of the CNN network and the knowledge-based features MK. Then, the new concatenated feature vector Xmerg will be used as the fused feature map for fault classification, expressed as:

[0084] X merg =[M D ,M K

[0085] In the proposed method, a dropout layer is used, which randomly discards some information in the training data to prevent overfitting of the deep network. The dropout rate is usually set between 0 and 1, and the generalization ability of the model is improved by reducing the co-adaptation of the training data. Subsequently, Softmax is used as the activation function for bearing diagnosis. If the label set contains g categories, the probability of the m-th category occurring can be expressed as:

[0086]

[0087] where w is the weighted vector. For multi-class classification, a separate loss is calculated for each class label and the results are summed:

[0088]

[0089] where Labeli is the true label indicator for the i-th class and Probi is the Softmax probability for the i-th class.

[0090] Step 3, the purpose of the adaptive network parameter optimization based on the constrained Gaussian process model is to quickly obtain the architecture parameters of the deep network with high generalization ability. As Figure 3 shown, the optimization of the learnable parameters in the deep network is obtained using the backpropagation method. At the same time, the architecture parameters of the deep network are determined using an adaptive constrained Gaussian process.

[0091] The generalization ability of the knowledge-based deep network varies with different architecture parameters and learnable parameters. Therefore, its structure design can be expressed as an optimization problem, expressed as

[0092]

[0093] ​In the formula, θ and P respectively represent the architecture parameters and learnable parameters of the deep network. σ represents a metric based on the generalization ability, which is used to accurately quantify the generalization ability of the deep network and can be defined as:

[0094] G e = A test - |A train - A test |

[0095] In the formula, A train and A test respectively represent the training accuracy and test accuracy of the deep network. It should be noted that the value range of the metric σ based on the generalization ability is [-1, 1]. Considering the constraints imposed by σ, an optimization method based on Constrained Gaussian Process (CGP) is adopted as the optimizer. CGP can be expressed as:

[0096] f|θ, G e , θ v , C(θ v ): = f|f(θ) + ε = G e , a(θ v ) ≤ ζf(θ v ) + ε v ≤ b(θ v )

[0097] In the formula, f ~ GP(μ(θ), K(θ, θ')) represents a Gaussian process with mean μ(θ) and covariance function K(θ, θ'). ζ is a linear operation on the realization of f ~ GP(μ(θ), K(θ, θ')). ε and ε v respectively represent a multivariate Gaussian distribution with a diagonal covariance matrix, where σ 2 and are its diagonal elements. is a finite set of virtual observation positions. C(θ v ) represents that the linear operation inequality constraint is satisfied for all points in θ v and can be defined as:

[0098]

[0099] It should be noted that the set θ v of virtual observation positions should be dense enough so that the constraints are satisfied with a high enough probability at any input position in a certain bounded set . Therefore, a target probability p target ∈ [0, 1) is specified to find the set θ v . When θ v satisfies the constraints at all virtual observation positions, the constraints hold for any θ *The probability of ∈Ω is at least p target , which can be expressed as:

[0100] p c (θ) = P(a(θ) - υ < ξ(θ, θ v ) < b(θ) + υ)

[0101] where, when θ v = φ, ξ(θ, θ v ) = ζf(θ * )|G e otherwise ξ(θ, θ v ) = ζf(θ * )|G e , C. Here, the non - negative fixed number v is used to ensure that p can be increased by observations with additional noise c , which can be expressed as:

[0102] v = max{σ v Φ -1 (p target ), 0}

[0103] Within the predefined design space, m random network architectures are generated using Latin - hypercube sampling The corresponding deep networks will be trained to obtain the corresponding generalization ability values G KIDN = [G e1 , G e2 ,..., G em T . Based on a set of initial samples Using the maximum - likelihood function to construct a constrained Gaussian process model, it can be expressed as:

[0104] L(η) = p(G e , C|η) = p(G e |η)p(C|G e , η)

[0105] where, the mean μ(θ|η) and covariance function K(θ, θ'|η) of the Gaussian process prior are both assumed to depend on η. After constructing the initial constrained Gaussian process model, the expected improvement metric is used to quantify the potential contribution of new sample points to the current model, which can be expressed as:

[0106] I(θ) = max(G emin - f(θ), 0)

[0107] where, G emin represents the global minimum of the estimated generalization ability in the current update iteration. The corresponding expected improvement E(I(θ)) can be expressed as:

[0108] ​

[0109] In the formula, Φ(·) represents the cumulative distribution function of the standard Gaussian distribution, and φ(·) represents the probability density function of the standard Gaussian distribution. Subsequently, the sample point with the maximum expected improvement value is selected to update the constrained Gaussian process model. The update process will be iteratively repeated until the maximum expected improvement value is lower than the critical threshold I c . Here, I c = |G emin | / 100 represents 1% of the absolute value of the current best global minimum estimate.

[0110] To verify the effectiveness of the method, taking the publicly available dataset of Case Western Reserve University as an example, the sampling frequency is set to 12 kHz, Figure 4 which is the real vibration signal that has been collected.

[0111] Using the vibration signal time-frequency domain feature extraction method based on prior physical knowledge proposed in step 1, the corresponding time-frequency domain features are extracted, including skewness, kurtosis, root mean square error, and fault spectrum energy. Subsequently, using the deep learning network construction method that fuses prior physical knowledge and data mentioned in step 2, a deep learning diagnosis model is constructed. The constructed model is a network with three convolutional layers, three pooling layers, and three fully connected layers. Among them, the range of structural parameters is shown in Table 1.

[0112] Table 1

[0113] Structural parameters <![CDATA[θ l > <![CDATA[θ h > Number of convolutional layer filters 1 50 Size of convolutional layer filters 10 100 Pooling window size 1 50 Number of neurons in the fully connected layer 1 50 Dropout rate 0.2 0.9

[0114] Furthermore, using 2340 samples, 70% of which are used for training and the rest for validation. The optimal structural parameters are determined by using the adaptive network parameter optimization based on the constrained Gaussian process model mentioned in step 3. The optimization process is as Figure 5 shown. It can be seen that the optimal structural parameters are obtained after 9 iterations. The number of filters in the three convolutional layers is 9, 8, and 10 respectively. The filter sizes of the three convolutional layers are 84, 90, and 41 respectively. The pooling window sizes are 4, 4, and 23 respectively. The dropout rate is 0.32, and the optimal number of neurons in the fully connected layers is 5, 29, and 22 respectively. Figure 6 is the training curve based on the validation set under the optimal structural parameters, and the validation accuracy in the last 5 training epochs is 100%. To verify the advantages of the proposed method compared with other similar methods, Table 2 shows the fault diagnosis accuracies of different methods. The results show that the proposed method can accurately identify faults.

[0115] Table 2,

[0116] Method KNN SVM ANN 1D-CNN LSTM Proposed method Accuracy 97.62% 98.51 98.90% 98.99% 98.48% 100%

[0117] In the second aspect of the present invention, a wireless vibration acquisition device integrating edge computing that can run a highly reliable fault diagnosis algorithm driven by physical-data joint optimization is provided. It consists of a vibration acquisition board, an edge computing circuit board, a power supply battery, a main control circuit board, and a liquid crystal display. Its overall structure is as shown in Figure 7 shown below. The functions of each component are described as follows:

[0118] (1) Vibration acquisition board: It is embedded with a vibration signal acquisition module, and an integrated electronic piezoelectric (IEPE) sensor is used to collect high-frequency vibration signals.

[0119] (2) Edge computing circuit board: This module is embedded with an edge computing module, which is mainly used to run a highly reliable fault diagnosis method driven by physical-data joint optimization to ensure the real-time and reliable evaluation of the operating state of the device.

[0120] (3) Power supply battery: It is used to supply power to the fault diagnosis device.

[0121] (4) Liquid crystal display: It is embedded with a result display module, which is used to display the calculation results and device status information in real time. In addition, the integrated small liquid crystal screen enables the sensor to have an independent display function. Even when wireless communication is restricted or remote monitoring cannot obtain data in time, users can still monitor the operation of the device in real time through the liquid crystal screen.

[0122] (5) Wireless transmitter: It is embedded with a wireless transmission module. In this module, wireless transmission technology is used for bearing vibration monitoring, and a wireless communication module is integrated to realize remote monitoring and analysis of vibration data and calculation results.

[0123] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0124] product.

[0125] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0126] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A high-reliability fault diagnosis method and related device based on physical-data joint optimization drive, including time-frequency domain feature extraction of vibration signals based on prior physical knowledge; construction of a deep learning network integrating prior physical knowledge and data; adaptive network parameter optimization based on a constrained Gaussian process model, characterized in that: A novel deep learning diagnosis framework that integrates prior physical knowledge and data is proposed. This framework extracts time domain and frequency domain features from vibration signals and combines them with key health information extracted from original monitoring data as prior physical knowledge, thereby enhancing the diagnostic model's ability to identify fault features and improving the accuracy of fault diagnosis. A strategy for designing an adaptive diagnostic model architecture based on a constrained Gaussian process model is proposed. This strategy improves the generalization ability of the fault diagnosis model while overcoming the shortcomings of traditional diagnostic model architecture adjustment, which relies heavily on subjective experience and is time-consuming. It provides an efficient and intelligent diagnostic model architecture adjustment solution. A high-frequency wireless vibration sensor integrated with edge computing is designed to break through the limitation of traditional vibration sensors that are limited to vibration signal collection. The sensor realizes functions such as high-frequency signal collection, signal processing, fault diagnosis and result display, integrates the complete function of equipment fault diagnosis, and provides a highly integrated single solution for real-time fault diagnosis of equipment.

2. According to the time-frequency domain feature extraction of vibration signals based on prior physical knowledge described in claim 1, the present invention mainly considers the time domain and frequency domain features contained in the bearing vibration signal. In terms of time domain features, the present invention mainly involves three features: skewness, kurtosis and root mean square error. The original vibration signal is processed using the Hilbert transform to obtain the corresponding envelope vibration signal. The vibration signal x(t) is coupled with the sample H(t) of the Hilbert transform signal at time t, and then the corresponding envelope signal a(t) is calculated. The envelope signal a(t) is then calculated using the fast Fourier transform method to obtain the envelope spectrum F(ω). Through the spectrum F(ω), select the specified range [f FCF -f se ,f FCF +f se ] to calculate the fault spectrum energy within the predefined sideband. In the present invention, the fault spectrum energy is defined as: Where, SE out , SE in and SE ball represent the fault spectrum energy of the outer raceway defect, inner raceway defect and ball defect, respectively, and f se Represents half of the passband range.

3. In the construction of a deep learning network that integrates prior physical knowledge and data as described in claim 1, it is characterized in that three dense layers are used in the proposed deep learning model to map the extracted feature vectors to the health status of the system. The data-driven features obtained from the dense layers are obtained through the learning process of the network (involving weight adjustment and nonlinear transformation), which provides a useful discriminative representation of the original data. A feature fusion layer is defined to fuse the data-driven features MD and the knowledge-based features MK extracted from the first dense layer of the CNN network, and then the new concatenated feature vector Xmerg will be used as a fused feature map for fault classification.

4. In the adaptive network parameter optimization based on the constrained Gaussian process model described in claim 1, the generalization ability of the knowledge-based deep network varies with the architecture parameters and the learnable parameters, characterized in that its structural design is represented as an optimization problem, σ represents a metric based on the generalization ability, which is used to accurately quantify the generalization ability of the deep network, and considering the constraints imposed by σ, an optimization method based on the constrained Gaussian process (CGP) is used as the optimizer. C(θ v ) represents the linear operation inequality constraint on θ v All points in satisfy, specify a target probability p target ∈[0,1), to find the set θ v When θ v When the constraint is satisfied at all virtual observation positions, the constraint is * The probability that ∈Ω is at least p target , which can be expressed as: p c (θ)=P(a(θ)-v<ξ(θ,θ v )<b(θ)+v) Wherein, when θ v = φ, ξ(θ, θ v ) = ζf(θ * )|G e Otherwise, ξ(θ, θ v ) = ζf(θ * )|G e , C 5. According to the method for building a deep network learning model described in claim 3, it is characterized by using a dropout layer, which randomly discards part of the information in the training data to prevent overfitting of the deep network. The dropout rate is usually set between 0 and 1 to improve the generalization ability of the model by reducing the coordinated adjustment of the training data. Subsequently, Softmax is used as the activation function for bearing diagnosis. If the label set contains g categories, the probability of the mth category appearing can be expressed as: Where w is a weight vector. For multi-classification, a separate loss is calculated for each class label and the results are summed: In the formula, Label i is the true label indicator for the i-th class, Prob i is the Softmax probability of the i-th class.

6. In the method according to claim 4, the optimization method based on the constrained Gaussian process (CGP) is used as the optimizer, characterized in that A non-negative fixed number υ is used to ensure that p can be increased by observations with additive noise. c , in the predefined design space, Latin hypercube sampling was used to generate m random network architectures The corresponding deep network will be trained to obtain the corresponding generalization ability value G KIDN =[G e1 ,G e2 ,...,G em ] T Based on an initial set of samples Using the maximum likelihood function to construct a constrained Gaussian process model, we can express it as: L(η)=p(G e ,C|η)=p(G e |η)p(C|G e ,or) Wherein, the mean μ(θ|η) and covariance function K(θ,θ'|η) of the Gaussian process prior are both assumed to depend on η. After constructing the initial constrained Gaussian process model, the expected improvement index is used to quantify the potential contribution of new sample points to the current model. Subsequently, the sample point with the largest expected improvement value is selected to update the constrained Gaussian process model. The update process will be repeated iteratively until the maximum expected improvement value is lower than the critical threshold I c Here, I c =|G emin | / 100 means 1% of the absolute value of the current best global minimum estimate.

7. The high reliability fault diagnosis method based on physical-data joint optimization drive according to claim 1 is characterized in that Fully combine the time domain and frequency domain characteristics of vibration signals, and use them as prior physical knowledge to integrate key health information in condition monitoring data to improve the accuracy of the diagnostic model; It effectively overcomes the shortcomings of the traditional model architecture parameter adjustment process, which relies heavily on subjective experience and is time-consuming, and provides an efficient and intelligent solution for fault diagnosis of complex equipment. Vibration signal acquisition, signal processing and fault diagnosis, result display and wireless transmission are integrated into one, and real-time fault diagnosis is performed directly on the front end of the equipment through edge computing technology, which significantly improves the response speed and data processing efficiency, and provides real-time technical guarantee for the accurate diagnosis and effective prevention of early faults.

8. A complete device diagnostic method, characterized in that: include: Vibration data collection, using IEPE technology to collect high-frequency vibration data; Signal processing and fault diagnosis: a fault diagnosis method based on knowledge-data joint optimization is proposed. This method fully integrates the prior physical knowledge such as the time domain and frequency domain characteristics of the vibration signal, as well as the key health information extracted from the original monitoring data. It is combined with an adaptive optimization algorithm based on a constrained Gaussian process model to quickly search for the optimal architecture parameters of the diagnosis model while improving the generalization ability of the diagnosis, ensuring the efficiency and reliability of equipment status assessment. Edge computing: AI computing chips are embedded in sensors to run highly reliable fault diagnosis algorithms driven by physical-data joint optimization, which can calculate and evaluate the operating status of equipment in real time; The results show that the built-in sensor LCD screen is used to display the calculation results and provide intuitive equipment status. The integrated wireless communication module wirelessly transmits the original vibration data and calculation results to the server to achieve remote monitoring and analysis of the data.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the high-reliability fault diagnosis method based on physical-data joint optimization drive as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the high-reliability fault diagnosis method based on physical-data joint optimization drive as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Fault diagnosis method based on rotating machinery

    CN112699722A

  • CNN feature fusion-based small sample transfer learning fault diagnosis method and system, computer and storage medium

    CN116702076A

  • Hydroelectric generating set vibration information fusion fault diagnosis method based on deep learning

    CN118940213A

Cited By

  • Illusion suppression method and device for power transmission line fault diagnosis model

    CN122364356A