A bearing fault prediction method based on multi-attention fuzzy cognitive map
By constructing a deep learning model based on multi-attention fuzzy cognitive graphs and combining STFCM, LSTM and residual connections, the shortcomings of existing bearing fault prediction technologies in terms of accuracy and comprehensiveness are solved, enabling early detection and automated monitoring of bearing faults and reducing equipment maintenance costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CAPITAL NORMAL UNIVERSITY
- Filing Date
- 2023-09-27
- Publication Date
- 2026-04-28
AI Technical Summary
Existing bearing failure prediction technologies suffer from low accuracy and comprehensiveness, particularly in handling non-stationary time series and sudden failures. Furthermore, existing methods are highly dependent on operator experience and are costly and complex.
We employ a multi-attention fuzzy cognitive graph-based approach, combining STFCM, LSTM, and residual connections. By upscaling data through kernel mapping and introducing self-attention and temporal attention mechanisms for multi-graph fusion, we construct a deep learning model to achieve accurate prediction of bearing faults.
It improves the accuracy and comprehensiveness of bearing failure prediction, enables early detection of minute changes, reduces equipment downtime and maintenance costs, is applicable to various working scenarios, and achieves automated monitoring and early warning.
Smart Images

Figure CN117390483B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bearing failure prediction technology. In particular, it relates to a bearing failure prediction method based on a multi-attention fuzzy cognitive graph. Background Technology
[0002] Bearings are common and crucial components in mechanical equipment, widely used in industrial production and transportation. However, due to long-term workloads and harsh working environments, bearings are prone to failure, leading to equipment downtime, production interruptions, and high maintenance costs. Bearings are key components commonly used in industrial equipment, supporting and rotating shafts and bearing mechanical loads. However, long-term operation and harsh working conditions can cause bearing wear, damage, or failure, thus affecting the performance and lifespan of the entire equipment. To achieve bearing health monitoring and fault prediction, reduce maintenance costs, and avoid unexpected downtime, bearing fault prediction methods have emerged. Bearing fault prediction technology has become particularly important for improving equipment reliability and reducing maintenance costs. Bearing fault prediction refers to using sensors to monitor bearing condition parameters and then analyzing, processing, and modeling them to predict bearing life and the likelihood of failure. Its goal is to achieve real-time monitoring and accurate prediction of bearing health, allowing for early maintenance or replacement measures to avoid unexpected downtime and production losses. Modern bearing fault prediction methods utilize sensors and advanced data analysis techniques to monitor bearing condition and identify potential fault signs. These methods allow equipment operators to take timely maintenance measures and prevent failures that could lead to serious damage.
[0003] The existing bearing failure prediction technologies mainly include:
[0004] Vibration analysis: This is one of the most commonly used bearing fault prediction techniques. By monitoring bearing vibration signals, fault frequencies and vibration characteristics can be analyzed to determine the bearing's operating condition. Acoustic analysis: Based on the acoustic signals generated during bearing operation, abnormal vibrations and noise are identified by analyzing the acoustic spectrum to predict the bearing's health. This method is more sensitive to early fault detection but is also sensitive to environmental noise, which may affect its accuracy.
[0005] Infrared thermography: This technique uses infrared thermography to monitor temperature changes in bearings. Faulty bearings typically generate abnormal heat, which can be detected promptly using infrared thermography. This method allows for remote, real-time monitoring and is suitable for high-temperature or inaccessible environments.
[0006] Oil analysis: By periodically collecting bearing lubricating oil samples and detecting changes in metal particles, impurities, and chemical composition, the health of the bearing can be assessed. This method can predict early wear and contamination, but it requires high-quality oil.
[0007] Machine learning methods: Machine learning algorithms are used to model and analyze bearing vibration, sound waves, and other monitoring data. This method can automatically learn bearing failure modes and has good predictive accuracy.
[0008] Fuzzy Cognitive Maps (FCMs) are a soft computation method with strong knowledge reasoning and representation capabilities, providing a new approach for interpretable time series forecasting. FCMs can be used for long-term time series forecasting, and higher-order fuzzy cognitive maps can also be constructed for time series forecasting.
[0009] Deep learning-based methods only learn local spatiotemporal correlations and cannot comprehensively learn the spatiotemporal correlations of time series. Multi-graph fusion research has achieved good results in predicting traffic flow datasets. However, these studies are only applied in the transportation field and not in other fields, limiting their applicability. Although they consider the long-term trends of time series, they do not comprehensively model the entire spatiotemporal relationship. Another problem with deep learning is that as network depth increases, gradient explosion and vanishing problems occur during backpropagation. Improvements in FCM are relatively successful in predicting stationary short-term time series, but their accuracy is poor in predicting non-stationary long-term series.
[0010] Many existing technologies rely on a single monitoring indicator, such as vibration or sound waves. The limitations of this single indicator can lead to incomplete predictions and difficulty in accurately assessing the complexity of bearing failures. Some methods require the regular collection of large amounts of monitoring data and complex data processing and analysis, increasing the cost and complexity of equipment maintenance. Certain technologies demand a high level of operator experience and skill, requiring experienced professionals for fault diagnosis, lacking universality and ease of use. Some methods are only suitable for predicting gradually developing bearing failures, performing poorly on sudden failures (such as those caused by external impacts). Some solutions present challenges in terms of energy consumption and real-time performance, making them unsuitable for resource-constrained scenarios or those requiring real-time prediction.
[0011] In real-world applications, existing deep learning methods suffer from low accuracy in predicting non-stationary time series. While existing bearing fault prediction technologies offer advantages in various scenarios, they also have limitations. Therefore, there is an urgent need to comprehensively utilize multiple monitoring and technical methods, combined with intelligent algorithms such as machine learning, to improve the accuracy and comprehensiveness of predictions. Furthermore, it is crucial to simplify data collection and processing in practical applications to reduce costs and barriers to entry, making the predictions more widespread and practical. Summary of the Invention
[0012] This invention provides a bearing fault prediction method based on multi-attention fuzzy cognitive graphs to address the limitations, low accuracy, and low comprehensiveness of existing technologies.
[0013] To achieve the above objectives, the present invention provides a bearing fault prediction method based on multi-attention fuzzy cognitive maps, comprising the following steps: presetting a first AFCM (Attention fuzzy cognitive maps), a second AFCM, and a third AFCM, wherein the first AFCM, the second AFCM, and the third AFCM are all composed of STFCM (spatio-temporal fuzzy cognitive maps), LSTM (Long Short Term Mermory network), and residual connections;
[0014] Obtain the normalized bearing target feature data;
[0015] The target feature data and its associated features are used as input data for the third AFCM to obtain the output result of the third AFCM;
[0016] The target feature data is up-dimensional by kernel mapping to obtain up-dimensional target data. The up-dimensional target data is used as input data for the first AFCM and the second AFCM to obtain the output results of the first AFCM and the second AFCM.
[0017] The outputs of the first AFCM, the second AFCM, and the third AFCM are combined with a preset time attention mechanism and a residual structure to perform multi-graph fusion to obtain a fusion result.
[0018] The fusion result is input into a preset 3-layer fully connected neural network to obtain the final output as the prediction result of the future target.
[0019] Preferably, the step of upscaling the target feature data to obtain the upscaled target data through kernel mapping includes linear functions and nonlinear functions. The linear function is: y = kx, where k is a coefficient; the non-target data linear function is: y = x i +1 , and Where l represents the parameter of the kurtosis of the mapping function, i = 1, 2, 3, ..., l, x is the target feature data, and y is the target data for dimensionality improvement.
[0020] Preferably, the preset first AFCM, second AFCM, and third AFCM specifically include:
[0021] A self-attention-based STFCM is pre-trained by inputting the input data into the STFCM, which then uses the self-attention mechanism to learn the correlations between nodes in the graph.
[0022] Calculate Q, K, and V, where Q = W q P(t)K=W k P(t)V=W v P(t)
[0023] Calculate the correlation between every two input vectors using the obtained Q and K:
[0024] A = K T Q
[0025] Perform a fuzzy mapping on A: P′(t)=VA′
[0026] Among them W q W k W v Here, d is the learning parameter, P(t) is the input, P′(t) is the output, f is the transfer function that maps the activation value to a specific range, Q is the content to be retrieved, K is the index, and V is the value to be retrieved. After learning the node relationship through the self-attention mechanism, the output node state value is updated to P′(t).
[0027] The updated node state value P′(t) is input into the improved attention mechanism; a time layer is expanded into multiple time layers to capture the relationship between the current node state value and the state value at previous time steps, where;
[0028]
[0029]
[0030] Among them W f1 W f2 , It is the learning parameter, B f1 B f2 For the bias term, α t It is a weight value. The hidden state and cell state of the transition-gated LSTM are represented by f, which is a transfer function that maps activation values to a specific range; after learning the time step information of the node, the node's state value is updated by z. t
[0031] The input data and updated node state values are then fed into the residual connection: The input data and updated node state values are then fed into the residual connection: in B is the weight of the linear transformation. qHere, z is the bias term, P(t) is the node state value, P″(t) is the node state after the residual structure, and z is the node state after the residual structure. t The node status value is updated after STFCM.
[0032] Preferably, in the first AFCM, second AFCM, and third AFCM, the node state value P″(t) after passing through the residual structure is input into the LSTM, and the specific information is controlled by three gates: the input gate that controls the input information. Forget gates that control forgotten information t c Output gate that retains information in:
[0033]
[0034]
[0035] Among them W c U c b c The parameters to be learned for all control gates used. P″(t-1) is the previously hidden state, σ(·) is the sigmoid function, tanh(·) is the hyperbolic tangent function, and ⊙ is element-wise multiplication.
[0036] Preferably, the step of fusing the outputs of the first AFCM, the second AFCM, and the third AFCM using a time attention mechanism and a residual structure to obtain a fusion result specifically includes:
[0037]
[0038]
[0039] in, This is the output of the first AFCM. This is the output of the second AFCM. For the output of the third AFCM, W m1 The parameters learned by the first linear transformation, W m2 These are the parameters learned by the second linear transformation. It is the parameter for learning, B m1 It is the bias term of the first linear transformation, B m2 It is the bias term of the second linear transformation, β tHere, P″′(t) represents the weight value, P″′(t) represents the node state value, f is the transfer function that maps the activation value to a specific range, T is the time range of the temporal attention mechanism (range 1-15), and N is the number of nodes. The updated node state P″′(t) is combined with the bearing target feature data P(t) of the initial node state to obtain the fusion result after the residual structure. in, B is the weight of the linear transformation. k This is a bias term.
[0040] Preferably, before acquiring the normalized bearing target feature data, the method further includes: training a MAFCM model using the Adam optimizer and stochastic gradient descent. The MAFCM model consists of the preset first AFCM, second AFCM, and third AFCM, a preset time attention mechanism, a preset residual structure, and the preset three-layer fully connected neural network. Specifically:
[0041]
[0042] Where Θ represents all parameters, For the prediction result of the future target, Pi is the input of the MAFCM model, M is the dimension of the kernel mapping dimensionality increase, and the input bearing feature training data includes one or more of the following: normal samples, outer ring damage samples, inner ring damage samples and rolling element damage samples; the damage point is simulated at least three positions relative to the bearing load area; the three data variables input to the prediction model are: drive end acceleration data, rotating end acceleration data and base acceleration data.
[0043] Secondly, the present invention also relates to a bearing fault prediction device based on a multi-attention fuzzy cognitive map, comprising:
[0044] The AFCM preset module is used to preset the first AFCM, the second AFCM and the third AFCM, wherein the first AFCM, the second AFCM and the third AFCM are all composed of STFCM, LSTM and residual connection;
[0045] The normalization module is used to obtain the normalized bearing target feature data;
[0046] The third AFCM input / output module is used to take the target feature data and its associated features as input data of the third AFCM and obtain the output result of the third AFCM.
[0047] The dimension-upgrading target data input / output module is used to upgrade the target feature data to obtain dimension-upgrading target data P1 and P2 through kernel mapping. P1 and P2 are used as input data for the first AFCM and the second AFCM, respectively, to obtain the output results of the first AFCM and the second AFCM.
[0048] The multi-image fusion module is used to fuse the outputs of the first AFCM, the second AFCM, and the third AFCM using a preset temporal attention mechanism and residual structure to obtain a fusion result; and
[0049] The neural network prediction module is used to input the fusion result into a preset 3-layer fully connected neural network to obtain the final output as the prediction result of the future target.
[0050] Preferably, the AFCM prediction module is specifically used for: pre-training an attention-based STFCM, inputting the input data into the STFCM, and using the self-attention mechanism to learn the correlation between nodes in the graph:
[0051] Calculate Q, K, and V, where Q = W q P(t)K=W k P(t)V=W v P(t)
[0052] Calculate the correlation between every two input vectors using the obtained Q and K:
[0053] A = K T Q
[0054] Perform a fuzzy mapping on A: P′(t)=VA′, where W q W is a parameter for learning the retrieved content. k W is the parameter for index learning. v Here are the parameters learned for the value to be retrieved, d is a constant, P(t) is the input, P′(t) is the output, f is the transfer function that maps the activation value to a specific range, Q is the content to be retrieved, K is the index, and V is the value to be retrieved. After learning the node relationship through the self-attention mechanism, the output node state value is updated to P′(t).
[0055] The updated node state value P′(t) is input into the improved attention mechanism; a time layer is expanded into multiple time layers to capture the relationship between the current state value of the node and the state value at previous time steps, as shown in the following formula;
[0056]
[0057]
[0058] Among them W f1 W f2 , It is the learning parameter, B f1 B f2 For the bias term, α t It is a weight value. The hidden state and cell state of the transition-gated LSTM are represented by f, which is a transfer function that maps activation values to a specific range; after learning the time step information of the node, the node's state value is updated by z. t ;
[0059] The input data and updated node state values are then fed into the residual connection: the input data and updated node state values are then fed into the residual connection, wherein... B is the weight of the linear transformation. q Here, z is the bias term, P(t) is the node state value, P″(t) is the node state after the residual structure, and z is the node state after the residual structure. t The node status value is updated after STFCM.
[0060] Preferably, the multi-image fusion module is specifically used for:
[0061]
[0062]
[0063] in, This is the output of the first AFCM. This is the output of the second AFCM. The output of the third AFCM, v t It is the result of the fusion of time attention mechanisms, W m1 The parameters learned by the first linear transformation, W m2 These are the parameters learned by the second linear transformation. It is the parameter for learning, B m1 B m2 All are bias terms, β t Here, P″′(t) represents the weight value, P″′(t) represents the node state value, f is the transfer function that maps the activation value to a specific range, T is the time range of the temporal attention mechanism (range 1-15), and N is the number of nodes. The updated node state P″′(t) is combined with the bearing target feature data P(t) of the initial node state to obtain the fusion result after the residual structure. in, B is the weight of the linear transformation. k This is a bias term.
[0064] Thirdly, the present invention also relates to a computer-readable storage medium storing instructions that, when executed, perform any of the above-described bearing fault prediction methods based on multi-attention fuzzy cognitive graphs.
[0065] The present invention relates to a bearing fault prediction method based on a multi-attention fuzzy cognitive map, which has the following advantages compared with the prior art:
[0066] Combining the framework of Fiber Optic Modeling (FCM) with the advantages of deep learning, a novel time series prediction model, MAFCM, is constructed. This model fuses multiple AFCMs and introduces self-attention and improved attention mechanisms to construct STFCM, learning the spatiotemporal relationship and non-stationary characteristics of the data, thus improving the model's prediction accuracy. To enhance the model's generality, kernel mapping is used to upgrade the target data as input to the AFCM. We introduce LSTM into the FCM framework to improve long-term prediction performance. For multi-graph fusion, a temporal attention mechanism is used to fuse the outputs of multiple AFCMs to learn the temporal dependencies of the time series. To address the gradient explosion and vanishing problems in the network, residual connections are added, achieving good results.
[0067] Deep learning-based bearing fault prediction methods can learn complex fault modes from multiple data sources, improving prediction accuracy. Compared to traditional methods, they are better able to capture subtle changes in bearing faults, enabling earlier fault detection.
[0068] The real-time predictive capabilities of deep learning models enable automated bearing fault monitoring, eliminating the need for manual intervention. Early warning systems can notify maintenance personnel immediately, helping to reduce losses and repair costs.
[0069] Deep learning models can integrate various monitoring data, resulting in more comprehensive and accurate predictions. Furthermore, this method is not limited by specific environments and is applicable to various types of bearings and operating scenarios. By predicting bearing failures in advance, the escalation of failures can be prevented, reducing equipment downtime and maintenance costs, and extending bearing life. Attached Figure Description
[0070] Figure 1 The following is a flowchart of a bearing fault prediction method based on a multi-attention fuzzy cognitive graph according to Embodiment 1 of the present invention. Figure 1 ;
[0071] Figure 2 These are MAFCM framework diagrams from Embodiments 1 and 2 of the present invention;
[0072] Figure 3 This is a framework diagram of the first AFCM, second AFCM, or third AFCM in Embodiment 1 and Embodiment 2 of the present invention;
[0073] Figure 4 This refers to the acceleration data of the driving end in the example of Embodiment 1 of the present invention;
[0074] Figure 5 This refers to the fan-end acceleration data in the example of Embodiment 1 of the present invention;
[0075] Figure 6 This refers to the base acceleration data in the example of Embodiment 1 of the present invention;
[0076] Figure 7 This is a schematic diagram of a bearing fault prediction device based on a multi-attention fuzzy cognitive graph according to Embodiment 2 of the present invention. Detailed Implementation
[0077] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0078] Example 1
[0079] A bearing fault prediction method based on multi-attention fuzzy cognitive graphs; please refer to [link / reference]. Figure 1-6 This includes the following steps:
[0080] S10: Preset first AFCM, second AFCM and third AFCM.
[0081] Time series often undergo sudden evolution. To address this, AFCM is proposed to learn the spatiotemporal characteristics of time series and simultaneously capture their non-stationary properties and long-term dependencies. This solves the problem of sudden events affecting prediction accuracy in different fields. The first, second, and third AFCMs are all composed of three parts: STFCM, LSTM, and residual connections.
[0082] Among them, STFCM: In order to better capture the spatiotemporal relationship and non-stationary characteristics of multivariate time series, this application pre-trains an attention-based STFCM. The input data is input into STFCM, and the correlation between nodes in the graph is learned by using the self-attention mechanism. The specific formulas (6), (7), (8), (9), (10), and (11) are as follows:
[0083] Calculate Q, K, and V:
[0084] Q = W q P(t) (6)
[0085] K = W k P(t) (7)
[0086] V = W v P(t) (8)
[0087] Calculate the correlation between every two input vectors using the obtained Q and K:
[0088] A = K T Q (9)
[0089] Perform a fuzzy mapping on A:
[0090]
[0091] P′(t)=VA′ (11)
[0092] Among them W q W is a parameter for learning the retrieved content. k W is the parameter for index learning. v Here are the parameters learned for the value to be retrieved, d is a constant, P(t) is the input, P′(t) is the output, f is the transfer function that maps the activation value to a specific range, Q is the content to be retrieved, K is the index, and V is the value to be retrieved. After learning the node relationship through the self-attention mechanism, the output node state value is updated to P′(t).
[0093] Then the updated node state value P′(t) is input into the improved attention mechanism as shown in Equation (15). In order to enhance the ability to predict long-term series, the current state value of a node is related to the previous time of this node. One time layer is expanded into multiple time layers to capture the relationship between the current state value of the node and the state value at the previous time, as shown in Equations (12), (13), and (14).
[0094]
[0095]
[0096]
[0097] Among them W f1 W f2 , It is the learning parameter, B f1 B f2 For the bias term, α t It is a weight value. The hidden state and cell state of the transition-gated LSTM are represented by f, which is a transfer function that maps activation values to a specific range; after learning the time step information of the node, the node's state value is updated by z. t .
[0098] Residual Connections: As the network depth increases in deep learning, gradient explosion and vanishing can occur during backpropagation. Based on the idea of residual networks, this application introduces residual connections into AFCM by referencing the hidden states of residual networks. To prevent the increase in network structure and the resulting computational complexity, residual information is added through linear transformation, which also accelerates the convergence speed. Inputting the input data and updated node state values into the residual connections: The input data and updated node state values are input into the residual connections, as shown in formula (15).
[0099]
[0100] in B is the weight of the linear transformation. q Here, z is the bias term, P(t) is the node state value, P″(t) is the node state after the residual structure, and z is the node state after the residual structure. t The node state values are updated for STFCM. Residual connections use a linear transformation to make the loss function approximate a state mapping, which does not increase training error and is more sensitive to sudden changes in data caused by unexpected situations.
[0101] LSTM: In the first, second, and third AFCMs, the node state values P″(t) after passing through the residual structure are input into the LSTM. The specific information is controlled by three gates: the input gate that controls the input information. Forget gates that control forgotten information t c Output gate that retains information The specific formulas (16), (17), and (18) are as follows:
[0102]
[0103]
[0104]
[0105] Among them W c U c b c The parameters to be learned for all control gates used. P″(t-1) represents the previous hidden state, σ(·) is the sigmoid function, tanh(·) is the hyperbolic tangent function, and ⊙ represents element-wise multiplication. LSTM is used to capture the contextual information of global nodes, thereby enhancing the predictive ability of AFCM for long-term series.
[0106] S20: Obtain the normalized bearing target feature data.
[0107] Specifically, the target characteristic data of the bearing is P = [P 1,t ,P 2,t ,P 3,t ,…,P N,t ], where N is the dimension of the target feature data and t is the time sequence number.
[0108] S30: Use the target feature data and its associated features as input data for the third AFCM to obtain the output result of the third AFCM.
[0109] S40: Upscale the target feature data by kernel mapping to obtain upscaled target data. Use the upscaled target data as input data for the first AFCM and the second AFCM to obtain the output results of the first AFCM and the second AFCM.
[0110] To improve the universality of the proposed model, kernel mapping is used to increase the dimensionality of the target features, extending it to other graphical models. This application constructs a set of mapping functions for increasing the dimensionality of target features, elevating them to a higher-dimensional space, called kernel mapping. The number of target features for the bearing. P = [P 1,t ,P 2,t ,P 3,t ,…,P N,t After mapping, the data is upgraded to higher dimensions: P 1 =P 2 =[P 1,t ,P 2,t ,P 3,t ,…,P M,t ], where M is the dimension of the feature data after dimensionality increase, t is the time sequence number, and the dimensionality increase data is used as the input data of the first AFCM and the second AFCM.
[0111] A linear function was used, as shown in formula (1).
[0112] y = kx(1), where k is a coefficient.
[0113] The nonlinear functions of 4xl are shown in formulas (2), (3), (4), and (5).
[0114] y = x i+1 (2)
[0115]
[0116]
[0117]
[0118] Where l represents the parameter of the kurtosis of the mapping function, i = 1, 2, 3, ..., l, x is the target feature data, and y is the target data for dimensionality improvement.
[0119] The linear function used in this paper preserves the distribution characteristics of the original time series, while the nonlinear function captures the implicit features present in the original time series. Therefore, kernel mapping can capture the implicit patterns of the original data, increasing the difficulty of predicting non-stationary time series, making it suitable for complex non-stationary time series, and improving the accuracy of model predictions.
[0120] S50: The outputs of the first AFCM, the second AFCM, and the third AFCM are combined with the residual structure using a preset time attention mechanism to perform multi-graph fusion to obtain the fusion result.
[0121] This application uses three AFCMs, resulting in three outputs: Simply weighting and fusing the outputs of the three graphs without considering the relationship between the current node and the state values of nodes at previous time steps leads to reduced prediction accuracy. Therefore, we use a temporal attention mechanism for multi-graph fusion to enhance the model's predictive power.
[0122] The previously hidden states of the LSTM are concatenated and then input into the temporal attention mechanism, as shown in the figure. The specific formulas (19), (20), and (21) are as follows:
[0123]
[0124]
[0125]
[0126] in, This is the output of the first AFCM. This is the output of the second AFCM. For the output of the third AFCM, W m1 The parameters learned by the first linear transformation, W m2 These are the parameters learned by the second linear transformation. It is the parameter for learning, B m1 It is the bias term of the first linear transformation, B m2 It is the bias term of the second linear transformation, β t Here, P″′(t) is the weight value, P″′(t) is the node state value, f is the transfer function that maps the activation value to a specific range, T is the time range of the time attention mechanism, ranging from 1 to 15, and N is the number of nodes. To further improve the accuracy of the model, we introduce residual connections again, combining the updated node state P″′(t) with the initial node state P(t), as shown in the following formula (22):
[0127]
[0128] in B is the weight of the linear transformation. k Here, P(t) is the bias term, and P(t) is the node state value. This represents the node state after the residual structure has been applied.
[0129] The S60 inputs the fusion result into a preset 3-layer fully connected neural network to obtain the final output as the prediction result of the future target.
[0130] Will The input is fed into a 3-layer fully connected neural network to obtain the final output. MAFCM employs an encoder-decoder architecture. The historical state value P(t) of a node is first input into the encoder, and after passing through three AFCMs, three outputs are obtained. The three outputs are concatenated with the node's historical state P(t) and then input into the decoder. The decoder mainly consists of a time attention mechanism and a residual structure. The final result is obtained by inputting the decoder into the output layer.
[0131] In some embodiments, before step S20, the method further includes: S21: training a MAFCM model using the Adam optimizer and stochastic gradient descent. The MAFCM model consists of a preset first AFCM, a second AFCM, and a third AFCM, a preset time attention mechanism, a preset residual structure, and a preset 3-layer fully connected neural network.
[0132] Specifically:
[0133] Where Θ represents all parameters, P is the prediction result of future goals. i The input to the MAFCM model is M, where M is the dimension of the kernel mapping. The input bearing feature training data includes one or more of the following: normal samples, outer ring damage samples, inner ring damage samples, and rolling element damage samples. The damage point is simulated at least three locations relative to the bearing load area. The three data variables input to the prediction model are: drive end acceleration data, rotating end acceleration data, and base acceleration data.
[0134] like Figure 4-6 As shown in the example experiment below, the prediction method of this application will be further explained and elaborated:
[0135] This invention utilizes a publicly available dataset from Western Reserve University, including four different data types: normal samples, outer ring damage samples, inner ring damage samples, and rolling element damage samples. The bearings under test support the motor shaft. The drive-end bearing is an SKF6205, with sampling frequencies of 12kHz and 48kHz; the fan-end bearing is an SKF6203, with a sampling frequency of 12kHz. The damaged bearings are single-point damages manufactured through electrical discharge machining. SKF bearings are used to simulate damage with diameters of 0.1778, 0.3556, and 0.5334 mm, while NTN bearings are used to simulate damage with diameters of 0.7112 and 1.016 mm. To consider the impact of different outer ring damage locations on the vibration response of the motor / bearing system, the damage point was simulated at different positions relative to the bearing load area, specifically at the 3 o'clock, 6 o'clock, and 12 o'clock positions. An accelerometer was placed above the bearing housings at both the fan end and drive end of the motor to collect the vibration acceleration signals of the faulty bearings. All vibration signals were acquired by a 16-channel data logger. Power and speed information were measured by a torque sensor / decoder. Therefore, the three data variables input to the prediction model are: DE drive-end acceleration data, FE fan-end acceleration data, and BA base acceleration data (data under normal conditions).
[0136] The following table shows a comparison with other methods:
[0137]
[0138]
[0139] Among them, MAE (mean absolute error) and RMSE (root mean square error) are commonly used evaluation metrics for prediction problems; HA is the historical average model; AR is the autoregressive model; ARIMA is the autoregressive combined moving average model; ES is the exponential smoothing model; VAR is the vector autoregressive model; MLP is the multilayer perceptron; FCM is the basic fuzzy cognitive graph; WFCM is wavelet transform and fuzzy cognitive graph; and BFCM is Bayesian ridge regression and fuzzy cognitive graph. T is the prediction iteration number, x(t) and The actual and predicted values for bearing failure prediction in the i-th cycle are calculated as follows:
[0140]
[0141]
[0142] As can be seen from the table, the MAE and RMSE of this invention are both minimal, proving that the method proposed in this invention is effective.
[0143] Example 2
[0144] like Figure 2-7 As shown, a bearing fault prediction device based on multi-attention fuzzy cognitive graph is presented. The bearing fault prediction device in this embodiment is applied in the bearing fault detection process. The device can be implemented as a control center with a central processing unit, such as a personal computer, host computer, server or other electronic device.
[0145] In the embodiments, such as Figure 3 As shown, the bearing fault prediction device includes: AFCM preset module 71, normalization module 72, third AFCM input / output module 73, upgraded target data input / output module 74, multi-image fusion module 75, and neural network prediction module 76.
[0146] AFCM preset module 71 is used to preset the first AFCM, the second AFCM and the third AFCM, wherein the first AFCM, the second AFCM and the third AFCM are all composed of STFCM, LSTM and residual connection;
[0147] Normalization module 72 is used to obtain the normalized bearing target feature data;
[0148] The third AFCM input / output module 73 is used to take the target feature data and its associated features as input data of the third AFCM and obtain the output result of the third AFCM.
[0149] The upscaling target data input / output module 74 is used to upscale the target feature data through kernel mapping to obtain upscaling target data P1 and P2. P1 and P2 are used as input data for the first AFCM and the second AFCM, respectively, to obtain the output results of the first AFCM and the second AFCM.
[0150] The multi-image fusion module 75 is used to fuse the outputs of the first AFCM, the second AFCM, and the third AFCM using a preset temporal attention mechanism and residual structure to obtain a fusion result; and
[0151] The neural network prediction module 76 is used to input the fusion result into a preset 3-layer fully connected neural network to obtain the final output as the prediction result of the future target.
[0152] In this embodiment, the AFCM pre-training module is specifically used for: pre-training an attention-based STFCM, inputting input data into STFCM, and using the self-attention mechanism to learn the correlation between nodes in the graph.
[0153] Calculate Q, K, and V, where Q = W q P(t) K=W kP(t) V=W v P(t)
[0154] Calculate the correlation between every two input vectors using the obtained Q and K:
[0155] A = K T Q
[0156] Perform a fuzzy mapping on A: P′(t)=VA′
[0157] Among them W q W is a parameter for learning the retrieved content. k W is the parameter for index learning. v Here are the parameters learned for the value to be retrieved, d is a constant, P(t) is the input, P′(t) is the output, f is the transfer function that maps the activation value to a specific range, Q is the content to be retrieved, K is the index, and V is the value to be retrieved. After learning the node relationship through the self-attention mechanism, the output node state value is updated to P′(t).
[0158] The updated node state value P′(t) is input into the improved attention mechanism; a time layer is expanded into multiple time layers to capture the relationship between the current state value of the node and the state value at previous time steps, as shown in the following formula;
[0159]
[0160]
[0161] Among them W f1 W f2 , It is the learning parameter, B f1 B f2 For the bias term, α t It is a weight value. The hidden state and cell state of the transition-gated LSTM are represented by f, which is a transfer function that maps activation values to a specific range; after learning the time step information of the node, the node's state value is updated by z. t
[0162] Input the input data and updated node state values into the residual join: Input the input data and updated node state values into the residual join: in B is the weight of the linear transformation. q Here, z is the bias term, P(t) is the node state value, P″(t) is the node state after the residual structure, and z is the node state after the residual structure. t The node status value is updated after STFCM.
[0163] Multi-image fusion module 75 is specifically used for:
[0164]
[0165]
[0166] in, This is the output of the first AFCM. This is the output of the second AFCM. For the output of the third AFCM, W m1 The parameters learned by the first linear transformation, W m2 These are the parameters learned by the second linear transformation. It is the parameter for learning, B m1 B m2 All are bias terms, β t Here, P″′(t) represents the weight value, P″′(t) represents the node state value, and f is the transfer function that maps the activation value to a specific range. The updated node state P″′(t) is combined with the bearing target feature data P(t) of the initial node state to obtain the fusion result after the residual structure. in, B is the weight of the linear transformation. k This is a bias term.
[0167] The bearing fault prediction device based on multi-attention fuzzy cognitive graph in this embodiment is the same as the bearing fault prediction method based on multi-attention fuzzy cognitive graph described in Embodiment 1 in terms of implementation process, method and effect, and will not be repeated here.
[0168] Example 3
[0169] This invention relates to a computer-readable storage medium storing instructions that, when executed, perform a bearing fault prediction method based on a multi-attention fuzzy cognitive graph according to Embodiment 1. The execution process and effects are the same as those described in Embodiment 1, and will not be repeated here.
[0170] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0171] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A bearing fault prediction method based on multi-attention fuzzy cognitive graphs, characterized in that, Includes the following steps: A first AFCM, a second AFCM, and a third AFCM are preset, wherein the first AFCM, the second AFCM, and the third AFCM are all composed of STFCM, LSTM, and residual connections; Obtain the normalized bearing target feature data; The target feature data and its associated features are used as input data for the third AFCM to obtain the output result of the third AFCM; The target feature data is up-dimensionalized by kernel mapping to obtain up-dimensional target data. The up-dimensional target data is used as input data for the first AFCM and the second AFCM to obtain the output results of the first AFCM and the second AFCM. The outputs of the first AFCM, the second AFCM, and the third AFCM are fused using a preset time attention mechanism and a residual structure to obtain a fusion result; specifically including: , , , in, This is the output of the first AFCM. This is the output of the second AFCM. This is the output of the third AFCM. These are the parameters learned by the first linear transformation. These are the parameters learned by the second linear transformation. These are the parameters for learning. It is the bias term of the first linear transformation. It is the bias term of the second linear transformation. It is a weight value. For node state values, It is a transfer function that maps activation values to a specific range. This refers to the time range of the time-based attention mechanism, which is 1-15. The number of nodes; update the node status. Bearing target feature data with initial node state By combining, we obtain the fusion result after the residual structure. : ,in, , They are all weights of a linear transformation. For bias terms; The fusion result is input into a preset 3-layer fully connected neural network to obtain the final output as the prediction result of the future target.
2. The bearing fault prediction method based on multi-attention fuzzy cognitive graphs according to claim 1, characterized in that, The upsizing of the target feature data through kernel mapping to obtain upsizing target data includes linear and nonlinear functions. The linear function is y = kx, where k is a coefficient. The non-target data linear function is: , , and Where l represents the parameter of the kurtosis of the mapping function, x represents the target feature data, and y represents the upgraded target data.
3. The bearing fault prediction method based on multi-attention fuzzy cognitive graphs according to claim 2, characterized in that, The preset first AFCM, second AFCM, and third AFCM specifically include: A self-attention-based STFCM is pre-trained by inputting the input data into the STFCM, which then uses the self-attention mechanism to learn the correlations between nodes in the graph. Calculate Q, K, and V. ; Use the obtained Q and K to calculate the correlation between every two input vectors: ; Perform a fuzzy mapping on A: , ,in , , All are learning parameters, where d is a constant. For input, For output, It is a transfer function that maps activation values to a specific range, where Q is the content to be retrieved, K is the index, and V is the value to be retrieved; after learning node relationships through the self-attention mechanism, the output node state value is updated to... ; Update the node state value Input is fed into an improved attention mechanism; a single time layer is expanded into multiple time layers to capture the relationship between the current state value of a node and its state values at previous time steps, where; , , , in , , These are all learning parameters. , All are bias terms. It is a weight value. This indicates the hidden state and cell state of the transition-gated LSTM. It is a transfer function that maps activation values to a specific range; after learning the time step information of a node, the node's state value is updated. ; The input data and updated node state values are then fed into the residual connection: The input data and updated node state values are then fed into the residual connection: ,in , They are all weights of a linear transformation. For bias terms, For node state values, The node state after the residual structure is applied. The node status value is updated after STFCM.
4. The bearing fault prediction method based on multi-attention fuzzy cognitive graphs according to claim 3, characterized in that, In the first AFCM, the second AFCM, and the third AFCM, the node state values after passing through the residual structure are... The information input to the LSTM is controlled by three gates: the input gate controls the input information. Forget gates that control forgotten information Output gate that retains information ,in: , , ; in , , These are all parameters that need to be learned in the control gates used. It was previously in a hidden state. It is in a hidden state. It represents the node state at the previous moment. It is information saved from the previous moment. It is the sigmoid function. It is the hyperbolic tangent function. It is element-wise multiplication.
5. The bearing fault prediction method based on multi-attention fuzzy cognitive graphs according to claim 1, characterized in that, Before obtaining the normalized bearing target feature data, the method further includes: training a MAFCM model using the Adam optimizer and stochastic gradient descent. The MAFCM model consists of the preset first AFCM, second AFCM, and third AFCM, a preset time attention mechanism, a preset residual structure, and the preset three-layer fully connected neural network. Specifically: ; in This represents all parameters. The prediction result for the stated future goal. The input to the MAFCM model is M, where M is the dimension of the kernel mapping upscaling. The input bearing feature training data includes one or more of the following: normal samples, outer ring damage samples, inner ring damage samples, and rolling element damage samples. The damage point is simulated at least three positions relative to the bearing load area. The three data variables input to the prediction model are: drive end acceleration data, rotating end acceleration data, and base acceleration data.
6. The bearing fault prediction device based on multi-attention fuzzy cognitive graph according to claim 1, characterized in that, include: The AFCM preset module is used to preset the first AFCM, the second AFCM and the third AFCM, wherein the first AFCM, the second AFCM and the third AFCM are all composed of STFCM, LSTM and residual connection; The normalization module is used to obtain the normalized bearing target feature data; The third AFCM input / output module is used to take the target feature data and its associated features as input data of the third AFCM and obtain the output result of the third AFCM. The dimension-upgrading target data input / output module is used to upgrade the target feature data to obtain dimension-upgrading target data P1 and P2 through kernel mapping. P1 and P2 are used as input data for the first AFCM and the second AFCM, respectively, to obtain the output results of the first AFCM and the second AFCM. The multi-image fusion module is used to fuse the outputs of the first AFCM, the second AFCM, and the third AFCM using a preset temporal attention mechanism and residual structure to obtain a fusion result; and The neural network prediction module is used to input the fusion result into a preset 3-layer fully connected neural network to obtain the final output as the prediction result of the future target.
7. The bearing fault prediction device based on multi-attention fuzzy cognitive graph according to claim 6, characterized in that, The AFCM prediction module is specifically used for: pre-training an attention-based STFCM, inputting the input data into the STFCM, and using the self-attention mechanism to learn the correlation between nodes in the graph. Calculate Q, K, and V. ; Use the obtained Q and K to calculate the correlation between every two input vectors: ; Perform a fuzzy mapping on A: , ; in These are parameters for learning the retrieved content. These are the parameters for index learning. These are the parameters learned from the value to be retrieved, where d is a constant. For input, For output, It is a transfer function that maps activation values to a specific range, where Q is the content to be retrieved, K is the index, and V is the value to be retrieved; after learning node relationships through the self-attention mechanism, the output node state value is updated to... ; Update the node state value The input is fed into an improved attention mechanism; a single time layer is expanded into multiple time layers to capture the relationship between the current state value of a node and its previous state values, as shown in the following formula; , , , in , , These are all learning parameters. , All are bias terms. It is a weight value. This indicates the hidden state and cell state of the transition-gated LSTM. It is a transfer function that maps activation values to a specific range; after learning the time step information of a node, the node's state value is updated. ; The input data and updated node state values are then fed into the residual connection: the input data and updated node state values are then fed into the residual connection, wherein... , These are the weights of the linear transformation. For bias terms, For node state values, The node state after the residual structure is applied. The node status value is updated after STFCM.
8. A bearing fault prediction device based on a multi-attention fuzzy cognitive map according to claim 6, characterized in that, The multi-image fusion module is specifically used for: , , ; in, This is the output of the first AFCM. This is the output of the second AFCM. This is the output of the third AFCM. It is the result of the fusion of time attention mechanisms. These are the parameters learned by the first linear transformation. These are the parameters learned by the second linear transformation. These are the parameters for learning. , All are bias terms. It is a weight value. For node state values, It is a transfer function that maps activation values to a specific range. This refers to the time range of the time-based attention mechanism, which is 1-15. The number of nodes; update the node status. Bearing target feature data with initial node state By combining, we obtain the fusion result after the residual structure. : ,in, , They are all weights of a linear transformation. This is a bias term.
9. A computer-readable storage medium, characterized in that: The storage medium stores instructions that, when executed, perform a bearing fault prediction method based on a multi-attention fuzzy cognitive graph as described in any one of claims 1-5.
Citation Information
Patent Citations
MR network signal intensity prediction method and system based on high-order fuzzy cognitive map
CN113365298A
Bearing fault prediction method based on fuzzy cognitive map
CN115901264A