Multi-sensor and cross-working-condition industrial fault diagnosis method

By constructing a 3D spatiotemporal co-current tensor and a lightweight multi-scale LightM-ConvKNet model, the problem of capturing the temporal sequence and local correlation of data in fault diagnosis under multiple sensors and cross-operating conditions was solved, and efficient fault diagnosis results were achieved.

CN120822033APending Publication Date: 2025-10-21NINGBO INST OF TECH ZHEJIANG UNIV ZHEJIANG
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510968322.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing multi-sensor and cross-condition industrial fault diagnosis methods, due to insufficient data acquisition and limitations of local characteristics of convolution kernels, struggle to fully capture the temporal and local correlations of sensor data, resulting in insufficient diagnostic accuracy and system stability.

Method used

A 3D spatiotemporal co-current tensor is constructed as the model input, and a lightweight multi-scale LightM-ConvKNet model is designed. Fault diagnosis is performed by combining transfer learning strategies.

Benefits of technology

It breaks through the limitation of convolution kernel size, fully obtains the local correlation between sensor data, improves the accuracy of fault diagnosis and system stability, and reduces computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822033A_ABST
    Figure CN120822033A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial fault diagnosis, in particular to a multi-sensor and cross-working-condition industrial fault diagnosis method. The method comprises the following steps: S1, acquiring equipment operation data of a plurality of sensors under different working conditions through a data acquisition system, and dividing the equipment operation data into a training set, a verification set and a test set; s2, constructing the original data into a 3D space-time collaborative tensor as model input; s3, CBT, LMSCB, ESRM and KAN-TD are embedded into a network architecture, LightM-ConvKNet intelligent fault diagnosis modeling is completed, then pre-training of the model is completed by using a source domain sample, and fine tuning is performed on the model based on a target domain sample; and S4, inputting the target domain test set into the fine-tuned model to generate a fault diagnosis result. By adopting the method, the limitation of the CNN convolution kernel size is broken through, the local correlation among sensor data can be fully acquired, the parameter quantity is reduced, the calculation efficiency is improved, and the calculation complexity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial fault diagnosis, and in particular to an industrial fault diagnosis method for multiple sensors and across working conditions. Background Art

[0002] With the accelerating pace of industrial intelligence, industrial equipment plays a critical role in productivity and safety. However, prolonged use and harsh environments can easily lead to failures, and lack of maintenance can threaten safety and reliability. Therefore, developing efficient industrial fault diagnosis algorithms is crucial. Fault diagnosis methods primarily fall into two categories: signal analysis-based and data-driven. Signal analysis-based methods suffer from a heavy reliance on domain experts and prior knowledge. In contrast, data-driven approaches offer greater flexibility by automatically extracting features from massive amounts of collected vibration data using statistical techniques to identify potential failure modes. The rapid development of deep learning in recent years, coupled with its powerful ability to extract deeper features, has provided new insights into intelligent fault diagnosis. Unlike traditional methods, deep learning automatically mines deep information from raw data by constructing nonlinear networks. This allows for the representation of complex nonlinear relationships without the need for manual feature selection, establishing a precise correspondence between signals and equipment status. Common methods include deep belief networks (DBNs), long short-term memory networks (LSTMs), and convolutional neural networks (CNNs). CNNs, in particular, have been widely used in the field of intelligent industry and achieved remarkable results due to their flexibility and powerful local perception capabilities. For example, multi-time channel CNNs (MC-CNNs) are used to solve the online prediction problem of multi-sampling rate quality variables; multi-layer wavelet attention CNNs (MWA-CNNs) achieve accurate diagnosis of industrial faults.

[0003] In actual industry, equipment fault signals often have multi-scale characteristics (such as low-frequency trends and high-frequency impulses), reflecting different working conditions. Researchers have found that single-receptive-field convolution models are only sensitive to specific frequency bands, while multi-scale convolution kernels can extract more comprehensive features. They have also proposed related work, such as using multi-scale convolutional neural networks to expand the network's receptive field and combining multiple attention mechanisms to enhance fault feature responses; and proposing a multi-scale attention CNN (MSACNN) to improve the prediction accuracy of soft sensing models. Although these methods effectively improve model performance, they also lead to model complexity, which limits their application in resource-constrained scenarios. For this reason, research on model lightweighting has attracted attention, such as the lightweight multi-scale feature fusion network (MSF-LightNet), which combines depthwise separable convolution (DSC) with multi-expansion rate dilated convolution to balance the model's lightweightness and generalization capabilities.

[0004] Furthermore, these CNN-based models excel at extracting spatial information from data, but struggle to fully reveal the temporal information inherent in sequential data. Therefore, some researchers have used time-delay techniques to capture the temporal dynamics of data and proposed dynamic CNNs. They have also introduced a sparsely modified multi-head self-attention mechanism based on CNNs and proposed the Convformer-NSE framework to model long-range dependencies between feature maps. Finally, they have combined adaptive multi-scale CNNs with bidirectional LSTMs to form the ACEL network, capturing both local and global information.

[0005] With the continuous development and increasing complexity of industrial technology, the information obtained by a single sensor is limited and susceptible to noise interference, resulting in reduced diagnostic accuracy and system instability. Multi-sensor systems, on the other hand, can provide richer information, reduce the limitations of a single sensor, and improve fault diagnosis accuracy and system stability. Therefore, multi-sensor fusion has become a trend in fault diagnosis research and has achieved certain results. For example, multi-sensor information is fused through a kurtosis-weighted strategy, and a one-dimensional coordinate attention module is introduced into a CNN to enhance feature extraction capabilities. Principal component analysis (PCA) is used to convert multi-sensor data into RGB images, and the MSF-RCNN model is proposed for diagnosis.

[0006] While the aforementioned CNN-based multi-sensor fusion methods have demonstrated their advantages, they still have certain limitations. First, they require sufficient raw data for model training and assume that the source and target data have the same distribution. In real-world industrial environments, however, due to variations in sensor sampling rates and complex and variable operating conditions, collecting rich data is difficult. Second, the local nature of convolutional kernels prevents them from fully capturing the contextual information in sensor data, resulting in a lack of temporal characterization of fault signatures. Existing research has addressed this issue by introducing sequence models such as LSTM and Transformer, but this has led to a dramatic increase in parameters and computational requirements, posing challenges for practical industrial applications. Furthermore, in multi-sensor systems, signals between sensors and acquisition channels exhibit coupling. In industry, the arrangement of different data types is typically based on the physical location of sensors or the order of acquisition channels. These signal data are not only closely correlated with data from adjacent acquisition channels but also exhibit correlations with nodes in distant topologies. Exploring the potential connections between these signals and uncovering and leveraging their correlated properties is crucial for enhancing model representation capabilities and improving diagnostic accuracy. However, as a classic local feature extractor, CNN is limited in its convolution kernel size, which restricts its ability to capture the correlation between distant topological structure data and thus cannot fully reveal the local correlation between multi-sensor data. Summary of the Invention

[0007] The technical solution to be solved by the present invention is: to provide an industrial fault diagnosis method for multiple sensors and cross-working conditions. The method first constructs a 3D spatiotemporal collaborative tensor as sample input, then designs a lightweight multi-scale LightM-ConvKNet model as the fault diagnosis backbone network, and finally uses a transfer learning strategy to realize industrial cross-working condition fault diagnosis.

[0008] The technical solution adopted by the present invention is: a multi-sensor and cross-working condition industrial fault diagnosis method, which includes the following steps:

[0009] S1. Use the data acquisition system to obtain equipment operating data from multiple sensors under different working conditions and divide it into training sets, validation sets, and test sets;

[0010] S2. Constructing the original data into a 3D spatiotemporal synergistic tensor as model input, which specifically includes the following steps:

[0011] S21, performing normalization preprocessing on the data;

[0012] S22, uses TDS technology to integrate time delay and feature information, and shuffles the two-dimensional samples multiple times;

[0013] S23, generate the target tensor by channel stacking and dimension raising;

[0014] S3: Embed CBT, LMSCB, ESRM, and KAN-TD into the network architecture to complete intelligent fault diagnosis modeling. Then, use source domain samples to complete the model pre-training, and then fine-tune the model based on target domain samples.

[0015] S4. Input the target domain test set into the fine-tuned model to generate fault diagnosis results.

[0016] Preferably, step S21 performs normalization preprocessing on the data, specifically: scaling the data using maximum-minimum normalization, the formula of which is as follows:

[0017]

[0018] Among them, X, X min and X max Represents the original data, minimum value and maximum value respectively, and the standard data X obtained after normalization normalized is scaled to the range [-1,1].

[0019] As a preferred embodiment, the TDS technology is used to integrate the time delay and characteristic information in step S22 as follows: Assume that the sample sample at the sampling time t of the data acquisition system is X(t) = [x1(t), x2(t), ..., x n(t)], where n represents the number of sensors. Through the time-delay displacement technology, the samples at time t and the previous T sampling times are expanded into two-dimensional dynamic samples:

[0020]

[0021] Where T is the sampling delay and f is the sampling frequency.

[0022] Preferably, the multiple shuffling of the two-dimensional samples in step S22 is specifically as follows:

[0023] Assume that the initial two-dimensional sample is recorded as:

[0024]

[0025] Then the two-dimensional sample obtained by the k-1th shuffle is recorded as:

[0026]

[0027] If X k is the original two-dimensional sample X 1 The data of the first column and the last column are exchanged in the shuffle operation, which can also be expressed as:

[0028]

[0029] Preferably, step S23 generates a target tensor by channel stacking and dimension raising as follows:

[0030] The two-dimensional samples obtained after multiple shuffling operations are superimposed with the original sequential two-dimensional samples to obtain a multi-channel 3D spatiotemporal coordination tensor.

[0031] Preferably, the intelligent fault diagnosis modeling in step S3 consists of two parts: a feature extractor and a classifier, wherein:

[0032] The feature extractor includes two CBT, two LMSCB and one ESRM modules;

[0033] Classifiers include KAN-TD.

[0034] Preferably, the LMSCB refers to a lightweight multi-scale convolution block, which specifically includes three parallel branches, each branch includes an ISDCRB, the three ISDCRBs use small, medium and large size convolution kernels respectively, and combine nonlinear activation and maximum pooling operations after convolution.

[0035] Preferably, the ISDCRB is an improved depth-separable convolution residual block, which is specifically:

[0036] In the main path, the input is first subjected to depthwise convolution, with the output channels corresponding one-to-one to the input channels. Subsequently, point convolution is used for cross-channel feature interaction. Batch normalization and activation functions are embedded in the two-step convolution. The entire process can be expressed as:

[0037] F(x)=Conv PW (Tanh(BN(Conv DW (x))))

[0038] Among them, Tanh represents the hyperbolic tangent activation function, BN is batch normalization, Conv DW For depth convolution, Conv PW It is a point-by-point convolution, the residual connection part is PW, and the input channel dimension is adjusted to match the output of the main path. The process can be expressed as:

[0039] r(x)=Conv PW (x)

[0040] Add the output of the main path and the residual connection as the final output of the module:

[0041] y=F(x)+r(x).

[0042] As an example, the ESRM refers to an efficient style attention module, which is specifically: style pooling is encoded by the mean and standard deviation, and the style feature F style It can be expressed as:

[0043] F style =GAP(c i )+GSP(c i ),i∈R C

[0044] where c i represents the i-th input channel, C is the number of input channels, GAP is global average pooling, GSP is global standard pooling, and then one-dimensional adaptive convolution is used to further integrate the style features of the device operation state, and the interaction range is determined by adaptively adjusting the convolution kernel size. In addition, the Sigmoid function is used to introduce nonlinear constraints to ensure that the generated channel weights are between [0,1]. The style feature F style The integration process can be expressed as:

[0045] ω=σ(C1D k (F style ))

[0046] Where ω represents channel attention information, C1D() represents a one-dimensional convolution operation, σ is the Sigmoid activation function, and the coverage of cross-channel information interaction, that is, the convolution kernel size k, is mapped to the number of feature channels C as follows:

[0047] C=φ(k)=2 (γ·k-b)

[0048] Therefore, the convolution kernel size k can be determined according to the number of feature channels C:

[0049]

[0050] Here, odd indicates that k can only be an odd number, and usually γ and b are 2 and 1 respectively.

[0051] Preferably, the KAN-TD is specifically composed of two layers of KANLinear, wherein the number of bottom network nodes and the number of top operators are both 3, and the SiLU activation function of its linear operation part is replaced by Tanh. In addition, the nonlinear part is kept stable in the KANLinear layer, and the Dropout mechanism is added to the linear feature extraction part.

[0052] Compared with the prior art, the present invention has the following advantages:

[0053] The original data is constructed into a 3D spatiotemporal coordination tensor as the model input. In this way, the time series data of the same sequence in the three-dimensional tensor may correspond to different sensors in different channels. Therefore, the three-dimensional target tensor has multiple domain topologies of multi-sensor data in its different channels, thus breaking the limitation of CNN convolution kernel size and facilitating the full acquisition of local correlation between sensor data.

[0054] The input tensor is processed by two layers of CBT to extract shallow fault features. CBT consists of convolutional layers, BN layers and Tanh activation functions. Batch normalization and Tanh activation accelerate convergence while improving the generalization and expression capabilities of the network. The two layers of LMSCB serve as the core of feature extraction, and multi-scale IDSCRB is used to extract spatial features at different resolutions in parallel, while reducing the number of parameters and improving computational efficiency. Subsequently, multi-branch features are integrated and input into the ESRM module to capture key fine-grained information in the global range. Finally, the classifier uses KAN-TD to enhance the nonlinear mapping capability of the network while reducing computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is an overall framework diagram of an industrial fault diagnosis method for multiple sensors and cross-working conditions.

[0056] Figure 2 This is a schematic diagram of time-delay displacement technology.

[0057] Figure 3 This is a schematic diagram of the dimensionality increase process based on multiple shuffling.

[0058] Figure 4 It is a model architecture diagram.

[0059] Figure 5 This is the ISCRB structure diagram.

[0060] Figure 6 This is the LMSCB structure diagram.

[0061] Figure 7 This is the structural diagram of ESRM.

[0062] Figure 8 This is the KAN-TD structure diagram.

[0063] Figure 9 (a) Experimental setup; (b) Bearing state type in the experiment.

[0064] Figure 10 is the confusion matrix for each number of channels in the experiment.

[0065] Figure 11 It is a radar chart of the performance of different methods in the experiment.

[0066] Figure 12 Comparison of accuracy and model complexity in the experiment: (a) number of parameters and accuracy; (b) floating-point operation performance and accuracy. DETAILED DESCRIPTION

[0067] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0068] Example 1:

[0069] A multi-sensor and cross-operation-condition industrial fault diagnosis method includes the following steps:

[0070] S1. Use the data acquisition system to obtain equipment operating data from multiple sensors under different working conditions and divide it into training sets, validation sets, and test sets;

[0071] S2, construct the original data into a 3D spatiotemporal synergistic tensor as model input,

[0072] In order to obtain the temporal sequence of the original data and effectively fuse multi-sensor data, this embodiment proposes the construction of a 3D spatiotemporal collaborative tensor. First, the original data is preprocessed to eliminate interference and accelerate model convergence. Specifically, the maximum-minimum normalization (Min-Max Normalization) is used to scale the data and convert it to a specified range to ensure that the data from different sensors have the same scale, thereby improving model efficiency. The formula is as follows:

[0073]

[0074] Among them, X, X min and X max Represents the original data, minimum value and maximum value respectively, and the standard data X obtained after normalization normalized It is scaled to the range of [-1, 1] to reduce feature differences.

[0075] In the data acquisition system, multiple sensors collect the operating data of industrial equipment at a certain sampling rate. Assume that the sampling sample at the sampling time t of a data acquisition system is X(t) = [x1(t), x2(t),…, x n (t)], where n represents the number of sensors, e.g. Figure 2 As shown in the figure, the samples at time t and the previous T sampling times are expanded into two-dimensional dynamic samples through the time delay shift technique (TDS):

[0076]

[0077] Where T is the sampling delay and f is the sampling frequency. By constructing two-dimensional dynamic samples, the integration of data delay information and feature information is achieved.

[0078] As can be seen, the sample values ​​in each column of the two-dimensional sample come from different sensors. We perform a uniform random column shuffle on the sample, randomly swapping the order of two columns each time, thereby changing the domain relationship between sensors. Multiple shuffles enrich the topological structure of multi-sensor data. Assume that the initial two-dimensional sample is denoted as:

[0079]

[0080] Then the two-dimensional sample obtained by the k-1th shuffle is recorded as:

[0081]

[0082] If X k is the original two-dimensional sample X 1 The data of the first column and the last column are exchanged in the shuffle operation, which can also be expressed as:

[0083]

[0084] Finally, the two-dimensional samples obtained after multiple shuffle operations are superimposed with the original sequential two-dimensional samples to obtain a multi-channel 3D spatiotemporal coordination tensor. In this way, when the sampling delay is T, the shape of a k-channel target three-dimensional tensor is (T+1)×k×n. Figure 3 The shuffling process is demonstrated. As can be seen, after shuffling and stacking the original two-dimensional samples, the time series data of the same sequence in the three-dimensional tensor may correspond to different sensors in different channels. Therefore, the three-dimensional target tensor has multiple domain topologies of multi-sensor data in its different channels, breaking through the limitations of CNN convolution kernel size and facilitating the full capture of local correlations between sensor data.

[0085] S3. Embed CBT, LMSCB, ESRM, and KAN-TD into the network architecture to complete the LightM-ConvKNet intelligent fault diagnosis modeling. Then, use source domain samples to complete the model pre-training, and then fine-tune the model based on target domain samples.

[0086] Among them Figure 4 As shown in Figure 2, the LightM-ConvKNet intelligent fault diagnosis model consists of two parts: a feature extractor and a classifier.

[0087] The feature extractor consists of two CBTs, two LMSCBs, and an ESRM module. The input tensor first undergoes two CBT layers to extract shallow fault features. The CBT, consisting of convolutional layers, batch normalization layers, and Tanh activation functions, accelerates convergence through batch normalization and Tanh activation, while improving the network's generalization and expressiveness. The two LMSCB layers serve as the core of feature extraction, utilizing multi-scale IDSCRBs to concurrently extract spatial features at different resolutions, while reducing the number of parameters and improving computational efficiency. Subsequently, the multi-branch features are integrated and fed into the ESRM module to capture critical, fine-grained information globally.

[0088] The classifier includes KAN-TD, whose special structure enhances the nonlinear mapping capability of the network while reducing computational complexity. Ultimately, KAN-TD maps high-dimensional features to the fault category space through linear and nonlinear transformations.

[0089] CBT, which stands for Convolution-BN-Tanh, consists of a convolutional layer, a BN layer, and a Tanh activation function. It accelerates convergence through batch normalization and Tanh activation, while improving the generalization and expression capabilities of the network.

[0090] LMSCB is a lightweight multi-scale convolution block. Traditional convolution operation integrates feature information through cross-channel convolution, but it needs to initialize the convolution kernel weight for each output channel, which increases the parameters and computational complexity of the model. Inspired by the depthwise separable convolution (DSC), the improved depthwise separable convolution residual block (IDSCRB) is proposed. Its structure is as follows: Figure 5 As shown in the figure, residual connections are introduced to promote cross-layer propagation of information in the network, alleviate the gradient vanishing problem, and compensate for the feature loss that may be caused by DSC, thereby improving model performance.

[0091] In the main path, the input is first subjected to depthwise convolution (DW), with the output channels corresponding one-to-one with the input channels, independently extracting spatial features and significantly reducing the number of parameters and computation. Subsequently, pointwise convolution (PW) is used to interact with cross-channel features, integrate information, and adjust the output dimension. Batch normalization (BN) and Tanh activation functions are embedded in the two-step convolution to reduce information loss and avoid the vanishing gradient problem, making PW computation more efficient. The entire process can be expressed as:

[0092] F(x)=Conv PW (Tanh(BN(Conv DW (x)))) (6)

[0093] Among them, Tanh represents the hyperbolic tangent activation function, BN is batch normalization, Conv DW For depth convolution, Conv PW is a point-by-point convolution. The residual connection part is PW, which adjusts the input channel dimension to match the output of the main path. The process can be expressed as:

[0094] r(x)=Conv PW (x) (7)

[0095] Add the output of the main path and the residual connection as the final output of the module:

[0096] y=F(x)+r(x) ​​(8)

[0097] In actual industrial applications, the equipment environment is complex, vibration and friction signals are often mixed with noise, and the differences in accuracy and range of different sensors increase the difficulty of fault feature extraction. To this end, a lightweight multi-scale convolutional block (LMSCB) is designed based on ISDCRB, which uses convolution kernels of different sizes to extract multi-level feature information. Figure 6 As shown in the figure, LMSCB uses three parallel branches, using small, medium, and large convolution kernels to capture local and global features and enhance fault pattern recognition capabilities. Convolution is combined with nonlinear activation and maximum pooling operations to extract more representative key features.

[0098] ESRM is an efficient style attention module. Style pooling combines global mean pooling (GAP) and global standard deviation pooling (GSP). Its core idea is to extract style information through GAP and GSP, and use it for style recalibration module (SRM) to optimize feature representation and improve model performance. This embodiment applies style pooling to cross-condition fault diagnosis and proposes a style-based efficient attention module (ESRM). Style pooling captures global trends and detail distribution through joint encoding of mean and standard deviation, reduces feature domain deviation, and enhances cross-domain adaptability. Figure 7 As shown, unlike SRM, we innovatively use element-by-element addition instead of feature dimension splicing for feature fusion, and the obtained style feature F style It can be expressed as:

[0099] F style =GAP(c i )+GSP(c i ),i∈R C (9)

[0100] where c i Represents the i-th input channel, and C is the number of input channels. This method directly combines the global trend (mean) and the detailed distribution (standard deviation) to make the feature expression more compact while reducing the feature dimension and computational complexity. Compared with SRM, this method effectively avoids redundant features and more efficiently encodes the distribution characteristics of cross-working condition signals. Subsequently, 1D adaptive convolution is used to further integrate the style features of the equipment operating status. This convolution models the correlation of style features, realizes cross-channel interaction of local features, and determines the interaction range by adaptively adjusting the convolution kernel size. In addition, the Sigmoid function is used to introduce nonlinear constraints to ensure that the generated channel weights are between [0,1]. Style feature F style The integration process can be expressed as:

[0101] ω=σ(C1D k (F style )) (10)

[0102] Where ω represents channel attention information, C1D() represents a one-dimensional convolution operation, and σ is the Sigmoid activation function. The coverage of cross-channel information interaction, that is, the convolution kernel size k, is mapped to the number of feature channels C as follows:

[0103] C=φ(k)=2 (γ·k-b) (11)

[0104] Therefore, the convolution kernel size k can be determined according to the number of feature channels C:

[0105]

[0106] Here, odd indicates that k can only be an odd number, and usually γ and b are 2 and 1 respectively.

[0107] KAN-TD stands for Tanh-activated KAN with Drop mechanism. Multilayer perceptrons (MLPs) are a fundamental building block of deep learning. Each neuron in the MLP is connected to all neurons in the previous layer, enabling it to capture global features. However, this leads to a large number of parameters in the fully connected layers, increasing training costs. Furthermore, since the training process is a near-black box, it often lacks interpretability. To address these issues, Kolmogorov-Arnold Networks (KANs) were developed. Their theoretical basis stems from the Kolmogorov-Arnold theorem, which states that any continuous multivariate function on a bounded domain can be expressed as a linear combination of a finite number of univariate functions. Unlike traditional MLPs, KANs use learnable univariate functions as activation functions on edges (weights) rather than fixed activation functions on nodes (neurons). This shift fundamentally changes the network's architecture and learning capabilities. Each edge is associated with a learnable univariate function parameterized in a spline form. This approach allows the network to dynamically adjust the activation function based on input data, improving flexibility and accuracy. While MLP relies on linear weight matrices, KAN replaces traditional weights with learnable single-variable functions represented by spline curves. This reduces reliance on linear transformations and enhances the ability to model complex functions. The following equation illustrates the relationship between neurons in a KAN layer:

[0108]

[0109] where x l,i is the input of any neuron number i in layer l, θ l,i,j (x l,i ) is the activation function, which is a linear combination of the basis function and the spline function:

[0110] θ(x)=ω b silu(x)+ω s spline(x) (14)

[0111] where ω b and ω s It is a learnable parameter that better controls the overall size of the activation function. silu(x) is a Sigmoid linear unit function, which is introduced into KAN as a basis function similar to the residual connection to stabilize the optimization process. spline(x) is parameterized as a weighted sum of B-splines:

[0112]

[0113] where ωp is the learnable factor, B p (x) is the basis spline, k and G are the grid size and order of the spline respectively. l The input neurons are transformed nonlinearly through a learnable B-spline combination, and then the j-th output x of the next layer is obtained through weighted summation. l+1,j The process can be expressed in matrix form as:

[0114]

[0115] Then the deep KAN can be expressed as a layer nested structure:

[0116] KAN(x)=(ψ l-1 ·ψ l-2 ·…·ψ2·ψ1)x (17)

[0117] The KAN architecture is essentially composed of multiple stacked KANLinear layers. Each KANLinear layer consists of learnable single-variable functions that replace the fixed weights in the MLP. In the inter-layer transformations of the KAN, key computations are performed by the operators at the bottom and top layers. The number of nodes in the bottom layer determines the scale of feature aggregation after cubic spline processing. The top and bottom layers, respectively, perform preliminary feature extraction and fine-tuning transformations using cubic spline functions to enhance the model's pattern capture capabilities. Their number, determined by the number of neurons between layers, is crucial for optimizing the KAN architecture and improving its diagnostic performance.

[0118] To address the problems and shortcomings of traditional MLP, this example introduces the concept of KAN into industrial fault diagnosis and proposes a Tanh-activated KAN with Drop mechanism (KAN-TD). It is then integrated into the existing model architecture as the final classification layer to output diagnostic results. Figure 8 is a KAN-TD with a structure of [2,3,1], consisting of two layers of KANLinear, where the number of bottom-layer network nodes and the number of top-layer operators are both 3. We replace the SiLU activation function in its linear operation part with Tanh to adapt to the symmetrical vibration signal data. Compared with SiLU, the Tanh function has a greater advantage in modeling symmetry and can more accurately capture the input data distribution centered on zero and with amplitude symmetry, thereby improving the model's ability to represent vibration data. Therefore, Equation (14) is updated as follows:

[0119] θ(x)=ω b Tanh(x)+ω s spline(x) (18)

[0120] In addition, the nonlinear part is kept stable in the KANLinear layer, and the Dropout mechanism is added to the linear feature extraction part to better regularize the model and prevent overfitting, while retaining the nonlinear part's ability to model complex relationships.

[0121] Specific experiments:

[0122] 1. Experimental setup,

[0123] The experimental data were collected on an industrial bearing test bench, which includes a drive motor, a rotor system, and a control system. Mechanical devices such as Figure 9 (a) As shown. The rated power of the drive motor is 2.2kW, and the speed range is 0-6000r / min. The test bench is equipped with multiple sensors to monitor the operating status of the bearing: the HY-YD-232 acceleration sensor is vertically installed on the bearing support seat to monitor the vibration acceleration signals in three directions; the WT series eddy current sensor detects displacement changes; the DYN-200 torque sensor detects torque and speed; and the CSM020GB series Hall current sensor collects electrical signals. Figure 9 As shown in (b), the bearing conditions include five different operating states: normal, outer ring failure, inner ring failure, cage fracture, and rolling element pitting. Four different radial forces (0 N, 500 N, 1000 N, and 1500 N) were applied to establish different operating conditions, resulting in 20 bearing state combinations, as detailed in Table 1. The OF1000 is used as an example to illustrate the outer ring failure condition under a 1000 N radial load. The experimental data were recorded on 12 channels using an HD9200 multi-channel data acquisition system at a speed of 1500 rpm and a sampling rate of 10,240 S / s. Table 2 summarizes the physical parameters of the dataset.

[0124] Table 1 Experimental bearing conditions

[0125]

[0126] Table 2 Data acquisition system variables

[0127]

[0128] 2. Data configuration

[0129] The original data is divided into a pre-training set and a fine-tuning set. Samples are generated by non-overlapping sliding window segmentation of the original signal to prevent data leakage.

[0130] The pre-training dataset contains 2,500 samples with a training-to-validation ratio of 4:1, covering only a single operating condition. The fine-tuning dataset contains 1,350 samples with a training-to-validation-to-test ratio of 4:1:4, consisting of mixed operating conditions to simulate complex industrial environments. To evaluate the cross-operational diagnostic capabilities of the proposed method, we set four transfer tasks, as shown in Table 3.

[0131] Table 3 Cross-condition task settings

[0132]

[0133] To evaluate the performance of the fault diagnosis model, we used accuracy, recall, precision, and F1-score as metrics. Furthermore, to ensure the stability and reliability of the experimental results, each set of experiments was repeated 10 times, and the average value was taken as the final result.

[0134] 3. Experimental process and result analysis

[0135] First, to verify the effectiveness of the shuffle operation in enhancing local correlations in multi-sensor data, we generated tensor samples with different numbers of channels by adjusting the number of shuffles and compared the model's diagnostic performance under different scenarios. Taking Task S1 as an example, the diagnostic results for each scenario are shown in Table 4, with the best performance highlighted in bold. The results show that the tensor without the shuffle operation is single-channel, resulting in a simple topological structure for the multi-sensor data. Limited by the convolution kernel size, it is difficult to fully exploit the local correlations between the multi-sensor data, resulting in the worst model diagnostic performance. After spatial dimensionality enhancement through the shuffle operation, the resulting multi-channel tensor is able to more fully exploit the complementary information in the sensor data, and model performance gradually improves with increasing channel numbers. The performance improvement is particularly significant when expanding from a single channel to a dual channel, with all metrics increasing by approximately 13%. This result demonstrates that spatial dimensionality enhancement based on shuffling can enrich the topological structure of multi-sensor data, enhance the ability to extract local correlations, and thus improve fault diagnosis performance.

[0136] Table 4 Experimental results under different channel numbers

[0137]

[0138] like Figure 10Shown are the confusion matrices for various channel numbers: (a) 1-channel Shuffle; (b) 2-channel Shuffle; (c) 3-channel Shuffle; and (d) 4-channel Shuffle. Performance peaks with 3 channels, with the confusion matrix indicating optimal classification. Across 10 trials, only 34 of the 15 target domain categories were misclassified, and 5,966 samples were correctly identified. With the exception of two categories with an accuracy of 98.5%, most classifications achieved accuracy exceeding 99%, and 100% accuracy for all three fault types. Increasing the number of channels beyond 4 slightly degrades performance, likely due to redundant information and noise, while also increasing computational complexity and cost. To balance performance and computational efficiency, we fixed the number of channels to 3 as the default configuration for subsequent experiments.

[0139] To quantitatively analyze the impact of each module on model performance, we used the proposed architecture as a baseline and gradually removed each module to construct a comparison model. We then conducted a systematic ablation experiment on Task S2. The specific settings are as follows:

[0140] (1) Model #1: Use one-dimensional convolution to directly process the original one-dimensional samples without spatiotemporal co-reconstruction;

[0141] (2) Model #2: Replace MSLCB with standard convolution;

[0142] (3) Model #3: Remove ESRM module;

[0143] (4) Model #4: To verify the role of KAN-TD in improving the model's expressiveness and lightweightness.

[0144] The various model structures are shown in Table 5, and the ablation experiment results are shown in Table 6. The results show that the baseline model performs best across all performance metrics. Compared to the baseline model, the accuracy of Model #1 drops significantly, reaching only 41.02%. Compared to Model #2, the baseline model reduces the amount of computation by approximately 18% while improving accuracy by 3.04%. Compared to Model #3, the baseline model further reduces the risk of missed diagnosis while maintaining similar model complexity. Model #4 uses a traditional MLP instead of KAN-TD, resulting in a significant decrease in expression efficiency, a 23-fold increase in parameters, a 17% increase in FLOPs, and a 1.54% decrease in F1 score. The experimental results fully validate the effectiveness of each module.

[0145] Table 5 Ablation experiment settings

[0146]

[0147] Table 6 Ablation experiment results

[0148]

[0149] To verify the superiority of LightMConvKNet in cross-operating-condition fault diagnosis, we compared it with four state-of-the-art related methods. All methods shared the same hyperparameter settings. The experimental results are shown in Table 7, with the best performance indicators highlighted in bold. The results demonstrate that the proposed method outperforms the other methods in all four cross-operating-condition diagnosis tasks, achieving average performance exceeding 99%.

[0150] Table 7 Comparative experimental results of different methods

[0151]

[0152]

[0153] Figure 11 A radar chart is used to visually compare the performance metrics of various methods. Each axis represents an evaluation metric, and the position of the dot corresponds to the metric value. Different methods are represented by different colors, and polygons are formed by connecting lines. The shape and size of the polygons clearly demonstrate the overall superiority of the proposed method in cross-condition diagnosis tasks.

[0154] Table 8 Relationship between accuracy and computational complexity

[0155]

[0156] Table 8 and Figure 12 The relationship between the average accuracy of each method and the computational complexity of the model in cross-operating condition diagnosis tasks is demonstrated. Compared with other methods, the proposed method has better diagnostic performance while maintaining lightweight and robustness. This comprehensive advantage is due to the synergistic effect of multi-sensor data fusion, multi-level feature learning and computational optimization mechanism. The method proposed in this patent effectively integrates data timing information and multi-sensor features by constructing a 3D spatiotemporal collaborative tensor. LightM-ConvKNet integrates multiple effective modules, and exhibits optimal diagnostic performance while achieving model lightweight.

[0157] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

[0158] Various changes and modifications will undoubtedly become apparent to those skilled in the art upon reading the foregoing description. Therefore, the appended claims should be construed to encompass all changes and modifications within the true intent and scope of the present invention. Any and all equivalents within the scope of the claims should be considered to be within the intent and scope of the present invention.

Claims

1. A multi-sensor and cross-operating-condition industrial fault diagnosis method, characterized in that: It includes the following steps: S1. Use the data acquisition system to obtain equipment operating data from multiple sensors under different working conditions and divide it into training sets, validation sets, and test sets; S2. Constructing the original data into a 3D spatiotemporal synergistic tensor as model input, which specifically includes the following steps: S21, performing normalization preprocessing on the data; S22, uses TDS technology to integrate time delay and feature information, and shuffles the two-dimensional samples multiple times; S23, generate the target tensor by channel stacking and dimension raising; S3. Embed CBT, LMSCB, ESRM, and KAN-TD into the network architecture to complete the LightM-ConvKNet intelligent fault diagnosis modeling. Then, use source domain samples to complete the model pre-training, and then fine-tune the model based on target domain samples. S4. Input the target domain test set into the fine-tuned model to generate fault diagnosis results.

2. The multi-sensor and cross-operating-condition industrial fault diagnosis method according to claim 1, characterized in that: Step S21 performs normalization preprocessing on the data, specifically: scaling the data using maximum-minimum normalization, the formula is as follows: Among them, X, X min and X max Represents the original data, minimum value and maximum value respectively, and the standard data X obtained after normalization normalized is scaled to the range [-1,1].

3. The multi-sensor and cross-operation-condition industrial fault diagnosis method according to claim 2, characterized in that: The specific steps of integrating the time delay and characteristic information by using TDS technology in step S22 are as follows: Assume that the sample of the data acquisition system at sampling time t is X(t) = [x1(t), x2(t), ..., x n (t)], where n represents the number of sensors. Through the time-delay displacement technology, the samples at time t and the previous T sampling times are expanded into two-dimensional dynamic samples: Where T is the sampling delay and f is the sampling frequency.

4. The method for multi-sensor and cross-operating-condition industrial fault diagnosis according to claim 2, characterized in that: The multiple shuffling of the two-dimensional samples in step S22 is specifically as follows: Assume that the initial two-dimensional sample is recorded as: Then the two-dimensional sample obtained by the k-1th shuffle is recorded as: If X k is the original two-dimensional sample X 1 The data of the first column and the last column are exchanged in the shuffle operation, which can also be expressed as:

5. The method for multi-sensor and cross-operating-condition industrial fault diagnosis according to claim 4, characterized in that: Step S23 generates the target tensor by channel stacking and dimension raising as follows: The two-dimensional samples obtained after multiple shuffling operations are superimposed with the original sequential two-dimensional samples to obtain a multi-channel 3D spatiotemporal coordination tensor.

6. The method for multi-sensor and cross-operating-condition industrial fault diagnosis according to claim 1, characterized in that: The intelligent fault diagnosis modeling in step S3 consists of two parts: feature extractor and classifier, where: The feature extractor includes two CBT, two LMSCB and one ESRM modules; Classifiers include KAN-TD.

7. The method for multi-sensor and cross-operating-condition industrial fault diagnosis according to claim 6, characterized in that: The LMSCB refers to a lightweight multi-scale convolutional block, which specifically includes three parallel branches, each of which includes an ISDCRB. The three ISDCRBs use small, medium, and large-sized convolution kernels, respectively, and combine nonlinear activation and maximum pooling operations after convolution.

8. The method for multi-sensor and cross-operating-condition industrial fault diagnosis according to claim 7, characterized in that: The ISDCRB is an improved depth-separable convolution residual block, which is specifically: In the main path, the input is first subjected to depthwise convolution, with the output channels corresponding one-to-one to the input channels. Subsequently, point convolution is used for cross-channel feature interaction. Batch normalization and activation functions are embedded in the two-step convolution. The entire process can be expressed as: F(x)=Conv PW (Tanh(BN(Conv DW (x)))) Among them, Tanh represents the hyperbolic tangent activation function, BN is batch normalization, Conv DW For depth convolution, Conv PW It is a point-by-point convolution, the residual connection part is PW, and the input channel dimension is adjusted to match the output of the main path. The process can be expressed as: r(x)=Conv PW (x) Add the output of the main path and the residual connection as the final output of the module: y=F(x)+r(x).

9. The method for multi-sensor and cross-operating-condition industrial fault diagnosis according to claim 6, characterized in that: The ESRM refers to the efficient style attention module, which is specifically: style pooling is encoded by the mean and standard deviation, and the style feature F style It can be expressed as: F style =GAP(c i )+GSP(c i ),i∈R C where c i represents the i-th input channel, C is the number of input channels, GAP is global average pooling, GSP is global standard pooling, and then one-dimensional adaptive convolution is used to further integrate the style features of the device operation state, and the interaction range is determined by adaptively adjusting the convolution kernel size. In addition, the Sigmoid function is used to introduce nonlinear constraints to ensure that the generated channel weights are between [0,1]. The style feature F style The integration process can be expressed as: ω=σ(C1D k (F style )) Where ω represents channel attention information, C1D() represents a one-dimensional convolution operation, σ is the Sigmoid activation function, and the coverage of cross-channel information interaction, that is, the convolution kernel size k, is mapped to the number of feature channels C as follows: C=φ(k)=2 (γ·k-b) Therefore, the convolution kernel size k can be determined according to the number of feature channels C: Here, odd indicates that k can only be an odd number, and usually γ and b are 2 and 1 respectively.

10. The method for multi-sensor and cross-operating-condition industrial fault diagnosis according to claim 6, characterized in that: Specifically, the KAN-TD consists of two layers of KANLinear, where the number of bottom-layer network nodes and the number of top-layer operators are both 3, and the SiLU activation function of its linear operation part is replaced by Tanh. In addition, the nonlinear part is kept stable in the KANLinear layer, and the Dropout mechanism is added to the linear feature extraction part.

Citation Information

Cited By

  • Energy consumption prediction model construction method, energy consumption prediction method and energy consumption prediction device

    CN121543843A