Fault information adaptive diagnosis method based on deep residual network
Through the gated residual enhancement module and the void convolution fusion module of the deep residual network, combined with multimodal dynamic attention fusion and adaptive update mechanism, the problems of weak signal feature expression and low model migration efficiency in early fault diagnosis of industrial equipment are solved, and efficient, real-time fault identification and lightweight deployment are achieved.
Patent Information
- Application Number
- CN202510739713.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies are unable to effectively express the weak signal characteristics of industrial equipment in the early stages of operation or in the stage of minor damage. In addition, traditional methods have low model migration efficiency and high retraining costs in industrial field environments, making it difficult to meet the needs of non-stop, real-time diagnosis. There is also a risk of model degradation caused by the accumulation of pseudo-label noise.
An adaptive fault information diagnosis method based on a deep residual network is adopted. Multi-scale features are extracted through a gated residual enhancement module and a dilated convolution fusion module. Combined with multimodal dynamic attention fusion and cross-layer connections, data distribution drift is monitored in real time and adaptive updates are triggered. Lightweight pruning and quantization operations are performed, and the system is deployed on edge computing devices.
The ability to characterize early micro-damage signals has been significantly enhanced, and real-time fault diagnosis with low latency and low power consumption has been achieved. The model can quickly adapt when the distribution drifts. After lightweighting, the number of parameters is reduced by 58.4% and the amount of calculation is reduced by 64.9%, meeting the real-time operation requirements of industrial edge devices.
Smart Images

Figure CN120670942A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault diagnosis, and in particular to a fault information adaptive diagnosis method based on a deep residual network. Background Art
[0002] With the continuous development of artificial intelligence and intelligent manufacturing technologies, industrial equipment operating status monitoring and fault diagnosis are gradually evolving from traditional manual inspections and rule-based threshold judgments to data-driven intelligent diagnostic models. In key industrial scenarios such as high-end manufacturing, energy equipment, and rail transit, the timeliness and accuracy of fault identification directly impact equipment safety and production efficiency. In recent years, convolutional neural networks, long-short-term memory networks, and ensemble learning methods have been widely applied to industrial signal processing and fault diagnosis tasks, achieving considerable success.
[0003] Early-stage faults in industrial equipment, such as those occurring during the initial stages of operation or minor damage, have weak signal energy and subtle features that are easily obscured by background noise. Existing modeling approaches based on traditional CNNs or standard deep networks struggle to effectively represent these weak signal characteristics. Furthermore, as the number of network layers increases, gradient degradation and feature attenuation are more likely to occur, leading to micro-faults being overlooked and low diagnostic sensitivity and accuracy.
[0004] Industrial field environments are complex and ever-changing. Sensor data characteristics of the same type of equipment can vary significantly across different workshops, operating times, or geographic locations. Traditional approaches face bottlenecks such as low model transfer efficiency, high retraining costs, and long offline update cycles, making them difficult to meet the production system's demand for non-stop, real-time diagnostics. Even some approaches that incorporate transfer learning or fine-tuning strategies often face the risk of insufficient pseudo-label quality control and model degradation caused by accumulated pseudo-label noise, limiting the practical application of online adaptive capabilities.
[0005] Therefore, it is urgent to propose a deep diagnosis method that has efficient feature capture capabilities, supports cross-domain adaptation, and can be embedded in edge operations to solve the above problems. Summary of the Invention
[0006] One purpose of the present invention is to propose a method for adaptive fault information diagnosis based on a deep residual network. The present invention provides a solid theoretical basis and a feasible engineering implementation path for a highly robust, highly adaptable, and embeddable industrial equipment fault diagnosis system.
[0007] A method for adaptive fault information diagnosis based on a deep residual network according to an embodiment of the present invention includes the following steps:
[0008] S1. Collect multimodal operation signal data of industrial equipment during operation, generate an original multimodal operation signal dataset, perform data preprocessing on the original multimodal operation signal dataset, and obtain a standardized multimodal operation tensor;
[0009] S2. Input the normalized multimodal operation tensor into a deep residual network consisting of a gated residual enhancement module and a dilated convolutional fusion module, and output a multi-scale fault feature map.
[0010] S3. Perform multimodal dynamic attention fusion based on multi-scale fault feature maps, and generate fused feature tensors using channel attention and cross-layer connection mechanism;
[0011] S4. Input the fused feature tensor into the fault classification head, establish an initial fault classification model and output the fault identification result;
[0012] S5. Monitor the difference between the distribution of the standardized multimodal running tensor and the initial training distribution during real-time inference, calculate the distribution drift index and compare it with the preset drift threshold. When the distribution drift index exceeds the drift threshold, trigger the online adaptive update process, use the cross-domain meta-tuning mechanism to quickly fine-tune the tail gated residual enhancement module of the deep residual network, and use the incremental pseudo-label distillation strategy to update the pseudo-label memory and model parameters to generate an adaptive optimization model.
[0013] S6. Perform pruning and quantization operations on the adaptive optimization model to generate a lightweight fault diagnosis model, and deploy the lightweight fault diagnosis model on the edge computing device to achieve real-time fault diagnosis;
[0014] S7. Output the final fault identification results and equipment health status assessment report based on the lightweight fault diagnosis model.
[0015] Optionally, the multimodal operation signal data includes vibration signals, acoustic signals, current signals and temperature signals, and the data preprocessing includes performing normalization, time-frequency domain enhancement and noise filtering operations.
[0016] Optionally, the S2 includes the following steps:
[0017] S21. Normalize the multimodal operation tensor X input The input is fed into a deep residual network, which consists of multiple stacked gated residual enhancement modules and dilated convolution fusion modules.
[0018] S22. The gated residual enhancement module receives the input tensor Where C represents the number of channels, H represents the tensor height, and W represents the tensor width. The convolutional mapping F is performed by the main branch. conv (X input ), and through the gating weight vector The main branch output is weighted by channel to obtain the channel enhanced feature tensor:
[0019] X gate =G⊙F conv (X input );
[0020] Among them, ⊙ represents the channel-dimensional element-by-element multiplication operation, represents the gated enhanced feature tensor;
[0021] S23. The gate weight vector G is obtained through the channel attention module:
[0022] G=σ(W2·δ(W1·GAP(X input )));
[0023] Among them, GAP(·) represents the global average pooling operation, and the output channel average vector is the learnable weight matrix, r is the compression rate hyperparameter, δ(·) is the ReLU activation function, and σ(·) is the Sigmoid activation function;
[0024] S24. Gated enhanced feature tensor X gate With the input tensor X input Perform residual connection to generate enhanced residual feature tensor;
[0025] S25. Enhance the residual feature tensor X res Input to the dilated convolution fusion module to perform multi-scale receptive field expansion. The dilated convolution operation is performed with different expansion rates d i Extract local and global features respectively, and output multi-scale feature tensors:
[0026]
[0027] in, The expansion rate is d i The dilated convolution operation, represents the feature dimension splicing operation, N is the number of multi-scale paths;
[0028] S26. Multi-scale feature tensor X dilated Input to the fusion convolution module, perform feature compression and nonlinear transformation, and output the final multi-scale fault feature map X fault .
[0029] Optionally, S3 includes the following steps:
[0030] S31. Receive the multi-scale fault feature map output by the dilated convolution fusion module Among them C ″ is the number of channels;
[0031] S32. Multi-scale fault feature map X fault Group by modal source and set the modal set to M = {m1,m2,...,m K}, where K represents the number of modes, corresponding to vibration signal, acoustic signal, current signal and temperature signal respectively. fault Split into modal feature sub-tensors
[0032] S33. For each modal feature sub-tensor Calculate its channel attention vector The channel attention vector is obtained through the modality-specific channel attention module:
[0033]
[0034] Among them, GAP(·) represents the global average pooling operation, and is the mode m k The attention weight matrix, δ(·) is the ReLU activation function, and σ(·) is the Sigmoid activation function;
[0035] S34. Using channel attention vector Perform weighted enhancement on the corresponding modal feature sub-tensor to obtain the modal enhancement tensor
[0036] S35. All modal enhancement tensors are concatenated in the channel dimension to form the modal fusion feature tensor X concat ;
[0037] S36. Fusion modality feature tensor X concat Input to the cross-layer connection module, which receives the shallow residual feature tensor X shallow ;
[0038] S37. Fusion feature tensor X fusion Serves as input to the fault classification submodule of the deep residual network.
[0039] Optionally, the S4 includes the following steps:
[0040] S41. Fusion feature tensor X fusion Perform spatial dimension compression operation and use global average pooling function GAP(·) to map it into fusion channel vector
[0041]
[0042] Among them, z fusion[c] represents the global average value of the fused feature tensor on the cth channel, reflecting the overall activation degree of the channel in the spatial dimension;
[0043] S42. Fusion channel vector z fusion Input fault classification fully connected network F cls (·), build the initial fault classification model and output the fault prediction vector
[0044] Among them, N c Indicates the total number of fault categories, is the fault classification weight matrix, is the bias vector, Softmax(·) is the normalized activation function used to convert the network output into the predicted probability of each fault category;
[0045] S43. Based on the fault prediction vector The category label corresponding to the maximum probability determines the initial fault type corresponding to the current input tensor, which is recorded as the predicted label y pred .
[0046] Optionally, the S5 includes the following steps:
[0047] S51. During the model inference phase, normalize the multimodal tensor for each frame Perform distribution coding and use global average pooling operation to compress its channel dimension to obtain the feature distribution vector of the current input frame
[0048] S52. The current input feature distribution vector and the reference distribution vector z extracted in the initial training phase ref For distribution comparison, Kullback-Leibler divergence is used as the distribution drift indicator
[0049] S53. Distribution Drift Index With the set drift threshold Compare, when satisfied It is determined that the characteristic distribution drift has occurred in the current equipment working condition, triggering the adaptive update mechanism;
[0050] S54. Start the cross-domain meta-tuning module to adjust the weight parameter Θ in the gated residual enhancement module at the end of the deep residual network tail Perform local quick fine-tuning:
[0051]
[0052] Among them, η meta is the meta-tuned learning rate, L supportis the support set loss function, are the tail model parameters after fine-tuning;
[0053] S55. Parallel execution of the incremental pseudo-label distillation mechanism, using the output confidence vector of the current inference sample As soft label input;
[0054] S56. Based on the distilled pseudo-label Build sample memory library M pseudo , using the confidence threshold γ conf ∈[0,1] to filter the pseudo labels. Then keep the current pseudo label to M pseudo and use it to update model parameters;
[0055] S57. Based on the current pseudo-label memory library M pseudo and fine-tuned model parameters Perform cross-distillation optimization and finally generate an optimized model parameter set Θ with distribution adaptive capability adaptive , and replace the tail parameters in the original fault classification model to update the model structure and form an adaptive optimization model.
[0056] Optionally, the S6 includes the following steps:
[0057] S61. Based on the adaptive optimization model parameter set Θ adaptive ,construct a lightweight fault diagnosis model for deployment on edge computing devices.,The lightweight process includes structure pruning operations and parameter quantization operations;
[0058] S62. Prune and evaluate the channel weights of each layer of parameters in the deep residual network, using the channel importance index I l,c Sort each channel c in each layer l, channel importance index:
[0059]
[0060] Among them, W l [c,i,j] represents the convolution kernel weight corresponding to the cth channel in the lth layer, H and W are the height and width of the convolution kernel respectively, I l,c Represents the average absolute weight of channel c, which is used to measure its contribution to feature extraction;
[0061] S63. Set the pruning ratio threshold α p ∈(0,1), the α ranked at the bottom of each layer p ×C l Channel clipping, where C l is the total number of original channels in layer l, generating the pruned network structure N pruned, the corresponding weight set is recorded as Θ pruned ;
[0062] S64. Model parameters Θ after pruning pruned Perform parameter quantization operation, introduce fixed bit width quantization mapping function Q(·), and transform the full precision weight tensor θ∈Θ pruned Mapping to fixed-point representation:
[0063]
[0064] Among them, θ q represents the quantized fixed-point weight, θ min and θ max Represent the minimum and maximum weight values in the current tensor respectively, b is the quantization bit width, Δ is the quantization step size, and the mapping process keeps the interval ratio of the original weight unchanged;
[0065] S65. Quantize the weights θ of all network layers q Replace the original floating-point parameters and construct the final lightweight fault diagnosis model N light , whose parameter set is denoted as Θ light ={θ q};
[0066] S66. Lightweight fault diagnosis model N light Perform edge deployment adaptation and generate model files suitable for embedded inference hardware. The model files contain parameter compression representation, compiled intermediate format, and execution graph structure description, supporting low-latency loading of heterogeneous edge chips.
[0067] S67. Lightweight fault diagnosis model N light Deployed on edge computing devices with limited computing resources, the model is loaded and initialized using the edge inference engine, and the online inference batch size is set to B = 1 to support single-sample real-time fault identification.
[0068] S68. Input normalized multimodal operation tensor Lightweight fault diagnosis model running on the edge light , output the fault prediction result at the current time t Complete real-time fault diagnosis tasks at the edge with low latency and low power consumption.
[0069] The beneficial effects of the present invention are:
[0070] Based on the deep residual network, this paper designs a channel-level gated residual enhancement module, which dynamically adjusts the main branch channel information weight through the gating factor, significantly enhancing the characterization ability of early micro-damage and micro-crack signals. It combines the void convolution fusion module to expand the receptive field, realizes multi-scale feature aggregation without increasing the number of parameters, and maintains the transmission integrity of high-frequency details through cross-layer connections, solving the problems of feature attenuation and gradient degradation of traditional residual structures in weak signal backgrounds.
[0071] During the diagnostic model inference process, the present invention monitors the KL divergence difference between the input data distribution and the initial training set distribution in real time, and uses this to drive the cross-domain meta-learning mechanism, and only quickly fine-tunes the gated residual enhancement module at the end of the network to ensure low-cost adaptation; at the same time, a confidence-driven pseudo-label distillation strategy is proposed, which adjusts the pseudo-label contribution through a dynamic temperature coefficient, effectively alleviating the problem of misleading small sample pseudo-labels.
[0072] To meet the real-time and lightweight requirements of diagnostic models in industrial sites, this paper, after obtaining a stable optimization model, performs channel importance-based structural pruning and fixed-point quantization operations to construct a low-latency diagnostic model. The pruning strategy ranks channels by importance based on their average absolute weights, and combines quantization step size control to achieve 8-bit fixed-point mapping while retaining key feature response paths. The resulting lightweight model has an average inference latency of less than 18ms, a 58.4% reduction in parameters, and a 64.9% decrease in computational effort. This significantly outperforms traditional uncompressed model deployment strategies and meets the real-time operation requirements of industrial edge devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0074] Figure 1 This is a flowchart of a fault information adaptive diagnosis method based on deep residual network proposed by the present invention. DETAILED DESCRIPTION
[0075] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0076] refer to Figure 1 , a fault information adaptive diagnosis method based on deep residual network, comprising the following steps:
[0077] S1. Collect multimodal operation signal data of industrial equipment during operation, generate an original multimodal operation signal dataset, perform data preprocessing on the original multimodal operation signal dataset, and obtain a standardized multimodal operation tensor;
[0078] S2. Input the normalized multimodal operation tensor into a deep residual network consisting of a gated residual enhancement module and a dilated convolutional fusion module, and output a multi-scale fault feature map.
[0079] S3. Perform multimodal dynamic attention fusion based on multi-scale fault feature maps, and generate fused feature tensors using channel attention and cross-layer connection mechanism;
[0080] S4. Input the fused feature tensor into the fault classification head, establish an initial fault classification model and output the fault identification result;
[0081] S5. Monitor the difference between the distribution of the standardized multimodal running tensor and the initial training distribution during real-time inference, calculate the distribution drift index and compare it with the preset drift threshold. When the distribution drift index exceeds the drift threshold, trigger the online adaptive update process, use the cross-domain meta-tuning mechanism to quickly fine-tune the tail gated residual enhancement module of the deep residual network, and use the incremental pseudo-label distillation strategy to update the pseudo-label memory and model parameters to generate an adaptive optimization model.
[0082] S6. Perform pruning and quantization operations on the adaptive optimization model to generate a lightweight fault diagnosis model, and deploy the lightweight fault diagnosis model on the edge computing device to achieve real-time fault diagnosis;
[0083] S7. Output the final fault identification results and equipment health status assessment report based on the lightweight fault diagnosis model.
[0084] In this embodiment, the multimodal operation signal data includes vibration signals, acoustic signals, current signals, and temperature signals, and the data preprocessing includes performing normalization, time-frequency domain enhancement, and noise filtering operations.
[0085] In this embodiment, S2 includes the following steps:
[0086] S21. Normalize the multimodal operation tensor X input The input is fed into a deep residual network, which consists of multiple stacked gated residual enhancement modules and dilated convolution fusion modules. The gated residual enhancement module is used to alleviate the vanishing gradient problem in multi-layer network structures and enhance the channel response of high-frequency details in scenarios where fault characteristics are weakly expressed.
[0087] S22. The gated residual enhancement module receives the input tensor Where C represents the number of channels, H represents the tensor height, and W represents the tensor width. The convolutional mapping F is performed by the main branch. conv (X input ), and through the gating weight vector The main branch output is weighted by channel to obtain the channel enhanced feature tensor:
[0088] X gate =G⊙F conv (X input );
[0089] Among them, ⊙ represents the channel-dimensional element-by-element multiplication operation, represents the gated enhanced feature tensor;
[0090] S23. The gate weight vector G is obtained through the channel attention module:
[0091] G=σ(W2·δ(W1·GAP(X input )));
[0092] Among them, GAP(·) represents the global average pooling operation, and the output channel average vector is the learnable weight matrix, r is the compression rate hyperparameter, δ(·) is the ReLU activation function, σ(·) is the Sigmoid activation function, and the gated weight vector G reflects the response strength of different channels to weak fault features;
[0093] S24. Gated enhanced feature tensor X gate With the input tensor X input Perform residual connection to generate enhanced residual feature tensor and enhance residual feature tensor It is used to explicitly enhance high-frequency weak fault information while maintaining the stability of gradient propagation;
[0094] S25. Enhance the residual feature tensor X res Input to the dilated convolution fusion module to perform multi-scale receptive field expansion. The dilated convolution operation is performed with different expansion rates d i Extract local and global features respectively, and output multi-scale feature tensors:
[0095]
[0096] in, The expansion rate is d i The dilated convolution operation, represents the feature dimension splicing operation, N is the number of multi-scale paths;
[0097] S26. Multi-scale feature tensor Input to the fusion convolution module, perform feature compression and nonlinear transformation, and output the final multi-scale fault feature map Among them C ′ with C ″ Represent the number of channels after splicing and compression respectively.
[0098] In this embodiment, S3 includes the following steps:
[0099] S31. Receive the multi-scale fault feature map output by the dilated convolution fusion module Among them C ″ is the number of channels, H is the feature map height, and W is the feature map width, which serves as the input tensor for multimodal dynamic attention fusion;
[0100] S32. Multi-scale fault feature map X fault Group by modal source and set the modal set to M = {m1,m2,...,m K}, where K represents the number of modes, corresponding to vibration signal, acoustic signal, current signal and temperature signal respectively. fault Split into modal feature sub-tensors
[0101] S33. For each modal feature sub-tensor Calculate its channel attention vector The channel attention vector is obtained through the modality-specific channel attention module:
[0102]
[0103] Among them, GAP(·) represents the global average pooling operation, and is the mode m k The attention weight matrix, δ(·) is the ReLU activation function, and σ(·) is the Sigmoid activation function;
[0104] S34. Using channel attention vector Perform weighted enhancement on the corresponding modal feature sub-tensor to obtain the modal enhancement tensor
[0105] S35. Concatenate all modal enhancement tensors in the channel dimension to form a modal fusion feature tensor
[0106] S36. Fusion modality feature tensor X concat Input to the cross-layer connection module, which receives the shallow residual feature tensor where X shallowThe output from the early stage of the gated residual enhancement module is used to compensate for the degradation of high-frequency information in the deep representation and perform cross-layer connection operations:
[0107] X fusion =φ(X concat ,X shallow );
[0108] Among them, φ(·) represents the feature fusion function, which uses the weighted summation or convolution fusion method after channel matching to output the fused feature tensor C f is the number of channels after fusion;
[0109] S37. Fusion feature tensor X fusion As the input of the fault classification submodule of the deep residual network, it is used to generate the final representation input of the fault recognition model.
[0110] In this embodiment, S4 includes the following steps:
[0111] S41. Receive the fused feature tensor output by the multimodal dynamic attention fusion module Among them C f is the number of channels after fusion;
[0112] S42. Fusion feature tensor X fusion Perform spatial dimension compression operation and use global average pooling function GAP(·) to map it into fusion channel vector
[0113]
[0114] Among them, z fusion [c] represents the global average value of the fused feature tensor on the cth channel, reflecting the overall activation degree of the channel in the spatial dimension;
[0115] S43. Fusion channel vector z fusion Input fault classification fully connected network F cls (·), build the initial fault classification model and output the fault prediction vector
[0116]
[0117] Among them, N c Indicates the total number of fault categories, is the fault classification weight matrix, is the bias vector, Softmax(·) is the normalized activation function used to convert the network output into the predicted probability of each fault category;
[0118] S44. Based on the fault prediction vector The category label corresponding to the maximum probability determines the initial fault type corresponding to the current input tensor, which is recorded as the predicted label y pred :
[0119]
[0120] Among them, y pred The final output is the initial fault identification result, which is used to determine the fault category corresponding to the current operating status of the industrial equipment.
[0121] In this embodiment, S5 includes the following steps:
[0122] S51. During the model inference phase, normalize the multimodal tensor for each frame Perform distribution coding and use global average pooling operation to compress its channel dimension to obtain the feature distribution vector of the current input frame
[0123] S52. The current input feature distribution vector and the reference distribution vector extracted in the initial training phase For distribution comparison, Kullback-Leibler divergence is used as the distribution drift indicator
[0124]
[0125] in, Measures the relative entropy between the current inference sample and the training sample in terms of channel feature expression, reflecting whether the multimodal features drift;
[0126] S53. Distribution Drift Index With the set drift threshold Compare, when satisfied It is determined that the characteristic distribution drift has occurred in the current equipment working condition, triggering the adaptive update mechanism;
[0127] S54. Start the cross-domain meta-tuning module to adjust the weight parameter Θ in the gated residual enhancement module at the end of the deep residual network tail Perform local fast fine-tuning. Fine-tuning adopts a nested optimization strategy based on a small number of samples. The optimization process is expressed as:
[0128]
[0129] Among them, η meta is the meta-tuned learning rate, L support is the support set loss function, are the tail model parameters after fine-tuning;
[0130] S55. Parallel execution of the incremental pseudo-label distillation mechanism, using the output confidence vector of the current inference sample As soft label input, the pseudo label confidence temperature adjustment function is defined as:
[0131]
[0132] Among them, T>1 is the distillation temperature coefficient, which is used to smooth the pseudo-label distribution. Adjust the pseudo-label for the temperature of the current sample to avoid overfitting the model to local anomalies;
[0133] S56. Based on the distilled pseudo-label Build sample memory library M pseudo , using the confidence threshold γ conf ∈[0,1] to filter the pseudo labels. Then keep the current pseudo label to M pseudo and use it to update model parameters;
[0134] S57. Based on the current pseudo-label memory library M pseudo and fine-tuned model parameters Perform cross-distillation optimization and finally generate an optimized model parameter set Θ with distribution adaptive capability adaptive , and replace the tail parameters in the original fault classification model to update the model structure and form an adaptive optimization model.
[0135] In this embodiment, S6 includes the following steps:
[0136] S61. Based on the adaptive optimization model parameter set Θ adaptive ,construct a lightweight fault diagnosis model for deployment on edge computing devices.,The lightweight process includes structure pruning operations and parameter quantization operations;
[0137] S62. Prune and evaluate the channel weights of each layer of parameters in the deep residual network, using the channel importance index I l,c Sort each channel c in each layer l, channel importance index:
[0138]
[0139] Among them, W l [c,i,j] represents the convolution kernel weight corresponding to the cth channel in the lth layer, H and W are the height and width of the convolution kernel respectively, I l,c Represents the average absolute weight of channel c, which is used to measure its contribution to feature extraction;
[0140] S63. Set the pruning ratio threshold α p∈(0,1), the α ranked at the bottom of each layer p ×C l Channel clipping, where C l is the total number of original channels in layer l, generating the pruned network structure N pruned , the corresponding weight set is recorded as Θ pruned ;
[0141] S64. Model parameters Θ after pruning pruned Perform parameter quantization operation, introduce fixed bit width quantization mapping function Q(·), and transform the full precision weight tensor θ∈Θ pruned Mapping to fixed-point representation:
[0142]
[0143] Among them, θ q represents the quantized fixed-point weight, θ min and θ max Represent the minimum and maximum weight values in the current tensor respectively, b is the quantization bit width, Δ is the quantization step size, and the mapping process keeps the interval ratio of the original weight unchanged;
[0144] S65. Quantize the weights θ of all network layers q Replace the original floating-point parameters and construct the final lightweight fault diagnosis model N light , whose parameter set is denoted as Θ light ={θ q};
[0145] S66. Lightweight fault diagnosis model N light Perform edge deployment adaptation and generate model files suitable for embedded inference hardware. The model files contain parameter compression representation, compiled intermediate format, and execution graph structure description, supporting low-latency loading of heterogeneous edge chips.
[0146] S67. Lightweight fault diagnosis model N light Deployed on edge computing devices with limited computing resources, the model is loaded and initialized using the edge inference engine, and the online inference batch size is set to B = 1 to support single-sample real-time fault identification.
[0147] S68. Input normalized multimodal operation tensor Lightweight fault diagnosis model running on the edge light , output the fault prediction result at the current time t Complete real-time fault diagnosis tasks at the edge with low latency and low power consumption.
[0148] Example 1: This example takes the intelligent diagnosis of early faults of the main bearings of wind turbines in a large-scale wind farm operation and maintenance center in a certain province as the scenario, and comprehensively demonstrates the application process, effects and technical advantages of the fault information adaptive diagnosis method based on the deep residual network of the present invention in actual engineering environments.
[0149] From May to August 2024, the operating conditions of multiple wind turbines at the Delta Wind Farm fluctuated for three consecutive months. Operation and maintenance engineers suspected that the main bearings of some units had early fatigue damage. However, due to the complex operating environment of the units, large changes in wind speed, temperature and humidity, and grid load, both traditional rule-based diagnosis and classic convolutional neural network diagnosis performed poorly, with frequent false alarms and missed alarms, and were unable to detect minor anomalies in time, resulting in some units being passively shut down for maintenance, causing direct economic losses.
[0150] To monitor the main bearing status, the wind farm collects data from four types of multimodal operating signals: vibration (acceleration and velocity), current (stator and rotor phases), acoustics (structure-borne noise and airborne noise), and main shaft temperature. The sampling frequency is 10kHz, and approximately 35GB of data is collected daily for each turbine. The collected data includes 20 turbines, of which A01, A03, and A07 are known to have early-stage damage, while A09, A14, and A18 are healthy.
[0151] To ensure the generalization ability of the model, 1,200 pieces of normal operation data and early damage data from May to June 2024 (each with a 10-second time window) were selected, totaling 2,400 training samples, and 2,000 pieces of unlabeled data under new working conditions in July 2024 were used as adaptive update test samples.
[0152] First, the raw multimodal signal data is synchronized through a unified acquisition platform. The vibration signal is pre-processed using pre-emphasis and high-pass filtering. The current signal is normalized to the [-1, 1] interval. The acoustic signal uses the STFT transform to extract time-frequency features. The spindle temperature signal uses a sliding window average to eliminate outliers. All data is then uniformly tensorized to form a standardized multimodal operation tensor, which is then fed into the proposed gated residual enhancement-atrous convolution fusion deep residual network.
[0153] In the feature extraction stage, the gated residual enhancement module of the present invention dynamically weights the features of each channel to effectively retain the micro-features of weak faults; the dilated convolution module realizes multi-scale feature extraction of tiny periodic shocks and broadband noise by setting the expansion rate to 3 / 5 / 7, and the cross-layer connection mechanism enhances the expression ability of high-frequency information.
[0154] Subsequently, the multimodal tensor is fused with channel attention and cross-layer residual features, and then fed into the classification head. After global average pooling and a fully connected layer, it outputs three prediction results (healthy, early injury, and severe injury). In early June 2024, the model was first deployed on three wind turbine edge operation and maintenance AI boxes, A01, A03, and A07. The measured inference latency was 15ms, and the average CPU utilization per machine was only 26%.
[0155] In July, after experiencing rapid fluctuations in ambient temperature and humidity, and peak-valley shifts in power grid dispatch, which caused signal distribution drift, the health false alarm rates of traditional ResNet and SVM models rose to 18.1% and 26.2%, respectively, while the false negative rates were 7.8% and 11.4%, respectively. The proposed method monitors the KL divergence between the input feature distribution and the initial training distribution in real time, automatically triggering an adaptive update process. Leveraging inference pseudo-labels from unlabeled new operating condition data, and employing a combined optimization approach of meta-tuning and incremental distillation, this method achieves rapid edge fine-tuning in just four minutes. After the model was adapted to the new operating conditions, the false positive rate dropped to 3.4% and the false negative rate to 1.9%.
[0156] Table 1 Comparative experimental results
[0157]
[0158] Table 2 Comparison of specific samples (10 representative samples from actual operation)
[0159]
[0160] Among them, samples 021, 143, and 603 were confirmed to be early-stage main bearing spalling after on-site manual inspection. Traditional models failed to provide timely warnings. However, the proposed method had inference probabilities of only 0.19, 0.28, and 0.24 for healthy samples, and 0.77, 0.68, and 0.73 for early-stage damage, directly providing early damage warnings in real-world, unlabeled samples. For healthy samples 089, 312, 507, and 812, the proposed method achieved a 0% false positive rate, while both SVM and traditional ResNet models produced 2-3 false positives for early-stage damage.
[0161] In terms of model deployment, the proposed method, after structural pruning and 8-bit quantization, reduces the model file size to just 14MB. It successfully runs in real time on an on-site ARM-based AI box (Cortex-A72, 1.8GHz, 2GB of memory) at a wind farm, achieving a stable single-frame inference latency of 15-17ms, significantly lower than both SVM and traditional ResNet. Under these new operating conditions, adaptive pseudo-label distillation enables rapid local fine-tuning with only 100 unlabeled data points, eliminating the need for equipment downtime and manual labeling. This truly achieves a closed-loop intelligent diagnostic system with zero downtime and low maintenance costs.
[0162] In summary, this embodiment, in a typical wind turbine health management scenario, achieves sensitive identification of weak faults, adaptive model migration, and lightweight edge real-time deployment through adaptive deep feature extraction and online optimization of distribution drift from multimodal complex signals. Extensive comparative data demonstrates that this invention significantly improves upon traditional methods in terms of early fault recognition rate, adaptability, latency, and resource consumption. This effectively addresses challenges faced by traditional methods in real-world industrial scenarios, such as insensitive early diagnosis, difficulty in model redeployment, and poor real-time performance.
[0163] Based on the deep residual network, this paper designs a channel-level gated residual enhancement module, which dynamically adjusts the main branch channel information weight through the gating factor, significantly enhancing the characterization ability of early micro-damage and micro-crack signals. It combines the void convolution fusion module to expand the receptive field, realizes multi-scale feature aggregation without increasing the number of parameters, and maintains the transmission integrity of high-frequency details through cross-layer connections, solving the problems of feature attenuation and gradient degradation of traditional residual structures in weak signal backgrounds.
[0164] During the diagnostic model inference process, the present invention monitors the KL divergence difference between the input data distribution and the initial training set distribution in real time, and uses this to drive the cross-domain meta-learning mechanism, and only quickly fine-tunes the gated residual enhancement module at the end of the network to ensure low-cost adaptation; at the same time, a confidence-driven pseudo-label distillation strategy is proposed, which adjusts the pseudo-label contribution through a dynamic temperature coefficient, effectively alleviating the problem of misleading small sample pseudo-labels.
[0165] To meet the real-time and lightweight requirements of diagnostic models in industrial sites, this paper, after obtaining a stable optimization model, performs channel importance-based structural pruning and fixed-point quantization operations to construct a low-latency diagnostic model. The pruning strategy ranks channels by importance based on their average absolute weights, and combines quantization step size control to achieve 8-bit fixed-point mapping while retaining key feature response paths. The resulting lightweight model has an average inference latency of less than 18ms, a 58.4% reduction in parameters, and a 64.9% decrease in computational effort. This significantly outperforms traditional uncompressed model deployment strategies and meets the real-time operation requirements of industrial edge devices.
[0166] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A fault information adaptive diagnosis method based on deep residual network, characterized in that: The steps include: S1. Collect multimodal operation signal data of industrial equipment during operation, generate an original multimodal operation signal dataset, perform data preprocessing on the original multimodal operation signal dataset, and obtain a standardized multimodal operation tensor; S2. Input the normalized multimodal operation tensor into a deep residual network consisting of a gated residual enhancement module and a dilated convolutional fusion module, and output a multi-scale fault feature map. S3. Perform multimodal dynamic attention fusion based on multi-scale fault feature maps, and generate fused feature tensors using channel attention and cross-layer connection mechanism; S4. Input the fused feature tensor into the fault classification head, establish an initial fault classification model and output the fault identification result; S5. Monitor the difference between the distribution of the standardized multimodal running tensor and the initial training distribution during real-time inference, calculate the distribution drift index and compare it with the preset drift threshold. When the distribution drift index exceeds the drift threshold, trigger the online adaptive update process, use the cross-domain meta-tuning mechanism to quickly fine-tune the tail gated residual enhancement module of the deep residual network, and use the incremental pseudo-label distillation strategy to update the pseudo-label memory and model parameters to generate an adaptive optimization model. S6. Perform pruning and quantization operations on the adaptive optimization model to generate a lightweight fault diagnosis model, and deploy the lightweight fault diagnosis model on the edge computing device to achieve real-time fault diagnosis; S7. Output the final fault identification results and equipment health status assessment report based on the lightweight fault diagnosis model.
2. The method for adaptive fault information diagnosis based on deep residual network according to claim 1, characterized in that: The multimodal operation signal data includes a vibration signal, an acoustic signal, a current signal, and a temperature signal, and the data preprocessing includes performing normalization, time-frequency domain enhancement, and noise filtering operations.
3. The method for adaptive fault information diagnosis based on deep residual network according to claim 1, characterized in that: The S2 comprises the following steps: S21. Normalize the multimodal operation tensor X input The input is fed into a deep residual network, which consists of multiple stacked gated residual enhancement modules and dilated convolution fusion modules. S22. The gated residual enhancement module receives the input tensor Where C represents the number of channels, H represents the tensor height, and W represents the tensor width. The convolutional mapping F is performed by the main branch. conv (X input ), and through the gating weight vector The main branch output is weighted by channel to obtain the channel enhanced feature tensor: X gate =G⊙F conv (X input ); Among them, ⊙ represents the channel-dimensional element-by-element multiplication operation, represents the gated enhanced feature tensor; S23. The gate weight vector G is obtained through the channel attention module: G=σ(W2·δ(W1·GAP(X input ))); Among them, GAP(·) represents the global average pooling operation, and the output channel average vector is the learnable weight matrix, r is the compression rate hyperparameter, δ(·) is the ReLU activation function, and σ(·) is the Sigmoid activation function; S24. Gated enhanced feature tensor X gate With the input tensor X input Perform residual connection to generate enhanced residual feature tensor; S25. Enhance the residual feature tensor X res Input to the dilated convolution fusion module to perform multi-scale receptive field expansion. The dilated convolution operation is performed with different expansion rates d i Extract local and global features respectively, and output multi-scale feature tensors: in, The expansion rate is d i The dilated convolution operation, represents the feature dimension splicing operation, N is the number of multi-scale paths; S26. Multi-scale feature tensor X dilated Input to the fusion convolution module, perform feature compression and nonlinear transformation, and output the final multi-scale fault feature map X fault .
4. The method for adaptive fault information diagnosis based on deep residual network according to claim 3, characterized in that: The S3 includes the following steps: S31. Receive the multi-scale fault feature map output by the dilated convolution fusion module Among them C ″ is the number of channels; S32. Multi-scale fault feature map X fault Group by modal source and set the modal set to M = {m1,m2,...,m K }, where K represents the number of modes, corresponding to vibration signal, acoustic signal, current signal and temperature signal respectively. fault Split into modal feature sub-tensors S33. For each modal feature sub-tensor Calculate its channel attention vector The channel attention vector is obtained through the modality-specific channel attention module: Among them, GAP(·) represents the global average pooling operation, and is the mode m k The attention weight matrix, δ(·) is the ReLU activation function, and σ(·) is the Sigmoid activation function; S34. Using channel attention vector Perform weighted enhancement on the corresponding modal feature sub-tensor to obtain the modal enhancement tensor S35. All modal enhancement tensors are concatenated in the channel dimension to form the modal fusion feature tensor X concat ; S36. Fusion modality feature tensor X concat Input to the cross-layer connection module, which receives the shallow residual feature tensor X shallow ; S37. Fusion feature tensor X fusion Serves as input to the fault classification submodule of the deep residual network.
5. The method for adaptive fault information diagnosis based on deep residual network according to claim 1, characterized in that: The S4 comprises the following steps: S41. Fusion feature tensor X fusion Perform spatial dimension compression operation and use global average pooling function GAP(·) to map it into fusion channel vector Among them, z fusion [c] represents the global average value of the fused feature tensor on the cth channel, reflecting the overall activation degree of the channel in the spatial dimension; S42. Fusion channel vector z fusion Input fault classification fully connected network F cls (·), build the initial fault classification model and output the fault prediction vector Among them, N c Indicates the total number of fault categories, is the fault classification weight matrix, is the bias vector, Softmax(·) is the normalized activation function used to convert the network output into the predicted probability of each fault category; S43. Based on the fault prediction vector The category label corresponding to the maximum probability determines the initial fault type corresponding to the current input tensor, which is recorded as the predicted label y pred .
6. The method for adaptive fault information diagnosis based on deep residual network according to claim 5, characterized in that: The S5 comprises the following steps: S51. During the model inference phase, normalize the multimodal tensor for each frame Perform distribution coding and use global average pooling operation to compress its channel dimension to obtain the feature distribution vector of the current input frame S52. The current input feature distribution vector and the reference distribution vector z extracted in the initial training phase ref For distribution comparison, Kullback-Leibler divergence is used as the distribution drift indicator S53. Distribution Drift Index With the set drift threshold Compare, when satisfied It is determined that the characteristic distribution drift has occurred in the current equipment working condition, triggering the adaptive update mechanism; S54. Start the cross-domain meta-tuning module to adjust the weight parameter Θ in the gated residual enhancement module at the end of the deep residual network tail Perform local quick fine-tuning: Among them, η meta is the meta-tuned learning rate, L support is the support set loss function, are the tail model parameters after fine-tuning; S55. Parallel execution of the incremental pseudo-label distillation mechanism, using the output confidence vector of the current inference sample As soft label input; S56. Based on the distilled pseudo-label Build sample memory library M pseudo , using the confidence threshold γ conf ∈[0,1] to filter the pseudo labels. Then keep the current pseudo label to M pseudo and use it to update model parameters; S57. Based on the current pseudo-label memory library M pseudo and fine-tuned model parameters Perform cross-distillation optimization and finally generate an optimized model parameter set Θ with distribution adaptive capability adaptive , and replace the tail parameters in the original fault classification model to update the model structure and form an adaptive optimization model.
7. The method for adaptive fault information diagnosis based on deep residual network according to claim 6, characterized in that: The S6 comprises the following steps: S61. Based on the adaptive optimization model parameter set Θ adaptive ,construct a lightweight fault diagnosis model for deployment on edge computing devices.,The lightweight process includes structure pruning operations and parameter quantization operations; S62. Prune and evaluate the channel weights of each layer of parameters in the deep residual network, using the channel importance index I l,c Sort each channel c in each layer l, channel importance index: Among them, W l [c,i,j] represents the convolution kernel weight corresponding to the cth channel in the lth layer, H and W are the height and width of the convolution kernel respectively, I l,c Represents the average absolute weight of channel c, which is used to measure its contribution to feature extraction; S63. Set the pruning ratio threshold α p ∈(0,1), the α ranked at the bottom of each layer p ×C l Channel clipping, where C l is the total number of original channels in layer l, generating the pruned network structure N pruned , the corresponding weight set is recorded as Θ pruned ; S64. Model parameters Θ after pruning pruned Perform parameter quantization operation, introduce fixed bit width quantization mapping function Q(·), and transform the full precision weight tensor θ∈Θ pruned Mapping to fixed-point representation: Among them, θ q represents the quantized fixed-point weight, θ min and θ max Represent the minimum and maximum weight values in the current tensor respectively, b is the quantization bit width, Δ is the quantization step size, and the mapping process keeps the interval ratio of the original weight unchanged; S65. Quantize the weights θ of all network layers q Replace the original floating-point parameters and construct the final lightweight fault diagnosis model N light , whose parameter set is denoted as Θ light ={θ q }; S66. Lightweight fault diagnosis model N light Perform edge deployment adaptation and generate model files suitable for embedded inference hardware. The model files contain parameter compression representation, compiled intermediate format, and execution graph structure description, supporting low-latency loading of heterogeneous edge chips. S67. Lightweight fault diagnosis model N light Deployed on edge computing devices with limited computing resources, the model is loaded and initialized using the edge inference engine, and the online inference batch size is set to B = 1 to support single-sample real-time fault identification. S68. Input normalized multimodal operation tensor Lightweight fault diagnosis model running on the edge light , output the fault prediction result at the current time t Complete real-time fault diagnosis tasks at the edge with low latency and low power consumption.
Citation Information
Cited By
Power equipment fault intelligent diagnosis method and system based on deep learning
CN121278535A
Intelligent diagnosis method and system for power equipment fault based on deep learning
CN121278535B
Lightweight model construction method for microseismic signal classification
CN121479439A
A lightweight model construction method for microseismic signal classification
CN121479439B
Ultrasonic water meter embedded self-diagnosis method based on lightweight AI and original signal real-time analysis
CN121702512A