An improved prototype network of coal mining machine cutting unit bearing fault diagnosis method
Patent Information
- Application Number
- CN202610889412.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-25
AI Technical Summary
[0008]针对现有技术的上述缺陷,本发明的目的在于提供一种改进型原型网络的采煤机截割部轴承故障诊断方法,解决小样本跨工况场景下域相关滋扰干扰特征有效性、原型网络固定原型新任务适应性不足的问题,通过因果解耦与快速适应机制的协同增益,提升采煤机截割部轴承故障诊断的准确率与鲁棒性
1、本发明基于因果推断理论设计双分支编码器,将振动特征拆分为故障本质特征与域滋扰特征;同时设计正交性损失、因果对比损失、域多样性正则损失形成动态平衡:正交性损失从高阶统计层面强制两类特征独立,实现信息分离;因果对比损失针对性优化故障特征的类内紧凑性与类间分离度,提升判别性;域多样性正则损失防止域滋扰特征坍缩,保障解耦稳定性。三重协同辅助损失配合双分支因果编码器的协同作用,可有效抑制采煤机转速、负载波动带来的域干扰,使模型捕捉到真正决定故障类型的本质特征,实现彻底的域滋扰解耦,显著提升跨工况鲁棒性。
Smart Images

Figure CN122818002A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of rotating machinery fault diagnosis and small sample element learning technology. Specifically, it relates to an improved prototype network-based method for fault diagnosis of bearings in the cutting section of a coal mining machine. It is applicable to fault diagnosis scenarios of bearings in the cutting section of a coal mining machine under various working conditions, where fault samples are scarce and working conditions are complex and variable in underground coal mines. Background Technology
[0002] Coal mining machines are core equipment in coal mining production systems. Their cutting section bearings operate under extremely harsh conditions of high impact, heavy load, variable speed, and high dust levels, making them one of the most critical components with the highest failure rate. A failure in the cutting section bearings can directly lead to machine shutdown, drum jamming, and even underground safety accidents, severely impacting coal mine production efficiency and personnel safety. Therefore, achieving accurate and efficient fault diagnosis of coal mining machine cutting section bearings is of significant engineering importance for ensuring safe coal mine production.
[0003] In recent years, deep learning has made significant progress in the field of rotating machinery fault diagnosis due to its powerful nonlinear feature extraction capabilities. However, these data-driven methods typically rely on a large number of labeled fault samples for training. In reality, the harsh underground environment of coal mines and the extremely high cost of fault reproduction result in a scarcity of labeled fault samples for cutting section bearings. Consequently, the diagnostic performance of traditional deep learning methods drops sharply in small-sample scenarios.
[0004] To address the problem of data scarcity, existing research has proposed various solutions, including data augmentation, transfer learning, and meta-learning. Meta-learning, through the concept of "learning how to learn," trains models to rapidly adapt to multiple source tasks, enabling them to quickly adapt to new tasks with limited samples, providing an effective path for small-sample fault diagnosis. Prototype networks, as a typical metric-based meta-learning method, achieve classification by calculating the distance between samples and class prototypes. They offer advantages such as high computational efficiency and simple implementation, and have been widely applied in small-sample fault diagnosis scenarios.
[0005] However, existing cross-condition fault diagnosis methods based on prototype networks still have two major drawbacks: First, the effectiveness of domain-related disturbance features. Changes in operating conditions (speed, load fluctuations) introduce a large amount of domain-related disturbance information unrelated to the fault into the vibration signal. Traditional feature extractors encode fault information and domain disturbance information together, leading to a decrease in the fault discriminative power of the features. Especially in small sample scenarios with limited source domain samples, mixed features severely weaken the model's cross-operating-condition generalization ability and robustness. Some existing studies have introduced causal decoupling to separate fault features and domain features, but they generally only use a single independence constraint, resulting in incomplete decoupling and a lack of targeted optimization for the discriminative power of fault features.
[0006] Second, fixed prototypes lack adaptability to new tasks. Traditional prototype networks rely on fixed-class prototypes trained in the source domain for classification, and cannot dynamically adjust according to the signal distribution of the target working condition. In cross-scenario situations with significant differences in working conditions, the distribution offset between the prototype and the target sample can lead to a sharp drop in diagnostic accuracy, making it difficult to adapt to the variable underground working conditions of coal mining machines. Existing improvement methods mostly directly modify the prototypes without achieving rapid adaptation at the feature extractor parameter level, resulting in limited adaptability.
[0007] In summary, existing technologies cannot simultaneously solve the problems of domain disturbance suppression and rapid adaptation to new tasks in small sample scenarios, nor can they form a collaborative optimization solution for the special working conditions of the bearings in the cutting section of the coal mining machine, making it difficult to meet the fault diagnosis needs of actual coal mine production. Summary of the Invention
[0008] To address the aforementioned shortcomings of existing technologies, the present invention aims to provide an improved method for diagnosing bearing faults in the cutting section of a coal mining machine using an improved prototype network. This method solves the problems of the effectiveness of domain-related disturbance features in small-sample cross-working-condition scenarios and the insufficient adaptability of the prototype network to new tasks with a fixed prototype. Through the synergistic gain of causal decoupling and rapid adaptation mechanisms, the accuracy and robustness of bearing fault diagnosis in the cutting section of a coal mining machine are improved.
[0009] To achieve the above objectives, the technical solution adopted by this invention is: an improved method for diagnosing bearing faults in the cutting section of a coal mining machine using a prototype network, the specific steps of which are as follows: S1. Collect vibration signal data of the bearing of the cutting part of the coal mining machine under source and target working conditions. Use the source working condition sample as the training set and the target working condition sample as the test set. Randomly sample the source working condition sample according to the N-way K-shot meta-learning method to construct a meta-training task set containing several meta-tasks. Each meta-task contains a support set and a query set.
[0010] S2. Construct a causal encoder with a dual-branch structure. Input the samples from the meta-training task set into the causal encoder: first, extract the basic features through the basic encoder, and then input the basic features into the causal head branch and the domain head branch respectively, and output the fault essence features and domain disturbance features accordingly.
[0011] S3. Perform T-step inner loop parameter optimization based on the fault essence features of the support set: In each inner loop step, calculate the class prototype of each category, measure the distance between the support set samples and the class prototypes by cosine similarity and calculate the inner loop classification loss, update the parameters of the base learner along the loss gradient, and obtain the task-specific parameters adapted to the current meta-task after T iterations.
[0012] S4. Based on the fault essence features and domain perturbation features output by the causal encoder, construct a collaborative auxiliary loss function; the collaborative auxiliary loss function includes orthogonality loss, causal contrast loss and domain diversity regularization loss, which are used to constrain the feature decoupling effect and optimize the discriminativeness of fault features.
[0013] S5. Input the query set into the model after inner loop optimization, calculate the query set classification loss, and sum the query set classification loss and the collaborative auxiliary loss in a weighted manner to obtain the total outer loop loss; backpropagate the gradient based on the total outer loop loss, update the global initial parameters of the meta-learner, and complete the model training.
[0014] S6. Use a small number of labeled samples of the target working condition as the support set, input them into the trained model, and perform T-step inner loop optimization to enable the model to quickly adapt to the target working condition distribution; then use the samples to be diagnosed in the target working condition as the query set to input into the model, and output the fault diagnosis results.
[0015] Furthermore, in step S1, given the source domain dataset... ,Include A finite sample, where Indicates input data, Fault category labels; target domain dataset Containing only a very small number of samples, the two domains share the same fault category space. The differences between the source and target operating conditions are reflected in variations in rotational speed, load, or both. The meta-learning task set is constructed using an N-way K-shot task format. Each training task includes a support set. and query set This constitutes the N-way K-shot task, which specifically involves sampling N categories from the dataset, selecting K samples from each category to form the support set, and M different samples to form the query set.
[0016] Furthermore, in step S2, the design of the causal encoder follows the causal inference theory, decomposing the characteristic representation of the vibration signal into two independent components: fault-related essential features directly related to the fault. And the non-fault-related characteristics of domain disturbances associated with operating conditions. This decomposition is based on the following assumption: In fault diagnosis, fault signals... Can be modeled as ,in The signal generated by the fault For domain-related interference signals, and They are respectively and The impulse response function, where * denotes convolution operation. And assuming fault correlation factors... Domain-related interference factors They do not contain each other.
[0017] Further, in step S2, the basic encoder The ResNet-18 one-dimensional convolutional network structure is adopted, and its basic output features are: ,in The basic feature dimension is then used; subsequently, the basic features are input into the causal head. Heyutou Feature extraction is performed; the causal head is mapped to the fault-based feature space through a two-layer fully connected network, and then output after batch normalization, ReLU activation, and Dropout regularization. Essential characteristics of faults Domain Header Through a fully connected layer combined with Batch Normalization (BN), ReLU, and Dropout, the output is... Domain Disturbance Characteristics The complete output of the causal encoder can be represented as: .in Used for prototyping and classification Used to assist in loss calculation. These are the basic features. Among them, the fault essence features are used for prototype calculation and classification diagnosis, while the domain disturbance features are only used to assist in loss constraints.
[0018] Furthermore, the specific sub-steps for inner loop optimization in step S3 are as follows: S3.1 Initialize the global initial parameters of the basic learner.
[0019] S3.2, will support set samples Enter the current parameter The causal encoder extracts the essential features of the support set faults. .
[0020] S3.3 Calculate the class prototype for each fault category, and the class prototype for the c-th category. The calculation formula is:
[0021] Where c is the category index. To support the sample set of class c in the set, To support the i-th sample in the set, To support the label of the i-th sample in the set, These are the current model parameters; For the sample The fault-related features output by the causal encoder; K is the number of samples contained in each fault category in the support set.
[0022] S3.4 Calculate the similarity between support set samples and various prototypes using cosine similarity. Introducing the temperature parameter τ to smooth the similarity distribution:
[0023] S3.5, The probability that a sample belongs to class c is:
[0024] in, This represents a sample of the query set. Represents a query set sample The predicted fault category label, where N represents the total number of fault categories in the current meta-task.
[0025] S3.6. Measure the distance between the support sample and the prototype using cosine similarity, and calculate the inner loop loss on the support set. :
[0026] S3.6 Update model parameters along the inner loop loss gradient: , where α is the inner loop learning rate; This represents the model parameters before the inner loop update in step t. Indicates the first Model parameters updated within the inner loop; The learning rate for the inner loop is used to control the step size for parameter updates; Indicates the current parameter The inner loop classification loss is calculated based on the support set S.
[0027] S3.8 Repeat S3.2 to S3.7 a total of T times to obtain the parameters adapted to the current task. .
[0028] Furthermore, in step S4, the orthogonality loss is constructed based on the Hilbert-Schmidt Independence Criterion (HSIC), and the specific calculation process is as follows: I. Orthogonality loss The Hilbert-Schmidt Independence Criterion (HSIC) is used to enforce statistical independence between the essential fault characteristics and the domain disturbance characteristics, thus ensuring information independence between the two types of characteristics. For a set of characteristics First, L2 normalization is performed on each feature vector:
[0029] II. The normalized characteristic matrix is denoted as and Calculate the Gram matrix. :
[0030] Each element Let represent the cosine similarity between sample i and sample j. Define the centering matrix. ,in for The identity matrix, It is a B-dimensional column vector of all 1s. Then, the Gram matrix is double-centered:
[0031] III. Definition of Orthogonality Loss for:
[0032] By minimizing this loss, the essential characteristics of the forced fault and the characteristics of domain disturbance are statistically independent, thus achieving information decoupling.
[0033] Furthermore, in step S4, the causal contrast loss and domain diversity regularization loss are calculated as follows: S4.1 Calculate the causal contrast loss Based on orthogonality, the ability to classify essential fault features is enhanced. First, the similarity matrix of essential fault features is calculated:
[0034] Define label mask Indicator mask to determine whether sample pairs belong to the same fault category Excluding self-comparison. The causal contrast loss is defined as:
[0035] The loss employs a supervised contrastive learning paradigm. For each sample i, it brings its feature representation closer to that of samples of the same class and pushes it further away from samples of different classes, thereby forming a compact and separate cluster structure in the embedding space and enhancing the fault detection capability.
[0036] S4.2 Computational Domain Diversity Regularization Loss To impose constraints on the domain perturbation features and prevent their collapse, data augmentation is first applied to the source domain samples using three transformations: Gaussian noise, amplitude scaling, and time shifting. For sample I, the domain labels are: 0 represents the original domain; 1 represents the data augmentation domain. Based on the fault category labels... and domain tags Group the samples. For each combination of category y and domain d, calculate the average of the domain perturbation features of all samples within that group as the domain perturbation feature center:
[0037] in Given the number of samples in this group, the domain diversity regularization loss is defined as:
[0038] This loss prevents domain-disturbing features from collapsing into identical representations by minimizing the similarity between centers of the same type but different domains.
[0039] S4.3, The total collaborative assistance loss is:
[0040] in , , These are the weighting coefficients for the corresponding loss terms.
[0041] Furthermore, the specific sub-steps for outer loop optimization in step S5 are as follows: S5.1 Input the query set samples into the model after T-step inner loop optimization, extract the essential fault features, and calculate the query set. The cosine similarity with various prototypes yields the query set classification loss. The specific formula is as follows:
[0042] S5.2 Calculate the collaborative assistance loss for the current batch. The total external circulation loss is:
[0043] S5.3 Calculate the gradient with respect to the global initial parameter θ based on the total external loop loss, and update the global parameters using the optimizer: Where β is the outer loop learning rate; through numerous iterations of inner and outer loops of meta-tasks, the model learns meta-knowledge that can be quickly transferred. This optimization strategy enables the model to learn a good set of initial parameters during the meta-training phase. This parameter enables excellent performance to converge quickly on new tasks with a small number of gradient steps.
[0044] Compared with the prior art, the present invention has the following advantages: 1. This invention designs a dual-branch encoder based on causal inference theory, decomposing vibration characteristics into fault essence characteristics and domain disturbance characteristics. Simultaneously, it designs orthogonality loss, causal contrast loss, and domain diversity regularization loss to form a dynamic balance: orthogonality loss forces the two types of features to be independent at a higher-order statistical level, achieving information separation; causal contrast loss specifically optimizes the intra-class compactness and inter-class separation of fault features, improving discriminative power; and domain diversity regularization loss prevents the collapse of domain disturbance characteristics, ensuring decoupling stability. The triple collaborative auxiliary loss, combined with the synergistic effect of the dual-branch causal encoder, can effectively suppress domain disturbances caused by fluctuations in coal mining machine speed and load, enabling the model to capture the essential characteristics that truly determine the fault type, achieving thorough domain disturbance decoupling, and significantly improving robustness across operating conditions.
[0045] 2. This invention introduces a dual-layer optimization mechanism with inner and outer loops within the prototype network framework: the inner loop utilizes a small number of support set samples from the target task to rapidly update the feature extractor parameters in a clean fault-based feature space, dynamically adjusting the prototype class to adapt to the target operating condition distribution; the outer loop optimizes global initial parameters through numerous meta-tasks, enabling the model to learn meta-knowledge that can be quickly transferred. Compared to traditional prototype networks with fixed prototypes, this mechanism achieves rapid task adaptation through inner loop parameter optimization, dynamically adjusting the feature space and prototype position according to the target operating condition, effectively overcoming the performance degradation caused by operating condition distribution shifts, and significantly improving the adaptability to new operating conditions with only a small number of labeled samples.
[0046] 3. The causal decoupling module of this invention provides a pure fault feature space free from domain interference for the inner loop optimization, avoiding the misleading gradient update direction by the operating condition disturbance information, and greatly improving the efficiency and stability of the inner loop parameter optimization, so that effective adaptation can be achieved with a small number of samples. Conversely, the rapid parameter adaptation of the inner loop further amplifies the discriminative advantage of causal features, so that the decoupled fault features can quickly fit the target operating condition distribution and give full play to their diagnostic value. This makes the causal decoupling and the rapid adaptation mechanism form a synergistic gain effect. Ablation experiments verify that the performance improvement of the combination of the two is significantly greater than the sum of the improvements of the individual modules.
[0047] 4. This invention is designed for the characteristics of underground working conditions such as strong impact, variable load, high noise, and scarce fault samples of the bearings of the coal mining machine cutting section. It can achieve accurate diagnosis across working conditions without the need for a large number of fault samples. It can be directly deployed in the coal mine equipment health monitoring system, providing technical support for the condition monitoring and fault early warning of the bearings of the coal mining machine cutting section. Attached Figure Description
[0048] Figure 1 This is a schematic diagram of the overall architecture of the prototype network based on causal decoupling for rapid adaptation as described in this invention.
[0049] Figure 2This is a schematic diagram of the network structure of the dual-branch causal encoder described in this invention.
[0050] Figure 3 The bar chart shows the comparison of diagnostic accuracy between the proposed method and existing meta-learning fault diagnosis methods on the PU dataset.
[0051] Figure 4 A visualization comparison of t-SNE features for each method under the H2→H3 cross-condition scenario. Detailed Implementation
[0052] The present invention will be further described below.
[0053] The present invention proposes a method for diagnosing bearing faults in the cutting section of a coal mining machine based on a Causal Disentangled Fast Adaptive Prototypical Network (CDFAPN). The overall architecture is as follows: Figure 1 As shown, the core consists of five parts: a meta-task construction module, a dual-branch causal encoder, a fast task adaptation inner loop module, a collaborative auxiliary loss module, and an outer loop global optimization module. The network structure of the causal encoder is as follows: Figure 2 As shown, a two-branch structure of "basic encoder + causal head + domain head" is used to achieve feature decoupling.
[0054] The following two specific embodiments, combined with publicly available datasets, verify the technical effects of the present invention. All embodiments were completed in the same hardware environment.
[0055] Example 1: Overload Fault Diagnosis of Coal Mining Machine Cutting Section Bearings Based on HUST Dataset This embodiment simulates the actual working condition of load fluctuation of the bearing in the cutting section of a coal mining machine. The HUST bearing dataset is used to verify the cross-load fault diagnosis performance of the present invention. This dataset is collected from the bearing test bench of the cutting section of the coal mining machine and is highly consistent with the characteristics of actual underground working conditions.
[0056] 1. Dataset and Experiment Setup The HUST dataset contains three load conditions: H1 (no load, 0W), H2 (medium load, 200W), and H3 (heavy load, 400W). Each condition includes seven health states: normal (N), inner race fault (I), outer race fault (O), rolling element fault (B), inner race-rolling element combined fault (IB), outer race-rolling element combined fault (OB), and inner and outer race combined fault (IO). Each fault type contains 30 labeled samples, each of which is a one-dimensional vibration signal with 1024 points and a sampling frequency of 12kHz.
[0057] This embodiment constructs six cross-load diagnostic scenarios: H1→H2, H2→H1, H1→H3, H3→H1, H2→H3, and H3→H2. The part before the arrow represents the source domain training scenario, and the part after the arrow represents the target domain testing scenario.
[0058] The meta-tasks employ a 7-way 1-shot setup: each meta-task randomly selects 7 fault categories, choosing 1 sample from each category to form the support set, and 5 unique samples from each category to form the query set. A total of 20,000 meta-tasks are constructed during the meta-training phase, and the experiment is repeated 100 times during the testing phase, with the average accuracy taken as the final result.
[0059] 2. Model parameter settings Causal encoder: The basic encoder uses a ResNet-18 one-dimensional convolutional network, takes a vibration signal of length 1024 as input, and outputs 256-dimensional basic features; the causal head is a two-layer fully connected network (256→128→64), followed by a batch normalization layer, a ReLU activation function, and a Dropout layer (dropout rate 0.2), outputting 64-dimensional fault-related features; the domain head is a one-layer fully connected network (256→32), followed by a batch normalization layer, a ReLU activation function, and a Dropout layer (dropout rate 0.2), outputting 32-dimensional domain perturbation features.
[0060] Optimized parameters: inner loop iteration steps T=5, inner loop learning rate α=0.01; outer loop learning rate β=0.001, Adam optimizer is used, batch size is 16; total training rounds are 200, and the learning rate decays to 0.8 times its original value every 50 rounds.
[0061] Loss function parameters: orthogonality loss weight λ1=0.1, causal contrast loss weight λ2=0.3, domain diversity regularization loss weight λ3=0.1; temperature parameter τ=0.1.
[0062] Data augmentation settings: The augmentation methods corresponding to the domain diversity regularization loss are: Gaussian noise (signal-to-noise ratio 0dB), amplitude scaling (scaling factor range 0.8~1.2), and time shift (shift length ±50 points). One augmentation method is randomly selected for each sample.
[0063] 3. Implementation Steps S1. Dataset partitioning and meta-task construction: Divide the source domain and target domain according to the above 6 cross-working scenarios, randomly sample the source domain samples, construct 20,000 7-way 1-shot meta-tasks, and form a meta-training task set.
[0064] S2. Feature Extraction and Causal Decoupling: The support set and query set samples in the meta-task are input into the dual-branch causal encoder. After the basic features are extracted by the ResNet basic encoder, the 64-dimensional fault essence features and 32-dimensional domain disturbance features are output through the causal head and domain head, respectively.
[0065] S3, Inner Loop Fast Task Adaptation: Initialize global model parameters and perform 5-step inner loop optimization on the support set: In each step, extract the essential features of the faults in the support set, calculate the class prototypes of 7 fault categories, calculate the classification loss of the support set through cosine similarity, and update the model parameters along the gradient direction; after 5 iterations, obtain the model parameters adapted to the current meta-task.
[0066] S4. Calculation of Collaborative Assistance Loss: Based on the fault characteristics and domain disturbance characteristics of all samples in the batch, calculate the orthogonality loss, causal contrast loss and domain diversity regularization loss respectively, and obtain the total collaborative assistance loss by weighting them according to their weights.
[0067] S5. Optimization of global parameters in the outer loop: Based on the optimized model parameters in the inner loop, extract the fault-related features of the query set and calculate the query set classification loss; add the classification loss and the collaborative auxiliary loss to obtain the total outer loop loss, and update the global initial parameters of the meta-learner through backpropagation; repeat the above meta-training process until 200 rounds of training are completed.
[0068] S6. Target Working Condition Testing and Diagnosis: Take 7 labeled samples (1 from each class) under the target working condition as the target support set, input them into the trained model, and perform 5-step inner loop optimization to enable the model to quickly adapt to the target load working condition; take the remaining samples to be diagnosed under the target working condition as the query set and input them into the model, and output the fault diagnosis result of each sample; calculate the proportion of correctly classified samples to obtain the diagnostic accuracy in this scenario.
[0069] 4. Implementation Results This embodiment achieves an average diagnostic accuracy of 92.3% across six different load scenarios, representing an improvement of 8.7 percentage points compared to the traditional prototype network (PN), 6.2 percentage points compared to the MAML method, and 4.5 percentage points compared to the prototype network containing only causal decoupling. In the scenario H1→H3, where the load conditions differ the most, this method still achieves an accuracy of 89.7%, significantly outperforming other comparative methods and demonstrating excellent cross-load robustness and small-sample adaptability.
[0070] Example 2: Multi-parameter cross-working-condition fault diagnosis of bearings in the cutting section of a coal mining machine based on the PU dataset This embodiment simulates a complex underground working condition where the speed, load, and radial force of the bearing in the cutting section of a coal mining machine fluctuate simultaneously. The invention's cross-condition diagnostic performance under multi-parameter variation scenarios is verified using the PU public bearing dataset.
[0071] 1. Dataset and Experiment Setup The PU dataset was collected from a rolling bearing fatigue test bench. In this embodiment, three typical combined working conditions were selected, and the parameters are shown in the table below:
[0072] Each operating condition includes 10 different bearing states with different fault types and severity. Each category contains 50 labeled samples, and each sample is a one-dimensional vibration signal with a length of 2048 points and a sampling frequency of 64kHz.
[0073] This embodiment constructs three cross-condition diagnostic scenarios: P1→P2, P1→P3, and P2→P3. Two meta-task modes are set: 10-way 1-shot (1 supporting sample per class) and 10-way 5-shot (5 supporting samples per class). The query set for each task contains 5 samples per class. During the testing phase, the experiment is repeated 100 times, and the average accuracy is calculated.
[0074] 2. Model parameter settings Causal encoder: The basic encoder uses a ResNet-18 one-dimensional convolutional network, takes a vibration signal of length 2048 as input, and outputs 512-dimensional basic features; the causal head is a two-layer fully connected network (512→256→128), followed by a batch normalization layer, a ReLU activation function and a Dropout layer (dropout rate 0.3), outputting 128-dimensional fault essence features; the domain head is a one-layer fully connected network (512→64), followed by a batch normalization layer, a ReLU activation function and a Dropout layer (dropout rate 0.3), outputting 64-dimensional domain perturbation features.
[0075] Optimization parameters: inner loop iteration steps T=5, inner loop learning rate α=0.01; outer loop learning rate β=0.001, using AdamW optimizer, weight decay coefficient is 1e-4, batch size is 8; total training epochs are 300, and the learning rate decays to 0.8 times the original value every 60 epochs.
[0076] Loss function parameters: orthogonality loss weight λ1=0.15, causal contrast loss weight λ2=0.25, domain diversity regularization loss weight λ3=0.1; temperature parameter τ=0.07.
[0077] Data augmentation settings: The augmentation methods corresponding to the domain diversity regularization loss are: Gaussian noise (signal-to-noise ratio -2dB~2dB), amplitude scaling (scaling factor range 0.7~1.3), and time shift (shift length ±100 points). One augmentation method is randomly selected for each sample.
[0078] 3. Implementation Steps The implementation process of this embodiment is the same as that of Embodiment 1. The core steps are meta-task construction, causal feature decoupling, rapid adaptation of the inner loop, collaborative auxiliary loss optimization, global update of the outer loop, and target condition testing. For both 1-shot and 5-shot settings, the number of support set samples was adjusted and the experiments were repeated.
[0079] 4. Implementation Results like Figure 3 As shown, the method of the present invention (CDFAPN) achieves the best performance compared with other existing methods in three cross-operating condition scenarios.
[0080] like Figure 4 The t-SNE feature visualization results show that the features of the comparative method have obvious domain distribution shifts, and the samples of the same type of fault are scattered due to different operating conditions. However, in the fault essence features extracted by this method, the samples of the same type are closely clustered, the samples of different types are clearly separated, and the samples of the same type under different operating conditions almost overlap, which verifies the effectiveness of causal decoupling and the performance advantages of the fast adaptation mechanism.
[0081] Ablation Experiment Verification: To verify the synergistic gain effect of the causal decoupling module and the fast adaptation mechanism in this invention, ablation experiments were conducted in the H2→H3 cross-load scenario and under a 7-way 1-shot setting. The diagnostic accuracy of the baseline prototype network, the network with only the causal decoupling module added, the network with only the inner loop fast adaptation added, and the complete method of this invention were tested. The results are shown in the table below:
[0082] The results show that when the causal decoupling module and the fast adaptation mechanism are used individually, they bring performance improvements of 3.6% and 3.3% respectively, with a total improvement of 6.9%. However, the complete method of this invention brings a performance improvement of 7.9%, which is significantly greater than the sum of the improvements of the two modules individually, verifying that there is a synergistic gain effect between the two.
[0083] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for diagnosing bearing faults in the cutting section of a coal mining machine using an improved prototype network, characterized in that, Includes the following steps: S1. Collect vibration signal data of the bearing of the cutting part of the coal mining machine under source and target conditions. Use the source condition sample as the training set and the target condition sample as the test set. Randomly sample the source condition sample according to the meta-learning method to construct a meta-training task set containing several meta-tasks. Each meta-task contains a support set and a query set. S2. Construct a causal encoder with a dual-branch structure. Input the samples in the meta-training task set into the causal encoder: first, extract the basic features through the basic encoder, and then input the basic features into the causal head branch and the domain head branch respectively, and output the fault essence features and domain disturbance features accordingly. S3. Perform T-step inner loop parameter optimization based on the fault essence features of the support set: In each inner loop step, calculate the class prototype of each category, measure the distance between the support set sample and the class prototype by cosine similarity and calculate the inner loop classification loss, update the parameters of the base learner along the loss gradient, and obtain the task-specific parameters adapted to the current meta-task after T iterations. S4. Based on the fault essence features and domain perturbation features output by the causal encoder, a collaborative auxiliary loss function is constructed. The collaborative auxiliary loss function includes orthogonality loss, causal contrast loss and domain diversity regularization loss, which are used to constrain the feature decoupling effect and optimize the discriminativeness of fault features. S5. Input the query set into the model after inner loop optimization, calculate the query set classification loss, and sum the query set classification loss and collaborative assistance loss in a weighted manner to obtain the total outer loop loss. Based on the backpropagation gradient of the total outer loop loss, the global initial parameters of the meta-learner are updated to complete the model training. S6. Use a small number of labeled samples of the target working condition as the support set, input them into the trained model, and perform T-step inner loop optimization to enable the model to quickly adapt to the target working condition distribution; then use the samples to be diagnosed in the target working condition as the query set to input into the model, and output the fault diagnosis results.
2. The method according to claim 1, characterized in that, In step S1, the difference between the source operating condition and the target operating condition is that the rotational speed is different, the load is different, or both the rotational speed and the load are different; the source domain dataset contains a certain number of labeled samples, and the target domain dataset contains fewer labeled samples than the source domain dataset, and the two share the same fault category space; the construction rule of the meta-task is: randomly select N fault categories from the dataset, select K samples for each category to form a support set, and select M samples that do not overlap with the support set to form a query set.
3. The method according to claim 1, characterized in that, In step S2, the causal encoder is constructed based on causal inference theory. The assumption of its characteristic decoupling is that the vibration signal of the bearing of the coal mining machine cutting section is formed by the convolution of the fault excitation signal and the working condition disturbance signal through their respective impulse response functions; the essential characteristics of the fault are only related to the fault type, and the domain disturbance characteristics are only related to the working condition changes, and the two information spaces are independent of each other.
4. The method according to claim 1, characterized in that, In step S2, the basic encoder adopts a ResNet-18 one-dimensional convolutional network structure, with an output dimension of... The basic features; the causal head branch is a two-layer fully connected network, sequentially connecting a batch normalization layer, a ReLU activation function, and a Dropout layer, with the final output dimension being... The essential characteristics of the fault The domain head branch is a fully connected layer that sequentially connects a batch normalization layer, a ReLU activation function, and a Dropout layer, ultimately outputting a dimension of... Domain harassment characteristics The fault-related features are used for prototype calculation and classification diagnosis, while the domain disturbance features are used to assist in loss constraints.
5. The method according to claim 1, characterized in that, In step S3, the specific sub-steps for inner loop optimization are as follows: S3.1 Initialize the global initial parameters θ of the basic learner; S3.2 Input the support set samples into the causal encoder under the current parameters to extract the essential features of the support set faults; S3.3 Calculate the class prototype for each fault category, and the class prototype for the c-th category. The calculation formula is: Where c is the category index. To support the sample set of class c in the set, To support the i-th sample in the set, To support the label of the i-th sample in the set, These are the current model parameters; For the sample The fault-related features output by the causal encoder; K is the number of samples contained in each fault category in the support set; S3.4 Calculate the similarity between support set samples and various prototypes using cosine similarity. Introducing the temperature parameter τ to smooth the similarity distribution: in, For sample i, the essential characteristics of the fault are... The fault-related essential characteristics of sample j; S3.5, The probability that a sample belongs to class c is: in, This represents a sample of the query set. Represents a query set sample The predicted fault category label, where N represents the total number of fault categories in the current meta-task; S3.
6. Measure the distance between the support sample and the prototype using cosine similarity, and calculate the inner loop loss on the support set. : S3.6 Update model parameters along the inner loop loss gradient: , where α is the inner loop learning rate; This represents the model parameters before the inner loop update in step t. Indicates the first Model parameters updated within the inner loop; The learning rate for the inner loop is used to control the step size for parameter updates; Indicates the current parameter Below, the inner loop classification loss is calculated based on the support set S; S3.8 Repeat S3.2 to S3.7 a total of T times to obtain the parameters adapted to the current task. .
6. The method according to claim 1, characterized in that, In step S4, the orthogonality loss is constructed based on the Hilbert-Schmidt independence criterion, and the specific calculation process is as follows: I. Orthogonality loss The Hilbert-Schmidt independence criterion is used to enforce statistical independence between the essential fault characteristics and the domain disturbance characteristics, thus ensuring information independence between the two types of characteristics; for a set of characteristics First, L2 normalization is performed on each feature vector: II. The normalized characteristic matrix is denoted as and Calculate the Gram matrix : Each element The cosine similarity between sample i and sample j is represented by the centering matrix. ,in for The identity matrix, Given a B-dimensional column vector of all 1s, the Gram matrix is then double-centered: III. Definition of Orthogonality Loss for: By minimizing this loss, the essential characteristics of the forced fault and the characteristics of domain disturbance are statistically independent, thus achieving information decoupling.
7. The method according to claim 1, characterized in that, In step S4, the causal contrast loss and domain diversity regularization loss are calculated as follows: S4.1 Causal Contrast Loss: A supervised contrastive learning paradigm is adopted to calculate the similarity matrix of essential fault features within a batch and define a label mask. Logical mask if and only if sample i and sample j belong to the same fault category Excluding comparisons with the samples themselves, the formula for causal comparison loss is: This loss brings similar characteristics closer together and pushes out dissimilar characteristics away, enhancing the discriminative power of the essential characteristics of the fault; S4.2 Domain Diversity Regularization Loss: Gaussian noise, amplitude scaling, and time shift are applied to the source domain samples to augment the data. Original samples are assigned domain label 0, and augmented samples are assigned domain label 1. After grouping by fault category and domain label, the domain disturbance feature center of each group is calculated. : in Given the number of samples in this group, the domain diversity regularization loss is defined as: By minimizing the similarity between feature centers of the same type but different domains, we prevent the collapse of domain-perturbed features and ensure the stability of decoupling. S4.3, The total collaborative assistance loss is: in , , These are the weighting coefficients for the corresponding loss terms.
8. The method according to claim 1, characterized in that, The specific sub-steps for outer loop optimization in step S5 are as follows: S5.1 Input the query set samples into the model after T-step inner loop optimization, extract the essential features of the fault, calculate the cosine similarity between the query set samples and various prototypes, and obtain the query set classification loss. ; S5.2 Calculate the collaborative assistance loss for the current batch. The total external circulation loss is: S5.3 Calculate the gradient with respect to the global initial parameter θ based on the total external loop loss, and update the global parameters using the optimizer: , where β is the outer loop learning rate.
9. The method according to claim 1, characterized in that, Step S6 specifically involves: S6.
1. Take N types of faults under the target working condition and K labeled samples of each type as the target support set, input them into the trained model, and perform T-step inner loop parameter update to make the model adapt to the feature distribution of the target working condition. S6.
2. Take the sample to be diagnosed under the target working condition as the target query set, input it into the adapted model, calculate the cosine similarity between the sample to be diagnosed and various prototypes, and take the category with the highest similarity as the final fault diagnosis result. S6.3 Calculate the percentage of correct classifications for all labeled samples to be diagnosed, and obtain the diagnostic accuracy under the target working condition.
10. The method according to claim 1, characterized in that, The failure types of the bearings in the cutting section of the coal mining machine include at least one of the following: normal condition, inner ring failure, outer ring failure, rolling element failure, inner ring-rolling element combined failure, outer ring-rolling element combined failure, and inner and outer ring combined failure.