Pumping unit fault diagnosis method and system based on channel attention and invariance learning
By constructing a fault diagnosis model of multi-scale channel attention and invariance learning, the feature extraction limitations and domain offset problems in oil pump fault diagnosis are solved, and high-accuracy cross-condition fault identification and real-time monitoring are achieved.
Patent Information
- Application Number
- CN202510854432.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The existing oil pump fault diagnosis methods have limitations in feature extraction, resulting in model failure or field offset, and insufficient interaction of single-scale features and mismatched conditional distributions, resulting in low accuracy of diagnostic results.
A fault diagnosis model is constructed using multi-scale channel attention unit, classifier, invariance feature learning unit and joint domain adaptation unit. The vibration signal characteristics are captured through multi-scale convolution kernels, and the gradient internal product penalty term constraint model is introduced to optimize direction consistency, and the alignment edge and conditional distribution is adapted through joint domains, and the conditional adversarial network and joint maximum mean difference are used to synchronize the distribution.
It improves the diagnostic accuracy of the fault diagnosis model, can identify faults of the pump motor and gearbox across operating conditions, realize real-time online monitoring, and has low false alarm rate. It is suitable for online monitoring systems for oil field oil pumps.
Smart Images

Figure CN120354252A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault diagnosis of pumping units, and particularly relates to a fault diagnosis method and system for pumping units based on channel attention and invariance learning. Background Art
[0002] With the improvement of the intelligent level of industrial equipment, the fault diagnosis technology of mechanical equipment has gradually shifted from traditional manual experience judgment to automated and intelligent analysis. As the core equipment for oil extraction in oilfields, the motors and speed reducers of pumping units operate under complex working conditions of long-term high load and variable speed, with frequent failures and great difficulty in diagnosis. The existing fault diagnosis technologies still have the following problems in practical applications: Traditional methods rely on expert experience and signal processing technologies (such as Fourier transform, wavelet analysis), and judge faults by manually extracting the time-frequency domain features of vibration signals (such as spectral peaks, envelope demodulation). For example, in rotating machinery, misalignment of couplings usually shows up as 2 times the rotation frequency (such as 80 Hz) and its harmonic components (such as 120 Hz), and bearing faults are accompanied by characteristic frequencies (such as the passing frequency of rolling elements). However, this method depends on manual experience for feature extraction, requires pre-defining fault feature frequency formulas, and is difficult to handle complex coupling faults (such as coexistence of bearing damage and gear wear). Moreover, its cross-condition generalization ability is poor. The vibration characteristics of the same fault are significantly different under different speeds and loads. For example, the spectral distribution offset under the working conditions of 600 r / min and 1200 r / min will cause the fault diagnosis model to fail, and specific fault information cannot be obtained, or the obtained fault information is incorrect or incomplete.
[0003] Although intelligent diagnosis methods based on convolutional neural networks (CNNs) and residual networks (ResNets) can automatically extract features, their performance depends on strict i.i.d. assumptions, requiring the training data (source domain) and test data (target domain) to follow the same distribution. In actual industrial scenarios, the working conditions of equipment are variable, resulting in domain shift. For example, the vibration signal distributions of a pumping unit motor under normal conditions (600 r / min) and fault conditions (1200 r / min) are significantly different, and it has strong label dependence. Supervised learning requires a large amount of labeled data, and the labeling cost of target domain data is high (such as the need to disassemble the equipment to verify the fault type).
[0004] In existing transfer learning, researchers have proposed unsupervised domain adaptation (UDA) methods to alleviate domain shift. However, there are still deficiencies in single-scale feature interaction. Traditional channel attention mechanisms (such as ECA) use a single convolutional kernel to extract channel weights, making it difficult to capture multi-scale fault features, such as the weak high-frequency components of early bearing damage and the low-frequency impacts of severe wear. Moreover, there is a lack of invariant feature learning. Existing methods only reduce domain differences through distribution alignment and do not explicitly constrain the model to learn invariant features shared across working conditions, such as the envelope energy features of bearing faults. There is also conditional distribution mismatch. Most adversarial training methods (such as DANN) only align the marginal distributions of the source domain and the target domain, ignoring the differences in class conditional distributions, such as the spectral morphology differences of the same fault at different speeds. Summary of the Invention
[0005] To address the problems that the feature extraction of existing pumping unit fault diagnosis methods has limitations or actual deviations from the extraction requirements, resulting in the failure of the fault diagnosis model or domain shift, and the single-scale feature interaction is insufficient and the conditional distribution is mismatched, ultimately leading to low accuracy of the diagnosis results, the present invention further proposes a pumping unit fault diagnosis method and system based on channel attention and invariant learning.
[0006] The technical solution adopted by the present invention is as follows: It includes the following steps: S1. Collect the multi-condition vibration signals of the pumping unit motor and the multi-condition vibration signals of the speed reducer respectively, and construct a labeled source domain dataset and an unlabeled target domain dataset according to all the collected multi-condition vibration signals.
[0007] S2. Preprocess the labeled source domain dataset and the unlabeled target domain dataset respectively to obtain a labeled source domain dataset III and an unlabeled target domain dataset III.
[0008] S3. Construct a fault diagnosis model, input the labeled source domain dataset III and the unlabeled target domain dataset III into the fault diagnosis model for training, output the fault category and location information, and obtain a trained fault diagnosis model.
[0009] The fault diagnosis model includes a multi-scale channel attention unit, a classifier, an invariant feature learning unit, and a joint domain adaptation unit.
[0010] S4. Collect the multi-condition vibration signals of the motor or the multi-condition vibration signals of the speed reducer to be diagnosed, input the collected multi-condition vibration signals into the trained fault diagnosis model in S3, and output the corresponding fault category and location information.
[0011] Further, when collecting the multi-condition vibration signals of the pumping unit motor and the multi-condition vibration signals of the speed reducer in S1, the collection parameters are: Sampling frequency: 20 kHz.
[0012] Length of the sampling signal: Each single sample contains 1024 data points, corresponding to a duration of 0.0512 seconds.
[0013] Operating condition label: The source domain data includes rotational speed, load, and fault labels. The fault labels include outer ring damage of the bearing and broken teeth of the gear.
[0014] Further, in S2, the labeled source domain dataset and the unlabeled target domain dataset are preprocessed respectively to obtain the labeled source domain dataset III and the unlabeled target domain dataset III. The specific process is as follows: S21: For each data sample in the labeled source domain dataset and the unlabeled target domain dataset, wavelet thresholding is used to eliminate high-frequency noise, obtaining the labeled source domain dataset I and the unlabeled target domain dataset I. The wavelet threshold uses the DB4 wavelet basis and is decomposed into 3 layers.
[0015] S22: Each data sample in the labeled source domain dataset I and the unlabeled target domain dataset I is normalized by the maximum and minimum values, obtaining the labeled source domain dataset II and the unlabeled target domain dataset II.
[0016] S23: Gaussian noise and time-domain random cropping are added to each data sample in the labeled source domain dataset II and the unlabeled target domain dataset II, obtaining the labeled source domain dataset III and the unlabeled target domain dataset III. The signal-to-noise ratio SNR of the Gaussian noise is 20 dB.
[0017] Further, in S3, a fault diagnosis model is constructed. The labeled source domain dataset III and the unlabeled target domain dataset III are input into the fault diagnosis model for training, and the fault category and location information are output to obtain the trained fault diagnosis model. The specific process is as follows: S31: The labeled source domain dataset III and the unlabeled target domain dataset III are input into the multi-scale channel attention unit for feature extraction, and the features corresponding to each data sample are output, obtaining the feature set of the labeled source domain dataset III and the feature set of the unlabeled target domain dataset III.
[0018] S32: The feature set of the labeled source domain dataset III and the feature set of the unlabeled target domain dataset III obtained in S31 are input into the classifier, and the predicted probabilities of the source domain data and the predicted probabilities of the target domain data are output respectively. The pseudo-labels of the corresponding target domain data are generated according to the predicted probabilities of the target domain data.
[0019] S33. Input the feature set of the labeled source domain dataset Ⅲ and the feature set of the unlabeled target domain dataset Ⅲ obtained in S31, as well as the predicted probability of the source domain data and the pseudo-label of the target domain data output in S32 into the invariance feature learning unit. Calculate the source domain cross-entropy loss according to the feature set of the labeled source domain dataset Ⅲ and the predicted probability of the source domain data, calculate the target domain pseudo-label loss according to the feature set of the unlabeled target domain dataset Ⅲ and the pseudo-label of the target domain data, and calculate the gradient inner product according to the source domain cross-entropy loss and the target domain pseudo-label loss.
[0020] S34. Input the feature set of the labeled source domain dataset Ⅲ and the feature set of the unlabeled target domain dataset Ⅲ obtained in S31, as well as the predicted probability of the source domain data, the predicted probability of the target domain data and the pseudo-label of the target domain data output in S32 into the joint domain adaptation unit. Use the conditional adversarial network to obtain the optimal domain discriminator loss function, generate the conditional distribution, and perform distribution alignment using the joint maximum mean discrepancy.
[0021] S35. The fault diagnosis model further includes a multi-objective joint optimization unit. Integrate the source domain cross-entropy loss and the gradient inner product obtained in S33, as well as the optimal domain discriminator loss function and the distribution alignment obtained in S34 into the multi-objective joint optimization unit, and use the integrated result as the total loss function of the fault diagnosis model.
[0022] S36. According to the predicted probability of the target domain data output in S32, add an entropy weighting strategy to screen the high-confidence target domain prediction results, and output the fault category and location information. The specific process is as follows: Calculate the entropy value according to the predicted probability of the target domain data : where is the predicted probability of the target domain data.
[0023] Generate the weight according to the entropy value , and screen the high-confidence target domain prediction results.
[0024] So far, the trained fault diagnosis model is obtained.
[0025] Furthermore, the multi-scale channel attention unit in S31 sequentially includes a convolutional layer 1, a batch normalization layer 1 (BN), a ReLU activation layer 1, a max pooling layer 1, an Msk-ECA module, a convolutional layer 2, a batch normalization layer 2, a ReLU activation layer 2, a max pooling layer 2, an Msk-ECA module, a convolutional layer 3, a batch normalization layer 3, a ReLU activation layer 3, a max pooling layer 3, an Msk-ECA module, a convolutional layer 4, a batch normalization layer 4, a ReLU activation layer 4 and a max pooling layer 4.
[0026] The Msk-ECA module successively includes a global average pooling layer, a multi-scale convolutional interaction layer, a feature weighted fusion layer, and a Sigmoid activation layer.
[0027] The global average pooling layer compresses the input feature map along the time dimension to generate a channel descriptor.
[0028] The multi-scale convolutional interaction layer is based on the channel descriptor and uses one-dimensional convolutional kernels of three different sizes to extract channel interaction features in parallel.
[0029] The feature weighted fusion layer adds the channel interaction features to obtain the added channel interaction features.
[0030] The Sigmoid activation layer generates the final channel weights by Sigmoid activation of the added channel interaction features, weights the feature map input to the global average pooling layer using the final channel weights, outputs a channel weighted feature map, and obtains the corresponding features according to the channel weighted feature map.
[0031] Furthermore, in S33, the feature set of the labeled source domain dataset III and the feature set of the unlabeled target domain dataset III obtained in S31, as well as the predicted probability of the source domain data and the pseudo-labels of the target domain data output in S32, are input into the invariance feature learning unit. The source domain cross-entropy loss is calculated according to the feature set of the labeled source domain dataset III and the predicted probability of the source domain data, the target domain pseudo-label loss is calculated according to the feature set of the unlabeled target domain dataset III and the pseudo-labels of the target domain data, and the gradient inner product is calculated according to the source domain cross-entropy loss and the target domain pseudo-label loss. The specific process is as follows: Calculate the source domain cross-entropy loss according to the feature set of the labeled source domain dataset III obtained in S31 and the predicted probability of the source domain data output in S32 : Among them, represents the number of source domain data samples in the current batch, represents the total number of fault categories, represents the one-hot encoding of the true label of the i-th sample in the source domain dataset in the c-th class, represents the predicted probability that the classifier predicts that the i-th data sample belongs to the c-th class.
[0032] Calculate the target domain pseudo-label loss according to the feature set of the unlabeled target domain dataset III obtained in S31 and the pseudo-labels of the target domain data output in S32 : Among them, represents the number of target domain samples in the current batch, is the pseudo-label of the j-th sample in the target domain dataset for the c-th class, indicating the predicted probability that the classifier assigns the -th sample to the c-th class.
[0033] Calculate the source domain cross-entropy loss and the target domain pseudo-label loss for the inner product of the gradients of the fault diagnosis model parameters : where, represents the i-th trainable parameter in the fault diagnosis model.
[0034] Furthermore, in S34, the feature set of the labeled source domain dataset III obtained in S31, the feature set of the unlabeled target domain dataset III, as well as the predicted probabilities of the source domain data, the predicted probabilities of the target domain data, and the pseudo-labels of the target domain data output in S32 are input into the joint domain adaptation unit. Using a conditional adversarial network, the optimal domain discriminator loss function is obtained, and a conditional distribution is generated. The distribution alignment is performed using the joint maximum mean discrepancy. The specific process is as follows: According to the features in the feature set of the labeled source domain dataset III obtained in S31, the corresponding feature vectors are obtained. According to the features in the feature set of the unlabeled target domain dataset III obtained in S31, the corresponding feature vectors are obtained. The two parts of the obtained feature vectors, the predicted probabilities of the source domain data, and the predicted probabilities of the target domain data output in S32 are input into the conditional adversarial network. The outer product of each feature vector and the corresponding predicted probability is calculated as the corresponding conditional feature, and all conditional features are obtained. Based on all conditional features, the domain discriminator is trained using a gradient reversal layer to obtain the optimal domain discriminator, and the loss function of the optimal domain discriminator is calculated and a conditional distribution for each data sample is generated.
[0035] where, is the feature vector of the i-th source domain data feature, is the feature vector of the j-th target domain data feature, is the predicted probability of the i-th source domain data, is the predicted probability of the j-th target domain data, is the conditional feature of the i-th source domain data, is the conditional feature of the i-th target domain data, and D(·) is the output of the domain discriminator.
[0036] Meanwhile, based on the feature set of the labeled source domain dataset III, the feature set of the unlabeled target domain dataset III, the pseudo-labels of the target domain data output by S32, and the conditional distribution of each data sample, the marginal distribution and conditional distribution of the same-class data samples in the source domain dataset and the target domain dataset are synchronously aligned by using the joint maximum mean discrepancy: Among them, represents distribution alignment, represents the set of samples belonging to the c-th class in the source domain dataset, represents the sample belonging to the c-th class in the source domain dataset, represents the set of samples belonging to the c-th class in the target domain dataset, represents the sample belonging to the c-th class in the target domain dataset, represents the number of samples of the c-th class in the source domain dataset, represents the number of samples of the c-th class in the target domain dataset, represents the Gaussian kernel mapping function, represents the norm in the RKHS space.
[0037] Furthermore, the total loss function of the fault diagnosis model in S35 is: Among them, , , are trade-off coefficients, = 0.5, = 1, = 0.3, is the penalty factor, = 0.1.
[0038] Furthermore, the alarm logic of the trained fault diagnosis model in S3 is: When the maximum prediction probability and the entropy value , a first-level alarm is triggered.
[0039] When or , a second-level early warning is triggered.
[0040] A pumping unit fault diagnosis system based on channel attention and invariance learning includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, any step of a pumping unit fault diagnosis method based on channel attention and invariance learning is implemented.
[0041] The beneficial effects of the present invention are: The multi - condition vibration signals of the pumping unit motor or the multi - condition vibration signals of the speed reducer after pre - treatment in the present invention are input into a self - constructed fault diagnosis model. The fault diagnosis model includes a multi - scale channel attention unit, a classifier, an invariance feature learning unit, a joint domain adaptation unit, and a multi - objective joint optimization unit. The multi - scale channel attention unit is used to extract the features of the multi - condition vibration signals, capture the fault features in different frequency bands of the vibration signals through 3 / 5 / 7 multi - scale convolution kernels, and overcome the problem of insufficient feature coverage of the traditional single - kernel attention mechanism. The classifier outputs the prediction probability according to the features. The invariance feature learning unit introduces a gradient inner - product penalty term, maximizes the gradient similarity, and constrains the consistency of the optimization direction of the fault diagnosis model in the source domain and the target domain to ensure the cross - condition invariance of the fault features. The joint domain adaptation unit fuses JMMD and conditional adversarial training, synchronously aligns the marginal distribution (overall feature distribution) and the conditional distribution (same - type fault feature distribution), and solves the problem of mis - matching caused by traditional methods ignoring the category structure. The multi - objective joint optimization unit integrates the source - domain cross - entropy loss, gradient inner - product, distribution alignment, and the best domain discriminator loss function obtained in the present invention as the total loss function of the fault diagnosis model, and obtains the final fault diagnosis model for diagnosing the faults of the pumping unit motor and the speed reducer, and outputs the fault category and location information. Through the above settings, the present invention improves the diagnostic accuracy of the fault diagnosis model in multiple aspects.
[0042] The present invention adopts lightweight 1DCNN and parallel attention calculation, and the single - sample inference time is less than 10 ms. The present invention can be connected to the online monitoring system of the pumping unit to achieve real - time alarm. Description of the Drawings
[0043] Figure 1 is a schematic diagram of the multi - scale channel attention unit in the fault diagnosis model; Figure 2 is an example of the spectrum of the motor vibration fault signal Figure 1 ; Figure 3 is an example of the spectrum of the motor vibration fault signal Figure 2 ; Figure 4 is an example of the spectrum of the motor vibration fault signal Figure 3 . Detailed Embodiments
[0044] Detailed Embodiment 1: In combination with Figures 1-4 This embodiment is described. The method for diagnosing faults of a pumping unit based on channel attention and invariance learning in this embodiment includes the following steps: S1. Collect the multi - condition vibration signals of the pumping unit motor and the multi - condition vibration signals of the speed reducer respectively. Construct a labeled source - domain dataset and an unlabeled target - domain dataset based on all the collected multi - condition vibration signals.
[0045] Deploy 1 vibration acceleration sensor in the axial and radial directions at the driving end of the pumping unit motor, and deploy 1 vibration acceleration sensor in the axial and radial directions at the non - driving end. The frequency response range of all vibration acceleration sensors is 5Hz - 20kHz. Collect the multi - condition vibration signals of the motor through all the vibration acceleration sensors installed on the motor. At the same time, deploy 1 vibration acceleration sensor in the radial direction of the input shaft (high - speed end) and the output shaft (low - speed end) of the speed reducer, and also deploy 1 vibration acceleration sensor on the top of the speed reducer housing. Collect the multi - condition vibration signals of the speed reducer through all the vibration acceleration sensors installed on the speed reducer. Divide all the collected multi - condition vibration signals into a labeled source - domain dataset and an unlabeled target - domain dataset. All the above - mentioned vibration acceleration sensors are fixed on the motor and the speed reducer through magnetic - adsorption bases to ensure that the axis of the vibration acceleration sensor is consistent with the vibration direction of the motor or the speed reducer. The signal line uses a shielded cable to prevent electromagnetic interference.
[0046] When collecting the multi - condition vibration signals above, the present invention sets the sampling frequency to 20kHz, so as to cover the upper limit of the bearing fault characteristic frequency of 10kHz in the motor and the speed reducer, thereby achieving comprehensiveness, reliability, convenience, and rapidity of diagnosis. Sampling signal length: A single sample contains 1024 data points, corresponding to a duration of 0.0512 seconds. Condition marking: The source - domain data needs to record the rotational speed (r / min), load (kW), and fault labels. The fault labels include outer - ring damage of the bearing and broken teeth of the gear.
[0047] S2. Pre - process the labeled source - domain dataset and the unlabeled target - domain dataset respectively to obtain the labeled source - domain dataset III and the unlabeled target - domain dataset III. The specific process is as follows: S21. For each data sample in the labeled source - domain dataset and the unlabeled target - domain dataset, use wavelet thresholding to eliminate high - frequency noise to obtain the labeled source - domain dataset I and the unlabeled target - domain dataset I. The wavelet threshold uses the DB4 wavelet basis and is decomposed into 3 layers.
[0048] S22. Perform maximum - minimum normalization on each data sample in the labeled source - domain dataset I and the unlabeled target - domain dataset I to obtain the labeled source - domain dataset II and the unlabeled target - domain dataset II: Among them, represents the data samples in the labeled source - domain dataset I and the unlabeled target - domain dataset II; Represent data samples of the labeled source domain dataset Ⅰ and the unlabeled target domain dataset Ⅱ after normalization processing.
[0049] S23. Add Gaussian noise and time-domain random cropping to each data sample in the labeled source domain dataset Ⅱ and the unlabeled target domain dataset Ⅱ to obtain the labeled source domain dataset Ⅲ and the unlabeled target domain dataset Ⅲ. The signal-to-noise ratio SNR of the Gaussian noise is 20 dB. The labeled source domain dataset Ⅲ is an augmented labeled source domain dataset, and the unlabeled target domain dataset Ⅲ is an augmented unlabeled target domain dataset.
[0050] S3. Build a fault diagnosis model, input the labeled source domain dataset Ⅲ and the unlabeled target domain dataset Ⅲ into the fault diagnosis model for training, and output the fault category and location information to obtain a trained fault diagnosis model. The specific process is as follows: The fault diagnosis model includes a multi-scale channel attention unit, a classifier, an invariance feature learning unit, a joint domain adaptation unit, and a multi-objective joint optimization unit.
[0051] S31. Input the labeled source domain dataset Ⅲ and the unlabeled target domain dataset Ⅲ into the multi-scale channel attention unit for feature extraction, and output the features corresponding to each data sample to obtain the feature set of the labeled source domain dataset Ⅲ and the feature set of the unlabeled target domain dataset Ⅲ.
[0052] The multi-scale channel attention unit is composed of 4 layers of one-dimensional convolutional layers (1DCNN) and a multi-scale kernel efficient channel attention module (Msk-ECA) stacked alternately. The specific structure includes a convolutional layer 1, a batch normalization layer 1 (BN), a ReLU activation layer 1, a max pooling layer 1, an Msk-ECA module, a convolutional layer 2, a batch normalization layer 2, a ReLU activation layer 2, a max pooling layer 2, an Msk-ECA module, a convolutional layer 3, a batch normalization layer 3, a ReLU activation layer 3, a max pooling layer 3, an Msk-ECA module, a convolutional layer 4, a batch normalization layer 4, a ReLU activation layer 4, and a max pooling layer 4. The specific parameters of the structure are as follows: Convolutional layer 1: The convolutional kernel size is 15, the number of channels is 16, and the stride is 2.
[0053] Convolutional layer 2: The convolutional kernel size is 3, the number of channels is 32, and the stride is 1.
[0054] Max pooling layer 2: The pooling kernel is 2.
[0055] Convolutional layer 3: The convolutional kernel size is 3, the number of channels is 64, and the stride is 1.
[0056] Convolutional layer 4: The convolutional kernel size is 3, the number of channels is 128, and the stride is 1.
[0057] Max pooling layer 4: The pooling kernel is 2.
[0058] The Msk-ECA module includes a global average pooling layer, a multi-scale convolutional interaction layer, a feature weighted fusion layer and a Sigmoid activation layer in sequence.
[0059] 1. Global average pooling layer: input feature map Compress along the time dimension to generate channel descriptors .
[0060] 2. Multi-scale convolution interaction layer: Based on the channel descriptor, three one-dimensional convolution kernels of different sizes are used ( Figure 1 k1, k2, k3) in parallel to extract channel interaction features , realizing multi-scale channel interaction, the channel interaction feature is the multi-scale feature.
[0061] Among them, the convolution kernel size k is adaptively adjusted according to the number of channels C: Among them, γ and b are empirical parameters, γ=2, b=1, and odd means taking the nearest odd number.
[0062] 3. Feature weighted fusion layer: channel interaction features Add them together to get the added channel interaction features.
[0063] 4. Sigmoid activation layer: The added channel interaction features are activated by Sigmoid to generate the final channel weights : Using the final channel weights The feature map of the global average pooling layer input Weighted, output channel weighted feature map : in, is the Sigmoid function, Represents channel weighting. The corresponding features are obtained according to the channel weighted feature map.
[0064] The multi-scale channel attention layer of the present invention captures the fault features in different frequency bands of the multi-condition vibration signals of each data sample through 3 / 5 / 7 multi-scale convolutional kernels, such as the high-frequency resonance of early bearing damage and the low-frequency impact of severe wear. The multi-scale features obtained in this way can better represent the vibration signals and generate more accurate feature maps, where the features are high-dimensional features of 256 dimensions. This step of the present invention overcomes the problem of insufficient feature coverage of the traditional single-core attention mechanism.
[0065] S32. Input the feature set of the labeled source domain dataset III and the feature set of the unlabeled target domain dataset III obtained in S31 into the classifier, and respectively output the prediction probabilities of the source domain data and the prediction probabilities of the target domain data . According to the prediction probabilities of the target domain data , use the threshold to screen out the high-confidence predictions to generate the pseudo-labels corresponding to the target domain data . The threshold is set to the lower limit of confidence ≥0.8, and the upper limit of entropy ≤0.3.
[0066] S33. Input the feature set of the labeled source domain dataset III and the feature set of the unlabeled target domain dataset III obtained in S31, as well as the prediction probabilities of the source domain data and the pseudo-labels of the target domain data output in S32 into the invariance feature learning unit (IFL) to calculate the gradient inner product penalty term, and obtain the gradient inner product . The specific process is as follows: Calculate the source domain cross-entropy loss according to the feature set of the labeled source domain dataset III obtained in S31 and the prediction probabilities of the source domain data output in S32 : Among them, represents the number of source domain data samples in the current batch, represents the total number of fault categories, represents the one-hot encoding of the true label of the i-th sample in the source domain dataset in the c-th category, represents the prediction probability that the classifier predicts that the i-th data sample belongs to the c-th category.
[0067] Calculate the target domain pseudo-label loss according to the feature set of the unlabeled target domain dataset III obtained in S31 and the pseudo-labels of the target domain data output in S32 : Among them, represents the number of target domain samples in the current batch, is the pseudo-label of the j-th sample in the target domain dataset in the c-th category, Indicates the predicted probability that the classifier assigns the th sample to the c-th class.
[0068] Calculate the source domain cross-entropy loss and the target domain pseudo-label loss for the inner product of the gradients of the parameters of the fault diagnosis model : where, represents the i-th trainable parameter in the fault diagnosis model.
[0069] In this step of the present invention, by introducing the gradient inner product penalty term, the gradient similarity is maximized, and the optimization directions of the fault diagnosis model in the source domain and the target domain are constrained to be consistent, thereby enhancing the feature invariance and ensuring the cross-operating condition invariance of the fault features, such as the envelope energy consistency of the bearing outer ring fault at different rotational speeds. The present invention adds the said gradient inner product as a regularization term to the total loss function, and the weight coefficient is 0.1.
[0070] S34. Input the feature set of the labeled source domain dataset III obtained in S31, the feature set of the unlabeled target domain dataset III, the predicted probabilities of the source domain data, the predicted probabilities of the target domain data, and the pseudo-labels of the target domain data output in S32 into the joint domain adaptation unit, use the conditional adversarial network to obtain the optimal domain discriminator loss function, and generate the conditional distribution of each data sample. Specifically, the process is as follows: Obtain the corresponding feature vectors according to the features in the feature set of the labeled source domain dataset III obtained in S31, obtain the corresponding feature vectors according to the features in the feature set of the unlabeled target domain dataset III obtained in S31, input the two parts of the obtained feature vectors, the predicted probabilities of the source domain data, and the predicted probabilities of the target domain data output in S32 into the conditional adversarial network, calculate the outer product of each feature vector and the corresponding predicted probability as the corresponding conditional feature, obtain all conditional features, and use the gradient reversal layer (GRL) to train the domain discriminator according to all conditional features to obtain the optimal domain discriminator, and calculate the loss function of the optimal domain discriminator,
[0071] where, is the feature vector of the i-th source domain data feature, is the feature vector of the j-th target domain data feature, is the predicted probability of the i-th source domain data, is the predicted probability of the j-th target domain data, is the conditional feature of the i-th source domain data, is the conditional feature of the i-th target domain data, and D(·) is the output of the domain discriminator. The conditional adversarial network adopts classic methods such as CDAN (an extension of DANN).
[0072] Meanwhile, according to the feature set of the labeled source domain dataset III and the feature set of the unlabeled target domain dataset III obtained in S31, the pseudo-labels of the target domain data output in S32, and the conditional distribution of each data sample, the marginal distribution and conditional distribution of the same-class samples in the source domain dataset and the target domain dataset are synchronously aligned using Joint Maximum Mean Discrepancy (JMMD) to reduce the feature distribution difference between the same-class samples in the source domain dataset and the target domain dataset, and the distribution alignment The expression is: where, represents the set of samples belonging to the c-th class in the source domain dataset, represents a sample belonging to the c-th class in the source domain dataset, represents the set of samples belonging to the c-th class in the target domain dataset, represents a sample belonging to the c-th class in the target domain dataset, represents the number of samples of the c-th class in the source domain dataset, represents the number of samples of the c-th class in the target domain dataset, represents the Gaussian kernel mapping function that maps the original features to the Reproducing Kernel Hilbert Space (RKHS) for facilitating the measurement of distribution differences, represents the norm in the RKHS space for measuring the distribution difference.
[0073] S35. Input the source domain cross-entropy loss, gradient inner product, distribution alignment, and the best domain discriminator loss function obtained above into the multi-objective joint optimization unit for integration, and use the integrated result as the total loss function of the fault diagnosis model .
[0074] where, , , are trade-off coefficients, = 0.5, = 1, = 0.3, = 0.1 is the penalty factor.
[0075] S36. Entropy weighting strategy: Calculate the entropy value according to the prediction probability of the target domain data , and generate weights , suppress the interference of low-confidence samples, screen the high-confidence target domain prediction results, and obtain a trained fault diagnosis model.
[0076] The present invention inputs the labeled source domain dataset III and the unlabeled target domain dataset III together into the constructed fault diagnosis model. The batch size of the input data each time is 64, that is, 64 samples are randomly selected from all the data samples of the two datasets each time to form a batch. The fault diagnosis model uses the Adam optimizer, and the initial learning rate of the Adam optimizer is 1e3, and the weight decay is 1e5.
[0077] The present invention sets the alarm logic of the trained fault diagnosis model as follows: When the maximum prediction probability and the entropy value , a first-level alarm is triggered.
[0078] When or , a second-level early warning is triggered.
[0079] The present invention uses the gradient inner product penalty term to constrain the consistency of the gradient directions of the loss functions of the source domain and the target domain, and forces the fault diagnosis model to learn cross-condition invariant features. The joint maximum mean discrepancy (JMMD) and the conditional adversarial network are used to align the marginal distributions and conditional distributions of the source domain and the target domain. Based on the entropy weighting strategy, the high-confidence target domain prediction results are screened, and the fault category and location information are output. The present invention improves the diagnostic accuracy of the fault diagnosis model from multiple perspectives, and makes up for the shortcomings of manual feature extraction and intelligent feature extraction.
[0080] S4. Collect the multi-condition vibration signals of the motor or the multi-condition vibration signals of the reducer box to be diagnosed, and input the collected multi-condition vibration signals into the trained fault diagnosis model in S3 to output the corresponding fault category and location information.
[0081] The present invention can meet the diagnosis of the oilfield pumping unit directly using unlabeled target domain data (such as the 1200r / min working condition data of unknown faults) during fault diagnosis (unsupervised cross-condition ability), and can also identify the coupling faults of components such as bearings, gears, and couplings at the same time, and ensure that the online diagnosis system applying the present invention outputs the results within 1 second, and the false alarm rate is less than 5%, with real-time performance and reliability.
[0082] Specific embodiment two: Combine Figures 1-4To describe this embodiment, the pumping unit fault diagnosis system based on channel attention and invariance learning described in this embodiment specifically includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, any step of the pumping unit fault diagnosis method based on channel attention and invariance learning is implemented.
[0083] Specific Embodiment 3: Refer to Figures 1-4 To describe this embodiment, this embodiment further limits the pumping unit fault diagnosis system based on channel attention and invariance learning described in Specific Embodiment 2. In this embodiment, the alarm logic of the pumping unit fault diagnosis system based on channel attention and invariance learning is as follows: When the maximum prediction probability and the entropy value trigger a first-level alarm.
[0084] When or trigger a second-level early warning.
[0085] Example 1 Fault diagnosis of the pumping unit motor bearing across working conditions: (1) Scenario description: Take the data of a pumping unit motor in a certain oil field under normal working conditions and at a rotational speed of 600 r / min as the labeled source domain data, and set the data at a rotational speed of 1200 r / min as the unlabeled target domain data. At a certain moment, the vibration value of the motor vibration signal suddenly increases to 7.5 mm / s (ISO 10816 3 alarm threshold), and this section of vibration signal is acquired.
[0086] (2) Diagnostic process of the fault diagnosis model: 1. Input the acquired vibration signal into the multi-scale channel attention unit to extract features. The output result shows that the last Msk-ECA module is significantly activated at channels 32, 64, and 128, and this phenomenon corresponds to the fault characteristic frequency band of the bearing outer ring.
[0087] 2. Perform invariance analysis through the invariance feature learning unit, and the gradient inner product value reaches 0.85 (the threshold is 0.8), indicating the existence of cross-working-condition invariant features.
[0088] 3. Perform distribution alignment through the joint domain adaptation unit, and the JMMD loss drops from the initial 0.62 to 0.18, and the conditional adversarial loss converges to 0.25.
[0089] 4. Output result of the fault diagnosis model: "Damage to the bearing outer ring" (confidence level 98.7%), and the fault is located at the motor drive end.
[0090] (3)Disassembly verification: There is a spalling area (size 3×5mm) on the outer ring of the bearing of the motor, which is consistent with the diagnostic result.
[0091] Embodiment 2 Compound fault diagnosis of the reduction gearbox gear and the coupling: (1)Scenario description: The gear material of a certain pumping unit reduction gearbox is 20CrMnTi, and the coupling type is elastic pin type. Vibration acceleration sensors are deployed in the radial direction of the input shaft (high-speed end) of the reduction gearbox, with a frequency response range of 0.5Hz - 10kHz, and vibration acceleration sensors are deployed in the radial direction of the output shaft (low-speed end), and a low-frequency vibration sensor is deployed on the top of the housing, with a frequency response range of 0.1Hz - 2kHz). The reduction gearbox rotates at 600r / min under normal working conditions, and the load power is 50kW. Set the data at a rotational speed of 900r / min of the reduction gearbox as the unlabeled target domain data. At this time, multiple frequency components appear in the vibration signal of the reduction gearbox: 2×rotation frequency 60Hz, 3×rotation frequency 240Hz, and obtain this section of vibration signal.
[0092] (2)Diagnostic process of the fault diagnosis model: 1. Input the obtained vibration signal into the multi-scale channel attention unit to extract features. The output result shows that the 3-core branch (high-frequency analysis) of the last Msk-ECA captures the gear meshing frequency of 820Hz, the amplitude suddenly increases to 2.5mm / s, and is accompanied by sidebands of ±rotation frequency (900 / 60 = 15Hz) (i.e., 820±15Hz), indicating local tooth breakage of the gear. Under normal circumstances, the amplitude of the gear meshing frequency of 820Hz is stable at 0.5 - 1.0mm / s without sideband modulation. The 5-core branch (low-frequency analysis) extracts the coupling misalignment feature as 2x rotation frequency 60Hz, the amplitude increases to 1.8mm / s, and the 3x rotation frequency 90Hz amplitude increases to 1.2mm / s, indicating radial deviation of the coupling. When the coupling is properly aligned under normal circumstances, the amplitude of 2x rotation frequency (2×900 / 60 = 30Hz) ≤ 0.3mm / s.
[0093] 2. The output results of the fault diagnosis model: "Local tooth breakage of the gear, tooth breakage length 3mm" (confidence level 92.3%, exceeding the threshold of 90%) and "Radial deviation of the coupling 0.15mm" (confidence level 88.5%).
[0094] (3)Maintenance measures: After the gear is replaced, the amplitude of 820Hz drops to 0.7mm / s, and the sideband energy disappears.
[0095] After adjusting the coupling alignment, the amplitude of 2x rotation frequency drops to 0.3mm / s, and the total RMS vibration value drops to 2.1mm / s (meeting the ISO Class 1 standard).
[0096] The above calculation examples of the present invention are only for explaining in detail the calculation model and calculation process of the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made on the basis of the above description. It is impossible to list all the implementation manners here. Any obvious changes or variations derived from the technical solutions of the present invention still fall within the protection scope of the present invention.
Claims
1. A fault diagnosis method for pumping units based on channel attention and invariance learning, characterized in that: It includes the following steps: S1. Collect the multi - condition vibration signals of the pumping unit motor and the multi - condition vibration signals of the speed reducer respectively. Construct a labeled source - domain dataset and an unlabeled target - domain dataset according to all the collected multi - condition vibration signals; S2. Pre - process the labeled source - domain dataset and the unlabeled target - domain dataset respectively to obtain the labeled source - domain dataset III and the unlabeled target - domain dataset III; S3. Construct a fault diagnosis model. Input the labeled source - domain dataset III and the unlabeled target - domain dataset III into the fault diagnosis model for training, and output the fault category and location information to obtain a trained fault diagnosis model; The fault diagnosis model includes a multi - scale channel attention unit, a classifier, an invariance feature learning unit and a joint domain adaptation unit; S4. Collect the multi - condition vibration signals of the motor or the multi - condition vibration signals of the speed reducer to be diagnosed. Input the collected multi - condition vibration signals into the trained fault diagnosis model in S3, and output the corresponding fault category and location information.
2. The fault diagnosis method for pumping units based on channel attention and invariance learning according to claim 1, wherein: When collecting the multi - condition vibration signals of the pumping unit motor and the multi - condition vibration signals of the speed reducer in S1, the collection parameters are: Sampling frequency: 20 kHz; Sampling signal length: A single sample contains 1024 data points, corresponding to a duration of 0.0512 seconds; Condition label: The source - domain data includes rotational speed, load and fault labels. The fault labels include outer - race damage of the bearing and broken teeth of the gear.
3. The fault diagnosis method for pumping units based on channel attention and invariance learning according to claim 1, wherein: When pre - processing the labeled source - domain dataset and the unlabeled target - domain dataset respectively in S2 to obtain the labeled source - domain dataset III and the unlabeled target - domain dataset III, the specific process is: S21. For each data sample in the labeled source - domain dataset and the unlabeled target - domain dataset, use wavelet threshold to eliminate high - frequency noise to obtain the labeled source - domain dataset I and the unlabeled target - domain dataset I. The wavelet threshold uses the DB4 wavelet basis and is decomposed by 3 layers; S22. Perform maximum - minimum normalization on each data sample in the labeled source - domain dataset I and the unlabeled target - domain dataset I to obtain the labeled source - domain dataset II and the unlabeled target - domain dataset II; S23. Add Gaussian noise and time - domain random cropping to each data sample in the labeled source - domain dataset II and the unlabeled target - domain dataset II to obtain the labeled source - domain dataset III and the unlabeled target - domain dataset III. The signal - to - noise ratio SNR of the Gaussian noise is 20 dB.
4. The fault diagnosis method for pumping units based on channel attention and invariance learning according to claim 1, characterized in that: When constructing the fault diagnosis model in S3, input the labeled source - domain dataset III and the unlabeled target - domain dataset III into the fault diagnosis model for training, and output the fault category and location information to obtain a trained fault diagnosis model. The specific process is: S31. Input the labeled source - domain dataset III and the unlabeled target - domain dataset III into the multi - scale channel attention unit for feature extraction, and output the features corresponding to each data sample to obtain the feature set of the labeled source - domain dataset III and the feature set of the unlabeled target - domain dataset III; S32. Input the feature sets of the labeled source domain dataset III and the unlabeled target domain dataset III obtained in S31 into the classifier, and respectively output the predicted probabilities of the source domain data and the target domain data. Generate pseudo-labels for the corresponding target domain data according to the predicted probabilities of the target domain data. S33. Input the feature sets of the labeled source domain dataset III and the unlabeled target domain dataset III obtained in S31, as well as the predicted probabilities of the source domain data and the pseudo-labels of the target domain data output in S32, into the invariance feature learning unit. Calculate the source domain cross-entropy loss according to the feature set of the labeled source domain dataset III and the predicted probabilities of the source domain data, calculate the target domain pseudo-label loss according to the feature set of the unlabeled target domain dataset III and the pseudo-labels of the target domain data, and calculate the gradient inner product according to the source domain cross-entropy loss and the target domain pseudo-label loss. S34. Input the feature sets of the labeled source domain dataset III and the unlabeled target domain dataset III obtained in S31, as well as the predicted probabilities of the source domain data, the predicted probabilities of the target domain data, and the pseudo-labels of the target domain data output in S32, into the joint domain adaptation unit. Use the conditional adversarial network to obtain the optimal domain discriminator loss function, generate the conditional distribution, and perform distribution alignment using the joint maximum mean discrepancy. S35. The fault diagnosis model further includes a multi-objective joint optimization unit. Input the source domain cross-entropy loss and the gradient inner product obtained in S33, as well as the optimal domain discriminator loss function and the distribution alignment obtained in S34, into the multi-objective joint optimization unit for integration, and use the integrated result as the total loss function of the fault diagnosis model. S36. According to the predicted probabilities of the target domain data output in S32, add an entropy weighting strategy to screen the high-confidence target domain prediction results, and output the fault category and location information. The specific process is as follows: Calculate the entropy value based on the predicted probability of the target domain data : Among them, is the predicted probability of the target domain data; According to the entropy value Generate weights , and screen the prediction results of the target domain with high confidence; So far, the trained fault diagnosis model is obtained.
5. The fault diagnosis method for pumping units based on channel attention and invariance learning according to claim 4, characterized in that: In S31, the multi-scale channel attention unit sequentially includes a convolutional layer 1, a batch normalization layer 1, a ReLU activation layer 1, a max pooling layer 1, an Msk-ECA module, a convolutional layer 2, a batch normalization layer 2, a ReLU activation layer 2, a max pooling layer 2, an Msk-ECA module, a convolutional layer 3, a batch normalization layer 3, a ReLU activation layer 3, a max pooling layer 3, an Msk-ECA module, a convolutional layer 4, a batch normalization layer 4, a ReLU activation layer 4, and a max pooling layer 4. The Msk-ECA module sequentially includes a global average pooling layer, a multi-scale convolutional interaction layer, a feature weighted fusion layer, and a Sigmoid activation layer. The global average pooling layer compresses the input feature map along the time dimension to generate a channel descriptor. The multi-scale convolutional interaction layer parallelly extracts channel interaction features based on the channel descriptor using three one-dimensional convolutional kernels with different sizes. The feature weighted fusion layer adds the channel interaction features to obtain the added channel interaction features. The Sigmoid activation layer generates the final channel weights by Sigmoid activating the channel interaction features after addition, weights the feature map input to the global average pooling layer using the final channel weights, outputs a channel-weighted feature map, and obtains corresponding features based on the channel-weighted feature map.
6. The fault diagnosis method for pumping units based on channel attention and invariance learning according to claim 4, wherein: In S33, the feature set of the labeled source domain dataset III and the feature set of the unlabeled target domain dataset III obtained in S31, as well as the predicted probability of the source domain data and the pseudo-labels of the target domain data output in S32, are input into the invariance feature learning unit. The source domain cross-entropy loss is calculated based on the feature set of the labeled source domain dataset III and the predicted probability of the source domain data, the target domain pseudo-label loss is calculated based on the feature set of the unlabeled target domain dataset III and the pseudo-labels of the target domain data, and the gradient inner product is calculated based on the source domain cross-entropy loss and the target domain pseudo-label loss. The specific process is as follows: Calculate the source domain cross-entropy loss based on the feature set of the labeled source domain dataset Ⅲ obtained from S31 and the predicted probability of the source domain data output by S32 : Among them, represents the number of source domain data samples in the current batch, represents the total number of fault categories, represents the one-hot encoding of the true label of the i-th sample in the source domain dataset in the c-th class, represents the predicted probability that the classifier predicts that the i-th data sample belongs to the c-th class; Calculate the target domain pseudo-label loss based on the feature set of the unlabeled target domain dataset Ⅲ obtained from S31 and the pseudo-labels of the target domain data output by S32 : Among them, represents the number of target domain samples in the current batch, is the pseudo-label of the j-th sample in the target domain dataset for the c-th class, represents the prediction probability that the classifier assigns the -th sample to the c-th class; Calculate the source domain cross-entropy loss and the target domain pseudo-label loss for the fault diagnosis model parameters inner product of gradients : Among them, represents the i-th trainable parameter in the fault diagnosis model.
7. The fault diagnosis method for pumping units based on channel attention and invariance learning according to claim 4, characterized in that: In S34, the feature set of the labeled source domain dataset III and the feature set of the unlabeled target domain dataset III obtained in S31, as well as the predicted probability of the source domain data, the predicted probability of the target domain data, and the pseudo-labels of the target domain data output in S32, are input into the joint domain adaptation unit. The optimal domain discriminator loss function is obtained using a conditional adversarial network, and a conditional distribution is generated. The joint maximum mean discrepancy is used for distribution alignment. The specific process is as follows: Obtain the corresponding feature vectors according to the features in the feature set of the labeled source domain dataset Ⅲ obtained in S31, obtain the corresponding feature vectors according to the features in the feature set of the unlabeled target domain dataset Ⅲ obtained in S31, input the two parts of the obtained feature vectors, the predicted probability of the source domain data output by S32, and the predicted probability of the target domain data into the conditional adversarial network, calculate the outer product of each feature vector and the corresponding predicted probability as the corresponding conditional feature, obtain all conditional features, and use the gradient reversal layer to train the domain discriminator according to all conditional features to obtain the best domain discriminator, and calculate the loss function of the best domain discriminator and generate the conditional distribution of each data sample; Among them, is the feature vector of the i-th source domain data feature, is the feature vector of the j-th target domain data feature, is the predicted probability of the i-th source domain data, is the predicted probability of the j-th target domain data, is the conditional feature of the i-th source domain data, is the conditional feature of the i-th target domain data, and D(·) is the output of the domain discriminator; Meanwhile, based on the feature set of the labeled source domain dataset III and the feature set of the unlabeled target domain dataset III obtained in S31, the pseudo-labels of the target domain data output in S32, and the conditional distribution of each data sample, the marginal distribution and the conditional distribution of the same-class data samples in the source domain dataset and the target domain dataset are synchronously aligned using the joint maximum mean discrepancy: Among them, represents distribution alignment, represents the set of samples belonging to the c-th class in the source domain dataset, represents a sample belonging to the c-th class in the source domain dataset, represents the set of samples belonging to the c-th class in the target domain dataset, represents a sample belonging to the c-th class in the target domain dataset, represents the number of samples of the c-th class in the source domain dataset, represents the number of samples of the c-th class in the target domain dataset, represents the Gaussian kernel mapping function, represents the norm in the RKHS space.
8. A fault diagnosis method for pumping units based on channel attention and invariance learning according to claim 4, characterized in that: In S35, the total loss function of the fault diagnosis model is: Among them, , , are weighing coefficients, = 0.5, = 1, = 0.3, is a penalty factor, = 0.
1.
9. The fault diagnosis method for pumping units based on channel attention and invariance learning according to claim 1, wherein: In S3, the alarm logic of the trained fault diagnosis model is: When the maximum predicted probability and the entropy value are met, a first-level alarm is triggered; When or is true, a secondary warning is triggered.
10. A pumping unit fault diagnosis system based on channel attention and invariance learning, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-9.
Citation Information
Patent Citations
Mechanical fault diagnosis method under variable working condition based on nuclear sensitivity alignment network
CN115935187A
Rolling bearing fault diagnosis method based on multi-scale convolutional shrinkage network and unsupervised domain adaptation
CN118670721A
Gearbox cross-domain fault diagnosis method based on branch attention contrast transfer learning
CN118758594A
Cross-working-condition rotor unknown fault diagnosis method
CN120046002A
Method and system of machine fault classification using label-consistent convolutional dictionary learning
US20250068892A1