Fault diagnosis method for wet clutch shifting system of agricultural machinery based on closed continuous time unit

By constructing a diagnostic model FaDCfCNet based on closed continuous time units, and combining experimental data simulation and denoising processing, the complexity and data limitations of existing diagnostic methods are solved, enabling efficient fault identification of wet clutch shifting systems in agricultural machinery and ensuring the safety and reliability of unmanned operation.

CN121743679BActive Publication Date: 2026-04-28SHANDONG AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG AGRICULTURAL UNIVERSITY
Filing Date
2026-02-26
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing fault diagnosis methods for wet clutch shifting systems in agricultural machinery suffer from problems such as complex modeling, feature redundancy and modal aliasing, and data-limited diagnostic performance. These methods make it difficult to identify and handle faults in a timely manner in an unmanned driving environment, thus affecting operational safety and efficiency.

Method used

A test bench for a wet clutch shifting system was built, and data was obtained through fault simulation. A diagnostic model FaDCfCNet based on closed continuous time units was constructed by using a fully ensemble empirical mode decomposition adaptive noise and improved wavelet threshold fusion denoising method. The denoised data was used for training and testing to achieve accurate identification of fault modes.

Benefits of technology

It improves the accuracy and efficiency of fault diagnosis, ensures the operational safety of unmanned agricultural machinery, and enhances the reliability and adaptability of the wet clutch shifting system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121743679B_ABST
    Figure CN121743679B_ABST
Patent Text Reader

Abstract

The application discloses a fault diagnosis method for a wet clutch shifting system of an agricultural machine based on a closed continuous time unit, relates to the technical field of intelligent monitoring of agricultural machines, and provides a reliable data basis for diagnosis by building a special test bench to obtain accurate fault simulation data; a complete set of empirical mode decomposition adaptive noise and an improved wavelet threshold fusion denoising method are used to realize efficient denoising of test data and improve the quality of feature data. The constructed FaDCfCNet diagnosis model can accurately identify the fault mode of the wet clutch shifting system by relying on the closed continuous time unit. The method effectively solves the problems of complex modeling, easy occurrence of feature redundancy modal aliasing and data limitation of diagnosis performance of existing diagnosis methods, greatly improves the accuracy and efficiency of fault diagnosis of the wet clutch shifting system, effectively guarantees the operation safety of unmanned agricultural machines, and improves the reliability of the wet clutch shifting system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent monitoring technology for agricultural machinery, specifically to a fault diagnosis method for a wet clutch shifting system of agricultural machinery based on a closed continuous time unit. Background Technology

[0002] Agricultural machinery is the core power equipment in modern agricultural production, widely used in all stages of operations including tillage, sowing, plant protection, harvesting, and transportation. This agricultural machinery is not only suitable for the transmission systems of heavy-duty equipment such as high-horsepower tractors, but also for small and medium-sized agricultural machinery such as tillers used in hilly and mountainous conditions, demonstrating good versatility and adaptability. With the development of agricultural mechanization, modern agricultural production places higher demands on operational efficiency and quality, driving the continuous development of agricultural machinery towards large-scale, automated, and intelligent production. Against this backdrop, new agricultural production models, represented by unmanned farms, are gradually emerging, giving rise to driverless agricultural machinery.

[0003] However, current research on unmanned agricultural machinery focuses primarily on positioning and automatic navigation technologies, neglecting the reliability of the transmission system itself. Modern agricultural machinery utilizes wet clutch shifting systems, which are highly integrated electromechanical-hydraulic systems. Failures in these systems can not only threaten operational safety but also lead to significant economic losses due to delays in planting seasons. Furthermore, in unmanned driving mode, faults are difficult to identify and address promptly using human experience. This makes intelligent fault diagnosis of the shifting system a critical challenge that urgently needs to be overcome in agricultural unmanned driving technology.

[0004] Currently, fault diagnosis methods for gear shifting systems employ the following approaches: The first is a physical model-based approach, which relies on establishing accurate dynamic differential equations. This requires a high level of understanding of the system mechanism, involves complex modeling, and has limited application scope. The second is a signal processing-based approach, utilizing time-frequency analysis tools such as short-time Fourier transform and Wigner-Ville distribution to extract fault features. However, this approach is prone to feature redundancy and mode aliasing. The third is a data science-based intelligent approach, which uses deep learning architectures such as convolutional neural networks, recurrent neural networks, and Transformers to automatically learn fault representations. This approach has strong adaptability to complex operating conditions, but its diagnostic performance is still limited by the completeness and annotation quality of the training data. Summary of the Invention

[0005] In order to solve the above-mentioned technical problems, this application proposes the following technical solution:

[0006] In a first aspect, embodiments of this application provide a fault diagnosis method for a wet clutch shifting system in agricultural machinery based on a closed continuous time unit, including:

[0007] A test bench for a wet clutch shifting system was built, and clutch hydraulic test data under normal and typical fault modes were obtained through fault simulation.

[0008] The collected test data were denoised using a denoising method that combines full set empirical mode decomposition with adaptive noise and improved wavelet threshold fusion.

[0009] A diagnostic model FaDCfCNet based on closed continuous time units is constructed. The FaDCfCNet is trained and tested using denoised test data. The trained FaDCfCNet can accurately identify fault modes.

[0010] In one possible implementation, the construction of a wet clutch shifting system test bench, and the acquisition of clutch hydraulic test data under normal and typical fault modes through fault simulation, includes:

[0011] A test bench for a wet clutch shifting system was constructed, comprising a wet clutch assembly, a hydraulic system, a measurement and control system, and a fault simulation module. The wet clutch assembly includes a first clutch responsible for starting and the low-speed hydrostatic drive range, and a second clutch responsible for the power split range under normal operating conditions. The hydraulic system includes an oil tank, an oil pump, and a control valve assembly. The measurement and control system includes a data acquisition card and an oil pressure sensor. The fault simulation module includes a throttle valve, a speed control valve, and a gate valve.

[0012] Typical fault modes, including oil leakage, oil passage blockage, clutch piston sticking, and solenoid valve spool sticking, were simulated by adjusting the throttle valve, speed control valve, gate valve, and inserting shims.

[0013] The pressure signal during clutch engagement is acquired by a data acquisition card, and sample data under normal and four fault modes are obtained during the complete gear shifting process from clutch engagement to disengagement.

[0014] In one possible implementation, the denoising process of the acquired test data using a fusion denoising method of full set empirical mode decomposition adaptive noise and improved wavelet threshold includes:

[0015] First, the original signal is processed using the fully ensemble empirical mode decomposition adaptive noise CEEMDAN. The decomposition yields the IMF components and residuals:

[0016]

[0017] in, , This is the noise amplitude coefficient. It is a Gaussian white noise sequence;

[0018] Perform Empirical Mode Decomposition (EMD) on each noisy signal to obtain the th... The IMF component, the first The IMFs are:

[0019]

[0020] in: The set number refers to the total number of times EMD decomposition is performed after adding different noise levels. For the first After adding white noise for the first time, the result obtained by EMD is the first... One IMF component;

[0021] Repeat the decomposition process until the remaining signal can no longer be decomposed into an IMF. The original signal can be represented as:

[0022]

[0023] in: For IMF quantity, These are the residual components;

[0024] Calculate the original signal Cross-correlation coefficients with various IMFs ,filter IMF values ​​greater than a preset value retain valid features:

[0025]

[0026]

[0027] in: For time delay parameters, For the delay range, For cross-correlation function, This indicates the absolute value operation;

[0028] Finally, an improved wavelet threshold function is used to denoise the filtered IMF components to obtain a high-fidelity hydraulic signal. The improved wavelet threshold function is as follows:

[0029]

[0030] in, For the threshold, It is an exponential decay factor. For symbolic functions, their definition is as follows:

[0031] .

[0032] In one possible implementation, the FaDCfCNet includes: a multi-scale wavelet Kolmogorov-Arnold theorem layer (MWKAN), a wavelet parameterized closed continuous-time unit, a continuous-time dynamic encoder, an inverted-dimensional decoder, and a physically-aware optimized target loss function.

[0033] In one possible implementation, the learnable wavelet transform path uses the Mexican Hat wavelet as the kernel function, defined as:

[0034]

[0035] in: The value of the Mexican Hat wavelet kernel function. For continuous time or dimensionless independent variables after scaling, Represented by natural constant The base is an exponential function, which is used to describe the exponential decay characteristics of wavelet functions;

[0036] For the input vector After being extended to a high-dimensional feature space by linear transformation, normalized coordinates are defined as follows:

[0037]

[0038] in: For normalized coordinates, These are intermediate features after linear projection. For learnable translation parameters, For learnable scale parameters, To prevent the minimum value where the denominator is zero;

[0039] No. The neuron in the first Activation response at wavelet scale for:

[0040]

[0041] Final wavelet path output It is a weighted sum of the responses of all basis functions:

[0042]

[0043] in: The wavelet activation function maps the normalized coordinates to the activation intensity of neurons, directly affecting the output calculation of subsequent wavelet paths. It is a wavelet activation tensor. It outputs a mixing matrix;

[0044] A parallel residual basis path based on the SiLU activation function is introduced. :

[0045]

[0046] in: This represents the basic output weight matrix. The SiLU activation function;

[0047] Through a linear superposition mechanism, two function spaces with different properties are orthogonally projected and reorganized:

[0048]

[0049] in: Let be the scale of the wavelet basis functions. For learnable weights, This indicates element-wise multiplication. It is a linear transformation matrix. For the first Learnable translation parameters of a wavelet basis For the first Learnable scale parameters of a small wavelet basis.

[0050] In one possible implementation, the state update of the wavelet parameterized closed continuous-time unit is as follows:

[0051]

[0052] in: and Processing the joint state vector by MWKAN get, After activation by the Sigmoid function, it becomes a time decay gate. After Tanh activation, it becomes a candidate equilibrium state. It is the Sigmoid activation function. This is the input feature vector from the previous time step. This is the Tanh activation function.

[0053] In one possible implementation, the continuous-time dynamic encoder is composed of multiple wavelet-parameterized closed continuous-time units stacked together, for the Layer, moment Hidden state Evolution follows a closed-form solution:

[0054]

[0055] in, Represents the composite function operator. Indicates the first Layer at time Input, Indicates the previous moment The hidden state, This represents the time interval between two sampling points, enabling the model to adapt to non-uniformly sampled data. After layer stacking, the output contains a rich sequence of local dynamic features. .

[0056] In one possible implementation, the inverted dimension decoder outputs the encoder. Execution dimension transpose Each decoding layer contains a grouped query sparse cue channel self-attention mechanism (GQA-SPCS) and a phantom feedforward neural network (GhostFFN), where:

[0057] The GQA-SPCS estimates channel importance through a gating function and adaptively retains key channels. The attention weights are calculated as follows:

[0058]

[0059] in, For learnable temperature parameters, Represents a sparse mask. For queries on attention mechanisms, For the healthy development of attention mechanisms, For each attention head, the feature dimensions, It is a normalized exponential function;

[0060] Attention output is Then on Perform residual connection to obtain ,in: This is the value matrix in the attention mechanism. To Features after residual connection This represents a random deactivation regularization operation;

[0061] Then, GhostFFN generates high-dimensional features through the Ghost Module, and the output is:

[0062]

[0063] in: This represents a lightweight feature extension module that efficiently expands the number of feature channels without increasing computational cost by generating a small number of intrinsic features through standard convolutions and phantom features through inexpensive depthwise convolutions. It can be used to replace the computationally intensive extension layers in traditional feedforward neural networks (FFNs). For linear projection, Linear bias;

[0064] Finally, after passing through two decoding layers, the data is sent to the classification head, which then performs adaptive average pooling, layer normalization, and linear feature mapping to obtain the final class prediction. :

[0065]

[0066] in: The linear projection weight matrix of the output layer. The linear projection weight matrix of the bottleneck layer. This is a fixed-length context vector after layer normalization. For the bias of the bottleneck layer, This is the bias for the output layer.

[0067] In one possible implementation, the physical perception optimization target loss function for:

[0068]

[0069]

[0070]

[0071]

[0072]

[0073] in: For classification cross-entropy loss, For frequency domain consistency constraints, For wavelet sparse regularization, For EWC regularization, , and To balance the hyperparameters of the various loss weights, This represents the total number of fault categories. The first output of the classifier The predicted value of Logits. For indicator functions, Mean pooling of features The spectrum in The amplitude at the frequency point, Represents the original input signal The spectrum in The amplitude at the frequency point, and They are respectively and The maximum amplitude of the spectrum is used to normalize the amplitude to eliminate scale differences. For the number of neurons, The number of wavelet basis functions corresponding to each neuron These are the diagonal elements of the Fisher information matrix. These are the current model parameters. These are the optimal parameters after convergence in the previous stage.

[0074] Secondly, embodiments of this application provide a fault diagnosis device for an agricultural machinery wet clutch shifting system based on a closed continuous time unit, comprising:

[0075] The data acquisition module is used to build a test bench for wet clutch shifting system and obtain clutch hydraulic test data under normal and typical fault modes through fault simulation.

[0076] The signal denoising module is used to denoise the acquired test data using a fusion denoising method of adaptive noise reduction and improved wavelet thresholding with complete set empirical mode decomposition.

[0077] The fault diagnosis module is used to construct a diagnostic model FaDCfCNet based on closed continuous time units. The FaDCfCNet is trained and tested using denoised test data. The trained FaDCfCNet can accurately identify fault modes.

[0078] In this embodiment, a dedicated test bench is built to obtain accurate fault simulation data, providing a reliable data foundation for diagnosis. A fully ensemble empirical mode decomposition adaptive noise and improved wavelet threshold fusion denoising method is employed to achieve efficient denoising of the test data and improve the quality of feature data. The constructed FaDCfCNet diagnostic model, based on closed continuous time units, can accurately identify fault modes in wet clutch shifting systems. This method effectively solves the problems of complex modeling, easy feature redundancy mode aliasing, and data-limited diagnostic performance in existing diagnostic methods, significantly improving the accuracy and efficiency of fault diagnosis in agricultural machinery wet clutch shifting systems, effectively ensuring the safety of unmanned agricultural machinery operations, and enhancing the reliability of wet clutch shifting systems. Attached Figure Description

[0079] Figure 1 A flowchart illustrating a fault diagnosis method for an agricultural machinery wet clutch shifting system based on a closed continuous time unit, provided in an embodiment of this application;

[0080] Figure 2 A comparison diagram of clutch pressure signal waveforms under four typical fault modes and normal conditions provided for embodiments of this application;

[0081] Figure 3 This is a schematic diagram illustrating how CEEMDAN, combined with an improved wavelet threshold, filters and denoises noise in the acquired data, as provided in an embodiment of this application.

[0082] Figure 4 This is a schematic diagram of the overall architecture of the FaDCfCNet model provided in the embodiments of this application;

[0083] Figure 5 A schematic diagram of the multi-scale wavelet Kolmogorov-Arnold theorem layer structure provided for the embodiments of this application;

[0084] Figure 6 This is a schematic diagram of the state update of a wavelet parameterized closed continuous-time unit provided in an embodiment of this application;

[0085] Figure 7 This is a schematic diagram of the inverted dimension decoder structure provided in an embodiment of this application;

[0086] Figure 8 This is a schematic diagram of a fault diagnosis device for an agricultural machinery wet clutch shifting system based on a closed continuous time unit, provided in an embodiment of this application. Detailed Implementation

[0087] The present solution will now be described in conjunction with the accompanying drawings and specific embodiments.

[0088] See Figure 1 The fault diagnosis method for agricultural machinery wet clutch shifting system based on closed continuous time unit provided in this application includes:

[0089] S101, a test bench for a wet clutch shifting system was built, and clutch hydraulic test data under normal and typical fault modes were obtained through fault simulation.

[0090] The test bench is based on the wet clutches C1 (controlling the start / low speed hydrostatic range HY) and C2 (controlling the normal operation power split range HM) of the general agricultural machinery gearbox. It is equipped with a hydraulic system (oil tank, oil pump, control valve group), wet clutch assembly (C1 and C2 clutches), measurement and control system (LabVIEW program, NI USB-6009 data acquisition card, oil pressure sensor) and fault simulation module (throttle valve, speed control valve, gate valve, etc.), which can accurately simulate typical faults in actual operation.

[0091] Simulate 5 system states:

[0092] Normal state (H): No fault intervention, acquiring reference pressure signal during clutch engagement;

[0093] Oil leakage (F1): Adjust the throttle valve opening to allow some of the clutch input hydraulic oil to flow back to the oil tank, simulating different degrees of leakage;

[0094] Oil passage blockage (F2): Adjust the speed control valve to reduce the cross-sectional area of ​​the pipe, change the clutch oil flow rate, and simulate different degrees of blockage;

[0095] Clutch piston jamming (F3): Close the gate valve to block the clutch inlet, simulating the closed volume pressure characteristics of a completely jammed piston;

[0096] Solenoid valve spool jamming (F4): Inserting shims of different thicknesses between the solenoid valve return spring and the limit block changes the spool limit position, simulating spool jamming.

[0097] The shift signal is triggered by a LabVIEW program, and the oil pressure sensor collects the pressure signal of the clutch engagement process at a frequency of 100Hz. The signal is then stored as a table file by the NI USB-6009 data acquisition card. A total of 15,000 samples were collected (3,000 samples for each state), covering the three stages of "piston movement - contact point - working pressure establishment".

[0098] This embodiment focuses on the wet clutch shifting system of agricultural machinery gearboxes, and builds a fault diagnosis test bench that can simulate real working conditions. The specific configuration is as follows:

[0099] Hydraulic system: A 20L high-pressure oil tank and a gear pump with a rated pressure of 16MPa are selected, along with a DN10 throttle valve (simulating oil leakage), a speed control valve (simulating oil passage blockage), a gate valve (simulating clutch piston jamming) and sealed pipelines.

[0100] Wet clutch assembly: It adopts C1 and C2 wet clutches (rated torque 250 N·m) matched with the gearbox. C1 controls the hydrostatic range of starting / low speed, and C2 controls the power split range of normal operation.

[0101] Measurement and control system: The LabVIEW 2021 program is deployed as the control core, and it is equipped with an NI USB-6009 data acquisition card (12-bit sampling accuracy) and a diffused silicon oil pressure sensor with an accuracy of 0.5%FS (range 0~20MPa). The sensor is installed at the clutch oil inlet.

[0102] Fault simulation: Three levels of oil leakage are simulated by adjusting the throttle valve opening (20% / 40% / 60%), three levels of oil passage blockage are simulated by adjusting the speed control valve opening (30% / 50% / 70%), the gate valve is closed to block the C1 clutch inlet to simulate complete piston jamming, and 0.5mm / 1.0mm / 1.5mm PTFE gaskets are inserted between the solenoid valve return spring and the limit block to simulate three levels of valve core jamming.

[0103] During the data acquisition process, the ambient temperature was set to 25℃, and anti-wear hydraulic oil with a viscosity grade of 40mm² / s was selected. The shift control signal was sent through the LabVIEW program, and the oil pressure sensor collected the pressure signal of the clutch throughout the entire shift process at a frequency of 100Hz. Each sample contained pressure data at 64 time points.

[0104] See Figure 2 Pressure sample data were collected under normal conditions (H) and four fault modes (F1~F4), with 3,000 sets for each condition, for a total of 15,000 sets of samples. The data were stored in CSV format files, covering the differences in operating conditions under different fault levels.

[0105] S102, the collected test data is denoised using a fusion denoising method of fully ensemble empirical mode decomposition adaptive noise and improved wavelet threshold.

[0106] See Figure 3 A denoising method combining CEEMDAN and an improved wavelet threshold fusion was used to preprocess the acquired pressure signal. The specific steps are as follows:

[0107] First, the original signal is processed using the fully ensemble empirical mode decomposition adaptive noise CEEMDAN. The decomposition yields the IMF components and residuals:

[0108]

[0109] in, , This is the noise amplitude coefficient. It is a Gaussian white noise sequence.

[0110] Perform Empirical Mode Decomposition (EMD) on each noisy signal to obtain the th... The IMF component, the first The IMFs are:

[0111]

[0112] in: The set number refers to the total number of times EMD decomposition is performed after adding different noise levels. For the first After adding white noise for the first time, the result obtained by EMD is the first... One IMF component.

[0113] Repeat the decomposition process until the remaining signal can no longer be decomposed into an IMF. The original signal can be represented as:

[0114]

[0115] in: For IMF quantity, These are the residual components.

[0116] Calculate the original signal Cross-correlation coefficients with various IMFs ,filter IMF values ​​greater than a preset value retain valid features:

[0117]

[0118]

[0119] in: For time delay parameters, For the delay range, For cross-correlation function, This indicates the absolute value operation;

[0120] Finally, an improved wavelet threshold function is used to denoise the filtered IMF components. An exponential decay factor is introduced to construct the improved threshold function, which solves the defects of traditional soft / hard thresholds and obtains a high-fidelity hydraulic signal. The improved wavelet threshold function is as follows:

[0121]

[0122] in, For the threshold, It is an exponential decay factor; This is the sign function. This function is continuous at the threshold point, has no fixed deviation, and balances noise suppression with signal detail preservation. Wavelet transform is applied to the filtered IMFs, the wavelet coefficients are processed using the improved threshold function, and then the signal is reconstructed through inverse transform.

[0123] S103, construct a diagnostic model FaDCfCNet based on closed continuous time unit, train and test the FaDCfCNet using denoised test data, and the trained FaDCfCNet accurately identifies fault modes.

[0124] See Figure 4 In this embodiment, the FaDCfCNet includes: a multi-scale wavelet Kolmogorov-Arnold theorem layer (MWKAN Layer), a wavelet parameterized closed continuous-time unit, a continuous-time dynamic encoder, an inverted-dimensional decoder, and a physically-aware optimized target loss function.

[0125] The trained FaDCfCNet model was input with 3000 sets of test data. Accuracy, Precision, Recall and F1-score were used as performance evaluation metrics. Seven cutting-edge methods were selected for comparison, including DDHGCN, TFN, MSCNN-LSTM-CBAM-SE, LiConvFormer, MRSCNN, WaveletKernelNet and MSDC-Swin-T.

[0126] Specifically, DDHGCN is a deep, dynamic, high-order graph convolutional network composed of a dynamic graph learning module, an adaptive high-order graph convolutional module, and a residual convolutional module. TFN is an interpretable neural network that incorporates time-frequency transformation techniques. MSCNN-LSTM-CBAM-SE is a deep ensemble network that combines the multi-scale features of MSCNN with the temporal features of LSTM. LiConvFormer is a lightweight fault diagnosis model. MRSCNN employs soft thresholding techniques for noise reduction. WaveletKernelNet replaces the first convolutional layer of a standard CNN with consecutive wavelet convolutional layers. MSDC-Swin-T is a multi-scale dilated convolutional method that achieves the fusion of global and local features.

[0127] In the above comparison methods, only the input sample size was adjusted, while the network structure and optimal hyperparameter settings remained consistent with those in the original papers. Each method was run independently ten times, and the average results were taken to ensure statistical stability. The random seed was randomly selected from the range of 0 to 100.

[0128] Referring to Table 1, the average accuracies of DDHGCN, TFN, MSCNN-LSTM-CBAM-SE, LiConvFormer, MRSCNN, WaveletKernelNet, and MSDC-Swin-T on the test set are 96.07%, 92.93%, 92.83%, 95.29%, 88.80%, 93.39%, and 93.52%, respectively, all of which are lower than the proposed FaDCfCNet method.

[0129] Furthermore, FaDCfCNet maintains a diagnostic accuracy of 98.06% even under multi-source environmental noise interference. Even without employing the Ceemdan combined wavelet thresholding denoising method, its accuracy remains at 94.64%, fully validating the effectiveness of the proposed denoising method.

[0130] Table 1 Comparison of average diagnostic performance on different model test sets

[0131]

[0132] To address the limitations of spectral bias in deep networks (such as MLPs) and the insufficient sensitivity of B-spline basis functions in KANs to high-frequency transient features, we propose the Multi-Scale Wavelet Kolmogorov-Arnold Theorem Layer (MWKAN Layer). This layer achieves two key innovations: first, it replaces the B-spline basis functions in the original KAN with Mexican Hat wavelets, which possess excellent time-frequency localization properties; second, it designs a dual-stream residual complementarity mechanism, which overcomes the potential gradient vanishing risk of pure wavelet networks by decoupling low-frequency manifolds from high-frequency details, significantly enhancing the model's feature extraction capability and training stability.

[0133] like Figure 5 First, the B-spline basis functions in the original KAN are... It is parameterized as a linear combination of a set of multi-scale wavelet basis functions. Specifically, for the input vector... We designed a dual-path architecture to ensure both the richness of feature extraction and the stability of gradient flow.

[0134] To capture local abrupt changes and multi-frequency components of the signal, we use the Mexican Hat wavelet as the kernel function. Unlike the fixed scale and translation parameters in traditional wavelet transforms, we set them as learnable parameters. The Mexican Hat wavelet is defined as:

[0135]

[0136] in: The value of the Mexican Hat wavelet kernel function. For continuous time or dimensionless independent variables after scaling, Represented by natural constant The base is an exponential function, which is used to describe the exponential decay characteristics of wavelet functions.

[0137] In the forward propagation of the network, input First, the space is expanded to a higher dimension feature space through a linear transformation. For each hidden layer neuron and each small wavelet We define normalized coordinates :

[0138]

[0139] in: For normalized coordinates, These are intermediate features after linear projection. For learnable translation parameters, For learnable scale parameters, To prevent the minimum value where the denominator is zero, therefore, the first... The neuron in the first Activation response at wavelet scale for:

[0140]

[0141] Training using the above formulas It can adaptively adjust the size of the receptive field to match fault characteristics at different frequencies, and This allows the activation function to move within the feature space to locate key information. The final wavelet path output... It is a weighted sum of the responses of all basis functions:

[0142]

[0143] in: The wavelet activation function maps the normalized coordinates to the activation intensity of neurons, directly affecting the output calculation of subsequent wavelet paths. It is a wavelet activation tensor. It outputs a mixing matrix;

[0144] To alleviate the gradient vanishing problem in deep networks and preserve low-frequency background information, we introduce a parallel residual basis path based on the SiLU activation function. :

[0145]

[0146] in: This represents the basic output weight matrix. The activation function is SiLU. SiLU was chosen over ReLU because of its smooth non-monotonicity and differentiability everywhere, which facilitates gradient propagation on the manifold. This path is not merely a simple skip connection; mathematically, it acts as a global approximator. It ensures gradient flow even in the early stages of training, when the wavelet basis has not yet converged or the input has drifted drastically. It is still possible to pass The path is backpropagated without loss, thus maintaining the training stability of the entire network.

[0147] The final output of the MWKAN layer is not a simple feature superposition, but a complementary fusion of local transient features and global semantic features. Through a linear superposition mechanism, we orthogonally project and reorganize two function spaces with different properties:

[0148]

[0149] in: Let be the scale of the wavelet basis functions. For learnable weights, This indicates element-wise multiplication. It is a linear transformation matrix. For the first Learnable translation parameters of a wavelet basis For the first Learnable scale parameters of a small wavelet basis.

[0150] Finally, to ensure that the wavelet basis can cover the distribution range of the input data in the initial stage, we adopt a grid initialization strategy. Translation parameters Initialized to be uniformly distributed in Grid points of the interval:

[0151]

[0152] scale parameter Initialize to This ensures that the initial wavelet has a standard bandwidth.

[0153] at discrete time step In this case, the state update of a closed continuous-time network is represented as an explicit gating form:

[0154]

[0155] here, It is a time decay gate, which controls the degree to which the system's memory is retained; It is a candidate equilibrium state, representing the target that the system tends to under the current input.

[0156] in, and Typically, this process is generated by a shared-parameter MLP controller. Because MLPs employ a fully connected structure, it is difficult to effectively distinguish the influence of different time scales or frequency components on system dynamics. Furthermore, this process lacks effective physical interpretability. Therefore, we utilize the proposed MWKAN to directly parameterize this process. Specifically, we will use the current input... and the hidden state of the previous moment Concatenate the data to construct a joint state vector. Then input it into the MWKAN backbone network:

[0157]

[0158] Then, output vector It is divided into two parts along the channel dimension, which are used to model the time decay dynamics and the equilibrium state dynamics, respectively:

[0159]

[0160] Finally, we use the Sigmoid and Tanh functions to map these two parts to intervals with explicit physical meaning, thus enabling direct parameterization. and :

[0161]

[0162] Based on the above derivation, the final forward propagation formula for the wavelet parameterized closed continuous-time unit is:

[0163]

[0164] in: and Processing the joint state vector by MWKAN get, After activation by the Sigmoid function, it becomes a time decay gate. After Tanh activation, it becomes a candidate equilibrium state. It is the Sigmoid activation function. This is the input feature vector from the previous time step. This is the Tanh activation function. See also... Figure 6 This is a schematic diagram of the state update of a wavelet parameterized closed continuous-time unit.

[0165] This formula essentially achieves a higher-precision function approximation than previous gating methods by using a set of learnable wavelet bases. Finally, to prevent covariate shift caused by increasing layer depth and to enhance training stability, we apply layer normalization and Dropout to the output:

[0166]

[0167] Given sensor time series input (in For batch size, (where is the length of the time series), for the Layer, moment Hidden state Evolution follows a closed-form solution:

[0168]

[0169] in, Represents the composite function operator. Indicates the first Layer at time Input, Indicates the previous moment The hidden state, This represents the time interval between two sampling points, enabling the model to adapt to non-uniformly sampled data. After layer stacking, the output contains a rich sequence of local dynamic features. .(here (This represents the dimension of the extracted hidden layer feature vector). The proposed encoder effectively solves the gradient vanishing problem in traditional networks when processing complex nonlinear input signals, while preserving the physical continuity of the signal.

[0170] To better utilize the output At every moment In obtaining the feature vectors, we fully considered the limitations of the Transformer in processing time series. Embedding multiple variables at the same timestamp into a single time token often weakens or even loses the correlation between variables, causing the attention weights to focus more on the relationship between different time positions. However, the correlation between different variables throughout the entire time series often better reflects a failure mode. For example... Figure 7 It uses an inverted-dimensional decoder, which is based on the latent feature tensor of the encoder output. Perform dimension transpose operation:

[0171]

[0172] In this inverted structure, dimension Redefined as the number of tokens, dimensions It is redefined as the feature embedding dimension of each token. Then it goes through a series of inverted dimension decoding layers. Each layer mainly consists of two parts: Grouped Query Sparse Hint Channel Self-Attention Mechanism (GQA-SPCS) and Phantom Feedforward Neural Network (GhostFFN).

[0173] In GQA-SPCS, the input features are first... Perform a linear mapping to obtain a query, key, and value matrix:

[0174]

[0175] Search here use Attention head, key-value pair , Only use Size. , , Rearranged into a multi-head form:

[0176]

[0177] here These are the feature dimensions of each attention head, which are then copied to make... In the head dimension and Alignment is used to achieve multi-head attention with shared parameters. However, considering that all tokens participate in the calculation along the channel dimension, a large number of unrelated or weakly correlated channels can introduce noise. Furthermore, early sparse attention methods typically set a fixed value. Before keeping The attention weight is %. Therefore, to suppress redundant channels, a gating function is introduced to estimate channel importance:

[0178]

[0179] here This represents an averaging operation on the absolute values ​​over the time dimension. This represents the Sigmoid activation function, used to generate normalized channel importance weights. The number of key channels retained for attention calculation is adaptively determined based on these channel weights.

[0180]

[0181] here, This represents the average operation over the channel dimension. To round down, This constraint is used to limit the minimum number of participating channels to ensure the stability of the attention calculation process. Based on this, the attention weight calculation is redefined:

[0182]

[0183] in, For learnable temperature parameters, Represents a sparse mask. For queries on attention mechanisms, For the healthy development of attention mechanisms, For each attention head, the feature dimensions, It is a normalized exponential function.

[0184] Attention output is Then on Perform residual connection to obtain ,in: This is the value matrix in the attention mechanism. To Features after residual connection This represents the random deactivation regularization operation.

[0185] right After layer normalization, the data is fed into GhostFFN. In the standard Transformer architecture, FFN accounts for the majority of the model parameters and computational complexity. To significantly reduce computational overhead while maintaining feature expressiveness, GhostFFN replaces the first-layer extended mapping in FFN with Ghost Modules, decoupling the high-dimensional feature generation process into two parts, specifically:

[0186] Convolution operation through Each convolutional kernel directly generates features. The Ghost Module, however, first uses convolutions to generate a small number of features. Each inherent feature map The formula is expressed as:

[0187]

[0188] in: This represents convolution or matrix multiplication. It is a standard convolutional kernel, but it has been compressed in terms of dimensions. This is the generated intrinsic feature map. Compression ratio is introduced. Inherent number of channels The calculation becomes This means that the computational cost becomes the original .

[0189] In order to obtain all Each channel, Ghost Module for the aforementioned inherent feature map Each channel in the process applies a series of inexpensive linear operations. For The first in Each inherent feature map Through depthwise convolution generate A new Ghost feature:

[0190]

[0191] Final output It is composed of intrinsic feature maps and Ghost feature maps:

[0192]

[0193] Based on the above theory, we designed GhostFFN, replacing the first computationally intensive extension layer in FFN with a Ghost Module to reduce the number of parameters. The formula becomes:

[0194]

[0195] in: This design represents a lightweight feature extension module that efficiently expands the number of feature channels without increasing computational cost by generating a small number of inherent features through standard convolution and generating phantom features through inexpensive depthwise convolution. This module can be used to replace the computationally intensive extension layers in traditional feedforward neural networks. For linear projection, Linear bias;

[0196] Subsequently, layer normalization is performed to stabilize the feature distribution and accelerate convergence:

[0197]

[0198] in, For learnable affine parameters, For input The mean across the feature dimensions, For input Variance in the feature dimension Pick The normalized features are then mapped through a bottleneck MLP to generate the final class prediction. This process integrates dimensionality reduction projection, the SiLU activation function for smoothing nonlinearities, and the Dropout mechanism for regularization:

[0199]

[0200] in: The linear projection weight matrix of the output layer. The linear projection weight matrix of the bottleneck layer. This is a fixed-length context vector after layer normalization. For the bias of the bottleneck layer, This is the bias for the output layer.

[0201] To ensure the model not only performs well on data-driven classification tasks but also learns physically interpretable time-frequency features, we designed a hybrid optimization objective function. This objective function consists of four parts: classification cross-entropy loss, frequency domain consistency constraint, wavelet sparsity regularization, and elastic weight consolidation (EWC) regularization. The total loss function is... The definition is as follows:

[0202]

[0203] in: For classification cross-entropy loss, For frequency domain consistency constraints, For wavelet sparse regularization, For EWC regularization, , and Hyperparameters are used to balance the weights of various losses.

[0204] For the multi-class classification task of fault diagnosis, we employ the standard cross-entropy loss function to minimize the difference between the predicted probability distribution and the true label. For a given input sample... and its corresponding real tags The classification loss is defined as:

[0205]

[0206] in: This represents the total number of fault categories. The first output of the classifier The predicted value of Logits. This is an indicator function.

[0207] The core physical features of mechanical fault signals are often concentrated on specific frequency components. To prevent deep neural networks from losing crucial spectral information during feature extraction, we introduce a physically-aware frequency domain consistency loss.

[0208] For time step and number of feature channels This loss forces latent space characteristics Spectral distribution and original input signal The main frequency components remain consistent. Specifically, we first calculate the mean pooling of the features. and the original input signal Fast Fourier Transform:

[0209]

[0210] Subsequently, we select the index set of the first R dominant frequencies based on the spectral energy of the original signal. Frequency domain consistency loss is defined as the mean square error of the amplitude at these key frequency points:

[0211]

[0212] in: express The spectrum in The amplitude at the frequency point, express The spectrum in The amplitude at the frequency point, and They are respectively and The maximum amplitude of the spectrum is used to normalize the amplitude to eliminate scale differences. Through this constraint, the model is forced to preserve the physical periodicity and impulsive characteristics of the original signal in the feature space, thereby improving its robustness to noise.

[0213] Learnable wavelet scaling parameters were introduced earlier. Translation parameters According to wavelet theory, an effective signal representation should be sparse. To improve the interpretability of the model and prevent overfitting, we apply certain parameters to both sets of parameters. Sparse regularization:

[0214]

[0215] in: For the number of neurons, The number of wavelet basis functions corresponding to each neuron. This regularization term encourages the model to use only the most necessary wavelet basis functions to fit the data, so that the finally learned wavelet parameters better reflect the true local time-frequency structure of the fault signal.

[0216] To enhance the model's stability in continuous learning scenarios, prevent catastrophic forgetting of important parameters in the later stages of training, or lock in key features in small sample situations, we introduce the Elastic Weight Consolidation (EWC) strategy. This method is based on a Bayesian framework and utilizes the Fisher Information Matrix (FIM) to estimate the importance of each parameter.

[0217] Assumption These are the current model parameters. These are the optimal parameters after convergence in the previous stage. The EWC loss is defined as a weighted quadratic penalty for parameter changes:

[0218]

[0219] in These are the diagonal elements of the Fisher information matrix, approximately the second moment of the log-likelihood gradient:

[0220]

[0221] In implementation, when the model reaches a high accuracy rate, we will calculate and update... This mechanism is equivalent to introducing an "elastic spring" into the parameter space. For parameters that are extremely sensitive to classification results (high Fisher information content), strong constraints are imposed to keep them stable; while for parameters that are not sensitive, they are allowed to adapt freely to new data changes.

[0222] Corresponding to the above-described method for diagnosing faults in an agricultural machinery wet clutch shifting system based on a closed continuous time unit, this application also provides an embodiment of a device for diagnosing faults in an agricultural machinery wet clutch shifting system based on a closed continuous time unit.

[0223] See Figure 8 A fault diagnosis device 20 for agricultural machinery wet clutch shifting system based on closed continuous time unit includes:

[0224] The data acquisition module 201 is used to build a test bench for a wet clutch shifting system and acquire clutch hydraulic test data under normal and typical fault modes through fault simulation.

[0225] The signal denoising module 202 is used to denoise the acquired test data using a fusion denoising method of adaptive noise reduction and improved wavelet thresholding based on complete set empirical mode decomposition.

[0226] The fault diagnosis module 203 is used to construct a diagnostic model FaDCfCNet based on a closed continuous time unit. The FaDCfCNet is trained and tested using denoised test data. The trained FaDCfCNet can accurately identify fault modes.

[0227] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects have an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0228] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A fault diagnosis method for a wet clutch shifting system in agricultural machinery based on a closed continuous time unit, characterized in that, include: A test bench for a wet clutch shifting system was built, and clutch hydraulic test data under normal and typical fault modes were obtained through fault simulation. The collected test data were denoised using a denoising method that combines full set empirical mode decomposition with adaptive noise and improved wavelet threshold fusion. A diagnostic model FaDCfCNet based on closed continuous time units is constructed. The FaDCfCNet is trained and tested using denoised test data. The trained FaDCfCNet can accurately identify fault modes. The FaDCfCNet includes: a multi-scale wavelet Kolmogorov-Arnold theorem layer (MWKAN layer), a wavelet parameterized closed continuous-time unit, a continuous-time dynamic encoder, an inverted-dimensional decoder, and a physically-aware optimized target loss function. The MWKAN Layer includes: a learnable wavelet transform path, a residual basis path, and a complementary aggregation module, wherein: The learnable wavelet transform path uses the Mexican Hat wavelet as the kernel function, defined as follows: in: The value of the Mexican Hat wavelet kernel function. For continuous time or dimensionless independent variables after scaling, Represented by natural constant The base is an exponential function, which is used to describe the exponential decay characteristics of wavelet functions; For the input vector After being extended to a high-dimensional feature space by linear transformation, normalized coordinates are defined as follows: in: For normalized coordinates, These are intermediate features after linear projection. For learnable translation parameters, For learnable scale parameters, To prevent the minimum value where the denominator is zero; No. The neuron in the first Activation response at wavelet scale for: Final wavelet path output It is a weighted sum of the responses of all basis functions: in: The wavelet activation function maps the normalized coordinates to the activation intensity of neurons, directly affecting the output calculation of subsequent wavelet paths. It is a wavelet activation tensor. It outputs a mixing matrix; A parallel residual basis path based on the SiLU activation function is introduced. : in: This represents the basic output weight matrix. The SiLU activation function; Through a linear superposition mechanism, two function spaces with different properties are orthogonally projected and reorganized: in: Let be the scale of the wavelet basis functions. For learnable weights, This indicates element-wise multiplication. It is a linear transformation matrix. For the first Learnable translation parameters of a wavelet basis For the first Learnable scaling parameters of a small wavelet basis; The state update of the wavelet parameterized closed continuous-time unit is as follows: in: and Processing the joint state vector by MWKAN get, After activation by the Sigmoid function, it becomes a time decay gate. After Tanh activation, it becomes a candidate equilibrium state. It is the Sigmoid activation function. This is the input feature vector from the previous time step. Use the Tanh activation function; The continuous-time dynamic encoder is composed of multiple wavelet-parameterized closed continuous-time units stacked together. For the first... Layer, moment Hidden state Evolution follows a closed-form solution: in, Represents the composite function operator. Indicates the first Layer at time Input, Indicates the previous moment The hidden state, This represents the time interval between two sampling points, enabling the model to adapt to non-uniformly sampled data. After layer stacking, the output contains a rich sequence of local dynamic features. ; The inverted dimension decoder outputs the encoder. Execution dimension transpose Each decoding layer contains a grouped query sparse cue channel self-attention mechanism (GQA-SPCS) and a phantom feedforward neural network (GhostFFN), where: The GQA-SPCS estimates channel importance through a gating function and adaptively retains key channels. The attention weights are calculated as follows: in, For learnable temperature parameters, Represents a sparse mask. For queries on attention mechanisms, For the healthy development of attention mechanisms, For each attention head, the feature dimensions, It is a normalized exponential function; Attention output is Then on Perform residual connection to obtain ,in: This is the value matrix in the attention mechanism. To Features after residual connection This represents a random deactivation regularization operation; Then, GhostFFN generates high-dimensional features through the Ghost Module, and the output is: in: This design represents a lightweight feature extension module that efficiently expands the number of feature channels without increasing computational cost by generating a small number of inherent features through standard convolution and generating phantom features through inexpensive depthwise convolution. This module can be used to replace the computationally intensive extension layers in traditional feedforward neural networks. For linear projection, Linear bias; Finally, after passing through two decoding layers, the data is sent to the classification head, which then performs adaptive average pooling, layer normalization, and linear feature mapping to obtain the final class prediction. : in: The linear projection weight matrix of the output layer. The linear projection weight matrix of the bottleneck layer. This is a fixed-length context vector after layer normalization. For the bias of the bottleneck layer, For the bias of the output layer; The physical perception optimization objective loss function for: in: For classification cross-entropy loss, For frequency domain consistency constraints, For wavelet sparse regularization, For EWC regularization, , and To balance the hyperparameters of the various loss weights, This represents the total number of fault categories. The first output of the classifier The predicted value of Logits. For indicator functions, Mean pooling of features The spectrum in The amplitude at the frequency point, Represents the original input The spectrum in The amplitude at the frequency point, and They are respectively and The maximum amplitude of the spectrum is used to normalize the amplitude to eliminate scale differences. For the number of neurons, The number of wavelet basis functions corresponding to each neuron These are the diagonal elements of the Fisher information matrix. These are the current model parameters. These are the optimal parameters after convergence in the previous stage.

2. The fault diagnosis method for agricultural machinery wet clutch shifting system based on closed continuous time unit according to claim 1, characterized in that, The aforementioned test bench for the wet clutch shifting system was constructed to obtain clutch hydraulic test data under normal and typical fault modes through fault simulation, including: A test bench for a wet clutch shifting system was constructed, comprising a wet clutch assembly, a hydraulic system, a measurement and control system, and a fault simulation module. The wet clutch assembly includes a first clutch responsible for starting and the low-speed hydrostatic drive range, and a second clutch responsible for the power split range under normal operating conditions. The hydraulic system includes an oil tank, an oil pump, and a control valve assembly. The measurement and control system includes a data acquisition card and an oil pressure sensor. The fault simulation module includes a throttle valve, a speed control valve, and a gate valve. Typical fault modes, including oil leakage, oil passage blockage, clutch piston sticking, and solenoid valve spool sticking, were simulated by adjusting the throttle valve, speed control valve, gate valve, and inserting shims. The pressure signal during clutch engagement is acquired by a data acquisition card, and sample data under normal and four fault modes are obtained during the complete gear shifting process from clutch engagement to disengagement.

3. The fault diagnosis method for agricultural machinery wet clutch shifting system based on closed continuous time unit according to claim 1, characterized in that, The method of denoising the collected test data using a fusion of adaptive noise from complete set empirical mode decomposition and improved wavelet thresholding includes: First, the original signal is processed using the fully ensemble empirical mode decomposition adaptive noise CEEMDAN. The decomposition yields the IMF components and residuals: in, , This is the noise amplitude coefficient. It is a Gaussian white noise sequence; Perform Empirical Mode Decomposition (EMD) on each noisy signal to obtain the th... The IMF component, the first The IMFs are: in: The set number refers to the total number of times EMD decomposition is performed after adding different noise levels. For the first After adding white noise for the first time, the result obtained by EMD is the first... One IMF component; Repeat the decomposition process until the remaining signal can no longer be decomposed into an IMF. The original signal can be represented as: in: For IMF quantity, These are the residual components; Calculate the original signal Cross-correlation coefficients with various IMFs ,filter IMF values ​​greater than a preset value retain valid features: in: For time delay parameters, For the delay range, For cross-correlation function, This indicates the absolute value operation; Finally, an improved wavelet threshold function is used to denoise the filtered IMF components to obtain a high-fidelity hydraulic signal. The improved wavelet threshold function is as follows: in, For the threshold, It is an exponential decay factor; It is a symbolic function.

4. A fault diagnosis device for a wet clutch shifting system in agricultural machinery based on a closed continuous time unit, characterized in that, The method based on claim 1 includes: The data acquisition module is used to build a test bench for wet clutch shifting system and obtain clutch hydraulic test data under normal and typical fault modes through fault simulation. The signal denoising module is used to denoise the acquired test data using a fusion denoising method of adaptive noise reduction and improved wavelet thresholding with complete set empirical mode decomposition. The fault diagnosis module is used to construct a diagnostic model FaDCfCNet based on closed continuous time units. The FaDCfCNet is trained and tested using denoised test data. The trained FaDCfCNet can accurately identify fault modes.

Citation Information

Patent Citations

  • Fault feature extraction and diagnosis method for power gear shifting system of unmanned tractor

    CN118328143A

  • Deep learning method for realizing mechanical fault diagnosis

    CN121524841A