Gearbox fault diagnosis method based on improved meta-learning network under small sample

CN117232826BActive Publication Date: 2026-09-25NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311343957.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-17
Publication Date
2026-09-25
Estimated Expiration
2043-10-17

AI Technical Summary

Technical Problem

生成器和判别器之间的对抗性优化使得训练过程容易陷入模式崩溃和震荡等

Benefits of technology

[0044]本发明考虑了工程实际中故障样本稀少的影响,首先建立了齿轮-转子系统的高保真度动力学模型,在该模型中引入了不同的啮合刚度参数,以生成不同齿轮失效类型的高质量模拟信号,为智能诊断模型提供了支持集作为泛化的基础。然后构建了一种新的基于元学习的故障诊断方法,该方法通过情景训练机制实现了对新任务的快速学习和参数微调,这种能力使模型能够更好地泛化到新任务,并仅用少量样本快速适应。最后采用LS正则化技术来减少模拟信号和真实信号之间数据特征分布的差异,使模型的输出更接近真实的概率分布,提高了模型的泛化能力。本发明所提出的方法能够显著提高小样本下齿轮故障诊断的精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117232826B_ABST
    Figure CN117232826B_ABST
Patent Text Reader

Abstract

The application discloses a gearbox fault diagnosis method based on an improved meta-learning network under small samples, comprising: establishing a high-fidelity gear-rotor dynamics model, generating simulation vibration signals of various fault gears by introducing different meshing stiffness; collecting gear fault measured vibration signals by using a gear fault simulation test bench; converting the simulation vibration signals and the measured vibration signals into corresponding energy graphs by using continuous wavelet transform respectively, and constructing a small sample data set; constructing a feature extraction model based on an improved meta-learning network, and training the feature extraction model by using support set data; inputting query set data into the trained feature extraction model for feature extraction, and calculating the prototype representation of each type of sample; calculating the distance between the query set data and the prototype representation, converting the distance into a probability distribution, and outputting the predicted fault category. The application improves the accuracy of the gearbox fault diagnosis result under the condition of insufficient fault data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fault diagnosis of rotating equipment such as gears, and more specifically, to a zero-sample gearbox fault diagnosis method based on an improved meta-learning network under small sample conditions. Background Technology

[0002] Gear transmission systems, due to their compact structure and high load-bearing capacity, are widely used in industries such as manufacturing, metallurgy, and aerospace. However, these large mechanical devices often operate in harsh environments with high temperatures and pressures, making their critical components prone to failure. As a whole, the performance of a gear transmission system is affected by the performance status of its individual components. If a critical component fails, it may affect the normal operation of other components, or even cause the entire system to fail, resulting in severe economic losses and, in severe cases, personal injury or death.

[0003] In recent years, deep learning models have been widely applied in fault diagnosis with good accuracy due to their powerful nonlinear processing and adaptive extraction capabilities. However, the high accuracy of these models relies on a large number of labeled samples as datasets. In real-world industrial scenarios, however, machinery typically operates under normal conditions, and in the event of a malfunction, the system is often shut down immediately for safety reasons. Furthermore, some rare gear fault types, such as tooth surface adhesion, are difficult to simulate through manual processing, resulting in a limited number of collected fault samples. As a typical data-driven model, when the dataset lacks sufficient samples, the model will overfit, significantly reducing diagnostic accuracy.

[0004] Currently, the mainstream methods for solving the small sample problem are mainly divided into the following two categories: (1) Data-level methods: generating more training samples by performing a series of transformations and expansions on existing small sample data. For example, expanding the sample set through image rotation, scaling, cropping, etc., or learning the data distribution and generating new samples through generative adversarial network (GAN) models. (2) Algorithm-level methods: improving recognition accuracy by improving the algorithm without changing the number of datasets. For example, transfer learning models are pre-trained on large-scale datasets, and then the model's weights are used as initial parameters for further training on small sample data, accelerating the model's learning process. For data-level methods, some operations may cause the loss or alteration of key information in the image. For example, cropping may crop out areas rich in fault information in the time-frequency image. When there is little sample data, there is a problem of training instability in the training process of GAN. The adversarial optimization between the generator and the discriminator makes the training process prone to mode collapse and oscillation. The success of transfer learning also benefits from the large amount of data collected in the laboratory as the training set. When the number of samples in the source domain is also small, its feature representation may not be rich or accurate enough, resulting in insufficient generalization ability of the model in the target domain. Meanwhile, the negative migration effect caused by the difference in the distribution characteristics of fault signals simulated in the laboratory and signals collected in actual engineering leads to a decrease in the diagnostic performance of the model.

[0005] To address the aforementioned problems, it is crucial to develop an effective intelligent fault diagnosis method for small sample or even zero sample conditions. Summary of the Invention

[0006] In view of this, the present invention provides a zero-sample gearbox fault diagnosis method based on an improved meta-learning network under small sample conditions, so as to improve the accuracy of gear fault diagnosis under small sample conditions.

[0007] Therefore, the present invention provides the following technical solution:

[0008] This invention provides a gearbox fault diagnosis method based on an improved meta-learning network under small sample conditions, comprising:

[0009] The gears are considered as concentrated mass points, and the meshing relationship between gears is simulated using spring-damping elements. A high-fidelity gear-rotor dynamic model is established, and simulation vibration signals of various faulty gears are generated by introducing different meshing stiffnesses.

[0010] The actual vibration signals of gear faults were collected using a gear fault simulation test bench.

[0011] The simulated vibration signal and the measured vibration signal are converted into corresponding energy maps using continuous wavelet transform, respectively, to construct a small sample dataset; the dataset includes: a support set for model feature extraction and a query set for evaluating model performance;

[0012] A feature extraction model based on an improved meta-learning network is constructed, and the feature extraction model is trained using support set data; query set data is input into the trained feature extraction model for feature extraction, and the prototype representation of each class of samples is calculated; the improved meta-learning network includes: optimizing the meta-learning network using a label smoothing strategy;

[0013] Calculate the distance between the query set data and the prototype representation, convert the distance into a probability distribution, and output the predicted fault category.

[0014] Furthermore, a high-fidelity gear-rotor dynamics model is established, including:

[0015] The kinematic differential equations of the spur gear pair established using the lumped parameter method are as follows:

[0016]

[0017] Among them, F 12 and X 12 These are the external load vector and the displacement vector, respectively; M 12 C is the mass matrix, containing the mass of the gear and the mass of the shaft; G is the damping matrix, using proportional damping; G is the gyroscope matrix, containing the gyroscope matrices of the gear and the shaft; K 12 This is the stiffness matrix, which includes the bearing stiffness and the meshing stiffness between the gear pairs.

[0018] Further, the simulated vibration signal and the measured vibration signal are converted into corresponding energy maps using continuous wavelet transform, including:

[0019] The amplitude of the vibration signal is scaled down to the range of [-1, 1] using a normalization operation;

[0020] The vibration signal after normalization is divided into multiple signal segments using the sliding window algorithm;

[0021] The signal segment was converted into an energy map using continuous wavelet transform;

[0022] A small sample dataset is constructed based on the energy map.

[0023] Furthermore, the signal segment is converted into an energy map using continuous wavelet transform, and the calculation formula is as follows:

[0024]

[0025]

[0026] Where α represents the scale parameter, ξ represents the translation parameter, and ψ α,ξ Let ψ represent the wavelet basis function. *(·) denotes the mother wavelet function.

[0027] Furthermore, feature extraction models based on meta-learning networks include:

[0028] Multiple convolutional modules, each including: convolutional layer, batch normalization layer, ReLU layer and max pooling layer.

[0029] Further, the test data is input into the trained feature extraction model for feature extraction, and the prototype representation of each class of samples is calculated, including:

[0030] The test data is mapped to a two-dimensional space using the feature extraction model, and the average embedding of each category sample is calculated as the category prototype. The calculation formula is as follows:

[0031]

[0032] Where f φ (x i (x) represents the features mined by the feature extraction module. i ,y i ) represents the sample feature vector and label, |S k | indicates the number of samples in the support set.

[0033] Furthermore, the meta-learning network is optimized using a label smoothing strategy, including:

[0034] A label smoothing strategy is used to mitigate the difference in data feature distribution between simulated and measured signals. The calculation formula is as follows:

[0035] Where α is the smoothing factor, K is the total number of categories, and y k This represents the original One-Hot encoded vector. It is the smoothed probability distribution;

[0036] The model parameters are fine-tuned using the cross-entropy loss function, and the formula is as follows:

[0037]

[0038] in, For real labels, p k K represents the probability predicted by the neural network, where K is the total number of categories.

[0039] Further, the distance is converted into a probability distribution to output the predicted fault category, including:

[0040] The softmax function is used to transform the predicted categories into a probability distribution, and the category with the highest probability is the fault category predicted by the model.

[0041]

[0042] Where, x i Let k be the input feature of the i-th sample, and k be the number of fault types. Let be the predicted probability of the j-th sample in the k-th class.

[0043] Advantages and positive effects of the present invention:

[0044] This invention considers the impact of scarce fault samples in practical engineering. First, a high-fidelity dynamic model of the gear-rotor system is established, in which different meshing stiffness parameters are introduced to generate high-quality simulated signals for different gear failure types, providing a support set as the basis for generalization of the intelligent diagnostic model. Then, a novel meta-learning-based fault diagnosis method is constructed. This method achieves rapid learning and parameter fine-tuning for new tasks through a scenario-based training mechanism. This capability enables the model to better generalize to new tasks and adapt quickly with only a small number of samples. Finally, LS regularization is used to reduce the difference in data feature distribution between simulated and real signals, making the model's output closer to the true probability distribution and improving the model's generalization ability. The method proposed in this invention can significantly improve the accuracy of gear fault diagnosis with small sample sizes. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a flowchart of a meta-learning network in an embodiment of the present invention;

[0047] Figure 2 This is a diagram of the dynamic model of the gear-rotor system in an embodiment of the present invention;

[0048] Figure 3 This is a flowchart comparing simulated signals and measured signals in an embodiment of the present invention;

[0049] Figure 4 This is a flowchart of the sliding window sampling method in an embodiment of the present invention;

[0050] Figure 5 This is a flowchart illustrating the use of wavelet transform to convert a vibration signal into a two-dimensional time-frequency graph in an embodiment of the present invention;

[0051] Figure 6 This is a flowchart of the network structure of the feature extraction module in an embodiment of the present invention;

[0052] Figure 7 This is a flowchart of the label smoothing method in an embodiment of the present invention;

[0053] Figure 8 This is a flowchart illustrating the distance between the feature vector and the prototype vector of a query point in an embodiment of the present invention.

[0054] Figure 9 This is a comparison chart of the diagnostic accuracy of the model in this embodiment of the invention with other models;

[0055] Figure 10 This is a comparison of confusion matrices for different models in the embodiments of the present invention. Detailed Implementation

[0056] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0057] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0058] like Figure 1 As shown in the figure, an embodiment of the present invention provides a zero-shot gearbox fault diagnosis method based on an improved meta-learning network under small-sample conditions. The method includes:

[0059] S1. Establish a high-fidelity gear-rotor dynamics model and generate simulation signals for various faulty gears by introducing different meshing stiffnesses.

[0060] The dynamic model diagram of the gear-rotor system is shown below. Figure 2 As shown. In specific implementation, S1 includes the following steps:

[0061] S11. Considering the gears as concentrated mass points, the meshing relationship between the gears is simulated using spring-damped elements. Taking into account the six degrees of freedom in two concentrated mass points, the established dynamic model is as follows:

[0062]

[0063] In the formula, m1 and m2 are the masses of the driving wheel and the driven wheel, respectively, and I x1 I y1 I z1 I x2 ,、I y2 and I z2 Let T1 and T2 represent the moments of inertia of its x, y, and z coordinate axes, respectively, and let k be the load torque. 12 (t) and c 12 (t) represent the stiffness and damping of the unit along the meshing line, respectively, Ω1 and Ω2 are the rotational angular velocities of the driving and driven gears, r b1 r b2 These are the base circle radii of the driving and driven gears, respectively, and θ y1 and θ y2 Let θ represent the rotational displacements of the driving and driven wheels about the y-axis, respectively. x1 and θ x2 Z1 and Z2 are the rotational displacements of the driving and driven gears about the x-axis, respectively, and the displacements of the driving and driven gears along the z-axis, respectively.

[0064] p 12 The mathematical expression for the relative displacement of the gear along the line of action is as follows:

[0065] p 12 (t)=(-x1sinψ 12 +x2sinψ 12 +y1cosψ 12 -y2cosψ 12 +σ×r b1 θ z1 +σ×r b2 θ z2 )-e 12 (t) (2)

[0066] x1 and x2 are the translational displacements of the driving wheel and driven wheel along the x-direction, respectively; y1 and y2 are the translational displacements of the driving wheel and driven wheel along the y-direction, respectively; θ z1 and θ z2 α represents the rotational displacement of the driving and driven wheels about the z-axis, respectively; α is the pressure angle at the pitch circle; r b1 and r b2 These are the base circle radii of the driving and driven gears, respectively, ψ 12 e is the angle between the gear's line of action and the positive y-axis. 12(t) represents the no-load transmission error of the gear pair.

[0067] S12. Substituting equation (2) into equation (1), the following differential equation of motion for the spur gear pair is obtained:

[0068]

[0069] Among them, F 12 and X 12 These are the external load vector and the displacement vector, respectively; M 12 C is the mass matrix, containing the mass of the gear and the mass of the shaft; G is the damping matrix, using proportional damping; G is the gyroscope matrix, containing the gyroscope matrices of the gear and the shaft; K 12 This is the stiffness matrix, which includes the bearing stiffness and the meshing stiffness between the gear pairs.

[0070] S2. Collect actual gear fault signals using a gear fault simulation test bench.

[0071] Figure 3 The image shows a comparison between the simulated and measured signals generated by the dynamic model after incorporating meshing stiffness sets for wear and spalling faults. It can be seen that the dynamic model can simulate the frequency domain images of the corresponding fault frequencies and their harmonic components under different fault conditions. The simulation data and experimental data are highly consistent, providing high-quality simulation signals for subsequent analysis.

[0072] S3. Use continuous wavelet transform to convert the one-dimensional vibration signal into an energy map and construct a dataset.

[0073] In its implementation, S3 includes the following steps:

[0074] S31. First, use the normalization method to scale the signal amplitude to the range [-1, 1]. The mathematical formula is as follows:

[0075]

[0076] Where x is the signal segment, x i Let x′ be the i-th data point of this signal segment. i These are the standardized data points.

[0077] S32. Use the sliding window algorithm to divide the samples into multiple signal segments;

[0078] The process of sliding window sampling is as follows: Figure 4As shown, the sliding window algorithm is used to uniformly divide the signal into signal segments containing 2048 data points. Specifically, using the sliding window algorithm, with a window size of 2048 and a slip amount of 30, sliding sampling is performed on the normalized signal. Assuming the number of data points in the original signal is M, the sample length is N, and the sampling slip amount is ΔK, the number of sample data that can be generated from this sequence is Y, as shown in the following formula:

[0079]

[0080] S33. Use wavelet transform to convert the divided signal segments into a two-dimensional time-frequency diagram;

[0081] The process of converting a vibration signal into a two-dimensional time-frequency graph using wavelet transform is as follows: Figure 5 As shown.

[0082] The specific formula is as follows:

[0083]

[0084]

[0085] Where α represents the scaling parameter, controlling the scaling of the mother wavelet function; ξ represents the translation parameter, controlling the translation of the mother wavelet function on the time axis; s(t) is the time signal; ψ α,ξ Let ψ represent the wavelet basis function. * (·) represents the mother wavelet function.

[0086] S34. Construct a small sample dataset;

[0087] Simulated signals are used as training data, and signals collected by the experimental platform are used as test data. The number of samples for each category does not exceed 30.

[0088] The details are shown in Table 1.

[0089] Table 1

[0090]

[0091] S4. Construct a feature extraction model based on an improved meta-learning network, and input the dataset into the meta-learning model optimized by label smoothing for feature extraction, and calculate the prototype representation of each class of samples.

[0092] like Figure 6 As shown, the meta-learning model optimized by label smoothing mainly consists of 6 convolutional modules, each of which contains 4 layers: a convolutional layer, a batch normalization layer, a ReLU layer, and a max pooling layer.

[0093]

[0094] in and Let x represent the weights and biases of the i-th convolutional kernel in the l-th layer. l (j) represents the j-th local region of the l-th layer, and * represents the dot product operation;

[0095]

[0096] Where, γ (i) and β (i) E[x] represents the learnable scaling parameter and offset parameter, respectively. (i) ] represents the mean, Var[x (i) ] represents the variance, and ε is a very small constant used to avoid division problems when the variance is zero;

[0097] f(x) = max(0,x) (10)

[0098] Where x represents the input value and f(x) represents the output value;

[0099]

[0100] in, Let represent the value of the t-th neuron in the i-th channel of the l-th layer, and W be the width of the pooling kernel. This represents the corresponding value of the neuron in the (l+1)th layer after the pooling operation.

[0101] The specific parameters of the model are shown in Table 2.

[0102] Table 2

[0103]

[0104] All samples in the dataset are mapped to the metric space through the model, and then the prototype representation C of each class is calculated by vector summation and averaging. k :

[0105]

[0106] Among them, f φ (x i (x) represents the features mined by the feature extraction module. i ,y i ) represents the sample feature vector and label, |S k | indicates the number of samples in the support set.

[0107] Figure 7 This is a label smoothing method that uses a weighted average of hard labels and a uniform distribution on the labels as soft labels. It reduces over-reliance on a single label and forces the model to consider the possibility of other categories, thus enabling it to generalize better to unseen data.

[0108] Specifically, a label smoothing strategy is used to mitigate the differences in data feature distribution between simulated and measured signals, thereby enhancing the model's generalization ability on unseen samples. The calculation formula is as follows:

[0109]

[0110] Where α is a smoothing factor (usually a very small value, such as 0.1 or 0.01), K is the total number of categories, and y k This represents the original One-Hot encoded vector. It is the smoothed probability distribution.

[0111] The model is fine-tuned using the label-smoothed cross-entropy loss function. The mathematical formula for the cross-entropy loss function is as follows:

[0112]

[0113] in, For actual labels, To predict the label value, n is the number of samples.

[0114] S5. Calculate the distance between the query set samples and the prototype representation, and apply softmax to convert these distances into a probability distribution to output the predicted fault category.

[0115] S51. For each sample in the query set, calculate the distance between it and the prototype of each category;

[0116] d(q j C i )=||f(q j )-C k ||2 (16)

[0117] Where ||·|| represents the L2 norm.

[0118] S52. Further, based on the distance calculated in S51, assign it to the category of the prototype representation that is closest to it;

[0119]

[0120] in, Is it a query sample q? j The prediction category.

[0121] S53. Apply the softmax function to transform the predicted category into a probability distribution, and the highest probability is the fault category predicted by the model.

[0122]

[0123] Where, x i Let k be the input feature of the i-th sample, and k be the number of fault types. Let be the predicted probability of the j-th sample in the k-th class.

[0124] In the meta-learning model, the dataset is defined as a support set and a query set. The support set is used for feature extraction by the model, providing generalization support. The query set is used to evaluate the performance of the meta-learning model. In this invention, the support set refers to simulation data, and the query set refers to the data actually collected on the experimental platform.

[0125] Figure 8 Euclidean distance is used to measure the distance between the feature vector and the prototype vector of the query point. Finally, Softmax is used to convert the distance into a probability distribution and output the predicted fault type.

[0126] Figure 9 Comparing the proposed model with other models, it can be seen that traditional deep learning models have low diagnostic accuracy under zero-shot conditions, with all three models achieving less than 60% accuracy. Among transfer learning models, the ResNet model combined with pre-training fine-tuning achieves a diagnostic accuracy of 85.33%, while the other two models with domain adaptation modules, DANN and DSAN, achieve accuracies of 68.55% and 62%, respectively. The improvement compared to traditional deep learning models is not significant, indicating that even with domain adaptation modules, models struggle to generalize with limited data. In contrast, the other four models based on few-shot learning all achieve diagnostic accuracy above 90% and exhibit good robustness. The proposed model, in particular, achieves 98.9% accuracy, demonstrating that few-shot learning still offers good diagnostic performance even with limited available samples.

[0127] Figure 10 This is the confusion matrix for different models of this invention. Figure 10 (a) shows the classification results of the WDCNN model, which can be seen that it exhibits varying degrees of confusion across all fault categories. Figure 10 (d) represents the confusion matrix of the Siamese model, which shows that it mainly misclassifies the tooth breakage and wear fault categories. Figure 10 (f) is the confusion matrix of the model proposed in this paper. It can be seen that the model proposed in this paper only misclassifies the wear fault category, while other fault samples can be correctly identified.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A gearbox fault diagnosis method based on an improved meta-learning network under small sample conditions, characterized in that, include: The gears are considered as concentrated mass points, and the meshing relationship between gears is simulated using spring-damping elements. A high-fidelity gear-rotor dynamic model is established, and simulation vibration signals of various faulty gears are generated by introducing different meshing stiffnesses. The actual vibration signals of gear faults were collected using a gear fault simulation test bench. The simulated vibration signal and the measured vibration signal are converted into corresponding energy maps using continuous wavelet transform, respectively, to construct a small sample dataset; the dataset includes: a support set for model feature extraction and a query set for evaluating model performance; A feature extraction model based on an improved meta-learning network is constructed and trained using support set data to calculate the prototype representation of each class of samples. Query set data is input into the trained feature extraction model to calculate the similarity between query points and the prototype representation of each class. The improved meta-learning network includes: optimizing the meta-learning network using a label smoothing strategy. The feature extraction model based on the meta-learning network includes: multiple convolutional modules, each convolutional module including: a convolutional layer, a batch normalization layer, a ReLU layer, and a max-pooling layer. Test data is input into the trained feature extraction model for feature extraction, and the prototype representation of each class of samples is calculated, including: mapping the test data to a two-dimensional space using the feature extraction model, and calculating the average embedding of each class of samples as the class prototype, the calculation formula of which is as follows: ; in This represents the features mined by the feature extraction module. Represents the sample feature vector and label. Indicates the number of samples in the support set; optimizes the meta-learning network using a label smoothing strategy, including: A label smoothing strategy is used to mitigate the difference in data feature distribution between simulated and measured signals. The calculation formula is as follows: ; in, It is a smoothing factor. The total number of categories, This represents the original One-Hot encoded vector. It is the smoothed probability distribution; The model parameters are fine-tuned using the cross-entropy loss function, and the formula is as follows: ; in, To predict probabilities for neural networks, Total number of categories; Calculate the distance between the query set data and the prototype representation, convert the distance into a probability distribution, and output the predicted fault category.

2. The gearbox fault diagnosis method based on an improved meta-learning network under small sample conditions according to claim 1, characterized in that, Establish a high-fidelity gear-rotor dynamics model, including: The kinematic differential equations of the spur gear pair established using the lumped parameter method are as follows: ; in, F 12 and X 12 These are the external load vector and the displacement vector, respectively. M 12 This is a mass matrix, containing the mass of the gears and the mass of the shafts; The damping matrix is ​​used, and proportional damping is employed. This is a gyroscope matrix, containing the gyroscope matrices for gears and shafts; K 12 This is the stiffness matrix, which includes the bearing stiffness and the meshing stiffness between the gear pairs.

3. The gearbox fault diagnosis method based on an improved meta-learning network according to claim 1, characterized in that, The simulated vibration signal and the measured vibration signal are converted into corresponding energy maps using continuous wavelet transform, including: The amplitude of the vibration signal is scaled down to the range of [-1, 1] using a normalization operation; The vibration signal after normalization is divided into multiple signal segments using the sliding window algorithm; The signal segment was converted into an energy map using continuous wavelet transform; A small sample dataset is constructed based on the energy map.

4. The gearbox fault diagnosis method based on an improved meta-learning network according to claim 3, characterized in that, The signal segment is converted into an energy map using continuous wavelet transform, and the calculation formula is as follows: ; ; in, Indicates the scale parameter. Indicates the translation parameter. Describe the wavelet basis functions. Represents the mother wavelet function; It is a time signal; Indicates time.

5. The gearbox fault diagnosis method based on an improved meta-learning network under small sample conditions according to claim 1, characterized in that, The distance is converted into a probability distribution to output the predicted fault category, including: The softmax function is used to transform the predicted categories into a probability distribution, and the category with the highest probability is the fault category predicted by the model. ; in, Indicates the first One query sample, The number of fault types, For the first The first sample Predicted probability of class samples.