Defect data expansion method for helicopter transmission system
Through the embedded smooth weighted CTGAN generation model and multi-dimensional data integration, the problem of scarcity and uneven distribution of helicopter transmission system defect data is solved, high-fidelity data expansion is achieved, structural fatigue strength design and life prediction are optimized, and flight safety and operation and maintenance efficiency are improved.
Patent Information
- Application Number
- CN202510430944.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-25
AI Technical Summary
The defect data of the helicopter transmission system is scarce and unevenly distributed. It is difficult for existing data expansion technologies to effectively model the nonlinear interaction between multi-dimensional data, resulting in insufficient generalization performance and reliability of defect prediction models, and it is difficult to support differentiated maintenance strategies.
The embedded smooth weighted CTGAN generation model is used for data expansion, combining multi-dimensional defect data integration and dynamic weight regulation, and high-fidelity expansion of defect data is carried out through one-hot encoding and size data normalization processing.
It realizes high-quality expansion of helicopter transmission system defect data, provides reliable data support, optimizes structural fatigue strength design and life prediction, and improves flight safety and operation and maintenance cost efficiency.
Smart Images

Figure CN120372384A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of helicopter structure fatigue data mining and analysis, and particularly relates to a method for augmenting defect data for a helicopter transmission system. Background Art
[0002] As a core power transmission component, the helicopter transmission system is in complex service environments such as deserts, gobi, and coastal areas for a long time. The formation and evolution of its component defects are significantly affected by multi-dimensional non-geometric factors such as service history, regional environment, and maintenance cycle. For example, in desert areas, the high-speed impact of sand dust particles easily leads to a significant increase in the proportion of surface scratch defects of transmission components, while the high-salt fog environment in coastal areas accelerates the propagation of corrosion fatigue cracks in metal materials. In addition, dimensional information such as the material properties of components, installation positions (such as the main rotor drive shaft and the tail drive shaft), and maintenance records (inspection intervals) are strongly correlated with the defect types (impact, scratch, corrosion, etc.) and their geometric features (depth, radius, position).
[0003] The acquisition of helicopter transmission system defect data is limited by high detection costs, strict safety standards, and long-cycle service characteristics, resulting in scarce and unevenly distributed samples. Although existing data augmentation techniques (such as Bootstrap resampling, Gaussian mixture model, Monte Carlo simulation, synthetic minority over-sampling technique SMOTE, and adaptive synthetic sampling ADASYN) have made some progress in general scenarios, traditional methods are difficult to effectively model the non-linear interaction relationships among multi-dimensional data such as defect size, component type, service time, material properties, and environmental conditions; existing methods are prone to generating outliers beyond the physical meaning range; the dynamic changes in service environment and maintenance strategies require the generation model to have the ability to adaptively adjust weights, while traditional static weight allocation mechanisms are difficult to achieve precise matching between the generated distribution and the target requirements. The above problems seriously restrict the generalization performance and reliability of the defect prediction model, making the formulation of differential maintenance strategies lack data support.
[0004] Therefore, there is an urgent need for a data augmentation method that can integrate multi-dimensional features and dynamically regulate the generated distribution to reveal the potential correlation laws between defect patterns and environmental factors. It is of great engineering significance to optimize the structural strength design of the transmission system coupling and the fatigue life prediction model, and at the same time, it can provide key technical support for reducing operation and maintenance costs and improving flight safety. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for augmenting defect data for a helicopter transmission system. The core lies in: establishing defect data of a helicopter transmission system with defect tolerance design, and using a CTGAN generation model with embedded smooth weighting for data augmentation to solve the problem of insufficient defect samples of the helicopter transmission system.
[0006] To achieve the above object, the present invention provides a method for expanding defect data of a helicopter transmission system, comprising the following steps:
[0007] S1: Establish a defect data set of a helicopter transmission system with defect tolerance design;
[0008] S2: Calculate the frequency distribution of each category of data, fit the frequency distribution of size data, and calculate and determine the corresponding category weights;
[0009] S3: Classify data using one-hot encoding and normalize the size data;
[0010] S4: Perform data expansion based on the CTGAN generation model with embedded smooth weighting;
[0011] S5: Conduct a spearman correlation data test;
[0012] S6: Denormalize the numerical columns and restore the encoded categorical values to the original labels
[0013] Preferably, in step S1, the defect data includes six dimensions: defect size, component parts, time, material, environment, and information source.
[0014] Preferably, the defect size dimension specifically includes depth, radius, and feature distance data; the component parts dimension specifically includes component, component part-related information, and position description; the time dimension includes inspection date, flight time, and last inspection date; the material dimension mainly refers to alloy grade; the environment dimension specifically includes service region and service environment; the information source specifically includes statisticians
[0015] Preferably, in step S2, the frequency of category k is:
[0016]
[0017] where N k is the number of samples of category k, and N is the total number of samples;
[0018] The weight distribution of the category is:
[0019]
[0020] ∈ is the smoothing factor;
[0021] Then perform normalized weights:
[0022]
[0023] Preferably, in step S3, assuming that the categorical data variable C has K mutually exclusive categories, the one-hot encoding mapping is:
[0024]
[0025] Normalize the dimension data:
[0026]
[0027] where x scale is the normalized data, x is the data before normalization, x max is the maximum value in the data, x min is the minimum value of the data.
[0028] Preferably, in step S4, a smoothing weighting condition is embedded in the GAN generation model. First, a class weight vector w is appended to the noise vector of the generator G, or the class weight w in the discriminator D is used as an additional feature and concatenated with the data:
[0029] G(z, w), D(x, w)
[0030] where is the normalized weight for each class;
[0031] Set the objective function of the generator G with embedded weights:
[0032]
[0033] where, is the expectation operator, sampling according to the probability distribution p z for all possible noise inputs z; y.G(z) / the predicted class label of the generated sample G(z), D.G(z) / the discrimination probability of the discriminator for the generated sample G(z), and λ is a hyperparameter controlling the weight of the KL divergence term, balancing the quality of the generated sample and the strength of the distribution matching;
[0034]
[0035] where, is the class distribution of the generated sample, is the target class distribution, and the KL divergence forces the generated distribution p gen to approximate the target distribution p target ;
[0036] Set the objective function of the discriminator D with embedded weights:
[0037]
[0038] where y(x) is the class label of the sample x, is the weight of the class of the sample x, p data is the real sample distribution, and D(x) is the prediction probability of the discriminator for the real sample x.
[0039] Preferably, in step S4, when updating the weights in each round, the dynamic weights will be automatically updated:
[0040]
[0041] where α is the smoothing coefficient.
[0042] Preferably, in step S5, for the expanded data, the reliability of the data is further tested by using the spearman correlation coefficient for the relationship between the data.
[0043]
[0044] where d i represents the difference in the rank values of the i-th data pair, and n represents the total number of observed samples. and the average rank.
[0045] Compared with the prior art, the advantages of the present invention are as follows: By integrating multi-dimensional defect data, optimizing the design of the generation model, and dynamically controlling the data distribution, the present invention solves the problem of insufficient defect samples in the helicopter transmission system, and provides high-reliability data support for the fatigue strength design and life prediction of the defect tolerance structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is the overall flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0047] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The specific implementation steps are as follows:
[0048] Step S1: Sort out the existing data of design institutes, production and manufacturing units, maintenance units, and user units, etc., and collect the defects of the active helicopter transmission system in the field to establish a defect data set of the helicopter transmission system for defect tolerance design. The data covers six dimensions of information. One is the defect size dimension, specifically including data such as depth, radius, and feature distance; the second is the component and part dimension, including component and part-related information and position description; the third is the time dimension, involving inspection date, flight time, and last inspection date; the fourth is the material dimension, mainly referring to the alloy grade; the fifth is the environment dimension, including service area and service environment; the sixth is the information source dimension, that is, the statistician. Some data are shown in Table 1.
[0049] Table 1
[0050]
[0051]
[0052] Step S2, Category Frequency Calculation: Divide the data according to the defect type (impact, scratch, corrosion, etc.) or service environment category, and calculate the frequency of each category k:
[0053]
[0054] where N k is the number of samples in category k, and N is the total number of samples;
[0055] Introduce a smoothing factor ∈ = 1e -5 , and calculate the initial weights:
[0056]
[0057] After normalization, the target weights are obtained:
[0058]
[0059] Step S3, perform one-hot encoding on discrete variables (such as defect type, service environment); assume that the categorical variable C has K mutually exclusive categories, then the one-hot encoding mapping is:
[0060]
[0061] Perform min-max normalization on continuous variables (such as defect depth, radius):
[0062]
[0063] Step S4, adopt a conditional generative adversarial network (CTGAN), where both the generator G and the discriminator D are based on a multi-layer perceptron (MLP), the hidden layer dimension is set to 256, and the activation function is LeakyReLU; use the normalized weight vector as the conditional input; append the category weight vector w to the noise vector of the generator G, or use the category weight w in the discriminator D as an additional feature and concatenate it with the data:
[0064] G(z, w), D(x, w)
[0065] Set the objective function of the generator G with embedded weights:
[0066]
[0067] where,[[]] is the expected value operator, according to the probability distribution p zSampling is performed for all possible noisy inputs z; the predicted class label of the generated sample G(z) is y.G(z) / , the discrimination probability of the discriminator D.G(z) / for the generated sample G(z), and λ = 0.3 is the weight for controlling the KL divergence term, which is used to constrain the generated distribution p gen Approximating the target distribution
[0068]
[0069] wherein, is the class distribution of the generated sample, is the target class distribution;
[0070] The setting of the objective function of the discriminator D with embedded weights is as follows:
[0071]
[0072] where y(x) is the class label of the sample x, is the weight of the class of the sample x, p data is the true sample distribution, and D(x) is the prediction probability of the discriminator for the true sample x; then, after each round of training, according to the generated data distribution the weights are updated, and the α smoothing coefficient is adjusted to balance the influence of the historical weights and the current generated distribution.
[0073] In step S5, calculate the Spearman correlation coefficient between the generated data and the original data in each dimension:
[0074]
[0075] If ρ > 0.7, it is considered that the correlation of the generated data meets the requirements.
[0076] In step S6, restore the generated data to the original range:
[0077] x = x scale ·(x max - x min ) + x min
[0078] According to the one-hot encoding mapping table, restore the encoding vector of the generated data to the original class label (such as "desert", "ocean").
[0079] The core of the present invention lies in achieving high-fidelity expansion of the defect data of the helicopter transmission system through an embedded smoothed weighted CTGAN generation model, combined with multi-dimensional data preprocessing and dynamic weight regulation. As Figure 1As shown in the figure, it is the overall flowchart of the method of the present invention, covering the entire processes of data collection, preprocessing, model training, generation and verification. Through the above-mentioned implementation manners, the present invention can efficiently expand high-quality data conforming to the defect distribution characteristics of the helicopter transmission system, providing reliable support for structural fatigue strength design and life prediction.
[0080] The above only expresses the technical solutions and implementation manners of the present invention. The description is relatively specific and detailed, but it cannot be considered that the above design limits the scope of the patent of the present invention. For those skilled in the art, without departing from the design concept of the present invention, it should be able to realize that the equivalent replacements and obvious changes made by using the content of the specification and drawings of the present invention can also make many changes and improvements. These acts are all within the scope protected by the present invention and should all be included in the protection scope of the present invention.
Claims
1. A method for defect data augmentation of a helicopter transmission system, characterized in that It includes the following steps: S1: Establish a defect data set for the helicopter transmission system with defect tolerance design; S2: Calculate the frequency distribution of each category of data, fit the frequency distribution of size data, and calculate and determine the corresponding category weights; S3: Use one-hot encoding to classify data and normalize the size data; S4: Perform data augmentation based on the CTGAN generation model with embedded smooth weighting; S5: Conduct a spearman correlation data test; S6: Reverse normalize the numerical columns and restore the encoded categorical values to the original labels.
2. The method for augmenting defect data for a helicopter transmission system according to claim 1, wherein In step S1, the defect data includes six dimensions: defect size, component parts, time, material, environment, and information source.
3. According to the method for defect data augmentation for a helicopter transmission system described in claim 1, the defect size dimension specifically includes depth, radius, and feature distance data; the component parts dimension specifically includes components, component-related information, and position descriptions; the time dimension includes inspection date, flight time, and last inspection date; the material dimension mainly refers to the alloy grade; the environment dimension specifically includes service region and service environment; the information source specifically includes statisticians.
4. A method for defect data augmentation of a helicopter transmission system according to claim 1, characterized in that In step S2, the frequency of its category k is: where N k is the number of samples in class k, and N is the total number of samples; The weight assignment for the category is: ∈ is the smoothing factor; Then perform normalized weights:
5. A method for expanding defect data for a helicopter transmission system according to claim 1, characterized in that In step S3, assuming that the categorical data variable C has K mutually exclusive categories, the one-hot encoding mapping is: Normalize the size data: where x scale is the normalized data, x is the data before normalization, and x max is the maximum value in the data, and x min is the minimum value of the data.
6. The method for expanding defect data for a helicopter transmission system according to claim 1, wherein In step S4, embed the smooth weighting condition in the GAN generation model. First, append the category weight vector w to the noise vector of the generator G, or use the category weight w in the discriminator D as an additional feature and concatenate it with the data: G(z, w), D(x, w) wherein is the normalized weight for each category; Set the objective function of the generator G with embedded weights: Among them, is the expected value operator, sampling according to the probability distribution p z For all possible noise inputs z; the predicted class label of the generated sample G(z) / is generated by sampling y.G(z) / , D.G(z) / is the discrimination probability of the discriminator for the generated sample G(z), and λ is a hyperparameter that controls the weight of the KL divergence term, balancing the quality of the generated sample and the strength of the distribution matching; Among them, is the class distribution of the generated samples, is the target class distribution, and the KL divergence forces the generated distribution p gen to approximate the target distribution p target ; Set the objective function of the discriminator D with embedded weights: where y(x) is the class label of sample x, is the weight of the class of sample x, p data is the true sample distribution, and D(x) is the predicted probability of discriminator for true sample x.
7. A method for expanding defect data of a helicopter transmission system according to claim 1, characterized in that In step S4, when updating the weights in each round, the dynamic weights will be automatically updated: Where α is the smoothing coefficient.
8. The method for expanding defect data for a helicopter transmission system according to claim 1, characterized in that In step S5, for the augmented data, further use the spearman correlation coefficient to test the reliability of the data for the relationship between the data. where d i represents the difference in the rank values of the i-th data pair, n represents the total number of observed samples, and the average rank.