An industrial machine coupling fault diagnosis method based on adaptive feature clustering under low-labeled samples
By adaptive normalization of heterogeneous signal kernel density and mining of three-dimensional weighted primary and secondary features, combined with iterative fusion of channel attention weights, the problems of high dependence on labeled samples and poor signal feature fusion effect in coupled fault diagnosis of industrial equipment are solved, and high-precision fault identification and source localization under complex working conditions are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN INST OF TECH
- Filing Date
- 2026-03-19
- Publication Date
- 2026-06-19
Smart Images

Figure CN122241270A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent fault diagnosis technology for industrial equipment, and in particular to a method for diagnosing coupled faults in industrial machinery based on adaptive feature clustering under low-labeled samples. Background Technology
[0002] Industrial equipment is the core equipment for industrial production and engineering operations. The stability of its operating status directly determines operational efficiency, production safety, and maintenance costs. Its core electromechanical and hydraulic systems consist of multiple deeply coupled subsystems. During equipment operation, a failure in a single component can easily trigger a chain reaction, forming coupled faults caused by multiple factors. These faults are characterized by weak signal characteristics, numerous interfering factors, and complex evolution patterns. Furthermore, in actual industrial scenarios, while monitoring data for normal equipment operation is abundant, fault data is scarce, labeled samples are limited, and sample distribution is prone to imbalance, posing a significant challenge to fault diagnosis.
[0003] Existing industrial equipment fault diagnosis methods struggle to effectively address the core pain point of coupled fault diagnosis, with each method exhibiting significant limitations: Traditional data-driven deep learning methods, while possessing strong feature extraction capabilities, are highly dependent on labeled samples, resulting in a sharp decline in diagnostic accuracy under conditions of insufficient labeled data, and they fail to fully consider the magnitude differences in multi-sensor signals and cross-channel coupling characteristics; Classical unsupervised clustering methods rely solely on the distance similarity between data points, depending entirely on the quality of the original signal features without label constraints, making them highly sensitive to changes in data distribution and unable to effectively uncover the intrinsic correlation features of coupled faults; Semi-supervised learning methods, while combining a small amount of labeled data with a large amount of unlabeled data, struggle to maintain consistency between task features and target data during cross-channel feature fusion, resulting in insufficient ability to locate and trace coupled faults.
[0004] Furthermore, monitoring signals collected by multiple sensors in industrial settings generally exhibit significant differences in magnitude and inconsistent data distribution. Traditional signal normalization methods (such as maximum-minimum normalization and standardization) are prone to masking or ignoring subtle fault characteristics, further reducing the accuracy of coupled fault identification. Simultaneously, existing methods show poor adaptability and robustness under real-world industrial conditions such as sample imbalance, limited sensor channel numbers, and strong noise interference. Moreover, they are mostly black-box models, failing to provide interpretable analysis of coupled faults and hindering the provision of targeted guidance for precise operation and maintenance of on-site equipment.
[0005] To address the aforementioned technical challenges, there is an urgent need to develop an intelligent diagnostic method that can adaptively process heterogeneous signals from multiple sensors, efficiently utilize limited labeled data, and deeply mine the intrinsic characteristics of coupled faults. Summary of the Invention
[0006] This invention aims to provide an industrial machinery coupling fault diagnosis method based on adaptive feature clustering under low-labeled sample conditions. Through a progressive technical chain of heterogeneous signal kernel density adaptive normalization, three-dimensional weighted primary and secondary feature mining, channel attention weight iterative fusion, and fault feature quantification and source tracing analysis, it solves the technical problems of existing fault diagnosis methods such as high dependence on labeled data, poor multi-source signal feature fusion effect, insufficient deep feature mining of coupled faults, and weak adaptability to complex working conditions. It achieves accurate identification, source location, and evolution analysis of coupled faults in industrial equipment.
[0007] A method for diagnosing coupled faults in industrial machinery based on adaptive feature clustering under low-labeled samples, comprising the following steps: S1. Collect monitoring signals from multiple sensors of industrial machinery, and perform heterogeneous signal kernel density adaptive normalization processing on the monitoring signals of multiple sensors of industrial machinery to eliminate the magnitude difference and noise interference of the multi-sensor signals, while amplifying the identification of weak coupling fault characteristics. S2. Perform three-dimensional weighted primary and secondary feature adaptive mining on the normalized industrial machinery multi-sensor monitoring signals to generate pre-trained weight matrices with physical constraints for each sensor channel. S3. Based on the pre-trained weight matrix of each channel, the channel attention weight iterative fusion is completed on the normalized industrial machinery multi-sensor monitoring signal. Combined with a small number of labeled samples, the deep collaborative fusion of multi-channel coupled fault features is realized, a coupled fault diagnosis clustering model is constructed, and the classification and identification of coupled faults are completed. S4. Optimize the model parameters and verify the performance of the coupled fault diagnosis clustering model constructed in steps S1-S3. Optimize the core hyperparameters and verify the stability, generalization ability and robustness of the model under different industrial conditions. S5. Conduct quantitative source analysis of coupled fault characteristics on the coupled fault diagnosis results of industrial machinery output by the coupled fault diagnosis clustering model, quantify the contribution of each sensor characteristic to the coupled fault, locate the source of the coupled fault, and reveal the evolution and propagation law of the fault.
[0008] Furthermore, S1 includes the following steps: S11. Collect S-dimensional sensor monitoring signals from complex industrial machinery and construct a signal set. ,in For the first A sequence of continuously sampled data from a sensor. The sampling length of a single-channel signal is set according to the monitoring needs of industrial machinery; S12, Calculation of the first based on improved kernel density estimation Each data point in the sensor signal The distribution density is introduced by incorporating data distribution weight coefficients calculated from the data outlier degree. The formula is as follows:
[0009] in To improve the bandwidth parameter of kernel density estimation, the optimal value is determined by grid search method; These are the data distribution weighting coefficients. For the first Each sensor channel signal The mean, For the first Each sensor channel signal Standard deviation; For the first Each sensor channel signal The The value at each position, , Let be the Cauchy kernel density function, expressed as: , For general variables; S13. Calculate the area under the kernel density curve using the 10th-order Gauss-Legendre adaptive quadrature method to determine if the constraints are satisfied. Effective upper and lower boundaries of the signal ,in The effective data percentage constraint threshold is calculated as follows: ; S14. Introduce a non-linear scaling factor. The signals from each sensor are nonlinearly scaled to a unified interval of [0,1] to achieve adaptive normalization of heterogeneous signals. The transformation formula is as follows: , in For the first Each sensor channel signal The first after adaptive normalization The value at each position, .
[0010] Furthermore, S2 includes the following steps: S21. Determine the weight matrix scale through cross-validation based on the actual monitoring needs of the equipment. Initialize the single-channel weight matrix The initial value is generated by the signal mean bias; S22, Arbitrary sensor channel signals of the normalized multi-sensor monitoring signals for industrial machinery. The normalized channel signal Set as core analysis parameter the remaining The normalized signals from each channel are subjected to principal component analysis for dimensionality reduction, generating a one-dimensional auxiliary parameter set. , build ,in Weighting of auxiliary parameters based on correlation coefficients enhances the characteristics of parameters with high correlation. For signal standard deviation One-dimensional auxiliary parameter set for the signal standard deviation The sensor channel signal after S1 adaptive normalization is divided by the normalized signal. Other A multi-channel sensor signal matrix for and covariance; S23. Calculate the Mahalanobis-weighted Euclidean distance between the auxiliary weighting parameters and the weight matrix to locate the best matching unit (BMU), as shown in the following formula:
[0011] in The spatial coordinates of the best matching unit in the three-dimensional weight matrix. The corresponding position in the weight matrix The element value at that position, For the first Each sensor channel signal The first after adaptive normalization The value at each position, , ; S24. Using the best matching unit as the center, set a dynamic neighborhood radius that decreases linearly with the number of iterations. Determine the neighborhood coordinate set:
[0012] in , The initial neighborhood radius is preset. This represents the current iteration number. This is the preset maximum number of iterations.
[0013] Furthermore, following S24, the process also includes feature similarity quantification and weight matrix iterative update of the normalized industrial machinery multi-sensor monitoring signals: S25. Calculate the cosine distance of the core parameters respectively. Chebyshev weighted distance with auxiliary parameters The comprehensive feature similarity between data is obtained by fusing data using an improved Softmax function. The formula is as follows:
[0014]
[0015]
[0016] in In order to be with the first Normalized signal of each sensor channel The same-dimensional neighborhood mask signal sequence, centered on the currently processed sampling point, retains the same signal within a preset neighborhood range. Consistent signal values, values outside the neighborhood range are taken as the signal value of the current sampling point, used to match... Extracting local signal difference features by subtracting element by element; For the first Sensor channel neighborhood mask signal sequence The extracted local neighborhood feature set consists of statistical features, dimensionality reduction features, or model learning features of the signal within the neighborhood, and is used to characterize the local operating state features within the neighborhood of the current sampling point. For the first Sensor channel neighborhood mask signal sequence The number of valid sampling points included is used to characterize the size of the neighborhood range, and its value is adaptively determined according to the dynamic response characteristics of the industrial equipment. S26. Combining the Manhattan distance of the neighborhood spatial coordinates with the similarity of the comprehensive data features, calculate the gradient descent optimization amount of the weight matrix unit. The weight matrix is iteratively updated using stochastic gradient descent, as shown in the following formula:
[0017]
[0018] in, The learning rate; The L2 regularization coefficient; For the number of iterations, For a one-dimensional auxiliary parameter set Location The corresponding element value, For the first Each sensor channel signal The first after adaptive normalization The value of each position; S27. All normalized industrial machinery multi-sensor monitoring signals The unsupervised mining operation is performed sequentially on each sensor channel to adaptively generate a pre-trained weight matrix for each channel. This enables in-depth mining and extraction of single-channel coupling fault characteristics.
[0019] Furthermore, S3 includes the following steps: S31. Merge the pre-trained weight matrices from all sensor channels to construct a global weight matrix. The optimal matching unit between the normalized multi-sensor monitoring signals of industrial machinery and the global weight matrix is calculated to generate the initial data membership matrix. ; S32. Initialize cluster categories using a small number of labeled samples, and construct a clustering system containing... Category matrix of data points Cluster centers for each category are determined through max pooling convolution operations. ,in To indicate the number of labeled samples; S33. Based on the normalized multi-sensor monitoring signals of industrial machinery, update the global weight matrix with the current cluster center, recalculate the data membership matrix and category matrix, and set the convergence threshold. If the change in cluster centers between two consecutive iterations is less than If the iteration terminates, the optimized global fusion weight matrix and coupled fault diagnosis clustering model are obtained. S34. For each full-dimensional sensor signal sequence of the normalized industrial machinery multi-sensor monitoring signal, calculate its relationship with each category matrix. distance The signal sequence is assigned to the category with the smallest distance to generate industrial mechanical coupling fault diagnosis results. ,in For the first Each sensor channel signal Corresponding fault diagnosis labels; Furthermore, in S3, the convergence threshold The value is When initializing cluster categories using a small number of labeled samples, the labeled samples are selected from the normal operation status of industrial machinery, the fault status of hydraulic system, and the fault status of power system, with a selection quantity of 10 / 30 / 50 / 100.
[0020] Furthermore, S4 includes the following steps: S41. Optimize the weight matrix scale based on the actual operating conditions of industrial machinery. Neighborhood radius calculate, This represents the actual number of test samples. Take the positive integer value of the calculation result; S42. The measured data of the multi-sensor monitoring signals of industrial machinery are divided into training set and test set in a ratio of 7:3. The 10-fold cross-validation method is used to verify the stability and generalization ability of the coupled fault diagnosis clustering model, and to ensure the effectiveness of unsupervised adaptive feature mining and multi-channel information fusion. S43. Test the diagnostic accuracy and F1 score of the coupled fault diagnosis clustering model under different labeled sample numbers, different sensor channel numbers, and different sample distribution conditions to verify the robustness of the model under actual industrial conditions such as sample imbalance, limited labeling, and limited channels. The diagnostic accuracy is the ratio of the number of correctly diagnosed samples to the total number of test samples, and the F1 score is the harmonic mean of precision and recall. Furthermore, S5 includes the following steps: S51. Visualize the membership data distribution of the feature clusters after multi-channel information fusion of the multi-sensor monitoring signals of industrial machinery, generate a heat map, and intuitively distinguish the normal operating status of the equipment from the coupled fault feature clusters of different types of industrial machinery. S52. Statistically analyze key features such as the average value of multi-sensor signals and the proportion of fault-related components for each feature cluster, quantitatively analyze the signal feature patterns of different types of industrial machinery coupling faults, and combine the results of unsupervised adaptive feature mining to locate the correlation between features and industrial machinery coupling faults. S53. By combining the actual physical topology and component connection relationships of complex industrial machinery, the characteristic clusters of coupled faults of each industrial machinery are mapped to the actual components and subsystems of the industrial machinery. The source components of coupled faults of industrial machinery are located, the cross-subsystem evolution and propagation laws of faults are revealed, and targeted hierarchical precision operation and maintenance and fault handling suggestions are formed.
[0021] A storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for diagnosing industrial mechanical coupling faults based on adaptive feature clustering for low-labeled samples.
[0022] A computer device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method for diagnosing industrial mechanical coupling faults based on adaptive feature clustering under low-labeled samples.
[0023] Compared with the prior art, the present invention achieves significant beneficial effects through the above technical solution: (1) Low dependence on labeled data, suitable for actual industrial scenarios: Through the dual-layer framework design of unsupervised adaptive feature mining and multi-channel information fusion, high-precision diagnosis of coupled faults can be achieved with only a small number of labeled samples. Under the condition of scarce labeled data, the diagnostic performance is significantly better than traditional deep learning and semi-supervised learning methods, effectively solving the problem of high difficulty and high cost of labeling fault samples in actual industry.
[0024] (2) Strong ability to mine and fuse coupled fault features and high diagnostic accuracy: First, unsupervised adaptive feature mining is used to achieve deep extraction of single-channel coupled fault features. Then, multi-channel information fusion is used to complete the collaborative learning of cross-channel features, which fully captures the multi-dimensional correlation features of coupled faults. In the fault diagnosis of electromechanical hydraulic equipment, industrial bearings and other equipment, the diagnostic accuracy and F1 score are significantly better than mainstream diagnostic methods such as multi-channel integrated convolutional network, multi-scale cascaded deep belief network and two-stage semi-self-supervised method.
[0025] (3) Strong adaptability and robustness to working conditions, adaptable to complex industrial environments: The unsupervised adaptive feature mining process does not require manual intervention and can adapt to the signal features of different sensors. The multi-channel information fusion process can effectively integrate limited channel data. Therefore, under various actual industrial working conditions such as sample imbalance, limited number of sensor channels, and strong noise interference, the model can still maintain stable diagnostic performance. Compared with existing methods, the adaptability to complex and ever-changing industrial environments is greatly improved.
[0026] (4) Excellent scalability and generalization ability, and wide applicability: No large-scale reconstruction of the model is required. Only the number of sensor channels and the scale of the weight matrix need to be adjusted according to the equipment type. This method can flexibly adapt to the unsupervised adaptive feature mining and multi-channel information fusion requirements of different equipment. This method can be extended to engineering machinery and the coupled fault diagnosis of general industrial equipment such as industrial bearings, and has good engineering application prospects.
[0027] When this invention is applied to the coupled fault diagnosis of electromechanical hydraulic systems, it utilizes multi-dimensional sensors to monitor signals. Under actual working conditions where labeled samples are scarce and sample distribution is unbalanced, the model's diagnostic accuracy is significantly better than other comparative methods after unsupervised adaptive feature mining and multi-channel information fusion. Attached Figure Description
[0028] Figure 1 The overall architecture diagram of the multi-channel adaptive feature clustering fault diagnosis method; Figure 2 A radar chart comparing the F1 and ACC performance of various diagnostic models based on a publicly available dataset; Figure 3 This is a heatmap showing the membership distribution of fault feature clusters. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Reference Figures 1-3 As shown, an industrial mechanical coupling fault diagnosis method based on adaptive feature clustering for low-labeled samples includes the following steps: S1. Collect monitoring signals from multiple sensors of industrial machinery, and perform heterogeneous signal kernel density adaptive normalization processing on the monitoring signals of multiple sensors of industrial machinery to eliminate the magnitude difference and noise interference of the multi-sensor signals, while amplifying the identification of weak coupling fault characteristics. S2. Perform three-dimensional weighted primary and secondary feature adaptive mining on the normalized industrial machinery multi-sensor monitoring signals to generate pre-trained weight matrices with physical constraints for each sensor channel. S3. Based on the pre-trained weight matrix of each channel, the channel attention weight iterative fusion is completed on the normalized industrial machinery multi-sensor monitoring signal. Combined with a small number of labeled samples, the deep collaborative fusion of multi-channel coupled fault features is realized, a coupled fault diagnosis clustering model is constructed, and the classification and identification of coupled faults are completed. S4. Optimize the model parameters and verify the performance of the coupled fault diagnosis clustering model constructed in steps S1-S3. Optimize the core hyperparameters and verify the stability, generalization ability and robustness of the model under different industrial conditions. S5. Conduct quantitative source analysis of coupled fault characteristics on the coupled fault diagnosis results of industrial machinery output by the coupled fault diagnosis clustering model, quantify the contribution of each sensor characteristic to the coupled fault, locate the source of the coupled fault, and reveal the evolution and propagation law of the fault.
[0031] Specifically, this invention proposes a fault diagnosis method based on label-assisted self-supervised clustering. By normalizing the constraint boundaries of heterogeneous signals, mining the master parameter clustering of single-channel coupling features, and fusion of cross-channel features with label-assisted self-supervised clustering, it can achieve accurate diagnosis and interpretable source tracing of coupled faults in complex industrial machinery and equipment, effectively making up for the shortcomings of existing technologies.
[0032] Furthermore, S1 includes the following steps: S11. Collect S-dimensional sensor monitoring signals from complex industrial machinery and construct a signal set. ,in For the first A sequence of continuously sampled data from a sensor. The sampling length of a single-channel signal is set according to the monitoring needs of industrial machinery; S12, Calculation of the first based on improved kernel density estimation Each data point in the sensor signal The distribution density is introduced by incorporating data distribution weight coefficients calculated from the data outlier degree. The formula is as follows:
[0033] in To improve the bandwidth parameter of kernel density estimation, the optimal value is determined by grid search method; These are the data distribution weighting coefficients. For the first Each sensor channel signal The mean, For the first Each sensor channel signal Standard deviation; For the first Each sensor channel signal The The value at each position, , Let be the Cauchy kernel density function, expressed as: , For general variables; S13. Calculate the area under the kernel density curve using the 10th-order Gauss-Legendre adaptive quadrature method to determine if the constraints are satisfied. Effective upper and lower boundaries of the signal ,in The effective data percentage constraint threshold is calculated as follows: ; S14. Introduce a non-linear scaling factor. The signals from each sensor are nonlinearly scaled to a unified interval of [0,1] to achieve adaptive normalization of heterogeneous signals. The transformation formula is as follows: , in For the first Each sensor channel signal The first after adaptive normalization The value at each position, .
[0034] Specifically, this invention calculates the sensor signal distribution density by improving kernel density estimation and combining it with weighting coefficients obtained from data outlier calculation. It also uses a Cauchy kernel density function that is more suitable for the heavy-tailed distribution characteristics of industrial monitoring data, making the calculated signal distribution density more consistent with actual industrial conditions. Furthermore, it uses a 10th-order Gauss-Legendre adaptive quadrature method to accurately determine the effective upper and lower boundaries of the signal that meet the effective data ratio constraint, ensuring that the effective interval can include the vast majority of normal data and achieve accurate differentiation between fault feature points and anomalies. The subsequently introduced nonlinear scaling factor can nonlinearly map the signals of each sensor to a unified interval, effectively eliminating the magnitude differences and noise interference between multi-sensor monitoring signals, while amplifying the identification of weakly coupled fault features. This effectively solves the technical problem that traditional signal normalization methods easily cause weak fault features to be masked or ignored.
[0035] Furthermore, S2 includes the following steps: S21. Determine the weight matrix scale through cross-validation based on the actual monitoring needs of the equipment. Initialize the single-channel weight matrix The initial value is generated by the signal mean bias; S22, Arbitrary sensor channel signals of the normalized multi-sensor monitoring signals for industrial machinery. The normalized channel signal Set as core analysis parameter the remaining The normalized signals from each channel are subjected to principal component analysis for dimensionality reduction, generating a one-dimensional auxiliary parameter set. , build ,in Weighting of auxiliary parameters based on correlation coefficients enhances the characteristics of parameters with high correlation. For signal standard deviation One-dimensional auxiliary parameter set for the signal standard deviation The sensor channel signal after S1 adaptive normalization is divided by the normalized signal. Other A multi-channel sensor signal matrix for and covariance; S23. Calculate the Mahalanobis-weighted Euclidean distance between the auxiliary weighting parameters and the weight matrix to locate the best matching unit (BMU), as shown in the following formula:
[0036] in The spatial coordinates of the best matching unit in the three-dimensional weight matrix. The corresponding position in the weight matrix The element value at that position, For the first Each sensor channel signal The first after adaptive normalization The value at each position, , ; S24. Using the best matching unit as the center, set a dynamic neighborhood radius that decreases linearly with the number of iterations. Determine the neighborhood coordinate set:
[0037] in , The initial neighborhood radius is preset. This represents the current iteration number. This is the preset maximum number of iterations.
[0038] Specifically, this invention determines the scale of the weight matrix through cross-validation and initializes the single-channel three-dimensional weight matrix using signal mean bias, enabling the weight matrix to accurately adapt to the actual monitoring needs of industrial machinery. This provides a working-condition-appropriate framework for single-channel coupled fault feature mining. The normalized signal of the single channel is set as the core analysis parameter, while the signals of other channels are used as auxiliary parameters after dimensionality reduction through principal component analysis. A parameter set containing the correlation coefficient weights of the auxiliary parameters is also constructed, which can effectively strengthen the feature information with high correlation between the core parameters and auxiliary parameters, highlight the coupled fault association features within a single channel, and filter out interference from irrelevant features. Furthermore, the optimal matching unit is accurately located by calculating the Mahalanobis weighted Euclidean distance, allowing the mining of coupled fault features to accurately anchor the corresponding unit in the weight matrix, improving the targeting of feature mining. At the same time, a dynamic neighborhood radius that decays linearly with the number of iterations is set, and the corresponding neighborhood coordinate set is determined, realizing global mining of coupled fault features in the early stage of iteration and fine optimization in the later stage of iteration, taking into account both the breadth and depth of feature mining. The entire process can adaptively carry out coupled fault association feature mining within a single channel without manual annotation, effectively extracting the key features of coupled faults in a single channel.
[0039] Furthermore, following S24, the process also includes feature similarity quantification and weight matrix iterative update of the normalized industrial machinery multi-sensor monitoring signals: S25. Calculate the cosine distance of the core parameters respectively. Chebyshev weighted distance with auxiliary parameters The comprehensive feature similarity between data is obtained by fusing data using an improved Softmax function. The formula is as follows:
[0040]
[0041]
[0042] in In order to be with the first Normalized signal of each sensor channel The same-dimensional neighborhood mask signal sequence, centered on the currently processed sampling point, retains the same signal within a preset neighborhood range. Consistent signal values, values outside the neighborhood range are taken as the signal value of the current sampling point, used to match... Extracting local signal difference features by subtracting element by element; For the first Sensor channel neighborhood mask signal sequence The extracted local neighborhood feature set consists of statistical features, dimensionality reduction features, or model learning features of the signal within the neighborhood, and is used to characterize the local operating state features within the neighborhood of the current sampling point. For the first Sensor channel neighborhood mask signal sequence The number of valid sampling points included is used to characterize the size of the neighborhood range, and its value is adaptively determined according to the dynamic response characteristics of the industrial equipment. S26. Combining the Manhattan distance of the neighborhood spatial coordinates with the similarity of the comprehensive data features, calculate the gradient descent optimization amount of the weight matrix unit. The weight matrix is iteratively updated using stochastic gradient descent, as shown in the following formula:
[0043]
[0044] in, The learning rate; The L2 regularization coefficient; For the number of iterations, For a one-dimensional auxiliary parameter set Location The corresponding element value, For the first Each sensor channel signal The first after adaptive normalization The value of each position; S27. All normalized industrial machinery multi-sensor monitoring signals The unsupervised mining operation is performed sequentially on each sensor channel to adaptively generate a pre-trained weight matrix for each channel. This enables in-depth mining and extraction of single-channel coupling fault characteristics.
[0045] Specifically, this invention calculates the cosine distance of the core parameters and the Chebyshev weighted distance of the auxiliary parameters respectively, which can accurately quantify the feature similarity of different types of parameters. Then, the two types of distances are fused by an improved Softmax function to obtain a comprehensive feature similarity, effectively balancing the contribution of core parameters and auxiliary parameters in feature similarity determination. This makes the similarity results more consistent with the actual correlation rules of coupled fault features within a single channel, avoiding the limitations of single distance calculation in feature similarity determination. Subsequently, the gradient descent optimization amount of the weight matrix unit is calculated by combining the Manhattan distance of the neighborhood spatial coordinates and the comprehensive feature similarity of the data. The weight matrix is iteratively updated through stochastic gradient descent. At the same time, the learning rate is dynamically adjusted by cosine annealing and an L2 regularization coefficient is introduced to prevent... Overfitting makes the iterative optimization of the weight matrix more accurate and stable, ensuring the convergence efficiency of the iteration process while avoiding overfitting issues in the weight matrix. It can accurately adapt to the mining requirements of single-channel coupled fault features. Finally, the above unsupervised mining operation is performed sequentially on all sensor channels, which can adaptively generate pre-trained weight matrices for each channel. No manual annotation is required throughout the process. It can deeply mine the coupled fault correlation features within each single channel, giving the generated pre-trained weight matrix physical constraints and accurately reflecting the coupled fault feature patterns of each channel. This provides a high-quality and targeted single-channel weight foundation for the construction of the global weight matrix in the subsequent cross-channel attention weight iterative fusion process, effectively improving the accuracy of subsequent deep collaborative fusion of multi-channel coupled fault features.
[0046] Furthermore, S3 includes the following steps: S31. Merge the pre-trained weight matrices from all sensor channels to construct a global weight matrix. The optimal matching unit between the normalized multi-sensor monitoring signals of industrial machinery and the global weight matrix is calculated to generate the initial data membership matrix. This lays the data foundation for the fusion of multi-channel information; S32. Initialize cluster categories using a small number of labeled samples, and construct a clustering system containing... Category matrix of data points Cluster centers for each category are determined through max pooling convolution operations. ,in To indicate the number of labeled samples; S33. Based on the normalized multi-sensor monitoring signals of industrial machinery, update the global weight matrix with the current cluster center, recalculate the data membership matrix and category matrix, and set the convergence threshold. If the change in cluster centers between two consecutive iterations is less than If the iteration terminates, the optimized global fusion weight matrix and coupled fault diagnosis clustering model are obtained. S34. For each full-dimensional sensor signal sequence of the normalized industrial machinery multi-sensor monitoring signal, calculate its relationship with each category matrix. distance The signal sequence is assigned to the category with the smallest distance to generate industrial mechanical coupling fault diagnosis results. ,in For the first Each sensor channel signal The corresponding fault diagnosis label.
[0047] Specifically, this invention integrates the pre-trained weight matrices of each sensor channel to construct a global weight matrix. It then generates an initial data membership matrix by combining the optimal matching unit between the full-dimensional sensor signals and the global weight matrix. This effectively integrates the coupled fault features mined from each single channel, establishing a unified and condition-appropriate framework for the deep fusion of multi-channel coupled fault features. This allows multi-channel information fusion to be carried out based on high-quality single-channel feature weights. By using a small number of labeled samples to initialize cluster categories and determining cluster centers through max-pooling convolution, it accurately matches the actual working conditions of scarce fault labeling samples in industrial settings. This effectively solves the technical pain point of traditional fault diagnosis methods' high dependence on a large number of labeled samples, significantly reducing labeling costs. Based on the full-dimensional input signals and the current cluster centers, the global weight matrix, data membership matrix, and category matrix are iteratively updated until the change in cluster centers is less than the convergence threshold. This allows the global fusion weight matrix and the coupled fault diagnosis clustering model to be continuously optimized during iteration, achieving deep collaborative fusion of multi-channel coupled fault features. This fully captures the cross-channel multi-dimensional correlation features of coupled faults, overcoming the shortcomings of existing methods such as poor cross-channel feature fusion and difficulty in mining deep features of coupled faults. By calculating the distance between each full-dimensional sensor signal sequence and each category matrix and assigning it to the category with the smallest distance, different types of coupled faults in industrial machinery can be accurately distinguished, generating reliable diagnostic results for coupled faults in industrial machinery. This effectively improves the accuracy of coupled fault classification and identification under low-labeled samples. The coupled fault diagnosis clustering model constructed throughout the process organically combines single-channel coupled fault feature mining with deep fusion of multi-channel features, fully leveraging the value of multi-sensor monitoring signals. This allows fault diagnosis to not only conform to the feature patterns of single channels but also capture cross-channel coupling correlations. It provides a reliable model foundation and diagnostic result support for subsequent model parameter optimization and performance verification, as well as coupled fault feature quantification and source tracing analysis. At the same time, it significantly improves the engineering applicability and fault diagnosis accuracy of the method under low-labeled sample conditions.
[0048] Furthermore, in S3, the convergence threshold The value is When initializing cluster categories using a small number of labeled samples, the labeled samples are selected from the normal operation status of industrial machinery, the fault status of hydraulic system, and the fault status of power system, with a selection quantity of 10 / 30 / 50 / 100.
[0049] Specifically, this invention sets the convergence threshold of the clustering iteration to be... It can set precise termination criteria for the iterative training of coupled fault diagnosis clustering models, effectively avoiding problems such as insufficient model iteration and inadequate optimization of global fusion weight matrix and cluster centers due to excessively large convergence thresholds. It can also prevent excessive iteration and waste of computational resources caused by excessively small thresholds, allowing the model to ensure the deep fusion effect of coupled fault features while taking into account iterative efficiency, and ensuring that the final coupled fault diagnosis clustering model has high-precision fault classification capabilities. During cluster category initialization, labeled samples are selected from the normal operating state of industrial machinery, the fault state of hydraulic systems, and the fault state of power systems. This precisely matches the core application scenario of this invention for coupled fault diagnosis of electromechanical-hydraulic systems, and matches the core fault distribution types of industrial machinery. This allows the cluster category initialization to anchor the most representative equipment operating state in the industrial field, improving the accuracy and fit of the cluster center initialization. Selecting different numbers of labeled samples (10 / 30 / 50 / 100) not only fully adapts to the actual working conditions where fault labeled samples are scarce in industrial fields, meeting the model initialization requirements under low labeled sample conditions, but also adapts to industrial scenarios with different labeled costs and data labeled conditions. At the same time, setting different numbers of labeled samples can also verify the diagnostic performance of the model under various limited labeled working conditions, further strengthening the robustness and adaptability of the coupled fault diagnosis clustering model under different low labeled sample scenarios. This makes the cluster category initialization process more in line with the actual industrial application needs, laying an accurate and reliable category foundation for subsequent model iteration training and accurate classification and identification of coupled faults, and effectively improving the stability and accuracy of coupled fault diagnosis under low labeled sample conditions.
[0050] Furthermore, S4 includes the following steps: S41. Optimize the weight matrix scale based on the actual operating conditions of industrial machinery. Neighborhood radius calculate, This represents the actual number of test samples. Take the positive integer value of the calculation result; S42. The measured data of the multi-sensor monitoring signals of industrial machinery are divided into training set and test set in a ratio of 7:3. The 10-fold cross-validation method is used to verify the stability and generalization ability of the coupled fault diagnosis clustering model, and to ensure the effectiveness of unsupervised adaptive feature mining and multi-channel information fusion. S43. Test the diagnostic accuracy (ACC) and F1 score of the coupled fault diagnosis clustering model under different numbers of labeled samples, different numbers of sensor channels, and different sample distribution conditions to verify the robustness of the model under actual industrial conditions such as sample imbalance, limited labeling, and limited channels.
[0051] Specifically, this invention optimizes core hyperparameters such as weight matrix scale, neighborhood radius, and maximum iteration count by combining them with the actual operating conditions of industrial machinery. The optimal weight matrix scale is calculated according to a specific formula, ensuring that the settings of these core hyperparameters accurately match the actual number of test samples and the real operating conditions of industrial equipment. This effectively avoids the problem of poor model adaptability caused by blindly setting hyperparameters, making the parameter system of the coupled fault diagnosis clustering model more closely aligned with the actual diagnostic needs of industrial sites. The measured data of multi-sensor monitoring signals from industrial machinery are divided into training and testing sets in a 7:3 ratio, providing reasonable and sufficient data support for model performance verification. Combined with 10-fold cross-validation, the model's stability and generalization ability are verified, effectively eliminating the randomness of single-validation tests and comprehensively and accurately verifying the actual effectiveness of the model's unsupervised adaptive feature mining and multi-channel information fusion stages, ensuring the reliability of the model's feature processing and fusion capabilities. The diagnostic accuracy and F1 score of the model were tested under different numbers of labeled samples, different numbers of sensor channels, and different sample distributions. This allowed for targeted verification of the model's performance under various real-world industrial conditions, such as sample imbalance, limited labeling, and limited number of sensor channels. The robustness of the model was fully verified, effectively addressing the technical pain point of poor adaptability and robustness of existing fault diagnosis methods under complex and changing industrial conditions. The entire process of model parameter optimization and performance verification achieved comprehensive optimization and multi-dimensional testing of the coupled fault diagnosis clustering model, resulting in better parameter settings and more stable performance. This ensures that the model maintains high-precision coupled fault diagnosis capabilities under various real-world industrial conditions, providing a solid performance guarantee for the model's engineering implementation and practical industrial applications.
[0052] Furthermore, S5 includes the following steps: S51. Visualize the membership data distribution of the feature clusters after multi-channel information fusion of the multi-sensor monitoring signals of industrial machinery, generate a heat map, and intuitively distinguish the normal operating status of the equipment from the coupled fault feature clusters of different types of industrial machinery. S52. Statistically analyze key features such as the average value of multi-sensor signals and the proportion of fault-related components for each feature cluster, quantitatively analyze the signal feature patterns of different types of industrial machinery coupling faults, and combine the results of unsupervised adaptive feature mining to locate the correlation between features and industrial machinery coupling faults. S53. By combining the actual physical topology and component connection relationships of complex industrial machinery, the characteristic clusters of coupled faults of each industrial machinery are mapped to the actual components and subsystems of the industrial machinery. The source components of coupled faults of industrial machinery are located, the cross-subsystem evolution and propagation laws of faults are revealed, and targeted hierarchical precision operation and maintenance and fault handling suggestions are formed.
[0053] Specifically, this invention visualizes the membership data distribution of feature clusters obtained by multi-channel information fusion of multi-sensor monitoring signals from industrial machinery and generates heat maps. This allows for a clear and intuitive distinction between the normal operating status of equipment and different types of coupled fault feature clusters in industrial machinery, making the feature distribution patterns of coupled faults visible and breaking through the limitations of traditional black-box fault diagnosis models that lack interpretability. By statistically analyzing key features such as the average multi-sensor signals of each feature cluster and the proportion of fault-related components, the signal feature patterns of coupled faults in different types of industrial machinery are quantitatively analyzed. Combined with unsupervised adaptive feature mining results, the correlation between features and coupled faults in industrial machinery is located, enabling precise quantification of the contribution of each sensor feature to the coupled fault. This establishes a precise mapping relationship between fault features and coupled faults, elevating the analysis of coupled faults from a qualitative to a quantitative level, thus improving the scientific rigor and accuracy of fault feature analysis. By combining the actual physical topology and component connections of complex industrial machinery, the characteristic clusters of coupled faults in various industrial machines are mapped to the actual components and subsystems of the machinery. This enables precise location of the source component of the coupled fault, while effectively revealing the cross-subsystem evolution and propagation patterns of the fault. This overcomes the shortcomings of existing fault diagnosis methods in the ability to locate and trace coupled faults, solving the technical problem of difficulty in tracing the causes of coupled faults in industrial settings. The resulting targeted, hierarchical, and precise operation and maintenance and fault handling recommendations ensure that industrial machinery operation and maintenance is no longer done blindly, but rather has clear quantitative basis and specific handling directions. It effectively transforms the results of coupled fault diagnosis into practical guidance for equipment operation and maintenance in industrial settings, achieving end-to-end implementation from precise identification of coupled faults to source location and then to the output of operation and maintenance recommendations. This significantly enhances the engineering application value of this invention in real-world industrial scenarios, effectively reduces the chain reactions caused by coupled faults, and lowers equipment operation and maintenance costs and production safety risks.
[0054] A storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for diagnosing industrial mechanical coupling faults based on adaptive feature clustering for low-labeled samples.
[0055] A computer device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method for diagnosing industrial mechanical coupling faults based on adaptive feature clustering under low-labeled samples.
[0056] Specifically, this invention stores the corresponding computer program on a storage medium, and the computer device carries and executes the program. This transforms the industrial machinery coupled fault diagnosis method based on adaptive feature clustering under low-labeled samples from an algorithmic level into a feasible and executable hardware and software solution. This effectively overcomes the limitations of pure algorithms in practical industrial applications, enabling the engineering deployment and practical application of this diagnostic method in industrial scenarios. The storage medium, with its storable, transmissible, and reusable characteristics, allows the diagnostic method to be deployed, replicated, and promoted quickly across various computer devices in different industrial sites, eliminating the need for repeated development and significantly improving the efficiency and ease of application. The computer device, integrating memory and processor, can read and process multi-sensor monitoring signals from industrial machinery in real time, achieving online and automated diagnosis and tracing of coupled faults. This meets the actual needs of real-time monitoring of equipment operating status and timely fault handling in industrial production, eliminating the inefficient mode of manual signal analysis and fault diagnosis. Simultaneously, after the storage medium and computer equipment execute the program, they can fully reproduce all the technical effects of the diagnostic method of this invention. Even under actual industrial conditions such as scarce labeled samples, heterogeneous multi-sensor signals, and unbalanced sample distribution, it can still achieve accurate identification, source location, and evolutionary pattern analysis of coupled faults. It can also provide quantitative fault tracing evidence and targeted handling suggestions for on-site operation and maintenance, fully leveraging the advantages of this invention's low labeling dependence, high diagnostic accuracy, strong operational robustness, and good scalability. Only simple parameter adjustments are needed according to the actual needs of different industrial machinery to adapt it for coupled fault diagnosis scenarios in various equipment such as engineering machinery and industrial bearings. Furthermore, the design of the storage medium and computer equipment allows the coupled fault diagnosis method of this invention to be deeply integrated into intelligent monitoring systems for industrial equipment, promoting the intelligent operation and maintenance upgrade of industrial production, effectively reducing cascading equipment failures caused by coupled faults, lowering the operation and maintenance costs of industrial equipment, ensuring operational efficiency and production safety in industrial production, and further expanding the engineering application value and practical application scope of this invention.
[0057] The following is a specific embodiment of the present invention: Using four publicly available bearing fault datasets—CWRU, IMS, OU, and PU—as verification cases, this paper details the complete implementation process of the present invention and verifies the effectiveness and generalization ability of the algorithm in mechanical fault diagnosis scenarios through multiple datasets and different limited annotation conditions.
[0058] Step 1: Multi-dataset signal acquisition and preprocessing Step 11: Dataset Selection and State Division Four typical publicly available datasets of mechanical failures were selected as monitoring objects, covering different failure types, operating conditions, and data collection scenarios: CWRU bearing dataset: Collects vibration signals of bearings at the motor drive end and fan end, and classifies the system status into four categories: normal state, rolling element fault, inner ring fault, and outer ring fault; IMS bearing dataset: Collects vibration signals of multiple bearings operating synchronously. The system status is divided into two categories: normal state and bearing fault, focusing on the progressive fault and coupling characteristics under long-term operation. OU Gearbox Dataset: Collects gearbox housing vibration signals, and classifies the system status into three categories: normal state, gear wear fault, and bearing rolling element fault, covering typical coupling faults of the transmission system; PU bearing dataset: Collects bearing vibration signals under different loads and speeds. The system status is divided into three categories: normal state, inner ring fault, and outer ring fault, focusing on the robustness verification of fault diagnosis under variable working conditions.
[0059] Steps 1 and 2: Heterogeneous signal normalization processing The normalization method proposed in this invention is used to process heterogeneous sensor signals from various datasets. All sensor signals are scaled to a unified range of [0,1] through linear transformation, eliminating amplitude and magnitude differences between different datasets and sensors, thus completing the standardization of heterogeneous signals and preparing data for subsequent adaptive feature mining and cross-channel fusion.
[0060] Step 2: Single-channel unsupervised adaptive feature mining.
[0061] Step 2: Determine the scale of the weight matrix through multiple sets of tests and optimizations. K =150, initialize the single-channel weight matrix Set neighborhood radius Maximum number of iterations Initial learning rate =0.01, L2 regularization coefficient =0.001.
[0062] Step 22: Sequentially use the sensor signal of one dimension as the core analysis parameter, and reduce the dimensionality of the normalized signals of other dimensions to one-dimensional auxiliary parameters through principal component analysis, and construct core-auxiliary parameter data pairs for each channel. .
[0063] Steps 2 and 3: Calculate the Euclidean distance between the core-auxiliary parameter data pairs of each channel and the corresponding weight matrix to locate the best matching unit; set the distance weight adjustment threshold. =0.6, calculate the cosine distance of the core parameters respectively. Chebyshev weighted distance with auxiliary parameters The data comprehensive feature similarity is obtained by fusing data using an optimized Softmax function. .
[0064] Step 24: Combining the Manhattan distance of the neighborhood spatial coordinates with the similarity of the comprehensive data features, calculate the gradient descent optimization amount of the weight matrix unit. The weight matrix is iteratively updated by adjusting it through stochastic gradient descent, thus adaptively obtaining the pre-trained weight matrix for each channel. .
[0065] Step 3: Cross-channel multi-sensor information fusion.
[0066] Step 3: First, merge the pre-trained weight matrices from each channel to construct the global weight matrix. Calculate the optimal matching unit between the full-dimensional sensor signals and the global weight matrix to generate the initial data membership matrix. N To build a basic framework for multi-channel information fusion.
[0067] Step 3.2: From the three states of normal state, hydraulic system failure, and power system failure, randomly select 10 / 30 / 50 / 100 labeled samples respectively to complete the clustering category initialization and construct the category matrix. The initial cluster centers for each category are determined through max pooling convolution operations. .
[0068] Step 3: Set the iterative convergence threshold Based on the input signal and the current cluster center, the global weight matrix is iteratively updated, and the data membership matrix and category matrix are recalculated to achieve deep fusion and collaborative learning of the coupled fault features of each channel. The iteration ends when the change in the cluster center is less than the convergence threshold.
[0069] Steps 3 and 4: Calculate the distance between each sensor signal sequence in the test set and each category matrix. The signal sequence is assigned to the cluster category with the smallest distance, and after multi-channel information fusion, the fault diagnosis result is output.
[0070] Step 4: Model Validation and Parameter Optimization.
[0071] Step 4: First, divide the preprocessed sensor dataset into training and testing sets in a 7:3 ratio. Use 10-fold cross-validation to verify the model's stability. At the same time, optimize hyperparameters such as learning rate and batch size using a grid search method to ensure the efficiency and accuracy of unsupervised adaptive feature mining and multi-channel information fusion.
[0072] Step 42: Train and test the proposed method and the comparative method on four public datasets respectively, test the diagnostic accuracy (ACC) and F1 score of the model, and fully verify the robustness and adaptability of the model in unsupervised adaptive feature mining and multi-channel information fusion under real working conditions.
[0073] Step 4.3, Test Results Figure 2 As shown, the proposed model (Our) outperforms both FS-learning and S3M in both F1 score and accuracy across all four datasets, forming a complete outermost contour in the radar image, demonstrating the algorithm's comprehensive diagnostic performance advantage across multiple scenarios. FS-learning, based on self-supervised learning to mine unlabeled signal features, transfers these features to an improved Siamese network, combining unlabeled representation learning with few-shot learning to enhance the model's robustness and generalization in few-shot scenarios. S3M is a two-stage semi-self-supervised model. It first learns global and local features of unlabeled samples through a dual-extractor and interface task, then freezes the feature extractor and trains the classifier with limited labeled samples.
[0074] Step 5: Fault Feature Cluster Localization and Interpretability Verification The clustering model output after multi-channel information fusion is visualized to generate a heatmap of feature cluster membership distribution, such as... Figure 3 As shown, the feature clusters corresponding to different states are intuitively distinguished on four public datasets: CWRU dataset: Clearly distinguishes normal state feature clusters, inner circle fault feature clusters, and outer circle fault feature clusters (OF-C). The feature clusters of different fault modes have high spatial separation and no obvious overlap. IMS dataset: Intuitively presents the normal state feature cluster and bearing fault feature cluster. The fault cluster is concentrated in the high-dimensional feature space, reflecting the typical feature pattern of bearing progressive failure under long-term operation. OU Gearbox Dataset: Effectively separates normal state feature clusters, gear wear fault feature clusters, and rolling element fault feature clusters. The boundaries of gear and bearing fault feature clusters are clear, and the fault features of different components of the transmission system can be distinguished. PU dataset: It accurately divides the normal state feature cluster, inner circle fault feature cluster, and outer circle fault feature cluster, and maintains good inter-cluster separation under varying operating conditions, verifying the robustness of the algorithm to operating condition disturbances.
[0075] This invention has good applicability to coupled fault diagnosis of engineering machinery and general industrial equipment such as industrial bearings. Without departing from the spirit and essence of this invention, those skilled in the art can flexibly adjust the relevant parameters of unsupervised adaptive feature mining and multi-channel information fusion according to actual application scenarios, making corresponding adjustments and modifications to this invention. All such adjustments and modifications should fall within the protection scope of the appended claims.
[0076] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for diagnosing coupled faults in industrial machinery based on adaptive feature clustering under low-labeled samples, characterized in that, Includes the following steps: S1. Collect monitoring signals from multiple sensors of industrial machinery, and perform heterogeneous signal kernel density adaptive normalization processing on the monitoring signals of multiple sensors of industrial machinery to eliminate the magnitude difference and noise interference of the multi-sensor signals, while amplifying the identification of weak coupling fault characteristics. S2. Perform three-dimensional weighted primary and secondary feature adaptive mining on the normalized industrial machinery multi-sensor monitoring signals to generate pre-trained weight matrices with physical constraints for each sensor channel. S3. Based on the pre-trained weight matrix of each channel, the channel attention weight iterative fusion is completed on the normalized industrial machinery multi-sensor monitoring signal. Combined with a small number of labeled samples, the deep collaborative fusion of multi-channel coupled fault features is realized, a coupled fault diagnosis clustering model is constructed, and the classification and identification of coupled faults are completed. S4. Optimize the model parameters and verify the performance of the coupled fault diagnosis clustering model constructed in steps S1-S3. Optimize the core hyperparameters and verify the stability, generalization ability and robustness of the model under different industrial conditions. S5. Conduct quantitative source analysis of coupled fault characteristics on the coupled fault diagnosis results of industrial machinery output by the coupled fault diagnosis clustering model, quantify the contribution of each sensor characteristic to the coupled fault, locate the source of the coupled fault, and reveal the evolution and propagation law of the fault.
2. The industrial mechanical coupling fault diagnosis method based on adaptive feature clustering under low-labeled samples according to claim 1, characterized in that, S1 includes the following steps: S11. Collect S-dimensional sensor monitoring signals from complex industrial machinery and construct a signal set. ,in For the first A sequence of continuously sampled data from a sensor. The sampling length of a single-channel signal is set according to the monitoring needs of industrial machinery; S12, Calculation of the first based on improved kernel density estimation Each data point in the sensor signal The distribution density is introduced by incorporating data distribution weight coefficients calculated from the data outlier degree. The formula is as follows: in To improve the bandwidth parameter of kernel density estimation, the optimal value is determined by grid search method; These are the data distribution weighting coefficients. For the first Each sensor channel signal The mean, For the first Each sensor channel signal Standard deviation; For the first Each sensor channel signal The The value at each position, , Let be the Cauchy kernel density function, expressed as: , For general variables; S13. Calculate the area under the kernel density curve using the 10th-order Gauss-Legendre adaptive quadrature method to determine if the constraints are satisfied. Effective upper and lower boundaries of the signal ,in The effective data percentage constraint threshold is calculated as follows: ; S14. Introduce a non-linear scaling factor. The signals from each sensor are nonlinearly scaled to a unified interval of [0,1] to achieve adaptive normalization of heterogeneous signals. The transformation formula is as follows: , in For the first Each sensor channel signal The first after adaptive normalization The value at each position, .
3. The industrial mechanical coupling fault diagnosis method based on adaptive feature clustering under low-labeled samples according to claim 1, characterized in that, S2 includes the following steps: S21. Determine the weight matrix scale through cross-validation based on the actual monitoring needs of the equipment. Initialize the single-channel weight matrix The initial value is generated by the signal mean bias; S22, Arbitrary sensor channel signals of the normalized multi-sensor monitoring signals for industrial machinery. The normalized channel signal Set as the core analysis parameter, the rest The normalized signals from each channel are subjected to principal component analysis for dimensionality reduction, generating a one-dimensional auxiliary parameter set. , build ,in Weighting of auxiliary parameters based on correlation coefficients enhances the characteristics of parameters with high correlation. For signal standard deviation One-dimensional auxiliary parameter set for the signal standard deviation The sensor channel signal after S1 adaptive normalization is divided by the normalized signal. Other A multi-channel sensor signal matrix for and covariance; S23. Calculate the Mahalanobis-weighted Euclidean distance between the auxiliary weighting parameters and the weight matrix to locate the best matching unit (BMU), as shown in the following formula: in The spatial coordinates of the best matching unit in the three-dimensional weight matrix. The corresponding position in the weight matrix The element value at that position, For the first Each sensor channel signal The first after adaptive normalization The value at each position, , ; S24. Using the best matching unit as the center, set a dynamic neighborhood radius that decreases linearly with the number of iterations. Determine the neighborhood coordinate set: in , The initial neighborhood radius is preset. This represents the current iteration number. This is the preset maximum number of iterations.
4. The industrial machinery coupling fault diagnosis method based on adaptive feature clustering under low-labeled samples according to claim 3, characterized in that, Following S24, the process also includes feature similarity quantification and weight matrix iterative update of the normalized industrial machinery multi-sensor monitoring signals: S25. Calculate the cosine distance of the core parameters respectively. Chebyshev weighted distance with auxiliary parameters The comprehensive feature similarity between data is obtained by fusing data using an improved Softmax function. The formula is as follows: in In order to be with the first Normalized signal of each sensor channel The same-dimensional neighborhood mask signal sequence, centered on the currently processed sampling point, retains the same signal within a preset neighborhood range. Consistent signal values, values outside the neighborhood range are taken as the signal value of the current sampling point, used to match... Extracting local signal difference features by subtracting element by element; For the first Sensor channel neighborhood mask signal sequence The extracted local neighborhood feature set consists of statistical features, dimensionality reduction features, or model learning features of the signal within the neighborhood, and is used to characterize the local operating state features within the neighborhood of the current sampling point. For the first Sensor channel neighborhood mask signal sequence The number of valid sampling points included is used to characterize the size of the neighborhood range, and its value is adaptively determined according to the dynamic response characteristics of the industrial equipment. S26. Combining the Manhattan distance of the neighborhood spatial coordinates with the similarity of the comprehensive data features, calculate the gradient descent optimization amount of the weight matrix unit. The weight matrix is iteratively updated using stochastic gradient descent, as shown in the following formula: in, The learning rate; The L2 regularization coefficient; For the number of iterations, For a one-dimensional auxiliary parameter set Location The corresponding element value, For the first Each sensor channel signal The first after adaptive normalization The value of each position; S27. All normalized industrial machinery multi-sensor monitoring signals The unsupervised mining operation is performed sequentially on each sensor channel to adaptively generate a pre-trained weight matrix for each channel. This enables in-depth mining and extraction of single-channel coupling fault characteristics.
5. The industrial mechanical coupling fault diagnosis method based on adaptive feature clustering under low-labeled samples according to claim 3 or 4, characterized in that, S3 includes the following steps: S31. Merge the pre-trained weight matrices from all sensor channels to construct a global weight matrix. The optimal matching unit between the normalized multi-sensor monitoring signals of industrial machinery and the global weight matrix is calculated to generate the initial data membership matrix. ; S32. Initialize cluster categories using a small number of labeled samples, and construct a clustering system containing... Category matrix of data points Cluster centers for each category are determined through max pooling convolution operations. ,in To indicate the number of labeled samples; S33. Based on the normalized multi-sensor monitoring signals of industrial machinery, update the global weight matrix with the current cluster center, recalculate the data membership matrix and category matrix, and set the convergence threshold. If the change in cluster centers between two consecutive iterations is less than If the iteration terminates, the optimized global fusion weight matrix and coupled fault diagnosis clustering model are obtained. S34. For each full-dimensional sensor signal sequence of the normalized industrial machinery multi-sensor monitoring signal, calculate its relationship with each category matrix. distance The signal sequence is assigned to the category with the smallest distance to generate industrial mechanical coupling fault diagnosis results. ,in For the first Each sensor channel signal The corresponding fault diagnosis label.
6. The industrial mechanical coupling fault diagnosis method based on adaptive feature clustering under low-labeled samples according to claim 5, characterized in that, In S3, the convergence threshold The value is When initializing cluster categories using a small number of labeled samples, the labeled samples are selected from the normal operation status of industrial machinery, the fault status of hydraulic system, and the fault status of power system, with a selection quantity of 10 / 30 / 50 / 100.
7. The industrial mechanical coupling fault diagnosis method based on adaptive feature clustering under low-labeled samples according to claim 6, characterized in that, S4 includes the following steps: S41. Optimize the weight matrix scale based on the actual operating conditions of industrial machinery. Neighborhood radius Core hyperparameters, where the optimal weight matrix is scaled according to... calculate, This represents the actual number of test samples. Take the positive integer value of the calculation result; S42. The measured data of the multi-sensor monitoring signals of industrial machinery are divided into training set and test set in a ratio of 7:
3. The 10-fold cross-validation method is used to verify the stability and generalization ability of the coupled fault diagnosis clustering model, and to ensure the effectiveness of unsupervised adaptive feature mining and multi-channel information fusion. S43. Test the diagnostic accuracy and F1 score of the coupled fault diagnosis clustering model under different numbers of labeled samples, different numbers of sensor channels, and different sample distribution conditions to verify the robustness of the model under actual industrial conditions such as sample imbalance, limited labeling, and limited channels. The diagnostic accuracy is the ratio of the number of correctly diagnosed samples to the total number of test samples, and the F1 score is the harmonic mean of precision and recall.
8. The industrial mechanical coupling fault diagnosis method based on adaptive feature clustering under low-labeled samples according to claim 1, characterized in that, In S5, the following steps are included: S51. Visualize the membership data distribution of the feature clusters after multi-channel information fusion of the multi-sensor monitoring signals of industrial machinery, generate a heat map, and intuitively distinguish the normal operating status of the equipment from the coupled fault feature clusters of different types of industrial machinery. S52. Statistically analyze key features such as the average value of multi-sensor signals and the proportion of fault-related components for each feature cluster, quantitatively analyze the signal feature patterns of different types of industrial machinery coupling faults, and combine the results of unsupervised adaptive feature mining to locate the correlation between features and industrial machinery coupling faults. S53. By combining the actual physical topology and component connection relationships of complex industrial machinery, the characteristic clusters of coupled faults of each industrial machinery are mapped to the actual components and subsystems of the industrial machinery. The source components of coupled faults of industrial machinery are located, the cross-subsystem evolution and propagation laws of faults are revealed, and targeted hierarchical precision operation and maintenance and fault handling suggestions are formed.
9. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the industrial mechanical coupling fault diagnosis method based on adaptive feature clustering for low-labeled samples as described in any one of claims 1-8.
10. A computer device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, the processor executing the program to implement the industrial mechanical coupling fault diagnosis method based on adaptive feature clustering for low-labeled samples as described in any one of claims 1-8.