A method for identifying mechanical equipment fault data based on artificial intelligence

Through adaptive normalization, mixed energy entropy and Gaussian kernel feature enhancement strategies, the problems of traditional methods being sensitive to outliers and ignoring energy correlation between frequency bands are solved, the accuracy and stability of mechanical equipment fault diagnosis are improved, and the adaptability of the training process is optimized.

CN120508812BActive Publication Date: 2025-09-16SICHUAN JINHUA HEDIAN TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510998239.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-09-16
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Among the existing mechanical equipment fault diagnosis methods, traditional normalization methods are sensitive to outliers, feature extraction methods ignore the energy correlation between frequency bands, feature enhancement methods cannot fully utilize clustering information, noise samples interfere with model training, and fixed learning rates are difficult to adapt to dynamic fault modes.

Method used

An adaptive normalization method based on energy density is adopted, combined with the inter-band energy jump penalty term and mixed energy entropy, to construct a feature enhancement strategy of adaptive Gaussian kernel. A dual-threshold activation function and gradient-directed correction mechanism are used to construct a neural network and optimize the learning rate.

Benefits of technology

It effectively solves the problem of insufficient non-stationary signal processing capabilities, improves sensitivity to impact-type faults, strengthens the feature distinguishability under complex fault modes, enhances the robustness and diagnostic accuracy of the model, and optimizes the adaptability and convergence speed of the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508812B_ABST
    Figure CN120508812B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for mechanical equipment fault data identification based on artificial intelligence, and belongs to the field of data processing and artificial intelligence technology. It solves the problems existing in the prior art in fault diagnosis, such as insensitive signal processing, insufficient feature extraction, significant noise interference, and low model training efficiency. The method proposes an adaptive normalization method based on energy density to effectively suppress the influence of non-stationary signals and retain key features; by mixing energy entropy feature extraction with inter-band energy jump penalty terms, it significantly enhances the sensitivity to complex fault modes; constructs a fault feature enhancement strategy based on Gaussian potential well to strengthen the response of fault feature cluster areas; adopts weighted loss function and gradient directional correction mechanism to improve model robustness and accuracy; combines entropy weighted learning rate and covariance scaling strategy to achieve adaptive training and improve convergence speed and adaptability. The method improves the accuracy and efficiency of mechanical fault diagnosis and provides support for the development of industrial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing and artificial intelligence technology, and in particular to a method for identifying mechanical equipment fault data based on artificial intelligence. Background Art

[0002] During long-term operation, mechanical equipment often experiences various failures due to factors such as wear, aging, or improper operation. These failures typically cause abnormalities in equipment vibration, noise, temperature, and other signals. Therefore, accurately identifying and diagnosing mechanical equipment failures is crucial for improving equipment reliability, extending service life, reducing downtime, and lowering maintenance costs.

[0003] The Chinese invention patent with publication number CN120008926A proposes a method and device for fan bearing fault warning based on vibration characteristics. Vibration signals are collected in real time by vibration sensors installed at key parts of the fan, and the signals are denoised and filtered. Vibration signal characteristics are extracted through frequency domain analysis, and the operating data of the fan is fused with the vibration signal to improve diagnostic accuracy. Fault diagnosis is performed on the vibration characteristics in combination with machine learning algorithms to accurately identify the type, location and severity of bearing faults. Warning information is fed back to operation and maintenance personnel through the user interface to support timely maintenance. The device includes a vibration signal acquisition module, a signal processing module, a data fusion and feature extraction module, a fault diagnosis module and a warning module, which can provide efficient and accurate fan bearing fault warnings, ensure the stable operation of the fan, and reduce maintenance costs.

[0004] A Chinese invention patent, publication number CN119848674B, proposes a generator fault diagnosis method and related products based on vibration trend prediction, belonging to the technical field of generator fault diagnosis. The invention provides a generator fault diagnosis method based on vibration trend prediction. By decomposing the generator's historical vibration signal sequence into multiple subsequences and performing random function generation and residual function calculation on each subsequence, the invention can more carefully capture the characteristic information in the vibration signal, thereby improving the accuracy of fault diagnosis. The invention also utilizes a neural network model for prediction, combined with an iterative loop and residual function minimization strategy, to gradually approximate the actual vibration trend and improve diagnostic accuracy. Through iterative loops and multiple random function generation, and selecting the optimal result, the invention overcomes the nonlinearity and non-periodicity of the vibration signal, enhancing the stability and reliability of the prediction. The introduction of the residual sequence smoothes the vibration signal, reduces noise interference, and improves the stability of the prediction.

[0005] The existing methods for diagnosing mechanical equipment faults still have problems that need to be solved, including:

[0006] 1. Existing traditional normalization methods, such as minimum-maximum normalization and standard deviation normalization, are usually very sensitive to outliers or non-stationary signals, which may cause the loss of useful signal features, thereby affecting the accuracy of subsequent fault diagnosis.

[0007] 2. Traditional feature extraction methods such as wavelet entropy and spectral entropy fail to fully consider the energy correlation between different frequency bands, resulting in the inability to effectively extract high-precision features when dealing with impact-type faults.

[0008] 3. Existing feature enhancement methods cannot fully utilize clustering information to strengthen the response of fault features, especially in complex fault modes, and cannot more clearly display different fault features in the feature space.

[0009] 4. In the traditional training process, the presence of noise samples often leads to a deviation in the gradient update direction, reducing the accuracy of the model.

[0010] 5. Traditional training processes often use a fixed learning rate. For fault modes with large dynamic feature changes, it is difficult to adapt to the changes in their feature space, resulting in inefficient and unstable training processes. Summary of the Invention

[0011] The purpose of the present invention is to provide a method for mechanical equipment fault data identification based on artificial intelligence to solve the technical problems raised in the prior art, including the traditional normalization method being sensitive to outliers, the feature extraction method ignoring the energy correlation between frequency bands, the feature enhancement method being unable to fully utilize clustering information, the noise samples interfering with model training, and the fixed learning rate being difficult to adapt to dynamic fault modes.

[0012] In order to solve the above technical problems, the specific technical solutions of the present invention are as follows:

[0013] A method for identifying mechanical equipment fault data based on artificial intelligence, the method comprising the following steps:

[0014] Step S1: collecting vibration data of mechanical equipment, storing the collected vibration data in floating point format and retaining complete time domain amplitude information and timestamp sequence;

[0015] Step S2: Denoising the mechanical equipment vibration data to obtain denoised vibration data, removing the sensor's own electronic noise, environmental interference, and random fluctuations;

[0016] Step S3: labeling the denoised vibration data to obtain labeled vibration data, and labeling the denoised vibration data as a normal state or a corresponding fault type;

[0017] Step S4: Adopting an adaptive normalization method based on energy density to adaptively normalize the marked vibration data to obtain normalized vibration data;

[0018] Step S5: enhancing the sensitivity to shock-type faults by combining the inter-band energy jump penalty term, and then calculating the mixed energy entropy;

[0019] Step S6: constructing a feature enhancement strategy based on an adaptive Gaussian kernel to strengthen the response strength of the clustering area in the feature space;

[0020] Step S7: constructing a fault data recognition neural network;

[0021] Step S8: training the neural network to obtain a trained neural network;

[0022] Step S9: Use the trained neural network to identify mechanical equipment fault data.

[0023] Furthermore, the energy density-based adaptive normalization method in step S4 is expressed as:

[0024]

[0025] in, Indicates the normalized vibration data at the The first sample The value of each sampling point; Indicates the vibration data after annotation. Sample No. The value of each sampling point; Indicates the vibration data after annotation. Sample No. The local window mean of the samples at each sampling point; The vibration data after marking is Sample No. The local standard deviation of the sampling points; To prevent division by zero constant; Indicates the The global energy of the samples; represents the local window energy.

[0026] Furthermore, step S5 includes the following steps:

[0027] Step S51: performing wavelet packet decomposition on the normalized vibration data to obtain wavelet packet coefficient sequences of multiple sub-bands;

[0028] Step S52: Based on the wavelet packet coefficient sequence of each sub-band, the square sum of the absolute values ​​of all coefficients in the wavelet packet coefficient sequence is calculated to obtain the sub-band energy value;

[0029] Step S53: Based on the energy values ​​of all sub-bands, the energy value of each sub-band is divided by the sum of the energy values ​​of all sub-bands to obtain a sub-band relative energy value;

[0030] Step S54: Based on the relative energy value of each sub-band, the Shannon entropy is calculated and an inter-band energy jump penalty term is constructed to obtain the mixed energy entropy.

[0031] Furthermore, the mixed energy entropy in step S54 is expressed as:

[0032]

[0033] Where, Indicates the The mixing energy entropy of samples; is the correlation factor; Indicates the The first sample relative energy of the sub-bands; Indicates the The first sample relative energy of the sub-bands; is a logarithmic function.

[0034] Furthermore, step S6 includes the following steps:

[0035] Step S61: performing clustering processing on the relative energy entropy feature vectors of all samples to obtain the fault feature cluster center vector;

[0036] Step S62: Calculate the intra-cluster standard deviation based on the fault feature cluster center vector and the sample set belonging to the cluster;

[0037] Step S63: constructing a Gaussian kernel function based on the obtained cluster center vector and the intra-cluster standard deviation;

[0038] Step S64: multiply the relative energy entropy feature vector by the corresponding Gaussian kernel function output vector element by element to obtain an enhanced feature vector.

[0039] Furthermore, step S7 includes the following steps:

[0040] Step S71: defining a neural network structure, where the neural network consists of a fully connected deep neural network consisting of an input layer, three hidden layers, and an output layer;

[0041] Step S72: Initialize the weight matrix based on the enhanced feature vector so that the initial weight points to the direction of fault feature difference;

[0042] Step S73: defining a dual threshold activation function of the neural network;

[0043] Step S74: Define the hybrid energy entropy weighted loss of the neural network.

[0044] Furthermore, step S72 includes the following steps:

[0045] Step S721: Calculating the covariance matrix of the enhanced eigenvector based on the enhanced eigenvector;

[0046] Step S722: performing a singular value decomposition operation on the covariance matrix of the enhanced eigenvector and multiplying it by the scaling factor to obtain a basic weight matrix;

[0047] Step S723: Calculate the difference between the enhanced feature mean vector of each fault category sample and the global enhanced feature mean vector, and perform normalization constraints based on the Frobenius norm to obtain a fault sensitivity correction term matrix;

[0048] Step S724: Add the basic weight matrix and the fault-sensitive correction term matrix to obtain the initial weight matrix of the neural network.

[0049] Furthermore, the dual threshold activation function in step S73 is expressed as follows:

[0050]

[0051] in, Represents the input value of the activation function; is a dual threshold activation function; Represents the absolute value of the input value of the activation function; is a low threshold; is the slope attenuation coefficient in the low amplitude region; is the high threshold; is the slope attenuation coefficient in the high amplitude region; is a symbolic function.

[0052] Furthermore, step S74 includes the following steps:

[0053] Step S741: Calculate the average mixed energy entropy of all fault categories;

[0054] Step S742: Calculate the fault category weighting factor based on the average mixed energy entropy of the fault sample and the average mixed energy entropy of all fault categories;

[0055] Step S743: Construct a hybrid energy entropy weighted loss.

[0056] Furthermore, the mixed energy entropy weighted loss is expressed as:

[0057]

[0058] in, represents the mixed energy entropy weighted loss; Indicates that the sample belongs to The true label of the fault class; Indicates that the neural network predicts that the sample belongs to probability of class failure; is the total number of fault categories.

[0059] Compared with the prior art, the present invention has the following beneficial technical effects:

[0060] 1) The adaptive normalization method effectively solves the problem of insufficient processing ability of traditional normalization for non-stationary signals and retains the transient characteristics of the impact component.

[0061] 2) Hybrid energy entropy feature extraction significantly improves the sensitivity to shock-type faults through the inter-band energy jump penalty term.

[0062] 3) The Gaussian potential well enhancement strategy strengthens the response intensity of the fault feature clustering area and improves the feature distinguishability under complex fault modes.

[0063] 4) The dual-threshold activation function and gradient-directed correction mechanism reduce noise interference and enhance the robustness and diagnostic accuracy of the model.

[0064] 5) Adaptive learning rate and covariance scaling strategies optimize the training process and improve the model’s adaptability to dynamic failure modes and convergence speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0066] Figure 1 It is a schematic diagram of the overall process of the present invention.

[0067] Figure 2 This is a schematic diagram of the processing results of non-stationary signals using the adaptive normalization method of the present invention.

[0068] Figure 3 Schematic diagram of the fault separability comparison of different feature extraction methods of the present invention.

[0069] Figure 4 Schematic diagram comparing the fault recognition accuracy of different feature extraction methods of the present invention.

[0070] Figure 5 Schematic diagram of the effect of noise sample ratio on model robustness in the present invention. DETAILED DESCRIPTION

[0071] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0072] The present invention proposes a method for identifying mechanical equipment fault data based on artificial intelligence, such as Figure 1 As shown, the method includes the following steps:

[0073] Step S1: Collect vibration data of mechanical equipment, store the collected vibration data in floating point format and retain complete time domain amplitude information and timestamp sequence.

[0074] High-precision vibration sensors are deployed at key locations on mechanical equipment, and data acquisition cards are used to continuously acquire raw vibration data at a preset sampling frequency. Key locations for high-precision vibration sensors are typically motor drive terminals, equipment bases, or areas with high end cap strength.

[0075] For rotating equipment, the speed pulse signal is collected synchronously to achieve periodic interception. For impact equipment, the peak trigger acquisition mode is used to ensure the integrity of transient characteristics.

[0076] The collection process covers the entire operating range of the equipment, including no-load, rated load, and overload conditions. Time domain waveform data for no less than ten minutes is continuously collected for each condition, and the corresponding equipment health status label is recorded as a benchmark for subsequent supervised learning.

[0077] The vibration data is stored in floating point format, preserving complete time domain amplitude information and time stamp sequence.

[0078] Step S2: Denoise the vibration data of the mechanical equipment to obtain denoised vibration data, remove the electronic noise of the sensor itself, environmental interference and random fluctuations, and reduce the impact of noise on the vibration data.

[0079] Since the collected vibration data is inevitably mixed with noise, this noise may come from the electronic noise of the sensor itself, environmental interference, or random fluctuations during equipment operation.

[0080] To improve data quality and the accuracy of subsequent fault diagnosis, the vibration data is denoised. This invention employs commonly used methods, including low-pass filtering, wavelet transform denoising, and other denoising methods. Low-pass filtering removes high-frequency noise components while retaining the low-frequency characteristic signals of equipment vibration. Wavelet transform denoising utilizes the multi-resolution analysis characteristics of wavelet basis functions to decompose the signal into details and approximate coefficients at different scales. Noise detail coefficients are removed through thresholding, and the signal is reconstructed to achieve denoising.

[0081] Through data denoising, the impact of noise on vibration data can be effectively reduced, making the fault characteristics more clearly apparent.

[0082] Step S3: labeling the denoised vibration data to obtain labeled vibration data, and labeling the denoised vibration data as a normal state or a corresponding fault type.

[0083] The denoised vibration data is labeled to obtain labeled vibration data. Combined with equipment operation records, maintenance history and expert experience, manual labeling is used to mark the denoised vibration data as normal status or corresponding fault type.

[0084] To enable vibration data to be used for training and validating fault diagnosis models, the denoised vibration data is labeled. This process identifies the vibration data as either normal or corresponding to a fault type based on the equipment's actual operating status and known fault information. Fault types include bearing wear, broken gear teeth, shaft misalignment, wheel imbalance, and loose base.

[0085] The labeling process needs to be implemented manually by combining the equipment's operating records, maintenance history, and the experience of field experts. For example, if a bearing wear failure occurs in a certain time period, the vibration data collected during that time period will be labeled as bearing wear failure data.

[0086] The labeled data serves as training samples for supervised learning, helping the fault diagnosis model learn the vibration characteristic patterns under different states, thereby improving the model's ability to identify unknown fault data.

[0087] Step S4: Adopting an energy density-based adaptive normalization method to adaptively normalize the labeled vibration data to obtain normalized vibration data.

[0088] Equipment vibration data has problems such as large amplitude range differences and significant non-stationary characteristics. The conventional min-max normalization method is more sensitive to outliers. The presence of outliers can easily cause the normalized data to lose its original energy distribution characteristics, thereby affecting the accuracy of subsequent fault diagnosis.

[0089] The present invention adopts an adaptive normalization method based on energy density. By combining local statistics with the global energy ratio, it suppresses the influence of non-stationarity and retains the energy distribution characteristics. Based on the sliding window mechanism, the local mean, local standard deviation and energy term are calculated to adapt to the non-stationary characteristics of the signal.

[0090] Specifically, the adaptive normalization method based on energy density uses a sliding window to calculate the local window mean, and then combines the local standard deviation and the ratio of global energy to local window energy to normalize the labeled vibration data to obtain normalized vibration data, thereby achieving effective normalization of non-stationary vibration data. The adaptive normalization method based on energy density is expressed as:

[0091]

[0092] Indicates the normalized vibration data at the The first sample The value of sampling points; each sample contains a total of sampling points; Indicates the vibration data after annotation. Sample No. The value of each sampling point; Indicates the vibration data after annotation. Sample No. The local window mean of the samples at each sampling point; The vibration data after marking is Sample No. The local standard deviation of the sampling points; To prevent division by zero constant, used to avoid the denominator being zero, such as, ; Indicates the The global energy of the samples; Represents the local window energy, which means the vibration data after annotation is The first sample The signal energy in the window centered at the sampling point.

[0093] No. The set of vibration data sampling points after sample normalization It is expressed as follows:

[0094] s ̂ n =[ s ˆ 1 n ,…, s ˆ i n ,…, s ˆ L n ]

[0095] The vibration data after marking is Sample No. The local window mean of the samples at each sampling point The calculation method is expressed as:

[0096]

[0097] Among them, the summation interval covers the first A window centered at a sampling point; is the sliding window length, which defines the size of the local calculation interval; The vibration data after marking is The first sample Sampling point values; is a positive integer; Is a positive integer.

[0098] The vibration data after marking is Sample No. The local standard deviation of the sampling points Characterizes the signal fluctuation intensity within the window, and the calculation method is expressed as:

[0099]

[0100] The global energy summation covers all sampling points, The global energy of the samples The calculation method is expressed as:

[0101]

[0102] in, is the total number of sampling points for each sample; The vibration data after marking is Sample No. The value of the sampling point.

[0103] No. The local window energy of samples The calculation method is expressed as:

[0104]

[0105] It should be noted that The term characterizes the energy density ratio, which is used to preserve the energy distribution characteristics of non-stationary signals and maintain global energy consistency by scaling the normalized results.

[0106] In one embodiment, the processing effect of the adaptive normalization technique on non-stationary signals is analyzed, such as Figure 2As shown, three time-series subgraphs intuitively compare the processing results of the original vibration signal, the conventional normalization method and the adaptive normalization of the present invention. The original vibration signal contains obvious impact components (red marked areas). These transient characteristics are crucial in equipment fault diagnosis. Experiments show that the traditional minimum and maximum value normalization and standard deviation normalization methods do not produce obvious fault differences in the impact area, resulting in distortion of the impact characteristic waveform. The adaptive normalization of the present invention completely retains the transient characteristics of the original signal in the impact area and maintains a good normalization effect in the stable area. This shows that the sliding window energy density ratio calculation mechanism of the present invention can dynamically adjust the normalization strength according to the local characteristics of the signal, effectively solving the problem of insufficient processing ability of conventional methods for non-stationary signals.

[0107] Step S5: The sensitivity to shock-type faults is enhanced by combining the inter-band energy jump penalty term, and then the mixed energy entropy is calculated.

[0108] Fault features in vibration data are often overwhelmed by noise and difficult to extract directly. Traditional wavelet entropy methods ignore the energy correlation between frequency bands when processing, resulting in inaccurate and incomplete extraction of fault features.

[0109] The present invention enhances the sensitivity to shock-type faults by combining the inter-band energy jump penalty term, and then calculates the mixed energy entropy. The specific steps are as follows:

[0110] Step S51: performing wavelet packet decomposition on the normalized vibration data to obtain wavelet packet coefficient sequences of multiple sub-bands.

[0111] By performing wavelet packet decomposition on the normalized vibration data, the wavelet packet coefficient sequences of multiple sub-bands are obtained, which can be expressed as:

[0112]

[0113] in, Indicates the After wavelet packet decomposition of samples Sub-band coefficients; is a positive integer, It is a wavelet packet decomposition operation, which decomposes the input signal into coefficient sequences of multiple sub-bands through wavelet basis functions.

[0114] Step S52: Based on the wavelet packet coefficient sequence of each sub-band, the square sum of the absolute values ​​of all coefficients of the wavelet packet coefficient sequence is calculated to obtain the sub-band energy value.

[0115] Based on the wavelet packet coefficient sequence of each sub-band, the sub-band energy value representing the signal strength of the sub-band is obtained by calculating the square sum of the absolute values ​​of all coefficients in the sequence, which is expressed as:

[0116]

[0117] Where, Indicates the Sample No. The energy value of the sub-band, Indicates the number of wavelet decomposition layers; For absolute value calculation, the absolute value symbol in a complex number represents the modulus of the complex number and is used to calculate the magnitude of the complex coefficient.

[0118] It should be noted that The term covers all coefficients of the sub-band by summing up, reflecting the The energy of each frequency band is accumulated for 4 samples.

[0119] Step S53: Based on the energy values ​​of all sub-bands, the energy value of each sub-band is divided by the sum of the energy values ​​of all sub-bands to obtain a sub-band relative energy value.

[0120] Based on the energy values ​​of all sub-bands, the relative energy value representing the energy proportion of each sub-band is obtained by dividing the energy value of each sub-band by the sum of the energy values ​​of all sub-bands, forming the probability distribution of energy in each sub-band, which is expressed as:

[0121]

[0122] Where, Indicates the The first sample Relative energy values ​​of sub-bands; is a positive integer; is the total number of sub-bands decomposed by wavelet packet.

[0123] It should be noted that Item represents the The total energy of all sub-bands of samples.

[0124] Step S54: Based on the relative energy value of each sub-band, the Shannon entropy is calculated and an inter-band energy jump penalty term is constructed to obtain the mixed energy entropy.

[0125] Based on the relative energy values ​​of each sub-band, the Shannon entropy is calculated and the inter-band energy jump penalty term is constructed. Based on the energy distribution disorder and the sensitivity of adjacent band energy mutation, the mixed energy entropy is obtained, which is expressed as:

[0126]

[0127] Where, Indicates the The mixed energy entropy of samples characterizes the complexity of the signal energy distribution; the mixed energy entropy is calculated based on the relative energy values ​​of different sub-bands, so the mixed energy entropy is a multi-scale mixed energy entropy; is the correlation factor, which adjusts the correlation strength between frequency bands, e.g. ; Indicates the The first sample relative energy of the sub-bands; For logarithmic functions, the default base is 10.

[0128] It should be noted that The term is the calculation method of Shannon entropy, which characterizes the disorder of energy distribution and quantifies the disorder of energy distribution.

[0129] It should also be noted that The term represents the penalty term for energy jumps between frequency bands, which enhances the sensitivity to energy mutations in adjacent frequency bands.

[0130] The average mixed energy entropy of all samples is calculated to obtain the average mixed energy entropy, which is expressed as follows:

[0131]

[0132] in, represents the average mixing energy entropy, Indicates the The mixed energy entropy of samples, is the total number of samples.

[0133] Step S6: Construct a feature enhancement strategy based on an adaptive Gaussian kernel to strengthen the response strength of the clustering area in the feature space.

[0134] Fault features are locally clustered in the feature space, and conventional enhancement methods are difficult to strengthen the feature response of the clustered area. This paper constructs a feature enhancement strategy based on an adaptive Gaussian kernel to strengthen the response strength of the clustered area in the feature space. The specific steps are as follows:

[0135] Step S61: performing clustering processing on the relative energy entropy feature vectors of all samples to obtain the fault feature cluster center vector.

[0136] The relative energy entropy feature vectors of all samples in the training set are clustered using the K-means clustering algorithm to obtain the fault feature cluster center vector representing the fault feature cluster center, which is expressed as:

[0137]

[0138] Where, Represents the fault feature cluster center vector, which is a vector representing One of the internal cluster center vectors; Indicates the The relative energy entropy feature vector of the samples; Represents the K-means clustering algorithm, which calculates cluster centers based on the distribution of eigenvectors.

[0139] No. The relative energy entropy feature vector of the samples , is through the The time window size is The numerical set of mixed energy entropy obtained within is expressed as follows:

[0140] F in n =[ H mix n- N cws 2 ,…, H mix n ,…, H mix n+ N cws 2 ]

[0141] in, Indicates the The mixing energy entropy of samples; Indicates the The mixing energy entropy of the samples.

[0142] For example, when When it is equal to 5, it is expressed as:

[0143] F in n =[ H mix n-2 , H mix n-1 , H mix n , H mix n+1 , H mix n+2 ]

[0144] Step S62: Calculate the intra-cluster standard deviation based on the fault feature cluster center vector and the sample set belonging to the cluster.

[0145] Based on the clustering results and the sample set belonging to the cluster, the intra-class standard deviation of the degree of fault feature clustering is obtained by calculating the square root of the average square of the Euclidean distance between the relative energy entropy feature vectors of these samples and the fault feature cluster center vector, which is expressed as:

[0146]

[0147] Where, represents the within-class standard deviation; Represents the sample set belonging to the current cluster; For the The number of samples in each cluster; Represents the L2 norm, which is the same as the Euclidean norm.

[0148] Step S63: constructing a Gaussian kernel function based on the obtained cluster center vector and the intra-cluster standard deviation.

[0149] Based on the obtained cluster center vector and the intra-class standard deviation, the Gaussian kernel function output vector is calculated to achieve the effect of forming a Gaussian potential well distribution in the feature space with the cluster center as the peak and the width controlled by the intra-class standard deviation. The Gaussian kernel function is expressed as:

[0150]

[0151] Where, Indicates the The Gaussian kernel function output vector of samples; Indicates the The relative energy entropy feature vector of the samples; represents the natural exponential function; the numerator Represents the square of the distance from the relative energy entropy feature vector to the cluster center; the denominator Controls the width of the Gaussian kernel, The larger it is, the flatter the potential well is.

[0152] It should be noted that the output vector of the Gaussian kernel function strengthens the response intensity of the fault feature cluster area through the Gaussian potential well. The Gaussian potential well is an adaptive amplifier based on the probability distribution of the feature space. It is constructed through the Gaussian kernel function, with the fault cluster center as the gravitational core and the intra-class standard deviation as the range controller. The balance between feature enhancement and noise suppression is achieved through exponential decay.

[0153] Step S64: multiply the relative energy entropy feature vector by the corresponding Gaussian kernel function output vector element by element to obtain an enhanced feature vector.

[0154] By multiplying the relative energy entropy feature vector by the corresponding Gaussian kernel function output vector element by element, the enhanced feature vector is obtained, thereby enhancing the fault feature response intensity close to the cluster center, which is expressed as:

[0155]

[0156] Where, Indicates the The enhanced feature vector of samples.

[0157] Step S7: Constructing a fault data recognition neural network.

[0158] Step S71: Define the neural network structure. The neural network consists of a fully connected deep neural network consisting of an input layer, three hidden layers and an output layer.

[0159] The number of neurons in the input layer is strictly aligned with the dimension of the enhanced feature vector;

[0160] The first hidden layer is designed to be wide, with the number of neurons approximately equal to 4 times the input dimension. The number of neurons in the subsequent two hidden layers is gradually reduced, with a reduction ratio of 0.7.

[0161] The number of neurons in the output layer is equal to the total number of fault categories, and the Softmax activation function is used to output the category probability distribution.

[0162] A dense connection strategy is adopted between layers, and the output of each hidden layer serves as the full input of the next layer to ensure the integrity of fault features across layers.

[0163] Step S72: Initialize the weight matrix based on the enhanced feature vector so that the initial weight points to the direction of fault feature difference.

[0164] Traditional random initialization causes the initial weight direction to be inconsistent with the direction of fault feature difference, resulting in low training efficiency.

[0165] The present invention initializes the weight matrix based on the statistical characteristics of hybrid energy entropy, so that the initial weight points to the direction of fault feature difference. The specific steps are as follows:

[0166] Step S721: Based on the enhanced eigenvector, calculate the covariance matrix of the enhanced eigenvector.

[0167] Based on the enhanced feature vectors of all sampling points in the training set, the covariance matrix representing the statistical correlation between the enhanced feature vectors is obtained by calculating the outer product of the enhanced feature vectors and averaging them, which is expressed as:

[0168]

[0169] Where, represents the covariance matrix of the enhanced eigenvector; Indicates the The enhanced feature vector of samples; express The transpose of .

[0170] Step S722: performing a singular value decomposition operation on the covariance matrix of the enhanced eigenvector and multiplying it by the scaling factor to obtain a basic weight matrix.

[0171] By performing a singular value decomposition operation on the covariance matrix of the enhanced eigenvector, taking the left singular vector corresponding to the maximum singular value and multiplying it with the scaling factor, we get the basic weight matrix, which is expressed as:

[0172]

[0173] Where, represents the basic weight matrix; The dimension is ; Represents the singular value decomposition operation, taking the left singular vector corresponding to the maximum singular value; is the scaling factor that controls the weight amplitude range, such as, .

[0174] Step S723: Calculate the difference between the enhanced feature mean vector of each fault category sample and the global enhanced feature mean vector, and perform normalization constraints based on the Frobenius norm to obtain a fault-sensitive correction term matrix.

[0175] By calculating the difference between the enhanced feature mean vector of each fault category sample and the global enhanced feature mean vector and performing normalization constraints based on the Frobenius norm, the fault sensitive correction term matrix is ​​obtained, which is expressed as:

[0176]

[0177] Where, represents the fault-sensitive correction term matrix, The dimension is ; is the total number of fault categories; is a positive integer; Indicates the The number of samples of class fault; Indicates the Feature mean vector of class fault samples; represents the global feature mean vector; represents the Frobenius norm, which is used to normalize the amplitude of the correction term; for The transpose of for The transpose of .

[0178] Step S724: Add the basic weight matrix and the fault-sensitive correction term matrix to obtain the initial weight matrix of the neural network.

[0179] The initial weight matrix of the neural network is obtained by adding the basic weight matrix to the fault sensitive correction term matrix, so that the initial weight contains the principal component direction and fault discrimination direction information, which is expressed as:

[0180]

[0181] Where, Represents the initial weight matrix of the neural network.

[0182] Step S73: Define a dual-threshold activation function of the neural network.

[0183] The response amplitude of weak fault features is low, and the conventional ReLU activation function causes gradient truncation in the near-zero region, resulting in feature loss.

[0184] The present invention adopts a dual-threshold activation function to protect weak fault characteristics and suppress strong noise by setting two high and low thresholds. The dual-threshold activation function is expressed as follows:

[0185]

[0186] Where, Represents the input value of the activation function; is a dual threshold activation function; Represents the absolute value of the input value of the activation function; is a low threshold; is the slope attenuation coefficient in the low amplitude region; is the high threshold; is the slope attenuation coefficient in the high amplitude region; is a symbolic function, Characterization Time output , Time output .

[0187] When the absolute value of the activation function input value is lower than the set low threshold, the input value is multiplied by a low amplitude area slope attenuation coefficient less than 1 for output, thereby retaining the weak fault feature gradient to a certain extent while suppressing the noise in the area. Protect weak fault characteristics from being filtered out, such as ; Slope attenuation coefficient in low amplitude area Suppress noise while preserving weak features, e.g. , characterizes when the input value is in the interval When , the output is the input value times, retain to avoid feature disappearance.

[0188] When the absolute value of the activation function input value is between the set low threshold and high threshold, the input value is directly used as the output value to achieve the transmission of the main fault characteristics without attenuation. Suppress strong noise interference, such as , characterizes when the input value is in the interval [ θ low , θ high ] or [ - θ high , - θ low ] When , the output is equal to the input value, ensuring that the main fault characteristics are transmitted without attenuation.

[0189] When the absolute value of the activation function input value exceeds the set high threshold, the difference between the input value and the product of its sign function and the high threshold is subtracted, and the output value is obtained based on the high amplitude area slope attenuation coefficient to greatly compress the strong noise response intensity and limit the output value to the vicinity of the high threshold. The high amplitude area slope attenuation coefficient Used to significantly compress the response strength of strong noise, such as ;

[0190] It should be noted that the high amplitude region saturation suppression operation compresses the input value toward zero, for example, when When the output value is limited to Avoid strong noise interference.

[0191] Step S74: Define the hybrid energy entropy weighted loss of the neural network.

[0192] The feature impurities of different fault categories vary significantly, and traditional loss treats all samples equally.

[0193] The present invention constructs a weighted loss based on mixed energy entropy, and improves the training weight of difficult samples by weighting the entropy value. The specific steps are as follows:

[0194] Step S741: Calculate the average mixed energy entropy of all fault categories.

[0195] By calculating the average mixed energy entropy of each type of fault sample in the training set, and then averaging the average mixed energy entropy of all fault categories, the average mixed energy entropy of all fault categories is obtained, which is expressed as:

[0196]

[0197] Where, represents the average mixed energy entropy of all fault categories; is the total number of fault categories; Indicates the The average mixed energy entropy of the fault-like samples; Is a positive integer.

[0198] It should be noted that the Average mixed energy entropy of fault-like samples According to the calculation method of average mixed energy entropy, the All samples of this type of fault are extracted separately, decomposed by wavelet packets, and the mixed energy entropy is calculated. Then the average is calculated to obtain the first The average mixed energy entropy of class fault samples.

[0199] Step S742: Calculate a fault category weighting factor based on the average mixed energy entropy of a certain type of fault samples and the average mixed energy entropy of all fault categories.

[0200] Based on the average mixed energy entropy of a certain type of fault samples and the average mixed energy entropy of all fault categories, the weighting factor of this type of fault is obtained, giving a greater weight to the fault category with higher mixed energy entropy, which is expressed as:

[0201]

[0202] Where, Indicates the Weighting factor for class faults; Indicates the The average mixed energy entropy of class fault samples.

[0203] It should be noted that the higher the average mixed energy entropy of each category of fault samples, the more mixed the fault characteristics of this category are, and the more difficult it is to distinguish samples. The weighting factor , to increase its training weight.

[0204] Step S743: Construct a hybrid energy entropy weighted loss.

[0205] On the basis of the standard cross entropy loss, the final mixed energy entropy weighted loss is formed by multiplying the loss term of each category by the weighting factor of the corresponding category, and then focusing on optimizing the difficult samples, which is expressed as:

[0206]

[0207] Where, represents the mixed energy entropy weighted loss; Indicates that the sample belongs to The true label of the fault class, in one-hot encoding format, with a value of 0 or 1; Indicates that the neural network predicts that the sample belongs to The probability of class failure. is the total number of fault categories.

[0208] In one embodiment, Figure 3As shown in the figure, a heat map is used to evaluate the fault differentiation ability of the hybrid energy entropy feature. The experiment compares the separability performance of traditional feature extraction methods such as wavelet entropy, spectral kurtosis, and envelope spectral entropy with the hybrid energy entropy of the present invention on six typical fault types. The heat map intuitively displays the differentiation ability of different methods for various faults with color depth, and annotates the specific separability index in each cell. The experimental results show that the hybrid energy entropy of the present invention exhibits the highest separability for all fault types, especially in complex fault modes such as gear pitting and bearing wear. The advantage is most significant, indicating that the innovative inter-band energy jump penalty term design of the hybrid energy entropy enhances the sensitivity to impact-type faults.

[0209] In this embodiment, the universality and effectiveness of the hybrid energy entropy feature extraction technology are verified, and a grouped histogram is also used, such as Figure 4 As shown, the horizontal axis shows five typical mechanical fault types, including bearing wear, gear tooth breakage, shaft misalignment, rotor imbalance, and base loosening, and the vertical axis represents the recognition accuracy. The traditional wavelet entropy, spectral entropy, and energy statistics methods are compared with the hybrid energy entropy technology proposed in the present invention. Each group of columns clearly shows that the method of the present invention (the rightmost column) achieves the highest accuracy in all fault types, and its advantage is more prominent in fault types with obvious frequency band energy mutations such as gear tooth breakage and shaft misalignment, indicating that the unique inter-band energy jump penalty term design in the hybrid energy entropy enhances the sensitivity to impact-type faults and solves the defect of traditional methods that ignore the energy correlation between frequency bands.

[0210] Step S8: Train the neural network to obtain a trained neural network.

[0211] Step S81: forward propagation.

[0212] The enhanced feature vector is input into the neural network, and the linear transformation of the weight matrix and the nonlinear mapping of the activation function are performed in sequence. The first layer receives the enhanced feature vector, and after weighted summation, it is output by the double threshold activation function. Each subsequent layer uses the output of the previous layer as input, repeats the linear transformation and activation process, and the final layer outputs the unnormalized probability value of each category, which is converted into a probability distribution by the Softmax function.

[0213] Step S82: Back propagation.

[0214] Based on the hybrid energy entropy weighted loss output, the gradient is calculated backward along the network layers.

[0215] Starting from the output layer, we first calculate the partial derivative of the loss with respect to the predicted probability. Then, we derive the gradient components of the loss with respect to the weights and biases of each layer layer by layer according to the chain rule. In the calculation of the hidden layer gradient, we combine the piecewise derivative characteristics of the dual-threshold activation function to ensure that the gradient backpropagation strictly matches the activation behavior. Finally, we obtain the gradient tensor of all weight parameters, which represents the parameter optimization direction driven by the current batch data.

[0216] Step S83: Gradient orientation correction.

[0217] The gradient interference of noise samples causes the weight update to deviate from the fault-sensitive direction, reducing the diagnostic robustness.

[0218] The present invention proposes a gradient correction mechanism based on intra-class distribution density and fault sensitivity matrix. The specific steps are as follows:

[0219] Step S831: Calculate the gradient confidence weight.

[0220] Based on the clustering results of feature enhancement, the sample gradient confidence is defined to reflect the degree of deviation between the feature and the cluster center, providing key weight information for gradient correction, which is expressed as:

[0221]

[0222] Where, Indicates the Gradient confidence weight of each sample; Indicates the The enhanced feature vector of samples; represents the fault feature cluster center vector; represents the within-class standard deviation.

[0223] It should be noted that the Gradient confidence weight of samples The value range is (0, 1]. The closer it is to 1, the closer the sample feature is to the cluster center and the higher the gradient reliability.

[0224] Step S832: Construct a fault-sensitive projection matrix.

[0225] The fault-sensitive correction matrix is ​​used to construct a projection matrix to align the fault discrimination direction and enhance the discriminability of the gradient update, which is expressed as:

[0226]

[0227] Where, represents the fault-sensitive projection matrix; represents the fault sensitive correction term matrix; represents the Frobenius norm; To prevent division by zero constant, used to avoid the denominator being zero, such as, .

[0228] Step S833: Perform gradient orientation correction.

[0229] The original gradient is confidence-weighted and projected to correct it, and the final gradient is output. The gradient of high-confidence samples is strengthened in the fault-sensitive direction, and the gradient of low-confidence samples is scaled by the inverse of the standard deviation within the class to suppress noise interference. At the same time, the gradient vector is projected into the fault-sensitive subspace through the fault-sensitive projection matrix to further enhance the discriminative update direction and improve the fault diagnosis performance of the neural network, which is expressed as:

[0230]

[0231] Where, represents the corrected gradient matrix; represents the original gradient matrix, which is obtained by backpropagation of the mixed energy entropy weighted loss; is an adaptive damping matrix used to suppress the gradient amplitude of low-confidence samples. The calculation method is expressed as , is the time window size;

[0232] For the The standard deviation of the class of the dimension feature is a scalar, that is, is the intra-class standard deviation of the first dimension feature, For the The within-class standard deviation of the dimensional feature.

[0233] It should be noted that the Intra-class standard deviation of dimensional features It is the local discreteness calculated for a single feature dimension, quantifying the The fluctuation intensity of the dimension feature within the cluster is calculated as follows: ,in, yes No. eigenvalues ​​(which are scalars), Indicates the The relative energy entropy feature vector of the samples (is a vector), is the fault feature cluster center vector in the The eigenvalue of dimension, For the The number of samples in a cluster.

[0234] It should be noted that The term strengthens the gradient of high confidence samples in the fault-sensitive direction, The gradient of the low-confidence sample is scaled by the inverse of the standard deviation within the class to suppress noise interference.

[0235] It should also be noted that the fault-sensitive projection matrix The gradient vector is projected into the fault-sensitive subspace to enhance the discriminative update direction.

[0236] Step S84: Neural network parameters are updated.

[0237] The feature distribution evolves dynamically with training, and fixed learning rates and update strategies are difficult to adapt to non-stationary feature spaces.

[0238] The present invention adopts an adaptive update strategy based on covariance scaling and entropy weighting, adjusts the learning rate through feature stability, and improves the convergence efficiency of the fault mode. The specific steps are as follows:

[0239] Step S841: Calculate the entropy weighted learning rate.

[0240] Based on the hybrid energy entropy weighting factor, the learning rate is dynamically scaled to focus on high entropy fault categories. Increasing the learning rate accelerates the weight update of difficult-to-classify samples and improves the learning efficiency of the neural network for complex fault features. It is expressed as:

[0241]

[0242] Where, represents the entropy-weighted learning rate; is the basic learning rate, such as, ; Indicates the Weighting factor for class faults, Indicates the Weighting factor for class faults, Indicates the Weighting factor for class faults; is the decay exponent, controlling the weighted intensity, e.g. .

[0243] It should be noted that the entropy weighted learning rate When the weighting factor is large, it indicates that the category is a high entropy category, so the learning rate is increased to accelerate the weight update of difficult-to-classify samples.

[0244] Step S842: Construct a covariance scaling factor.

[0245] The covariance matrix of the enhanced eigenvector is used to construct a scaling factor to adapt the feature correlation and characterize the discreteness of the feature distribution. The more dispersed the features are, the smaller the update step size is to ensure the stability of the parameter update, which is expressed as:

[0246]

[0247] Where, represents the covariance scaling factor; is the covariance matrix of the enhanced eigenvector; represents the matrix trace operation; Represents the largest eigenvalue of a matrix.

[0248] It should be noted that the covariance scaling factor Characterizes the degree of discreteness of feature distribution. The larger the value, the more dispersed the features are, and the smaller the update step size needs to be.

[0249] Step S843: Execute adaptive momentum update.

[0250] Combining the output of gradient-directed correction, entropy-weighted learning rate, and covariance scaling factor, the momentum method is used to update the weights. The integration of entropy weighting and covariance scaling ensures the correctness of the parameter update direction and the rationality of the step size. The smoothing correction term of historical momentum reduces high-frequency oscillations, improving the convergence speed and stability of the neural network. Ultimately, adaptive parameter updates based on covariance scaling and entropy weighting are achieved, improving the convergence efficiency of the fault mode, which can be expressed as:

[0251]

[0252]

[0253] Where, Indicates the current time ( (times) iteration momentum vector; Indicates the last time ( (times) iteration momentum vector; is the momentum decay factor, such as, ; represents a symbolic function; is the oscillation suppression coefficient, such as, ; is the momentum smoothing index, such as, ; Indicates the current time ( The weight matrix of the neural network for each iteration; Indicates the last time ( The weight matrix of the neural network for each iteration.

[0254] It should be noted that Term integration entropy weighting and covariance scaling; The term adds a smoothing correction to the historical momentum to reduce high frequency oscillations.

[0255] Step S85: Stop iterative condition judgment.

[0256] Set up a dual convergence monitoring mechanism:

[0257] The first condition is that the accuracy of the validation set does not fluctuate by more than 1% over 10 consecutive training cycles;

[0258] The secondary condition is that the loss decrease rate is continuously lower than 1‰ and the number of training rounds exceeds the minimum iteration threshold, such as 100 rounds.

[0259] When any of the conditions are met, the training is terminated immediately and the early stopping callback is started to save the optimal weights.

[0260] In one embodiment, the model stability of the gradient directional correction mechanism under noise interference is analyzed, such as Figure 5 As shown, the horizontal axis sets the nine-level noise sample ratio from 0% to 40%, and the vertical axis measures the test accuracy. The four optimization strategies of random gradient descent, adaptive moment estimation, root mean square propagation and gradient directional correction of the present invention are compared. The experimental results show that as the proportion of noise samples increases, the accuracy of all methods decreases, but the method of the present invention always maintains the highest level with the smallest decline, indicating that the gradient confidence weight based on the Gaussian kernel screens reliable samples, the fault-sensitive projection matrix aligns the discrimination direction, and the adaptive damping matrix suppresses the interference of low-confidence samples, all of which have the effect of improving the robustness of the model. The present invention effectively solves the core problem of parameter update deviation caused by noise samples through the gradient directional correction mechanism, thereby improving the reliability of the fault diagnosis system in actual industrial environments.

[0261] Step S9: Use the trained neural network to identify mechanical equipment fault data, and finally output a structured diagnosis report containing fault type, confidence level, and uncertainty flag to support equipment operation and maintenance decision-making.

[0262] The newly collected vibration data of the device to be diagnosed is processed in sequence through the complete processing flow from step S1 to step S6, and the enhanced feature vector is obtained and then input into the trained neural network.

[0263] The forward propagation outputs the probability distribution of each category, and the category corresponding to the highest probability is selected as the preliminary diagnosis result.

[0264] At the same time, the difference between the second highest probability and the highest probability is calculated. When the difference is lower than the set threshold, such as 0.3, the uncertainty warning is activated, requiring additional multi-sensor data joint analysis.

[0265] The final output includes a structured diagnostic report with fault type, confidence level, and uncertainty flags to support equipment operation and maintenance decision-making.

[0266] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying mechanical equipment fault data based on artificial intelligence, characterized in that: The method comprises the following steps: Step S1: collecting vibration data of mechanical equipment, storing the collected vibration data in floating point format and retaining complete time domain amplitude information and timestamp sequence; Step S2: Denoising the mechanical equipment vibration data to obtain denoised vibration data, removing the sensor's own electronic noise, environmental interference, and random fluctuations; Step S3: labeling the denoised vibration data to obtain labeled vibration data, and labeling the denoised vibration data as a normal state or a corresponding fault type; Step S4: Adopting an adaptive normalization method based on energy density to adaptively normalize the marked vibration data to obtain normalized vibration data; Step S5: enhancing the sensitivity to shock-type faults by combining the inter-band energy jump penalty term, and then calculating the mixed energy entropy; Step S6: constructing a feature enhancement strategy based on an adaptive Gaussian kernel to strengthen the response strength of the clustering area in the feature space; Step S7: constructing a fault data recognition neural network; Step S8: training the neural network to obtain a trained neural network; Step S9: using the trained neural network to identify mechanical equipment fault data; Step S5 includes the following steps: Step S51: performing wavelet packet decomposition on the normalized vibration data to obtain wavelet packet coefficient sequences of multiple sub-bands; Step S52: Based on the wavelet packet coefficient sequence of each sub-band, the square sum of the absolute values ​​of all coefficients in the wavelet packet coefficient sequence is calculated to obtain the sub-band energy value; Step S53: Based on the energy values ​​of all sub-bands, the energy value of each sub-band is divided by the sum of the energy values ​​of all sub-bands to obtain a sub-band relative energy value; Step S54: Based on the relative energy value of each sub-band, the Shannon entropy is calculated and the inter-band energy jump penalty term is constructed to obtain the mixed energy entropy; Step S6 includes the following steps: Step S61: performing clustering processing on the relative energy entropy feature vectors of all samples to obtain the fault feature cluster center vector; Step S62: Calculate the intra-cluster standard deviation based on the fault feature cluster center vector and the sample set belonging to the cluster; Step S63: constructing a Gaussian kernel function based on the obtained cluster center vector and the intra-cluster standard deviation; Step S64: multiply the relative energy entropy feature vector by the corresponding Gaussian kernel function output vector element by element to obtain an enhanced feature vector.

2. The method for identifying mechanical equipment fault data based on artificial intelligence according to claim 1, characterized in that: The adaptive normalization method based on energy density in step S4 is expressed as: in, Indicates the normalized vibration data at the The first sample The value of each sampling point; Indicates the vibration data after annotation. Sample No. The value of each sampling point; Indicates the vibration data after annotation. Sample No. The local window mean of the samples at each sampling point; The vibration data after marking is Sample No. The local standard deviation of the sampling points; To prevent division by zero constant; Indicates the The global energy of the samples; represents the local window energy.

3. The method for identifying mechanical equipment fault data based on artificial intelligence according to claim 2, characterized in that: The mixed energy entropy in step S54 is expressed as: Where, Indicates the The mixing energy entropy of samples; is the correlation factor; Indicates the The first sample relative energy of the sub-bands; Indicates the The first sample relative energy of the sub-bands; is a logarithmic function.

4. The method for identifying mechanical equipment fault data based on artificial intelligence according to claim 1, characterized in that: Step S7 includes the following steps: Step S71: defining a neural network structure, where the neural network consists of a fully connected deep neural network consisting of an input layer, three hidden layers, and an output layer; Step S72: Initialize the weight matrix based on the enhanced feature vector so that the initial weight points to the direction of fault feature difference; Step S73: defining a dual threshold activation function of the neural network; Step S74: Define the hybrid energy entropy weighted loss of the neural network.

5. The method for identifying mechanical equipment fault data based on artificial intelligence according to claim 4, characterized in that: Step S72 includes the following steps: Step S721: Calculating the covariance matrix of the enhanced eigenvector based on the enhanced eigenvector; Step S722: performing a singular value decomposition operation on the covariance matrix of the enhanced eigenvector and multiplying it by the scaling factor to obtain a basic weight matrix; Step S723: Calculate the difference between the enhanced feature mean vector of each fault category sample and the global enhanced feature mean vector, and perform normalization constraints based on the Frobenius norm to obtain a fault sensitivity correction term matrix; Step S724: Add the basic weight matrix and the fault-sensitive correction term matrix to obtain the initial weight matrix of the neural network.

6. The method for identifying mechanical equipment fault data based on artificial intelligence according to claim 5, characterized in that: The dual threshold activation function in step S73 is expressed as follows: in, Represents the input value of the activation function; is a dual threshold activation function; Represents the absolute value of the input value of the activation function; is a low threshold; is the slope attenuation coefficient in the low amplitude region; is the high threshold; is the slope attenuation coefficient in the high amplitude region; is a symbolic function.

7. The method for identifying mechanical equipment fault data based on artificial intelligence according to claim 6, characterized in that: Step S74 includes the following steps: Step S741: Calculate the average mixed energy entropy of all fault categories; Step S742: Calculate the fault category weighting factor based on the average mixed energy entropy of the fault sample and the average mixed energy entropy of all fault categories; Step S743: Construct a hybrid energy entropy weighted loss.

8. The method for identifying mechanical equipment fault data based on artificial intelligence according to claim 7, characterized in that: The mixed energy entropy weighted loss is expressed as: in, represents the mixed energy entropy weighted loss; Indicates that the sample belongs to The true label of the fault class; Indicates that the neural network predicts that the sample belongs to probability of class failure; is the total number of fault categories.

Citation Information

Patent Citations

  • A generator fault diagnosis method based on vibration trend prediction and related products

    CN119848674B

  • Fan bearing fault early warning method and device based on vibration characteristics

    CN120008926A

  • Wavelet energy entropy detecting method for recognizing faults of ultra-high voltage direct-current transmission line

    CN102156246A

  • AI-based power distribution network fault identification method and related device

    CN119312068A