A generator set fault early warning method
By constructing an adaptive reference trajectory for the operating domain and a channel fault sensitivity coefficient, combined with differentiated enhancement and a multi-scale time-series feature network, the problem of misjudgment in fault identification of multi-channel data of generator sets was solved, and more accurate and stable fault early warning was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN DATANG BAOCHANG GAS POWER GENERATION
- Filing Date
- 2026-05-29
- Publication Date
- 2026-07-31
AI Technical Summary
Existing generator set fault early warning methods fail to effectively distinguish between normal fluctuations and fault deviations when processing multi-channel time-series data, leading to misjudgment or obscuring of real faults. Furthermore, they do not fully utilize the coupling relationship between channels, resulting in insufficient classification stability.
By constructing an adaptive reference trajectory for the operating domain and channel fault sensitivity coefficients, performing differentiated nonlinear enhancement and local impact gating, and combining a multi-scale temporal feature network with fault prototype constraints, differentiated processing and coupling relationship learning of multi-channel data are achieved.
It effectively separates normal fluctuations and fault anomalies caused by load and speed changes, improves the accuracy and stability of fault identification, highlights the discrimination characteristics, enhances the response to different fault modes, and improves the reliability of classification.
Smart Images

Figure CN122333227B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of generator set technology, specifically a generator set fault early warning method. Background Technology
[0002] As core equipment in industrial power systems, generator sets' operational reliability directly impacts production safety and economic efficiency. With the development of condition monitoring technology, an increasing number of generator sets are equipped with multi-physical channel online monitoring systems, continuously collecting multi-dimensional time-series data such as vibration velocity, temperature, current, speed, and power. This data contains rich information on equipment health status, providing a practical basis for data-driven fault early warning. However, effectively extracting fault characteristics from massive amounts of multi-source monitoring data and achieving accurate and timely fault identification remains a core technological challenge in this field. Existing methods have significant shortcomings in the following key aspects: 1. Existing methods generally use global mean or uniform normalization to process multi-channel time series data. They do not consider that the normal distribution of each physical channel of the generator set will drift under different load and speed combinations. This makes it easy for normal fluctuations under changing operating conditions to be misjudged as fault deviations, or for real faults to be masked within the range of operating condition changes.
[0003] 2. Existing methods typically apply the same intensity of preprocessing or enhancement to all physical channels without differentiating the fault response level of each channel in the current sample. This results in normal fluctuations in channels with inconspicuous fault characteristics being amplified synchronously, which not only introduces redundant interference but may also drown out the information of truly sensitive channels.
[0004] 3. Existing methods neglect the essential differences in the response patterns of different faults in the physical channels during the time-series signal preprocessing stage. Short-time spikes representing impacts and steady harmonics in the vibration channel are often treated the same, and slow trends and instantaneous noise in the temperature and current channels are not processed separately, resulting in insufficient prominence of the characteristics of early gradual faults and local impact faults.
[0005] 4. Existing methods either fail to utilize multi-channel coupling relationships in the feature extraction and classification stages, treating each channel as an independent input dimension, or fail to introduce class center constraints in the embedding space, relying solely on linear classification boundaries. This results in channel-to-channel co-operational anomalies being difficult to capture, and insufficient classification stability in regions with blurred class boundaries and for classes with few samples. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a generator set fault early warning method to solve the technical problems mentioned in the background section.
[0007] To achieve the above objectives, this invention provides the following technical solution: a generator set fault early warning method, comprising the following steps: Data from five physical channels—vibration velocity, temperature, current, speed, and power—are collected during the generator set's operation. These five physical channels are aligned on the time axis and labeled; the labels represent the generator set's fault state category. The operating domain is divided according to the speed channel and the power channel. A normal reference trajectory and a normal robust scale are constructed for each operating domain. The channel fault sensitivity coefficient of the current original generator set multivariate time series sample relative to the current operating domain reference trajectory is calculated. Differential nonlinear enhancement is performed on the five physical channels. Local impact gating is performed on the vibration velocity channel. Trend enhancement is performed on the temperature channel and the current channel to obtain the processed generator set multivariate time series matrix. Construct and train a fault diagnosis model, which includes a channel-gated multi-scale temporal feature network and fault prototype constraints. The target data of the device to be diagnosed is input into the fault diagnosis model for fault diagnosis, and the fault diagnosis results are output.
[0008] Furthermore, the labels include normal, unbalanced, misaligned, bearing fault, winding fault, and cooling system fault.
[0009] Furthermore, the operating domain is divided according to the speed channel and power channel, and a normal reference trajectory and a normal robust scale are constructed for each operating domain, specifically as follows: Extract the speed channel and power channel from the normal category samples in the training set, calculate the average speed channel and the average power channel for each normal category sample, and obtain the running domain description vector; Clustering is performed on the runtime description vectors of all normal category samples to obtain multiple runtimes; For the For each operating domain, all normal category samples belonging to that operating domain are collected. Within these normal category samples, the normal reference median is calculated for each time step and each physical channel, yielding the [missing value]. The runtime reference trajectory matrix of each runtime domain; For the For each operating domain, the normal robust scale is calculated for each physical channel in all normal category samples belonging to that operating domain, resulting in the normal robust scale vector of the operating domain. For the original generator set multivariate time series sample to be processed, calculate its average speed channel value and average power channel value to obtain the current operating domain description vector. Compare the distance between the current operating domain description vector and the center of each operating domain, and select the operating domain with the closest distance as the current operating domain.
[0010] Furthermore, the channel fault sensitivity coefficient of the current original generator set multivariate time-series sample relative to the current operating domain reference trajectory is calculated, specifically as follows: Read the current running domain index The corresponding runtime reference trajectory matrix and runtime normal robust scaling vector; Multivariate time-series samples of the original generator set The standardized deviation is calculated for each time step and each physical channel in the process; The absolute value of the standardized deviation for each physical channel is taken to obtain the channel deviation sequence; In the channel deviation sequence of each physical channel, calculate the 95th percentile deviation value and the average deviation value; The channel anomaly response value is calculated based on the 95th percentile deviation value and the average deviation value; The abnormal response values of the five physical channels are normalized into channel fault sensitivity coefficients.
[0011] Furthermore, differentiated nonlinear enhancements are performed on the five physical channels, specifically as follows: The obtained standardized bias is used as the input for nonlinear enhancement. Based on the channel fault sensitivity coefficient, a hyperbolic tangent nonlinear mapping is performed on each physical channel to obtain the multivariate time sequence matrix of the enhanced generator set. The positive and negative directions of the enhancement values are retained. When the enhancement value is positive, it indicates that the corresponding physical channel is higher than the normal reference trajectory of the current operating domain; when the enhancement value is negative, it indicates that the corresponding physical channel is lower than the normal reference trajectory of the current operating domain.
[0012] Furthermore, differentiated nonlinear enhancements are performed on the five physical channels, specifically as follows: Extract the vibration velocity channel enhancement sequence from the multivariate time-series matrix of the enhanced generator set; Mirror padding is applied to both ends of the vibration velocity channel enhancement sequence; Establish a local sliding window centered on each time step; Converting local excess kurtosis into local impact gating values, the first... The local impact gating value at each time step is denoted as ; By using local impact gating values to modulate the vibration velocity channel enhancement sequence step by step, an impact-fidelity vibration sequence is obtained.
[0013] Furthermore, the enhanced execution trend in both the temperature and current channels is specifically as follows: Extract temperature channel enhancement sequences and current channel enhancement sequences from the multivariate time-series matrix of the enhanced generator set; A positive exponential moving average was applied to both the temperature channel and the current channel to obtain a positive trend sequence. By performing inverse exponential moving averages on the temperature channel and the current channel respectively, an inverse trend sequence is obtained; The two-way trend sequence is obtained by averaging the positive trend sequence and the negative trend sequence. Calculate the residual components of the temperature channel and the current channel; The trend enhancement sequence is obtained by weighting and combining the two-way trend components and the residual components. The five physical channels were recombined to obtain the processed multi-element timing matrix of the generator set.
[0014] Furthermore, the construction and training of the fault diagnosis model specifically involves: A local coupling matrix is constructed near each time step, and channel gating values are generated from the local coupling matrix to enhance physical channels related to fault coupling. Multi-scale temporal convolution is used to extract fault features at different time scales, so as to automatically cover various fault temporal forms from short-term impacts and periodic fluctuations to long-term trends with a unified network structure. Attention aggregation and segmented statistical augmentation are employed to obtain multi-granular temporal representation vectors; The multi-granular temporal representation vector is projected into a compact embedding vector, and a category prototype vector is set for each fault category, so that the final classification considers both the linear classification result and the prototype distance in the embedding space. The model is trained using class-weighted cross-entropy loss, same-class compaction loss, and prototype interval loss. Class-weighted cross-entropy loss is used to address the imbalance of the number of fault class samples, same-class compaction loss is used to make the embedding vectors of the same class samples close to the corresponding class prototype vectors, and prototype interval loss is used to keep the prototype vectors of different classes at the minimum distance.
[0015] Furthermore, a local coupling matrix is constructed near each time step, and channel gating values are generated from the local coupling matrix to enhance physical channels related to fault coupling, specifically: Five-channel input vectors are extracted from the processed generator set multivariate time-series matrix according to time steps, where the first... The five-channel input vector at each time step is denoted as... ; With the first A coupled statistics window is established centered on a time step, and the half-width of the coupled statistics window is denoted as . The example value is 5, which corresponds to a window length of 11 time steps. For positions at both ends of the sequence that are less than 11 time steps, a mirror padding method is used to fill in the gaps. Calculate the local coupling matrix within the coupling statistics window; The local coupling matrix is straightened into a 25-dimensional local coupling vector in a fixed order; Convert the channel gate vector into a residual gate vector, where the residual gate vector = 0.5 + the channel gate vector; The residual gating vector is multiplied element-wise with the five-channel input vector to obtain the gated five-channel input vector.
[0016] Furthermore, attention aggregation and segmented statistical enhancement are employed to obtain multi-granularity temporal representation vectors, specifically: Set trainable query vectors ; Temporal hidden feature matrix At each time step, an attention score is calculated, which is obtained by combining the trainable query vector with the temporal hidden feature vector. The inner product is calculated; Softmax normalization is performed on the attention scores at 1000 time steps to obtain 1000 attention weights; The global attention aggregation vector is obtained by weighting and summing the 1000 temporal hidden feature vectors using attention weights. The 1000 time steps are equally divided into 5 consecutive segments, each containing 200 time steps; The segment mean vector and segment standard deviation vector of 5 consecutive segments are concatenated in a fixed order to obtain the segmented statistical vector; By concatenating the global attention aggregation vector with the segmented statistical vector, a multi-granularity temporal representation vector is obtained.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention proposes an adaptive reference construction method for the operating domain. Based on two physical channels, speed and power, normal samples are clustered. A reference trajectory matrix and robust scale vector are independently established for each typical operating condition region. This ensures that subsequent deviation calculations always use the normal baseline corresponding to the current operating condition as a reference. This separates normal fluctuations caused by load and speed changes from fault anomalies at the source, avoiding misjudgments caused by changing operating conditions.
[0018] 2. This invention adopts a differentiated enhancement mechanism driven by channel fault sensitivity. By calculating the standardized deviation of each channel relative to the reference trajectory of the operating domain, and combining its high quantile deviation and average deviation, the abnormal response value of the channel is obtained. Then, different enhancement coefficients are assigned to each channel to control the strength of subsequent nonlinear mapping. This allows channels with strong fault responses to be significantly enhanced, while irrelevant channels remain smooth, effectively avoiding the synchronous amplification of normal fluctuations in irrelevant channels and interference with identification.
[0019] 3. In view of the differences in fault response modes of different physical channels, this invention introduces two dedicated processing modules: local impact gating and bidirectional trend enhancement. The vibration velocity channel uses local excess kurtosis to construct a gating value, adaptively retaining the impact characteristics and attenuating the stable harmonic components. The temperature and current channels use phase-lag-free trend extraction and weighted recombination of positive and negative exponential moving averages to differentiate and enhance the slow drift trend and short-time residual, making the discrimination features more prominent.
[0020] 4. This invention constructs a coupled channel-gated multi-scale temporal feature network and a fault prototype constraint classification framework. It dynamically learns the synergistic relationship between the five channels and generates channel gating values through the local coupling matrix inside the network. At the same time, it uses multi-scale convolutional branches to cover different fault temporal forms from short-term impact to long-term trend. Finally, it introduces learnable class prototype vectors into the embedding space and constrains the embedding distribution with similar compact loss and prototype spacing loss, so that the class boundary is more stable and the classification of early faults with fewer samples is more reliable. Attached Figure Description
[0021] Figure 1 This is a flowchart of the present invention; Figure 2 It is a raw multi-dimensional time-series sample collected from the generator set under bearing failure conditions, containing the variation curves of five physical channels—vibration velocity, temperature, current, speed, and power—over 1000 time steps. Figure 3 This is a localized impact gating process, showing the complete process of localized impact gating executed by the vibration velocity enhancement sequence; Figure 4 It is a decomposition diagram showing the bidirectional trend enhancement of the temperature channel; Figure 5 This is a distribution diagram of the aggregation weights for fault attention. Detailed Implementation
[0022] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined in this application.
[0023] like Figure 1 As shown, this invention proposes a generator set fault early warning method, the main contents of which are as follows: S1. Collect data from five physical channels during generator operation: vibration velocity, temperature, current, speed, and power. Align the five physical channels on the time axis and label them. The labels represent the generator fault state categories. Specifically, this invention collects data from five physical channels—vibration velocity, temperature, current, speed, and power—during generator operation, and divides the continuous operating data into fixed-length time-series samples. Since different fault types exhibit different response patterns on different physical channels—for example, imbalance and misalignment faults are typically related to changes in vibration velocity and speed; bearing faults usually manifest as localized impacts in the vibration velocity channel; winding faults typically manifest as abnormal current and temperature increases; and cooling system faults typically manifest as slow temperature drift and slight power changes—it is necessary to ensure that the five physical channels are aligned on the time axis before training and to establish accurate fault labels for each sample. The specific steps are as follows: 1> Vibration velocity sensors, temperature sensors, current sensors, speed sensors, and power acquisition modules are installed at key monitoring locations of the generator set to collect continuous monitoring data for the vibration velocity channel, temperature channel, current channel, speed channel, and power channel, respectively.
[0024] Furthermore, when the original sampling frequencies of the five physical channels are inconsistent, the data from each physical channel are uniformly resampled to the same diagnostic time axis. For the vibration velocity channel with a higher sampling frequency, the root mean square value, peak value, or average vibration velocity value can be calculated within each diagnostic time step; for the temperature channel, current channel, speed channel, and power channel, linear interpolation, zero-order hold, or synchronous buffering methods can be used to align them to the same diagnostic time axis. After this processing, each of the five physical channels has one value at each time step.
[0025] 2> Extract raw multi-dimensional time-series samples of the generator sets from the continuous monitoring data using a sliding window method. These raw multi-dimensional time-series samples of the generator sets are denoted as... , The size is ,in, This indicates the number of time steps contained in a single sample, with a value of 1000. This represents the total number of physical channels, with a value of 5. This represents the physical channel index, with a value range of 0 to 4, corresponding to the vibration velocity channel, temperature channel, current channel, rotational speed channel, and power channel, respectively. This represents the time step index, with a value range from 1 to 1000; Represents the original generator set multivariate time series samples The Middle The physical channel in the first The value at each time step.
[0026] In one implementation, for the first... The physical channel in the first The value at each time step, when (Vibration velocity channel) The unit is millimeters per second (mm / s). (Temperature channel) The unit is degrees Celsius (°C). (Current path) The unit is ampere (A). (In the speed channel) unit, the unit is revolutions per minute (r / min). (Power channel) The unit is kilowatt (kW).
[0027] In one implementation, the sliding window length is 1000 time steps, and the sliding step size can be 100 or 200 time steps. The smaller the sliding step size, the more adjacent samples overlap, which can increase the number of training samples; the larger the sliding step size, the stronger the independence of adjacent samples, which can reduce duplicate samples.
[0028] 3. Based on maintenance records, manual inspection records, shutdown inspection results, fault simulation test records, and operation logs, label the fault category for each original generator set multivariate time-series sample. The total number of fault categories is denoted as [missing information]. , Take 6, and denote the category index as . The value ranges from 0 to 5, corresponding to normal, unbalanced, misaligned, bearing fault, winding fault, and cooling system fault, respectively.
[0029] Among them, the normal category corresponds to samples of generator sets with no obvious abnormalities and all 5 physical channels are in stable operating range; the unbalanced category corresponds to samples with enhanced periodic vibration caused by uneven rotor mass distribution; the misalignment category corresponds to samples with abnormal vibration and speed coupling caused by coupling or transmission shaft deviation; the bearing failure category corresponds to samples with impact vibration caused by local bearing damage, rolling element impact, or abnormal lubrication; the winding failure category corresponds to samples with abnormal current, local heating, or abnormal electrical insulation; and the cooling system failure category corresponds to samples with decreased cooling efficiency, slow temperature rise, or slow temperature recovery rate.
[0030] 4> Combine all original generator set multivariate time series samples into a generator set multivariate time series training sample set, denoted as . , which includes Sample, This represents the number of training samples. The data portion of the generator set's multivariate time-series training sample set can be represented as a dataset of size [size missing]. The three-dimensional tensor, the label part can be represented as a length of The category label vector.
[0031] 5. Divide the generator set's multivariate time-series training sample set into training, validation, and test sets. The training set is used to optimize model parameters, the validation set is used to select hyperparameters and save the optimal model, and the test set is used for final evaluation. Normal class samples are also used to subsequently construct the operating domain reference trajectory and robust scale.
[0032] In one embodiment, such as Figure 2 As shown, the original 5-channel multivariate time series sample (including bearing failure impact) is analyzed, and an original multivariate time series sample collected by the generator set under bearing failure condition is displayed, which includes the change curves of five physical channels, namely vibration velocity, temperature, current, speed and power, within 1000 time steps. Figure 2 The vibration velocity channel showed obvious short-term spikes at multiple locations, which are local impact characteristics caused by bearing failure. The multi-physical channel data can synchronously record different response modes of the fault. The horizontal axis is the time step (unit: step), and the vertical axis is the vibration velocity (unit: mm / s), temperature (unit: ℃), current (unit: A), rotational speed (unit: r / min), and power (unit: kW).
[0033] S2. Divide the operating domain according to the speed channel and power channel, construct a normal reference trajectory and normal robust scale for each operating domain, calculate the channel fault sensitivity coefficient of the current original generator set multivariate time series sample relative to the current operating domain reference trajectory, perform differentiated nonlinear enhancement on the 5 physical channels, perform local impact gating on the vibration velocity channel, and perform trend enhancement on the temperature channel and current channel to obtain the processed generator set multivariate time series matrix. Specifically, under different loads, speeds, and output power, even without a fault, the five physical channels of vibration speed, temperature, current, speed, and power of a generator set will exhibit different normal operating levels. If a global average or a uniform normalization method is used directly, the changes in operating conditions will be mistakenly identified as fault deviations, or the actual fault deviations will be classified as normal fluctuations.
[0034] This invention first divides the operating domain according to the speed channel and the power channel, and constructs a normal reference trajectory and a normal robust scale for each operating domain; then, it calculates the channel fault sensitivity coefficient of the current original generator set multivariate time series sample relative to the current operating domain reference trajectory; next, it performs differentiated nonlinear enhancement on the five physical channels; finally, it performs local impact gating on the vibration velocity channel and trend enhancement on the temperature and current channels to obtain the processed generator set multivariate time series matrix. The specific steps are as follows: S201, Construction of Operating Domain Reference Trajectory and Normal Robust Scale The normal operation data of the generator set will vary with the load and speed. If only one normal reference trajectory is established, the normal samples under high load may be misjudged as abnormal, and the fault samples under low load may also be masked.
[0035] This invention constructs operating domains based on speed and power channels, enabling each operating domain to have a corresponding normal reference trajectory, thereby reducing the interference of operating condition differences on fault identification. The specific steps are as follows: 1> Extract the speed channel and power channel from the normal category samples in the training set, calculate the average value of the speed channel and the average value of the power channel for each normal category sample, and obtain the running domain description vector.
[0036] Among them, the The runtime description vector of a normal category sample is denoted as... , This is a 2D vector, where the first component represents the average value of the speed channel for this sample, and the second component represents the average value of the power channel for this sample. This represents the index of the normal category samples.
[0037] In practice, the average value of the speed channel is obtained by calculating the arithmetic mean of the speed channel values over 1000 time steps in the normal category samples; the average value of the power channel is obtained by calculating the arithmetic mean of the power channel values over 1000 time steps in the normal category samples. To avoid clustering bias caused by the different dimensions of speed and power, the average values of the speed channel and power channel for all normal category samples can be normalized to the range of 0 to 1.
[0038] 2> Perform clustering on the runtime description vectors of all normal category samples to obtain multiple runtime domains. The number of runtime domains is denoted as . For example, we can choose 4, which corresponds to the low-load low-speed operating domain, the medium-load operating domain, the high-load operating domain, and the high-speed operating domain, respectively.
[0039] In one implementation, the clustering method uses K-means clustering to apply the runtime description vector of all normal categories. Perform clustering, number of clusters This represents the number of operating domains. The algorithm iteratively updates the cluster centers by minimizing the sum of the squared Euclidean distances from each vector to its respective cluster center until convergence. Each cluster represents an operating domain, described by the center vector of that cluster.
[0040] It should be noted that the operating domain refers to a region on a two-dimensional plane defined by the average values of the speed channel and the power channel. This region represents a typical combination of operating conditions for the generator set (such as low load and low speed). Within the same operating domain, the normal fluctuation patterns of each physical channel are similar.
[0041] 3> For the first Each runtime domain collects all normal category samples belonging to that runtime domain. This represents the run domain index, with values ranging from 1 to... Then, in these normal category samples, the normal reference median is calculated for each time step and each physical channel, yielding the first normal reference median. Run-domain reference trajectory matrix of each run-domain Operating domain reference trajectory matrix The size is ,in, Indicates the first The first operating domain The physical channel in the first The normal reference median for each time step.
[0042] In one embodiment, for example, let the number of running domains be... Select the second run domain ( ); Assume there are 50 normal class samples in this domain; at time step Physical channels At the (vibration velocity channel), the 50 values are [1.2, 1.0, 1.3, 1.1, 0.9, ...]. These values are sorted from smallest to largest. Since the sample size is an even number (50), the median is the average of the 25th and 26th values after sorting. If this average is 1.1, then... mm / s; for all and Repeating this operation yields a runtime reference trajectory matrix of size 1000×5. .
[0043] It should be noted that this invention uses the median instead of the arithmetic mean to construct the normal reference trajectory because normal category samples may still contain slight sensor noise or occasional operational disturbances. The median is less sensitive to extreme values and can form a more stable normal reference trajectory.
[0044] 4> For the first For each operating domain, the normal robustness scale is calculated for each physical channel in all normal category samples belonging to that operating domain, resulting in the operating domain normal robustness scale vector. Operating domain normal robust scaling vector The dimension is 5, where Indicates the first The first operating domain Normal robust scalars for each physical channel.
[0045] In practical implementation, the first... The first normal category sample in the nth operating domain After merging the values from each physical channel, the 75th and 25th percentiles are calculated, and the difference between them is taken as the interquartile range. The interquartile range is then divided by 1.349 to obtain an approximate standard deviation, which is used as the standard deviation scale. When the interquartile range is 0 or extremely small, Set to the preset minimum scale, for example This is to avoid abnormal amplification of subsequent normalized values.
[0046] In one embodiment, for example, in the first In each operating domain, all normal category samples are collected in the physical channel. All values for the (temperature channel); assuming the 75th percentile is 60.0℃ and the 25th percentile is 55.0℃, then the interquartile range is... ℃, calculate normal robust scale ℃; if the calculated interquartile range is 0, then Set to the preset minimum scale .
[0047] 5> For the original generator set multivariate time series samples to be processed Calculate the average values of the speed channel and the power channel to obtain the current operating domain description vector; then compare the distance between the current operating domain description vector and the center of each operating domain, and select the operating domain with the closest distance as the current operating domain. The index of the current operating domain is denoted as [index missing]. , This represents the current original generator set multivariate time series sample. The corresponding runtime domain.
[0048] In one embodiment, as an example, suppose four operating domains are obtained through clustering, and their center vectors are as follows: the first center vector The second center vector The third center vector The fourth center vector For a raw sample to be processed Its normalized runtime description vector is Calculate the Euclidean distances between this vector and the four center vectors respectively, and find that... The closest distance; therefore, the current operating domain index of this sample. It was determined to be 3.
[0049] In one implementation, if the distance between the current running domain description vector and the center of the nearest running domain is less than a preset distance threshold, then the running domain reference trajectory matrix corresponding to the nearest running domain is directly used. and the normal robust scaling vector of the operating domain If the current running domain description vector is located between two running domains, the reference trajectory and robust scale of the two nearest running domains can be weighted and fused using the inverse distance weight to obtain the fused reference trajectory and fused robust scale corresponding to the current sample.
[0050] In one embodiment, for example, if the number of running domains Take 4, the current original generator set multivariate time series sample If the average value of the speed channel is close to the rated speed and the average value of the power channel is close to 80% of the rated power, then the current operating domain can be classified as the high-load operating domain. Subsequent channel deviations are no longer compared with normal samples under low load, but with the operating domain reference trajectory matrix under high load, thus avoiding misinterpreting normal current increases under high load as winding faults.
[0051] It should be noted that this invention separates the normal operating condition changes caused by speed and power from the abnormal deviations caused by faults by using the operating domain reference trajectory matrix and the operating domain normal robust scale vector, so that the subsequent channel fault sensitivity coefficient mainly reflects the strength of fault response rather than the difference in load size.
[0052] S202, Calculation of Channel Fault Sensitivity Coefficient Different faults have different response intensities in the five physical channels. Bearing faults may only cause sparse impacts in the vibration velocity channel, cooling system faults may mainly cause continuous drift in the temperature channel, and winding faults may affect both the current channel and the temperature channel at the same time. If the same enhancement intensity is applied to the five physical channels, it is easy to amplify normal fluctuations in unrelated physical channels.
[0053] This invention calculates the channel fault sensitivity coefficient for each physical channel based on the deviation of the current original generator set multivariate time-series samples from the current operating domain reference trajectory. The specific steps are as follows: 1> Read the current runtime domain index Corresponding runtime reference trajectory matrix and the normal robust scaling vector of the operating domain Among them, the running domain reference trajectory matrix Used to provide a normal time variation benchmark under the current operating domain, operating domain normal robust scaling vector Used to provide the normal fluctuation range for different physical channels.
[0054] 2> Multivariate time-series samples of the original generator set The standardized deviation is calculated for each time step and each physical channel. The standardized deviation is denoted as . The calculation method is expressed as follows: ; in, Indicates the first The physical channel in the first The standardized deviation of each time step relative to the normal reference trajectory of the current operating domain; This represents a very small constant to prevent the denominator from being zero. Examples of its values are: ; Indicates the first The first operating domain The physical channel in the first The normal reference median for each time step; Indicates the first The first operating domain Normal robust scalars for each physical channel.
[0055] It should be noted that a positive standardization deviation indicates that the current value is higher than the normal reference trajectory of the current operating domain; a negative standardization deviation indicates that the current value is lower than the normal reference trajectory of the current operating domain; the larger the absolute value of the standardization deviation, the more obvious the deviation from the normal operating state at that position.
[0056] 3> Take the absolute value of the standardized deviation for each physical channel to obtain the channel deviation sequence. Specifically, the first... The channel bias sequence of each physical channel is composed of Composition, used to describe the first The degree of abnormal deviation of each physical channel within the entire diagnostic window.
[0057] Furthermore, to reduce the impact of a single acquisition glitch on subsequent results, a standard one-dimensional signal deglitch processing method is adopted. First, a median filter of length 5 is performed on the channel deviation sequence. The median filter of length 5 means taking 5 deviation values before and after the current time step as the center, and using the median of these 5 deviation values as the filtering result of the current time step.
[0058] In practical implementation, a one-dimensional signal de-glitch processing method is adopted for time steps. The deviation value of five consecutive time steps centered on it is taken. , , , , Sort these 5 values by size and take the median of the sorted values as the time step. The filtered deviation value is then used to remove isolated spikes and glitches while preserving the step change trend of the deviation sequence.
[0059] 4> Calculate the 95th percentile deviation value and the average deviation value in the channel deviation sequence of each physical channel.
[0060] Among them, the The 95th percentile deviation of each physical channel is denoted as This is used to characterize anomalous deviations in the physical channel that are strong but not controlled by single-point extreme spikes; The average deviation of each physical channel is denoted as This is used to characterize the overall deviation of the physical channel.
[0061] In the specific implementation, the first The 1000 deviation values of each physical channel are sorted in ascending order, and the deviation value at the 95th percentile after sorting is taken as... The arithmetic mean of 1000 deviation values is used as the... Among them, the 95th percentile deviation value is more suitable for representing local anomalies such as bearing impact and misalignment fluctuations, while the average deviation value is more suitable for representing continuous anomalies such as cooling system failure and winding failure.
[0062] 5> Calculate the channel anomaly response value based on the 95th percentile deviation value and the average deviation value. The abnormal response value of each physical channel is denoted as: The calculation method is expressed as follows: for ; in, Indicates the first The response intensity is obtained by integrating local strong anomalies and overall deviations from each physical channel.
[0063] It should be noted that in the fault response of generator sets, local impact faults (such as bearing failures) will cause the deviation sequence to show a high peak value but a low average value, while continuous drift faults (such as cooling system failures) will show a medium peak value but a relatively high average value. Sensitive to local strong anomalies For generator sets that are sensitive to overall continuous deviations, using weights of 0.7 and 0.3 is merely one approach for ease of implementation. If the generator set is primarily concerned with impact faults, the weight of the 95th percentile deviation value can be increased; if the generator set is primarily concerned with thermal and electrical faults, the weight of the average deviation value can be increased.
[0064] 6> Normalize the abnormal response values of the five physical channels into channel fault sensitivity coefficients, thereby enabling differentiated enhancement based on the response intensity of each channel to faults.
[0065] In specific implementation, the first The channel fault sensitivity coefficient of each physical channel is denoted as: ,according to The calculation is performed in the following manner; where, This represents the arithmetic mean of the channel anomaly response values of the five physical channels. This represents the sensitivity amplification factor, with values ranging from 0.5 to 2.0. This represents a very small constant to prevent the denominator from being zero. Examples of its values are: .
[0066] Furthermore, to avoid excessive enhancement, the channel fault sensitivity coefficient is limited to between 1 and 6; that is, when the calculated value is... When the value is less than 1, take 1; when the calculated value is less than 1, take 1. When the value is greater than 6, take 6; the larger the channel fault sensitivity coefficient, the higher the value of the first channel fault. Each physical channel in the current original generator set multi-dimensional time-series sample The responses from China and Vietnam may include obvious aspects related to the fault.
[0067] In one embodiment, as an example, in a certain bearing failure sample, the 95th percentile deviation of the vibration velocity channel The average deviation is 5.0. If the value is 1.2, then the abnormal response value of the vibration velocity channel is... If the average value of the abnormal response values of the 5 physical channels The sensitivity amplification factor is 1.5. If we take 1.0, then the channel fault sensitivity coefficient of the vibration velocity channel is approximately This indicates that the vibration velocity channel should receive stronger enhancement in subsequent nonlinear enhancement. The superscript index 0 indicates the physical channel index.
[0068] In another embodiment, as an example, in a certain cooling system failure sample, the 95th percentile deviation of the temperature channel is 2.2, and the average deviation is 1.9. Both are relatively large, indicating that the temperature channel is not a single-point impact, but a continuous deviation from the normal reference trajectory. The channel failure sensitivity coefficient calculated based on the abnormal response value of the channel will be increased, so that the slow drift of the temperature channel will be preserved and enhanced in subsequent steps.
[0069] S203, Differential Nonlinear Enhancement Driven by Channel Fault Sensitivity The five physical channels in the original generator set multivariate time series sample have different dimensions and normal fluctuation ranges. Directly amplifying the original values can easily lead to the high-dimensional channels dominating the model input.
[0070] This invention first uses the current operating domain reference trajectory and normal robust scale to transform each physical channel into a comparable standardized deviation, and then uses the channel fault sensitivity coefficient to control the nonlinear mapping intensity to obtain the enhanced generator set multivariate time series matrix. The specific steps are as follows: 1> The standardized deviation obtained in step S202 Standardized deviation serves as a nonlinear enhancement input. The difference in dimensions between the operating domain reference trajectory and the physical channel has been eliminated, and it can directly represent the degree of deviation of the current position from the normal state of the current operating domain.
[0071] 2> Perform hyperbolic tangent nonlinear mapping on each physical channel based on the channel fault sensitivity coefficient to obtain the enhanced generator set multivariate time-series matrix. Enhance the multi-dimensional time-series matrix of generator sets The size is ,in Indicates the first The physical channel in the first The enhancement value at each time step is calculated as follows: ; in, Represents the hyperbolic tangent function; This represents the global enhancement sharpness coefficient, used to control the overall scaling of the hyperbolic tangent function input, with values ranging from 0.3 to 0.8. Indicates the first Channel fault sensitivity coefficient for each physical channel.
[0072] It should be noted that when the channel fault sensitivity coefficient of a certain physical channel is large, The nonlinear mapping is steeper near 0, and even small but persistent abnormal deviations will be amplified; when the channel fault sensitivity coefficient of a physical channel is close to 1, the nonlinear mapping is relatively gentle, and normal fluctuations will not be excessively amplified. Since the hyperbolic tangent function output range is -1 to 1, the enhancement result is less likely to experience numerical explosion.
[0073] 3. Retain the positive and negative directions of the enhancement values. Specifically, a positive enhancement value indicates that the corresponding physical channel is above the normal reference trajectory of the current operating domain; a negative enhancement value indicates that the corresponding physical channel is below the normal reference trajectory of the current operating domain. This symbol information is meaningful for distinguishing fault types. For example, a winding fault may manifest as a simultaneous increase in current and temperature, while an imbalance fault may manifest as increased vibration velocity but no significant temperature change.
[0074] In one embodiment, for example, if the standardized deviation of the vibration velocity channel at a certain time step... The channel fault sensitivity coefficient for the vibration velocity channel is 1.0. The global sharpness enhancement factor is 3.5. Taking 0.5, the enhanced input becomes 1.75, after hyperbolic tangent mapping. The value is approximately 0.94, indicating that the vibration deviation is significantly amplified; if the normalized deviation of the power channel is also 1.0, but the channel fault sensitivity coefficient of the power channel is... If the value is 1.1, the enhanced input is 0.55, which is approximately 0.50 after mapping, and the power channel maintains a relatively smooth response.
[0075] S204, Local Impact Gating and Harmonic Suppression of Vibration Velocity Channel The vibration velocity channel simultaneously contains fault impact, rotational fundamental frequency, harmonic components, and random noise. Bearing faults and mechanical shock faults usually manifest as short-term sharp waveforms, while the rotational fundamental frequency and its harmonics usually manifest as relatively stable periodic waveforms. Fixed-band filtering requires prior knowledge of rotational frequency-related frequencies and is prone to failure under varying operating conditions.
[0076] This invention utilizes local kurtosis to identify impact patterns and performs local impact gating on the vibration velocity channel. The specific steps are as follows: 1> Enhance the multi-element timing matrix of generator sets The vibration velocity channel enhancement sequence is extracted from the vibration velocity channel enhancement sequence, which is composed of... Composition, where the superscript 0 indicates the vibration velocity channel.
[0077] 2> Perform mirror padding at both ends of the vibration velocity channel enhancement sequence.
[0078] Specifically, mirror filling refers to inserting several values in the reverse order of the beginning of the sequence at the left end of the sequence and several values in the reverse order of the end of the sequence at the right end of the sequence, so that a complete local window can also be formed near the first time step and the 1000th time step.
[0079] 3> Establish a local sliding window centered on each time step, and denote the half-width of the local sliding window as . The example value is 15, which corresponds to a window length of 31 time steps. The local sliding window is used to cover short-term impact areas while avoiding mixing long-term operating condition changes into local statistics.
[0080] Furthermore, within each local sliding window, the local mean, local second-order central moment, and local fourth-order central moment are calculated, and the local excess kurtosis is calculated accordingly. The local super-kurtosis at each time step is denoted as , Used to characterize the The sharpness of the vibration waveform near a time step. The larger the local excess kurtosis, the more likely there is an impulsive waveform within the window; when the local excess kurtosis is close to 0, it indicates that the waveform within the window is closer to a stationary periodic waveform or an approximate Gaussian fluctuation.
[0081] In the specific implementation, the arithmetic mean of the enhanced values of the 31 vibration velocity channels within the local sliding window is first calculated to obtain the local mean. Then, the squared average and fourth power average of the 31 values relative to the local mean are calculated respectively. Finally, the fourth power average is divided by the square of the squared average, and 3 is subtracted to obtain the local excess kurtosis. To prevent the squared average from approaching 0 and causing calculation anomalies, a very small constant is added to the denominator. An example value is shown below. .
[0082] In one embodiment, for example, let half the window width be... (Corresponding window length 5); at time step At that time, the five vibration velocity enhancement values within the window were: Then, calculate the local mean. Calculate the deviation of each value from the mean. Calculate the squared value of the deviation. Its average (squared mean) is Calculate the fourth power of the deviation. Its average value (fourth power mean) is Then, calculate the local excess kurtosis. This result is close to 0, indicating that the waveform within this window is relatively smooth and there are no strong impacts.
[0083] 4> Convert the local excess kurtosis into a local impact gating value, the first The local impact gating value at each time step is denoted as The value ranges from 0 to 1.
[0084] In one implementation, a sigmoid function can be used to implement the gated transition when... When the value is greater than the kurtosis threshold, the local impact gating value approaches 1; when When the value is less than the kurtosis threshold, the local impact gating value is close to 0. The kurtosis threshold can be set to 1.5, and the gating steepness parameter can be set to 6 to 12.
[0085] In one embodiment, for example, the kurtosis threshold is set to 1.5 and the gate steepness parameter is set to 10; when At that time, the gating value calculated using the Sigmoid function is close to 1, indicating a strong impact; the conversion method is expressed as... This value is very close to 1; among them, Indicates the first Local impact gating value at each time step.
[0086] 5> The vibration velocity channel enhancement sequence is modulated step-by-step using local impact gating values to obtain an impact-fidelity vibration sequence.
[0087] Specifically, the first The impact fidelity vibration values at each time step are denoted as . , can be according to It is calculated in the following way, where, Indicates the vibration velocity channel number The value after local impact gating at each time step; This represents the minimum retention factor, with an example value of 0.2. Indicates the first Local impact gating value at each time step; Indicates the vibration velocity channel number The enhancement value at each time step.
[0088] It should be noted that although the low kurtosis region may mainly contain harmonic components, it may still retain some fault background information. If the low kurtosis region is directly set to 0, it may destroy the overall timing structure of the vibration velocity channel. After adopting the minimum retention coefficient, the high kurtosis impact segment is basically retained, and the low kurtosis harmonic segment is attenuated but not completely deleted.
[0089] In one embodiment, as an example, when a bearing rolling element impact occurs within a certain local sliding window, the local excess kurtosis... It may reach 5.0. If the kurtosis threshold is 1.5 and the gating steepness parameter is 10, the local impact gating value is close to 1, and the impact fidelity vibration value is basically equal to the enhancement value. When the main component in a certain local sliding window is the rotational fundamental frequency harmonic, the local excess kurtosis is close to 0, and the local impact gating value is close to 0. At this time, the impact fidelity vibration value is about 20% of the enhancement value, and the harmonic effect is significantly weakened.
[0090] It should be noted that this step is not based on fixed-frequency filtering, but rather on adaptive gating based on the sharpness of the local waveform. For generator sets with significant speed variations, the fundamental frequency and harmonic positions will shift with the speed, requiring frequent parameter adjustments for fixed-band filtering; local impact gating does not rely on a fixed frequency band and can more stably preserve the time-domain characteristics of bearing faults and mechanical shock faults.
[0091] S205, Bidirectional trend enhancement of temperature channel and current channel and multi-channel reorganization Fault information in the temperature and current channels is usually manifested as a slow rise, a slow fall, a phased drift, or a continuous shift. If only short-term sudden changes are considered, early signals of cooling system faults and winding faults are easily overlooked.
[0092] This invention extracts bidirectional smoothing trends from the temperature and current channels, and recombines the trend components with the residual components to obtain the processed multivariate time-series matrix of the generator set. The specific steps are as follows: 1> Enhance the multi-element timing matrix of generator sets Extract the temperature channel enhancement sequence and the current channel enhancement sequence, where the temperature channel corresponds to... Current channel corresponding .
[0093] 2> Perform positive exponential moving averages on the temperature channel and the current channel respectively to obtain a positive trend sequence.
[0094] Specifically, for the first One physical channel Choose 1 or 2, and the positive trend sequence is denoted as... The first time step command Starting from the second time step, the current enhancement value is proportionally weighted with the positive trend value from the previous time step to obtain the current positive trend value. The smoothing factor is denoted as... The values range from 0.03 to 0.10. The smaller the value, the smoother the trend sequence.
[0095] In practical implementation, the smoothing factor Used to control how quickly a positive trend sequence tracks real-time data; The smaller the value, the smaller the weight assigned to the new data, and the trend series The smoother the time step, the less sensitive it is to short-term random fluctuations, but this also means a slower response to new trends; for time steps ( ), positive trend sequence The calculation method is as follows ,in, Indicates the previous time step ( The positive trend value of the time step.
[0096] In one embodiment, as an example, suppose , (Temperature channel); If ,but ;when At that time, if ,but The new data of 0.3 only had a 5% weighting effect on the trend, and the trend value only slowly changed from 0.1 to 0.11.
[0097] 3> Perform reverse exponential moving averages on the temperature channel and the current channel respectively to obtain the reverse trend sequence.
[0098] Specifically, the inverse exponential moving average is recursively applied from the 1000th time step to the 1st time step to offset the time lag caused by the unidirectional exponential moving average; the inverse trend sequence is denoted as... ,in Choose 1 or 2.
[0099] In its implementation, the inverse exponential moving average starts from the end of the sequence ( ) To the beginning of the sequence ( The recursion is similar to the forward recursion, but it smooths out the reverse temporal relationship; specifically, let... ,for Decrementing from 999 to 1, calculate .
[0100] In one embodiment, as an example, suppose (Current channel), smoothing factor To maintain consistency with positive smoothing, for example, 0.05; the enhancement value at step 1000 is 1.2, then... If the enhancement value in step 999 is 1.1, then proceed with the reverse recursion. .
[0101] 4> Averaging the positive trend sequence and the negative trend sequence yields a bidirectional trend sequence.
[0102] Specifically, the first The physical channel in the first The bidirectional trend value at each time step is denoted as... ,in Choose 1 or 2; bidirectional trend sequences have a smaller phase delay than unidirectional trend sequences, and can more accurately pinpoint the time period in which temperature drift or current drift occurs.
[0103] In practice, the bidirectional trend value is obtained by taking the arithmetic mean of the positive and negative trend values at the same time step, in order to offset the time lag caused by unidirectional smoothing. That is, for a time step... , .
[0104] In one embodiment, following the example above, in Step, if the value of the positive trend sequence in this step The value of the reverse trend sequence is 1.08. If the value is 1.195, then the bidirectional trend value for this step is... .
[0105] 5> Calculate the residual components of the temperature channel and the current channel.
[0106] Specifically, the first The physical channel in the first The residual components at each time step are denoted as... ,in Choose 1 or 2; the residual component is obtained by subtracting the bidirectional trend value from the enhanced value, and is used to preserve short-term fluctuations that are still diagnostically significant in the temperature and current channels.
[0107] In practical implementation, the residual component is obtained by subtracting the bidirectional trend value from the augmented value, thereby capturing short-term fluctuations and noise surrounding the long-term trend. .
[0108] In one embodiment, following the example above, in Step, enhance numerical values The value is 1.1, representing a two-way trend. If the value is 1.1375, then the residual components of this step are... .
[0109] 6>Weigh and combine the bidirectional trend components and residual components to obtain the trend enhancement sequence.
[0110] Specifically, the first The physical channel in the first The trend enhancement value at each time step is denoted as ,in Take 1 or 2; the trend enhancement value can be set according to... It is calculated in the following way, where, This indicates the trend weight, with values ranging from 1.2 to 1.8. This represents the residual weight, with values ranging from 0.8 to 1.2; hyperbolic tangent function. Used to limit the results to a stable range.
[0111] In one embodiment, for example, if the temperature channel shows a sustained slight increase in the early stages of a cooling system failure, the bidirectional trend sequence will remain positive across multiple consecutive time steps. Trend weights are then applied. Residual weights At that time, the continuous upward trend in the temperature channel is enhanced, while isolated random noise is still mainly located in the residual components and will not dominate the final trend enhancement result.
[0112] 7> The five physical channels are recombined to obtain the processed generator set multi-element timing matrix. The processed generator set multivariate time series matrix The size is ,in, Indicates the first The physical channel in the first The processed values at each time step.
[0113] Specifically, the vibration velocity channel uses the impact-fidelity vibration sequence obtained in step S204, i.e. The temperature channel uses a temperature channel trend enhancement sequence, i.e. The current channel adopts a current channel trend enhancement sequence, i.e. The speed channel retains the speed channel enhancement value in the multi-dimensional time-series matrix of the enhanced generator set, that is... The power channel retains the power channel enhancement values in the multi-element time-series matrix of the enhanced generator set, i.e. .
[0114] It should be noted that the processed generator set multi-element time series matrix It is not a simple normalization result, but a result that integrates adaptive reference of the operating domain, enhanced channel fault sensitivity, vibration velocity channel impact gating, and enhanced temperature and current trends. This makes the discrimination features of different fault types in the corresponding physical channels more prominent, while reducing the interference of normal fluctuations in irrelevant physical channels on subsequent models.
[0115] In one embodiment, such as Figure 3 As shown, the local impact gating process is analyzed, demonstrating the complete process of local impact gating executed by the vibration velocity enhancement sequence. Figure 3 Figure (a) in the image compares the enhanced sequence (blue) and the impact-gated impact fidelity sequence (red), with the impact segment preserved. Figure 3 Figure (b) shows the local excess kurtosis at each time step, with the peak position far exceeding the kurtosis threshold (1.5, dashed line), corresponding to the shock zone. Figure 3 Figure (c) shows the gating values obtained through the S-shaped function mapping. The gating value is close to 1 in the strong impact zone and close to 0 in the gentle impact zone. The horizontal axis represents time steps (unit: steps); the vertical axis represents... Figure 3 In figure (a), the amplitude (dimensionless) is shown. Figure 3 Figure (b) shows the local over-optimal kurtosis (dimensionless). Figure 3 In Figure (c), the gating values are (0-1, dimensionless).
[0116] In one embodiment, Figure 4 The bidirectional trend enhancement decomposition of the temperature channel is analyzed, decomposing the enhancement sequence into a positive trend, a negative trend, a bidirectional smoothed trend, and a residual component. The red solid line represents the bidirectional trend, with a very small phase delay, accurately tracking slow temperature changes; the residual component (cyan) preserves short-term fluctuations. The bidirectional exponential moving average effectively extracts the trend, smoothing random noise while preserving slowly changing fault information, providing a foundation for subsequent trend enhancement. The horizontal axis represents time steps (unit: steps), and the vertical axis represents numerical values (dimensionless).
[0117] S3. Construct and train a fault diagnosis model, which includes a channel-gated multi-scale temporal feature network and fault prototype constraints. Specifically, the processed generator set multi-element time-series matrix While channel differentiation enhancement has been completed, generator set fault identification still requires further utilization of the coupling relationships between different physical channels. For example, whether abnormal vibration velocity is accompanied by changes in rotational speed helps distinguish between imbalance and bearing faults; whether abnormal temperature is accompanied by abnormal current helps distinguish between cooling system faults and winding faults.
[0118] This invention constructs a coupled channel-gated multi-scale temporal feature network and introduces fault prototype constraints to make similar samples more compact in the embedding space and dissimilar samples more separated. The specific steps are as follows: S301, Locally Coupled Channel Gated Input Reconstruction Conventional time-series networks typically use the five physical channels as ordinary input dimensions, without explicitly expressing the cooperative relationships between these channels. This invention constructs a local coupling matrix near each time step and generates channel gating values from this matrix to enhance the interaction between physical channels and faults. The specific steps are as follows: 1> From the processed generator set multivariate time series matrix The five-channel input vector is extracted according to the time step, where the first... The five-channel input vector at each time step is denoted as... , It is a 5-dimensional vector, containing vibration velocity channel, temperature channel, current channel, rotational speed channel and power channel in sequence. The processed values at each time step.
[0119] In one embodiment, as an example, suppose the processed generator set multivariate timing matrix In the The channel values for each time step are as follows: vibration velocity channel 0.94, temperature channel 0.21, current channel 0.55, rotational speed channel 0.88, and power channel 0.10. Therefore, the five-channel input vector for the 10th time step... for .
[0120] 2> with the first A coupled statistics window is established centered on a time step, and the half-width of the coupled statistics window is denoted as . The example value is 5, which corresponds to a window length of 11 time steps. For positions at both ends of the sequence that are less than 11 time steps, a mirror padding method is used to fill in the gaps.
[0121] 3> Calculate the local coupling matrix within the coupling statistics window.
[0122] Specifically, the first The local coupling matrix at each time step is denoted as... , The size is ; the first in the local coupling matrix Line 1 Column elements are statistically coupled within the statistical window. The physical channel and the first The average of the products of the physical channels is obtained, where, and Both represent physical channel location indices, with values ranging from 0 to 4.
[0123] In one embodiment, as an example, take Coupled statistical window half-width (Window length 11), window coverage time steps 15 to 25; with the first Line (vibration velocity channel) and the first Taking the element of the column (speed channel) as an example, this element is the average of the product of the vibration velocity channel value and the speed channel value for 11 time steps within time steps 15 to 25. The calculation method is as follows: This value reflects the intensity of the coordinated change between vibration velocity and rotational speed within this local window.
[0124] It should be noted that the diagonal elements of the local coupling matrix represent the energy intensity of a single physical channel within the current local window, while the off-diagonal elements represent the cooperative relationship between two different physical channels within the current local window. For example, the off-diagonal elements between the vibration velocity channel and the rotational speed channel can indicate whether vibration anomalies are synchronized with changes in rotational speed; the off-diagonal elements between the temperature channel and the current channel can indicate whether heating anomalies are accompanied by changes in current.
[0125] 4> Local coupling matrix Straighten the vectors into 25-dimensional local coupling vectors in a fixed order. The fixed order can be achieved by straightening the vectors row by row, and this should be consistent between the training and inference phases. The local coupling vectors are used in the input channel gating mapping layer.
[0126] Furthermore, the 25-dimensional local coupling vector is input into the channel gating mapping layer to obtain a 5-dimensional channel gating vector; where, the ______ The channel gating vector at each time step is denoted as . , The dimension is 5, and the value range is from 0 to 1.
[0127] In one implementation, the channel-gated mapping layer consists of a fully connected layer and a cascaded sigmoid activation function; its input is a straightened 25-dimensional local coupling vector, and the weight matrix of the fully connected layer has a size of [missing information]. The bias vector length is 5; the 5-dimensional output of the fully connected layer, after being processed by the Sigmoid function, yields a 5-dimensional vector with values ranging from 0 to 1, which is the channel gating vector. .
[0128] 5> Convert the channel gating vector into a residual gating vector, denoted as . ,pass The residual gating vector is calculated in this way. Each component has a value ranging from 0.5 to 1.5, representing the percentage of preservation or enhancement of the corresponding physical channel.
[0129] It should be noted that the residual gating vector Using a range of 0.5 to 1.5 for each component can prevent a physical channel from being completely shut down, while allowing important physical channels to be moderately amplified.
[0130] 6> Multiply the residual gating vector element-wise with the five-channel input vector to obtain the gated five-channel input vector, where the first... The gated five-channel input vector at each time step is denoted as... , The dimension is 5.
[0131] Furthermore, this operation is repeated for all 1000 time steps to obtain the gated generator multivariate time sequence matrix. , The size is .
[0132] In one embodiment, as an example, when the multi-element timing matrix of the generator set is processed... When a local mode of increased vibration velocity but relatively stable rotational speed occurs, the vibration velocity channel in the local coupling matrix itself is enhanced, and the coupling relationship between the vibration velocity channel and the rotational speed channel is different from the synchronous change in the imbalance fault. The channel gating mapping layer can learn this difference and increase the gating ratio of the vibration velocity channel, thereby enhancing the bearing fault-related characteristics.
[0133] It should be noted that the local coupling channel gating of the present invention does not simply increase the number of input channels, but explicitly converts the pairwise collaborative relationship between the five physical channels into a channel enhancement ratio, thereby enabling the subsequent temporal feature network to simultaneously focus on single-channel anomalies and multi-channel coupling anomalies.
[0134] S302, Multi-scale Temporal Convolution Feature Extraction The duration of generator set faults varies along the time axis. Bearing impact may only last for a few time steps, misalignment and imbalance may exhibit periodic fluctuations, and cooling system faults may show a slow trend over a longer time scale. The length of a single convolution kernel is difficult to cover all of the above fault types simultaneously.
[0135] This invention employs multi-scale temporal convolution to extract fault features at different time scales, enabling the automatic coverage of various fault temporal morphologies, from short-term impacts and periodic fluctuations to long-term trends, using a unified network structure. This avoids manually designing features or selecting time windows for different fault types. The specific steps are as follows: 1> Gated generator set multi-element timing matrix The input consists of three parallel one-dimensional temporal convolution branches. The kernel lengths of the three parallel one-dimensional temporal convolution branches are 7, 31 and 101 respectively. The number of input channels is 5 for each branch and the number of output channels can be 32 for each branch.
[0136] Among them, the one-dimensional temporal convolution branch with a kernel length of 7 is used to extract local shocks and short-term abrupt changes; the one-dimensional temporal convolution branch with a kernel length of 31 is used to extract short-period fluctuations and local oscillations; and the one-dimensional temporal convolution branch with a kernel length of 101 is used to extract long-term trends such as temperature drift, current drift and power change.
[0137] In the actual implementation, the three convolutional branches have the same structure, differing only in the length of the convolutional kernel. There are differences. Taking 7, 31, and 101 respectively, each branch contains: a> One-dimensional convolutional layer: 5 input channels, 32 output channels, kernel length is... The step size is 1, and the padding method is mirror padding to keep the output time length 1000. b> Batch normalization layer: Used to normalize the 32 channels of the convolution output; c>ReLU activation function layer: used to add nonlinearity.
[0138] In the specific implementation, boundary padding is used for each parallel one-dimensional temporal convolution branch to keep the convolution output time length at 1000. Boundary padding can be zero padding or mirror padding, with mirror padding preferred to reduce abrupt changes in the convolution boundary at both ends of the sequence.
[0139] In the specific implementation, the convolution output of each parallel one-dimensional temporal convolution branch is sequentially processed by batch normalization and ReLU activation function to obtain temporal feature matrices at three scales; the size of the temporal feature matrix at each scale is [size missing]. .
[0140] 2> Concatenate the three-scale temporal feature matrices along the channel dimension to obtain a multi-scale temporal feature matrix. The size of the multi-scale temporal feature matrix is... The 96 feature channels are obtained by splicing together 32 feature channels at each of the three scales.
[0141] 3> Input the multi-scale temporal feature matrix into a one-dimensional convolutional layer of length 1 to compress the channel dimension from 96 to 64, thus obtaining the temporal hidden feature matrix. Temporal hidden feature matrix The size is ,in, Indicates the first A 64-dimensional temporal hidden feature vector for each time step.
[0142] In the specific implementation, the one-dimensional convolutional layer adopts the standard one-dimensional convolutional layer with the following parameters: 96 input channels (the number of channels after concatenation), 64 output channels, and a convolutional kernel length of 1. Then, the 96 multi-scale features at each time step are linearly combined and dimensionality reduced, without mixing information across time steps.
[0143] It should be noted that a one-dimensional convolutional layer of length 1 is used to fuse features output from convolutional branches of different scales without changing the time length; through this layer, the model can automatically learn the combined relationships between short-term impact features, periodic fluctuation features, and trend features.
[0144] It should also be noted that multi-scale temporal convolutional feature extraction can take into account impact faults, periodic faults and trend faults in the same network, which can avoid manually selecting different filters or different time windows for different fault types and improve the adaptability of the model in the identification of multiple types of generator set faults.
[0145] S303, Fault Attention Aggregation and Segmented Statistics Enhancement Temporal hidden feature matrix It contains hidden features across 1000 time steps, but key information about different faults may appear in different time periods. Bearing faults may be concentrated in a few impact time steps, while cooling system faults may gradually become apparent in the latter half.
[0146] This invention employs both attention aggregation and segmented statistical enhancement to obtain multi-granularity temporal representation vectors. The specific steps are as follows: 1> Set trainable query vectors , The dimension is 64; the trainable query vector It is a randomly initialized 64-dimensional vector, which is a trainable parameter that is optimized and updated through backpropagation during training. It is used to learn "which hidden features at which time steps are more conducive to fault classification" during training.
[0147] 2> Temporal hidden feature matrix Attention scores are calculated at each time step.
[0148] Specifically, the first The attention score at each time step is denoted as Through trainable query vectors With temporal hidden feature vectors The inner product is calculated; to avoid excessively high attention scores, the inner product result can be divided by... (i.e., hidden features) The square root of dimension 64), calculated as follows: ,in, Represents a trainable query vector The transpose of .
[0149] 3> Perform Softmax normalization on the attention scores of 1000 time steps to obtain 1000 attention weights.
[0150] Specifically, the first The attention weights at each time step are denoted as follows: , The value ranges from 0 to 1, and the sum of the attention weights of 1000 time steps is 1; the larger the attention weight, the greater the contribution of that time step to the final fault judgment.
[0151] 4> Use attention weights to perform a weighted summation of 1000 temporal hidden feature vectors to obtain a global attention aggregation vector. Global attention aggregation vector The dimension is 64, which is used to represent the most discriminative time-series fault information in the entire sample.
[0152] 5> Divide the 1000 time steps into 5 consecutive segments, each containing 200 time steps; the first consecutive segment corresponds to time steps 1 to 200, the second consecutive segment corresponds to time steps 201 to 400, and so on, with the fifth consecutive segment corresponding to time steps 801 to 1000.
[0153] Furthermore, the mean and standard deviation of the temporal hidden feature vector within each consecutive segment are calculated dimension by dimension. The segment mean vector of a series of consecutive segments is denoted as... , No. The standard deviation vector of a series of consecutive segments is denoted as . ,in, This represents the index of a continuous segment, with values ranging from 1 to 5. The segment mean vector represents the average level of fault characteristics within that continuous segment, while the segment standard deviation vector represents the degree of fluctuation of fault characteristics within that continuous segment.
[0154] 6> Concatenate the segment mean vector and segment standard deviation vector of 5 consecutive segments in a fixed order to obtain the segmented statistical vector. Since each continuous segment contains a 64-dimensional segment mean vector and a 64-dimensional segment standard deviation vector, the five continuous segments yield a total of... Dimensional statistical features, therefore piecewise statistical vectors The dimension is 640.
[0155] 7> Aggregate the global attention vector Piecewise statistical vectors By concatenating the vectors, a multi-granularity temporal representation vector is obtained. Multi-granularity temporal representation vector The dimension is 704, of which 64 dimensions come from the global attention aggregation vector and 640 dimensions come from the segmented statistical vector.
[0156] In one embodiment, for example, a cooling system failure may not be obvious in the first 200 time steps of a sample, but the temperature gradually increases in the following 600 time steps. At this time, the segment mean vectors of the 3rd to 5th consecutive segments will gradually change, and the segmented statistical vector can reflect this stage-by-stage change; at the same time, the attention weights will gradually shift towards the time steps where the temperature drift is more obvious, making the global attention aggregation vector more focused on the later failure signals.
[0157] In one embodiment, Figure 5 The distribution of attention aggregation weights in the fault attention model (bearing fault samples) is analyzed, showing the attention weight distribution learned by the fault attention module over 1000 time steps on simulated bearing fault samples. The weights exhibit significant spikes near several time points of the pre-injected impact, while the weights are low and uniform in the remaining stable segments. The model can automatically focus on the local impact time periods most important for fault identification, demonstrating the effectiveness of the attention aggregation mechanism. Figure 5 The red dashed line marks the injection impact location. The horizontal axis represents the time step (unit: step), and the vertical axis represents the attention weight (dimensionless, summing to 1).
[0158] S304, Embedded Vector Generation and Fault Prototype Augmentation Classification Multi-granularity temporal representation vector With higher dimensionality, direct classification would increase the number of parameters, and the boundaries between adjacent fault categories might be unstable.
[0159] This invention projects multi-granularity temporal representation vectors into compact embedding vectors and sets a category prototype vector for each fault category, so that the final classification considers both the linear classification result and the prototype distance in the embedding space. The specific steps are as follows: 1> Multi-granularity temporal representation vector Inputting a fully connected projection layer yields a 128-dimensional projection vector.
[0160] In its implementation, the fully connected projection layer consists of a single fully connected layer, and its weight matrix (trainable parameters) has a size of [size missing]. The bias vector (trainable parameter) has a length of 128, so the input dimension of the fully connected projection layer is 704 and the output dimension is 128.
[0161] 2> Perform ReLU activation and batch normalization sequentially on the 128-dimensional projection vector to obtain the embedding vector. Embedded vectors The dimension is 128, used to represent the original generator set multivariate time series samples. The feature locations in the embedding space after preprocessing and feature extraction.
[0162] 3> Set up category prototype vectors for each of the 6 fault categories. The category prototype vectors are initialized along with the network parameters. Initially, for each category... A 128-dimensional vector is randomly sampled from a normal distribution with a mean of 0 and a standard deviation of 1, and used as its initial class prototype vector.
[0163] Specifically, the first The prototype vector of each fault category is denoted as... , The dimension is 128. The value ranges from 0 to 5; the category prototype vector is used to represent the center position of the corresponding fault category in the embedding space.
[0164] 4> Calculate embedding vectors using a linear classifier The corresponding linear classification score, the linear classifier consists of a fully connected layer, without bias terms, and its weight matrix has a size of... Therefore, the linear classifier has an input dimension of 128 and an output dimension of 6, where the first... The linear classification score for each fault category represents the embedding vector. Under the conventional linear classification boundary, it belongs to the first... The tendency of each fault category.
[0165] 5> Calculate the embedding vector The negative squared Euclidean distance to the prototype vector of each category.
[0166] Specifically, the first The prototype distance adjustment value for each fault category is denoted as: , represented as ,in, This represents the L2 norm; the larger the prototype distance adjustment value, the stronger the embedding vector. The closer to the first A category prototype vector for each fault category.
[0167] 6> Fuse the linear classification score with the prototype distance adjustment value to obtain the category logits vector. Category logits vector The dimension is 6, where the first... element Indicates the first The final unnormalized classification score for each fault category.
[0168] In practice, the fusion method involves multiplying the linear classification score by a coefficient. The prototype distance adjustment values are added element by element, and the calculation method is expressed as follows: ,in, It is a linear classifier for the first... The class output score, This represents the prototype fusion coefficient, with values ranging from 0.05 to 0.2.
[0169] 7> For the category logits vector Perform Softmax normalization to obtain the predicted probabilities of 6 fault categories. The fault category with the highest predicted probability is taken as the generator set fault classification result for the current sample.
[0170] It should be noted that prototype augmentation classification can provide embedding space center constraints for linear classifiers; when an early fault sample is near the boundary of an adjacent class, if the embedding vector of the sample is closer to the prototype vector of the true fault class, the prototype distance adjustment value will improve the final classification score of that class, thereby enhancing classification stability.
[0171] S305, Joint Loss Training and Dynamic Update of Class Prototype Vectors To simultaneously improve classification accuracy and embedding space separability, this invention employs class-weighted cross-entropy loss, similarity compaction loss, and prototype margin loss to jointly train the model. Class-weighted cross-entropy loss addresses the imbalance of fault category samples; similarity compaction loss brings the embedding vectors of similar samples closer to their corresponding class prototype vectors; and prototype margin loss minimizes the distance between prototype vectors of different categories. The specific steps are as follows: 1> Count the number of samples for the six fault categories in the training set.
[0172] Specifically, the first The number of samples for each fault category is denoted as The largest number of samples among all fault categories is denoted as . No. The category loss weights for each fault category are denoted as: According to the specific It is calculated in the following way, where, This represents a constant to prevent the denominator from being zero; the example value is 1. The smaller the sample size of a fault category, the greater the corresponding category loss weight.
[0173] 2> In each training batch, the original generator set multivariate time-series samples within the batch are sequentially processed through steps S2 and S3 to obtain the class logits vector and embedding vector for each sample; the batch size is denoted as . Examples of possible values are 32 or 64; The embedding vector of a sample is denoted as The true category label is denoted as ,in Indicates the sample index within the batch. The value ranges from 0 to 5.
[0174] 3> Calculate the class-weighted cross-entropy loss based on the class logits vector, the true class label, and the class loss weights; the class-weighted cross-entropy loss is used to improve the true class prediction probability and apply greater training weights to minority class samples.
[0175] In one embodiment, for example, suppose there are 2 samples in a batch, and the total number of categories is... ; Sample True Category 1 (Normal), its class loss weights The probability of class 0 in its predicted probability is ; Sample True Category 2 (If not, assume it's a minority class), its class loss weights The probability of class 2 in its predicted probability is The category-weighted cross-entropy loss for this batch is: The loss of minority class samples is weighted Magnification makes the model focus more on correctly classifying the sample.
[0176] 4> Calculate the class compaction loss, which is obtained by calculating the squared Euclidean distance between the embedding vector of each sample and the prototype vector of the true class.
[0177] Specifically, for the first Sample, calculate ,in, Represents the true category The corresponding category prototype vector; the average of all samples in the batch is used to obtain the same category compact loss; the smaller the same category compact loss, the more concentrated the samples of the same fault category are in the embedding space.
[0178] In one embodiment, as an example, suppose there are two samples in the batch with a true class of 2 (misalignment), and their corresponding embedding vectors are respectively and The current category prototype vector is Similar compact loss = This loss term drives samples with similar physical meanings to cluster in the feature space.
[0179] 5> Calculate the prototype interval loss.
[0180] Specifically, pairwise comparisons are performed on the prototype vectors of the six categories. When the squared Euclidean distance between two prototype vectors of different categories is less than the preset prototype interval, the comparison is made accordingly. A penalty is applied when the squared Euclidean distance between two different class prototype vectors is greater than or equal to the preset prototype interval. No penalty is incurred at this time. The preset prototype interval is specified. The value can be 1.0, and the prototype interval loss weight can be between 0.05 and 0.2.
[0181] In one embodiment, as an example, the prototype interval is set. prototype vectors for category 0 and category 1 , Calculate its squared Euclidean distance; if Then a penalty item will be generated. (i.e., the calculated prototype interval loss); if the distance is 1.5, the penalty term is 0 (i.e., the calculated prototype interval loss), thus forcing the center points of different fault categories to maintain a sufficient distance in the feature space.
[0182] 6> The class-weighted cross-entropy loss, similarity compactness loss, and prototype margin loss are summed in weights to obtain the total model training loss. The total model training loss is denoted as... , represented as: ; in, This represents the category-weighted cross-entropy loss; Indicates similar compact losses; Indicates prototype interval loss; This represents the weight of the similar compact loss, with an example value of 0.1. This represents the prototype interval loss weight, with an example value of 0.05.
[0183] 7> Update the trainable parameters in the network using the Adam optimizer.
[0184] Specifically, trainable parameters include channel-gated mapping layer parameters, multi-scale temporal convolution branch parameters, one-dimensional convolutional layer parameters of length 1, trainable query vector, fully connected projection layer parameters, batch normalization parameters, linear classifier parameters, batch normalization layer parameters after the fully connected projection layer, and class prototype vectors, etc.
[0185] In one implementation, the initial learning rate can be set to... The batch size can be 64, and the number of training rounds can be 100 to 200. After each training round, the macro average F1 score is calculated using the validation set, and the model parameters at the highest macro average F1 score are saved.
[0186] 8> After each training batch completes the network parameter update, perform an exponential moving average update on the class prototype vectors.
[0187] Specifically, for the first For each fault category, first find the true class label in the current training batch that equals... The samples; if samples of this fault category exist in the current training batch, then calculate the average of the embedding vectors of these samples, denoted as . Then follow The category prototype vector is updated in the manner described above, where, Indicates the updated number The prototype vectors of each fault category; Indicates the number before the update The prototype vectors of each fault category; This represents the prototype update momentum coefficient, with values ranging from 0.95 to 0.99. If the current training batch does not contain the [previous batch name]... For the sample of the fault category, then the first fault category... The category prototype vectors for each fault category remain unchanged.
[0188] In one embodiment, for example, the prototype vector of bearing fault categories may not be near the center of the embedding vector of bearing fault samples in the early stage of training. As training batches are continuously input, the average embedding vector of bearing fault samples is gradually injected into the prototype vector of bearing fault categories through exponential moving average, so that the prototype vector of bearing fault categories gradually moves to the vicinity of the center of the bearing fault embedding cluster; at the same time, the prototype margin loss promotes the bearing fault category prototype vector to maintain distance from the prototype vector of unbalanced category and misaligned category, thereby reducing the confusion between adjacent mechanical fault categories.
[0189] It should be noted that the training objective of this invention is not only to improve the final classification probability, but also to improve the classification stability of early fault samples and minority class fault samples by constraining the embedding space structure through category prototype vectors. This is especially suitable for generator set mechanical fault identification scenarios with similar features such as imbalance, misalignment, and bearing faults.
[0190] S4. Input the target data of the device to be diagnosed into the fault diagnosis model to perform fault diagnosis, and output the fault diagnosis results. Specifically, after model training is completed, the runtime reference trajectory matrix, runtime normal robust scale vector, trained coupled channel gated multi-scale temporal feature network parameters, and category prototype vector are deployed to the generator set online monitoring system. The online monitoring system continuously receives data from five physical channels: vibration velocity, temperature, current, speed, and power, and completes fault identification according to the same process as in the training phase. The specific steps are as follows: 1> The online monitoring system continuously reads data from the vibration velocity channel, temperature channel, current channel, speed channel, and power channel according to a fixed diagnostic time axis. The system maintains a sliding diagnostic window with a length of 1000 time steps. Whenever new time step data arrives, the oldest time step data is discarded and the latest time step data is added to form a multi-dimensional time sequence sample of the original generator set to be identified.
[0191] 2. Perform operating domain determination on the multi-dimensional time-series samples of the original generator set to be identified. Specifically, calculate the average speed channel value and the average power channel value of the multi-dimensional time-series samples of the original generator set to be identified to form the current operating domain description vector; compare the current operating domain description vector with the pre-saved operating domain centers to determine the current operating domain index. .
[0192] 3> Index based on the current running domain Read the corresponding runtime reference trajectory matrix and the normal robust scaling vector of the operating domain Following steps S202 to S205, the following steps are performed sequentially on the original generator set multivariate time series sample to be identified: standardization deviation calculation, channel fault sensitivity coefficient calculation, differential nonlinear enhancement, vibration velocity channel local impact gating, temperature channel and current channel bidirectional trend enhancement, and multi-channel reorganization, to obtain the generator set multivariate time series matrix after the identification process.
[0193] 4> Input the multivariate time-series matrix of the generator set to be identified and processed into the trained coupled channel-gated multi-scale time-series feature network. The network first constructs a local coupling matrix and generates channel-gated vectors, and performs local coupling-gated reconstruction on the five physical channels; then, it extracts short-term impact, periodic fluctuation, and long-term trend features through three parallel one-dimensional time-series convolutional branches; then, it generates multi-granularity time-series representation vectors through fault attention aggregation and segmented statistical enhancement; finally, it generates embedding vectors and combines them with a linear classifier and category prototype vectors to output the predicted probabilities of six fault categories.
[0194] 5> The fault category with the highest predicted probability is used as the identification result of the current sliding diagnostic window. If the maximum predicted probability is greater than the confidence threshold, the corresponding fault alarm is output; the confidence threshold can be set to 0.7. To reduce occasional false alarms, it can be required that the same fault category meets the confidence threshold in three consecutive sliding diagnostic windows before triggering a formal alarm.
[0195] 6. The online monitoring system also outputs auxiliary diagnostic information, including the current operating domain index, channel fault sensitivity coefficients for the five physical channels, the fault category corresponding to the maximum predicted probability, the predicted probability for each fault category, and the fault occurrence time window. The channel fault sensitivity coefficients help maintenance personnel determine the primary response channel for a fault. For example, when a bearing fault alarm occurs, the channel fault sensitivity coefficient for the vibration velocity channel is usually high; similarly, when a cooling system fault alarm occurs, the channel fault sensitivity coefficient for the temperature channel is usually high.
[0196] It should be noted that the online identification phase strictly reuses the operating domain reference trajectory, normal robust scale, channel enhancement method, and deep classification network parameters from the training phase to avoid identification bias caused by inconsistencies between the training and inference processes. By continuously executing the above process through a sliding diagnostic window, real-time identification and early warning of generator set normal operation, imbalance, misalignment, bearing faults, winding faults, and cooling system faults can be achieved.
[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method of early warning of failure of a genset, characterized in that, Includes the following steps: Data from five physical channels—vibration velocity, temperature, current, speed, and power—are collected during the generator set's operation. These five physical channels are aligned on the time axis and labeled; the labels represent the generator set's fault state category. The operating domain is divided according to the speed channel and the power channel. A normal reference trajectory and a normal robust scale are constructed for each operating domain. The channel fault sensitivity coefficient of the current original generator set multivariate time series sample relative to the current operating domain reference trajectory is calculated. Differential nonlinear enhancement is performed on the five physical channels. Local impact gating is performed on the vibration velocity channel. Trend enhancement is performed on the temperature channel and the current channel to obtain the processed generator set multivariate time series matrix. Construct and train a fault diagnosis model, which includes a channel-gated multi-scale temporal feature network and fault prototype constraints. Input the target data of the device to be diagnosed into the fault diagnosis model to perform fault diagnosis, and output the fault diagnosis results. The channel fault sensitivity coefficient of the current original generator set multivariate time-series sample relative to the current operating domain reference trajectory is calculated as follows: reading a current running domain index a corresponding running domain reference trajectory matrix and a running domain normal robust scale vector Computing a standardized deviation for each time step, each physical channel in the original generator set multivariate time series samples The absolute value of the standardized deviation for each physical channel is taken to obtain the channel deviation sequence; In the channel deviation sequence of each physical channel, calculate the 95th percentile deviation value and the average deviation value; The channel anomaly response value is calculated based on the 95th percentile deviation value and the average deviation value; The abnormal response values of the five physical channels are normalized into channel fault sensitivity coefficients. Differential nonlinear enhancements are performed on the five physical channels, specifically as follows: The obtained standardized bias is used as the input for nonlinear enhancement. Based on the channel fault sensitivity coefficient, a hyperbolic tangent nonlinear mapping is performed on each physical channel to obtain the multivariate time sequence matrix of the enhanced generator set; The positive and negative directions of the enhancement values are retained. When the enhancement value is positive, it indicates that the corresponding physical channel is higher than the normal reference trajectory of the current operating domain; when the enhancement value is negative, it indicates that the corresponding physical channel is lower than the normal reference trajectory of the current operating domain.
2. The method of claim 1, wherein, The labels include normal, unbalanced, misaligned, bearing fault, winding fault, and cooling system fault.
3. The method of claim 1, wherein, The operating domain is divided according to the speed channel and the power channel. A normal reference trajectory and a normal robust scale are constructed for each operating domain, specifically as follows: Extract the speed channel and power channel from the normal category samples in the training set, calculate the average speed channel and the average power channel for each normal category sample, and obtain the running domain description vector; Clustering is performed on the runtime description vectors of all normal category samples to obtain multiple runtimes; For the For each operating domain, all normal category samples belonging to that operating domain are collected. Within these normal category samples, the normal reference median is calculated for each time step and each physical channel, yielding the [missing value]. The runtime reference trajectory matrix of each runtime domain; For the first Continue to calculate the normal robust scale of each physical channel in all normal class samples belonging to the running domain to obtain a running domain normal robust scale vector; For the original generator set multivariate time series sample to be processed, calculate its average speed channel value and average power channel value to obtain the current operating domain description vector. Compare the distance between the current operating domain description vector and the center of each operating domain, and select the operating domain with the closest distance as the current operating domain.
4. The method of claim 1, wherein, Differential nonlinear enhancements are performed on the five physical channels, specifically as follows: Extract the vibration velocity channel enhancement sequence from the multivariate time-series matrix of the enhanced generator set; Mirror padding is applied to both ends of the vibration velocity channel enhancement sequence; Establish a local sliding window centered on each time step; Converting local excess kurtosis into local impact gating values, the first... The local impact gating value at each time step is denoted as ; By using local impact gating values to modulate the vibration velocity channel enhancement sequence step by step, an impact-fidelity vibration sequence is obtained.
5. The generator set fault early warning method according to claim 4, characterized in that, The enhanced execution trend in temperature and current channels is specifically as follows: Extract temperature channel enhancement sequences and current channel enhancement sequences from the multivariate time-series matrix of the enhanced generator set; A positive exponential moving average was applied to both the temperature channel and the current channel to obtain a positive trend sequence. By performing inverse exponential moving averages on the temperature channel and the current channel respectively, an inverse trend sequence is obtained; The two-way trend sequence is obtained by averaging the positive trend sequence and the negative trend sequence. Calculate the residual components of the temperature channel and the current channel; By weighting and combining the bidirectional trend sequence and the residual components, a trend-enhancing sequence is obtained. The five physical channels were recombined to obtain the processed multi-element timing matrix of the generator set.
6. The generator set fault early warning method according to claim 1, characterized in that, The specific steps for constructing and training the fault diagnosis model are as follows: A local coupling matrix is constructed near each time step, and channel gating values are generated from the local coupling matrix to enhance physical channels related to fault coupling. Multi-scale temporal convolution is used to extract fault features at different time scales, enabling the automatic coverage of various fault temporal forms, from short-term impacts and periodic fluctuations to long-term trends, using a unified network structure. Attention aggregation and segmented statistical augmentation are employed to obtain multi-granular temporal representation vectors; The multi-granular temporal representation vector is projected into a compact embedding vector, and a category prototype vector is set for each fault category, so that the final classification considers both the linear classification result and the prototype distance in the embedding space. The model is trained using class-weighted cross-entropy loss, same-class compaction loss, and prototype interval loss. Class-weighted cross-entropy loss is used to address the imbalance of the number of fault class samples, same-class compaction loss is used to make the embedding vectors of the same class samples close to the corresponding class prototype vectors, and prototype interval loss is used to keep the prototype vectors of different classes at the minimum distance.
7. The method of claim 6, wherein, A local coupling matrix is constructed near each time step, and channel gating values are generated from the local coupling matrix to enhance physical channels related to fault coupling. Specifically: Five-channel input vectors are extracted from the processed generator set multivariate time-series matrix according to time steps, where the first... The five-channel input vector at each time step is denoted as... ; With the first A coupled statistics window is established centered on a time step, and the half-width of the coupled statistics window is denoted as . The example value is 5, which corresponds to a window length of 11 time steps. For positions at both ends of the sequence that are less than 11 time steps, a mirror padding method is used to fill in the gaps. Calculate the local coupling matrix within the coupling statistics window; The local coupling matrix is straightened into a 25-dimensional local coupling vector in a fixed order; Convert the channel gate vector into a residual gate vector, where the residual gate vector = 0.5 + the channel gate vector; The residual gating vector is multiplied element-wise with the five-channel input vector to obtain the gated five-channel input vector.
8. The method of claim 6, wherein, Attention aggregation and segmented statistical augmentation are employed to obtain multi-granularity temporal representation vectors, specifically: Setting trainable query vectors ; Temporal hidden feature matrix At each time step, an attention score is calculated, which is obtained by combining the trainable query vector with the temporal hidden feature vector. The inner product is calculated; Softmax normalization is performed on the attention scores at 1000 time steps to obtain 1000 attention weights; The global attention aggregation vector is obtained by weighting and summing the 1000 temporal hidden feature vectors using attention weights. The 1000 time steps are equally divided into 5 consecutive segments, each containing 200 time steps; The segment mean vector and segment standard deviation vector of 5 consecutive segments are concatenated in a fixed order to obtain the segmented statistical vector; By concatenating the global attention aggregation vector with the segmented statistical vector, a multi-granularity temporal representation vector is obtained.