Container generator set multi-modal data fusion fault early warning method and system

CN122548575APending Publication Date: 2026-08-11LIAONING HEYU IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供集装箱发电机组多模态数据融合的故障预警方法及系统,解决了现有集装箱发电机组监测系统因模态特征孤立割裂、时序基准错位以及静态预警阈值刚性固化,导致在复杂波动工况下多物理场早期隐微故障特征难以被精准提取与因果关联,进而造成故障预警严重滞后与误报频发的技术问题

Benefits of technology

[0063] This invention overcomes the limitations of isolated information from a single mode by synchronously acquiring and processing multimodal physical quantities such as vibration, temperature, oil, acoustic, and electrical parameters. Using vibration sequences as a benchmark, it employs cross-correlation calculations and nonlinear phase compensation to achieve precise alignment of cross-physical field signals in the time dimension, laying a unified spatiotemporal benchmark for subsequent fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548575A_ABST
    Figure CN122548575A_ABST
Patent Text Reader

Abstract

This invention relates to the field of industrial equipment condition monitoring and fault diagnosis technology, and discloses a fault early warning method and system for multimodal data fusion of container generator sets. The method includes: synchronously acquiring and preprocessing multimodal data; using vibration sequences as anchor points, performing cross-physical field benchmark alignment through cross-correlation and nonlinear phase compensation; extracting feature vectors of each mode and calculating modal confidence weight vectors; performing two-stage fusion, in the first stage splicing L2 normalized features, and in the second stage adding the weight vectors and features and then performing time-series weighting using a multi-head self-attention mechanism; using a dual-branch network to capture local impacts and long-term degradation trends, and controlling feature cross-fusion by a dynamic gating coefficient based on real-time operating parameters; finally mapping to fault prediction probability, and triggering multi-level protection actions based on dynamically adjusted early warning thresholds. This invention improves the accuracy and robustness of fault early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial equipment condition monitoring and fault diagnosis technology, specifically to a fault early warning method and system for multimodal data fusion of container generator sets. Background Technology

[0002] As a core underlying infrastructure ensuring absolute continuity of energy supply in data centers, communication base stations, ocean-going vessels, and special industrial production scenarios, container generator sets are widely deployed in unattended operating spaces in remote locations or extremely harsh environments. The robustness of their operation directly determines the continuous operational stability of the upper-level business systems.

[0003] Traditional container generator set condition monitoring paradigms mostly rely on discretely arranged single-modal sensors to collect fragmented local physical quantities, and trigger alarm logic at the system edge through rigid hard boundary threshold comparisons. Container generator sets are essentially high-dimensional complex giant systems with deeply coupled mechanical rotor shaft systems, stator electromagnetic windings, cooling fluid networks, and combustion chambers. Any material fatigue or gap wear at the microscopic level will be synchronously radiated and dissipated in the macroscopic form of extremely weak energy into multiple physical dimensions such as mechanical vibration, acoustic pressure, thermodynamic gradient, and electromagnetic field distortion.

[0004] Single-dimensional threshold monitoring not only fails to capture the hidden early fault evolution chain that crosses physical boundaries, but is also easily overwhelmed by the complex and ever-changing background noise of the environment and the normal fluctuations brought about by transient operating conditions such as unit start-up and shutdown and nonlinear load switching. This results in the real fault signal being in a long-term monitoring blind zone before it occurs, and the frequent occurrence of invalid false alarms further weakens the engineering reliability of the monitoring system, making it unable to meet the high-dimensional operation and maintenance requirements of modern industry for the continuous high reliability of equipment. Summary of the Invention

[0005] The purpose of this invention is to provide a fault early warning method and system for multimodal data fusion of container generator sets. It solves the technical problems of existing container generator set monitoring systems, which suffer from isolated and fragmented modal features, misaligned time series references, and rigid static early warning thresholds. These problems make it difficult to accurately extract and causally correlate early hidden fault features in multi-physics fields under complex fluctuating operating conditions, resulting in serious delays in fault early warning and frequent false alarms.

[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0007] A fault early warning method based on multimodal data fusion for container generator sets includes the following steps:

[0008] Step 1: Synchronously acquire raw multimodal data of the container generator set under operating conditions, and perform filtering, resampling and dimensional normalization preprocessing on the raw multimodal data, and output the preprocessed data of each channel; wherein, the raw multimodal data includes at least vibration sequence, temperature sequence, oil parameter sequence, acoustic signal sequence and electrical parameter sequence;

[0009] Step 2: Using the vibration sequence in each preprocessed channel data as the reference anchor point, calculate the time delay of the remaining sequences relative to the reference anchor point, and perform nonlinear phase compensation alignment on the remaining sequences based on the time delay to output modal data with a unified time reference.

[0010] Step 3: For each modal data of the unified time reference, perform time-domain and / or frequency-domain feature extraction independently according to the physical modal type, and output the feature vector of each modality respectively;

[0011] Step 4: Calculate the modality confidence weight vector corresponding to each modality feature vector and perform two-stage fusion: The first stage of fusion performs L2 normalization on each modality feature vector at the same time and then concatenates them in a preset physical order to obtain an early fusion feature vector; The second stage of fusion first projects the modality confidence weight vector and adds it to the early fusion feature vector, and then uses a multi-head self-attention mechanism to perform temporal weighting on the time series obtained after addition, and outputs the fusion feature vector.

[0012] Step 5: The first and second time-series branches are used in parallel to receive the fused feature vector, capturing local impact features and long-term degradation trends respectively; at the same time, dynamic gating coefficients are generated based on the real-time operating parameters of the container generator set, and the dynamic gating coefficients are used to control the feature cross-fusion between the two branches, outputting the interactively enhanced spliced ​​vector.

[0013] Step six: Map the enhanced splicing vector to the predicted probability of the fault mode, and dynamically adjust the warning threshold according to the real-time operating parameters; when the predicted probability exceeds the warning threshold, trigger the corresponding control and protection action.

[0014] Furthermore, step one involves performing filtering on the original multimodal data, specifically including: using an adaptive filter to clean the background noise of the acoustic signal sequence, wherein the weight coefficient update equation of the adaptive filter is:

[0015]

[0016] In the formula, This is the updated filter weight coefficient vector; This is the current filter weight coefficient vector; The convergence factor; The error deviation between the desired reference signal and the actual output of the adaptive filter is given by the desired reference signal, which is taken from the background noise signal collected by an auxiliary reference microphone located far from the noise source. is the acoustic signal input vector.

[0017] Furthermore, the dimensional normalization mentioned in step one specifically includes: performing zero-mean standardization on the resampled channel data:

[0018]

[0019] In the formula, For the first The standardized values ​​for each channel; This is the real signal value after resampling of this channel; and The first The historical baseline mean and standard deviation of each channel are pre-calculated under the fault-free baseline operating condition of the generator set.

[0020] Furthermore, step two involves calculating the time delay of the remaining sequences relative to the reference anchor point, specifically including: calculating the cross-correlation function between the sequence to be aligned and the reference anchor point sequence.

[0021]

[0022] in, This represents the total number of sampling points within the sliding time window. This is a discrete-time delay index, representing the corresponding physical time delay. , To ensure a uniform sampling rate, it is necessary to iterate through... The cross-correlation sequence was obtained.

[0023] Furthermore, step two, which involves performing nonlinear phase compensation alignment based on the time delay, specifically includes: retrieving the time delay estimate that maximizes the cross-correlation function value. ;

[0024] A cubic spline interpolation polynomial is used to perform nonlinear phase compensation on the sequence to be aligned. This polynomial operates within the interval... Expanded above:

[0025]

[0026] In the formula, The result is the interpolation result. These are the polynomial coefficients; It is a time variable; This is the starting point of the time interval.

[0027] Furthermore, the feature extraction described in step three specifically includes:

[0028] Extracting the temporal kurtosis coefficient of the vibration sequence :

[0029]

[0030] In the formula, The amplitude of the vibration. To calculate the mean amplitude of vibrations within the window, the Mel frequency cepstral coefficients of the acoustic signal sequence are extracted. The Mel frequency mapping equation is as follows:

[0031]

[0032] In the formula, The Mel frequency value, The linear physical frequency value; the rate of temperature rise of the extracted temperature sequence. and temperature fluctuation variance Extract the three-phase imbalance from the electrical parameter sequence. Total Harmonic Distortion Extract the Pearson correlation coefficient between the oil temperature series and the oil pressure series. :

[0033]

[0034] In the formula, , These are the sequences for lubricating oil temperature and pressure, respectively. , Its mean; , Its standard deviation.

[0035] Furthermore, step four involves calculating the modal confidence weight vector, specifically including: constructing and decomposing the covariance matrix of the historical eigenvectors for each modality to obtain the eigenvalues. ; Calculate the modal quality factor :

[0036]

[0037] In the formula, For the first The feature vector dimension of each modality; All modalities' The model is concatenated and mapped to the output modal confidence weight vector through a single-layer perceptron and a sigmoid activation function.

[0038] Furthermore, the L2 normalization of the first-stage fusion in step four is as follows:

[0039]

[0040] In the formula, The original feature vector, The normalized feature vectors, for The L2 norm.

[0041] Furthermore, the multi-head self-attention mechanism used in the second stage of fusion in step four is calculated as follows:

[0042]

[0043]

[0044] In the formula, For the first The query, key, and value matrix of the size. For the dimension of the head, For the number of heads, This is for outputting the projection matrix.

[0045] Furthermore, in step five, the first temporal branch in parallel is a temporal convolutional network composed of multiple one-dimensional convolutional layers; the second temporal branch is a recurrent network composed of multiple bidirectional long short-term memory networks.

[0046] Furthermore, the mathematical intervention model for controlling the feature cross-fusion between the two branches in step five is as follows:

[0047]

[0048] In the formula, To enhance the features of the first branch, Features of the original first branch For dynamic gating coefficients, The cross-attention weight matrix is ​​generated from the features of the second branch. The Hadamard product is represented by the dynamic gating coefficient. ,in It is a high-dimensional operating condition sensing vector generated based on real-time operating condition parameter encoding.

[0049] Furthermore, the dynamically adjusted warning threshold in step six... The calculation formula is:

[0050]

[0051] In the formula, For the first Level benchmark threshold; This represents the load change rate. Functions for the runtime phase; Ambient temperature; For reference temperature; This is the adjustment coefficient.

[0052] Furthermore, the temporal convolutional network in the first temporal branch adopts a method based on the rotor rotation frequency. A dynamic convolutional architecture that adaptively adjusts the dilation rate. The calculation formula is:

[0053]

[0054] In the formula, The rated reference frequency, For the first The scaling constant of the layer; The reconstructed one-dimensional convolution operator is:

[0055]

[0056] Furthermore, in the process of mapping the enhanced concatenated vector to the predicted probability, a dynamic distribution entropy penalty mechanism is introduced to output the predicted probability after confidence correction. :

[0057]

[0058] In the formula, For the first Initial logical score for the type of fault; To adjust the intensity hyperparameter; The global uncertainty scalar is calculated from the modal confidence weight vector.

[0059] In addition, the present invention also discloses a fault early warning system for multimodal data fusion of container generator sets, which is used to execute the fault early warning method for multimodal data fusion of container generator sets as described above. The system includes the following nodes connected in sequence: multimodal physical quantity acquisition array node, digital signal filtering and cleaning node, cross-physical field benchmark alignment node, independent physical feature deconstruction node, two-stage spatiotemporal aggregation operation node, two-branch trend memory and feature intervention node, and adaptive physical boundary defense and control node.

[0060] The dual-stage spatiotemporal aggregation operation node is used to execute the dual-stage fusion operation described in step four of claim 1.

[0061] The dual-branch trend memory and feature intervention node is used for the parallel feature extraction and cross-fusion operation described in step five.

[0062] Compared with the prior art, the present invention has the following beneficial effects:

[0063] This invention overcomes the limitations of isolated information from a single mode by synchronously acquiring and processing multimodal physical quantities such as vibration, temperature, oil, acoustic, and electrical parameters. Using vibration sequences as a benchmark, it employs cross-correlation calculations and nonlinear phase compensation to achieve precise alignment of cross-physical field signals in the time dimension, laying a unified spatiotemporal benchmark for subsequent fusion.

[0064] In the first stage, this invention concatenates multimodal feature vectors according to specific physical logic, preserving the structural information of micro-subsystems. In the second stage, it innovatively introduces a confidence weight vector generated based on modal quality factors, injecting it as a quality mask into temporal features to dynamically suppress distorted data interference. Then, a multi-head self-attention mechanism is used to capture causal relationships across time periods, greatly enhancing the robustness and accuracy of fault feature extraction. A dual-branch temporal network is employed to process fused features in parallel, capturing local impacts and long-term degradation trends separately. Dynamic gating coefficients generated based on real-time operating parameters are used to control the intensity of cross-fusion of information between the two branches, effectively decoupling and utilizing fault representations at different time scales.

[0065] This invention maps the enhanced interactive features to fault probabilities and dynamically adjusts the warning threshold in real time based on the load change rate, operating stage, and ambient temperature, replacing the static rigid threshold and significantly reducing the false alarm rate.

[0066] This invention also compensates for the temporal distortion caused by rotor slip through a time-series network with dynamic dilated convolutions, ensuring complete capture of transient impact waveforms. Furthermore, by introducing a distributed entropy penalty mechanism based on global uncertainty at the probability output, the probability output is forced to flatten under extreme environments such as strong electromagnetic interference, fundamentally eliminating false triggers caused by data distortion and ensuring the absolute reliability of physical protection actions. In addition, the online incremental learning mechanism enables the model to adapt to feature drift caused by long-term equipment service, maintaining the long-term stability of early warning accuracy. Attached Figure Description

[0067] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.

[0068] Figure 1 This is a flowchart of the method described in this invention.

[0069] Figure 2 This is a flowchart of the dynamic dilated convolution process of the present invention.

[0070] Figure 3 This is a flowchart of the distributed entropy penalty subprocess of the present invention. Detailed Implementation

[0071] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the embodiments of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0072] The following is in conjunction with the appendix Figures 1-3 The embodiments of the present invention will be described in detail below.

[0073] Example 1: This example discloses a fault early warning system for multimodal data fusion of container generator sets. The system includes the following nodes connected in sequence: multimodal physical quantity acquisition array node, digital signal filtering and cleaning node, cross-physical field benchmark alignment node, independent physical feature deconstruction node, two-stage spatiotemporal aggregation operation node, two-branch trend memory and feature intervention node, and adaptive physical boundary defense and control node.

[0074] The dual-stage spatiotemporal aggregation operation node is used to execute the dual-stage fusion operation described in step four of claim 1.

[0075] The dual-branch trend memory and feature intervention node is used for the parallel feature extraction and cross-fusion operation described in step five.

[0076] This embodiment also discloses a fault early warning method for multimodal data fusion of container generator sets. The method is executed by a high-performance computing cluster or industrial control computer deployed at the edge of the industrial site. The entire data flow strictly follows a forced closed-loop physical causal chain, ensuring that every continuous waveform collected from the physical field of nature can undergo lossless dimensionality reduction, alignment mapping, and feature decoding, ultimately transforming it into specific command outputs for the servo control mechanism of the container generator set. The method includes the following core technical sub-steps in sequence.

[0077] S1: Sensing and synchronous acquisition of high-dimensional multi-physics state space;

[0078] To overcome information silos of a single dimension, a dense multimodal physical quantity acquisition array node is constructed on the physical surface and interior cavity of the container generator set to synchronously capture raw multimodal data during unit operation. This multimodal physical quantity acquisition array node contains five categories of key sensing elements, each corresponding to a different physical radiation channel.

[0079] The vibration sensor array used to characterize mechanical dynamic degradation consists of piezoelectric triaxial accelerometers orthogonally arranged in space at the main bearing, auxiliary bearing end caps, and high-rigidity base. To fully capture high-frequency transient impact pulses generated during rolling element spalling or early gear meshing, the synchronous sampling frequency is [frequency value missing]. ,in This is the synchronous sampling frequency of the vibration sensor group. This high-frequency threshold fully satisfies the Nyquist sampling theorem's requirement for aliasing-free mapping of the bearing's high-frequency resonant band.

[0080] The acoustic sensor array used to detect internal aerodynamic and combustion knocking noises consists of an omnidirectional electret condenser microphone array suspended inside the sealed container. The microphone array is installed at the acoustic standing wave node positions determined by finite element simulation.

[0081] In practice, the standing wave node is located at a spatial point 1.5 m from the outer shell of the container generator set and 2 m from the interior floor. A microphone array provides omnidirectional sound field coverage within the cabin, with a sampling frequency of [missing information]. ,in The sampling frequency of the acoustic sensor array is used to capture extremely weak ultrasound precursor frequencies above the human hearing threshold.

[0082] The temperature sensor array used to detect thermodynamic anomalies consists of platinum resistance temperature sensors (PT100 class, high temperature measurement accuracy) deeply embedded inside the generator stator winding slots, the bearing housing metal lubrication interface, the main coolant circulation return pipe, and the lubricating oil pan. The temperature physical quantity exhibits significant inertia and time delay characteristics, and its state evolution along the time axis is extremely slow. To match its slowly changing physical nature and avoid data redundancy, the sampling frequency of the temperature sensor group is [missing information]. ,in This is the sampling frequency of the temperature sensor group, which is one temperature data point collected per second.

[0083] The electrical parameter sensor group used to reveal electromagnetic conversion efficiency and air gap eccentricity distortion consists of a high-frequency Hall current transformer and a voltage transformer connected in parallel to the three-phase output bus of the generator. The sampling frequency of the electrical parameters is... ,in The sampling frequency of the electrical parameter sensor group is used to facilitate subsequent analysis of the stator current high-order harmonic modulation phenomenon caused by rotor eccentricity.

[0084] A sensor array for tracking fluid dynamics and frictional wear conditions consists of a diffused silicon pressure transmitter and an immersion oil temperature sensor used in the engine's main lubrication circuit. Its data stream is... The frequency is continuously transmitted back, among which This is the sampling frequency of the oil parameter sensor group.

[0085] The vibration sensor group, acoustic sensor group, temperature sensor group, electrical parameter sensor group, and oil parameter sensor group included in the above-mentioned multimodal physical quantity acquisition array node are aggregated to the central acquisition module through the industrial Ethernet bus and the Precision Time Protocol (IEEE 1588 PTP protocol). The converted and output results are collectively referred to as raw multimodal data.

[0086] S2: Noise reduction and reconstruction of the original signal and multi-dimensional standardization processing;

[0087] Due to the significant differences in the working mechanisms, measurement ranges, and sensitivity to environmental background noise among different types of sensors, directly injecting unmodified heterogeneous data into deep networks will lead to severe collapse of gradient calculations. Therefore, after receiving the raw multimodal data, the digital signal filtering and cleaning node must perform refined filtering and dimensional unification operations at the signal preprocessing node.

[0088] To address the physical challenge of acoustic signals being easily masked by broadband turbulent noise from the cabin cooling fan, this method innovatively incorporates an adaptive filter in the time domain for dynamic background noise suppression. Compared to a fixed-coefficient filter with static frequency band clipping, this adaptive filter can be optimized in real time based on the error surface gradient. The weight coefficient update equation is defined as:

[0089] ;

[0090] In the formula, Discrete time nodes The updated filter weight coefficient vector; Discrete time nodes The current filter weight coefficient vector at the current position; In this embodiment, the convergence factor, which controls the balance between the algorithm's convergence speed and steady-state misalignment, is calibrated as... ; For the first The expected reference signal is the error deviation between the desired reference signal and the actual output of the adaptive filter. The desired reference signal is taken from the background noise signal collected by an auxiliary reference microphone installed inside the container hold, away from the noise source of the container generator set. This auxiliary reference microphone is, for example, installed inside the air inlet, a distance from the outer casing of the container generator set. Place; The first one fed into the adaptive filter The acoustic signal input vector at time step.

[0091] After adaptive dynamic cleaning, to eliminate the uneven data point density caused by differences in Nyquist sampling boundaries across different physical channels, a resampling engine needs to be constructed to align all array channels to the same frequency grid plane. All channel data is then guided into a group of anti-aliasing low-pass digital filters with steep cutoff frequency characteristics. The cutoff frequency of this anti-aliasing low-pass filter is... Stopband attenuation After removing high-frequency aliasing components, forced resampling is performed. A unified benchmark sampling rate.

[0092] Faced with the order-of-magnitude gap in various physical dimensions (e.g., voltage can reach kilovolts, while vibration acceleration amplitude is only in the single digits), the system performs zero-mean standardization on the resampled channel data. This calculation, by eliminating the mean bias of the original distribution and compressing its dispersion, forces all channel data to converge to a dimensionless probability space with zero mean and unit variance. The calculation formula is as follows:

[0093]

[0094] In the formula, For the first Standardized values ​​for each physical measurement channel; This is the real signal value output after resampling of this channel; and Representing the first The channel is pre-calculated based on the historical baseline mean and standard deviation under the fault-free baseline operating conditions of the container generator set. Among them, and The specific values ​​are determined by ensuring the container generator set operates continuously and stably in a fault-free state before leaving the factory. Hours, collect this The data sequences for each channel within an hour are processed, and the arithmetic mean and standard deviation of the acquired sequences for each channel are calculated. The calculation results are then stored in the parameter storage area of ​​the edge computing node. After this process is completed, the digital signal filtering and cleaning node stably outputs the preprocessed data for each channel.

[0095] S3: Time-phase alignment and mesh reconstruction of cross-modal physical signals;

[0096] Because the propagation speeds of longitudinal waves from mechanical vibrations within container generator sets, sound from air, and heat from the fluid boundary layer differ by orders of magnitude, and because different industrial sensors employ varying inherent hardware delays in their hardware bus transmission protocols, directly fusing pre-processed data from each channel can lead to severe feature misalignment and causal inversion. To address the time reference tearing problem in multiphysics observation systems, it is necessary to deploy cross-physics reference alignment nodes to mathematically align the time axes of all physical quantities.

[0097] The cross-physics benchmark alignment node designates the vibration channel output by a piezoelectric triaxial accelerometer mounted at the main bearing of the container generator set as the benchmark anchor point, defined as the reference channel. The main bearing is closest to the core rotational work region, and its vibration response exhibits minimal phase hysteresis in capturing early mechanical excitation forces. The data sequence from the reference channel within a sliding time window is extracted, the length of which is... sampling points (at Corresponding to uniform sampling rate (seconds), sliding step size is The sampling points are used to calculate the cross-correlation function of all other observation nodes (channels to be aligned) of the container generator set relative to the reference channel. The mathematical logic for calculating the cross-correlation function is defined as follows:

[0098]

[0099] In the formula, The cross-correlation function value reflects the time delay between the two physical channel signals. Waveform similarity under the following conditions; This is a mathematical expectation operator used to calculate the statistical average of a sample set within a specified time window. For reference channel at time The signal amplitude; For the channel to be aligned at time The signal amplitude; This represents the time delay in the continuous domain. For discrete sampled sequences, the cross-correlation function described above is estimated using the cross-correlation formula for finite-length sequences in practical calculations:

[0100]

[0101] in This represents the total number of sampling points within the sliding time window. This is a discrete-time delay index, representing the corresponding physical time delay. , To standardize the sampling rate, and by... Within the predefined search area ( ,correspond The cross-correlation sequence is obtained by traversing within the maximum physical delay.

[0102] After establishing the cross-correlation function range matrix, the cross-physics benchmark alignment nodes are searched within the time window boundary by an optimization algorithm to find the extreme points where the cross-correlation function reaches its global maximum value, and the time delay estimate is extracted. :

[0103]

[0104] In the formula, To find the objective function Maximizing the independent variable Operators; This is to extract the optimal time delay estimate.

[0105] Since the different channels are fixed as discrete digital grids after resampling in step S2, the calculated time delay estimate is... These values ​​often fall at non-integer positions between two discrete sampling points. To faithfully reconstruct the translated signal at the sub-pixel level, cubic spline interpolation is used across the physical field reference alignment nodes to perform nonlinear phase compensation on the data of the alignment channels, employing natural boundary conditions (the second derivative is zero at both ends of the interval). The cubic spline interpolation method approximates the real physical waveform by constructing piecewise smooth polynomial curves, setting the discrete sampling point sequence as... Time interval The cubic spline polynomial equation on is expanded as follows:

[0106]

[0107] In the formula, Output the value for the constructed polynomial function; , , , All are intervals The specific coefficient parameters of the polynomial to be solved; It is a time variable; The starting point of the time interval is defined as [reference point]. Solving for the specific coefficients of the polynomial relies on the continuity constraints of the zeroth, first, and second derivatives of the function at the connection points between adjacent intervals, as well as natural boundary conditions. Through rigorous mathematical solving and phase shifting, the time-series waveforms of all sensors are tightly anchored to the same time scale, minimizing the maximum amplitude error of the cubic spline interpolation. Time delay error At this point, the cross-physics benchmark alignment node outputs modal data with a unified time reference.

[0108] S4: Deconstruction of local spatial high-dimensional features of multimodal heterogeneous signals;

[0109] Even after phase alignment and a unified time base, the modal data still appear as lengthy and low-information-density raw waveform sequences. To provide dense information with physical discriminative significance to deep time-series networks, independent physical feature deconstruction nodes must be invoked to perform independent high-dimensional mappings on each physical field signal. Since the fault mechanisms represented by different physical quantities are fundamentally different, the feature extraction algorithm is divided into multiple independent and parallel deconstruction operators based on modal properties.

[0110] Vibration signal feature extraction: Targeting the impact echo phenomenon implicit in the main bearing vibration signal, the independent physical feature deconstruction node focuses on extracting the fourth-order central moment feature, which is most sensitive to transient impacts, from the time-domain statistics. The time-domain kurtosis coefficient is used to quantify the probability distribution density of large-amplitude abnormal pulses in the vibration waveform; the mathematical calculation model is defined as follows:

[0111]

[0112] In the formula, is a dimensionless kurtosis coefficient; This represents the current discrete amplitude of the vibration signal; To calculate the statistical mean amplitude of the vibration signal within the time window, the calculation time window length is... Each sampling point, adjacent window overlap rate When the container generator set is in healthy operating condition, the vibration amplitude follows a Gaussian normal distribution, and the kurtosis coefficient stably converges to a constant. Nearby. When microscopic spalling occurs at the metal contact surface, inducing periodic transient impacts, the tail of the time-domain waveform thickens, and the kurtosis coefficient exhibits an exponential increase.

[0113] Acoustic signal feature extraction: For acoustic signals excited by aerodynamic flow and mechanical friction, the low-frequency components are mostly dominated by the steady-state rotation noise of the fan, while the high-frequency singular acoustic patterns reflecting friction and wear are often masked. Independent physical feature deconstruction nodes are used to introduce cepstral domain analysis techniques to extract Mel-Frequency Cepstral Coefficients (MFCCs). After performing a short-time Fourier transform on the acoustic frame to obtain the power spectrum, it is mapped to a Mel-scale filter bank that simulates the nonlinear perception characteristics of human hearing. The nonlinear mapping boundary equation between the linear true frequency and the Mel-scale frequency is:

[0114]

[0115] In the formula, The Mel frequency value obtained after mapping, in units of ; The linear absolute physical frequency value of the acoustic signal, in units of . Subsequently, the logarithm of the energy within the filter bank is taken, and a discrete cosine transform is performed to decouple the correlation between frequency components. The low-order coefficients are then extracted to form the acoustic signature matrix characterizing the local resonance properties of the sound field.

[0116] Temperature signal feature extraction: Extraction of independent physical feature deconstruction nodes from temperature sequences collected by a group of temperature sensors. Rate of temperature rise within the minute sliding window and temperature fluctuation variance :

[0117]

[0118] In the formula, The rate of temperature rise, in units of ; for Temperature value at a given second, in units of ; for Temperature value at any given time, in units of ;constant The number of sampling points corresponding to the sliding window (because the sampling rate is ).

[0119]

[0120] In the formula, This represents the temperature fluctuation variance, in units of... ; for The total number of sampling points within minutes; For the first Temperature values ​​at each sampling point; This is the arithmetic mean of the temperature series within the sliding window.

[0121] Electrical Parameter Signal Feature Extraction: For the three-phase current and voltage sequences collected by the electrical parameter sensor array, three-phase imbalance is extracted by deconstructing nodes based on independent physical features. Total Harmonic Distortion as well as , , Harmonic amplitude at twice the rotational frequency:

[0122]

[0123] In the formula, Three-phase imbalance, unit: ; , , These are the effective values ​​of the three-phase currents, in units of... ; This is a function to find the maximum value. The function is for finding the minimum value; This is a function for taking the arithmetic mean.

[0124]

[0125] In the formula, Total harmonic distortion (THD) is expressed in units of 1000 m / s. ; The highest harmonic order; For the first The amplitude of the second harmonic voltage, in units of ; The amplitude of the fundamental voltage, in units of .

[0126] Oil parameter signal feature extraction: For oil parameter signals characterizing fluid dynamics, a single drop in oil pressure or rise in oil temperature is insufficient to diagnose the lubrication failure boundary. Independent physical feature deconstruction nodes focus on capturing the thermodynamic hysteresis coupling relationship between oil temperature trends and oil pressure fluctuations, calculating the Pearson correlation coefficient between the two (observation window length). Second):

[0127]

[0128] In the formula, The Pearson correlation coefficient is the coupling coefficient between oil temperature and oil pressure. The time series distribution of lubricating oil temperature, in units of ; The time series distribution of lubricating oil pressure, in units of ; To observe the average oil temperature inside the window, the unit is... ; To observe the statistical average oil pressure inside the window, the unit is... ; This represents the standard deviation of the oil temperature series, in units of... ; This represents the standard deviation of the oil pressure series, in units of... As the dynamic viscosity of lubricating oil decreases due to high-temperature carbonization, the original linear relationship between hot and pressure will be broken, and the Pearson correlation coefficient will deviate significantly.

[0129] After each physical signal branch undergoes specialized mathematical deconstruction in its corresponding field, the extracted scattered feature values ​​are repackaged according to a preset physical index, and the independent physical feature deconstruction nodes uniformly output the feature vectors of each mode.

[0130] S5: Modal uncertainty assessment and dynamic confidence perception; In complex industrial environments, some sensors may experience severe distortion in their output modal feature vectors within a specific time slice due to intense electromagnetic interference, probe loosening, or instantaneous packet loss in the signal transmission bus. To prevent degraded data from poisoning subsequent deep fusion networks, the system incorporates modal uncertainty assessment logic after the independent physical feature deconstruction node and before the two-stage spatiotemporal aggregation operation node, dynamically calculating the relative information quality of each modality at the current moment.

[0131] Extract the modal feature vectors of each dimension within the sliding window's history, and construct the covariance matrix of the corresponding physical mode. Specifically, for the ... Take the nearest modality. Feature vectors within a time window Each feature vector has a dimension of Calculate its covariance matrix:

[0132]

[0133] In the formula, For the first The covariance matrix of each mode; The number of feature vectors within the sliding window; For the first Within the first time window Feature vectors of each modality; This is the sample mean vector; This is the matrix transpose operator.

[0134] Subsequently, eigenvalue orthogonal decomposition is performed on each covariance matrix to calculate the set of spectral eigenvalues ​​representing the energy distribution along the principal directions of the feature space. A modal quality factor based on a Shannon information entropy variant is introduced to quantify the information abundance and data discrete stability of each physical field feature.

[0135]

[0136] In the formula, For the calculated first Modal quality factor for a specific mode; For the first The absolute dimension of the feature vector for extracting features from a specific modality; For the first The first specific modal covariance matrix obtained after orthogonal decomposition is... The magnitude of the principal component eigenvalues.

[0137] When the sensor channel is completely flooded with white noise or experiences a constant dead value, the eigenvalue energy distribution tends to become extremely flat or drops sharply to zero, causing a drastic decrease in the modal quality factor. The system concatenates all parallel modal quality factors to obtain a dimension... The vector (in this embodiment, electrical parameters are divided into two independent modes, current and voltage, along with six physical modes including vibration, acoustics, temperature, and oil) is input into a single-layer perceptron to perform a linear projection transformation. The single-layer perceptron contains a weight matrix. and a bias vector Output After function activation, a modal confidence weight vector with strictly normalized boundaries (each element value is between 0 and 1) is obtained. Modal confidence weight vector As a key quality mask, it will be directly fed into the second stage of fusion in the subsequent two-stage spatiotemporal aggregation operation nodes.

[0138] S6: Early Spatial Stitching and Normalization of Cross-Physical Features (First-Stage Fusion); The first-stage fusion node in the two-stage spatiotemporal aggregation operation receives the modal feature vectors output by the independent physical feature deconstruction nodes and performs a one-time spatial aggregation of complementary information within the modality. Since the feature vectors extracted from different physical modes differ significantly in numerical distribution range and Euclidean space span, direct stitching would cause the deep network to be dominated by high-energy features during backpropagation, resulting in gradient shift. The first-stage fusion node separately feeds the modal feature vectors extracted at the same time into the L2 normalization operator for spatial projection. The L2 normalization operation forces the feature vectors of different physical fields to be mapped onto a high-dimensional unit hypersphere, ensuring that features of all dimensions have the same initial gradient contribution weight in subsequent network layers. The normalization mathematical equation is defined as:

[0139]

[0140] In the formula, This represents the feature vector after normalized projection; The feature vectors representing each modality of the original input; The L2 norm scalar value of the original input vector in multidimensional Euclidean space is represented by the following formula: ,in The dimension of the feature vector. The eigenvector of the eigenvector Each component.

[0141] After spatial projection is completed, the first-stage fusion nodes strictly follow the transmission order of the physical coupling of the container generator set: namely, the specific arrangement rules of vibration characteristics, acoustic characteristics, temperature characteristics, current characteristics, voltage characteristics, and oil pressure characteristics; and perform connection operations on all normalized feature vectors along the feature channel dimension. This stage outputs an early fusion feature vector that is strictly aligned according to physical logic. Early fusion of feature vectors It preserves the original structural information of each physical subsystem at the microscopic level to the greatest extent possible.

[0142] S7: Dynamic weighting and global mapping of time-dimensional evolution sequences (second-stage fusion); the failure evolution of container generator sets is not an isolated time slice, but a dynamic physical process with time memory effect.

[0143] In order to establish causal relationships between fluctuations of different physical variables on the time axis, the second-stage fusion node in the two-stage spatiotemporal aggregation operation node adopts a multi-head self-attention mechanism to perform dynamic weighted fusion of cross-time-series correlation information.

[0144] Set the sequence data block of the input second-stage fusion node as ,in Representing the The early fusion feature vector output by the first-stage fusion node at each time step; The absolute time step length representing the sliding time window (set to continuous). (sampling points at each time point); Represents the total absolute feature dimension of the early fused feature vector; express A real vector space.

[0145] To compensate for the positional blind spots of the self-attention mechanism in the time series and to inject data quality awareness information, the system will use the modal confidence weight vector output by the modal uncertainty assessment node. (Its original dimension is equal to the number of modes) (Through a learnable linear projection matrix) Mapped to the same dimension as the earlier fused feature vectors The expanded confidence vector is obtained. Then, the expanded confidence vector is combined with the early fused feature vector at each time step in the sequence. Perform element-wise addition to inject supplementary location-encoded information, including data quality awareness capabilities, into the input sequence. ,in The feature vector after injecting confidence information is directly used as input for subsequent self-attention calculations. express OK The space of real matrix columns.

[0146] For the self-attention mechanism Each computational head is independent, and the system allocates three sets of independent linear projection weight matrices. The input sequence is projected onto the query subspace, key subspace, and value subspace respectively using matrix multiplication to generate the query matrix. Key matrix Sum matrix Among them, the projection dimension of each attention head. Set as , For the total number of attention heads, This represents the total absolute feature dimension of the early fused feature vectors. To ensure divisibility, the actual value is taken as... ; express OK The column is a real matrix space. Then, within the subspace, the attention similarity score between different time steps is calculated, the th... The equation for the output feature representation of an attention head is:

[0147]

[0148] In the formula, For the first A weighted matrix of the outputs of each attention head; The normalized exponential function is defined as follows: ; For the first A query matrix for each head; For the first The transpose of the key matrix of the head; The projection dimension within the independent calculation head; For the first The value matrix of each attention head. The system sets the total number of attention heads. This drives different attention heads to focus on distinct physical dynamic patterns within a time series, such as high-frequency impacts and low-frequency trends. Ultimately, this is achieved through an output global projection matrix. The results of all parallel attention heads are concatenated and mapped back to the original feature space dimension:

[0149]

[0150] In the formula, This is the final output of the multi-head self-attention mechanism; This is the vector concatenation operator; for The output matrix of each attention head; The output global projection matrix is ​​used to map the concatenation result of multiple attention heads back to the original feature space dimension. After the second-stage fusion node completes the joint weighting of the spatiotemporal dimensions, it outputs the fused feature vector to the next level link.

[0151] S8: Dual-stream parallel temporal deconstruction of local distortion and long-term degradation; dual-branch trend memory and feature intervention node deployment of dual-stream parallel temporal modeling nodes to receive fused feature vectors. Unit operation includes both high-frequency, high-energy impacts such as sudden bearing surface fractures and long-term slow degradation caused by insulation aging or progressive gear wear. To overcome the limitations of a single network scale, this method designs two physically parallel feature extraction network branches.

[0152] The first timing branch is dedicated to capturing high-frequency partial fault pulse patterns, configured to be... A temporal convolutional network consisting of cascaded one-dimensional convolutional layers and global average pooling layers.

[0153] Wherein, the kernel size is Step size is The padding method is "same," and each layer is followed by batch normalization and ReLU activation. The one-dimensional convolution operator slides across the time dimension to accurately extract sudden, sharp waveform features from the time series.

[0154]

[0155] In the formula, Represents the first sequential convolutional network Layer at time Output of local features of the settlement; Represents the receptive field length of the convolution kernel; The first element in the convolution kernel weight parameter matrix obtained by the network learning is the... One element; For the first Layer at time Input features; For the first Layer bias scalar; This represents the batch normalization operation that eliminates internal covariance bias. The linear rectifier function that introduces nonlinear activation characteristics is defined as follows: After the first time-series branch completes its calculation, it outputs the first time-series output sequence. .

[0156] The second temporal branch addresses the degradation trend of memories with long-span dependencies, configured as follows: A bidirectional LSTM network is constructed by stacking layers of bidirectional Long Short-Term Memory (LSTM) networks, with each hidden layer having a dimension of [missing information]. Dropout rate This network unit controls the retention and discarding of physical information along the timeline through a highly sophisticated three-gating mechanism:

[0157]

[0158]

[0159]

[0160]

[0161]

[0162] In the formula, The input gate activation vector; The activation vector for the forget gate; The output gate activation vector; This represents the cell state vector; The hidden state vector; This is the input vector at the current moment; This is the hidden state vector from the previous time step; This is the cell state vector from the previous time step; The input weight matrix; This is a cyclic weight matrix; It is the bias vector; The Sigmoid activation function is defined as follows: ; Let hyperbolic tangent activation function be defined as follows: ; This is the Hadamard product (element-by-element multiplication) operator. The hidden state information from forward and backward propagation is orthogonally concatenated at the end of the node. The second temporal branch outputs the second temporal hidden state sequence, carrying the direction of the system's macroscopic evolution. .

[0163] S9: Cross-branch cross-attention intervention based on operating condition perception vectors; container generator sets exhibit distinctly different fault-dominant characteristics during variable load startup and steady-state cruise phases. Dual-branch trend memory and feature intervention nodes deploy cross-branch cross-attention nodes to receive the first time-series output sequence. With the second time-series hidden state sequence An information flow interoperability mechanism is established between the two independent computing branches. The system reads the load rate data transmitted back in real time by the container generator set. and ambient temperature data High-dimensional condition-aware vectors are generated through fully connected network encoding. This fully connected network consists of two hidden layers: the input layer receives the normalized load rate value. and ambient temperature value (unit The first hidden layer contains... The second hidden layer contains neurons with ReLU activation function; There are *n* neurons, with ReLU activation function; the output layer has an output dimension of *n*. Working condition sensing vector .

[0164] The condition-aware vector is used to extrapolate the intensity of feature fusion intervention at the current moment in real time, i.e., to calculate the dynamic gating coefficient. :

[0165]

[0166] In the formula, To deduce the dynamic gating coefficients output; It is a Sigmoid activation function with nonlinear truncation properties; The working condition mapping weight parameter matrix; This is a high-dimensional operating condition sensing vector generated by fully connected encoding based on load rate data and ambient temperature data; This represents the bias vector mapped to the operating conditions. The gating coefficients are also included. It directly affects the cross-branch attention interaction computation equation. Taking the absorption of long-term global features from the first time-series output sequence as an example:

[0167]

[0168] In the formula, This represents the output feature matrix of the current layer of a temporal convolutional network that enhances interaction after absorbing global information over a long span. This represents the original first-order output sequence matrix; The element-wise Hadamard product operator for matrices; This represents the cross-attention weight matrix generated by mapping the second temporal hidden state sequence. The specific generation method is as follows: First, the second time-series hidden state sequence is generated. Through a linear layer Projection, making its feature dimensions equal to The feature dimensions are the same, then the attention weight matrix is ​​calculated:

[0169]

[0170] in This is the attention scaling factor, with a value of [value missing]. The feature dimensions. It is a learnable weight matrix. Through deep computation, the features of the two branches are combined and encapsulated, and the output is a concatenated vector with enhanced interaction.

[0171] S10: Decoding mapping from high-dimensional topology space to fault probability space; the enhanced concatenated vector is pushed to the prediction probability mapping sub-node in the adaptive physical boundary defense and control node. This sub-node is configured with two neurons, with the following numbers of nodes: and The network employs hidden fully connected layers and Dropout regularization layers to prevent overfitting. The fully connected layers non-linearly combine high-level features scattered across different dimensions and define classification surfaces. The final network layers utilize a softmax normalized exponential function to perform probability distribution transformation.

[0172]

[0173] In the formula, Given input Under the condition that the sample belongs to the first Posterior probabilities of each fault category; Represents mapping to the first Specific fault categories (covering bearing raceway spalling, stator insulation breakdown, lubricating oil deterioration, cooling system blockage, fuel injector blockage, and air gap eccentricity) The magnitude of the nonnormalized logic score of a known pattern; Represents the absolute total number of fault categories monitored by default; It is a natural exponential function; the denominator is... This is a normalization factor that ensures the sum of the output probabilities for all classes is equal to 1. The predicted probability mapping child node eventually outputs a normalized value. arrive The predicted probability matrix for each fault category.

[0174] S11: Multi-dimensional adaptive boundary construction and physical protection closed-loop linkage; isolated probability data must be converted into mechanical actuation logic. The adaptive physical boundary defense and control node triggers hierarchical early warning and control linkage sub-nodes, receiving the predicted probability matrix for each fault category and equipment operating parameter streams. The system constructs a three-dimensional joint environmental perception equation based on the real-time load change rate. , Operational phase characteristic functions and transient ambient temperature Calculate the boundary of the three-level early warning defense line with dynamic elastic floating. Among them, the characteristic function of the operational phase... Defined as discretized value: Cold start stage (before unit startup) (minutes) takes the value Steady-state cruise phase (unit operation exceeding...) minutes and load fluctuation less than The value can be... Hot shutdown phase (after unit shutdown) The value within minutes) is taken as The adaptive threshold calculation model is defined as follows:

[0175]

[0176] In the formula, For a moment The first dynamic settlement Level-adaptive early warning boundary threshold These correspond to Level 1, Level 2, and Level 3 warnings, respectively. The offline calibration of the first Level absolute physical reference thresholds, respectively , , ; The load change rate coefficient of the main power grid for container generator sets is read in real time; To characterize the discretized operation phase feature function of container generator sets in the cold start, steady cruise or hot shutdown phases; The absolute temperature is the actual measured temperature of the external environment, in units of . ; The reference temperature is calibrated in the standard operating condition laboratory. This is the absolute value operator; , , All of these are balancing adjustment coefficients used to control the weighting of various operating condition variables.

[0177] Based on a microsecond-level high-speed polling mechanism, the system drives the numerical comparison logic between the predicted probability matrix of each fault category and the adaptive dynamic threshold:

[0178] When the predicted probability exceeds the first warning line (threshold) When this occurs, the system backend generates a first-level warning signal and initiates a communication push action to the remote control hub;

[0179] When the predicted probability breaks through the second early warning line (threshold) When this occurs, the system generates a second-level warning signal and simultaneously injects a physical load reduction control frame through the programmable logic controller (PLC) interface to limit the fuel injection pulse width and excitation current output of the container generator set, thereby locking the absolute output power at a certain percentage of the rated peak power. Within the safety envelope;

[0180] When the predicted probability reaches the third early warning line (threshold) When this occurs, the system generates the highest-level Level 3 warning signal, which activates the emergency shutdown protection relay contacts within milliseconds, cutting off the fuel supply solenoid valve and locking the air intake duct, physically isolating the equipment from irreversible catastrophic disaster. Thus, the container generator set achieves full-stack physical closed-loop protection from multi-dimensional physical field observation to hardware terminal action execution.

[0181] The model pre-training method provides supplementary explanations of all neural network parameters (including the multi-head attention weight matrix) involved in Example 1 above. The weights of convolutional kernels in temporal convolutional networks, the gate weights of bidirectional LSTMs, and the weights of fully connected layers are obtained through the following pre-training process:

[0182] collection Taiwan same model Operating data of container generator sets under the following conditions:

[0183] Each unit collects data. Hours, covering scenarios such as cold start, steady-state cruise, hot shutdown, and variable load;

[0184] Preset fault simulation data: Simulates bearing raceway spalling, stator insulation breakdown, lubricating oil deterioration, cooling system blockage, fuel injector blockage, and air gap eccentricity, respectively. Data is collected for each type of fault per unit. The defects were caused by replacing components with known man-made faults or injecting deteriorating media on healthy units. For example, bearing raceway spalling was caused by replacing bearings with micro-drilling pits from wire cutting; stator insulation breakdown was caused by creating localized weak points in the winding insulation layer; and lubricating oil deterioration was caused by adding a certain proportion of oxidized and deteriorated lubricating oil. The total raw data size was approximately [missing information]. Hours were manually labeled according to fault type to form the training set (approximately 10% of the total). ), validation set (approximately) ), test set (approximately) The samples are evenly divided according to the time length ratio to ensure a balanced distribution of unit operating conditions and fault types within the three datasets.

[0185] The training configuration is as follows: The Adam optimizer is used, with an initial learning rate of... Batch size The loss function is Categorical Cross-Entropy, and the training... Round. Early stopping strategy: Validation set loss is continuous. Training stopped when the wheel did not decrease. The final validation set accuracy reached [percentage missing]. Test set accuracy .

[0186] After training, all network weights except those in the top fully connected layer of the prediction probability mapping node are fixed as read-only parameters, along with... , Standardized parameters are burned together into the parameter storage area of ​​the edge computing node for real-time inference.

[0187] On-site testing verified the effectiveness of the test. tower Container generator sets are being developed One month of on-site testing (each unit operates continuously, with a total runtime of approximately [number] hours). (hours), comparing the method in this embodiment with the traditional single-modal threshold monitoring method:

[0188] This embodiment provides the average early warning time for bearing-related failures. Hours, traditional methods average Hours, the advance warning time has increased times;

[0189] The false alarm rate (number of false alerts / total operating hours) of this method is lower than The false alarm rate of traditional methods is approximately The false alarm rate has decreased. ;

[0190] This method successfully alerted all. The actual failure (including bearing peeling) Insulation deterioration Starting, lubricating oil deterioration Start-up, cooling blockage fuel injector blockage Air gap eccentricity (Starting from), no missed reports, while traditional methods have missed reports. rise.

[0191] Example 2: This example is a further optimization based on Example 1. In this example, over the long service life of container generator sets (which can last for several years), changes in mechanical clearances can cause a slow drift in the underlying physical characteristics, leading to a decline in the accuracy of model warnings. This example provides a deep online incremental fine-tuning architecture.

[0192] The system continuously records past events at the edge computing core. The predicted probability output distribution under healthy working conditions within a day. A relative entropy (Kullback-Leibler Divergence, or KL divergence) measurement logic is introduced to monitor the real-time predicted probability mean distribution. Initial baseline distribution of equipment after production line The difference in KL divergence between them.

[0193] When the judgment condition Upon establishment, the system determines that a significant mechanical parameter drift has occurred and automatically triggers an online incremental learning stream on the edge side.

[0194] in, Here is the formula for calculating KL divergence. Incremental training samples are from the past... Operating data deemed healthy by the system within a day (without fault labels) is used only for fine-tuning the model to accommodate feature drift.

[0195] To prevent catastrophic forgetting, the incremental learning flow locks all convolutional kernel parameters of the underlying feature extraction network, only opening the fully connected layer weights of the top-level prediction probability mapping nodes for small gradient updates. This small gradient update uses a stochastic gradient descent optimizer with a learning rate set to [value missing]. Batch size set to Only iterative training after each trigger One epoch.

[0196] Simultaneously, the system periodically calculates the maximum Youden index point in the Receiver Operating Characteristic Curve (ROC) space to find a new equilibrium point, and employs... The inertia weight coefficients are smoothly updated using existing thresholds. The Yoden index is defined as follows: Sensitivity is the true class rate, and specificity is the true negative class rate.

[0197] Example 3: This example is based on the same principle as Example 1. The difference is that in this example, the edge node computing power of some older model container generator sets is insufficient and cannot bear the huge matrix dot product attention calculation pressure.

[0198] This embodiment provides a lightweight dynamic gating alternative avoidance network for cross-branch feature fusion. Upon receiving the first time-series output sequence (labeled as a variable matrix)... ) and the second temporal hidden state sequence (labeled as a variable matrix) Afterwards, the original self-attention similarity calculation system was abandoned, and a learnable gating function model composed of fully connected parameters was constructed to calculate the element-wise distribution parameters between the features of the two types of sequences. :

[0199]

[0200] In the formula, The parameter vector is distributed element by element; This is the weight matrix for a lightweight fully connected network. It is the bias vector; This indicates that two sequences are concatenated along the feature dimension; The activation function is Sigmoid. The weight parameters of the lightweight gating are transferred from the full model in Example 1 using a knowledge distillation method: the soft labels output by the full model (teacher network) guide the training of the student network (lightweight gating model), and the distillation temperature... Distillation loss weight is Subsequently, a complementary superposition operation is performed on the temporal features generated by the two different network branches to generate a fused output feature matrix. :

[0201]

[0202] In the formula, To fuse and output the feature matrix; The parameter vector is distributed element by element; The first time-series output sequence matrix; This is the second temporal hidden state sequence matrix; This is the Hadamard product operator. This lightweight dimensionality reduction avoidance scheme greatly reduces the number of multiply-accumulate operations and the memory footprint within the edge control chip, providing a robust and feasible backup protection implementation path for hardware bases with extreme resource constraints.

[0203] Example 4: This example reconstructs the structural underlying computation based on the dual-branch trend memory and feature intervention nodes in Example 1. When container generator sets encounter extreme nonlinear conditions such as sudden full loads or grid short-circuit crossings, the reverse drag torque from the main grid forces the mechanical rotor shaft system to undergo transient acceleration and deceleration oscillations. In this physical process, the macroscopic mechanical characteristics are directly manifested as rapid slippage of rotational speed. The mechanical excitation points attached to the microscopic damage interface of the main bearing or gearbox generate transient high-frequency impact pulses each time they mesh or roll over the defect area, which exhibit non-stationary stretching or compression variations on the absolute scale of the time series.

[0204] In the original first-order branch, the fixed temporal convolutional network relies on continuous memory reads between adjacent physical sampling points. This constant local receptive field span is highly susceptible to truncation or missed detection of early weak distortion waveforms due to spatiotemporal misalignment of the sampling window when processing non-stationary time series distorted by physical slip. The original physical logic of cross-physical field feature splicing will be completely destroyed. This embodiment abandons the conventional fixed-step traversal mechanism and forcibly implants dynamic hole sensing logic into the bottom-level operators of each layer of the temporal convolutional network. The system utilizes the rigid coupling characteristics of the electromagnetic field to reshape the boundary range of the bottom-level temporal receptive field by tracking the frequency shift law of the stator current.

[0205] The modal data with a unified time base is continuously pushed into the computation stream.

[0206] The system extracts a sequence of electrical parameters characterizing the electromagnetic state and imports the three-phase current waveforms into a digital zero-crossing detection algorithm module. This algorithm module operates within an extremely narrow... Within a millisecond sliding microscopic time window, the alternating node where the current amplitude crosses the zero potential axis is captured. Two consecutive positive and negative amplitude switches are used as the basis for determining the effective zero crossing point, and the time interval between adjacent zero crossing points is locked.

[0207] The edge computing unit performs high-frequency measurement of the fundamental frequency based on this time interval, and converts and outputs the current instantaneous rotor rotation frequency of the container generator set. ,in For a moment The instantaneous rotor rotation frequency, in units of The physical change in rotor speed is the driving force behind the periodic oscillations of the entire fluid piping network and mechanical rotor shaft system. Instantaneous rotor rotation frequency. As a core variable, it is directly injected into the deep computational graph structure of the dual-branch trend memory and feature intervention nodes.

[0208] Arranged within the first time sequence branch Cascaded temporal convolutional networks (CCNNs) at different depths undertake multi-dimensional feature deconstruction tasks. The lower layers focus on extracting transient spikes with extremely short durations and high energy; the deeper layers are responsible for piecing together the fragmented waves from the lower layers into a wideband envelope spanning multiple mechanical cycles. The system assigns an independent receptive field scaling and attenuation constant to each computational layer. In the parameter storage area of ​​the edge controller, the first... Layer to the first corresponding to the layer The parameter matrix is ​​hard-coded and fixed as .

[0209] Within each microsecond-level execution cycle triggered by forward inference, the system calls the space expansion mapping equation to calculate the dynamic void ratio scalar. Offline calibrated rotor reference rotational frequency under rated operating conditions. Fixed as .

[0210]

[0211] In the formula, For the first Layer at time The dynamic void ratio scalar; For the first Layer-specific receptive field scaling attenuation constant; The reference rotational frequency of the rotor under rated operating conditions, calibrated offline; For a moment The instantaneous rotor rotation frequency; This is the floor operation operator; To find the maximum value function. During the steady-state cruise phase, the rotor of the container generator set operates at its rated speed. Approaching At this point, the dynamic void ratio scalar of each layer... It basically maintains the default compact sampling mode. When the unit encounters a sudden increase in heavy load, the instantaneous rotor rotation frequency... Dropped to within 100 milliseconds At that time, the division terms in the equation surge. Deep networks rely on higher... Weight configuration, and the resulting dynamic void ratio scalar It exhibits non-linear amplification and jumping.

[0212] In edge-side memory addressing operations, when the physical convolutional kernel extracts slices of the underlying input sequence, its memory addressing pointer will be adjusted according to the dynamic dilatation rate scalar. The value is read in a skip-style manner within the circular data buffer. This code-level memory span stretching perfectly compensates for the widening of the excitation point time interval caused by the decrease in rotational speed in physical space.

[0213] The dynamic void ratio scalar obtained by calculation Reconstructing the one-dimensional convolution operator:

[0214]

[0215] In the formula, For a moment Temporal convolutional networks Local feature output of layer settlement; A slice of the underlying input sequence extracted across physical sampling points based on a dynamic porosity scalar; is the absolute receptive field length of the physical convolution kernel; and The first convolutional kernel weight parameter matrix generated by the network learning is respectively the first... Each element and bias term scalar; Batch normalization to eliminate internal covariance bias; A linear rectified function with nonlinear activation characteristics is introduced. Discrete feature aggregation is performed across physical sampling points in the time dimension, eliminating the high-frequency energy loss problem associated with traditional time-series downsampling techniques.

[0216] The time-domain distribution matrix wrapped by the first time-series output sequence always maintains a close engagement with the actual physical excitation phase angle of the unit.

[0217] The system testing department is at a rated power of A real-ship-class simulation experiment of sudden load application and unloading was performed on the hydraulic loading rig of the container generator set. On the surface of the outer ring metal layer of the raceway of the generator's rear main bearing, a section with a width of [missing information] was artificially laser-etched. Microscopic spalling defects. The steady-state load factor of the main power grid of the unit has been maintained at the rated power for a long time. During the test, a high-impedance dummy load array was periodically applied, forcing the output power to... Within an extremely short time window, the load surged to full capacity. The data acquisition system intercepted data from continuous processes using the static convolutional architecture of Example 1 and the dynamic dilated architecture of this example. The internal feature tensors under the submutation shock slice. Table 1 records the quantitative indicators of the deconstruction capability of the dual-path parallel architecture.

[0218] Table 1. Comparison of temporal network capture performance under sudden load shock;

[0219]

[0220] In Example 1, the network experienced severe sampling window drift during the rotor slip phase, truncating more than half of the effective impact tailwaves. The low-energy, incomplete waveforms failed to activate deep neurons, leading to weight dispersion in the cross-attention calculation. This example, relying on an adaptive receptive field dynamic stretching mechanism based on the fundamental frequency, maintained extremely high waveform envelope integrity, stably delivering high-dimensional physical representations to the interaction nodes, and completely filling the monitoring blind spot under transient conditions.

[0221] Example 5: This example is basically the same as Example 1, except that in this example, in complex physical scenarios such as the engine room of a large ocean-going vessel or the power supply side of a supercomputing center, the hard switching action of a heavy-duty nonlinear frequency converter group or the corona discharge of the high-voltage bus will radiate highly destructive broadband electromagnetic pulses into space. Under the dual pressure of strong radiation and severe mechanical broadband vibration, the physical communication bus will experience instantaneous packet loss and bit flipping. The tiny analog signals of the front-end sensor array are easily and completely masked by Gaussian white noise surges.

[0222] When some or all observation channels fall into a blind zone due to physical limitations, uncontrolled white noise will be fed into the multi-layer neural network. The high-dimensional spatial feature stream extracted by the fully connected layer will completely fall into a state of disordered diffusion. The normalized exponential function (softmax) accessed at the terminal network layer in conventional architectures exhibits mathematical properties as an extreme rigid classifier. This function only focuses on the relative difference between each input component and lacks the ability to determine the absolute physical limits of the purity of the information in the input space itself. Under the stimulation of pure white noise, the multiplication and addition operations inside the classifier will inevitably randomly and forcibly increase the order of magnitude of the non-normalized logical score of a certain preset failure mode. Once the distorted and drastically fluctuating forgery probability breaks through the physical load reduction defense line at the back end, it will trigger a devastating large-scale service interruption.

[0223] Between the enhanced concatenated vector and the probabilistic classifier, the system intercepts the feature stream and forcibly enters the dynamic distribution entropy penalty module. The triggering criterion for the penalty mechanism originates from the modal confidence weight vector defined in Example 1.

[0224] The system scheduling calculation kernel extracts the absolute confidence projection values ​​of all channels within the current time slice. Perform multi-channel inverse weighted averaging to calculate a global uncertainty scalar that reflects the degree of chaos in the fused feature space. :

[0225]

[0226] In the formula, For a moment The global uncertainty is a scalar, dimensionless; The total number of physical modes (covering vibration, acoustics, temperature, voltage, current, and hydraulic pressure) that are simultaneously activated for container generator sets. The first element in the modal confidence weight vector The absolute confidence projection value of a specific mode is a dimensionless dimensionless representation of the effective physical information abundance remaining after a single physical mode passes through front-end adaptive filtering and zero-mean compression.

[0227] The concatenated vector, enhanced by interaction, is fed into the deep hidden units of the probabilistic classifier. The number of nodes in the two neurons are respectively... and The fully connected layer completes the non-linear dimension mapping, and the output contains The initial nonnormalized logical score matrix for each predicted node value. Before performing the final probability space transformation, the underlying logic calls the nonlinear temperature scaling adjustment equation:

[0228]

[0229] In the formula, For the given input after reliability correction Under the condition that the sample belongs to the first Posterior probabilities of each fault category; and Represents mapping to the first The and the first The magnitude of the initial nonnormalized logic score for each specific fault category; This represents the absolute total number of pre-set monitored fault categories; The constraint calibration hyperparameters for controlling the intensity of nonlinear temperature scaling are hard-coded and solidified on the edge side; For a moment The global uncertainty scalar; It is a natural exponential function. When the unit is in a healthy operating state without external interference, the signals in each channel are clear, and the global uncertainty scalar... Attenuation and approximation At this point, the temperature term in the denominator of the equation degenerates into a constant. The confidence-corrected prediction probability distribution maintains a steep gradient curve that matches the magnitude of the initial non-normalized logic score, ensuring that the deep network maintains an extremely keen sense for hidden faults in micro-evolution.

[0230] When a high-power radio frequency radiation interference event occurs in external space, the transmission noise floor of the front-end piezoelectric accelerometer and platinum resistance temperature sensor increases by an order of magnitude. The energy distribution of each dimension within the modal confidence weight vector collapses instantaneously. Global uncertainty scalar. A steep step response occurs, rapidly climbing to... The above extreme danger zone. At this point, the denominator of the temperature regulation equation expands sharply to... The magnitude of the initial nonnormalized logic score from the fully connected layer output. All were harvested by the indiscriminate depth division suppression operator.

[0231] Due to the forced reduction of the exponential input, the outputs of the various probability classification nodes, which might have previously sputtered randomly, were forced to converge.

[0232] The predicted probability matrices for each fault category exhibit extremely flat and smooth probability distribution characteristics. Even under the most severe chaotic feature flow, the predicted probability of any single fault mode, after penalty, cannot exceed [the specified threshold]. The physical red line. Based on the three-tiered defense rule defined by adaptive physical boundary defense and control nodes, the absolute physical baseline threshold of the first early warning defense line is... The highly flattened dynamic probability boundary is strictly and rigidly confined within the underlying security envelope. This severs the catastrophic path at the algorithm's root, preventing environmental interference from triggering misleading physical control protection commands from the edge controller to the downstream programmable logic controller.

[0233] To assess the engineering immunity of the distributed entropy penalty mechanism under harsh electromagnetic environments, the research team additionally deployed a broadband electrostatic discharge generator and a radio frequency radiation interference antenna array inside a sealed containerized test chamber. The tested equipment was in a hot shutdown phase, and fuel supply had been cut off. The test source subjected the generator set's surface sensor cluster to a long-term... Hours of continuous high-frequency, high-energy pulse burst bombardment. Table 2 details the core defense effectiveness data of the credibility penalty system.

[0234] Table 2 Comparison of False Trigger Protection Suppression Tests under Strong Electromagnetic Interference Environment

[0235]

[0236] As shown in the table above, the probability space temperature scaling adjustment method and the multi-level hard early warning boundary achieve deep physical synergy. The deterministic dynamic collapse of the prediction probability boundary during disturbances completely absorbs and dissolves the system-level risk exposure caused by the distortion of underlying data.

[0237] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0238] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A fault early warning method based on multimodal data fusion for container generator sets, characterized in that, Includes the following steps: Step 1: Synchronously acquire raw multimodal data of the container generator set under operating conditions, and perform filtering, resampling and dimensional normalization preprocessing on the raw multimodal data, and output the preprocessed data of each channel; wherein, the raw multimodal data includes at least vibration sequence, temperature sequence, oil parameter sequence, acoustic signal sequence and electrical parameter sequence; Step 2: Using the vibration sequence in each preprocessed channel data as the reference anchor point, calculate the time delay of the remaining sequences relative to the reference anchor point, and perform nonlinear phase compensation alignment on the remaining sequences based on the time delay to output modal data with a unified time reference. Step 3: For each modal data of the unified time reference, perform time-domain and / or frequency-domain feature extraction independently according to the physical modal type, and output the feature vector of each modality respectively; Step 4: Calculate the modal confidence weight vector corresponding to each modal feature vector and perform two-stage fusion: The first stage of fusion performs L2 normalization on each modal feature vector at the same time and then splices them in a preset physical order to obtain the early fusion feature vector. The second stage of fusion first projects the modality confidence weight vector and adds it to the early fusion feature vector. Then, it uses a multi-head self-attention mechanism to perform temporal weighting on the time series obtained after addition and outputs the fusion feature vector. Step 5: The first and second time-series branches are used in parallel to receive the fused feature vector, capturing local impact features and long-term degradation trends respectively; at the same time, dynamic gating coefficients are generated based on the real-time operating parameters of the container generator set, and the dynamic gating coefficients are used to control the feature cross-fusion between the two branches, outputting the interactively enhanced spliced ​​vector. Step six: Map the enhanced splicing vector to the predicted probability of the fault mode, and dynamically adjust the warning threshold according to the real-time operating parameters; when the predicted probability exceeds the warning threshold, trigger the corresponding control and protection action.

2. The fault early warning method for multimodal data fusion of container generator sets according to claim 1, characterized in that, Step one involves filtering the original multimodal data, specifically including: using an adaptive filter to clean the background noise of the acoustic signal sequence, wherein the weight coefficient update equation of the adaptive filter is: In the formula, This is the updated filter weight coefficient vector; This is the current filter weight coefficient vector; The convergence factor; The error deviation between the desired reference signal and the actual output of the adaptive filter is given by the desired reference signal, which is taken from the background noise signal collected by an auxiliary reference microphone located far from the noise source. is the acoustic signal input vector.

3. The fault early warning method for multimodal data fusion of container generator sets according to claim 1, characterized in that, The dimensional normalization mentioned in step one specifically includes: performing zero-mean standardization on the resampled channel data: In the formula, For the first The standardized values ​​for each channel; This is the real signal value after resampling of this channel; and The first The historical baseline mean and standard deviation of each channel are pre-calculated under the fault-free baseline operating condition of the generator set.

4. The fault early warning method for multimodal data fusion of container generator sets according to claim 1, characterized in that, Step two involves calculating the time delay of the remaining sequences relative to the reference anchor point, specifically including: calculating the cross-correlation function between the sequence to be aligned and the reference anchor point sequence. in, This represents the total number of sampling points within the sliding time window. This is a discrete-time delay index, representing the corresponding physical time delay. , To ensure a uniform sampling rate, it is necessary to iterate through... The cross-correlation sequence was obtained.

5. The fault early warning method for multimodal data fusion of container generator sets according to claim 4, characterized in that, Step two involves performing nonlinear phase compensation alignment based on the time delay, specifically including: retrieving the time delay estimate that maximizes the cross-correlation function value. ; A cubic spline interpolation polynomial is used to perform nonlinear phase compensation on the sequence to be aligned. This polynomial operates within the interval... Expanded above: In the formula, The result is the interpolation result. These are the polynomial coefficients; It is a time variable; This is the starting point of the time interval.

6. The fault early warning method for multimodal data fusion of container generator sets according to claim 1, characterized in that, The feature extraction described in step three specifically includes: Extracting the temporal kurtosis coefficient of the vibration sequence : In the formula, The amplitude of the vibration. To calculate the mean amplitude of vibrations within the window, the Mel frequency cepstral coefficients of the acoustic signal sequence are extracted. The Mel frequency mapping equation is as follows: In the formula, The Mel frequency value, The linear physical frequency value; the rate of temperature rise of the extracted temperature sequence. and temperature fluctuation variance Extract the three-phase imbalance from the electrical parameter sequence. Total Harmonic Distortion Extract the Pearson correlation coefficient between the oil temperature sequence and the oil pressure sequence. : In the formula, , These are the sequences for lubricating oil temperature and pressure, respectively. , Its mean; , Its standard deviation.

7. The fault early warning method for multimodal data fusion of container generator sets according to claim 1, characterized in that, Step four involves calculating the modal confidence weight vector, specifically including: constructing and decomposing the covariance matrix of the historical eigenvectors for each modality to obtain the eigenvalues. ; Calculate the modal quality factor : In the formula, For the first The feature vector dimension of each modality; All modalities' The model is concatenated and mapped to the output modal confidence weight vector through a single-layer perceptron and a sigmoid activation function.

8. The fault early warning method for multimodal data fusion of container generator sets according to claim 1, characterized in that, The L2 normalization of the first stage fusion in step four is as follows: In the formula, The original feature vector, The normalized feature vectors, for The L2 norm.

9. The fault early warning method for multimodal data fusion of container generator sets according to claim 1, characterized in that, In step five, the first temporal branch in parallel is a temporal convolutional network composed of multiple one-dimensional convolutional layers; the second temporal branch is a recurrent network composed of multiple bidirectional long short-term memory networks.

10. A fault early warning system for container generator sets based on multimodal data fusion, used to execute the fault early warning method for container generator sets based on multimodal data fusion as described in any one of claims 1 to 9, characterized in that, The system comprises, in sequence, the following nodes connected in sequence: a multimodal physical quantity acquisition array node, a digital signal filtering and cleaning node, a cross-physical field benchmark alignment node, an independent physical feature deconstruction node, a two-stage spatiotemporal aggregation operation node, a two-branch trend memory and feature intervention node, and an adaptive physical boundary defense and control node. The dual-stage spatiotemporal aggregation operation node is used to execute the dual-stage fusion operation described in step four of claim 1. The dual-branch trend memory and feature intervention node is used for the parallel feature extraction and cross-fusion operation described in step five.