A cloud-edge collaborative diagnosis method and system based on multi-source asynchronous data
Patent Information
- Application Number
- CN202611073635.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]现有技术存在以下不足:首先,变电站各类监测传感器的采样频率和触发机制各不相同,现有技术中将多源数据同步融合的方法,导致各参量原有的时间尺度信息丢失或失真;其次,现有的边缘计算架构中,边缘节点通常仅实现数据采集、协议转换和阈值判断功能,由于数据传输延迟和带宽限制,难以对突发性故障作出及时响应;此外,现有技术的诊断模型经部署后无法利用云端数据进行持续优化,云端的分析模型缺乏对边缘诊断结果的有效验证调整,边缘端与云端独立工作,没有形成真正的协同闭环流程
[0053]本发明通过构建多通道连续时间递归网络,将各通道的时间常数分别设置为与油中溶解气体、振动信号、放电脉冲信号的采样周期相对应,使得网络模型能够以各自合适的时间尺度连续、异步地处理多源数据,既能完整保留放电脉冲的瞬态信息,又能捕捉气体产出的慢变趋势;其次利用脉冲神经网络对放电脉冲信号进行稀疏事件检测,有效平衡快尺度通道的响应及时性与计算开销,增强对局部放电模式的表征能力;将异步状态序列进行跨尺度注意力计算,使慢尺度通道的状态能够在诊断时刻从历史放电事件记忆中检索最相关的放电模式,从而建立跨物理场的因果关联,提升对潜伏性故障早期微弱征兆的识别能力;此外,在不确定度超过预设阈值时触发数字孪生模型仿真和增量训练,既解决了因设备老化导致的仿真偏差问题,又避免频繁无效的模型更新,节省计算与通信资源,实现模型的稳定优化。
Smart Images

Figure CN122839282A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent operation and maintenance technology for substations, specifically to a cloud-edge collaborative diagnostic method and system based on multi-source asynchronous data. Background Technology
[0002] As a critical node in the power system, the safe and stable operation of substations directly affects the reliability and economic efficiency of the power grid. With the vigorous promotion of smart grid construction, a large number of intelligent monitoring devices have been deployed in substations. The monitoring data from these devices exhibits characteristics of large volume, diverse types, and varying sampling periods, resulting in multi-source asynchronous processing. Among power transformer faults, winding and insulation problems are the main sources of failure, and the types of faults are becoming increasingly concealed and chain-like. Traditional operation and maintenance diagnostic methods are no longer sufficient to meet the requirements of real-time performance and accuracy.
[0003] Currently, various sensors are widely used in field deployments to monitor various status parameters during equipment operation in real time, determine the health status of the equipment, and identify potential faults. Common monitoring methods for transformers include dissolved gas analysis in oil, partial discharge detection, and vibration monitoring. These methods assess the insulation aging state, discharge characteristics caused by insulation defects, and vibration characteristics of the mechanical structure of oil-immersed equipment, respectively. These methods reflect the equipment's operating status from different physical fields such as heat, force, and electricity, relying on a single threshold for judgment. Later, to avoid the limitations of a single parameter, some technical solutions integrated multiple monitoring parameters, often using fixed time window alignment or interpolation resampling methods for time synchronization.
[0004] With the development of the information age, edge computing technology has been gradually applied to the field of substation equipment condition monitoring, aiming to solve problems such as difficulty in accessing heterogeneous sensors, large data processing latency, and high data transmission bandwidth occupancy. Existing solutions have designed a cloud-edge collaborative service architecture for substation equipment condition monitoring, realizing signal processing, risk assessment, and data preprocessing on the edge side. However, most solutions are still limited to data acquisition and simple feature extraction, and no adaptive deep diagnostic models have been deployed on the edge side.
[0005] The existing technology has the following shortcomings: First, the sampling frequencies and triggering mechanisms of various monitoring sensors in substations are different. The existing method of synchronously fusing multi-source data leads to the loss or distortion of the original time-scale information of each parameter. Second, in the existing edge computing architecture, edge nodes usually only implement data acquisition, protocol conversion and threshold judgment functions. Due to data transmission delay and bandwidth limitations, it is difficult to respond to sudden faults in a timely manner. In addition, the diagnostic models of the existing technology cannot be continuously optimized using cloud data after deployment. The cloud analysis model lacks effective verification and adjustment of edge diagnostic results. The edge and the cloud work independently and do not form a true collaborative closed-loop process. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention aims to provide a cloud-edge collaborative diagnostic method and system based on multi-source asynchronous data. First, multi-source asynchronous data from the target device is acquired through edge computing nodes. A multi-channel continuous-time recursive network is constructed based on the sampling period of the multi-source asynchronous data to obtain the asynchronous state sequence corresponding to each multi-source asynchronous data. Then, cross-scale attention calculation is performed to obtain fused correlation features. Next, a diagnostic classifier generates fault diagnosis results and uncertainties. When the uncertainty exceeds a preset threshold, updated parameters are obtained through a digital twin model and incremental training, completing the system optimization and update. This invention achieves deep collaborative diagnosis of multi-source asynchronous data and provides confidence assessment and model self-updating mechanisms, significantly improving the accuracy and reliability of substation equipment fault diagnosis.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] In a first aspect, the present invention provides a cloud-edge collaborative diagnostic method based on multi-source asynchronous data, comprising:
[0009] Multi-source asynchronous data of the target device is acquired through edge computing nodes. The multi-source asynchronous data includes dissolved gas data in oil, vibration signals, and discharge pulse signals.
[0010] A multi-channel continuous-time recursive network is constructed, wherein the time constant of each channel in the multi-channel continuous-time recursive network corresponds one-to-one with the sampling period of each multi-source asynchronous data.
[0011] Based on the discharge pulse signal, sparse event detection is performed. When a valid discharge event is detected, a sparse feature vector is obtained. The dissolved gas data in the oil, the vibration signal, and the sparse feature vector are input into the multi-channel continuous-time recursive network to obtain the corresponding asynchronous state sequences.
[0012] Cross-scale attention computation is performed based on the asynchronous state sequence to obtain fused correlation features; based on the fused correlation features, a fault diagnosis result and uncertainty are generated through a diagnostic classifier;
[0013] When the uncertainty exceeds a preset threshold, simulation is performed using the cloud-based digital twin model of the target device, and then incremental training is used to obtain updated parameters to update the multi-channel continuous-time recursive network and the diagnostic classifier.
[0014] Preferably, the specific steps for sparse event detection based on the discharge pulse signal include:
[0015] The discharge pulse signal is input into a preset pulse neural network. The pulse neural network includes a threshold comparison unit and a residual coding unit. When the amplitude of the discharge pulse signal reaches a preset amplitude threshold, the threshold comparison unit generates a sparse feature vector.
[0016] The residual coding unit is configured with a sliding time window to statistically analyze the mean amplitude and mean pulse interval of the discharge pulse signal, and to map the mean amplitude and mean pulse interval into a subthreshold residual vector through a preset embedding matrix.
[0017] The subthreshold residual vector is used as a bias term and superimposed on the hidden state update equation of the corresponding channel in the multi-channel continuous-time recursive network.
[0018] Preferably, the input to the multi-channel continuous-time recursive network also includes a cross-channel state real-time update process, specifically including the following steps:
[0019] The sparse feature vectors are simultaneously input into a pre-trained mapping network to generate an incremental vector representing the impact on the slow-scale channel.
[0020] The influence increment vector is added element-wise to the hidden state vector of the slow-scale channel of the multi-channel continuous-time recursive network to obtain the updated hidden state.
[0021] Preferably, after obtaining the asynchronous state sequence, the method further includes correcting the time constant of the multi-channel continuous-time recursive network, specifically including:
[0022] The initial time constant of each channel is determined based on the sampling period of each multi-source asynchronous data, and the state change rate of each channel is calculated based on the historical asynchronous state sequence.
[0023] The first cross-source difference vector is calculated based on the state change rate of the first channel and the state change rate of the second channel; the second cross-source difference vector is calculated based on the state change rate of the second channel and the state change rate of the third channel.
[0024] The first cross-source difference vector and the second cross-source difference vector are concatenated and input into the fully connected layer, and the output time constant compensation vector is obtained. The initial time constant of each channel is multiplied element by element by the time constant compensation vector to obtain the corrected time constant.
[0025] Preferably, the specific steps for performing cross-scale attention computation based on the asynchronous state sequence include:
[0026] Obtain the current diagnostic request time, and during the current asynchronous state sequence output process, the updated hidden state vector and update timestamp for each channel;
[0027] For each channel, using the hidden state vector as the initial value, the diagnostic request time is subtracted from the update timestamp to obtain the integral duration;
[0028] Based on the initial value and the integration duration, the differential equations of the neurons in each channel are numerically integrated under the condition of no external input, and the final state of the integration is used as the alignment state vector corresponding to each channel.
[0029] Preferably, the specific steps for performing cross-scale attention computation based on the asynchronous state sequence further include:
[0030] For each channel, the alignment state vector is input into a preset confidence estimation network, and the corresponding modal confidence is output.
[0031] The alignment state vector is multiplied by the modal confidence to obtain the query vector for each channel, and the historical discharge mode features are used as the key vector and value vector.
[0032] Based on the query vector, key vector, and value vector of all channels, the fused association features are obtained by scaling dot product attention calculation.
[0033] Preferably, the specific steps for generating fault diagnosis results and uncertainties through a diagnostic classifier include:
[0034] The fused and associated features are input into the joint diagnostic classifier, which outputs the first probability distribution of each fault type. The fault type with the highest probability in the first probability distribution is taken as the fault diagnosis result.
[0035] Each aligned feature vector is input into the corresponding single-modal classifier to obtain its own second probability distribution. The Jensen-Shannon divergence between each pair of the second probability distributions is calculated, and the average value is then used to obtain the modal conflict degree.
[0036] The uncertainty is obtained by adding the information entropy of the first probability distribution to the modal conflict degree.
[0037] Preferably, the specific steps for obtaining the updated parameters include:
[0038] When the uncertainty exceeds a preset threshold, the average value of the historical asynchronous state sequence output by each channel is obtained as the state statistical feature of the target device;
[0039] Using the state statistical features as the observation vector and the preset aging parameters as the variables to be estimated, the Bayesian inference method is used to perform maximum a posteriori estimation to obtain the estimated values of the aging parameters.
[0040] The digital twin model is corrected based on the estimated aging parameters, and then the corrected digital twin model is simulated using the fault type as the boundary condition to generate simulation samples.
[0041] Preferably, the specific steps for obtaining the updated parameters further include:
[0042] The simulation samples are input into the cloud-based diagnostic model, and simulation fusion correlation features are obtained through asynchronous state sequence calculation and cross-scale attention calculation, which serve as the simulation feature set.
[0043] Extract discharge mode features and corresponding historical diagnostic labels, combine the discharge mode features and the simulation feature set into a hybrid feature set, and combine the historical diagnostic labels and simulation fault type labels into a label set;
[0044] The mean squared gradient of each weight parameter of the cloud diagnostic model on the mixed feature set is calculated as the importance weight; the cross-entropy loss of the label set and the product of the squared change of each weight parameter and the corresponding importance weight are summed to form the loss function.
[0045] With the goal of minimizing the loss function, the cloud-based diagnostic model is trained, and all weight parameters after training are extracted and sent to the edge computing node as update parameters.
[0046] Secondly, the present invention provides a cloud-edge collaborative diagnostic system based on multi-source asynchronous data, comprising:
[0047] The data acquisition module is used to acquire multi-source asynchronous data of the target device through edge computing nodes. The multi-source asynchronous data includes dissolved gas data in oil, vibration signals, and discharge pulse signals.
[0048] The network construction module is used to construct a multi-channel continuous-time recursive network, wherein the time constant of each channel in the multi-channel continuous-time recursive network corresponds one-to-one with the sampling period of each multi-source asynchronous data.
[0049] The sequence generation module is used to perform sparse event detection based on the discharge pulse signal. When a valid discharge event is detected, a sparse feature vector is obtained. The dissolved gas data in the oil, the vibration signal, and the sparse feature vector are input into the multi-channel continuous-time recursive network to obtain the corresponding asynchronous state sequences.
[0050] The diagnostic output module is used to perform cross-scale attention calculation based on the asynchronous state sequence to obtain fused correlation features; based on the fused correlation features, a diagnostic classifier is used to generate fault diagnosis results and uncertainties.
[0051] The model update module is used to update the multi-channel continuous-time recurrent network and the diagnostic classifier by simulating the target device's cloud-based digital twin model and obtaining update parameters through incremental training when the uncertainty exceeds a preset threshold.
[0052] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0053] This invention constructs a multi-channel continuous-time recursive network, setting the time constant of each channel to correspond to the sampling period of dissolved gas in oil, vibration signals, and discharge pulse signals. This allows the network model to process multi-source data continuously and asynchronously at appropriate time scales, preserving the transient information of discharge pulses while capturing the slow-changing trend of gas production. Secondly, it utilizes a spiking neural network for sparse event detection of discharge pulse signals, effectively balancing the timeliness of fast-scale channels with computational overhead, enhancing the representation of partial discharge modes. Furthermore, it performs cross-scale attention calculations on asynchronous state sequences, enabling the slow-scale channels to retrieve the most relevant discharge modes from historical discharge event memories at diagnostic time, thereby establishing causal relationships across physical fields and improving the ability to identify early, weak signs of latent faults. Finally, it triggers digital twin model simulation and incremental training when uncertainty exceeds a preset threshold, solving the simulation bias problem caused by equipment aging, avoiding frequent and ineffective model updates, saving computational and communication resources, and achieving stable model optimization. Attached Figure Description
[0054] Figure 1 A flowchart of a cloud-edge collaborative diagnostic method based on multi-source asynchronous data provided for this embodiment;
[0055] Figure 2 A sparse event detection flowchart is provided for this embodiment;
[0056] Figure 3 This is a schematic diagram of the structure of a cloud-edge collaborative diagnostic system based on multi-source asynchronous data, provided for this embodiment. Detailed Implementation
[0057] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof.
[0058] Example 1:
[0059] like Figure 1 As shown, this embodiment provides a cloud-edge collaborative diagnostic method based on multi-source asynchronous data, including:
[0060] S1. Obtain multi-source asynchronous data of the target device through edge computing nodes. The multi-source asynchronous data includes dissolved gas data in oil, vibration signals, and discharge pulse signals.
[0061] S2. Construct a multi-channel continuous-time recursive network, wherein the time constant of each channel in the multi-channel continuous-time recursive network corresponds one-to-one with the sampling period of each multi-source asynchronous data;
[0062] S3. Based on the discharge pulse signal, perform sparse event detection. When a valid discharge event is detected, obtain a sparse feature vector. Input the dissolved gas data in the oil, the vibration signal, and the sparse feature vector into the multi-channel continuous-time recursive network to obtain the corresponding asynchronous state sequences.
[0063] S4. Perform cross-scale attention calculation based on the asynchronous state sequence to obtain fused correlation features; generate fault diagnosis results and uncertainties through a diagnostic classifier based on the fused correlation features;
[0064] S5. When the uncertainty exceeds a preset threshold, simulation is performed using the cloud-based digital twin model of the target device, and then incremental training is used to obtain updated parameters to update the multi-channel continuous-time recursive network and the diagnostic classifier.
[0065] This embodiment provides a cloud-edge collaborative diagnostic method based on multi-source asynchronous data, which is applied to the intelligent operation and maintenance system of substations. It involves deploying edge computing nodes at the substation site and a cloud server located in the central control center.
[0066] Regarding step S1:
[0067] In one alternative implementation, the target device mainly refers to a transformer. The edge computing node is deployed in the secondary equipment room within the substation and is connected to various sensors installed near the transformer via fiber optic or wireless sensor networks. The sensors include a dissolved gas monitoring device in oil, a vibration acceleration sensor, and a partial discharge ultra-high frequency sensor, which are used to sense the chemical state of the transformer's oil-paper insulation system, the mechanical state of the winding core, and the discharge state of the insulation structure, respectively.
[0068] Specifically, edge computing nodes acquire three data streams monitored by various sensors through data acquisition and communication interfaces. The first stream is dissolved gas data in the oil. A chromatograph quantitatively analyzes characteristic gases separated from the transformer oil, generating concentration values for each component, including hydrogen, ethane, and carbon dioxide. Since the generation and diffusion of dissolved gases in the oil are relatively slow physicochemical processes, the sampling period is on the order of hours, forming a slowly updated concentration time series. This reflects the degree of oil-paper pyrolysis caused by overheating, discharge, or insulation aging inside the transformer. Its slow change characteristics give the carried fault information cumulative and trend-like characteristics. The second stream is vibration signals, continuously sampled at a higher frequency by accelerometers adsorbed on the surface of the oil tank or at fixed locations on the transformer. The sampling period is on the order of milliseconds. This can capture the mechanical and dynamic responses caused by winding loosening, core loosening, or cooling system abnormalities. Its change frequency is more rapid than the dissolved gas data in the oil, reflecting instantaneous fluctuations in equipment operating status on a second-level timescale. The third channel is the discharge pulse signal, captured by a high-frequency sensor from the high-frequency electromagnetic waves excited by partial discharge. Due to the randomness and suddenness of the discharge pulse occurrence, the time interval between two effective discharge events can vary drastically from microseconds to hours, exhibiting sparsity and non-uniformity in time. Using a microsecond-level sampling period, the arrival time, amplitude, and phase distribution of the pulse signal can be accurately recorded. The three data channels differ significantly in sampling mechanism, time scale, and update rhythm: dissolved gas data in oil is a slowly varying sequence, vibration signals are a rapidly varying sequence, and discharge pulse signals are event-driven burst pulse sequences, collectively constituting multi-source asynchronous data. After acquiring the multi-source asynchronous data, the edge computing node performs timestamp marking and protocol parsing, but does not perform forced time point alignment or interpolation resampling. This preserves the original asynchronous temporal characteristics of each modality's data, fully retaining the inherent speed differences in fault response to different physical fields and the causal precedence relationships between events, providing distortion-free raw input for subsequent construction of a multi-timescale fusion diagnostic model.
[0069] Regarding step S2:
[0070] If a traditional fixed-step recurrent neural network is used to uniformly process the aforementioned multi-source asynchronous data, it will lead to severe distortion of fast-scale vibration signals and discharge pulse signals due to downsampling, or cause meaningless high-frequency filling of slow-scale dissolved gas data in oil, thus introducing false trends. Furthermore, the static time discretization processing method cannot adapt to the autonomous evolution of each physical process within the sampling interval, causing the causal relationship of early fault symptoms to be severed at different scales. Therefore, to fundamentally solve the problem of collaborative modeling of asynchronous data at multiple time scales, this step constructs a multi-channel continuous-time recurrent network. The continuous-time characteristic of this network lies in the fact that the dynamic evolution of network neurons is described by ordinary differential equations, and the network state can be solved at any time, no longer constrained by discrete fixed time steps, and possessing the ability to process irregular sampling and event-driven data.
[0071] In one optional implementation, the output of the multi-channel continuous-time recurrent network is an asynchronous state sequence generated by three channels. The core of each channel is a continuous-time recurrent neuron, and the evolution of the hidden state vector of each channel follows an ordinary differential equation: , where h is the hidden state vector of the channel, Let be the derivative of the hidden state with respect to time, and let x be the input vector of the network. The time constant of the channel. The state self-feedback weight matrix, Let b be the input weight matrix and b be the bias vector. For non-linear activation functions, the hyperbolic tangent function or the sigmoid function is usually chosen. , b is determined through random initialization. This ordinary differential equation ensures the continuous-time dynamics of the network, and the hidden state vector can be solved by numerical integration at any time. In this embodiment, the hidden state vector represents the device state of the physical field corresponding to each channel at the corresponding time. The dissolved gas data in the oil corresponds to the thermal field, the vibration signal corresponds to the stress field, and the discharge pulse signal corresponds to the electric field.
[0072] In this embodiment, the multi-channel continuous-time recursive network establishes independent dynamic channels for the three types of multi-source asynchronous data. Each channel maintains its own unique internal state sequence. During network construction, the time constant of each channel determines the rate at which the neuron's state autonomously decays or remains unchanged under conditions of no external input. This is equivalent to the channel's memory duration of historical information and its response speed to new input. Specifically, the time constant of each channel is set to correspond one-to-one with the sampling period of the multi-source asynchronous data, because the sampling period essentially reflects the minimum change interval that the corresponding data can be effectively observed by the sensor, and is also a direct manifestation of the signal bandwidth. Specifically, setting the time constant of the first channel to a slow-scale value corresponding to the sampling period of dissolved gas data in the oil allows its state to decay gradually between two gas samplings, smoothly preserving the long-term memory of chemical faults. Setting the time constant of the second channel to a medium-scale value corresponding to the sampling period of vibration signals allows its state to update at an appropriate rate between adjacent vibration signal sampling points, tracking mechanical response without amplifying instantaneous noise. Setting the time constant of the third channel to a fast-scale value corresponding to the sampling period of discharge pulse signals allows it to quickly return to a resting state between two effective discharge events, thus maintaining high sensitivity to the next sudden pulse. In this embodiment, the time constant of each channel is obtained by multiplying each sampling period by a different preset scaling factor. The preset scaling factor is determined by searching the network hyperparameters on the historical validation set, with the search objective being to minimize the mean square error between the predicted value of each channel's state during periods without input and the actual gradual change trend of the sensor.
[0073] Regarding step S3:
[0074] Since only a very small proportion of discharge pulse signals actually carry effective insulation defect information, and the vast majority of sampling points record background noise or zero values, the multi-channel continuous-time recursive network constructed according to the above implementation method will cause the fast-scale channel to be continuously updated at high frequency if the high-frequency sampled discharge pulse signal is directly input into the third channel, consuming a large amount of computational resources. At the same time, a large amount of useless noise input will drown out the true discharge characteristics, causing the network state to be driven by noise and deviate from the correct representation of the insulation state. If a simple threshold binary classification is directly performed, weak discharge signals that have not reached the threshold but have deterioration trend information will be lost, causing the network to lose its ability to perceive the gradual change process of the insulation state. Therefore, it is necessary to compress information and filter noise from the original discharge pulse signal. This step introduces a sparse event detection mechanism for the discharge pulse signal.
[0075] A preferred embodiment, such as Figure 2 As shown, the specific steps for sparse event detection based on the discharge pulse signal include:
[0076] S301. The discharge pulse signal is input into a preset pulse neural network. The pulse neural network includes a threshold comparison unit and a residual coding unit. When the amplitude of the discharge pulse signal reaches a preset amplitude threshold, the threshold comparison unit generates a sparse feature vector.
[0077] S302. The residual coding unit is configured with a sliding time window to statistically analyze the average amplitude and average pulse interval of the discharge pulse signal, and to map the average amplitude and average pulse interval into a subthreshold residual vector through a preset embedding matrix.
[0078] S303. The subthreshold residual vector is used as a bias term and superimposed on the hidden state update equation of the corresponding channel in the multi-channel continuous-time recursive network.
[0079] Specifically, in this embodiment, the pulse neural network includes two branches: a threshold comparison unit and a residual encoding unit. The threshold comparison unit makes real-time judgments on the amplitude of the input discharge pulse signal. When the amplitude of the discharge pulse signal exceeds a preset amplitude threshold, an output pulse is emitted, indicating that a valid discharge event has been detected. At this time, a sparse feature vector is output. The preset amplitude threshold is set based on the statistical amplitude distribution of background noise under normal operating conditions of the transformer. The sparse feature vector is composed of the phase feature and normalized amplitude feature of the valid discharge event. The phase feature refers to the position of the discharge pulse signal on the power frequency voltage cycle, represented by the angle relative to the zero-crossing point of the power frequency. The amplitude feature is the logarithmic transformation value of the ratio of the discharge pulse peak value to the background noise baseline, used to characterize the discharge intensity. Subsequently, the dissolved gas data in the oil is fed into the first channel of the multi-channel continuous-time recursive network, the vibration signal is fed into the second channel, and the sparse feature vector is fed into the third channel. Under the driving force of their respective inputs and the action of their own ordinary differential equations, each channel updates its hidden state vector. The updated hidden state vectors are arranged in chronological order to obtain the corresponding asynchronous state sequences, including the first asynchronous state sequence, the second asynchronous state sequence, and the third asynchronous state sequence, thus preserving the evolution trajectory of each multi-source asynchronous data on a non-uniform time scale.
[0080] Furthermore, in actual substation operation, discharge pulses slowly rise in the early stages of a fault, reflecting the evolution of the fault from its nascent stage to its outbreak stage. Partial discharge typically evolves from weak, intermittent subthreshold discharges to high-energy discharges. While subthreshold discharges do not trigger alarms, they contain early signs of insulation degradation and cannot be discarded directly. Therefore, a residual coding unit is introduced to process discharge pulse signals that have not reached the preset amplitude threshold and may carry fault trend information. Specifically, the sliding time window length of the residual coding network can be set to include 5-10 power frequency cycles to ensure sufficient pulse statistical samples. First, the mean amplitude and mean pulse arrival time interval of all discharge pulse signals within the sliding time window are statistically analyzed. These two statistics reflect the noise floor intensity and frequency of subthreshold discharge activity within the sliding time window. Then, combined with a preset embedding matrix (essentially a linear transformation layer), the mean amplitude and mean pulse interval are mapped to a linear transformation layer. The subthreshold residual vector with the same dimension as the hidden state of the continuous-time recurrent network is embedded in a matrix obtained through supervised learning using statistics of pulseless events labeled with normal periods and known early fault periods in a historical database. The resulting subthreshold residual vector retains key information that distinguishes different degradation states and is superimposed on the hidden state update equation of the third channel of the multi-channel continuous-time recurrent network. That is, the bias vector in the ordinary differential equation of the third channel is added element by element to the subthreshold residual vector, so that even when there is no effective discharge event to trigger the state update, the bias input reflecting the change in the activity of the current discharge noise floor will be continuously received, thereby realizing the perception of the trend of weak discharge activity.
[0081] According to the above implementation, in the operation of the multi-channel continuous-time recursive network, the third channel triggers state updates through effective discharge events and subthreshold residual vectors, while the first and second channels update their states according to the multi-source asynchronous data sampling period, maintaining the old state for a long time. However, when an effective discharge event occurs, it can cause chemical damage to the transformer oil-paper insulation, thus affecting subsequent gas production, and can also excite transient mechanical responses in the tank or windings. If the first and second channels wait for the next sampled data to arrive before updating their states, there is a delay between this effective discharge event and its multiphysics consequences, indicating that in the subsequent cross-scale attention fusion, it is impossible to correctly establish the instantaneous causal relationship between the effective discharge event and the vibration response and gas trend.
[0082] In a preferred embodiment, while inputting the multi-channel continuous-time recursive network, a cross-channel state real-time update process is also included, specifically comprising the following steps:
[0083] S311. Simultaneously input the sparse feature vector into the pre-trained mapping network to generate an incremental vector representing the influence of the slow-scale channel.
[0084] S312. Add the influence increment vector to the hidden state vector of the slow-scale channel of the multi-channel continuous-time recursive network element by element to obtain the updated hidden state.
[0085] Specifically, in this embodiment, when the sparse feature vector is fed into the third channel, it is also fed into a pre-trained mapping network. This mapping network is a feedforward neural network structure, with the sparse feature vector as input and a two-dimensional influence increment vector as output. The slow-scale channel is the first channel and the second channel, and the influence increment vector contains the hidden state increments of the first channel and the second channel. During the training process of the mapping network in this embodiment, a time period in historical operating data where an effective discharge event occurred and subsequent correlation changes were observed in vibration signal and dissolved gas data monitoring in oil is selected. The sparse feature vector of this effective discharge event is used as input, and the changes in the hidden states of the first and second channels before and after this event are used as supervision labels. The mapping network parameters are trained through regression, enabling it to learn the ability to predict the amplitude and direction of the impact on the slow-scale channel state from the discharge pulse features. Further, the hidden state increment corresponding to the first channel in the influence increment vector is added element-wise to the current hidden state vector of the first channel to obtain the updated hidden state of the first channel; the hidden state increment corresponding to the second channel in the influence increment vector is added element-wise to the current hidden state vector of the second channel to obtain the updated hidden state of the second channel. This instantaneous update process is executed immediately upon the arrival of a valid discharge event, without waiting for the next sampling time of the first and second channels. This allows the state of the slow-scale channel to respond instantly to the expected impact of the discharge time on the thermal and mechanical physical fields, achieving causal synchronization of the three channels. It should be noted that this process does not change the time constant update rhythm of each channel itself. The first and second channels continue to update their states according to the original time constants when the next sampling period of the corresponding multi-source asynchronous data arrives. This implementation utilizes the influence increment vector to modify the internal state of the slow-scale channel, simulating the physical impact caused by partial discharge.
[0086] In the process of constructing the multi-channel continuous-time recursive network described above, the time constant of each channel is statically set based on the sampling period of each multi-source asynchronous data. However, as transformer equipment ages, the dielectric loss of the insulation material increases, the winding stiffness decreases, and the thermal conductivity of the oil decreases, which leads to changes in the evolution rate of the fault physical process. For example, the partial discharge repetition rate and amplitude growth slope of aging transformers are higher than those of new equipment, and the gas generation rate accelerates after the oil-paper insulation deteriorates. In this case, the set static time constant will gradually deviate from the actual physical response rhythm of the equipment. Therefore, the time constant needs to be dynamically adjusted according to the changes in equipment status.
[0087] In a preferred embodiment, after obtaining the asynchronous state sequence, the method further includes correcting the time constant of the multi-channel continuous-time recursive network, specifically including:
[0088] S321. Determine the initial time constant of each channel based on the sampling period of each multi-source asynchronous data, and calculate the state change rate of each channel according to the historical asynchronous state sequence;
[0089] S322. Calculate the first cross-source difference vector based on the state change rate of the first channel and the state change rate of the second channel; calculate the second cross-source difference vector based on the state change rate of the second channel and the state change rate of the third channel.
[0090] S323. Concatenate the first cross-source difference vector and the second cross-source difference vector, input them into the fully connected layer, and output the time constant compensation vector; multiply the initial time constant of each channel element by the time constant compensation vector to obtain the corrected time constant.
[0091] Specifically, according to the above implementation method, the sampling period of each multi-source asynchronous data is multiplied by the corresponding scaling factor to obtain the initial time constant of each channel; secondly, the edge computing node calculates the average value of the element-wise difference between the current output asynchronous state sequence and the previous output asynchronous state sequence of each channel from the asynchronous state sequence continuously output by the multi-channel continuous-time recursive network, as the state change rate of each channel; subsequently, the absolute value of the difference between the state change rate of the first channel and the state change rate of the second channel is calculated as the first cross-source difference vector, and the absolute value of the difference between the state change rate of the second channel and the state change rate of the third channel is calculated as the second cross-source difference vector; the first cross-source difference vector and the second cross-source difference vector are concatenated and input into the fully connected layer to output the time constant compensation vector; wherein, the fully connected layer is set with 1 or 2 hidden layers, adopts the ReLU activation function, the output dimension is the same as the number of channels, and the weight parameters are obtained through end-to-end training to realize the output of the time constant scaling factor required by each channel according to the cross-modal speed difference. Specifically, each element in the time constant compensation vector corresponds to a time constant scaling factor for each of the three channels. The initial time constant of each channel is multiplied by its corresponding scaling factor to obtain the corrected time constant for each channel. This corrected time constant is then applied to the corresponding neuron state ordinary differential equation for the next state update. It is understood that this implementation does not calculate the absolute value of the difference between the state change rates of the first and third channels. This is because the two differ significantly in time scale and do not have an instantaneous physical coupling relationship. Forcibly adding cross-source difference vectors for the first and third channels would introduce physically meaningless noise components. The coupling between the first and second channels, and between the second and third channels, is the main causal path for fault propagation. Calculating cross-source difference vectors between adjacent scales can effectively capture the rate differences in fault propagation between multiple physics fields.
[0092] Regarding step S4:
[0093] In the online monitoring of substation transformers, the asynchronous state sequence output by the multi-channel continuous-time recursive network is non-uniform in time scale. Each of the three channels updates its hidden state according to the arrival time of its corresponding input data. The operation and maintenance system may trigger diagnostic requests through scheduled inspections or manual intervention. There is a different time difference between the diagnostic request time and the latest state update time of each channel. If calculations are directly performed based on the asynchronous state sequence representing the equipment state at different physical moments, time misalignment will occur, leading to equipment diagnostic errors. Therefore, cross-scale attention calculation for state-time alignment is required.
[0094] In a preferred embodiment, the specific steps for performing cross-scale attention computation based on the asynchronous state sequence include:
[0095] S401. Obtain the current diagnostic request time, and the updated hidden state vector and update timestamp of each channel during the current asynchronous state sequence output process;
[0096] S402. For each channel, using the hidden state vector as the initial value, subtract the update timestamp from the diagnostic request time to obtain the integral duration;
[0097] S403. Based on the initial value and the integration duration, perform numerical integration on the neuron differential equation of each channel under the condition of no external input, and use the integration endpoint state as the alignment state vector corresponding to each channel.
[0098] Specifically, in this embodiment, the diagnostic request time is represented by a timestamp determined by the timer of the operation and maintenance system, the abnormal alarm trigger signal, or the manual operation command. After determining the diagnostic request time, the hidden state vector of each channel, which was most recently updated due to input data or event-driven processes, and the update timestamp corresponding to this update are read during the current asynchronous state sequence output. The update timestamp of the first channel is the completion time of the most recent sampling of dissolved gas data in oil, the update timestamp of the second channel is the arrival time of the most recent vibration signal, and the update timestamp of the third channel is the time when the most recent effective discharge event was detected. For each channel, the diagnostic request time is subtracted from its respective update timestamp to obtain the integration duration. Using the read hidden state vector and the integration duration, the Euler method is used to numerically integrate the ordinary differential equation corresponding to the channel. The integration step size is set by those skilled in the art based on the time constant of each channel. The larger the time constant, the larger the integration step size. This integration process infers the most likely state of each channel at the diagnostic request time without new input data. The state at the end of the integration is the aligned state vector corresponding to each channel. This step yields three aligned state vectors, each physically corresponding to the state at the moment of the diagnostic request, thus eliminating the timestamp misalignment problem in the asynchronous state sequence.
[0099] In a preferred embodiment, the specific steps for performing cross-scale attention computation based on the asynchronous state sequence further include:
[0100] S404. For each channel, the alignment state vector is input into a preset confidence estimation network, and the corresponding modal confidence is output.
[0101] S405. Multiply the alignment state vector by the modal confidence to obtain the query vector for each channel, and use the historical discharge mode features as the key vector and value vector.
[0102] S406. Based on the query vector, the key vector, and the value vector of all channels, the fused association features are obtained by scaling dot product attention calculation.
[0103] Due to differences in sensor conditions and noise environments, the diagnostic reliability of different monitoring data varies significantly under different operating conditions, and their contribution to the final decision differs at the time of diagnostic request. Therefore, a confidence estimation network is introduced to determine the reliability of the current aligned state vector. Specifically, the confidence estimation network consists of two fully connected layers and a sigmoid activation function. The input is the aligned state vector, and the output is a scalar with values ranging from 0 to 1, representing the modal confidence. When a channel experiences abnormal fluctuations due to sensor failure, strong electromagnetic interference, or mechanical shock, the corresponding aligned state vector deviates from its normal range, and the confidence estimation network outputs a modal confidence approaching zero. Then, the aligned state vector of each channel is multiplied by its corresponding modal confidence to obtain the query vector for that channel. Since the historical evolution information of dissolved gas data and vibration signals in the oil is already contained in the hidden states of the corresponding channel's continuous-time recursive network, the corresponding aligned state vector, as the query vector, already carries the corresponding long-term trend. Therefore, the key vector and value vector only come from historical valid discharge events. Effective discharge events are sparse and sudden, lacking long-term trends. In this embodiment, the essence of cross-scale attention computation is to retrieve causally related discharge patterns from the query vector formed by dissolved gas data and vibration signals in the oil. Specifically, historical discharge pattern features are encoded representations of phase and amplitude features in historical sparse feature vectors after dimensional compression. These historical discharge pattern features are simultaneously used as key and value vectors, with one feature corresponding to one key and value vector. All historical discharge pattern features are extracted from the memory storage module to form a set of key and value vectors. Then, scaled dot product attention computation is performed. Specifically, the dot product of the query vector for each channel and all key vectors is calculated, then divided by the square root of the key vector dimension to avoid gradient explosion. The attention weight for each historical discharge pattern is obtained by normalization using the softmax function. The attention weights are then multiplied sequentially by their corresponding value vectors, and finally summed to obtain the fused correlation features. In this embodiment, the query vector dimension for each of the three channels is equal to the aligned feature vector dimension, and both are equal to the dimension of the hidden state of each channel.
[0104] In a preferred embodiment, the specific steps for generating fault diagnosis results and uncertainties through a diagnostic classifier include:
[0105] S411. Input the fused correlation features into the joint diagnostic classifier, output the first probability distribution of each fault type, and take the fault type corresponding to the highest probability in the first probability distribution as the fault diagnosis result.
[0106] S412. Input each aligned feature vector into the corresponding single-modal classifier to obtain their respective second probability distributions. Calculate the Jensen-Shannon divergence between each pair of second probability distributions and then calculate the average value to obtain the modal conflict degree.
[0107] S413. Add the information entropy of the first probability distribution to the modal conflict degree to obtain the uncertainty.
[0108] Specifically, in this embodiment, the joint diagnostic classifier and the single-modal classifier adopt a unified network structure, consisting of fully connected layers and a softmax activation function. The difference lies in the number of input nodes. The number of input nodes for each diagnostic classifier is equal to the dimension of its respective input vector, and the number of output nodes is equal to the total number of predefined fault types. Each node represents the estimated probability of a fault type. In this embodiment, fault types are defined as state levels characterizing the degree of degradation of equipment operating status, including four state levels: "Normal," "Attention," "Abnormal," and "Severe." Therefore, the number of output nodes is set to four. The normal state indicates that the equipment operating indicators are normal, and no additional maintenance measures are required. The attention state indicates that the monitoring data deviates slightly from the normal baseline, and there may be an early degradation trend, requiring a shorter inspection cycle and strengthened monitoring. The abnormal state indicates that the monitoring data has deteriorated significantly, and identifiable fault characteristics have appeared, requiring immediate repair or confirmation. The severe state indicates that the equipment has a rapidly developing fault, and continued operation will face high risks, requiring immediate measures. The weight parameters of each diagnostic classifier are obtained through backpropagation training based on multi-source asynchronous data labeled with the true state levels. The fused and correlated features are input into the joint diagnostic classifier. The four output nodes output the estimated probability of each fault type, forming a first probability distribution. The fault type corresponding to the highest probability among the four estimated probabilities is taken as the final fault diagnosis result output by the joint diagnostic classifier. Three channels correspond to three single-modal classifiers. Each aligned state vector is input into its corresponding single-modal classifier, and each single-modal classifier independently outputs a second probability distribution, representing the estimated probability of the four state levels based on single-modal information. There are three sets of second probability distributions in total.
[0109] Furthermore, the Jensen-Shannon divergence between each pair of the three sets of second probability distributions is calculated. The Jensen-Shannon divergence is a symmetric measure of probability distribution difference; the smaller the divergence value between two second probability distributions, the more consistent the judgments of the two single-modal classifiers. After calculating the three divergence values, the arithmetic mean is taken as the modal conflict degree, characterizing the consistency of the multi-source physical field's judgment on the equipment status level. Then, according to the formula for calculating information entropy, the information entropy of the first probability distribution is calculated. The value of the information entropy reflects the concentration of each estimated probability in the first probability distribution. This information entropy is added to the modal conflict degree to obtain the final diagnostic uncertainty.
[0110] Regarding step S5:
[0111] According to the above implementation method, when the uncertainty of the diagnostic result is too high, it will cause misjudgment, leading to missed detection of early faults or overreaction to fluctuations in normal conditions. The cloud system builds a digital twin model of the target device based on the device design parameters. However, as the device's operating years increase, the actual physical parameters of the device gradually deviate from the initial design values, and the samples generated by the digital twin model through simulation will have deviations. Therefore, it is necessary to correct the model and introduce a cloud-edge collaborative update process to solve the model evolution problem when the uncertainty exceeds the limit. Specifically, the preset threshold is a configurable parameter determined based on the statistical characteristics of the uncertainty of historically confirmed diagnostic results. When the uncertainty calculated by the edge computing node exceeds the preset threshold, it indicates that the credibility of the current diagnostic result is insufficient. At this time, the cloud-edge collaborative process is triggered, and the relevant information of the current diagnosis is packaged and sent to the cloud. The digital twin model is used for targeted simulation, and then updated parameters are obtained through incremental training. The updated parameters are then sent to the edge computing nodes to replace the corresponding parameters of the multi-channel continuous-time recurrent network and the diagnostic classifier, completing the self-evolution update of the model and improving the diagnostic confidence.
[0112] In a preferred embodiment, the specific steps for obtaining the updated parameters include:
[0113] S501. When the uncertainty exceeds a preset threshold, the average value of the historical asynchronous state sequence output by each channel is obtained as the state statistical feature of the target device.
[0114] S502. Using the state statistical features as the observation vector and the preset aging parameters as the variables to be estimated, the Bayesian inference method is used to perform maximum a posteriori estimation to obtain the estimated values of the aging parameters.
[0115] S503. Based on the estimated aging parameters, the digital twin model is corrected, and then the fault type is used as the boundary condition to drive the corrected digital twin model for simulation to generate simulation samples.
[0116] Two problems exist when using digital twin models for cloud simulation: first, the actual equipment operating parameters will deviate from the design parameters in the digital twin model; second, uncertainty itself cannot reflect which type of fault is causing the diagnostic difficulty. Simulating all fault types one by one would waste cloud computing power and generate redundant training noise. Therefore, in this embodiment, when the uncertainty exceeds a preset threshold, the edge computing node reads the asynchronous state sequence output by each channel within a preset operating time period from the storage system. The preset operating time period is set by those skilled in the art to ensure that it includes a complete operating cycle of the target device before this diagnosis. Then, the average value of the asynchronous state sequence output by the three channels within the operating time period is calculated on the corresponding dimension. The three mean vectors are then concatenated into a long vector as the state statistical feature of the target device. This state statistical feature reflects the long-term operating state of the target device in the three physical fields, including information on physical state drift caused by equipment aging and environmental changes.
[0117] Furthermore, based on the implementation logic of Bayesian inference, a set of aging parameters is first preset as variables to be estimated in the digital twin model in the cloud. Aging parameters refer to physical parameters that play a decisive role in the deterioration of transformer oil-paper insulation, mechanical structure degradation, and thermal performance decay, including but not limited to insulation dielectric loss factor, winding stiffness coefficient, and oil thermal conductivity coefficient. For a given aging parameter, the corresponding expected state statistical characteristics can be calculated forward through the digital twin model. In this implementation, the prior distribution of the aging parameters and the likelihood function constitute the posterior probability density function. The prior distribution is determined based on historical aging statistics of similar equipment, and the likelihood function represents the probability of observing the current state statistical characteristics under given aging parameters.
[0118] In one optional implementation, the likelihood function is formed by mapping the aging parameters and state statistical features constructed through digital twin model simulation. Specifically, firstly, within the possible value range of aging parameters during the evolution of the target device from a healthy state to a severely aged state, different combinations of aging parameter sample points are generated using a uniform grid sampling method. For each set of aging parameter sample points, they are set into the digital twin model, driving the digital twin model to perform steady-state multiphysics simulation under fixed load and rated ambient temperature conditions, outputting corresponding simulated multi-source asynchronous data. Then, the simulated multi-source asynchronous data is input into a time recursive network deployed in the cloud. This time recursive network has the same structure as the multi-channel continuous time recursive network in the edge computing node, obtaining the simulated asynchronous state sequence of each channel. The mean vector of the simulated asynchronous state sequence is calculated and concatenated to form the simulated state statistical features corresponding to each set of aging parameters. The aging parameter values of all sample points and the simulated state statistical features constitute a mapping model. Furthermore, taking the state statistical features uploaded from the edge computing node as the observation vector, assuming that the observation noise follows a zero-mean multivariate normal distribution, and the noise covariance matrix is set according to the deviation statistics between the simulation output of the digital twin model and the actual monitoring data, the value of the multivariate normal distribution probability density function with the simulated state statistical features as the mean vector and the noise covariance matrix as the covariance matrix at the observation vector is the likelihood probability density of the current state statistical features, thus obtaining the likelihood function.
[0119] Furthermore, the maximum a posteriori (MAP) estimate is equivalent to maximizing the posterior probability density function. The posterior probability density function is proportional to the product of the prior distribution and the likelihood function. Therefore, maximizing the posterior estimate problem is transformed into minimizing the negative a posteriori estimate problem, i.e., taking the logarithm of the product of the prior distribution and the likelihood function, and then taking the negative of that logarithm to form the objective function. Using the current observation vector as input and the mean of the prior distribution of the variable to be estimated as the initial value, the gradient descent method is used to iteratively minimize the objective function. The variable to be estimated obtained after iterative convergence is the MAP estimate, which is used as the aging parameter estimate. The iterative convergence conditions include: the Euclidean norm of the change in the variable to be estimated after three consecutive iterations is less than a preset first threshold; the absolute value of the change in the objective function after three consecutive iterations is less than a preset second threshold; the Euclidean norm of the gradient of the objective function to the variable to be estimated is less than a preset third threshold; or the number of iterations reaches the maximum number of iterations. When any one of the convergence conditions is met, the iteration terminates and the current variable to be estimated is output. The first, second, and third thresholds are set by those skilled in the art according to the required accuracy, and the maximum number of iterations is set based on expert experience.
[0120] Furthermore, the physical parameters designed in the digital twin model are corrected using the estimated aging parameters to obtain a digital twin model calibrated with the current physical aging state of the equipment. Then, the fault type corresponding to the highest probability is estimated from the second probability distribution output by each single-modal classifier. The union of the fault types determined by the three second probability distributions is used to form the simulation boundary conditions, thereby achieving targeted simulation to generate simulation samples highly correlated with the current equipment and avoiding blindly simulating each fault one by one. Then, each fault type in the boundary conditions drives the corrected digital twin model to perform a simulation, simulating the complete process of the equipment evolving from the current state to the corresponding state level. The output simulation multi-source asynchronous data is used as simulation samples. The simulation samples formed by all simulations for the boundary conditions constitute the simulation sample set.
[0121] In a preferred embodiment, the specific steps for obtaining the updated parameters further include:
[0122] S504. Input the simulation sample into the cloud diagnostic model, and obtain the simulation fusion correlation features as the simulation feature set through asynchronous state sequence calculation and cross-scale attention calculation;
[0123] S505. Extract discharge mode features and corresponding historical diagnostic labels, combine the discharge mode features and the simulation feature set into a hybrid feature set, and combine the historical diagnostic labels and simulation fault type labels into a label set;
[0124] S506. Calculate the mean squared gradient of each weight parameter of the cloud diagnostic model on the mixed feature set, and use it as the importance weight; sum the cross-entropy loss of the label set and the product of the squared change of each weight parameter and the corresponding importance weight to form the loss function.
[0125] S507. With the goal of minimizing the loss function, train the cloud diagnostic model, extract all weight parameters after training, and send them as update parameters to the edge computing node.
[0126] Specifically, this embodiment deploys a cloud-based diagnostic model with the same structure as the edge computing nodes in the cloud. This includes a time-recurrent network with the same structure as the multi-channel continuous-time recurrent network, and a classifier network with the same structure as the joint diagnostic classifier and the single-modal classifier. Simulation samples are input into the cloud-based diagnostic model, and after asynchronous state sequence calculation and cross-scale attention calculation (similar to the edge side), simulation fusion correlation features are output to form a simulation feature set. In this embodiment, the fault diagnosis results output by the joint diagnostic classifier in the edge computing nodes are read from the memory storage module each time, serving as historical diagnostic labels corresponding to historical discharge mode features. The simulation fault type label refers to the fault type driving the digital twin model for simulation. The simulation feature set and discharge mode features are merged into a hybrid feature set, and the historical diagnostic labels and simulation fault type labels are merged into a hybrid label set. Further, the gradient squared mean is calculated through forward and backward propagation of the cloud-based diagnostic model on the hybrid feature set. The weight parameters of the cloud-based diagnostic model include the weight matrix and bias vector of the time-recurrent network, and the weight matrix and bias vector between each layer of the classifier network. Specifically, the calculation process of the mean squared gradient is as follows: First, the samples in the mixed feature set are input into the cloud diagnostic model, and forward propagation is performed. Each sample is processed by a time recursive network and cross-scale attention to obtain fused associated features, which are then output by the allocator network to produce a predicted probability distribution. Second, the cross-entropy loss value between the predicted probability distribution of each sample and the corresponding label in the label set is calculated. Third, the obtained cross-entropy loss value is backpropagated, and the partial derivative of the cross-entropy loss value with respect to each weight parameter in the cloud diagnostic model is calculated to obtain the gradient value corresponding to each sample. Fourth, for each weight parameter, the gradient value on all samples in the mixed feature set is counted, and the average value of each gradient value is obtained by squaring each gradient value. Then, the mean squared gradient value is used as the importance weight of each weight parameter, reflecting the degree of contribution of each weight parameter to the correct classification of historical samples and new samples.
[0127] Furthermore, incremental training aims to learn fault modes from new simulation samples to improve the diagnostic capability for current uncertain states, while also preserving knowledge learned from historical samples to avoid excessive forgetting. Therefore, this implementation uses the importance weight of each weight parameter as a basis, imposing stronger constraints on the more important weight parameters during training. Specifically, the cross-entropy loss values of all samples in the mixed feature set are summed and divided by the total number of samples to obtain the cross-entropy loss of the label set. Simultaneously, for each weight parameter, the square of the difference between the current value and the initial training value is calculated and multiplied by the corresponding importance weight to obtain a regularization penalty, where the initial training value is the weight parameter value of the cloud diagnostic model before incremental training begins. Then, the regularization penalties of all weight parameters are summed and added to the cross-entropy loss of the label set as the loss function for incremental training. Furthermore, by minimizing this loss function, the cloud diagnostic model is incrementally trained, enabling the model to adapt to new state diagnostic tasks while retaining its ability to discriminate historical state levels. Before training begins, the actual weight parameters of the current edge computing nodes are synchronized to the cloud diagnostic model as the starting value for training. The batch size is set to 16-256, the Adam optimizer is used, and the learning rate is set to 0.0001. The specific training process is as follows: First, the samples in the mixed feature set are divided into batches according to the batch size. The samples in each batch are input into the cloud diagnostic model. After asynchronous state sequence calculation and cross-scale attention calculation, the forward propagation of the expected classifier network is performed, and the predicted probability distribution is output. Second, the predicted probability distribution is substituted into the loss function to calculate the loss function value. Then, backpropagation is performed based on the loss function value to calculate the gradient of the loss function value with respect to all weight parameters. Then, the Adam optimizer is called to calculate the update amount of each weight parameter in this iteration of training based on the gradient of each weight parameter and the historical gradient statistics. The update amount is then superimposed with the starting value to output the current updated parameter. After each training epoch, the total loss function value of the validation set is calculated. Training terminates when the decrease in the total loss function value is lower than a preset convergence threshold for three consecutive epochs, or when the preset total number of training epochs is reached. The preset convergence threshold and preset total number of epochs are set by those skilled in the art. After training, all weight parameters in the cloud diagnostic model are extracted and used as update parameters for the multi-channel continuous-time recurrent network and diagnostic classifier in the edge computing nodes, and then distributed to the edge computing nodes. Upon receiving the updated parameters, the edge computing nodes replace the original corresponding parameters with the updated parameters, completing one cloud-edge collaborative self-evolutionary update.
[0128] Example 2:
[0129] like Figure 3 As shown, this embodiment provides a cloud-edge collaborative diagnostic system based on multi-source asynchronous data, including:
[0130] The data acquisition module is used to acquire multi-source asynchronous data of the target device through edge computing nodes. The multi-source asynchronous data includes dissolved gas data in oil, vibration signals, and discharge pulse signals.
[0131] The network construction module is used to construct a multi-channel continuous-time recursive network, wherein the time constant of each channel in the multi-channel continuous-time recursive network corresponds one-to-one with the sampling period of each multi-source asynchronous data.
[0132] The sequence generation module is used to perform sparse event detection based on the discharge pulse signal. When a valid discharge event is detected, a sparse feature vector is obtained. The dissolved gas data in the oil, the vibration signal, and the sparse feature vector are input into the multi-channel continuous-time recursive network to obtain the corresponding asynchronous state sequences.
[0133] The diagnostic output module is used to perform cross-scale attention calculation based on the asynchronous state sequence to obtain fused correlation features; based on the fused correlation features, a diagnostic classifier is used to generate fault diagnosis results and uncertainties.
[0134] The model update module is used to update the multi-channel continuous-time recurrent network and the diagnostic classifier by simulating the target device's cloud-based digital twin model and obtaining update parameters through incremental training when the uncertainty exceeds a preset threshold.
[0135] In summary, this invention constructs a multi-channel continuous-time recursive network, mapping the time constants of each channel to the sampling periods of multi-source asynchronous data, thus preserving the differences in time scales of the multi-source asynchronous data and avoiding information distortion caused by traditional time alignment. It utilizes a spiking neural network for sparse event detection, preserving weak discharge trends with subthreshold residual vectors, and updates the network structure to follow the physical state of the equipment through real-time cross-channel updates and dynamic correction of the time constants. During the diagnostic process, cross-scale attention calculations and uncertainty estimation are used to improve the accuracy and reliability of the diagnostic results. When the uncertainty exceeds the limit, continuous self-evolution of the edge model is achieved through digital twin model inversion simulation and incremental training. Therefore, this invention improves the accuracy of early latent fault diagnosis in substation transformers and the reliability of operation and maintenance decisions.
[0136] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A cloud-edge collaborative diagnosis method based on multi-source asynchronous data, characterized in that, The method comprises the following steps: obtaining multi-source asynchronous data of a target device through an edge computing node, wherein the multi-source asynchronous data comprises dissolved gas in oil data, vibration signals and discharge pulse signals; constructing a multi-channel continuous-time recurrent network, wherein the time constant of each channel of the multi-channel continuous-time recurrent network corresponds to the sampling period of each multi-source asynchronous data; performing sparse event detection based on the discharge pulse signals, and obtaining a sparse feature vector when a valid discharge event is detected; inputting the dissolved gas in oil data, the vibration signals and the sparse feature vector into the multi-channel continuous-time recurrent network to obtain corresponding asynchronous state sequences respectively; performing cross-scale attention calculation based on the asynchronous state sequences to obtain fusion correlation features; generating a fault diagnosis result and an uncertainty through a diagnostic classifier based on the fusion correlation features; when the uncertainty exceeds a preset threshold, simulating through a cloud digital twin model of the target device, and obtaining updated parameters through incremental training to update the multi-channel continuous-time recurrent network and the diagnostic classifier.
2. The cloud-edge collaborative diagnosis method based on multi-source asynchronous data according to claim 1, characterized in that, The specific steps of performing sparse event detection based on the discharge pulse signals comprise: inputting the discharge pulse signals into a preset pulse neural network, wherein the pulse neural network comprises a threshold comparison unit and a residual coding unit, and the threshold comparison unit generates a sparse feature vector when the amplitude of the discharge pulse signal reaches a preset amplitude threshold; the residual coding unit is configured with a sliding time window to calculate the amplitude mean and pulse interval mean of the discharge pulse signal, and map the amplitude mean and pulse interval mean into a sub-threshold residual vector through a preset embedding matrix; the sub-threshold residual vector is used as a bias term and is superimposed into the hidden state update equation of the corresponding channel of the multi-channel continuous-time recurrent network.
3. The cloud-edge collaborative diagnosis method based on multi-source asynchronous data according to claim 1, characterized in that, The inputting into the multi-channel continuous-time recurrent network also comprises a cross-channel state real-time updating process, and the specific steps comprise: inputting the sparse feature vector into a pre-trained mapping network to generate an influence increment vector representing the influence on the slow-scale channel; element-wise adding the influence increment vector and the hidden state vector of the slow-scale channel of the multi-channel continuous-time recurrent network to obtain an updated hidden state.
4. The cloud-edge collaborative diagnosis method based on multi-source asynchronous data according to claim 1, characterized in that, After obtaining the asynchronous state sequence, the time constant of the multi-channel continuous-time recurrent network is also modified, and the specific steps comprise: determining the initial time constant of each channel based on the sampling period of each multi-source asynchronous data, and calculating the state change rate of each channel according to the historical asynchronous state sequence; calculating a first cross-source difference vector based on the state change rate of the first channel and the state change rate of the second channel, and calculating a second cross-source difference vector based on the state change rate of the second channel and the state change rate of the third channel; concatenating the first cross-source difference vector and the second cross-source difference vector, inputting them into a fully connected layer, and outputting a time constant compensation vector; and multiplying the initial time constant of each channel with the time constant compensation vector element by element to obtain the modified time constant.
5. The cloud-edge collaborative diagnosis method based on multi-source asynchronous data according to claim 1, characterized in that, The specific steps for performing cross-scale attention computation based on the asynchronous state sequence include: Obtain the current diagnostic request time, and during the current asynchronous state sequence output process, the updated hidden state vector and update timestamp for each channel; For each channel, using the hidden state vector as the initial value, the diagnostic request time is subtracted from the update timestamp to obtain the integral duration; Based on the initial value and the integration duration, the differential equations of the neurons in each channel are numerically integrated under the condition of no external input, and the final state of the integration is used as the alignment state vector corresponding to each channel.
6. The cloud-edge collaborative diagnosis method based on multi-source asynchronous data according to claim 5, characterized in that, The specific steps for performing cross-scale attention computation based on the asynchronous state sequence also include: For each channel, the alignment state vector is input into a preset confidence estimation network, and the corresponding modal confidence is output. The alignment state vector is multiplied by the modal confidence to obtain the query vector for each channel, and the historical discharge mode features are used as the key vector and value vector. Based on the query vector, key vector, and value vector of all channels, the fused association features are obtained by scaling dot product attention calculation.
7. The cloud-edge collaborative diagnosis method based on multi-source asynchronous data according to claim 1, characterized in that, The specific steps for generating fault diagnosis results and uncertainties through the diagnostic classifier include: The fused and associated features are input into the joint diagnostic classifier, which outputs the first probability distribution of each fault type. The fault type with the highest probability in the first probability distribution is taken as the fault diagnosis result. Each aligned feature vector is input into the corresponding single-modal classifier to obtain its own second probability distribution. The Jensen-Shannon divergence between each pair of the second probability distributions is calculated, and the average value is then used to obtain the modal conflict degree. The uncertainty is obtained by adding the information entropy of the first probability distribution to the modal conflict degree.
8. The cloud-edge collaborative diagnosis method based on multi-source asynchronous data according to claim 1, characterized in that, The specific steps for obtaining the updated parameters include: When the uncertainty exceeds a preset threshold, the average value of the historical asynchronous state sequence output by each channel is obtained as the state statistical feature of the target device; Using the state statistical features as the observation vector and the preset aging parameters as the variables to be estimated, the Bayesian inference method is used to perform maximum a posteriori estimation to obtain the estimated values of the aging parameters. The digital twin model is corrected based on the estimated aging parameters, and then the corrected digital twin model is simulated using the fault type as the boundary condition to generate simulation samples.
9. The cloud-edge collaborative diagnosis method based on multi-source asynchronous data according to claim 8, characterized in that, The specific steps for obtaining the updated parameters also include: The simulation samples are input into the cloud-based diagnostic model, and simulation fusion correlation features are obtained through asynchronous state sequence calculation and cross-scale attention calculation, which serve as the simulation feature set. Extract discharge mode features and corresponding historical diagnostic labels, combine the discharge mode features and the simulation feature set into a hybrid feature set, and combine the historical diagnostic labels and simulation fault type labels into a label set; The mean squared gradient of each weight parameter of the cloud diagnostic model on the mixed feature set is calculated as the importance weight; the cross-entropy loss of the label set and the product of the squared change of each weight parameter and the corresponding importance weight are summed to form the loss function. With the goal of minimizing the loss function, the cloud-based diagnostic model is trained, and all weight parameters after training are extracted and sent to the edge computing node as update parameters.
10. A cloud-edge collaborative diagnosis system based on multi-source asynchronous data, used to implement the cloud-edge collaborative diagnosis method based on multi-source asynchronous data according to any one of claims 1-9, characterized in that, include: The data acquisition module is used to acquire multi-source asynchronous data of the target device through edge computing nodes. The multi-source asynchronous data includes dissolved gas data in oil, vibration signals, and discharge pulse signals. The network construction module is used to construct a multi-channel continuous-time recursive network, wherein the time constant of each channel in the multi-channel continuous-time recursive network corresponds one-to-one with the sampling period of each multi-source asynchronous data. The sequence generation module is used to perform sparse event detection based on the discharge pulse signal, and to obtain a sparse feature vector when a valid discharge event is detected. The dissolved gas data in the oil, the vibration signal, and the sparse feature vector are input into the multi-channel continuous-time recursive network to obtain the corresponding asynchronous state sequences. The diagnostic output module is used to perform cross-scale attention calculation based on the asynchronous state sequence to obtain fused correlation features; Based on the fused correlation features, a diagnostic classifier is used to generate fault diagnosis results and uncertainties. The model update module is used to update the multi-channel continuous-time recurrent network and the diagnostic classifier by simulating the target device's cloud-based digital twin model and obtaining update parameters through incremental training when the uncertainty exceeds a preset threshold.