High-altitude power transformation equipment defect dynamic diagnosis method and device based on cross-modal analysis

By using multimodal data fusion and uncertainty quantification, the problem of low diagnostic accuracy of power equipment in high-altitude environments has been solved, achieving higher diagnostic accuracy and robustness.

CN121744167BActive Publication Date: 2026-05-05STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST
Filing Date
2026-02-28
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In high-altitude environments, the accuracy of defect diagnosis for power equipment is low, and existing single-type data is insufficient to characterize the equipment status.

Method used

By acquiring multimodal monitoring data (video, audio, and text) of power equipment, uncertainty is quantified and fused. The intrinsic connectivity strength between modes is determined by using the uncertainty covariance matrix and the Riemannian manifold geodesic distance. Geometric modulation attention fusion is then performed to generate enhanced representations, which are then input into a pre-trained defect diagnosis model for diagnosis.

Benefits of technology

It improves the accuracy of defect diagnosis for power equipment in high-altitude environments, enhances the robustness and accuracy of the diagnostic structure, and can provide reliable diagnostic results in highly uncertain scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744167B_ABST
    Figure CN121744167B_ABST
Patent Text Reader

Abstract

This invention discloses a method and apparatus for dynamic diagnosis of defects in high-altitude substations using cross-modal analysis. It belongs to the field of power operation and maintenance technology. The method includes acquiring monitoring data for each mode of the substation; solving the uncertainty evolution equation for each mode using the monitoring data to determine the uncertainty covariance matrix of each mode; determining the intrinsic connectivity strength between modes based on the similarity of the uncertainty covariance matrices between different modes; determining the geometric modulation attention weights between modes based on the intrinsic connectivity strength and the content relevance between corresponding modes; performing weighted fusion of each mode using the geometric modulation attention weights to determine the enhanced representation of each mode; and inputting the enhanced representation into a pre-trained defect diagnosis model to obtain the diagnostic results of the substation defects. By quantifying and fusing uncertainty in multimodal data to determine the enhanced representation for defect diagnosis, the accuracy of substation defect diagnosis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power operation and maintenance technology, specifically to a method and apparatus for dynamic diagnosis of defects in high-altitude substations using cross-modal analysis. Background Technology

[0002] In the operation of power systems, the health status of substation equipment, as the core infrastructure for energy transmission, conversion and distribution, directly determines the stability of the power grid and the reliability of power supply.

[0003] Currently, defect diagnosis of substation equipment typically involves inputting a single type of data into a classification model for defect diagnosis, and then using the output of the classification model as the diagnostic result for the substation equipment defect. However, in high-altitude environments, due to environmental influences such as low air pressure, low temperature, and large temperature differences, single-type data is often insufficient to characterize the state of the substation equipment. Therefore, traditional defect diagnosis methods for substation equipment in high-altitude environments suffer from low diagnostic accuracy. Summary of the Invention

[0004] The purpose of this invention is to provide a method and apparatus for dynamic diagnosis of defects in high-altitude substations through cross-modal analysis. By acquiring multimodal monitoring data of the substations and performing uncertainty quantification and fusion on the multimodal data, an enhanced characterization for defect diagnosis is determined, thus solving the problem of low accuracy in defect diagnosis of substations in high-altitude environments.

[0005] This invention is achieved through the following technical solution:

[0006] The first aspect of this application provides a method for dynamic diagnosis of defects in high-altitude substations using cross-modal analysis, including:

[0007] Acquire monitoring data for each mode of the power equipment; the modes include video modes for monitoring the appearance of the power equipment, audio modes for collecting the acoustic signature of the power equipment, and text modes for describing relevant information about the power equipment.

[0008] The uncertainty evolution equation for each mode is solved using the monitoring data to determine the uncertainty covariance matrix of each mode;

[0009] The intrinsic connectivity strength between different modes is determined based on the similarity of the uncertainty covariance matrices; the similarity is based on the geodesic distance representation of the uncertainty covariance matrices on the Riemannian manifold.

[0010] Based on the intrinsic connection strength between the modes and the content correlation between the corresponding modes, the geometric modulation attention weights between the modes are determined.

[0011] The geometrically modulated attention weights are used to perform weighted fusion of each modality while preserving the Riemannian manifold geometry, in order to determine the enhanced representation of each modality that incorporates features from other modalities.

[0012] The enhanced representation is input into the pre-trained defect diagnosis model, causing the defect diagnosis and identification model to perform the following operations to output the diagnostic results of the substation equipment defects:

[0013] Based on the geometric modulation attention weights between modalities, the cognitive uncertainty coefficient is determined to characterize the overall cognitive uncertainty of each modality;

[0014] The quantitative risks of the data-driven branch and the physical model branch are weighted and fused by cognitive uncertainty coefficients to determine the risk indicators of substation equipment defects. The data-driven branch quantifies the defect risk of substation equipment from the data statistics level of each modality, while the physical model branch quantifies the defect risk of substation equipment from the physical mechanism level.

[0015] In one feasible implementation, the method further includes: determining the root cause type of the substation equipment defect risk through a sensitivity index based on the risk index of the substation equipment defect, and formulating countermeasures for the defect risk based on the determined root cause type;

[0016] The root cause types include environmental type, equipment type, and observation noise type, and each root cause type includes data of at least two characterizing attributes;

[0017] The sensitivity index characterizes the degree of influence of each root cause type on the risk indicator;

[0018] The sensitivity index includes the degree of independent impact of data changes based on one root cause type on the risk indicator, and the degree of total impact of data changes based on at least one root cause type on the risk indicator.

[0019] In one feasible implementation, the method further includes: determining an arbitration mechanism for consensus conflicts between video and audio modalities based on inter-modal geometric modulation attention weights; specifically including:

[0020] From the intermodal geometric modulation attention weights, obtain the first attention weight of the audio modality on the video modality and the second attention weight of the video modality on the audio modality.

[0021] The conflict index is determined based on the minimum value between the first attention weight and the second attention weight;

[0022] When the conflict index exceeds a preset threshold, the corresponding text modality data is acquired as arbitration evidence.

[0023] The arbitration conclusion is determined based on the aforementioned arbitration evidence; the arbitration conclusion includes the credibility of the defect diagnosis conclusion based on the video modality and the defect diagnosis conclusion based on the audio modality.

[0024] The process of quantifying defect risk is adjusted based on the arbitration conclusion, adjusting the data-driven branch and / or physical model branch.

[0025] In one feasible implementation, the method further includes determining an uncertainty score for the diagnostic result based on the cognitive uncertainty coefficient, and when the uncertainty score indicates that the prediction result is questionable, performing:

[0026] The enhanced representation, the uncertainty score, the diagnostic results, the timestamp, and the substation metadata tags are uploaded to the cloud to obtain the global metacognitive model fed back from the cloud; the global metacognitive model uses the data uploaded by each node as samples and performs training on the samples to determine the model.

[0027] The output of the global metacognitive model is used as a soft label to optimize the defect diagnosis and identification model of the local node.

[0028] In one feasible implementation, the method further includes: triggering tiered alarms based on the coupling relationship between comprehensive risk indicators and integrated uncertainty indicators; specifically including:

[0029] If the comprehensive risk index is less than the first preset risk value, a level one alarm is triggered;

[0030] If the comprehensive risk index exceeds the first risk preset value but is less than the second risk preset value, and the fused uncertainty index is less than the uncertainty preset value, then a level two alarm is triggered.

[0031] If the fusion uncertainty index exceeds the preset uncertainty value, a level three alarm will be triggered;

[0032] If the comprehensive risk index exceeds the second risk preset value and the fused uncertainty index is less than the uncertainty preset value, a level four alarm will be triggered.

[0033] The risk levels of the first-level alarm, second-level alarm, third-level alarm, and fourth-level alarm increase sequentially.

[0034] The fusion uncertainty index is the optimized result of the uncertainty score during cloud training; the comprehensive risk index is determined based on the uncertainty index and the risk index of the substation equipment defects.

[0035] In one feasible implementation, the method further includes constructing a causal sensing network to predict the operational risks of the power grid where the substation is located, specifically including:

[0036] A connection graph is constructed between substation equipment in the power grid; the connection graph uses substation equipment in the power grid as equipment nodes, and the edges of the connection graph are constructed based on the electrical topology and / or physical location relationships between the substation equipment; a time-series attribute is added to each equipment node, and the time-series attribute is determined based on the monitoring data;

[0037] Each variable in the time-series data of the device node is represented as a nonlinear function of the causal parent node;

[0038] The objective function is determined by minimizing the noise term in the nonlinear function and regularizing the input weights of the neural network used to fit the nonlinear function;

[0039] Using causal rules based on objective laws as constraints, the objective function is solved to determine the set of causal parent nodes for each variable;

[0040] Each node in the set of causal parent nodes connecting the power equipment and each variable of the power equipment is used as an edge of the causal perception graph. The causal perception graph is constructed by using each power equipment as a node of the causal perception graph.

[0041] Based on the causal perception map and the defect diagnosis results of the substation equipment, the risk of the power grid where the substation equipment is located is predicted.

[0042] In one feasible implementation, the weights of the edges in the causal sensing graph are determined in the following way:

[0043] Based on multidimensional driving factors, a time-varying attenuation coefficient is determined; the multidimensional driving factors include at least the determination of technological generation gap and equipment iteration coefficient; the technological generation gap represents the generational gap between the current technology and the technology on which the edge of the causal perception graph is based; the equipment iteration coefficient represents the degree of iteration of the substation equipment;

[0044] The weights of the edges in the causal sensing graph are determined based on the time-varying decay coefficient.

[0045] A second aspect of this application provides a dynamic diagnostic device for defects in high-altitude substations using cross-modal analysis, comprising:

[0046] A multimodal data acquisition unit is used to acquire monitoring data for each mode of the substation equipment; the modes include a video mode for monitoring the appearance of the substation equipment, an audio mode for collecting the acoustic signature of the substation equipment, and a text mode for describing the relevant information of the substation equipment.

[0047] The monitoring data statistical analysis unit solves the pre-constructed uncertainty evolution equation for each mode using the monitoring data, and determines the uncertainty covariance matrix of each mode;

[0048] An intrinsic connectivity strength determination unit determines the intrinsic connectivity strength between modes based on the similarity of the uncertainty covariance matrices between different modes; the similarity is based on the geodesic distance representation of the uncertainty covariance matrix on the Riemannian manifold.

[0049] The attention weight determination unit determines the geometric modulation attention weights between modes based on the inherent connection strength between the modes and the content correlation between the corresponding modes;

[0050] Feature representation unit: The geometrically modulated attention weights are used to perform weighted fusion of each modality while preserving the Riemannian manifold geometry, so as to determine the enhanced representation of each modality that has been fused with features of other modalities;

[0051] The model diagnostic unit is used to input the enhanced representation into a pre-trained defect diagnostic model, causing the defect diagnostic identification model to perform the following operations to output the diagnostic results of the substation equipment defects:

[0052] Based on the geometric modulation attention weights between modalities, the cognitive uncertainty coefficient is determined to characterize the overall cognitive uncertainty of each modality;

[0053] The quantitative risks of the data-driven branch and the physical model branch are weighted and fused by cognitive uncertainty coefficients to determine the risk indicators of substation equipment defects. The data-driven branch quantifies the defect risk of substation equipment from the data statistics level of each modality, while the physical model branch quantifies the defect risk of substation equipment from the physical mechanism level.

[0054] In one feasible implementation, the device further includes a defect tracing unit, which determines the root cause type of the defect risk of the power equipment based on the risk index of the power equipment defect and through a sensitivity index, and formulates countermeasures for the defect risk based on the determined root cause type.

[0055] The root cause types include environmental type, device type, and observation noise type, and each root cause type includes data of at least two characterizing attributes;

[0056] The sensitivity index characterizes the degree of influence of each root cause type on the risk indicator;

[0057] The sensitivity index includes the degree of independent impact of data changes based on one root cause type on the risk indicator, and the degree of total impact of data changes based on at least one root cause type on the risk indicator.

[0058] In one feasible implementation, the device further includes an arbitration unit that determines an arbitration mechanism when there is a consensus conflict between the video modality and the audio modality based on the geometric modulation attention weights between the modalities.

[0059] Specifically used for:

[0060] From the geometric modulation attention weights between modalities, the first attention weight of the audio modality to the video modality and the second attention weight of the video modality to the audio modality are obtained respectively; based on the minimum value of the first attention weight and the second attention weight, a conflict index is determined; when the conflict index exceeds a preset threshold, the data of the corresponding text modality is obtained as arbitration evidence; based on the arbitration evidence, an arbitration conclusion is determined; the arbitration conclusion includes the credibility of the defect diagnosis conclusion based on the video modality and the defect diagnosis conclusion based on the audio modality; based on the arbitration conclusion, the quantification process of defect risk by the data-driven branch and / or the physical model branch is adjusted.

[0061] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0062] This application embodiment quantifies the uncertainty of monitoring data from each modality to characterize the degree of variation in monitoring data of substation equipment under high-altitude environments due to interference from various factors. During multimodal monitoring data fusion, the quantified uncertainty is represented on a Riemannian manifold, and the intrinsic connectivity strength between modalities is determined using geodesic distance. Geometric modulation attention fusion is achieved by fusing intrinsic connectivity strength with content similarity. Subsequently, the monitoring data from each modality are weighted and fused while maintaining the geometric structure of the Riemannian manifold, generating an enhanced representation for each modality that incorporates global information. Since the geometric modulation attention fusion stage fuses the intrinsic connectivity strength reflecting the uncertainty formula with the content similarity determined based on the monitoring data—that is, it achieves the fusion of prior strength and actual monitoring data—it can improve the robustness of the diagnostic structure in high-uncertainty scenarios. Furthermore, the enhanced representation of each modality incorporates features from other modalities, further improving the accuracy of the diagnostic results. Attached Figure Description

[0063] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:

[0064] Figure 1 A flowchart illustrating a method for dynamic diagnosis of defects in high-altitude substations using cross-modal analysis, provided in an embodiment of this application.

[0065] Figure 2A schematic diagram of the conflict arbitration mechanism in a cross-modal analysis method for dynamic diagnosis of defects in high-altitude substations provided in this application embodiment;

[0066] Figure 3 A flowchart illustrating the metacognitive federated network guided by the uncertainty of edge-cloud collaboration in a cross-modal analysis method for dynamic diagnosis of defects in high-altitude substations provided in this application embodiment;

[0067] Figure 4 This is a schematic diagram of the results of graded early warning in a cross-modal analysis method for dynamic diagnosis of defects in high-altitude substations provided in an embodiment of this application.

[0068] Figure 5 A schematic diagram of the workflow of the causal sensing network in a cross-modal analysis method for dynamic diagnosis of defects in high-altitude substations provided in this application embodiment;

[0069] Figure 6 A schematic diagram of a dynamic diagnostic device for defects in high-altitude substations based on cross-modal analysis provided in this application embodiment;

[0070] Figure 7 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation

[0071] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments and accompanying drawings. The illustrative embodiments and descriptions of this invention are for explanation only and are not intended to limit the invention. All other embodiments obtained by those skilled in the art based on the embodiments in this application without creative effort are within the scope of protection of this application.

[0072] As will be known to those skilled in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0073] The terms “comprising” and “having”, and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such process, method, product, or apparatus.

[0074] Example 1:

[0075] Embodiment 1 of this application provides a dynamic diagnostic method for defects in high-altitude substations using cross-modal analysis, which aims to solve the problem of low accuracy in defect diagnosis of substations in high-altitude environments.

[0076] The subject executing this method can be any computing device capable of implementing the method, such as a server, mobile phone, personal computer, smart wearable device, smart robot, etc.

[0077] Furthermore, the embodiments of this application do not limit the execution order of different steps. When using the method provided in the embodiments of this application, the execution order of different steps can be adjusted according to actual needs.

[0078] For ease of description, the following uses a high-altitude substation defect dynamic diagnosis device with cross-modal analysis as the execution subject of this method to provide a detailed description of the method provided in this application embodiment.

[0079] like Figure 1 The diagram shown is a flowchart illustrating the specific implementation of a cross-modal analysis method for dynamic diagnosis of defects in high-altitude substations provided in this application, including the following steps 11-16:

[0080] Step 11: Obtain monitoring data for each mode of the power equipment.

[0081] The modalities in this embodiment include a video modal for monitoring the appearance of the power equipment, an audio modal for collecting the voiceprints of the power equipment, and a text modal for describing the relevant information of the power equipment.

[0082] Among them, the video modal monitoring data is a visual representation of the appearance and spatial status of the substation equipment, and is the basis for intuitively judging the defects of the substation equipment; the video modal monitoring data includes image sequences of the appearance of the substation equipment.

[0083] The monitoring data of audio modes is an acoustic characterization of the operating acoustic characteristics of power equipment and environmental noise; the monitoring data of audio modes includes the time-domain / frequency-domain waveforms of the equipment operating sound signals.

[0084] The monitoring data in the text modality is a semantic representation of equipment operation and maintenance semantic information and environmental description; including real-time inspection logs of substation equipment, environmental text reports, historical maintenance records and standard technical documents.

[0085] Step 12: Solve the pre-constructed uncertainty evolution equation for each mode using the monitoring data to determine the uncertainty covariance matrix of each mode.

[0086] In the extreme environment of high altitude, power equipment monitoring data exhibits complex and inherent random characteristics. For example, video data is affected by intermittent obstruction from sandstorms and raindrops, resulting in random-intensity, spatiotemporally correlated impulse noise that easily obscures key features. Voiceprint signals undergo nonlinear distortion due to the low-pressure environment, exhibiting frequency-selective attenuation with pressure fluctuations and interference from background wind noise, leading to confusion regarding the physical meaning of acoustic features. Text logs are constrained by the subjectivity of manual recording, resulting in inconsistent descriptive granularity, strong contextual dependence, and the mixing of plateau-specific terminology, easily causing semantic ambiguity and vagueness. Therefore, modal state variables are defined as follows: , represents the uncertainty index of the i-th mode at time t. The higher the index value, the higher the uncertainty.

[0087] The construction of dynamic evolution equations for each modal uncertainty index. Specifically, this includes:

[0088] For each mode, an observation model of SDE (Stochastic Differential Equation) is established.

[0089] For video modal SDE, a latent stochastic evolution field is designed to model the dynamic changes of spatially correlated uncertainties caused by occlusion. Specifically, the spatial correlation of occlusion noise is modeled using the latent stochastic evolution field, and a spatiotemporal decoupling uncertainty probe is designed to extract the uncertainties of the video modality, as specifically expressed as:

[0090] ;

[0091] From a physical perspective, this equation describes the uncertainty of the video. The evolution of. Among them Indicates the optical flow field of the current frame; This represents a binary occlusion mask. A value of 1 indicates a white area that is obscured by sand or dust, while a value of 0 indicates a clean black area. It is a mask (Binary occlusion mask, =1 indicates that the corresponding pixel is obscured by sand, raindrops, etc. =0 indicates the total variation of pixels without occlusion (clean area), which is used to quantify the spatial irregularity and fragmentation of the occluded area. The more fragmented the occluded area is and the longer the boundary is, the larger its TV value is. The reversion coefficients, characterizing the video mode, are derived from the optical flow field of the current frame. The decision is used to adjust the rate at which uncertainty regresses to the benchmark value; The benchmark value characterizing the video modal uncertainty represents the inherent reliability level of video data when there is no environmental interference (a preset constant based on historical normal data statistics). The noise intensity coefficient characterizes the video mode and quantifies the impact of random interference (such as sudden sandstorms) on video uncertainty in high-altitude environments (preset constant). The Wiener process (Brownian motion) increments characterizing video modalities follow a mean of 0 and a variance of . The normal distribution is used to simulate the random fluctuations of environmental disturbances.

[0092] For the speaker signature modal SDE, a spectral deformation-aware stochastic process is designed to model the dynamic changes in uncertainty caused by signal distortion due to ambient air pressure fluctuations. Specifically, the acoustic signal-to-noise ratio and frequency domain distortion are inferred in real time using a dynamic covariance kernel function, thereby decoupling and generating a physically interpretable time-varying uncertainty confidence band in the speaker signature modality, as specifically expressed as:

[0093] ;

[0094] From a physical perspective, voiceprint uncertainty It will revolve around a baseline value determined by altitude h. Fluctuations. Among them... It is the spectrum of the current audio frame. It is the reference spectrum of the equipment under normal conditions, measured under standard air pressure and noise-free conditions; SSC(a,b) is the spectral structure similarity coefficient, which is used to measure the degree of similarity between two spectra in shape; The recovery rate constant characterizing the voiceprint mode adjusts the rate at which uncertainty regresses to the altitude reference value (a preset constant, determined based on voiceprint signal stability testing). The noise intensity coefficient characterizes the acoustic signature mode and quantifies the impact of environmental wind noise and equipment vibration interference on the uncertainty of acoustic signature (preset constant). Wiener process increments characterize acoustic signature modes to simulate the volatility of random disturbances such as low-pressure fluctuations and sudden wind noise.

[0095] For text modal semantic evolution (SDE), the topic-aware jump-diffusion process is designed to model the sudden jumps in textual semantic uncertainty caused by topic conflicts or descriptive ambiguities. Specifically, it decouples topic-guided latent space perturbations and uncertainties through a topic-aware semantic entanglement quantifier, and then combines this with dynamic context evaluation to achieve refined tracing and quantification of semantic uncertainty in text modalities.

[0096] ;

[0097] From a physical perspective, textual uncertainty There is a natural tendency for uncertainty to decrease; for example, the clearer the context, the lower the uncertainty. It is the strength of The Poisson process, the rate at which jumps occur It is proportional to the topic distribution entropy of the current sentence; It is a description of the current text. The function that determines this: , It is directly proportional to the number of ambiguous words and the vagueness of reference in the description; (Preset constant) represents the natural decay coefficient of the text modality, adjusting the rate of decrease in uncertainty when there is no ambiguity: the clearer the semantics, the better. It decays faster over time.

[0098] The above establishes the uncertainty evolution equation for each mode.

[0099] Based on the monitoring data obtained in step 11, the uncertainty evolution equation for each mode is solved by combining dynamic evolution and statistical depth to obtain the dynamic evolution trajectory of uncertainty for each mode and its corresponding (posterior) covariance matrix. The result of this covariance matrix not only accurately quantifies the instantaneous unreliability of video, voiceprint, and text modes affected by random environmental disturbances in the plateau region, but also reveals the model's own level of confidence in the above scores, thus providing a fusion input that combines temporal dynamics and statistical depth for the uncertainty geometric transformation described below.

[0100] Step 13: Determine the intrinsic connectivity strength between modes based on the similarity of the uncertainty covariance matrices between different modes; the similarity is based on the geodesic distance representation of the uncertainty covariance matrix on the Riemannian manifold.

[0101] To further integrate the uncertainty information of each mode, this embodiment proposes an uncertainty geometric graph transformation. By defining the uncertainty covariance matrix of each mode as coordinate points on the Riemannian manifold, the essential differences in uncertainty between modes are quantified by calculating the geodesic distance on the manifold, and a geometric prior is constructed accordingly.

[0102] The intrinsic connectivity weights between modalities based on uncertainty similarity are calculated as follows:

[0103] For any two modes i and j, calculate the Riemann distance between their covariance matrices:

[0104] ;

[0105] Next, calculate the prior strength (intrinsic connectivity strength): ;

[0106] In the formula, The square geodesic distance on the Riemannian manifold represents the uncertainty covariance matrix of mode i and mode j at time t. It is used to quantify the essential difference in uncertainty between the two modes. The smaller the value, the higher the consensus between the two modes. The uncertainty covariance matrices of modes i and j at time t are respectively used to characterize the statistical distribution characteristics of mode uncertainty; The intrinsic connection strength (prior weight) between mode i and mode j at time t is represented by a value close to 1, indicating that the uncertainty distributions of the two modes are more similar and the consensus is higher. The distance sensitivity coefficient, which modulates the influence of Riemann distance on prior weights, is a preset constant.

[0107] Step 14: Determine the geometric modulation attention weights between modes based on the intrinsic connection strength between the modes and the content correlation between the corresponding modes.

[0108] Content correlation between modalities is generated through linear projection:

[0109] , , ;

[0110] In the formula, The query matrix of modality i at time t is used to actively match the key matrix of other modalities and capture modality i's need for "semantic association". The key matrix of mode j at time t is used to match the query matrix of other modes, representing the core semantic features of mode j. The value matrix of mode j at time t represents the specific feature information of mode j that matches semantic association. The learnable query linear transformation weight matrix maps the features of modality i to a query vector, which is used to capture the key feature requirements of modality i. The learnable key linear transformation weight matrix maps the features of mode j to key vectors, which are used to match the query requirements of mode i. Represents the learnable weights (values) linear transformation weight matrix; Represents the original feature vector of mode i at time t (such as the visual feature vector of a video, the spectral feature vector of a voiceprint, or the semantic embedding vector of a text). This represents the original eigenvector of mode j at time t.

[0111] Next, calculate the scaled dot product score: ;

[0112] In the formula, At time t, the semantic matching score between mode i and mode j is represented. The higher the score, the stronger the semantic association between the feature contents of the two. is the scaling factor, and is the key matrix. Dimensions ( The square root of ) is used to alleviate the problem of excessively large dot product results due to high dimensionality and gradient vanishing after passing through the Softmax activation function.

[0113] The standard Attention mechanism is used to capture the correlation between features at the content level between modalities. ,Right now:

[0114] ;

[0115] in, Represents the original eigenvectors of modes i and j at time t; This represents the learnable linear transformation weights.

[0116] By fusing the intrinsic connection strength between the two modalities and the corresponding content relevance between the two modalities, the geometric modulation attention weights between the two modalities are determined. Specifically, it multiplies the intrinsic connection strength by the content relevance, as expressed in the following expression:

[0117] ;

[0118] And standardize to ensure that for each mode i, the sum of the weights of all modes j is 1:

[0119] ;

[0120] In the formula, This represents the modality, with values ​​of v (video modality), a (audio modality), and t (text modality). The unstandardized weights of mode i and all modes (k=v,a,t) are represented in the denominator. Summation, as a normalization term, ensures that the sum of the weights of mode i with respect to all modes after standardization is 1.

[0121] The final weights reflect semantic relevance and uncertainty confidence, and contributions to high-uncertainty modes are automatically suppressed.

[0122] Step 15: Weighted fusion of each modality is performed using the geometrically modulated attention weights while preserving the Riemannian manifold geometry, to determine the enhanced representation of each modality that incorporates features from other modalities.

[0123] Calculate the Riemann mean: ;

[0124] Riemann Mean The Riemann mean (reference point) of the covariance matrix of all modal uncertainties on the Riemann manifold at time t is used to find the point on the Riemann manifold that minimizes the sum of squared geodesic distances to all modal covariance matrices, representing the geometric center of the multimodal uncertainty.

[0125] In the formula, Candidate covariance matrices (variables) representing the Riemannian manifold are used to find the optimal reference point; Characterize the uncertainty covariance matrix of mode i at time t; The uncertainty covariance matrix of mode i The geometric difference between the candidate covariance matrix Σ and the squared geodesic distance on the Riemannian manifold is quantified.

[0126] Furthermore, the uncertainty covariance matrix of each mode is projected from the Riemannian manifold to the tangent space of the reference point through a logarithmic mapping, i.e.:

[0127] ;

[0128] We obtain a symmetric matrix, which can be viewed as a vector in the tangent space.

[0129] In the formula, The covariance matrix of mode j at time t. After logarithmic mapping, at the reference point Tangent vectors in tangent space (symmetric matrix form); Characterization with The logarithmic mapping operator with a base point is a core geometric transformation tool on Riemannian manifolds, which projects a point on the manifold onto the tangent space of that point. The square root of the covariance matrix at the reference point.

[0130] In the Euclidean subspace of the tangent space, a secure linear weighted aggregation is performed using geometrically modulated attention weights. For the target mode i, the fusion result is calculated as follows:

[0131] ;

[0132] In the formula, At time t, the weighted aggregated fused tangent vector of target mode i contains the information contribution of all modes under the geometric modulation attention weight.

[0133] Finally, the aggregated tangent space vectors are mapped back from the tangent space to the original Riemannian manifold via an exponential mapping, yielding the final uncertainty representation after fusion:

[0134] ;

[0135] At time t, the final fusion uncertainty covariance matrix of target mode i, or the fusion uncertainty representation or the fusion enhancement representation, is the uncertainty covariance matrix after exponential mapping. It contains multimodal fusion information and maintains the geometric structure of the Riemannian manifold. The inverse square root of the covariance matrix of the reference point.

[0136] In summary, for each mode i, the output is a fused geometric modulation attention multimodal uncertainty representation that includes information about mode i itself, as well as supplementary information provided by other modes under geometric modulation attention weights, while maintaining the manifold structure.

[0137] This embodiment proposes a cross-modal uncertainty tracing approach in the modal fusion process, achieving accurate tracing of data noise and model cognitive uncertainty in video, voiceprint, and text modalities, and generating physically interpretable confidence scores. Furthermore, the proposed uncertainty geometric transformation represents uncertainty on a Riemannian manifold, and utilizes geodesic distance and parallel transmission mechanisms to accurately model the nonlinear time-delay correlation of cross-modal uncertainties.

[0138] Step 16: Input the enhanced representation into the pre-trained defect diagnosis model, so that the defect diagnosis and identification model performs the following steps 1601-1602 to output the diagnosis result of the substation defect.

[0139] In step 16, the enhanced representation determined in step 15 is transformed into quantifiable risk indicators and actionable decisions: by calculating the information entropy of the geometrically modulated attention weights, the overall cognitive uncertainty level of the system is dynamically assessed, and an adaptive aggregation coefficient is generated accordingly. Subsequently, a dual-stream risk assessment was run in parallel: the data-driven branch introduced a consensus-weighted CVaR mechanism, using attention weights to modulate the tail risk expectation, making it reflect the intermodal consensus; the physical model branch, through uncertainty propagation theory, derived the full probability density function of the fault electrical quantities and calculated their higher-order moment deviations. Both were then... After weighted fusion, a cognitive adaptive risk index R is generated. Specifically:

[0140] Step 1601: Determine the cognitive uncertainty coefficient based on the geometric modulation attention weights between modalities to characterize the overall cognitive uncertainty of each modality.

[0141] Quantify the overall uncertainty of the entire system at time t to adjust the trust ratio between the data-driven branch and the physical model branch. Specifically, this involves geometrically modulating the attention weights. Extract cognitive uncertainty and calculate the entropy value for each modality i:

[0142] ;

[0143] Entropy The larger the entropy value, the more dispersed the dependence of mode i on other modes during decision-making, indicating a higher overall cognitive uncertainty of the system. The average entropy values ​​of the three modes are then used to obtain the overall cognitive uncertainty entropy of the system. :

[0144] ;

[0145] In the formula, , , These represent the entropy values ​​of the video modality, audio modality, and text modality, respectively.

[0146] Mapping the entropy to [0,1] using a monotonically decreasing function yields the cognitive uncertainty coefficient. ,Right now:

[0147] ;

[0148] in, Represents the baseline entropy value. This represents the slope parameter. When the overall uncertainty of the system is high, the system relies more on data-driven branches; when the uncertainty is low, the decision-making trusts the explicit inferences of the physical model more. This is the slope coefficient.

[0149] Step 1602: Use cognitive uncertainty coefficients to weight and fuse the quantitative risks of the data-driven branch and the physical model branch to determine the risk indicators of substation defects.

[0150] Specifically, the data-driven branch quantifies the defect risk of substation equipment from the perspective of data statistics for each modality; the physical model branch quantifies the defect risk of substation equipment from the perspective of its physical mechanisms.

[0151] In the data-driven branch, consensus weighting is introduced, which uses attention weights to reweight the samples in the loss distribution, i.e.:

[0152] ;

[0153] In the formula, The consensus-weighted conditional value of risk (VoV) characterizes the data-driven branch, with α representing the confidence level (preset to 0.95). As a multimodal fusion feature, it represents the average loss after consensus weighting among samples whose loss exceeds the VaR threshold at confidence level α, and it primarily reflects the extreme defect risk of multimodal consensus in high-altitude environments; The consensus weighting coefficient for the k-th sample; Characterize the defect risk loss value corresponding to the k-th sample, and quantify the potential loss if the sample is misjudged; This is the indicator function. Consensus-based sample weights are added to the traditional CVaR (Conditional Value at Risk) calculation process, making the final risk value more biased towards the prediction results of modes with high consensus.

[0154] In the physical model branch, an analytical model of the device state is established based on the MMC (Modular Multilevel Converter SwitchingFunction) switching function analytical model:

[0155] Bridge arm current: ;

[0156] Capacitor voltage dynamic equation: = ;

[0157] Where N represents the total number of bridge arm submodules; (0 indicates excision, 1 indicates input). Characterizes the branch current of the k-th submodule; Thermal noise figure; Characterizing the total current in the bridge arm, it directly reflects the conduction status of the bridge arm; for example, a submodule of a power equipment may experience an IGBT (Insulated Gate Bipolar Transistor) failure due to high altitude and low temperature. (If it should be 1 but is 0), then Not included Abnormal attenuation may occur, and the faulty submodule can be deduced by checking the current deviation.

[0158] Through uncertainty propagation, the complete probability distribution of key monitored quantities (such as fault current) is obtained. Finally, the standardized third moment (skewness) of this distribution under normal conditions is calculated:

[0159] E[ ]= ;

[0160] Among them, E[ The mathematical expectation operator, i.e., the standardized third moment; Key monitoring parameters characterizing branches of the physical model; Characterizing key monitoring quantities The mean; Characterization submodule capacitor voltage The posterior probability density function is given by μ, where μ and σ are the mean and standard deviation, respectively. Using this as a physical risk indicator can sensitively detect early failures such as asymmetry.

[0161] By using a cognitive uncertainty coefficient, the quantitative risks of the data-driven branch and the physical model branch are weighted and fused to obtain a cognitive adaptive risk index R, which serves as a risk indicator characterizing defects in power equipment.

[0162] R= ;

[0163] in, This indicates that at time t, by The derived entropy of overall system cognition uncertainty.

[0164] In one feasible implementation, the method further includes: determining the root cause type of the power equipment defect risk based on the risk index of the power equipment defect through a sensitivity index, and formulating countermeasures for the defect risk based on the determined root cause type.

[0165] Based on the R-value, the Sobol index is used to conduct global sensitivity analysis to trace the dominant sources of uncertainty and trigger targeted mitigation strategies: for parameter uncertainty, stochastic predictive control with explicit optimization of the variance term is implemented; for observation noise, a neural ordinary differential equation filter that preserves manifold geometry is designed; for model mismatch, domain adaptive alignment with cognitive state distribution as the adversarial target is adopted. The final output provides a decision-making basis that combines quantification of risk and tracing of its causes.

[0166] The root cause types include environmental types, equipment types, and observation noise types, and each root cause type includes data on at least two characterizing attributes; the sensitivity index characterizes the degree of influence of each root cause type on the risk indicator.

[0167] Perform a global sensitivity analysis of Sobol: with risk indicators as output and uncertainty sources of each mode as input;

[0168] Input parameter X=[ ],in Representing environmental disturbance parameters, characterizing the root causes of environmental types, such as =[Wind speed (m / s), snowfall intensity (mm / h), air pressure (kPa)]; This indicates deviations in the equipment's inherent parameters, representing the root cause of the equipment's inherent type, such as... =[Insulation resistance reduction rate of key components (%) / year), loosening amount of mechanical structure (mm)]; This represents the text description quality parameters, characterizing the root cause of the observed noise type, such as... =[Log description length (number of words), percentage of technical terms in the description (%), description fuzziness score], where the fuzziness score can be obtained by scoring the sentence using an NLP (Natural Language Processing) model.

[0169] The input parameter X is sampled within its possible distribution range to generate N sample points. Random fluctuations in high-altitude environments (such as sudden drops in air pressure and equipment aging) are simulated, and SDE source tracing, geometric fusion, and risk calculation are performed on each sample point to obtain the corresponding risk index output.

[0170] Calculate the sensitivity index, first-order index:

[0171] ;

[0172] It measures a single attribute input based on a root type. The proportion of output variance caused by changes in ; The larger the value, the more significant the impact of individual changes in this type of parameter on the risk index R. For example, in high-altitude scenarios... (Environmental parameters) are usually at their maximum (0.4-0.6) because low air pressure and strong winds directly interfere with multimodal data. Represents the i-th type of parameter, Characterization fixed When the value is , Y is relative to The expected value of the condition; Characterization pairs Calculate the variance of the conditional expectation for all possible values; The total variance of all samples Y(k) is represented.

[0173] Total effect index: ;

[0174] It measures the input The proportion of the output variance caused by all contributions to the total variance; Include Individual contributions and Interaction contribution with other parameters (e.g.) low pressure and (The synergistic effect of decreased insulation resistance). For example, in high-altitude scenarios. The total environmental effect may reach 0.7-0.8, due to the significant interaction between environmental parameters and equipment parameters.

[0175] Based on the sensitivity index value ( or ), and propose targeted measures. If the environmental index is the highest, it is speculated that the system risk originates from random environmental disturbances at high altitudes, and consideration should be given to strengthening the system's anti-interference capabilities; if the equipment index is the highest, it is speculated that the equipment's performance is degrading, and consideration should be given to developing a preventive maintenance plan for the equipment; if the text index is the highest, it is speculated that the risk mainly stems from non-standard and inaccurate text records, and consideration should be given to optimizing the recording standards and quality inspection processes of the inspection logs.

[0176] In one feasible implementation, to improve the reliability of the diagnostic results output by the diagnostic device, this embodiment further includes:

[0177] Based on the geometric modulation attention weights between modalities, an arbitration mechanism is determined when consensus conflicts exist between video and audio modalities; for example... Figure 2 As shown, the conflict arbitration mechanism specifically includes:

[0178] The input layer uses geometric modulation attention weights based on continuous input, and the geometric modulation attention weights are derived from intermodal input. In the process, the first attention weights of the audio modality on the video modality are obtained respectively. And the second attention weights of the video modality on the audio modality. ;

[0179] Define a consensus index, monitoring the matrix for two elements representing the degree of mutual attention between video and voiceprint: video's attention to voiceprint, and voiceprint's attention to video; representing respectively:

[0180] Attention to voiceprints in video: ;

[0181] Voiceprint attention to video: ;

[0182] The lower the attention value, the more the modality considers the information provided by the other party to be unreliable.

[0183] The conflict index is determined based on the minimum value between the first attention weight and the second attention weight;

[0184] Conflict Indicators Represented as: ;

[0185] When the conflict index exceeds a preset threshold, arbitration is initiated by sending a high-priority interrupt to the SDE observation model of the text modality, causing it to switch from standby to full-power operation, so as to obtain the corresponding text modality data as arbitration evidence.

[0186] The diagnostic equipment will automatically retrieve textual evidence related to the current conflict event, including: real-time inspection log: the description of the substation equipment currently entered by the operator; historical maintenance records: past fault records and handling solutions for the substation equipment; substation equipment operating procedures: a standard state description of the substation equipment under specific operating conditions; and environmental reports: current environmental text reports such as wind speed, air pressure, and humidity.

[0187] The arbitration conclusion is determined based on the aforementioned arbitration evidence; the arbitration conclusion includes the credibility of the defect diagnosis conclusion based on the video modality and the defect diagnosis conclusion based on the audio modality.

[0188] Based on the acquired arbitration evidence, the text modality acts as the arbitrator, providing a judgment and its confidence level. When the system checks the reliability of the arbitration result itself, it resets and restores the device state; if the result is deemed uncertain, manual intervention is initiated. This embodiment, by introducing a conflict arbitration mechanism, constructs a more complete cross-modal confidence-based collaborative decision-making system.

[0189] The arbitration conclusion obtained through this implementation can adjust the quantification process of defect risk by the data-driven branch and / or physical model branch.

[0190] The proposed conflict arbitration mechanism based on uncertainty perception enables automatic triggering of text modality arbitration by real-time monitoring of the geometric attention weight matrix. This intelligently resolves contradictions in low-level perceptual modalities, rebuilds the consistency of system cognition, and ensures the reliability of decisions under abnormal conditions.

[0191] In one feasible implementation, to further improve the accuracy of diagnostic results, this embodiment also includes: designing a metacognitive federated network guided by uncertainty in edge-cloud collaboration, expanding the focus from single-point substation equipment to the entire distributed network, and realizing the feedback of collective wisdom gathered in the cloud to the defect diagnosis of single-point substation equipment.

[0192] like Figure 3 As shown, a metacognitive federated network architecture guided by uncertainty in edge-cloud collaboration is presented. While the arbitration mechanism effectively solves the immediacy problem of multimodal perception conflicts within a single device, its cognitive boundaries are still limited to local historical and real-time data.

[0193] The implementation of a metacognitive federated network architecture guided by the uncertainty of edge-cloud collaboration includes:

[0194] The uncertainty score of the diagnostic result is determined based on the cognitive uncertainty coefficient.

[0195] Each edge node uses its local multimodal model to perform real-time analysis of the substation equipment status, while the metacognitive module synchronously outputs the diagnostic results and its uncertainty score U_local. When U_local < 0, it indicates that the conclusion is reliable, the edge node independently completed this diagnosis, and the data can be used for subsequent local incremental learning; when U_local ≥ 0, it indicates that the conclusion is questionable.

[0196] When the uncertainty score characterizes a questionable prediction result, the following action is taken:

[0197] The enhanced representation, the uncertainty score, the diagnostic results, the timestamp, and the substation metadata tags are uploaded to the cloud to obtain a global metacognitive model from the cloud. The global metacognitive model uses the data uploaded by each node as samples and performs training on the samples to determine its results. The output of the global metacognitive model is used as soft labels to optimize the defect diagnosis and identification model of the local node.

[0198] To protect privacy and save bandwidth, edge nodes do not upload all data when uploading data (help packet) to the cloud. Instead, they extract a pre-determined enhanced representation F_local, encrypt it with {F_local, U_local, Y_local} (features, uncertainty score, local diagnostic results), add a timestamp and device type metadata tag, and upload it to the cloud as a help packet.

[0199] The cloud securely receives encrypted help packets from various edge nodes, decrypts them, and integrates them into the hard sample dataset D_global_hard. Based on this dataset, the cloud performs specialized training on its maintained global metacognitive model M_meta_global. This training aims to enable M_meta_global to accurately classify hard samples and learn their deep features and patterns. After training, the global metacognitive model shows a significant improvement in the accuracy of its discrimination of difficult samples, and the false positive rate is effectively reduced.

[0200] Training the global metacognitive model can be done as follows: Construct a hard sample dataset, and divide the dataset into a training set (for model parameter updates) and a validation set (for performance evaluation during training) in a ratio (e.g., 7:3), ensuring that the data covers diverse interference scenarios at high altitudes (such as a combination of sandstorms, strong winds, and low air pressure). Optimize and train based on existing defect diagnosis models. In each iteration, input the feature data of a training sample from the dataset (based on information extracted from the help package) into the global metacognitive model to be trained. Use the defect diagnosis results of the substation equipment corresponding to the same training sample, as standardized by manual inspection, as labels to obtain the actual output of the global metacognitive model to be trained. Determine the value of the loss function based on the actual output and labels. When the loss function converges or reaches the preset number of iterations, training ends, resulting in the trained global metacognitive model.

[0201] During the local model fusion phase, edge nodes employ intelligent strategies to integrate the global metacognitive model deployed from the cloud, rather than directly replacing it. Specifically, knowledge distillation technology can be used to treat the global metacognitive model M_meta_global as the teacher model, utilizing its output (especially uncertainty assessment results) as soft labels to guide and optimize the local model. This allows for the absorption of generalized knowledge while retaining adaptability to specific local conditions. Simultaneously, a weighted averaging method can be used to weight and fuse the parameters of the cloud and local models based on model performance metrics, achieving a smooth transition.

[0202] After integration, the nodes' metacognitive abilities evolve, including more accurate self-assessment capabilities, enabling them to identify their own uncertainties earlier and more accurately, and gaining stronger diagnostic capabilities. This process constructs a positive feedback loop from "raising questions" to "learning to answer questions" and then to "capability improvement," continuously strengthening the system's intelligence level.

[0203] In one feasible implementation, to achieve defect early warning, tiered alarms are triggered based on the diagnostic results of substation equipment defects obtained by diagnostic equipment. Specifically, this includes triggering tiered alarms based on the coupling relationship between comprehensive risk indicators and fused uncertainty indicators.

[0204] Data from multiple processing stages in the above implementation of this embodiment is obtained as input. Specifically, a comprehensive risk index is determined based on the uncertainty index and the risk index of the substation defect. A fused uncertainty index is determined based on the optimization results of the uncertainty score during cloud training.

[0205] The fusion uncertainty index is calculated as follows: R = (1 - P_normal) * (1 + U_fused);

[0206] Here, P_normal represents the probability that the model diagnoses the system as "normal"; (1-P_normal) constitutes the baseline risk value, and the lower this value, the higher the confidence that the system is in an abnormal state, and the greater the baseline risk; U_fused represents the overall uncertainty after fusion, and (1+U_fused) serves as the risk multiplier, reflecting the gain in decision risk caused by "unclear understanding". The higher the uncertainty, the more significantly the decision risk is amplified.

[0207] like Figure 4 As shown, in the dynamic triggering of graded alarms, the system does not rely solely on a single risk value for judgment. Instead, it comprehensively considers the coupling relationship between the comprehensive risk index R and the fused uncertainty U_fused, thereby achieving intelligent alarm decision-making. Specifically, this includes:

[0208] If the comprehensive risk index is less than the first preset risk value, a level one alarm is triggered (corresponding to...). Figure 4 Level 1 (and I), a Level 1 alarm can be a suggestive "attention" level;

[0209] If the comprehensive risk index exceeds the first risk preset value but is less than the second risk preset value, and the fused uncertainty index is less than the uncertainty preset value, it indicates that a definite minor anomaly has been detected, and a level two alarm is triggered (corresponding to...). Figure 4 Level 2 and II alarms can be abnormal alarms;

[0210] If the fusion uncertainty index exceeds the preset uncertainty value, regardless of whether the comprehensive risk index is in the medium or high range, a level three alarm will be triggered (corresponding to...). Figure 4 Level 3 and III alarms can prompt maintenance personnel to prioritize troubleshooting.

[0211] If the comprehensive risk index exceeds the second preset risk value and the fused uncertainty index is less than the preset uncertainty value, a level four alarm will be triggered (corresponding to...). Figure 4 Level 4 and IV alarms require immediate attention.

[0212] The risk levels of the level 1 alarm, level 2 alarm, level 3 alarm, and level 4 alarm increase sequentially, and in Figure 4 Different colors are used to distinguish them;

[0213] In one feasible implementation, to enable the prediction of the power grid's operating status, this embodiment further includes: constructing a causal sensing network to predict the operational risks of the power grid where the substation is located.

[0214] like Figure 5As shown, by deeply integrating real-time multimodal data, electrical topology, and physicochemical mechanisms, a leap from "current state diagnosis" to "future risk projection" is achieved, providing causal-level decision-making basis for predictive maintenance. Specifically, this includes steps 51-56:

[0215] Step 51: Construct a connection graph between substation equipment in the power grid; the connection graph uses substation equipment in the power grid as equipment nodes, and constructs the edges of the connection graph based on the electrical topology and / or physical location relationships between the substation equipment; add a time-series attribute to each equipment node, the time-series attribute being determined based on the monitoring data.

[0216] Graph nodes can be transformers, insulators, or other power transmission equipment; specific temporal attributes can be determined based on enhanced characterization, raw monitoring data, and environmental data such as temperature, humidity, and air pressure from external sensors. When initially constructing the connectivity graph, edge weights are determined by the connection distance between graph nodes.

[0217] Step 52: Characterize each variable in the timing data of the device node as a nonlinear function of the causal parent node.

[0218] Applying neural structural equation modeling to time-series attribute data, for each variable in the time-series data... Construct a structural equation and represent its value as the causal parent node PA( The nonlinear function of ), i.e.:

[0219] ;

[0220] in, It is a neural network used to fit nonlinear relationships (such as a multilayer perceptron, MLP). This represents a noise term that is independent of its parent node.

[0221] Step 53: Determine the objective function by minimizing the noise term in the nonlinear function and regularizing the neural network input weights used to fit the nonlinear function.

[0222] In the process of structure learning, the optimization problem is solved by introducing regularization of the neural network input weights:

[0223] +λ ;

[0224] Where G represents the causal graph structure to be discovered, implicitly defined by non-zero connection weights; λ It is an L1 regularization term used to suppress redundant connections in a neural network.

[0225] Step 54: Using the causal rules based on objective laws as constraints, solve the objective function to determine the set of causal parent nodes for each variable.

[0226] Objective laws can be known physical and chemical laws. Digitizing these known physical and chemical laws into causal rules essentially transforms prior qualitative knowledge and quantitative laws into machine-processable mathematical constraints, integrating them into the model's structure learning and parameter estimation processes. The main paradigms can be divided into two categories: hard constraints and soft constraints. Hard constraints include: directional prohibitions (if A cannot cause B, then the directed edge A→B is directly removed from the candidate edge set); functional form constraints (if the known relationship obeys a specific physical law, then the structural equation is forced to adopt a parameterized form); and conservation law embedding (for closed systems that follow conservation laws, constraints are directly introduced as equations into the optimization problem; for example, mass balance is represented as...). It must be strictly satisfied by the structural equations of the relevant variables. Soft constraints include: sign constraints, if the judgment... right If there is a monotonic positive effect, a penalty term can be added to the loss function.

[0227] Step 55: Connect the nodes in the causal parent node set of each variable of the power equipment and the power equipment as edges of the causal perception graph, and construct the causal perception graph with each power equipment as a node of the causal perception graph.

[0228] Step 56: Based on the causal perception map and the defect diagnosis results of the substation equipment, predict the risk of the power grid where the substation equipment is located. And the corresponding maintenance strategy. Specifically, this could involve inputting the defect diagnosis results of the causal sensing substation equipment into the STGNN (Spatio-Temporal Graph Neural Network) model to obtain the model's output prediction results. And the corresponding maintenance strategies; the prediction results may include future risks of the power grid where the substation is located.

[0229] To endow the causal sensing network with autonomous evolution capabilities, this implementation also designs a parameterized decay differential equation. Taking equipment iteration and technology updates as inputs, the decay rate of the edge weights of historical cases is dynamically solved.

[0230] The weights of the edges in a causal perception graph are determined in the following way:

[0231] Based on multidimensional driving factors, a time-varying attenuation coefficient is determined; the multidimensional driving factors include at least the technological generation gap and the determination of the equipment iteration coefficient; the technological generation gap represents the generational gap between the current technology and the technology on which the edges of the causal perception graph are based; the equipment iteration coefficient represents the iteration degree of the substation equipment; and the weights of the edges of the causal perception graph are determined based on the time-varying attenuation coefficient.

[0232] Specifically: assign a dynamic weight to each edge e in the causal sensing network. Its rate of change over time is governed by the following differential equation:

[0233] ;

[0234] in, The time-varying decay rate coefficient is represented by a multidimensional driving factor vector. Decide:

[0235] ;

[0236] in, (·)express A smooth ReLU function is used as the decay rate calculation function to avoid abrupt changes in the decay rate; Indicates a technological gap, Indicates the device iteration coefficient. This indicates the relevant policy status, and ϕ indicates decay behavior. , , , , etc. represent learnable decay behavior parameters.

[0237] By solving the differential equation, at each time step Automatically update the weights of all edges:

[0238] ;

[0239] Edges with weights below a preset threshold will be considered outdated knowledge and automatically downgraded or archived. This allows the system to prioritize calculations based on the latest and most relevant evidence chains in subsequent causal reasoning and risk simulation, ultimately achieving self-purification and continuous evolution of the knowledge base.

[0240] In a feasible implementation model, to further achieve fault early warning, this embodiment further designs a spatiotemporal graph sequence predictor to realize forward extrapolation of fault propagation paths and generation of personalized maintenance strategies. The core lies in using a causal sensing network as input, learning the latent representations of its nodes and edges through an attention-based spatiotemporal graph neural network (STGNN) model, and recursively predicting the graph state sequence for the next τ steps. Its objective function is to minimize the structural and attribute differences between the predicted graph and the real graph.

[0241] ;

[0242] Where H and A represent the node attribute matrix and adjacency matrix, respectively. Indicates the predictor parameters, and These are represented as loss functions for points and edges, respectively. Indicates the time step; , This represents the learnable parameters. Through this process, the model can identify the most likely failure propagation path in the current state. This enables early warning of faults.

[0243] Design a multimodal similar case matching function to perform similarity calculation in the latent space of the spatiotemporal evolution causal perception network. Specifically: Given the current graph state... Its representation in the latent space is as follows , and the historical case library D={( Search for the most similar cases. Furthermore, by calculating the difference map... Accurately pinpoint key differentiators.

[0244] The maintenance strategy generator will integrate all the above information: current status Future risk projection And the differences from historical cases In a defined action space In order to optimize its strategy Π, it will maximize long-term returns:

[0245] ;

[0246] in, For the optimized strategy, With γ as the reward function and γ as the discount factor, the resulting maintenance strategies are no longer uniform, but rather highly personalized and actionable decision recommendations that fully consider the dynamic evolution of the system, future risks, and individual differences, ultimately achieving a closed loop of predictive maintenance at the causal level.

[0247] This application embodiment quantifies the uncertainty of monitoring data from each modality to characterize the degree of variation in monitoring data of substation equipment under high-altitude environments due to interference from various factors. During multimodal monitoring data fusion, the quantified uncertainty is represented on a Riemannian manifold, and the intrinsic connectivity strength between modalities is determined using geodesic distance. Geometric modulation attention fusion is achieved by fusing intrinsic connectivity strength with content similarity. Subsequently, the monitoring data from each modality are weighted and fused while maintaining the geometric structure of the Riemannian manifold, generating an enhanced representation for each modality that incorporates global information. Since the geometric modulation attention fusion stage fuses the intrinsic connectivity strength reflecting the uncertainty formula with the content similarity determined based on the monitoring data—that is, it achieves the fusion of prior strength and actual monitoring data—it can improve the robustness of the diagnostic structure in high-uncertainty scenarios. Furthermore, the enhanced representation of each modality incorporates features from other modalities, further improving the accuracy of the diagnostic results.

[0248] Example 2:

[0249] To address the issue of low accuracy in diagnosing defects in power equipment at high altitudes, and based on the same inventive concept as Embodiment 1, this application also provides a dynamic diagnostic device for defects in high-altitude power equipment using cross-modal analysis.

[0250] A schematic diagram of the specific structure of the device is shown below. Figure 6 As shown, it includes the following functional units:

[0251] The multimodal data acquisition unit 61 is used to acquire monitoring data for each mode of the substation equipment; the modes include a video mode for monitoring the appearance of the substation equipment, an audio mode for collecting the soundprint of the substation equipment, and a text mode for describing the relevant information of the substation equipment.

[0252] The monitoring data statistical analysis unit 62 solves the pre-constructed uncertainty evolution equation for each mode using the monitoring data, and determines the uncertainty covariance matrix of each mode;

[0253] The intrinsic connectivity strength determination unit 63 determines the intrinsic connectivity strength between modes based on the similarity of the uncertainty covariance matrices between different modes; the similarity is based on the geodesic distance characterization of the uncertainty covariance matrix on the Riemannian manifold.

[0254] Attention weight determination unit 64 determines the geometric modulation attention weights between modes based on the inherent connection strength between the modes and the content correlation between the corresponding modes;

[0255] Feature representation unit 65: Weighted fusion of each modality while maintaining the Riemannian manifold geometry through the geometric modulation attention weight, to determine the enhanced representation of each modality that has fused the features of other modalities;

[0256] The model diagnostic unit 66 is used to input the enhanced representation into a pre-trained defect diagnostic model, causing the defect diagnostic identification model to perform the following operations to output the diagnostic results of the substation equipment defects:

[0257] Based on the geometric modulation attention weights between modalities, the cognitive uncertainty coefficient is determined to characterize the overall cognitive uncertainty of each modality;

[0258] The quantitative risks of the data-driven branch and the physical model branch are weighted and fused by cognitive uncertainty coefficients to determine the risk indicators of substation equipment defects. The data-driven branch quantifies the defect risk of substation equipment from the data statistics level of each modality, while the physical model branch quantifies the defect risk of substation equipment from the physical mechanism level.

[0259] In one feasible implementation, the diagnostic device of this embodiment further includes a defect tracing unit, which is used to determine the root cause type of the defect risk of the power equipment based on the risk index of the power equipment defect and through a sensitivity index, and to formulate countermeasures for the defect risk based on the determined root cause type.

[0260] The root cause types include environmental type, device type, and observation noise type, and each root cause type includes data of at least two characterizing attributes;

[0261] The sensitivity index characterizes the degree of influence of each root cause type on the risk indicator;

[0262] The sensitivity index includes the degree of independent impact of data changes based on one root cause type on the risk indicator, and the degree of total impact of data changes based on at least one root cause type on the risk indicator.

[0263] In one feasible implementation, the diagnostic device of this embodiment further includes an arbitration unit, which determines the arbitration mechanism when there is a consensus conflict between the video modality and the audio modality based on the geometric modulation attention weights between the modalities.

[0264] Specifically used for:

[0265] From the geometric modulation attention weights between modalities, the first attention weight of the audio modality to the video modality and the second attention weight of the video modality to the audio modality are obtained respectively; based on the minimum value of the first attention weight and the second attention weight, a conflict index is determined; when the conflict index exceeds a preset threshold, the data of the corresponding text modality is obtained as arbitration evidence; based on the arbitration evidence, an arbitration conclusion is determined; the arbitration conclusion includes the credibility of the defect diagnosis conclusion based on the video modality and the defect diagnosis conclusion based on the audio modality; based on the arbitration conclusion, the quantification process of defect risk by the data-driven branch and / or the physical model branch is adjusted.

[0266] In one feasible implementation, the diagnostic device of this embodiment further includes a cloud interaction module, specifically used for:

[0267] Based on the cognitive uncertainty coefficient, an uncertainty score is determined for the diagnostic result. When the uncertainty score indicates that the prediction result is questionable, the following steps are performed:

[0268] The enhanced representation, the uncertainty score, the diagnostic results, the timestamp, and the substation metadata tags are uploaded to the cloud to obtain a global metacognitive model fed back from the cloud. The global metacognitive model uses the data uploaded by each node as samples and performs training on the samples to determine the results. The output of the global metacognitive model is used as a soft label to optimize the defect diagnosis and identification model of the local node.

[0269] In one feasible implementation, the diagnostic device of this embodiment further includes an alarm unit that triggers a graded alarm based on the coupling relationship between the comprehensive risk index and the fusion uncertainty index.

[0270] Specifically used for:

[0271] If the comprehensive risk index is less than the first preset risk value, a Level 1 alarm is triggered; if the comprehensive risk index exceeds the first preset risk value but is less than the second preset risk value, and the fused uncertainty index is less than the preset uncertainty value, a Level 2 alarm is triggered; if the fused uncertainty index exceeds the preset uncertainty value, a Level 3 alarm is triggered; if the comprehensive risk index exceeds the second preset risk value and the fused uncertainty index is less than the preset uncertainty value, a Level 4 alarm is triggered; the risk levels of the Level 1, Level 2, Level 3, and Level 4 alarms increase sequentially; the fused uncertainty index is the optimized result of the uncertainty score during cloud training; the comprehensive risk index is determined based on the uncertainty index and the risk index of the substation equipment defects.

[0272] In one feasible implementation, the diagnostic device of this embodiment further includes a prediction unit for constructing a causal sensing network to predict the operational risks of the power grid where the substation is located.

[0273] Specifically used for:

[0274] A connection graph is constructed among substation equipment in the power grid. The connection graph uses substation equipment as equipment nodes and its edges are constructed based on the electrical topology and / or physical location relationships between the substation equipment. A time-series attribute is added to each equipment node, determined based on the monitoring data. Each variable in the time-series data of the equipment node is represented as a nonlinear function of a causal parent node. An objective function is determined by minimizing the noise term in the nonlinear function and regularizing the input weights of the neural network used to fit the nonlinear function. The objective function is solved using causal rules based on objective laws as constraints to determine the set of causal parent nodes for each variable. The nodes connecting the substation equipment and the set of causal parent nodes for each variable of the substation equipment are used as edges in a causal perception graph. A causal perception graph is constructed using each substation equipment as a node in the causal perception graph. Based on the causal perception graph and the defect diagnosis results of the substation equipment, the risk of the power grid where the substation equipment is located is predicted.

[0275] In one feasible implementation, the diagnostic device of this embodiment further includes a weight determination unit, used to determine the edge weights of the causal sensing network, specifically used to determine the time-varying attenuation coefficient based on multidimensional driving factors; the multidimensional driving factors include at least the technological generation gap and the equipment iteration coefficient determination; the technological generation gap represents the generational gap between the current technology and the technology on which the edges of the causal sensing graph are determined; the equipment iteration coefficient represents the iteration degree of the substation equipment; and the weights of the edges of the causal sensing graph are determined based on the time-varying attenuation coefficient.

[0276] This application embodiment quantifies the uncertainty of monitoring data from each modality to characterize the degree of variation in monitoring data of substation equipment under high-altitude environments due to interference from various factors. During multimodal monitoring data fusion, the quantified uncertainty is represented on a Riemannian manifold, and the intrinsic connectivity strength between modalities is determined using geodesic distance. Geometric modulation attention fusion is achieved by fusing intrinsic connectivity strength with content similarity. Subsequently, the monitoring data from each modality are weighted and fused while maintaining the geometric structure of the Riemannian manifold, generating an enhanced representation for each modality that incorporates global information. Since the geometric modulation attention fusion stage fuses the intrinsic connectivity strength reflecting the uncertainty formula with the content similarity determined based on the monitoring data—that is, it achieves the fusion of prior strength and actual monitoring data—it can improve the robustness of the diagnostic structure in high-uncertainty scenarios. Furthermore, the enhanced representation of each modality incorporates features from other modalities, further improving the accuracy of the diagnostic results.

[0277] Based on the same inventive concept as the foregoing embodiments of this application, this application also provides a computing device.

[0278] like Figure 7 As shown, the computing device includes a memory 71 and a processor 72. The memory 71 can be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device. The memory 71 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0279] The processor 72, coupled to the memory 71, is used to execute the computer program stored in the memory 71 to perform a cross-modal analysis method for dynamic diagnosis of defects in high-altitude substations as described in the foregoing embodiments.

[0280] When processor 72 executes the computer program to perform a dynamic diagnostic method for defects in high-altitude substations using cross-modal analysis, it quantifies the uncertainty of monitoring data from each modality to characterize the degree of variation in monitoring data of substations under high-altitude environments due to interference from various factors. During multimodal monitoring data fusion, the quantified uncertainty is represented on a Riemannian manifold, and the intrinsic connectivity strength between modalities is determined using geodesic distance. Geometric modulation attention fusion is achieved by fusing intrinsic connectivity strength with content similarity. Subsequently, the monitoring data from each modality are weighted and fused while maintaining the geometric structure of the Riemannian manifold, generating an enhanced representation for each modality that incorporates global information. Since the geometric modulation attention fusion stage fuses the intrinsic connectivity strength reflecting the uncertainty formula with the content similarity determined based on the monitoring data—that is, it achieves the fusion of prior strength and actual monitoring data—it can improve the robustness of the diagnostic structure in high-uncertainty scenarios. Furthermore, the enhanced representation of each modality incorporates features from other modalities, further improving the accuracy of the diagnostic results.

[0281] When the processor 72 executes the computer program in the memory 71, in addition to the functions described above, it can also perform other functions, as detailed in the descriptions of the preceding embodiments.

[0282] Furthermore, such as Figure 7 As shown, the computing device also includes other components such as a display 74, a communication component 73, a power supply component 75, and an audio component 76. Figure 7 The diagram only shows some components and does not mean that the computing device includes only these components. Figure 7 The components shown.

[0283] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a computer, can implement the methods provided in the above embodiments.

[0284] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0285] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.

[0286] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A dynamic diagnostic method for defects in high-altitude substations using cross-modal analysis, characterized in that, include: Acquire monitoring data for each mode of the power equipment; the modes include video modes for monitoring the appearance of the power equipment, audio modes for collecting the acoustic signature of the power equipment, and text modes for describing relevant information about the power equipment. The uncertainty evolution equation for each mode is solved using the monitoring data to determine the uncertainty covariance matrix of each mode; The intrinsic connection strength between modes is determined based on the similarity of the uncertainty covariance matrices between different modes. The similarity is based on the geodesic distance representation of the uncertainty covariance matrix on the Riemannian manifold; Based on the intrinsic connection strength between the modes and the content correlation between the corresponding modes, the geometric modulation attention weights between the modes are determined. The geometrically modulated attention weights are used to perform weighted fusion of each modality while preserving the Riemannian manifold geometry, in order to determine the enhanced representation of each modality that incorporates features from other modalities. The enhanced representation is input into the pre-trained defect diagnosis model, causing the defect diagnosis and identification model to perform the following operations to output the diagnostic results of the substation equipment defects: Based on the geometric modulation attention weights between modalities, the cognitive uncertainty coefficient is determined to characterize the overall cognitive uncertainty of each modality; By using cognitive uncertainty coefficients to weight and fuse the quantitative risks of the data-driven branch and the physical model branch, risk indicators for substation equipment defects can be determined. The data-driven branch quantifies the defect risk of substation equipment from the data statistics level of each modality; The physical model branch quantifies the defect risks of power equipment from the perspective of the physical mechanism of the equipment.

2. The method for dynamic diagnosis of defects in high-altitude substations using cross-modal analysis according to claim 1, characterized in that, The method further includes: determining the root cause type of power equipment defect risk based on the risk index of power equipment defects through a sensitivity index, and formulating countermeasures for defect risk based on the determined root cause type.

3. The method for dynamic diagnosis of defects in high-altitude substations using cross-modal analysis according to claim 1, characterized in that, The method further includes: determining an arbitration mechanism for consensus conflicts between video and audio modalities based on inter-modal geometric modulation attention weights; specifically including: From the intermodal geometric modulation attention weights, obtain the first attention weight of the audio modality on the video modality and the second attention weight of the video modality on the audio modality. The conflict index is determined based on the minimum value between the first attention weight and the second attention weight; When the conflict index exceeds a preset threshold, the corresponding text modality data is acquired as arbitration evidence. The arbitration conclusion is determined based on the aforementioned arbitration evidence; the arbitration conclusion includes the credibility of the defect diagnosis conclusion based on the video modality and the defect diagnosis conclusion based on the audio modality. The process of quantifying defect risk is adjusted based on the arbitration conclusion, adjusting the data-driven branch and / or physical model branch.

4. The method for dynamic diagnosis of defects in high-altitude substations using cross-modal analysis according to claim 1, characterized in that, The method further includes determining an uncertainty score for the diagnostic result based on the cognitive uncertainty coefficient, and when the uncertainty score indicates that the prediction result is questionable, performing the following: The enhanced representation, the uncertainty score, the diagnostic results, the timestamp, and the substation metadata tags are uploaded to the cloud to obtain the global metacognitive model fed back from the cloud; the global metacognitive model uses the data uploaded by each node as samples and performs training on the samples to determine the model. The output of the global metacognitive model is used as a soft label to optimize the defect diagnosis and identification model of the local node.

5. The method for dynamic diagnosis of defects in high-altitude substations using cross-modal analysis according to claim 4, characterized in that, The method further includes: triggering tiered alarms based on the coupling relationship between comprehensive risk indicators and fused uncertainty indicators; specifically including: If the comprehensive risk index is less than the first preset risk value, a level one alarm is triggered; If the comprehensive risk index exceeds the first risk preset value but is less than the second risk preset value, and the fused uncertainty index is less than the uncertainty preset value, then a level two alarm is triggered. If the fusion uncertainty index exceeds the preset uncertainty value, a level three alarm will be triggered; If the comprehensive risk index exceeds the second risk preset value and the fused uncertainty index is less than the uncertainty preset value, a level four alarm will be triggered. The risk levels of the first-level alarm, second-level alarm, third-level alarm, and fourth-level alarm increase sequentially. The fusion uncertainty index is the optimized result of the uncertainty score during cloud training; the comprehensive risk index is determined based on the uncertainty index and the risk index of the substation equipment defects.

6. The method for dynamic diagnosis of defects in high-altitude substations using cross-modal analysis according to claim 1, characterized in that, The method also includes constructing a causal sensing network to predict the operational risks of the power grid where the substation is located, specifically including: A connection graph is constructed between substation equipment in the power grid; the connection graph uses substation equipment in the power grid as equipment nodes, and the edges of the connection graph are constructed based on the electrical topology and / or physical location relationships between the substation equipment; a time-series attribute is added to each equipment node, and the time-series attribute is determined based on the monitoring data; Each variable in the time-series data of the device node is represented as a nonlinear function of the causal parent node; The objective function is determined by minimizing the noise term in the nonlinear function and regularizing the input weights of the neural network used to fit the nonlinear function; Using causal rules based on objective laws as constraints, the objective function is solved to determine the set of causal parent nodes for each variable; Each node in the set of causal parent nodes connecting the power equipment and each variable of the power equipment is used as an edge of the causal perception graph. The causal perception graph is constructed by using each power equipment as a node of the causal perception graph. Based on the causal perception map and the defect diagnosis results of the substation equipment, the risk of the power grid where the substation equipment is located is predicted.

7. The method for dynamic diagnosis of defects in high-altitude substations using cross-modal analysis according to claim 6, characterized in that, The weights of the edges in the causal perception graph are determined in the following way: Based on multidimensional driving factors, a time-varying decay coefficient is determined; the multidimensional driving factors include at least the technological generation gap and the determination of equipment iteration coefficients; the technological generation gap represents the generational gap between the current technology and the technology on which the edge of the causal perception graph is based; The equipment iteration coefficient characterizes the degree of iteration of the substation equipment; The weights of the edges in the causal sensing graph are determined based on the time-varying decay coefficient.

8. A dynamic diagnostic device for defects in high-altitude substations using cross-modal analysis, characterized in that, include: A multimodal data acquisition unit is used to acquire monitoring data for each mode of the substation equipment; the modes include a video mode for monitoring the appearance of the substation equipment, an audio mode for collecting the acoustic signature of the substation equipment, and a text mode for describing the relevant information of the substation equipment. The monitoring data statistical analysis unit solves the pre-constructed uncertainty evolution equation for each mode using the monitoring data, and determines the uncertainty covariance matrix of each mode; The intrinsic connectivity strength determination unit determines the intrinsic connectivity strength between modes based on the similarity of the uncertainty covariance matrices between different modes. The similarity is based on the geodesic distance representation of the uncertainty covariance matrix on the Riemannian manifold; The attention weight determination unit determines the geometric modulation attention weights between modes based on the inherent connection strength between the modes and the content correlation between the corresponding modes; Feature representation unit; The geometrically modulated attention weights are used to perform weighted fusion of each modality while preserving the Riemannian manifold geometry, in order to determine the enhanced representation of each modality that incorporates features from other modalities. The model diagnostic unit is used to input the enhanced representation into a pre-trained defect diagnostic model, causing the defect diagnostic identification model to perform the following operations to output the diagnostic results of the substation equipment defects: Based on the geometric modulation attention weights between modalities, the cognitive uncertainty coefficient is determined to characterize the overall cognitive uncertainty of each modality; By using cognitive uncertainty coefficients to weight and fuse the quantitative risks of the data-driven branch and the physical model branch, risk indicators for substation equipment defects can be determined. The data-driven branch quantifies the defect risk of substation equipment from the data statistics level of each modality; The physical model branch quantifies the defect risks of power equipment from the perspective of the physical mechanism of the equipment.

9. The high-altitude substation defect dynamic diagnosis device based on cross-modal analysis according to claim 8, characterized in that, The device also includes a defect tracing unit, which determines the root cause type of the defect risk of the power equipment based on the risk index of the power equipment defect and through the sensitivity index, and formulates countermeasures for the defect risk based on the determined root cause type. The root cause types include environmental type, device type, and observation noise type, and each root cause type includes data of at least two characterizing attributes; The sensitivity index characterizes the degree of influence of each root cause type on the risk indicator; The sensitivity index includes the degree of independent impact of data changes based on one root cause type on the risk indicator, and the degree of total impact of data changes based on at least one root cause type on the risk indicator.

10. The high-altitude substation defect dynamic diagnosis device based on cross-modal analysis according to claim 8, characterized in that, The device also includes an arbitration unit, which determines the arbitration mechanism when there is a consensus conflict between the video modality and the audio modality based on the geometric modulation attention weights between modalities. Specifically used for: From the intermodal geometric modulation attention weights, obtain the first attention weight of the audio modality on the video modality and the second attention weight of the video modality on the audio modality. The conflict index is determined based on the minimum value between the first attention weight and the second attention weight; When the conflict index exceeds a preset threshold, the corresponding text modality data is acquired as arbitration evidence. The arbitration conclusion is determined based on the aforementioned arbitration evidence; the arbitration conclusion includes the credibility of the defect diagnosis conclusion based on the video modality and the defect diagnosis conclusion based on the audio modality. The process of quantifying defect risk is adjusted based on the arbitration conclusion, adjusting the data-driven branch and / or physical model branch.

Citation Information

Patent Citations

  • Pumped storage unit fault diagnosis method based on multi-modal data fusion

    CN121117736A

  • Defect image enhancement method integrating reasoning and generation

    CN121414610A