Multimodal-based substation operation regulation and control method and apparatus
By employing a substation operation control method enhanced with multimodal data fusion and graph neural networks, the problem of low control accuracy caused by independent operation of monitoring equipment has been solved, achieving efficient multimodal adaptability and accurate control of substation operation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GUANGDONG POWER GRID CO LTD
- Filing Date
- 2025-02-15
- Publication Date
- 2026-04-30
AI Technical Summary
In existing substation operation and control methods, the independent operation of each monitoring device lacks coordination, resulting in low efficiency in comprehensive analysis of monitoring data and low accuracy in control.
A multimodal substation operation and control method is adopted. By acquiring acoustic features, visible light features and infrared features, attention weights are calculated and fused using a self-attention mechanism, and feature enhancement is performed based on a graph neural network. Combined with a pre-set multi-task loss training joint control model, a comprehensive analysis of multimodal data is achieved.
It improves the accuracy and multimodal adaptability of substation operation and control, meets different task requirements, and enhances the accuracy of control.
Smart Images

Figure CN2025077507_30042026_PF_FP_ABST
Abstract
Description
A multimodal substation operation control method and device Technical Field
[0001] This application relates to the field of substation operation and control, and in particular to a substation operation and control method and device based on multimodal operation. Background Technology
[0002] As a crucial node in the power system, the safe and stable operation of substations is vital to improving the overall reliability of the power grid. Therefore, it is necessary to improve the timeliness and accuracy of substation operation control. In existing technologies, substation operation control mainly relies on various monitoring devices within the substation, such as surveillance equipment, sensors, and acoustic signature detectors. However, these monitoring devices often operate independently, lacking coordination, resulting in low efficiency in the comprehensive analysis of monitoring data and low accuracy in substation control. Therefore, improving the accuracy of substation operation control remains a pressing technical problem that needs to be solved. Summary of the Invention
[0003] This application provides a substation operation control method and device based on multimodal operation to solve the technical problem of insufficient accuracy in existing substation control.
[0004] According to a first aspect of the embodiments of this application, a substation operation control method based on multimodal operation is provided, comprising:
[0005] The historical monitoring characteristics of the substation to be regulated are obtained; wherein, the historical monitoring characteristics include acoustic signature characteristics, visible light characteristics, and infrared characteristics;
[0006] Based on the self-attention mechanism, the attention weights of the voiceprint features, visible light features and infrared features are calculated respectively, and the voiceprint features, visible light features and infrared features are fused according to the attention weights to obtain multimodal fused features;
[0007] Based on a graph neural network, the multimodal fusion features are enhanced to obtain fusion graph features;
[0008] Based on the fusion graph features and the preset multi-task loss, multiple preset parallel models are trained to obtain a joint control model; wherein, the multiple parallel models are based on multiple preset tasks; the multiple tasks include equipment anomaly detection, equipment classification, and equipment status prediction.
[0009] The real-time monitoring data of the substation to be regulated is obtained, the real-time monitoring data is input into the joint regulation model to obtain the operation decision, and the substation to be regulated is regulated according to the operation decision.
[0010] This application first acquires the historical monitoring features of the substation to be regulated, calculates the attention weights of each feature based on a self-attention mechanism, and then fuses the features to obtain multimodal fusion features. This can comprehensively integrate monitoring features from different aspects and modes of the substation, improving the multimodal adaptability of substation operation and regulation. Then, it enhances the multimodal fusion features based on graph neural networks to obtain fusion graph features, which can further utilize graph structures to capture and enhance multimodal features. Finally, it trains multiple parallel models based on the fusion graph features and a preset multi-task loss to obtain a joint regulation model. By first dividing the substation operation and regulation into different situations and using different models for training, it can meet the different task requirements of substation operation and regulation, thereby improving the accuracy of substation operation and regulation when making operation decisions based on real-time monitoring data and regulating the substation.
[0011] In some embodiments of this application, obtaining the historical monitoring characteristics of the substation to be regulated specifically includes:
[0012] Acquire historical monitoring data of the substation to be regulated; wherein, the historical monitoring data includes acoustic signature data, visible light images, and infrared images;
[0013] Based on a bidirectional recurrent neural network, temporal features are extracted from the voiceprint data to obtain voiceprint features;
[0014] Based on a self-attention image encoder, image features are extracted from the visible light image to obtain visible light features;
[0015] Based on a convolutional neural network, image features are extracted from the infrared image to obtain infrared features.
[0016] This application first acquires historical monitoring data, and then extracts features from voiceprint data, visible light images, and infrared images based on different types of neural networks to obtain voiceprint features, visible light features, and infrared features. By selecting different feature extraction networks based on the data category, the corresponding features can be extracted more accurately, thereby improving the accuracy of the acquired features.
[0017] In some embodiments of this application, the step of fusing the voiceprint features, the visible light features, and the infrared features according to the attention weights to obtain multimodal fused features specifically includes:
[0018] Based on a convolutional neural network, the dimensions of the voiceprint feature and the infrared feature are unified to the dimension of the visible light feature to obtain the first voiceprint feature and the first infrared feature.
[0019] Based on the attention weights, and using a multi-layer self-attention encoder, the first voiceprint feature, the visible light feature, and the first infrared feature are spliced and fused to obtain a multimodal fused feature.
[0020] This application first unifies the dimensions of voiceprint features and infrared features into the dimensions of visible light features based on convolutional neural networks, which can ensure data consistency and prevent data dimension conflicts. Then, based on attention weights and multi-layer self-attention encoders, the first voiceprint features, visible light features and the first infrared features are spliced and fused to obtain multimodal fusion features, which can integrate monitoring data of different types and modes in substations and improve the multimodal adaptability of substation operation and control.
[0021] In some embodiments of this application, the enhancement of the multimodal fusion features based on graph neural networks to obtain fused graph features specifically includes:
[0022] Using the time step of the multimodal fusion feature as a graph node, a first multimodal fusion graph corresponding to the multimodal fusion feature is generated;
[0023] Based on a preset edge connection threshold, edges between graph nodes of the first multimodal fusion graph are removed to obtain the second multimodal fusion graph.
[0024] Based on the self-attention mechanism, the edge weights between each graph node in the second multimodal fusion graph are calculated, and the second multimodal fusion graph is enhanced based on the edge weights and a multilayer graph convolutional network to obtain fusion graph features.
[0025] This application first uses the time steps of the multimodal fusion features as graph nodes to generate a first multimodal fusion graph, which can capture multimodal features using graph structure. Then, based on a preset edge connection threshold, some edges are removed to obtain a second multimodal fusion graph, which can increase the weight of edges that meet the conditions. Then, edge weights are calculated based on a self-attention mechanism, and the second multimodal fusion graph is enhanced based on the edge weights and a multi-layer graph convolutional network. Through the graph convolutional network, the relationship between different modalities and different time steps can be captured, improving the matching degree between the output fusion graph features and the actual situation.
[0026] In some embodiments of this application, the step of training multiple preset parallel models based on the fusion graph features and a preset multi-task loss to obtain a joint control model specifically includes:
[0027] The output types of the multiple parallel models are set according to the task categories of the multiple tasks;
[0028] Using the fused graph features as model input, and under the premise of the output type setting, the multiple parallel models are trained based on the preset multi-task loss to obtain a joint control model;
[0029] The preset multi-task loss is specifically: L=λ1*L anomaly +λ2*Ldevice +λ3*L state ;
[0030] Where L is the multi-task loss, L anomaly L device, L state These represent equipment anomaly loss, equipment type loss, and equipment state loss, respectively, with λ1, λ2, and λ3 being the weights of equipment anomaly loss, equipment type loss, and equipment state loss, respectively.
[0031] This application first sets the output types of multiple parallel models according to the task categories of multiple tasks. By differentiating the models in terms of output type, a single model can specialize in a single task, which is beneficial to improving the accuracy of the model. Then, using the fusion graph features as the model input, multiple parallel models are trained based on the preset multi-task loss to obtain a joint control model, which can meet the different task requirements of substation operation control. In this way, when performing substation operation control based on the joint control model, the accuracy of substation operation control is improved.
[0032] According to a second aspect of the embodiments of this application, a substation operation control device based on multimodality is provided, including a feature acquisition module, a feature fusion module, a feature enhancement module, a model training module, and an operation control module;
[0033] The feature acquisition module is used to acquire historical monitoring features of the substation to be regulated; wherein, the historical monitoring features include acoustic signature features, visible light features, and infrared features;
[0034] The feature fusion module is used to calculate the attention weights of the voiceprint features, visible light features and infrared features respectively based on the self-attention mechanism, and fuse the voiceprint features, visible light features and infrared features according to the attention weights to obtain multimodal fused features;
[0035] The feature enhancement module is used to enhance the multimodal fusion features based on a graph neural network to obtain fusion graph features;
[0036] The model training module is used to train multiple preset parallel models based on the fusion graph features and a preset multi-task loss to obtain a joint control model; wherein, the multiple parallel models are based on multiple preset tasks; the multiple tasks include equipment anomaly detection, equipment classification, and equipment status prediction.
[0037] The operation and control module is used to acquire real-time monitoring data of the substation to be controlled, input the real-time monitoring data into the joint control model to obtain operation decisions, and control the substation to be controlled according to the operation decisions.
[0038] In some embodiments of this application, the feature acquisition module includes a historical data acquisition unit, a voiceprint feature extraction unit, a visible light feature extraction unit, and an infrared feature extraction unit;
[0039] The historical data acquisition unit is used to acquire historical monitoring data of the substation to be regulated; wherein, the historical monitoring data includes acoustic data, visible light images and infrared images;
[0040] The voiceprint feature extraction unit is used to extract temporal features from the voiceprint data based on a bidirectional recurrent neural network to obtain voiceprint features.
[0041] The visible light feature extraction unit is used to extract image features from the visible light image based on a self-attention image encoder to obtain visible light features;
[0042] The infrared feature extraction unit is used to extract image features from the infrared image based on a convolutional neural network to obtain infrared features.
[0043] In some embodiments of this application, the feature fusion module includes a dimension unification unit and a feature fusion unit;
[0044] The dimension unification unit is used to unify the dimension of the voiceprint feature and the dimension of the infrared feature into the dimension of the visible light feature based on a convolutional neural network, thereby obtaining the first voiceprint feature and the first infrared feature.
[0045] The feature fusion unit is used to splice and fuse the first voiceprint feature, the visible light feature, and the first infrared feature based on the attention weight and a multi-layer self-attention encoder to obtain a multimodal fused feature.
[0046] In some embodiments of this application, the feature enhancement module includes a fusion graph generation unit, a fusion edge culling unit, and a fusion graph enhancement unit;
[0047] The fusion graph generation unit is used to generate a first multimodal fusion graph corresponding to the multimodal fusion feature, using the time step of the multimodal fusion feature as a graph node.
[0048] The fusion edge removal unit is used to remove edges between graph nodes of the first multimodal fusion graph based on a preset edge connection threshold to obtain a second multimodal fusion graph.
[0049] The fusion graph enhancement unit is used to calculate the edge weights between each graph node in the second multimodal fusion graph based on a self-attention mechanism, and enhance the second multimodal fusion graph based on the edge weights and a multilayer graph convolutional network to obtain fusion graph features.
[0050] In some embodiments of this application, the model training module includes a type setting unit and a model training unit;
[0051] The type setting unit is used to set the output type of the multiple parallel models according to the task categories of the multiple tasks;
[0052] The model training unit is used to train the multiple parallel models based on the preset multi-task loss, with the fusion graph features as model input and the output type set, to obtain a joint control model.
[0053] The preset multi-task loss is specifically: L=λ1*L anomaly +λ2*L device +λ3*L state ;
[0054] Where L is the multi-task loss, L anomaly ,L device L state These represent equipment anomaly loss, equipment type loss, and equipment state loss, respectively, with λ1, λ2, and λ3 being the weights of equipment anomaly loss, equipment type loss, and equipment state loss, respectively.
[0055] This application first acquires the historical monitoring features of the substation to be regulated, calculates the attention weights of each feature based on a self-attention mechanism, and then fuses the features to obtain multimodal fusion features. This can comprehensively integrate monitoring features from different aspects and modes of the substation, improving the multimodal adaptability of substation operation and regulation. Then, it enhances the multimodal fusion features based on graph neural networks to obtain fusion graph features, which can further utilize graph structures to capture and enhance multimodal features. Finally, it trains multiple parallel models based on the fusion graph features and a preset multi-task loss to obtain a joint regulation model. By first dividing the substation operation and regulation into different situations and using different models for training, it can meet the different task requirements of substation operation and regulation, thereby improving the accuracy of substation operation and regulation when making operation decisions based on real-time monitoring data and regulating the substation. Attached Figure Description
[0056] Figure 1: A flowchart illustrating a multimodal substation operation control method according to certain embodiments of this application;
[0057] Figure 2: A modular structure diagram of a substation operation control device based on multimodal operation as shown in some embodiments of this application. Detailed Implementation
[0058] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below in conjunction with the accompanying drawings are exemplary and are only used to explain some embodiments of this application, and should not be construed as limiting the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments shown in this application without inventive effort are within the protection scope of this application.
[0059] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, unless otherwise explicitly specified, "a plurality of" or "several" means two or more.
[0060] Substation operation control primarily relies on various monitoring devices within the substation, such as surveillance equipment, sensors, and acoustic signature detectors. However, these devices often operate independently, lacking coordination, resulting in low efficiency in the comprehensive analysis of monitoring data. Furthermore, existing substation operation control methods typically focus only on monitoring data under a single mode, lacking the integration and analysis of multi-modal monitoring data, leading to low accuracy in substation operation control. Therefore, improving the accuracy of substation operation control remains a critical technical problem that needs to be addressed.
[0061] Based on the above technical background, please refer to Figure 1. This application provides a substation operation control method based on multi-mode operation, including steps S101 to S105, each step of which is as follows:
[0062] Step S101: Obtain the historical monitoring characteristics of the substation to be regulated; wherein the historical monitoring characteristics include acoustic signature characteristics, visible light characteristics and infrared characteristics.
[0063] In some embodiments of this application, obtaining the historical monitoring characteristics of the substation to be regulated specifically includes:
[0064] Acquire historical monitoring data of the substation to be regulated; wherein, the historical monitoring data includes acoustic signature data, visible light images, and infrared images;
[0065] Based on a bidirectional recurrent neural network, temporal features are extracted from the voiceprint data to obtain voiceprint features;
[0066] Based on a self-attention image encoder, image features are extracted from the visible light image to obtain visible light features;
[0067] Based on a convolutional neural network, image features are extracted from the infrared image to obtain infrared features.
[0068] Specifically, the voiceprint data in the historical monitoring data is represented as X[b,1,T], the visible light image is represented as V[b,3,h,w], and the infrared image is represented as I[b,1,h,w], where b is the amount of data processed in a single session, T is the audio duration, h and w are the height and width of the image, respectively, and the second digit in each data representation is the original dimension corresponding to each data.
[0069] In some embodiments of this application, the step of extracting temporal features from the voiceprint data X[b,1,T] based on a bidirectional recurrent neural network to obtain voiceprint features specifically involves:
[0070] Adaptive time-frequency analysis is performed on the voiceprint data X[b,1,T] to obtain the first voiceprint time-frequency feature;
[0071] Based on the mask, the first voiceprint time-frequency feature is augmented in frequency and time to obtain the second voiceprint time-frequency feature. Then, based on two-dimensional convolution, the local time-frequency feature of the second voiceprint time-frequency feature is extracted to obtain the third voiceprint time-frequency feature.
[0072] Based on a bidirectional recurrent neural network, temporal features are extracted from the third voiceprint time-frequency features to obtain the initial voiceprint features;
[0073] The initial voiceprint features are averaged and pooled to obtain voiceprint features F. x [b,D x ] = H avg [b,2*D].
[0074] Among them, D x denoted as the voiceprint feature dimension, and D as the hidden state dimension of the bidirectional recurrent neural network.
[0075] In some embodiments of this application, the bidirectional recurrent neural network includes, but is not limited to, bidirectional RNN and bidirectional LSTM, with bidirectional LSTM being the preferred embodiment.
[0076] In some embodiments of this application, the step of extracting image features from the visible light image based on a self-attention image encoder to obtain visible light features specifically involves:
[0077] Based on a self-attention image encoder, image features are extracted from the visible light image V[b,3,h,w] to obtain visible light features F. v [b,N,D v ].
[0078] Where N is the number of patches, and N is related to the height and width h and w of the image, as well as the preset patch size of the self-attention image encoder, D v Dimensions represent the features of a visible light image.
[0079] In some embodiments of this application, a preferred embodiment of the self-attention image encoder is the Vision Transformer (ViT).
[0080] In some embodiments of this application, the step of extracting image features from the infrared image based on a convolutional neural network to obtain infrared features specifically involves:
[0081] Based on a convolutional neural network, image features are extracted from the infrared image I[b,1,h,w] to obtain infrared features F. i [b,D i ].
[0082] Among them, D i This represents the feature dimension of the infrared image.
[0083] In some embodiments of this application, the preferred embodiment of the convolutional neural network is a CNN.
[0084] This application first acquires historical monitoring data, and then extracts features from voiceprint data, visible light images, and infrared images based on different types of neural networks to obtain voiceprint features, visible light features, and infrared features. By selecting different feature extraction networks based on the data category, the corresponding features can be extracted more accurately, thereby improving the accuracy of the acquired features.
[0085] Step S102: Based on the self-attention mechanism, calculate the attention weights of the voiceprint features, the visible light features, and the infrared features respectively, and fuse the voiceprint features, the visible light features, and the infrared features according to the attention weights to obtain multimodal fused features.
[0086] In some embodiments of this application, the step of fusing the voiceprint features, the visible light features, and the infrared features according to the attention weights to obtain multimodal fused features specifically includes:
[0087] Based on a convolutional neural network, the dimensions of the voiceprint feature and the infrared feature are unified to the dimension of the visible light feature to obtain the first voiceprint feature and the first infrared feature.
[0088] Based on the attention weights, and using a multi-layer self-attention encoder, the first voiceprint feature, the visible light feature, and the first infrared feature are spliced and fused to obtain a multimodal fused feature.
[0089] In some embodiments of this application, the step of unifying the dimensions of the voiceprint features and the infrared features based on a convolutional neural network to the visible light features to obtain the first voiceprint features and the first infrared features specifically involves:
[0090] Among them, F x-seq [b,N,D v ],F i-seq [b,N,D i These are the first voiceprint feature and the first infrared feature, respectively; F x [b,N,D x / N],F i [b,N,D i [ / N] represents the intermediate voiceprint feature and the intermediate infrared feature, respectively; resize represents the sampling operation; CNN represents the variable dimension operation based on convolutional neural network.
[0091] In some embodiments of this application, the step of splicing and fusing the first voiceprint feature, the visible light feature, and the first infrared feature based on the attention weight and a multi-layer self-attention encoder to obtain a multimodal fused feature specifically involves:
[0092] The first voiceprint feature F x-seq [b,N,D v The visible light feature F v [b,N,D v ] and the first infrared feature F i-seq [b,N,D i Then, concatenation is performed at dim=1 to obtain the multimodal concatenated feature F. cat [b,3*N,D v Based on a multi-layer self-attention encoder, multimodal coding features F are obtained. trans [b,3*N,D t ]; where D t For the hidden state dimension of a multi-layer self-attention encoder;
[0093] The attention weight {W x [b,1],W v [b,1],W i [b,1]} and the multimodal coding feature F trans [b,3*N,D t Corresponding multiplication yields the multimodal fusion feature F. fused [b,N,D t ].
[0094] This application first unifies the dimensions of voiceprint features and infrared features into the dimensions of visible light features based on convolutional neural networks, which can ensure data consistency and prevent data dimension conflicts. Then, based on attention weights and multi-layer self-attention encoders, the first voiceprint features, visible light features and the first infrared features are spliced and fused to obtain multimodal fusion features, which can integrate monitoring data of different types and modes in substations and improve the multimodal adaptability of substation operation and control.
[0095] Step S103: Based on the graph neural network, the multimodal fusion features are enhanced to obtain fusion graph features.
[0096] In some embodiments of this application, the enhancement of the multimodal fusion features based on graph neural networks to obtain fused graph features specifically includes:
[0097] Using the time step of the multimodal fusion feature as a graph node, a first multimodal fusion graph corresponding to the multimodal fusion feature is generated;
[0098] Based on a preset edge connection threshold, edges between graph nodes in the first multimodal fusion graph are removed to obtain a second multimodal fusion graph.
[0099] Based on the self-attention mechanism, the edge weights between each graph node in the second multimodal fusion graph are calculated, and the second multimodal fusion graph is enhanced based on the edge weights and a multilayer graph convolutional network to obtain fusion graph features.
[0100] In some embodiments of this application, the preferred value of the preset edge connection threshold k is N / 4.
[0101] In some embodiments of this application, the step of removing edges between graph nodes of the first multimodal fusion graph based on a preset edge connection threshold to obtain a second multimodal fusion graph specifically involves:
[0102] Based on a preset edge connection threshold k, only the top-k connection edges of each graph node in the first multimodal fusion graph are retained to obtain the second multimodal fusion graph.
[0103] It should be noted that the connecting edges described in this application are used to describe the degree of association between graph nodes, and can reflect the correlation and importance between different modal features. Specifically, the connecting edge strength is obtained based on the semantic similarity of multimodal features.
[0104] In some embodiments of this application, the connection edge strength is the size of the matrix product of the eigenvector of each graph node and the eigenvectors of other graph nodes for which the connection edge strength needs to be calculated.
[0105] Specifically, after obtaining the edge strength of all connected edges of each graph node based on the first multimodal fusion graph, all connected edges of each graph node are sorted in descending order of edge strength, and the top-k connected edges of each graph node are selected to construct the second multimodal fusion graph.
[0106] In some embodiments of this application, the step of calculating the edge weights between graph nodes in the second multimodal fusion graph based on a self-attention mechanism, and enhancing the second multimodal fusion graph based on the edge weights using a multi-layer graph convolutional network to obtain fusion graph features, specifically involves:
[0107] Based on the self-attention mechanism, the edge weights E[b,N,k,D] between each graph node in the second multimodal fusion graph are calculated. edge ]; where D edge The edge feature dimension;
[0108] Based on the edge weights, the second multimodal fusion graph is enhanced using a multi-layer graph convolutional network to obtain the fusion graph feature F. graph [b,N,D graph ]; where D graph is the feature dimension of the convolutional graph.
[0109] In some embodiments of this application, the preferred embodiment of the multilayer graph convolutional network is a graph convolutional network (GCN).
[0110] This application first uses the time steps of the multimodal fusion features as graph nodes to generate a first multimodal fusion graph, which can capture multimodal features using graph structure. Then, based on a preset edge connection threshold, some edges are removed to obtain a second multimodal fusion graph, which can increase the weight of edges that meet the conditions. Then, edge weights are calculated based on a self-attention mechanism, and the second multimodal fusion graph is enhanced based on the edge weights and a multi-layer graph convolutional network. Through the graph convolutional network, the relationship between different modalities and different time steps can be captured, improving the matching degree between the output fusion graph features and the actual situation.
[0111] Step S104: Based on the fusion graph features and the preset multi-task loss, train multiple preset parallel models to obtain a joint control model; wherein, the multiple parallel models are based on multiple preset tasks; the multiple tasks include equipment anomaly detection, equipment classification and equipment status prediction.
[0112] In some embodiments of this application, the step of training multiple preset parallel models based on the fusion graph features and a preset multi-task loss to obtain a joint control model specifically includes:
[0113] The output types of the multiple parallel models are set according to the task categories of the multiple tasks;
[0114] Using the fused graph features as model input, and under the premise of the output type setting, the multiple parallel models are trained based on the preset multi-task loss to obtain a joint control model;
[0115] The preset multi-task loss is specifically as follows:
[0116] L=λ1*L anomaly +λ2*L device +λ3*L state ;
[0117] Where L is the multi-task loss, L anomaly L device L state These represent equipment anomaly loss, equipment type loss, and equipment state loss, respectively, with λ1, λ2, and λ3 being the weights of equipment anomaly loss, equipment type loss, and equipment state loss, respectively.
[0118] In some embodiments of this application, the plurality of tasks include equipment anomaly detection, equipment classification, and equipment status prediction, and all of the plurality of tasks are classification tasks; specifically, equipment anomaly detection is a binary classification task, and equipment classification and equipment status prediction are multi-classification tasks.
[0119] In some embodiments of this application, setting the output type of the multiple parallel models according to the task categories of the multiple tasks specifically involves:
[0120] Set the output type of the parallel model corresponding to the task category of device anomaly monitoring to score. anomaly [b,1];
[0121] Set the output type of the parallel model corresponding to the task category "device" to "score". device [b,num device ]; where num device This represents the total number of equipment types.
[0122] Set the output type of the parallel model corresponding to the task category of device status prediction to score. state [b,num state ]; where num state This represents the total number of device states.
[0123] In some embodiments of this application, the model structure of the plurality of parallel models is that the fused graph features are pooled and then connected to a fully connected layer.
[0124] This application first sets the output types of multiple parallel models according to the task categories of multiple tasks. By differentiating the models in terms of output type, a single model can specialize in a single task, which is beneficial to improving the accuracy of the model. Then, using the fusion graph features as the model input, multiple parallel models are trained based on the preset multi-task loss to obtain a joint control model, which can meet the different task requirements of substation operation control. In this way, when performing substation operation control based on the joint control model, the accuracy of substation operation control is improved.
[0125] Step S105: Obtain real-time monitoring data of the substation to be regulated, input the real-time monitoring data into the joint regulation model to obtain the operation decision, and regulate the substation to be regulated according to the operation decision.
[0126] This application first acquires the historical monitoring features of the substation to be regulated, calculates the attention weights of each feature based on a self-attention mechanism, and then fuses the features to obtain multimodal fusion features. This can comprehensively integrate monitoring features from different aspects and modes of the substation, improving the multimodal adaptability of substation operation and regulation. Then, it enhances the multimodal fusion features based on graph neural networks to obtain fusion graph features, which can further utilize graph structures to capture and enhance multimodal features. Finally, it trains multiple parallel models based on the fusion graph features and a preset multi-task loss to obtain a joint regulation model. By first dividing the substation operation and regulation into different situations and using different models for training, it can meet the different task requirements of substation operation and regulation, thereby improving the accuracy of substation operation and regulation when making operation decisions based on real-time monitoring data and regulating the substation.
[0127] Corresponding to the aforementioned method, please refer to Figure 2. This application provides a substation operation control device based on multimodality, including a feature acquisition module 210, a feature fusion module 220, a feature enhancement module 230, a model training module 240, and an operation control module 250.
[0128] The feature acquisition module 210 is used to acquire historical monitoring features of the substation to be controlled; wherein, the historical monitoring features include acoustic features, visible light features and infrared features;
[0129] The feature fusion module 220 is used to calculate the attention weights of the voiceprint features, the visible light features and the infrared features respectively based on the self-attention mechanism, and fuse the voiceprint features, the visible light features and the infrared features according to the attention weights to obtain multimodal fused features;
[0130] The feature enhancement module 230 is used to enhance the multimodal fusion features based on a graph neural network to obtain fusion graph features;
[0131] The model training module 240 is used to train multiple preset parallel models based on the fusion graph features and a preset multi-task loss to obtain a joint control model; wherein, the multiple parallel models are based on multiple preset tasks; the multiple tasks include equipment anomaly detection, equipment classification, and equipment status prediction.
[0132] The operation and control module 250 is used to acquire real-time monitoring data of the substation to be controlled, input the real-time monitoring data into the joint control model to obtain operation decisions, and control the substation to be controlled according to the operation decisions.
[0133] In some embodiments of this application, the feature acquisition module 210 includes a historical data acquisition unit, a voiceprint feature extraction unit, a visible light feature extraction unit, and an infrared feature extraction unit;
[0134] The historical data acquisition unit is used to acquire historical monitoring data of the substation to be regulated; wherein, the historical monitoring data includes acoustic data, visible light images and infrared images;
[0135] The voiceprint feature extraction unit is used to extract temporal features from the voiceprint data based on a bidirectional recurrent neural network to obtain voiceprint features.
[0136] The visible light feature extraction unit is used to extract image features from the visible light image based on a self-attention image encoder to obtain visible light features;
[0137] The infrared feature extraction unit is used to extract image features from the infrared image based on a convolutional neural network to obtain infrared features.
[0138] In some embodiments of this application, the feature fusion module 220 includes a dimension unification unit and a feature fusion unit;
[0139] The dimension unification unit is used to unify the dimension of the voiceprint feature and the dimension of the infrared feature into the dimension of the visible light feature based on a convolutional neural network, thereby obtaining the first voiceprint feature and the first infrared feature.
[0140] The feature fusion unit is used to splice and fuse the first voiceprint feature, the visible light feature, and the first infrared feature based on the attention weight and a multi-layer self-attention encoder to obtain a multimodal fused feature.
[0141] In some embodiments of this application, the feature enhancement module 230 includes a fusion graph generation unit, a fusion edge culling unit, and a fusion graph enhancement unit;
[0142] The fusion graph generation unit is used to generate a first multimodal fusion graph corresponding to the multimodal fusion feature, using the time step of the multimodal fusion feature as a graph node.
[0143] The fusion edge removal unit is used to remove edges between graph nodes of the first multimodal fusion graph based on a preset edge connection threshold to obtain a second multimodal fusion graph.
[0144] The fusion graph enhancement unit is used to calculate the edge weights between each graph node in the second multimodal fusion graph based on a self-attention mechanism, and enhance the second multimodal fusion graph based on the edge weights and a multilayer graph convolutional network to obtain fusion graph features.
[0145] In some embodiments of this application, the model training module 240 includes a type setting unit and a model training unit;
[0146] The type setting unit is used to set the output type of the multiple parallel models according to the task categories of the multiple tasks;
[0147] The model training unit is used to train the multiple parallel models based on the preset multi-task loss, with the fusion graph features as model input and the output type set, to obtain a joint control model.
[0148] The preset multi-task loss is specifically as follows:
[0149] L=λ1*L anomaly +λ2*L device +λ3*L state ;
[0150] Where L is the multi-task loss, L anomely ,L device, L state These represent equipment anomaly loss, equipment type loss, and equipment state loss, respectively, with λ1, λ2, and λ3 being the weights of equipment anomaly loss, equipment type loss, and equipment state loss, respectively.
[0151] This application first acquires the historical monitoring features of the substation to be regulated, calculates the attention weights of each feature based on a self-attention mechanism, and then fuses the features to obtain multimodal fusion features. This can comprehensively integrate monitoring features from different aspects and modes of the substation, improving the multimodal adaptability of substation operation and regulation. Then, it enhances the multimodal fusion features based on graph neural networks to obtain fusion graph features, which can further utilize graph structures to capture and enhance multimodal features. Finally, it trains multiple parallel models based on the fusion graph features and a preset multi-task loss to obtain a joint regulation model. By first dividing the substation operation and regulation into different situations and using different models for training, it can meet the different task requirements of substation operation and regulation, thereby improving the accuracy of substation operation and regulation when making operation decisions based on real-time monitoring data and regulating the substation.
[0152] It should be understood that the device provided in the embodiments of this application corresponds to the aforementioned method. The substation operation control device based on multimode provided in the embodiments of this application can realize the substation operation control method based on multimode provided in any embodiment of this application.
[0153] Adaptively, embodiments of this application also provide a computer device and a computer-readable storage medium.
[0154] The computer device includes: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor;
[0155] The processor executes the computer program to implement a multimodal substation operation control method according to this application.
[0156] The computer-readable storage medium stores multiple instructions, which are adapted for a processor to load and execute a multimodal substation operation control method according to this application.
[0157] The above description represents some embodiments of this application, providing a further detailed explanation of the purpose, technical solution, and beneficial effects of this application. It should be understood that the above-described embodiments of this application should not be construed as limiting this application. In particular, any changes, modifications, equivalent substitutions, and variations made by those skilled in the art within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A substation operation and control method based on multimodal operation, characterized in that, include: The historical monitoring characteristics of the substation to be regulated are obtained; wherein, the historical monitoring characteristics include acoustic signature characteristics, visible light characteristics, and infrared characteristics; Based on the self-attention mechanism, the attention weights of the voiceprint features, visible light features and infrared features are calculated respectively, and the voiceprint features, visible light features and infrared features are fused according to the attention weights to obtain multimodal fused features; Based on a graph neural network, the multimodal fusion features are enhanced to obtain fusion graph features; Based on the fusion graph features and the preset multi-task loss, multiple preset parallel models are trained to obtain a joint control model; wherein, the multiple parallel models are based on multiple preset tasks; the multiple tasks include equipment anomaly detection, equipment classification, and equipment status prediction. The real-time monitoring data of the substation to be regulated is obtained, the real-time monitoring data is input into the joint regulation model to obtain the operation decision, and the substation to be regulated is regulated according to the operation decision.
2. The substation operation and control method based on multimodal operation according to claim 1, characterized in that, The acquisition of historical monitoring characteristics of the substation to be regulated specifically includes: Acquire historical monitoring data of the substation to be regulated; wherein, the historical monitoring data includes acoustic signature data, visible light images, and infrared images; Based on a bidirectional recurrent neural network, temporal features are extracted from the voiceprint data to obtain voiceprint features; Based on a self-attention image encoder, image features are extracted from the visible light image to obtain visible light features; Based on a convolutional neural network, image features are extracted from the infrared image to obtain infrared features.
3. The substation operation and control method based on multimodal operation according to claim 1, characterized in that, The process of fusing the voiceprint features, visible light features, and infrared features according to the attention weights to obtain multimodal fused features specifically includes: Based on a convolutional neural network, the dimensions of the voiceprint feature and the infrared feature are unified to the dimension of the visible light feature to obtain the first voiceprint feature and the first infrared feature. Based on the attention weights, and using a multi-layer self-attention encoder, the first voiceprint feature, the visible light feature, and the first infrared feature are spliced and fused to obtain a multimodal fused feature.
4. The substation operation and control method based on multimodal operation according to claim 1, characterized in that, The enhancement of the multimodal fusion features based on graph neural networks to obtain fused graph features specifically includes: Using the time step of the multimodal fusion feature as a graph node, a first multimodal fusion graph corresponding to the multimodal fusion feature is generated; Based on a preset edge connection threshold, edges between graph nodes in the first multimodal fusion graph are removed to obtain a second multimodal fusion graph. Based on the self-attention mechanism, the edge weights between each graph node in the second multimodal fusion graph are calculated, and the second multimodal fusion graph is enhanced based on the edge weights and a multilayer graph convolutional network to obtain fusion graph features.
5. The substation operation and control method based on multimodal operation according to claim 1, characterized in that, The step of training multiple preset parallel models based on the fusion graph features and a preset multi-task loss to obtain a joint control model specifically includes: The output types of the multiple parallel models are set according to the task categories of the multiple tasks; Using the fused graph features as model input, and under the premise of the output type setting, the multiple parallel models are trained based on the preset multi-task loss to obtain a joint control model; The preset multi-task loss is specifically as follows: L=λ*L anomaly +λ2*L device +λ3*L state ; Where L is the multi-task loss, L anomaly L device L state These represent equipment anomaly loss, equipment type loss, and equipment state loss, respectively, with λ1, λ2, and λ3 being the weights of equipment anomaly loss, equipment type loss, and equipment state loss, respectively.
6. A substation operation control device based on multimodal operation, characterized in that, It includes a feature acquisition module, a feature fusion module, a feature enhancement module, a model training module, and a runtime control module; The feature acquisition module is used to acquire historical monitoring features of the substation to be regulated; wherein, the historical monitoring features include acoustic signature features, visible light features, and infrared features; The feature fusion module is used to calculate the attention weights of the voiceprint features, visible light features and infrared features respectively based on the self-attention mechanism, and fuse the voiceprint features, visible light features and infrared features according to the attention weights to obtain multimodal fused features; The feature enhancement module is used to enhance the multimodal fusion features based on a graph neural network to obtain fusion graph features; The model training module is used to train multiple preset parallel models based on the fusion graph features and a preset multi-task loss to obtain a joint control model; wherein, the multiple parallel models are based on multiple preset tasks; the multiple tasks include equipment anomaly detection, equipment classification, and equipment status prediction. The operation and control module is used to acquire real-time monitoring data of the substation to be controlled, input the real-time monitoring data into the joint control model to obtain operation decisions, and control the substation to be controlled according to the operation decisions.
7. A substation operation control device based on multimodal operation according to claim 6, characterized in that, The feature acquisition module includes a historical data acquisition unit, a voiceprint feature extraction unit, a visible light feature extraction unit, and an infrared feature extraction unit; The historical data acquisition unit is used to acquire historical monitoring data of the substation to be regulated; wherein, the historical monitoring data includes acoustic data, visible light images and infrared images; The voiceprint feature extraction unit is used to extract temporal features from the voiceprint data based on a bidirectional recurrent neural network to obtain voiceprint features. The visible light feature extraction unit is used to extract image features from the visible light image based on a self-attention image encoder to obtain visible light features; The infrared feature extraction unit is used to extract image features from the infrared image based on a convolutional neural network to obtain infrared features.
8. A substation operation control device based on multimodal operation according to claim 6, characterized in that, The feature fusion module includes a dimension unification unit and a feature fusion unit; The dimension unification unit is used to unify the dimension of the voiceprint feature and the dimension of the infrared feature into the dimension of the visible light feature based on a convolutional neural network, thereby obtaining the first voiceprint feature and the first infrared feature. The feature fusion unit is used to splice and fuse the first voiceprint feature, the visible light feature, and the first infrared feature based on the attention weight and a multi-layer self-attention encoder to obtain a multimodal fused feature.
9. A substation operation control device based on multimodal operation according to claim 6, characterized in that, The feature enhancement module includes a fusion graph generation unit, a fusion edge removal unit, and a fusion graph enhancement unit; The fusion graph generation unit is used to generate a first multimodal fusion graph corresponding to the multimodal fusion feature, using the time step of the multimodal fusion feature as a graph node. The fusion edge removal unit is used to remove edges between graph nodes of the first multimodal fusion graph based on a preset edge connection threshold to obtain a second multimodal fusion graph. The fusion graph enhancement unit is used to calculate the edge weights between each graph node in the second multimodal fusion graph based on the self-attention mechanism, and enhance the second multimodal fusion graph based on the edge weights and a multilayer graph convolutional network to obtain fusion graph features.
10. A substation operation control device based on multimodal operation according to claim 6, characterized in that, The model training module includes a type setting unit and a model training unit; The type setting unit is used to set the output type of the multiple parallel models according to the task categories of the multiple tasks; The model training unit is used to train the multiple parallel models based on the preset multi-task loss, with the fusion graph features as model input and the output type set, to obtain a joint control model. The preset multi-task loss is specifically as follows: L=λ1*L anomaly +λ2*L device +λ3*L state Where L is the multi-task loss, L anomaly ,L device ,L state These represent equipment anomaly loss, equipment type loss, and equipment state loss, respectively, with λ1, λ2, and λ3 being the weights of equipment anomaly loss, equipment type loss, and equipment state loss, respectively.
Citation Information
Patent Citations
Zero-sample picture identification method based on multi-granularity fusion network
CN112488241A
Fusion method and device of infrared image and visible light image, and storage medium
CN116824319A
Transformer fault diagnosis method and system based on voiceprint and infrared feature fusion
CN117292716A
Multi-angle image data processing method and device and electronic equipment
CN118334371A
Substation operation regulation and control method and device based on multiple modes
CN119362706A
Cited By
Toy bag plastic layer online quality detection method based on multi-modal visual perception
CN122156213A