Intelligent transfer learning method and system for power grid multi-mode large model
By building multimodal large models of the power grid and introducing field adaptive networks and physical perception constraints, the problem of multimodal data fusion in the power grid system is solved, efficient multimodal transfer learning and strong generalization capabilities are achieved, and the ability of intelligent management of the power grid is improved.
Patent Information
- Application Number
- CN202510021478.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
AI Technical Summary
In power grid systems, the distribution deviation of multi-source heterogeneous data and complex physical constraints make it difficult for traditional single-modal or specific structure machine learning models to effectively fusion of multi-modal data, and the model's robustness or the results lack physical rationality, affecting migration and cross-domain capabilities.
Using the intelligent transfer learning method of the multimodal large grid model, the standardized processing and feature fusion of multimodal data are carried out by building a pre-trained grid model, including the basic modal embedding layer, the cross-modal interaction module and the fusion feature generation module. At the same time, special-purpose field adaptive networks and physical perception constraints are introduced to eliminate feature deviations between source data and target data distribution layer by layer, and ensure the physical rationality of features.
It realizes efficient multimodal transfer learning in complex power grid scenarios, improves the robustness and physical rationality of the model, enhances the ability to identify and predict power grid tasks, has strong generalization capabilities, and is suitable for the needs of intelligent power grid management.
Smart Images

Figure CN119940472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large power models, and in particular to an intelligent transfer learning method and system for a large multi-modal power grid model. Background Art
[0002] In modern power systems, with the continuous expansion of the scale of power grid equipment and the increasing complexity of the operating environment, multi-source heterogeneous data (such as text, images, and time series data) play an important role in power grid operation monitoring, fault diagnosis, and status assessment. However, due to the differences in the sources of different modal data and the diversity of power grid scenarios, these data often have obvious distribution deviations and complex physical constraints, which poses a huge challenge to intelligent analysis. Traditional single-modal or specific structure machine learning models are often difficult to effectively integrate multimodal data. At the same time, due to the lack of effective constraints on domain knowledge constraints (such as the law of power grid fault propagation or the law of power conservation), it is easy to lead to insufficient robustness of the model or lack of physical rationality of the results. In addition, when there is a deviation in the distribution of source data and target data (such as differences in power grid scenarios across regions and devices), the migration and cross-domain capabilities of the model are also limited. Summary of the invention
[0003] In order to solve the above problems, the purpose of the present invention is to provide an intelligent transfer learning method and system for a large multimodal model of a power grid, which can realize efficient multimodal transfer learning in complex power grid scenarios and meet the high requirements of intelligence and precision in the power grid field.
[0004] To achieve the above object, the present invention adopts the following technical solutions:
[0005] An intelligent transfer learning method for a large multi-modal model of a power grid includes the following steps:
[0006] S1: Acquire multimodal data of power grid, including text, image, and time series data;
[0007] S2: Build a pre-trained power grid model, including a basic modal embedding layer, a cross-modal interaction module, and a fusion feature generation module; standardize the multi-modal data through the basic modal embedding layer and extract the embedding vector of the modal data; the cross-modal interaction module uses the joint embedding space technology to perform cross-modal feature alignment on the features of each modal data, obtain the final aligned multi-modal embedding vector, and fuse it through the fusion feature generation module;
[0008] S3: Based on the pre-trained large power grid model, add task-specific output layers according to downstream tasks, partially freeze the pre-trained model weights, and only fine-tune the upper network or task head to transfer knowledge to specific power grid tasks;
[0009] S4: In the process of knowledge transfer, a dedicated domain adaptive network is introduced. Based on the layer-by-layer adversarial method, an adversarial domain classifier is added to each layer of features represented by the power grid data to eliminate the feature deviation between the source data and the target data distribution layer by layer. Physical perception constraints are used to limit the physical rationality of the features in the domain adaptation process.
[0010] Furthermore, the multimodal data is standardized, including text data processing, image data processing, and time series data processing;
[0011] The text data processing is specifically as follows:
[0012] Clear irrelevant fields, including extra spaces and templated strings, and use regular matching to remove extra symbols and repeated information;
[0013] Use the Chinese word segmentation tool Jieba to segment text, and use word tagging or term tagging methods for English fields;
[0014] Create a stop word vocabulary to remove useless entries; use Unicode to unify the character encoding of the text.
[0015] Use the pre-trained text embedding model to convert the cleaned text into a text embedding vector;
[0016] Image data processing is as follows:
[0017] Convert the collected images to JPEG format and adjust the image size to a fixed target;
[0018] Use bilinear interpolation to complete scaling to avoid pixel distortion, and use median filtering to denoise the image;
[0019] Normalize each image by channel mean and standard deviation;
[0020] Use a pre-trained visual feature extractor to convert images into image embedding vectors.
[0021] The time series data processing is specifically as follows:
[0022] Resample non-uniformly sampled time series data into fixed time intervals;
[0023] Use the sliding window method to detect and remove abnormal data points;
[0024] Normalize each time series segment;
[0025] Split the time series into time segments of fixed length;
[0026] The preprocessed fixed window segments are input into the time series feature extraction model to generate a time series embedding vector.
[0027] Furthermore, the cross-modal interaction module is as follows:
[0028] Embedding generation based on unimodal feature extraction, including text embedding vector T, image embedding vector I and time series embedding vector S, uses a unified projection mechanism to map each modality feature to a fixed dimension d E The shared feature space of:
[0029] E t =W t ·T+b t ,E i =W i I+b i ,E s =W s ·S+b s ;
[0030] Among them, W t ,W i ,W s is the linear transformation weight matrix, b t ,b i ,b s is the bias term;
[0031] For different modal feature combinations, the interaction between modal features is directly calculated through the attention mechanism.
[0032] Text and image interaction:
[0033] E t For query, Ei is key / value:
[0034] E t→i =Attention(Q t ,K i ,V i );
[0035] Among them, E t→i Q is the image embedding vector associated with the text description information; t K is the query matrix obtained by projecting the text embedding vector; i ,V i Get the key and value from the image embedding vector respectively;
[0036] Text and timing interaction:
[0037] E t For query, E s For key / value:
[0038] E t→s =Attention(Q t ,Ks ,V s );
[0039] Among them, E t→s K is the temporal embedding vector associated with the text description information; s ,V s The keys and values are projected from the time series embedding vector respectively;
[0040] Image and time series interaction:
[0041] E i For query, E s For key / value:
[0042] E i→s =Attention(Q i ,K s ,V s) ;
[0043] Among them, E i→s is the dynamic correlation between the image device and the timing fluctuation; Q i To obtain the query matrix from the image embedding vector projection;
[0044] The cross-modal pairwise interaction results E t→i ,E t→s ,E i→s Fusion into each modality, enhance its feature expression ability, and obtain a new text embedding vector E t , image embedding vector E i and the temporal embedding vector E s ;
[0045]
[0046] Among them, α1, α2, β1, β2, γ1, γ2 are weighted fusion coefficients;
[0047] Through a unified multimodal attention module, the comprehensive interaction relationship between all modalities is captured to form a global joint representation E out_all .
[0048] Furthermore, the pre-trained large power grid model is obtained by pre-training based on Multimodal Transformer, specifically:
[0049] The Transformer framework is input in the form of a joint sequence of multimodal representations, specifically:
[0050]
[0051] Add a learnable identification vector for each modality. Specifically, add [MODALITY_TEXT] for text modality, [MODALITY_IMAGE] for image modality, and [MODALITY_TIME] for time series modality. The input data format is:
[0052]
[0053] And add position information to the input embedding to distinguish the order relationship in the sequence;
[0054] The Multimodal Transformer extends the architecture of the standard Transformer model to learn unified fusion representations through multi-layer attention mechanisms and cross-modal feature modeling;
[0055] Pre-training captures contextual information of multimodal data through self-supervised learning mechanisms, improving the model's cross-modal understanding capabilities;
[0056] The last layer sequence representation generated by the multimodal Transformer is pooled to construct the final multimodal global embedding.
[0057] Furthermore, S3 is specifically as follows: load the pre-trained large power grid model, including the basic modal embedding layer, the cross-modal interaction module, and the fusion feature generation module; load only part of the model parameters according to the requirements of the specific task; design a task-specific output head for each specific task, use the classification head in the classification task, and use the regression head in the regression task; freeze the underlying modules of the pre-trained large power grid model, and only fine-tune the intermediate cross-modal interaction module and task head; select and unfreeze some intermediate layers to enable the model to further learn according to the specific task data.
[0058] Furthermore, a dedicated domain adaptive network is introduced. Based on the layer-by-layer adversarial method, an adversarial domain classifier is added to each layer of features represented by the power grid data to eliminate the feature deviation between the source data and the target data distribution layer by layer. The details are as follows:
[0059] An adversarial domain classifier is added after the feature representation of each layer of the network, and the difference in feature distribution between the source domain and the target domain is reduced through adversarial optimization;
[0060] Layer-by-layer adversarial network structure:
[0061] For the power grid task, suppose the feature extraction network has L layers, and for each layer of feature H l , add a domain classifier D l :
[0062] H l =f l (H l-1 ), D l(H l )=softmax(W l ·H l +b l );
[0063] Among them, D l The role of is to judge the feature H l From the source domain or the target domain;
[0064] Using adversarial training to make the domain classifier unable to distinguish the source and target domain features:
[0065] Train Dl to accurately predict the domain label y of the input data domain ;
[0066] Reverse optimization feature extraction network f l Make D l The classification error rate is as high as possible, thereby ignoring the domain information; the loss function is adversarial loss:
[0067]
[0068] The feature generator f l The optimization goal is to maximize the prediction error of the domain classifier by achieving adversarial effects through reverse gradient;
[0069] Through adversarial training, the domain characteristics of the input data are eliminated layer by layer, and finally the domain-free representation of the features is achieved.
[0070] Furthermore, physical perception constraints are used to limit the physical rationality of features in the domain adaptation process, as follows:
[0071] When aligning data features layer by layer, direct adversarial training may cause the physical meaning of the features to be lost, introduce physical constraints of the power grid, and limit the rationality of feature alignment;
[0072] Grid faults will dynamically propagate along the topology of device interconnection. Topology consistency constraints are expressed through the grid topology matrix A and the propagation neighbor characteristics F.
[0073] In topological propagation, if the node i represented by the network feature H fails, its neighbor features also show certain similarities; the topological neighbor consistency loss is introduced:
[0074]
[0075] Where (i, j) is the adjacent device in the fault propagation path;
[0076] And through the conservation of physical variables, constraints are imposed. Specifically, for node i in the topology diagram, according to the power distribution:
[0077]
[0078] Introducing the physical non-conservation constraint loss L physics :
[0079]
[0080] Get the final optimization goal:
[0081]
[0082] Among them, λ1,λ2 are weight hyperparameters.
[0083] A main-distribution-microgrid integrated real-time coordinated risk scheduling system comprises a processor, a memory and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the main-distribution-microgrid integrated real-time coordinated risk scheduling method as described above.
[0084] The present invention has the following beneficial effects:
[0085] 1. The present invention can effectively solve the distribution deviation problem and realize efficient migration of target domain tasks by constructing a unified large multi-modal model of the power grid and combining layer-by-layer domain adaptation and physical perception constraints;
[0086] 2. The pre-training method based on the Multimodal Transformer framework of the present invention makes full use of the semantic interaction and alignment characteristics between modalities such as text, image, and time series, and can effectively improve the performance of multimodal tasks in power grid scenarios;
[0087] 3. The present invention applies the pre-trained large multimodal model of the power grid to downstream tasks, greatly improving the recognition and prediction capabilities of cross-modal tasks. By combining the fine-tuning strategy with the task-specific head design, the model can be quickly adapted and efficiently deployed. At the same time, it has strong generalization capabilities when the labeled data is limited. The transfer learning solution can be widely used in the intelligent management of power grid scenarios, such as fault detection, operation prediction, and analysis needs of text and image combination, thereby helping the power grid to operate more intelligently and efficiently. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0089] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0090] refer to Figure 1 In the present invention, a method for intelligent transfer learning of a large multi-modal model of a power grid is provided, comprising the following steps:
[0091] S1: Obtain multi-modal power grid data including text (such as fault description information), images (such as power grid equipment images, fault detection results), and time-series data (such as load time-series).
[0092] S2; Construct a pre-trained large power grid model, including a basic modal embedding layer, a cross-modal interaction module, and a fusion feature generation module; Standardize the multi-modal data through the basic modal embedding layer and extract the embedding vectors of the modal data; The cross-modal interaction module uses the joint embedding space technology to perform cross-modal feature alignment on the features of each modal data, obtain the finally generated aligned multi-modal embedding vectors, and perform fusion through the fusion feature generation module.
[0093] S3: Based on the pre-trained large power grid model, add task-specific output layers (such as classification heads, regression heads, or similarity calculation modules) according to downstream tasks (such as fault classification, operating parameter prediction, text-image matching), partially freeze the weights of the pre-trained model, and only fine-tune the upper network or task head to transfer knowledge to specific power grid tasks.
[0094] S4: During the knowledge transfer process, introduce a dedicated domain adaptation network. Based on the method of Layer-wise DANN, add adversarial domain classifiers to the features of each layer of the power grid data representation to gradually eliminate the feature deviation between the source data and the target data distributions; Adopt physical perception constraints to limit the physical rationality of features during the domain adaptation process.
[0095] In this embodiment, standardizing the multi-modal data includes text data processing, image data processing, and time-series data processing.
[0096] The text data processing is as follows:
[0097] Remove irrelevant fields, including extra spaces and templated strings, and use regular matching to remove extra symbols and duplicate information.
[0098] Use the Chinese word segmentation tool Jieba for text segmentation, and use character marking or word marking methods for English fields.
[0099] Establish a stop word vocabulary list (including functional auxiliary words such as "de", "le", "shi", etc.) and remove useless entries; Use Unicode to perform unified character encoding processing on the text.
[0100] Use a pre-trained text embedding model (such as BERT, RoBERTa, or Chinese-based ERNIE) to convert the cleaned text into text embedding vectors.
[0101] The image data processing is as follows:
[0102] Convert the collected images to JPEG format and adjust the image size to a fixed target (such as 224×224 or 384×384);
[0103] Use bilinear interpolation to complete scaling to avoid pixel distortion, and use median filtering to denoise the image;
[0104] Normalize each image by channel mean and standard deviation;
[0105] Use a pre-trained visual feature extractor (e.g., ResNet, ViT) to convert images into image embedding vectors.
[0106] The time series data processing is specifically as follows:
[0107] Resample non-uniformly sampled time series data into fixed time intervals;
[0108] Use the sliding window method to detect abnormal data points (the abnormal points can be judged by the statistical rules of voltage / current fluctuations) and remove the abnormal data points;
[0109] Normalize each time series segment;
[0110] Split the time series into time segments of fixed length (e.g., 30 seconds of data is divided into six 5-second windows for modeling);
[0111] The preprocessed fixed window fragments are input into the time series feature extraction model (such as TCN, Transformer) to generate a time series embedding vector.
[0112] In this embodiment, the cross-modal interaction module is as follows:
[0113] Embedding generation based on unimodal feature extraction, including text embedding vector T, image embedding vector I and time series embedding vector S, uses a unified projection mechanism to map each modality feature to a fixed dimension d E The shared feature space of:
[0114] E t =W t ·T+b t ,E i =W i I+b i ,E s =W s ·S+b s ;
[0115] Among them, W t ,W i ,W s is the linear transformation weight matrix, b t ,b i ,bs is the bias term;
[0116] For different modal feature combinations, the interaction between modal features is directly calculated through the attention mechanism.
[0117] Text and image interaction:
[0118] E t For query, Ei is key / value:
[0119] E t→i =Attention(Q t ,K i ,V i );
[0120] Among them, E t→i Q is the image embedding vector associated with the text description information; t K is the query matrix obtained by projecting the text embedding vector; i ,V i Get the key and value from the image embedding vector respectively;
[0121] Text and timing interaction:
[0122] E t For query, E s For key / value:
[0123] E t→s =Attention(Q t ,K s ,V s );
[0124] Among them, E t→s K is the temporal embedding vector associated with the text description information; s ,V s The keys and values are projected from the time series embedding vector respectively;
[0125] Image and time series interaction:
[0126] E i For query, E s For key / value:
[0127] E i→s =Attention(Q i ,K s ,V s) ;
[0128] Among them, E i→s is the dynamic correlation between the image device and the timing fluctuation; Q i To obtain the query matrix from the image embedding vector projection;
[0129] The cross-modal pairwise interaction results E t→i ,E t→s ,E i→s Fusion into each modality, enhance its feature expression ability, and obtain a new text embedding vector E t , image embedding vector E i and the temporal embedding vector E s ;
[0130]
[0131] Among them, α1, α2, β1, β2, γ1, γ2 are weighted fusion coefficients;
[0132] Through a unified multimodal attention module, the comprehensive interaction relationship between all modalities is captured to form a global joint representation E out_all .
[0133] In this embodiment, the pre-trained large power grid model is obtained by pre-training based on Multimodal Transformer, specifically:
[0134] The Transformer framework is input in the form of a joint sequence of multimodal representations, specifically:
[0135]
[0136] Add a learnable identification vector for each modality. Specifically, add [MODALITY_TEXT] for text modality, [MODALITY_IMAGE] for image modality, and [MODALITY_TIME] for time series modality. The input data format is:
[0137]
[0138] And add position information to the input embedding to distinguish the order relationship in the sequence;
[0139] The Multimodal Transformer extends the architecture of the standard Transformer model to learn unified fusion representations through multi-layer attention mechanisms and cross-modal feature modeling;
[0140] Pre-training captures contextual information of multimodal data through self-supervised learning mechanisms, improving the model's cross-modal understanding capabilities;
[0141] The last layer sequence representation generated by the multimodal Transformer is pooled to construct the final multimodal global embedding.
[0142] In this embodiment, S3 is specifically:
[0143] Load the pre-trained large power grid model, including the basic modality embedding layer, the cross-modal interaction module, and the fusion feature generation module;
[0144] According to the requirements of specific tasks, only some model parameters are loaded;
[0145] For example, for a task that only requires text and timing, the image-related part may not be loaded;
[0146] Design task-specific output heads for each specific task, using classification heads for classification tasks and regression heads for regression tasks;
[0147] Freeze the underlying modules of the pre-trained Grid Big model (such as unimodal attention layers and feature extraction layers), and only fine-tune the intermediate cross-modal interaction modules and task heads;
[0148] Select to unfreeze some intermediate layers (such as the modal interaction layer) to allow the model to further learn based on specific task data.
[0149] In this embodiment, a dedicated domain adaptive network is introduced. Based on the layer-wise DANN method, an adversarial domain classifier is added to each layer of features represented by the power grid data to eliminate the feature deviation between the source data and the target data distribution layer by layer. The details are as follows:
[0150] An adversarial domain classifier is added after the feature representation of each layer of the network, and the difference in feature distribution between the source domain and the target domain is reduced through adversarial optimization;
[0151] Layer-by-layer adversarial network structure:
[0152] For the power grid task, suppose the feature extraction network has L layers, and for each layer of feature H l , add a domain classifier D l :
[0153] H l =f l (H l-1 ), D l (H l )=softmax(W l ·H l +b l );
[0154] Among them, D l The role of is to judge the feature H l From the source domain or the target domain;
[0155] Using adversarial training to make the domain classifier unable to distinguish the source and target domain features:
[0156] Train Dl to accurately predict the domain label y of the input data domain ;
[0157] Reverse optimization feature extraction network f l Make D l The classification error rate is as high as possible, thereby ignoring the domain information; the loss function is adversarial loss:
[0158]
[0159] The feature generator f l The optimization goal is to maximize the prediction error of the domain classifier by achieving adversarial effects through reverse gradient;
[0160] Through adversarial training, the domain characteristics of the input data are eliminated layer by layer, and finally the domain-free representation of the features is achieved.
[0161] In this embodiment, physical perception constraints are used to limit the physical rationality of features in the domain adaptation process, as follows:
[0162] When aligning data features layer by layer, direct adversarial training may cause the physical meaning of the features to be lost, introduce physical constraints of the power grid, and limit the rationality of feature alignment;
[0163] Grid faults will dynamically propagate along the topology of device interconnection. Topology consistency constraints are expressed through the grid topology matrix A and the propagation neighbor characteristics F.
[0164] In topological propagation, if the node i represented by the network feature H fails, its neighbor features also show certain similarities; the topological neighbor consistency loss is introduced:
[0165]
[0166] Where (i, j) is the adjacent device in the fault propagation path;
[0167] And through the conservation of physical variables, constraints are imposed. Specifically, for node i in the topology diagram, according to the power distribution:
[0168]
[0169] Introducing the physical non-conservation constraint loss L physics :
[0170]
[0171] Get the final optimization goal:
[0172]
[0173] Among them, λ1,λ2 are weight hyperparameters.
[0174] A main distribution microgrid integrated real-time coordinated risk dispatching system comprises a processor, a memory and a computer program stored in the memory. When the processor executes the computer program, the steps in the main distribution microgrid integrated real-time coordinated risk dispatching method as described above are specifically performed.
[0175] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0176] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0177] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0179] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.
Claims
1. An intelligent transfer learning method for a large multi-modal model of a power grid, characterized in that: The following steps are involved: S1: Acquire multimodal data of power grid, including text, image, and time series data; S2: Build a pre-trained large power grid model, including a basic modality embedding layer, a cross-modal interaction module, and a fusion feature generation module; The multimodal data is standardized through the basic modality embedding layer, and the embedding vector of the modal data is extracted; the cross-modal interaction module uses the joint embedding space technology to align the features of each modality data across modalities, and finally generates the aligned multimodal embedding vector, which is then fused through the fusion feature generation module; S3: Based on the pre-trained large power grid model, add task-specific output layers according to downstream tasks, partially freeze the pre-trained model weights, and only fine-tune the upper network or task head to transfer knowledge to specific power grid tasks; S4: In the process of knowledge transfer, a dedicated domain adaptive network is introduced. Based on the layer-by-layer adversarial method, an adversarial domain classifier is added to each layer of features represented by the power grid data to eliminate the feature deviation between the source data and the target data distribution layer by layer. Physical perception constraints are used to limit the physical rationality of the features in the domain adaptation process.
2. According to claim 1, a method for intelligent transfer learning of a large multi-modal model of a power grid, characterized in that: The standardized processing of the multimodal data includes text data processing, image data processing and time series data processing; The text data processing is specifically as follows: Clear irrelevant fields, including extra spaces and templated strings, and use regular matching to remove extra symbols and repeated information; Use the Chinese word segmentation tool Jieba to segment text, and use word tagging or term tagging methods for English fields; Create a stop word vocabulary to remove useless entries; use Unicode to unify the character encoding of the text. Use the pre-trained text embedding model to convert the cleaned text into a text embedding vector; Image data processing is as follows: Convert the collected images to JPEG format and adjust the image size to a fixed target; Use bilinear interpolation to complete scaling to avoid pixel distortion, and use median filtering to denoise the image; Normalize each image by channel mean and standard deviation; Use a pre-trained visual feature extractor to convert images into image embedding vectors. The time series data processing is specifically as follows: Resample non-uniformly sampled time series data into fixed time intervals; Use the sliding window method to detect and remove abnormal data points; Normalize each time series segment; Split the time series into time segments of fixed length; The preprocessed fixed window segments are input into the time series feature extraction model to generate a time series embedding vector.
3. The intelligent transfer learning method of a large multi-modal model of a power grid according to claim 2 is characterized in that: The cross-modal interaction module is as follows: Embedding generation based on unimodal feature extraction, including text embedding vector T, image embedding vector I and time series embedding vector S, uses a unified projection mechanism to map each modality feature to a fixed dimension d E The shared feature space of: E t =W t ·T+b t ,E i =W i ·I+b i ,E s =W s ·S+b s ; Among them, W t ,W i ,W s is the linear transformation weight matrix, b t ,b i ,b s is the bias term; For different modal feature combinations, the interaction between modal features and text and image interaction is directly calculated through the attention mechanism: E t For query, Ei is key / value: E t→i =Attention(Q t ,K i ,V i ); Among them, E t→i Q is the image embedding vector associated with the text description information; t K is the query matrix obtained by projecting the text embedding vector; i ,V i Get the key and value from the image embedding vector respectively; Text and timing interaction: E t For query, E s For key / value: E t→s =Attention(Q t ,K s ,V s ); Among them, E t→s K is the temporal embedding vector associated with the text description information; s ,V s The keys and values are projected from the time series embedding vector respectively; Image and time series interaction: E i For query, E s For key / value: E i→s =Attention(Q i ,K s ,V s) ; Among them, E i→s is the dynamic correlation between the image device and the timing fluctuation; Q i To obtain the query matrix from the image embedding vector projection; The cross-modal pairwise interaction results E t→i ,E t→s ,E i→s Fusion into each modality, enhance its feature expression ability, and obtain a new text embedding vector E t , image embedding vector E i and the temporal embedding vector E s ; Among them, α1, α2, β1, β2, γ1, γ2 are weighted fusion coefficients; Through a unified multimodal attention module, the comprehensive interaction relationship between all modalities is captured to form a global joint representation E out_all .
4. The intelligent transfer learning method of a large multi-modal model of a power grid according to claim 1 is characterized in that: The pre-trained large power grid model is obtained by pre-training based on Multimodal Transformer, specifically: The Transformer framework is input in the form of a joint sequence of multimodal representations, specifically: Add a learnable identification vector for each modality. Specifically, add [MODALITY_TEXT] for text modality, [MODALITY_IMAGE] for image modality, and [MODALITY_TIME] for time series modality. The input data format is: And add position information to the input embedding to distinguish the order relationship in the sequence; The Multimodal Transformer extends the architecture of the standard Transformer model to learn unified fusion representations through multi-layer attention mechanisms and cross-modal feature modeling; Pre-training captures contextual information of multimodal data through self-supervised learning mechanisms, improving the model's cross-modal understanding capabilities; The last layer sequence representation generated by the multimodal Transformer is pooled to construct the final multimodal global embedding.
5. The intelligent transfer learning method of a large multi-modal model of a power grid according to claim 1 is characterized in that: The S3 is specifically as follows: loading a pre-trained large power grid model, including a basic modal embedding layer, a cross-modal interaction module, and a fusion feature generation module; loading only part of the model parameters according to the requirements of specific tasks; designing a task-specific output head for each specific task, using a classification head in a classification task and a regression head in a regression task; freezing the underlying modules of the pre-trained large power grid model, and only fine-tuning the intermediate cross-modal interaction module and the task head; and selectively unfreezing some intermediate layers so that the model can further learn according to specific task data.
6. The intelligent transfer learning method of a large multi-modal model of a power grid according to claim 1, characterized in that: The introduction of a dedicated domain adaptive network, based on a layer-by-layer adversarial method, adds an adversarial domain classifier to each layer of features represented by the power grid data, and eliminates the feature deviation between the source data and the target data distribution layer by layer, as follows: An adversarial domain classifier is added after the feature representation of each layer of the network, and the difference in feature distribution between the source domain and the target domain is reduced through adversarial optimization; Layer-by-layer adversarial network structure: For the power grid task, suppose the feature extraction network has L layers, and for each layer of feature H l , add a domain classifier D l : H l =f l (H l-1 ),D l (H l )=softmax(W l ·H l +b l ); Among them, D l The role of is to judge the feature H l From the source domain or the target domain; Using adversarial training to make the domain classifier unable to distinguish the source and target domain features: Train Dl to accurately predict the domain label y of the input data domain ; Reverse optimization feature extraction network f l Make D l The classification error rate is as high as possible, thereby ignoring the domain information; the loss function is adversarial loss: The feature generator f l The optimization goal is to maximize the prediction error of the domain classifier by achieving adversarial effects through reverse gradient; Through adversarial training, the domain characteristics of the input data are eliminated layer by layer, and finally the domain-free representation of the features is achieved.
7. The intelligent transfer learning method of a large multi-modal model of a power grid according to claim 6 is characterized in that: The physical perception constraints are used to limit the physical rationality of the features in the domain adaptation process, as follows: When aligning data features layer by layer, direct adversarial training may cause the physical meaning of the features to be lost, introduce physical constraints of the power grid, and limit the rationality of feature alignment; Grid faults will dynamically propagate along the topology of device interconnection. Topology consistency constraints are expressed through the grid topology matrix A and the propagation neighbor characteristics F. In topological propagation, if the node i represented by the network feature H fails, its neighbor features also show certain similarities; the topological neighbor consistency loss is introduced: Where (i, j) is the adjacent device in the fault propagation path; And through the conservation of physical variables, constraints are imposed. Specifically, for node i in the topology diagram, according to the power distribution: Introducing the physical non-conservation constraint loss L physics : Get the final optimization goal: Among them, λ1,λ2 are weight hyperparameters.
8. A real-time coordinated risk dispatching system for integrated main and distribution microgrids, characterized in that: It includes a processor, a memory and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the real-time coordinated risk scheduling method for integrated main and distribution microgrids as described in any one of claims 1 to 7.
Citation Information
Cited By
Flow casting material effect prediction method and system based on multi-modal deep learning
CN120374200A
Transfer learning-based cross-regional electric meter anomaly detection method and system, and electric energy meter
CN120822157A
Power system scheduling method and device based on multi-modal data fusion
CN120879789A
Training method and device of video generation model
CN121056659A
Power load scheduling method and device based on large language model
CN121503946A