Power grid cloud edge cooperative computing optimization method and system based on cross-modal adversarial learning
By introducing coarse and fine granularity feature perturbations in the power grid cloud-edge collaborative architecture, constructing the video convex hull and strengthening the temporal correlation, the problem of poor cross-model migration of black box attacks is solved, and efficient adversarial sample generation and intelligent analysis of image-video heterogeneous models are achieved.
Patent Information
- Application Number
- CN202510809304.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
Under the power grid cloud-edge collaborative architecture, black-box adversarial attacks have poor cross-model transferability, resulting in inefficient adversarial sample generation. Existing technologies are also unable to effectively deal with adversarial attacks in image-video heterogeneous model collaborative scenarios.
Coarse-grained and fine-grained feature perturbations are introduced. The video convex hull is constructed through coarse-grained feature perturbations and long-term dependencies are modeled using global temporal constraints. Fine-grained feature perturbations focus on the local manifold structure of the video and strengthen short-term temporal correlation by constraining the correlation between two adjacent frames. A dual-granularity joint loss is established to achieve synergistic enhancement of global video frame deception and local adjacent frame perturbations.
It improves the intelligent analysis level of cloud-edge cross-modal models, lowers the implementation threshold of video adversarial attacks, is suitable for image-video heterogeneous model collaboration scenarios, and enhances the cross-model portability and robustness of adversarial samples.
Smart Images

Figure CN120708125A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power system optimization, and specifically relates to a power grid cloud-edge collaborative computing optimization method and system based on cross-modal adversarial learning. Background Art
[0002] At present, the new energy industry represented by distributed photovoltaics is experiencing explosive growth, and innovative business models such as virtual power plants and integrated energy services are emerging rapidly, indicating that the power system is undergoing the most profound technological changes since electrification.
[0003] As artificial power systems continue to expand in scale, become increasingly complex, and the pace of energy transition accelerates, the power grid, with its numerous geographically diverse sites, is experiencing a strong demand for personalized computing. This necessitates the use of artificial intelligence (AI) to comprehensively enhance the intelligence of system state perception, operational cognition, and decision-making and control. An effective solution involves optimizing cloud-edge model migration and sharing new knowledge based on a two-tier collaborative architecture: a cloud (grid company headquarters) and an edge (local power company). The cloud-side model is responsible for training and regularly sends updated model parameters to the edge-side model. The edge-side model performs local inference and transmits new collected samples to the cloud-side model, enabling cross-tier and cross-scenario collaboration of data, knowledge, and models. However, with this continuous advancement in technology come vulnerabilities. A phenomenon known as adversarial attacks poses a serious challenge to the accuracy and reliability of AI models. Adversarial attacks involve adding small perturbations to the input, causing the classifier to misclassify. In a cloud-edge collaborative architecture, adversarial attacks primarily manifest as a small adversarial perturbation on the edge side that can destabilize the cloud-side power system. For example, by tampering with distributed photovoltaic inverter data or sensor readings, a small perturbation can be injected, misleading cloud-side state estimation or scheduling decisions. Adversarial attack research has revealed the vulnerabilities of deep learning models and promoted the development of robust models.
[0004] Adversarial attacks were first systematically proposed by Szegedy et al. for image classification tasks. With the application of deep learning in video analysis, this technique has gradually expanded to the video field in recent years. Adversarial attacks can be broadly categorized into two types: white-box attacks based on complete model knowledge and black-box attacks that rely solely on input and output. In white-box attacks, the attacker possesses complete internal information about the target model, including its architecture, parameters, training data, and gradient calculation methods. This allows them to directly construct adversarial examples using backpropagation or optimization methods. In contrast, in black-box attacks, the attacker can only infer model behavior through input-output interactions, without knowledge of the model's internal structure, training parameters, or algorithms. Consequently, they typically rely on transfer learning or heuristic search to generate adversarial inputs through repeated trial and error. Compared to white-box testing, black-box testing is more challenging and more widely used in real-world scenarios because it requires no knowledge of the system's internal implementation and is more realistic for attack and defense. Black-box adversarial attacks are categorized into query-based and transfer-based attacks. Query-based attacks involve attackers repeatedly sending carefully crafted query samples to the target model and observing the model's output. By analyzing the correlations between different query samples and outputs, they infer the model's properties and vulnerabilities, thereby constructing effective adversarial samples. However, this approach requires a large number of model queries to obtain adversarial samples, resulting in low attack efficiency. Transfer-based attacks first use a white-box attack algorithm on an accessible white-box model to generate adversarial samples, and then transfer these adversarial samples to other black-box models for attack. The key challenge of this attack method is how to effectively improve the cross-model transferability of adversarial samples. Summary of the Invention
[0005] The purpose of the present invention is to address the problems in the above-mentioned prior art and provide a power grid cloud-edge collaborative computing optimization method and system based on cross-modal adversarial learning, introduce two types of feature disturbances of coarse and fine granularity, effectively combine the advantages of knowledge sedimentation in the power field, realize migration optimization and knowledge sharing between cross-modal cloud-edge models, and improve the intelligent analysis level of cloud-edge cross-modal models.
[0006] In order to achieve the above objectives, the present invention has at least the following beneficial effects:
[0007] In the first aspect, a power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning is provided, including:
[0008] Obtain the original video of the power grid station and divide the original video sequence into a frame set;
[0009] The method introduces both coarse-grained and fine-grained feature perturbations to a collection of frames for adversarial learning. The coarse-grained feature perturbation constructs the video convex hull, expands the single-frame perturbation into a spatiotemporal joint optimization problem, and uses global temporal constraints to model long-term dependencies. Fine-grained feature perturbations focus on the local manifold structure of the video, and constrain the correlation between adjacent frames to enhance short-term temporal correlations.
[0010] A dual-granularity joint loss is established for both coarse and fine-grained feature perturbations to achieve synergistic enhancement of global video frame deception and local adjacent frame perturbations, and adversarial samples are obtained based on the dual-granularity joint loss solution.
[0011] As a preferred solution, in the step of introducing coarse and fine granularity feature perturbations to the frame set for adversarial learning, the optimization variable of the adversarial learning is set to the corresponding perturbation on each frame in the original video sequence, and the search space dimension is the same as the video frame size. The expression is as follows:
[0012] X adv =X+Δ
[0013] f(X adv )≠y
[0014] st||Δ|| ∞ <ε
[0015] Where, For the original video, To combat disturbances, X adv represents the adversarial sample, f(·) represents the model output and adopts a black box model, y∈{1,2,...,K} is the category label; l of constraint Δ ∞ norm, ||·|| represents the size of the adversarial interference, and ε is the threshold for limiting the size of the adversarial interference.
[0016] As a preferred solution, in the step of constructing the video convex hull through coarse-grained feature perturbations, expanding the single-frame perturbation into a spatiotemporal joint optimization problem, and using global temporal constraints to model long-term dependencies, the perturbation of each point in the convex hull space carries the correlation information between frames, thereby achieving a temporally consistent adversarial attack. The mathematical expression is as follows:
[0017]
[0018] Where x i Represents the original video The i-th frame in represents a real number, T represents the number of video frames, H represents the height of the video frame, W represents the width of the video frame, C represents the number of channels of the video frame feature vector, α i Represents frame x i The weight coefficient of
[0019] Add interference to each point in the convex hull according to the following formula to generate an adversarial sample and construct the video convex hull:
[0020]
[0021] Where, δ i Represents frame x i adversarial disturbances;
[0022] By minimizing the cosine similarity between the original video sample and the adversarial sample, the model is induced to misclassify:
[0023]
[0024] Where E[·] represents the expected calculation, (u, v) represents the original sample and the adversarial sample pair, g represents the pre-trained image classification model, and g k (u) represents the feature vector of the kth layer of model g, g k (u) T represents the transpose of the feature vector of the kth layer of the original sample extracted by model g, g k (v) represents the feature vector of the k-th layer of the adversarial sample extracted by model g, and ||·|| specifically refers to the l2 norm.
[0025] As a preferred solution, in the step of focusing on the local manifold structure of the video through fine-grained feature perturbation and strengthening the short-term temporal correlation by constraining the correlation between two adjacent frames, the following loss function is designed to reduce the similarity between adjacent frames of the adversarial sample:
[0026]
[0027] Where g k (x i +δ i ) T represents the transpose of the feature vector of the kth layer of the adversarial sample frame extracted by model g.
[0028] As a preferred solution, in the step of establishing a dual-granularity joint loss for the two types of feature perturbations of coarse and fine granularity to achieve synergistic enhancement of global video frame deception and local adjacent frame perturbations, the mathematical expression of the dual-granularity joint loss is as follows:
[0029]
[0030] Where, L T represents the dual-granularity joint loss function, represents the minimization loss function for the parameters to be learned, and β represents the weight for balancing coarse and fine-grained features.
[0031] Secondly, a power grid cloud-edge collaborative computing optimization system based on cross-modal adversarial learning is provided, including:
[0032] The frame set acquisition module is used to obtain the original video of the power grid station and divide the original video sequence into frame sets;
[0033] The coarse-grained and fine-grained feature perturbation module is used to introduce both coarse-grained and fine-grained feature perturbations to a collection of frames for adversarial learning. The coarse-grained feature perturbation constructs the video convex hull, expands the single-frame perturbation into a spatiotemporal joint optimization problem, and uses global temporal constraints to model long-term dependencies. Fine-grained feature perturbation focuses on the local manifold structure of the video and constrains the correlation between adjacent frames to enhance short-term temporal correlations.
[0034] The collaborative enhancement module is used to establish a dual-granularity joint loss for both coarse and fine-grained feature perturbations to achieve collaborative enhancement of global video frame deception and local adjacent frame perturbations, and obtain adversarial samples based on the dual-granularity joint loss solution.
[0035] As a preferred solution, when the coarse-grained and fine-grained feature perturbation module introduces two types of coarse-grained and fine-grained feature perturbations to the frame set for adversarial learning, the optimization variable of the adversarial learning is set to the corresponding perturbation on each frame in the original video sequence, and the search space dimension is the same as the video frame size. The expression is as follows:
[0036] X adv =X+Δ
[0037] f(X adv )≠y
[0038] st||Δ|| ∞ <ε
[0039] Where, For the original video, To combat disturbances, X adv represents the adversarial sample, f(·) represents the model output and adopts a black box model, y∈{1,2,...,K} is the category label; l of constraint Δ ∞ norm, ||·|| represents the size of the adversarial interference, and ε is the threshold for limiting the size of the adversarial interference.
[0040] As a preferred solution, the coarse-grained and fine-grained feature perturbation module constructs the video convex hull through coarse-grained feature perturbation, expands the single-frame perturbation into a spatiotemporal joint optimization problem, and uses global temporal constraints to model long-term dependencies. In this way, the perturbation of each point in the convex hull space carries the correlation information between frames, thereby achieving a temporally consistent adversarial attack. The mathematical expression is as follows:
[0041]
[0042] Where x i Represents the original video The i-th frame in represents a real number, T represents the number of video frames, H represents the height of the video frame, W represents the width of the video frame, C represents the number of channels of the video frame feature vector, α i Represents frame x i The weight coefficient of
[0043] Add interference to each point in the convex hull according to the following formula to generate an adversarial sample and construct the video convex hull:
[0044]
[0045] Where, δ i Represents frame x i adversarial disturbances;
[0046] By minimizing the cosine similarity between the original video sample and the adversarial sample, the model is induced to misclassify:
[0047]
[0048] Where E[·] represents the expected calculation, (u, v) represents the original sample and the adversarial sample pair, g represents the pre-trained image classification model, and g k (u) represents the feature vector of the kth layer of model g, g k (u) T represents the transpose of the feature vector of the kth layer of the original sample extracted by model g, g k (v) represents the feature vector of the k-th layer of the adversarial sample extracted by model g, and ||·|| specifically refers to the l2 norm.
[0049] As a preferred solution, the coarse-grained and fine-grained feature perturbation module focuses on the local manifold structure of the video through fine-grained feature perturbation and strengthens the short-term temporal correlation by constraining the correlation between two adjacent frames. The following loss function is designed to reduce the similarity between adjacent frames of the adversarial sample:
[0050]
[0051] Where g k (x i +δ i ) T represents the transpose of the feature vector of the kth layer of the adversarial sample frame extracted by model g.
[0052] As a preferred solution, when the collaborative enhancement module establishes a dual-granularity joint loss for the two feature perturbations of coarse and fine granularity to achieve collaborative enhancement of global video frame deception and local adjacent frame perturbation, the mathematical expression of the dual-granularity joint loss is as follows:
[0053]
[0054] Where, L T represents the dual-granularity joint loss function, represents the minimization loss function for the parameters to be learned, and β represents the weight for balancing coarse and fine-grained features.
[0055] In a third aspect, an electronic device is provided, comprising a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning.
[0056] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning is implemented.
[0057] Compared with the prior art, the first aspect of the present invention has at least the following beneficial effects:
[0058] Compared to traditional adversarial attack methods, which are limited to single-modal attack scenarios (targeting only images or videos), this paper proposes a power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning. This method effectively addresses adversarial attack challenges in heterogeneous image-video model collaboration scenarios and is applicable to business scenarios where image models are used on the edge and video models are used on the cloud. Specifically, the adversarial learning method introduces both coarse-grained and fine-grained feature perturbations. Coarse-grained feature perturbations construct the video convex hull, extending single-frame perturbations into a spatiotemporal joint optimization problem and modeling long-term dependencies using global temporal constraints. Fine-grained feature perturbations focus on the local manifold structure of the video, enhancing short-term temporal correlations by constraining the high correlation between adjacent frames. This method effectively leverages the advantages of accumulated knowledge in the power sector, enabling transfer optimization and knowledge sharing between cross-modal cloud-edge models, and improving the intelligent analysis capabilities of cloud-edge cross-modal models. Furthermore, by replacing the traditional model-generated video adversarial samples with pre-trained images, this method significantly lowers the threshold for implementing video adversarial attacks and offers significant practical value.
[0059] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0061] Figure 1 A schematic diagram of a flow chart of a power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning in an embodiment of the present invention;
[0062] Figure 2 Structural block diagram of the power grid cloud-edge collaborative computing optimization system based on cross-modal adversarial learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0063] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0064] See also Figure 1 , an embodiment of the present invention proposes a power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning, which mainly includes the following steps:
[0065] S1. Obtain the original video of the power grid station and divide the original video sequence into a frame set;
[0066] S2. We introduce both coarse-grained and fine-grained feature perturbations to the frame set for adversarial learning. We construct the video convex hull through coarse-grained feature perturbations, expand the single-frame perturbation into a spatiotemporal joint optimization problem, and use global temporal constraints to model long-term dependencies. We focus on the local manifold structure of the video through fine-grained feature perturbations, and strengthen short-term temporal correlations by constraining the correlation between adjacent frames.
[0067] S3. A dual-granularity joint loss is established for both coarse and fine-grained feature perturbations to achieve synergistic enhancement of global video frame deception and local adjacent frame perturbations, and adversarial samples are obtained based on the dual-granularity joint loss solution.
[0068] In cloud-edge collaborative computing optimization, raw power grid video involves video data from various scenarios. This data is the basis for intelligent video inspection and monitoring of the power grid, including but not limited to the following:
[0069] 1. Overhead line video:
[0070] Real-time video captured by fixed inspection cameras installed on line towers is used to monitor the operating status of transmission lines.
[0071] The images captured by the camera at regular intervals are used for subsequent analysis.
[0072] 2. Drone inspection video:
[0073] The video captured by the drone flying over the key parts of the transmission lines and towers contains detailed information about the transmission lines.
[0074] 3. Substation video:
[0075] Videos captured by cameras installed in substations are used to monitor the operating status and safety of substation equipment.
[0076] 4. Distribution room video:
[0077] The video captured by the monitoring equipment installed in the power distribution room is used to monitor the equipment status and environmental conditions in the power distribution room in real time.
[0078] 5. Transmission tunnel gallery video:
[0079] Videos captured by monitoring equipment such as cameras and tunnel robots installed in transmission tunnel corridors are used to monitor the status and safety of facilities in the tunnel corridors.
[0080] Significant geographical differences lead to huge differences in data characteristics, equipment configuration, network conditions and business needs among sites, which in turn creates a strong demand for personalized computing.
[0081] Cloud-edge collaborative computing needs to balance computing efficiency and resource costs through edge node customization, cloud-edge collaborative strategy optimization, standardization and modular design to meet the differentiated needs of different sites.
[0082] In one possible implementation, an adversarial attack can be formalized as the process of generating adversarial examples and successfully deceiving a neural network into misclassifying them. This is essentially an optimization problem, with the core goal being to find an appropriate perturbation for each frame in a video sequence. Therefore, in step S2 of this embodiment of the present invention, the optimization variables for adversarial learning are set to the corresponding perturbations for each frame in the original video sequence. The search space dimension is the same as the video frame size, as expressed as follows:
[0083] X adv =X+Δ
[0084] f(X adv )≠y
[0085] st||Δ|| ∞ <ε
[0086] Where, For the original video, To combat disturbances, X adv represents the adversarial sample, f(·) represents the model output and adopts a black box model, y∈{1,2,...,K} is the category label; l of constraint Δ ∞ norm, ||·|| represents the size of the adversarial interference, and ε is the threshold for limiting the size of the adversarial interference.
[0087] In one possible implementation, video adversarial attacks require explicit modeling of temporal consistency constraints, while traditional methods (such as FGSM-Video) usually simply add perturbations to the video frame by frame independently, causing the generated adversarial samples to lose the inherent global temporal correlation of the video frames. Intuitively, establishing interactions between frames is equivalent to introducing cross-frame lateral perturbations for each frame, which can effectively enhance the robustness and diversity of gradient optimization. Therefore, step S2 of the embodiment of the present invention constructs a video convex hull through coarse-grained feature perturbations, expands single-frame perturbations into a spatiotemporal joint optimization problem, strengthens the mutual correlation between frames, and uses global temporal constraints to model long-term dependencies. In the convex hull space, the perturbation of each point carries the correlation information between frames, thereby achieving temporally consistent adversarial attacks. The mathematical expression is as follows:
[0088]
[0089] Where x i Represents the original video The i-th frame in represents a real number, T represents the number of video frames, H represents the height of the video frame, W represents the width of the video frame, C represents the number of channels of the video frame feature vector, α i Represents frame x i The weight coefficient of
[0090] Add interference to each point in the convex hull according to the following formula to generate an adversarial sample and construct the video convex hull:
[0091]
[0092] Where, δ i Represents frame x i adversarial disturbances;
[0093] By minimizing the cosine similarity between the original video sample and the adversarial sample, the model is induced to misclassify:
[0094]
[0095] Where E[·] represents the expected calculation, (u, v) represents the original sample and the adversarial sample pair, g represents the pre-trained image classification model, and gk (u) represents the feature vector of the kth layer of model g, g k (u) T represents the transpose of the feature vector of the kth layer of the original sample extracted by model g, g k (v) represents the feature vector of the k-th layer of the adversarial sample extracted by model g, and ||·|| specifically refers to the l2 norm.
[0096] In one possible implementation, coarse-grained feature perturbations model the long-term dependencies of adversarial samples through global temporal constraints, while fine-grained feature perturbations focus on the local manifold structure of the video and strengthen short-term temporal correlations through adjacent frame similarity constraints. In fact, a video can be regarded as a smooth extension of a still image in the temporal dimension, and there is a high correlation between adjacent frames. It is this temporal-local correlation that forms the unique manifold structure of the video and distinguishes the video from a collection of disordered images. Based on this, step S2 designs the following loss function to reduce the similarity between adjacent frames of the adversarial sample:
[0097]
[0098] Where g k (x i +δ i ) T represents the transpose of the feature vector of the kth layer of the adversarial sample frame extracted by model g.
[0099] In one possible implementation, step S3 establishes a mathematical expression for the dual-granularity joint loss as follows:
[0100]
[0101] Where, L T represents the dual-granularity joint loss function, represents the minimization loss function for the parameters to be learned, and β represents the weight for balancing coarse and fine-grained features.
[0102] The embodiment of the present invention solves the adversarial attack problem in the "image-video" heterogeneous model collaboration scenario based on the power grid cloud-edge collaborative computing optimization method of cross-modal adversarial learning, and is applicable to business scenarios where the edge side is an image model and the cloud side is a video model. The adversarial learning method proposed in the present invention introduces two types of feature perturbations, coarse and fine granularity. The coarse-grained feature perturbation constructs the video convex hull, expands the single-frame perturbation into a spatiotemporal joint optimization problem, and uses global temporal constraints to model long-term dependencies. The fine-grained feature perturbation focuses on the local manifold structure of the video and strengthens the short-term temporal correlation by constraining the high correlation between two adjacent frames. The present invention effectively combines the advantages of knowledge sedimentation in the power field, realizes migration optimization and knowledge sharing between cross-modal cloud-edge models, and improves the intelligent analysis level of cloud-edge cross-modal models. In addition, the present invention generates video adversarial samples by pre-training the image substitution model, thereby significantly reducing the implementation threshold of video adversarial attacks and has important practical value.
[0103] See also Figure 2 The embodiment of the present invention provides a power grid cloud-edge collaborative computing optimization system based on cross-modal adversarial learning, including:
[0104] The frame set acquisition module 210 is used to acquire the original video of the power grid station and divide the original video sequence into frame sets;
[0105] The coarse-grained and fine-grained feature perturbation module 220 is used to introduce both coarse-grained and fine-grained feature perturbations to a set of frames for adversarial learning. The coarse-grained feature perturbation constructs the video convex hull, expands the single-frame perturbation into a spatiotemporal joint optimization problem, and uses global temporal constraints to model long-term dependencies. Fine-grained feature perturbation focuses on the local manifold structure of the video and constrains the correlation between adjacent frames to enhance short-term temporal correlation.
[0106] The collaborative enhancement module 230 is used to establish a dual-granularity joint loss for the two types of feature perturbations of coarse and fine granularity to achieve collaborative enhancement of global video frame deception and local adjacent frame perturbations, and obtain adversarial samples based on the dual-granularity joint loss solution.
[0107] In one possible implementation, when the coarse-grained and fine-grained feature perturbation module 220 introduces both coarse-grained and fine-grained feature perturbations to a frame set for adversarial learning, the optimization variable of the adversarial learning is set to the corresponding perturbation on each frame in the original video sequence, and the search space dimension is the same as the video frame size, as expressed as follows:
[0108] X adv =X+Δ
[0109] f(X adv )≠y
[0110] st||Δ|| ∞ <ε
[0111] Where, For the original video, To combat disturbances, X adv represents the adversarial sample, f(·) represents the model output and adopts a black box model, y∈{1,2,...,K} is the category label; l of constraint Δ ∞ norm, ||·|| represents the size of the adversarial interference, and ε is the threshold for limiting the size of the adversarial interference.
[0112] In one possible implementation, the coarse-grained and fine-grained feature perturbation module 220 constructs the video convex hull through coarse-grained feature perturbation, expands the single-frame perturbation into a spatiotemporal joint optimization problem, and uses global temporal constraints to model long-term dependencies. The perturbation of each point in the convex hull space carries the correlation information between frames, thereby achieving a temporally consistent adversarial attack. The mathematical expression is as follows:
[0113]
[0114] Where x i Represents the original video The i-th frame in represents a real number, T represents the number of video frames, H represents the height of the video frame, W represents the width of the video frame, C represents the number of channels of the video frame feature vector, α i Represents frame x i The weight coefficient of
[0115] Add interference to each point in the convex hull according to the following formula to generate an adversarial sample and construct the video convex hull:
[0116]
[0117] Where, δ i Represents frame x i adversarial disturbances;
[0118] By minimizing the cosine similarity between the original video sample and the adversarial sample, the model is induced to misclassify:
[0119]
[0120] Where E[·] represents the expected calculation, (u, v) represents the original sample and the adversarial sample pair, g represents the pre-trained image classification model, and g k (u) represents the feature vector of the kth layer of model g, g k (u) T represents the transpose of the feature vector of the kth layer of the original sample extracted by model g, g k (v) represents the feature vector of the k-th layer of the adversarial sample extracted by model g, and ||·|| specifically refers to the l2 norm.
[0121] In one possible implementation, the coarse-grained feature perturbation module 220 focuses on the local manifold structure of the video through fine-grained feature perturbation and strengthens the short-term temporal correlation by constraining the correlation between two adjacent frames. The following loss function is designed to reduce the similarity between adjacent frames of the adversarial sample:
[0122]
[0123] Where g k (x i +δ i ) T represents the transpose of the feature vector of the kth layer of the adversarial sample frame extracted by model g.
[0124] In one possible implementation, when the collaborative enhancement module 230 establishes a dual-granularity joint loss for the coarse and fine granularity feature perturbations to achieve collaborative enhancement of global video frame deception and local adjacent frame perturbations, the mathematical expression of the established dual-granularity joint loss is as follows:
[0125]
[0126] Where, L T represents the dual-granularity joint loss function, represents the minimization loss function for the parameters to be learned, and β represents the weight for balancing coarse and fine-grained features.
[0127] Another embodiment of the present invention further proposes an electronic device, including a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning.
[0128] Another embodiment of the present invention further proposes a computer-readable storage medium, which stores at least one instruction. When the at least one instruction is executed by a processor, it implements the power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning.
[0129] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals. For ease of explanation, the above content only shows the part related to the embodiment of the present invention. For specific technical details not disclosed, please refer to the method part of the embodiment of the present invention. The computer-readable storage medium is non-transitory and can be stored in a storage device formed by various electronic devices, and can implement the execution process recorded in the method of the embodiment of the present invention.
[0130] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0131] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0132] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0133] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning, characterized by: include: Obtain the original video of the power grid station and divide the original video sequence into a frame set; The method introduces both coarse-grained and fine-grained feature perturbations to a collection of frames for adversarial learning. The coarse-grained feature perturbation constructs the video convex hull, expands the single-frame perturbation into a spatiotemporal joint optimization problem, and uses global temporal constraints to model long-term dependencies. Fine-grained feature perturbations focus on the local manifold structure of the video, and constrain the correlation between adjacent frames to enhance short-term temporal correlations. A dual-granularity joint loss is established for both coarse and fine-grained feature perturbations to achieve synergistic enhancement of global video frame deception and local adjacent frame perturbations, and adversarial samples are obtained based on the dual-granularity joint loss solution.
2. The power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning according to claim 1 is characterized in that: In the step of introducing coarse and fine granularity feature perturbations to the frame set for adversarial learning, the optimization variable of the adversarial learning is set to the corresponding perturbation on each frame in the original video sequence. The search space dimension is the same as the video frame size, and the expression is as follows: X adv =X+Δ f(X adv )≠y st||D|| ∞ <e Where, For the original video, To combat disturbances, X adv represents the adversarial sample, f(·) represents the model output and adopts a black box model, y∈{1,2,...,K} is the category label; l of constraint Δ ∞ norm, ||·|| represents the size of the adversarial interference, and ε is the threshold for limiting the size of the adversarial interference.
3. The power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning according to claim 2 is characterized in that: In the steps of constructing the video convex hull through coarse-grained feature perturbations, expanding the single-frame perturbation into a spatiotemporal joint optimization problem, and using global temporal constraints to model long-term dependencies, the perturbation of each point in the convex hull space carries the correlation information between frames, thereby achieving a temporally consistent adversarial attack. The mathematical expression is as follows: Where x i Represents the original video The i-th frame in represents a real number, T represents the number of video frames, H represents the height of the video frame, W represents the width of the video frame, C represents the number of channels of the video frame feature vector, α i Represents frame x i The weight coefficient of Add interference to each point in the convex hull according to the following formula to generate an adversarial sample and construct the video convex hull: Where, δ i Represents frame x i adversarial disturbances; By minimizing the cosine similarity between the original video sample and the adversarial sample, the model is induced to misclassify: Where E[·] represents the expected calculation, (u, v) represents the original sample and the adversarial sample pair, g represents the pre-trained image classification model, and g k (u) represents the feature vector of the kth layer of model g, g k (u) T represents the transpose of the feature vector of the kth layer of the original sample extracted by model g, g k (v) represents the feature vector of the k-th layer of the adversarial sample extracted by model g, and ||·|| specifically refers to the l2 norm.
4. The power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning according to claim 3 is characterized in that: In the step of focusing on the local manifold structure of the video through fine-grained feature perturbation and strengthening the short-term temporal correlation by constraining the correlation between two adjacent frames, the following loss function is designed to reduce the similarity between adjacent frames of the adversarial sample: Where g k (x i +δ i ) T represents the transpose of the feature vector of the kth layer of the adversarial sample frame extracted by model g.
5. The power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning according to claim 4 is characterized in that: In the step of establishing a dual-granularity joint loss for the two types of feature perturbations of coarse and fine granularity to achieve the synergistic enhancement of global video frame deception and local adjacent frame perturbations, the mathematical expression of the dual-granularity joint loss is as follows: Where, L T represents the dual-granularity joint loss function, represents the minimization loss function for the parameters to be learned, and β represents the weight for balancing coarse and fine-grained features.
6. A power grid cloud-edge collaborative computing optimization system based on cross-modal adversarial learning, characterized by: include: The frame set acquisition module is used to obtain the original video of the power grid station and divide the original video sequence into frame sets; The coarse-grained and fine-grained feature perturbation module is used to introduce both coarse-grained and fine-grained feature perturbations to a collection of frames for adversarial learning. The coarse-grained feature perturbation constructs the video convex hull, expands the single-frame perturbation into a spatiotemporal joint optimization problem, and uses global temporal constraints to model long-term dependencies. Fine-grained feature perturbation focuses on the local manifold structure of the video and constrains the correlation between adjacent frames to enhance short-term temporal correlations. The collaborative enhancement module is used to establish a dual-granularity joint loss for both coarse and fine-grained feature perturbations to achieve collaborative enhancement of global video frame deception and local adjacent frame perturbations, and obtain adversarial samples based on the dual-granularity joint loss solution.
7. The power grid cloud-edge collaborative computing optimization system based on cross-modal adversarial learning according to claim 6 is characterized in that: When the coarse-grained and fine-grained feature perturbation module introduces two types of feature perturbations to the frame set for adversarial learning, the optimization variable of the adversarial learning is set to the corresponding perturbation on each frame in the original video sequence. The search space dimension is the same as the video frame size, and the expression is as follows: X adv =X+Δ f(X adv )≠y st||D|| ∞ <e Where, For the original video, To combat disturbances, X adv represents the adversarial sample, f(·) represents the model output and adopts a black box model, y∈{1,2,…,K} is the category label; l of constraint Δ ∞ norm, ||·|| represents the size of the adversarial interference, and ε is the threshold for limiting the size of the adversarial interference.
8. The power grid cloud-edge collaborative computing optimization system based on cross-modal adversarial learning according to claim 7 is characterized in that: The coarse-grained and fine-grained feature perturbation module constructs the video convex hull through coarse-grained feature perturbation, expands the single-frame perturbation into a spatiotemporal joint optimization problem, and uses global temporal constraints to model long-term dependencies. The perturbation of each point in the convex hull space carries the correlation information between frames, thereby achieving a temporally consistent adversarial attack. The mathematical expression is as follows: Where x i Represents the original video The i-th frame in represents a real number, T represents the number of video frames, H represents the height of the video frame, W represents the width of the video frame, C represents the number of channels of the video frame feature vector, α i Represents frame x i The weight coefficient of Add interference to each point in the convex hull according to the following formula to generate an adversarial sample and construct the video convex hull: Where, δ i Represents frame x i adversarial disturbances; By minimizing the cosine similarity between the original video sample and the adversarial sample, the model is induced to misclassify: Where E[·] represents the expected calculation, (u, v) represents the original sample and the adversarial sample pair, g represents the pre-trained image classification model, and g k (u) represents the feature vector of the kth layer of model g, g k (u) T represents the transpose of the feature vector of the kth layer of the original sample extracted by model g, g k (v) represents the feature vector of the k-th layer of the adversarial sample extracted by model g, and ||·|| specifically refers to the l2 norm.
9. The power grid cloud-edge collaborative computing optimization system based on cross-modal adversarial learning according to claim 8 is characterized in that: The coarse-grained and fine-grained feature perturbation module focuses on the local manifold structure of the video through fine-grained feature perturbation and strengthens the short-term temporal correlation by constraining the correlation between two adjacent frames. The following loss function is designed to reduce the similarity between adjacent frames of the adversarial sample: Where g k (x i +δ i ) T represents the transpose of the feature vector of the kth layer of the adversarial sample frame extracted by model g.
10. The power grid cloud-edge collaborative computing optimization system based on cross-modal adversarial learning according to claim 9 is characterized in that: When the collaborative enhancement module establishes a dual-granularity joint loss for the two feature perturbations of coarse and fine granularity to achieve collaborative enhancement of global video frame deception and local adjacent frame perturbation, the mathematical expression of the established dual-granularity joint loss is as follows: Where, L T represents the dual-granularity joint loss function, represents the minimization loss function for the parameters to be learned, and β represents the weight for balancing coarse and fine-grained features.
11. An electronic device, characterized in that: It includes a processor and a memory, and the processor is used to execute a computer program stored in the memory to implement the power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning as described in any one of claims 1 to 5.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, it implements the power grid cloud-edge collaborative computing optimization method based on cross-modal adversarial learning as described in any one of claims 1 to 5.