Discrete continuous mixed action confrontation attack method based on deep reinforcement learning
By designing discrete continuous hybrid action adversarial attack methods in the edge computing system, generating adversarial perturbation and using alternative models to perform black box attacks, the security vulnerability problem of deep reinforcement learning models in task unloading and resource allocation scenarios is solved, and the effectiveness of the attack and the robustness of the system are improved.
Patent Information
- Application Number
- CN202510766838.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The prior art has failed to effectively study the adversarial attack problem of deep reinforcement learning models in edge computing systems under task offloading and resource allocation scenarios, especially the potential vulnerability under dynamic system state vector input, resulting in the inadequate disclosure of security risks.
A discrete continuous hybrid action adversarial attack method based on deep reinforcement learning is designed. By constructing an alternative model to approximate the target strategy, it generates adversarial perturbation and attacks under the black box attack framework, evaluates the impact of attacks on delay, energy consumption and service quality, and generates adversarial perturbation in combination with the migration paradigm to reveal potential security risks.
It realizes efficient attack resistance in edge computing scenarios, improves the concealment and destructive power of attacks, provides a robust framework, provides theoretical support for the security design of edge intelligent systems, and improves generalization capabilities.
Smart Images

Figure CN120277664A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence security, and particularly to a discrete and continuous hybrid action adversarial attack method based on deep reinforcement learning. Background Art
[0002] By integrating deep learning and reinforcement learning frameworks, deep reinforcement learning (DRL) can effectively capture the characteristics of dynamic environments, achieving a breakthrough in the powerful representation ability of high-dimensional state spaces and complex decision-making abilities. However, the neural network components it depends on are also vulnerable to adversarial perturbations. The sensitivity of deep reinforcement learning models to input perturbations may pose serious security risks - attackers can inject subtle adversarial noises to make the models make irrational decisions, thereby leading to a decline in the quality of system services. For example, in autonomous driving, adversarial attacks on environmental perception models may mislead vehicle trajectory planning; in the field of medical health, tampering with medical image inputs may cause diagnostic errors. Current research on adversarial attacks against deep reinforcement learning mostly focuses on vision-driven scenarios (such as image classification, autonomous driving), and its attack methods are usually designed for pixel-level static data. However, in the task offloading scenario based on edge computing, the input of DRL is essentially a dynamic system state vector, which is constrained by real physical rules and consists of a small number of parameters with clear physical meanings and there are explicit physical associations between the parameters, which is fundamentally different from images.
[0003] Therefore, it is very necessary to study the adversarial attack problem of deep reinforcement learning models in the task offloading and resource allocation scenarios of edge computing systems, revealing the potential vulnerability of DRL policies driven by dynamic system states under adversarial perturbations, and further injecting new vitality into the security research of deep reinforcement learning models in adversarial attack scenarios. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to design a discrete and continuous hybrid action adversarial attack method based on deep reinforcement learning for the task offloading and resource allocation scenarios in edge computing systems, evaluate the destruction intensity of the attack on core indicators such as delay, energy consumption, and quality of service, and finally design a black-box attack framework based on the transferability paradigm, approximate the decision-making behavior of the target policy by constructing a surrogate model, and generate adversarial perturbations using surrogate gradient information, revealing the potential security risks of DRL in task offloading and resource allocation in edge computing scenarios, and providing theoretical support for constructing a robust edge intelligent computing framework.
[0005] The technical solution adopted by the present invention to solve the above technical problems is that the attacker obtains the original state observed by the agent in the environment, uses the discrete and continuous hybrid action adversarial perturbation attack method based on deep reinforcement learning to generate adversarial perturbations, and superimposes the adversarial perturbations and the original system observations and sends them to the agent as the new observation state, causing the agent to make irrational decision-making actions. On this basis, further use the black-box attack framework based on the transferability paradigm to perform adversarial attacks on the DRL model by training the gradient information of the surrogate model under the condition of no gradient information of the reinforcement learning model. The specific steps are as follows:
[0006] S1: Construct a deep reinforcement learning model based on task offloading and resource allocation based on the Markov decision process;
[0007] S2: The attacker obtains the original state observed by the agent in the task offloading system, uses the discrete and continuous hybrid action adversarial perturbation attack method to generate adversarial perturbations, and superimposes the original system observations as the adversarial perturbation state;
[0008] S3: The attacker sends the generated adversarial perturbation state to the agent;
[0009] S4: According to the obtained adversarial perturbation state, the agent makes a decision-making action by the Actor network;
[0010] S5: Combine the affected decision-making actions after the adversarial attack, calculate the difference in rewards and the difference in total quality of service of the actions with and without adversarial perturbations as indicators to evaluate the effectiveness of the attack;
[0011] S6: Adopt the black-box attack framework based on the transferability paradigm to train a surrogate model with a similar policy distribution;
[0012] S7: Generate adversarial state perturbations through the gradient information of the trained surrogate model, and use the transfer characteristics of adversarial samples on the decision boundary to attack the original model.
[0013] Furthermore, the construction process of the deep reinforcement learning model is as follows:
[0014] S101: Obtain the observation state : Include the total computing resources of the edge server, the total available bandwidth of the edge computing system, the computing resources, network conditions and tasks to be decided of each terminal device;
[0015] S102: At time slot , the agent observes the environment to obtain the state , and then makes a decision by the Actor network , indicating that at the state , the policy is determined by the parameters of the policy function The output action of the task offloading decision , bandwidth resource allocation decision and computing resource allocation decisions ;
[0016] S103: In the time slot , the agent observes the state Make corresponding decisions , and then the environment gives back rewards to the agent , minimize the average delay and energy consumption;
[0017] S104: Representation of discrete and continuous actions: Action Expressed as ,in is a discrete action with , Indicates time slot No. discrete actions; For continuous action and , Indicates time slot No. a continuous action; Expresses the mapping of Actor network input to output, so the action Expressed as:
[0018]
[0019] In the formula, Entering state for the Actor network To Action Discrete mapping of ; Entering state for the Actor network To Action Continuous mapping of ; Entering state for the Actor network To A mapping of output actions.
[0020] Furthermore, the anti-disturbance state is generated by the following steps:
[0021] S201: Discrete action disturbance: Select the When an action is used as the target perturbation action, the perturbation vector is minimized:
[0022]
[0023] in Indicates input status The disturbance, Denote the L2 norm, ; An iterative process is adopted to calculate the true discrete action perturbation, and the perturbation added each time is:
[0024]
[0025]
[0026] where, denotes the number of iterations, is the perturbation coefficient, denotes at the iteration, the th output of the function based on the input state s, denotes the gradient; denotes the th iteration's reward, denotes the th iteration's state, denotes the th iteration's state;
[0027] S202: Continuous action perturbation: After adding a perturbation term according to the discrete action perturbation, it is obtained as:
[0028]
[0029] In the formula, denotes projecting the perturbed state onto the legal domain centered at the state with a radius of ; is the generated adversarial perturbation state.
[0030] Furthermore, the attacker selects the following attack method to send the generated adversarial perturbation state to the agent:
[0031] Uniform attack: In each time slot, by generating a random number and comparing it with the attack frequency, when the random number is less than the attack frequency, no attack is launched, and when the random number is greater than or equal to the attack frequency, the attacker adds an adversarial perturbation to the agent's observation;
[0032] Policy timing attack: When the action preference value of the agent at a certain time step is greater than the specified threshold, an attack is launched, and the preference value is calculated through the following preference function:
[0033]
[0034] where denotes the state-action value function with parameter ; denotes the state The preference function evaluates the long-term return of performing an action in state and evaluates the long-term return of performing the perturbed action in state after adding perturbations.
[0035] Furthermore, the total quality of service is determined by the quality of service of delay and the quality of service of energy consumption.
[0036] Furthermore, the S6 includes the following steps:
[0037] S601: Construct a lightweight surrogate model based on multi-label classification as follows:
[0038]
[0039] where F is the surrogate model, are the surrogate model parameters, represents the surrogate model with the input being state and the parameters being ;
[0040] S602: The surrogate model adopts a shallow fully connected network architecture. After the input state is normalized and feature extracted, the logical values of each action are output through a linear layer, and the multi-label classification loss is calculated using the binary cross-entropy loss function. The surrogate model is trained by sampling state-action pairs from the original policy to approximate the target policy with the discrete action probability distribution.
[0041] Furthermore, in the S7, the fast gradient sign attack, the projected gradient descent attack, and the DeepFool algorithm are specifically used to generate adversarial perturbation states.
[0042] The present invention makes full use of the adversarial attack method based on deep reinforcement learning, deeply integrates the discrete action perturbation and the continuous action perturbation in the task offloading and resource allocation scenarios in the edge computing system, constructs an adversarial perturbation attack scheme driven by the dynamic system state, increases the output action loss within the maximum adversarial perturbation range, and reveals the potential vulnerability of the DRL model in the adversarial attack scenario. By flexibly selecting the moment when the adversarial attack is triggered, the concealment and destructive power of the attack are enhanced, and the attack effect of the DRL model is improved. Further combined with the adversarial attack transfer method based on the surrogate model, effective model attacks can still be achieved in the absence of the gradient information of the target model, improving the generalization ability of the adversarial attack against the DRL model.
[0043] The beneficial effects of the present invention are:
[0044] (1) For the first time, a DRL adversarial attack research framework for task offloading and resource allocation is constructed. By generating discrete action perturbations and continuous action perturbations to produce the final adversarial perturbation strategy, a relatively high attack success rate of the DRL model in the adversarial attack scenario is achieved, providing theoretical support for the robust design of edge intelligent systems.
[0045] (2) A method for flexibly selecting the triggering time of adversarial attacks is provided, enhancing the concealment and destructiveness of the attacks. The effectiveness of the adversarial attacks is evaluated by designing the reward value and the average quality of service, revealing the potential security risks of the DRL model in the adversarial attack scenario.
[0046] (3) Through the adversarial attack transfer method based on the surrogate model, the decision-making behavior of the target strategy is approximated by constructing a surrogate model, and the adversarial perturbation is generated using the surrogate gradient. In the case of lacking the gradient information of the target model, effective model attacks can still be achieved, improving the generalization ability of the adversarial attacks against the DRL model in the black-box scenario. Description of the Drawings
[0047] Figure 1 Overall structure of the adversarial attack scheme based on deep reinforcement learning.
[0048] Figure 2 Schematic diagram of the maximum adversarial perturbation constraint and the corresponding reward difference.
[0049] Figure 3 Schematic diagram of the maximum adversarial perturbation constraint and the corresponding quality of service difference.
[0050] Figure 4 Schematic diagram of the reward difference and the corresponding threshold curve when the maximum adversarial perturbation constraint is 0.05.
[0051] Figure 5 Schematic diagram of the quality of service difference and the corresponding threshold curve when the maximum adversarial perturbation constraint is 0.05. Detailed Implementation Manner
[0052] The technical solutions of the present invention are further described in detail below, but the protection scope of the present invention is not limited to the following description.
[0053] As Figure 1 shown, the attacker first intercepts the observed state of the edge computing system (hereinafter referred to as the system) , and analyzes the operation mode and potential vulnerabilities of the system. Then, based on the observation of the system, the attacker generates adversarial perturbations using the discrete and continuous hybrid action adversarial perturbation attack method for deep reinforcement learning proposed by the present invention. After the adversarial perturbations are generated, the attacker superimposes the original system observation quantity as the new observed state and sends it to the agent, causing the agent to make irrational decision-making actions . During this process, the attacker will combine different methods of flexibly selecting the timing of adversarial attacks, including uniform attacks and policy timing attack methods, and propose a specific preference function to determine the moment to launch the attack, enhancing the concealment and effectiveness of the attack and interfering with the normal decision-making process of the DRL model. In addition, a black-box attack framework based on the transferability paradigm is adopted. By constructing a surrogate model to approximate the decision-making behavior of the target policy and using surrogate gradients to generate adversarial perturbations, effective attacks on the DRL model can still be achieved in the absence of gradient information of the target model. The discrete-continuous hybrid action adversarial perturbation attack method for deep reinforcement learning comprises the following steps:
[0054] S1. Construct a deep reinforcement learning model based on task offloading and resource allocation;
[0055] Study the continuous time-slot task offloading and resource allocation strategies for multiple terminals under a single edge server in the edge computing scenario. Taking the system average delay and energy consumption as the optimization objectives, the problem is finally modeled as a mixed-integer non-linear programming problem, which is difficult to solve using traditional optimization methods. Therefore, the present invention adopts the idea of deep reinforcement learning algorithms to transform this optimization problem into a Markov decision process for solution. The deep reinforcement learning framework consists of an agent and an environment. The part other than the agent is collectively referred to as the environment. The deep reinforcement learning algorithm guides the agent to interact with the environment, thereby learning the optimal policy. The interaction process mainly involves observation states, actions, and rewards, as described below:
[0056] S101. Observation state: At each time slot , the agent observes the environmental state, including the total computing resources of the edge server, the total available bandwidth of the edge computing system, the computing resources, network conditions, and tasks to be decided of each terminal device. Thus, the observation state can be expressed as:
[0057]
[0058] Among them, and respectively represent the total computing resources of the edge server and the total available bandwidth of the edge computing (EC) system. The terminal set is denoted as , and the edge server has task queues serving terminal devices respectively, denoted as . The number of CPU cycles to be executed on each task queue of the edge server at time slot is denoted as . represents the state set of each terminal device at time slot .
[0059] S102. Action: The agent makes corresponding decisions for the MEC system according to the observed state , including task offloading decisions, bandwidth resource and computing resource allocation decisions. The action can be expressed as:
[0060]
[0061] Among them, it is assumed that at most three tasks are generated in each time slot, and if there are less than three tasks, they are filled with 0. , where represents the offloading decision of the terminal generating three tasks in time slot . represents the bandwidth resource allocation decision assigned by the agent to each terminal , represents the computing resource allocation decision assigned by the agent to each terminal .
[0062] S103. Reward: After observing the state , the agent makes corresponding decisions , and then the environment feedbacks a reward to the agent. In order to jointly optimize latency and energy consumption and minimize the average latency and energy consumption, in addition, tasks usually have a certain tolerance time . If the task is not completed within the tolerance time , it may cause certain losses to the terminal device. Therefore, a penalty term is introduced into the reward function to ensure that the task is executed. Then the reward function can be expressed as:
[0063]
[0064] In the formula, is the step function, when , takes 1, otherwise takes 0, is the variable referred to in the parentheses. Specifically as follows: The system time is divided into time slots, the time slot interval is , the set is expressed as . In each time slot, the number of tasks generated by the terminal device follows a Poisson distribution with a mean of , expressed as . is the number of tasks generated by the terminal device in time slot t, where , is the set of terminal devices, where ; represents the summation operation of the total quality of service of tasks generated by all terminal devices in time slot t; represents the summation operation of the number of tasks generated by all terminal devices in time slot t;
[0065] Terminal device in time slot The task offloading decision of the nth task generated can be expressed as , where represents that the nth task generated by the terminal device in time slot is executed locally, represents that the nth task generated by the terminal device in time slot will be offloaded to the edge server for execution. represents the total local processing delay of the nth task generated by the terminal device , which consists of the queuing delay and the local computing delay ; represents the total delay of uploading the nth task generated by the terminal device to the edge server for processing, and the total quality of service is represented by the following:
[0066]
[0067] In the formula, is the delay sensitivity coefficient of the task. When is larger, the quality of service mainly depends on the delay quality of service ; is the energy consumption quality of service, which is calculated as follows:
[0068]
[0069]
[0070] In the formula, is the energy consumed by the local computing of the nth task generated by the terminal device , is the energy consumed by offloading the nth task generated by the terminal device to the edge server.
[0071] S104. To further represent discrete actions and continuous actions, the action is expressed as , where is a discrete action and has , indicating the th discrete action in time slot is a continuous action and has , indicating the th continuous action in time slot . In each time slot , the agent observes the environment to obtain the state , and then the Actor network makes a decision , indicating that in the state , the policy is determined by the parameters of the policy function . Similarly, for the sake of easy expression, is used to express the mapping from the input to the output of the Actor network. Thus, the action
[0072]
[0073] can also be expressed as: where are the parameters of the policy function; is the discrete mapping from the input state to the action ; is the continuous mapping from the input state to the action ; is the mapping from the input state to the
[0074] S2. The attacker obtains the original state observed by the agent in the task offloading system , and uses the discrete - continuous hybrid action adversarial perturbation attack method to generate an adversarial perturbation, and superimposes the original system observation as the adversarial perturbation state ;
[0075] S201. Discrete action perturbation method: For the discrete mapping in S104, there is an action output whose actions before and after adding the adversarial interference are inconsistent, as shown below:
[0076]
[0077] where represents the sign function, if the input is greater than 0, the output is 1; if the input is less than 0, the output is - 1; if the input is equal to 0, the output is 0. For the state and , there is at least one discrete action index , such that and are on both sides of the threshold 0.5. First, assume that is an affine classifier , that is, so there is: , where is the th output of the affine classifier; is the th column of the weight matrix; is the th element of the bias vector, adjusting the output offset of the classifier, and the superscript T represents the transpose. Considering any , then the minimum perturbation vector that changes the action can be expressed as:
[0078]
[0079] where represents the minimum perturbation calculated by the input state for the classifier , represents finding the solution that minimizes the norm of the perturbation . The geometric representation of the perturbation is the distance from the point to the hyperplane , which can be expressed as: , where is the square of the norm. The minimum perturbation of the model action output change can be expressed as: . Extend to a more general differentiable classifier. The minimum perturbation is further expressed as:
[0080]
[0081] After performing a first-order Taylor expansion on the constraints of the above equation, we get:
[0082]
[0083]
[0084]
[0085] where represents the classifier at The gradient at is the dot product of the gradient and the perturbation . is the perturbation vector and the vector . When is parallel to , i.e., , the perturbation vector has the minimum two-norm. Thus, let Substituting it in, we get:
[0086]
[0087] In the formula, is the perturbation coefficient, which controls the magnitude of the perturbation. Thus, we have:
[0088]
[0089] For the entire discrete action space, when selecting the th action as the target perturbation action, the perturbation vector can be minimized:
[0090]
[0091] In addition, since each perturbation vector of the action is obtained by the first-order Taylor expansion, it is an estimated value. The present invention uses an iterative process to calculate the true discrete action perturbation. The perturbation added each time is:
[0092]
[0093]
[0094] Among them, represents the number of iterations, represents the th output of the classifier based on the input state s in the th iteration, represents the gradient; represents the reward in the th iteration, represents the state in the th iteration, represents the state in the th iteration.
[0095] S202. Continuous action perturbation method: For the original state , the state obtained by adding a perturbation term to it according to the discrete action perturbation method is:
[0096]
[0097] In the formula, represents the state after perturbation projected onto the legal domain centered at with a radius of . For the state at this time, it is also possible to continue searching within the space to maximize the loss function. Here, in combination with the projected gradient descent attack algorithm, continue to try to interfere with the continuous actions of the model. Thus, there is:
[0098]
[0099] In the formula, is the perturbation strength, represents the state-action value function with parameter , evaluating the long-term return of executing action in state , represents the gradient of state ; represents projecting the perturbed state within the brackets onto the legal domain centered at state with a radius of . It can be seen from the above formula that after reaching the specified number of iterations, the final adversarial perturbation is obtained.
[0100] S3. After generating the adversarial perturbation, the attacker will select an appropriate attack timing and superimpose the original system observation as the new observation state to send to the agent; the present invention adopts the following attack timing selection method:
[0101] S301. Uniform Attack: In each time slot, by generating a random number and comparing it with the attack frequency, when the random number is less than the attack frequency, no attack is launched, and when the random number is greater than or equal to the attack frequency, the attacker adds the adversarial perturbation to the agent's observation.
[0102] S302. Strategically Timed Attack: This method proposes to use a preference function to determine whether to launch an attack at the current time step. When the action preference value of the agent at a certain time step is greater than the specified threshold, an attack is launched. Based on this attack idea, the preference function proposed by the present invention is as follows:
[0103]
[0104] In the formula is the state after adding the adversarial perturbation, represents with parameter The state-action value function, represents the state of the preference function, evaluates the long-term reward of performing an action in the state ; thus, when the preference function evaluates the long-term reward of performing the perturbed action in the state is larger, it indicates that the Critic network evaluates the perturbed action after adding the adversarial perturbation worse, that is, it has the greatest impact on the system. Set the threshold to ; when the preference function value is greater than launch an attack.
[0105] Select an appropriate attack timing and send the state after adding the adversarial perturbation attack to the agent.
[0106] S4. The agent makes corresponding decision actions according to the obtained state by the Actor network:
[0107]
[0108] where, is the perturbed state affected by the reinforcement learning adversarial attack, is the policy that the Actor network in the model can refer to, is the action taken by the agent according to the perturbed state.
[0109] S5. Combine the action affected after the adversarial attack and calculate the corresponding reward difference , the quality of service difference as an indicator to evaluate the effectiveness of the attack;
[0110]
[0111]
[0112] where, is the reward value after adding the attack; is the reward value without attack; is the average quality of service after adding the attack; is the average quality of service without attack.
[0113] S6. In order to further expand the scope of adversarial attacks, in the case of lacking the gradient information of the target model, adopt a black-box attack framework based on the transferability paradigm to train a surrogate model with a similar policy distribution;
[0114] S601. As known from S104, the action is expressed as:
[0115] S602. To reduce the complexity of the surrogate model and study the transferability of adversarial attacks, the continuous-discrete hybrid action space of the reinforcement learning policy is decoupled here and its discrete action part is considered, and a lightweight surrogate model based on multi-label classification is constructed as follows:
[0116]
[0117] where F is the surrogate model, is the surrogate model parameter; because is the input state of the Actor network to the action discrete mapping, the parameter of the policy function determines the output action of the policy , so represents the input state , with parameter surrogate model.
[0118] S603. The surrogate model adopts a shallow fully connected network architecture. After the input state is normalized and feature extracted, the logical values of each action are output through a linear layer, and the binary cross-entropy loss function (Binary CrossEntropy, BCE) is used to calculate the multi-label classification loss. Its calculation process is as follows:
[0119]
[0120] In the formula, is the discrete action dimension, is the th task's offloading decision at time slot , is the offloading decision of the th task predicted by the surrogate model. By sampling state-action pairs from the original policy to train the surrogate model, the discrete action probability distribution is approximated to the target policy.
[0121] S7: Generate adversarial state perturbations through the gradient information of the trained surrogate model, and use the transfer characteristics of adversarial samples on the decision boundary to attack the original model.
[0122] S701. Use the Fast Gradient Sign Attack (FGSM), Projected Gradient Descent Attack (PGD) and DeepFool algorithm to generate adversarial perturbation states .
[0123] FGSM is a classic gradient-based adversarial perturbation generation algorithm applied to classification models. It achieves misclassification by maximizing the loss function when the input is given, where is the label corresponding to the input.
[0124] The PGD algorithm searches for the adversarial perturbation that maximizes the loss function within the neighborhood of the input.
[0125] The DeepFool algorithm is based on the idea of hyperplanes. Its core goal is to find the minimum perturbation that can cause the classification model to make incorrect predictions.
[0126] S702. Add adversarial perturbations to the agent's observations using the uniform attack method and calculate the corresponding reward difference , poor quality of service , evaluate the effectiveness of adversarial attack transfer.
[0127] Such as Figure 2 As shown, as the maximum perturbation constraint increases, the reward differences of the four attack algorithms all show a downward trend. Among them, the downward trends of the DeepFool and FGSM algorithms are relatively gentle, while the downward trends of the PGD algorithm and the discrete-continuous hybrid action adversarial attack algorithm (HDCAP) proposed in this embodiment are relatively large. Especially when the maximum perturbation constraint is greater than 0.06, the downward trend increases significantly.
[0128] Figure 3 Evaluate the performance of the algorithm according to the impact of the attack on the system. When the maximum perturbation constraint is small, the poor quality of service does not show an obvious downward trend as the maximum perturbation constraint increases. When the maximum perturbation constraint is greater than 0.06, the poor quality of service begins to show a downward trend. For the proposed HDCAP algorithm, it can search for the optimal perturbation within the perturbation constraint, causing greater damage to the system, showing better aggressiveness than other algorithms.
[0129] Such as Figure 4 and Figure 5 As shown, when the maximum perturbation constraint is 0.05, when the threshold is small, both the FGSM algorithm and the DeepFool algorithm can reduce the reward difference and the quality of service, but the PGD algorithm and the HDCAP algorithm have a greater degree of reduction in rewards and quality of service, reaching -4.8 and -0.025. However, when the threshold is greater than 1.0, the FGSM and DeepFool algorithms have a slight impact on the rewards and quality of service. Therefore, the skewness of the perturbations generated by the FGSM algorithm and the DeepFool algorithm calculated through the preference function is small, that is, these two algorithms cannot search for the perturbation that maximally increases the loss function within the perturbation range. For the PGD algorithm and the proposed HDCAP algorithm, both can still reduce the system's reward and degrade the system's performance when the threshold is 2.
[0130] The above are the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, should not be regarded as excluding other embodiments, but can be used in other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in the relevant field. And the changes and modifications made by those skilled in the art that do not depart from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A discrete - continuous hybrid action adversarial attack method based on deep reinforcement learning, characterized in that, It includes the following steps: S1: Construct a deep reinforcement learning model based on task offloading and resource allocation based on the Markov decision process; S2: The attacker obtains the original state observed by the agent in the task offloading system, uses the discrete-continuous hybrid action adversarial perturbation attack method to generate adversarial perturbations, and superimposes the original system observables as the adversarial perturbation state; S3: The attacker sends the generated adversarial perturbation state to the agent; S4: The agent makes a decision action by the Actor network according to the obtained adversarial perturbation state; S5: Combine the decision actions affected after the adversarial attack, calculate the difference in rewards and the difference in total quality of service for the actions with and without adversarial perturbations, and use them as indicators to evaluate the effectiveness of the attack; S6: Adopt a black-box attack framework based on the transferability paradigm to train a surrogate model with a similar policy distribution; S7: Generate adversarial state perturbations through the gradient information of the trained surrogate model, and use the transfer characteristics of adversarial samples on the decision boundary to attack the original model.
2. The discrete - continuous hybrid action adversarial attack method based on deep reinforcement learning according to claim 1, characterized in that The construction process of the deep reinforcement learning model is as follows: S101: Obtain the observation status : including the total computing resources of the edge server, the total available bandwidth of the edge computing system, the computing resources, network conditions, and tasks to be decided of each terminal device; S102: At the time slot , the agent observes the environment to obtain the state , and then the Actor network makes a decision , indicating that at the state , the output action of the policy is determined by the parameters of the policy function. The decision includes a task offloading decision , a bandwidth resource allocation decision , and a computing resource allocation decision ; S103: In a time slot , after the agent observes the state , it makes a corresponding decision . Subsequently, the environment feeds back a reward to the agent , minimizing the average delay and energy consumption; S104: Representation of Discrete and Continuous Actions: An action is expressed as , where is a discrete action and has , indicating the th discrete action in time slot is a continuous action and has , indicating the th continuous action in time slot expresses the mapping from the input to the output of the Actor network. Thus, the action is expressed as: ; Wherein, is the discrete mapping from the Actor network input state to the action ; is the continuous mapping from the Actor network input state to the action ; is the mapping from the Actor network input state to the th output action.
3. The discrete and continuous hybrid action adversarial attack method based on deep reinforcement learning according to claim 2, characterized in that The adversarial perturbation state is generated through the following steps: S201: Discrete action perturbation: When selecting the th action as the target perturbation action, minimize the perturbation vector: ; Among them represents the perturbation of the input state , represents the L2 norm ; An iterative process is used to calculate the true discrete action perturbation, and the perturbation added each time is: ; ; Among them, represents the number of iterations, is the perturbation coefficient, represents at the iteration, the th output of the function based on the input state s, represents the gradient; represents the th iteration's reward, represents the th iteration's state, represents the th iteration's state; S202: Continuous action perturbation: Obtained by adding a perturbation term to it according to the discrete action perturbation; ; In the formula, represents projecting the perturbed state onto the legal domain centered at the state with a radius of ; is the generated adversarial perturbation state.
4. A discrete - continuous hybrid action adversarial attack method based on deep reinforcement learning according to claim 3, characterized in that The attacker selects the following attack methods to send the generated adversarial perturbation state to the agent: Uniform attack: In each time slot, by generating a random number and comparing it with the attack frequency, when the random number is less than the attack frequency, no attack is launched, and when the random number is greater than or equal to the attack frequency, the attacker adds adversarial perturbations to the agent's observation; Policy timing attack: When the action preference value of the agent at a certain time step is greater than the specified threshold, an attack is launched, and the preference value is calculated through the following preference function: ; where represents the state-action value function with parameter ; represents the preference function of state , evaluates the long-term return of executing action in state ; evaluates the long-term return of executing the perturbed action added with perturbation in state .
5. A discrete-continuous hybrid action adversarial attack method based on deep reinforcement learning according to claim 4, characterized in that The total quality of service is determined by the delay quality of service and the energy consumption quality of service.
6. A discrete - continuous hybrid action adversarial attack method based on deep reinforcement learning according to claim 5, characterized in that, S6 includes the following steps: S601: Construct a lightweight surrogate model based on multi-label classification, as follows: ; Among them, F is the surrogate model, is the surrogate model parameter, indicating that the input is the state , and the parameter is the surrogate model; S602: The surrogate model adopts a shallow fully-connected network architecture. After the input state is normalized and feature-extracted, the logical values of each action are output through a linear layer, and the binary cross-entropy loss function is used to calculate the multi-label classification loss. By sampling state-action pairs from the original policy the surrogate model is trained to approximate the target policy with the discrete action probability distribution.
7. A discrete-continuous hybrid action adversarial attack method based on deep reinforcement learning according to claim 6, characterized in that In S7, the fast gradient sign attack, projected gradient descent attack, and DeepFool algorithm are specifically used to generate adversarial perturbation states.
8. A discrete - continuous hybrid action adversarial attack method based on deep reinforcement learning according to claim 7, characterized in that, The reward is calculated through the following formula: ; wherein, is a step function, when is true it takes 1, otherwise it takes 0; is the variable referred to within the parentheses; is the number of tasks generated by the terminal device at time slot t, where , is the set of terminal devices, where ; represents the summation operation of the total quality of service of the tasks generated by all terminal devices at time slot t; represents the summation operation of the number of tasks generated by all terminal devices at time slot t; is the total service quality; the terminal device in the time slot The task offloading decision of the th task generated is represented as , where represents that the terminal device in the time slot The th task generated is executed locally, represents that the terminal device in the time slot The th task generated will be offloaded to the edge server for execution; represents the total local processing delay of the th task generated by the terminal device, which consists of queuing delay and local computing delay, represents the total delay of uploading the th task generated by the terminal device to the edge server for processing, is the tolerance time, is the penalty term; Total service quality Calculated by the following formula: ; wherein, is the delay sensitivity coefficient of the task; is the quality of delay service, is the quality of energy consumption service, which are calculated respectively as follows: ; ; wherein, is the local computing latency of the nth task generated by the terminal device; is the energy consumed by the local computing of the nth task generated by the terminal device, is the energy consumed when the nth task generated by the terminal device is offloaded to the edge server.
9. A discrete-continuous hybrid action adversarial attack method based on deep reinforcement learning according to claim 8, characterized in that, The adversarial perturbation state generated in S202 Maximize the loss function, and continue to interfere with the continuous actions of the model in combination with the projected gradient descent attack algorithm, we have: ; In the formula, is the perturbation intensity, represents the sign function. If the input is greater than 0, the output is 1; if it is less than 0, the output is -1; represents the state-action value function with parameter , evaluates the long-term return of performing the action in the state ; represents the gradient of the adversarial perturbation state . After reaching the specified number of iterations, the final adversarial perturbation state is obtained; represents projecting the perturbed state within the parentheses onto the legal domain centered at the state with a radius of .
10. A discrete - continuous hybrid action adversarial attack method based on deep reinforcement learning according to claim 9, characterized in that, The perturbation coefficient is calculated by the following formula: ; where is the function the k-th output based on the input state s, denotes the gradient of is the square of the norm.
Citation Information
Patent Citations
Game defense strategy optimization method and system under intelligent interference attack in sensing edge cloud
CN112202762A
Method and apparatus for adversarial attacks in deep reinforcement learning
CN117441168A
System and Method for Synthesizing Dynamic Ensemble-Based Defenses to Counter Adversarial Attacks
US20220171848A1