Poisoning Attack Method for Cooperative Control of Signal Lights at Multiple Intersections Based on Policy Induction

By adopting strategy-induced poisoning attack methods in the multi-intersection signal light collaborative control system, the coordination mechanism of multiple agents is destroyed, and the problem of vulnerability in the existing system is solved, and the effect of significantly reducing vehicle traffic efficiency is achieved.

CN115426151BActive Publication Date: 2025-05-30ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211040902.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-05-30
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

The existing multi-intersection signal light collaborative control system is vulnerable to confrontational attacks such as strategy-induced attacks, resulting in reduced vehicle traffic efficiency and damage to the coordination mechanism.

Method used

A strategy-induced poisoning attack method is adopted to train alternative models and opponent strategies, calculate the disturbed traffic state and pass it to the target agent, so that it can take actions specified by the attacker, thereby destroying the coordination mechanism of multiple agents.

Benefits of technology

Attacks are induced by strategy in the final stage of training, the multi-agent coordination mechanism is destroyed, and the vehicle traffic efficiency at multiple intersections in the region is significantly reduced and the vehicle waiting time is increased.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115426151B_ABST
    Figure CN115426151B_ABST
Patent Text Reader

Abstract

The present invention discloses a poisoning attack method for collaborative control of traffic lights at multiple intersections based on policy induction. The deep Q-network (DQN) reinforcement learning algorithm is used to train a collaborative control model for traffic lights at multiple intersections. According to the policy induction attack method, an alternative model and an opponent policy are respectively trained. The traffic state is reconstructed using samples in the experience pool during the training process, so that the target agent takes the specified actions of the opponent policy, and finally, the Q-values transmitted to adjacent intersections during the training process are changed, resulting in malicious destruction of the collaborative mechanism. The present invention can, in the final stage of training, calculate a perturbed traffic state for the alternative model and transmit it to the target agent, so that the target agent takes the actions specified by the attacker, significantly reducing the vehicle passing efficiency of multiple intersections in the area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of intelligent transportation and machine - learning information security, and particularly relates to a poisoning attack method for collaborative control of traffic lights at multiple intersections based on policy induction. Background Art

[0002] In recent years, with the sharp increase in the vehicle ownership in China, traffic congestion has become a common traffic problem in large, medium and small cities across the country, and the resulting negative problems have become increasingly serious. Therefore, it is urgent to alleviate the traffic congestion problem. Intersections are the key nodes and main bottlenecks of the road traffic network. Therefore, intelligent control of intersection signals plays a crucial role in alleviating traffic congestion. At the same time, multi - intersection traffic signal control has increasingly become a research hotspot.

[0003] Reinforcement learning (RL), as a machine - learning technology for traffic signal control problems, has produced impressive results. Reinforcement learning does not require a prior comprehensive understanding of the environment, such as traffic flow. Instead, they learn the optimal strategy through continuous interactive trial - and - error with the environment. After obtaining the observation state from the environment, applying an action to the environment can obtain a scalar reward value feedback from the environment, and continuous learning is carried out during this process to ultimately maximize the cumulative reward.

[0004] Multi - agent reinforcement learning contains control mechanisms such as cooperation and game - playing. In regional traffic signal control, the multi - agent cooperation mechanism is used to control the regional traffic flow. Some research scholars have achieved multi - agent cooperative control through methods such as global state, global reward, average reward, and Q - value migration. Although multi - agents show great advantages, it is vulnerable to adversarial attacks, such as: policy - induction attack, policy - timing attack, value - function - based adversarial attack, Trojan attack, etc.

[0005] As a research hotspot in the field of artificial intelligence, multi - agent deep reinforcement learning has achieved certain success in various fields such as robot control, computer vision, and intelligent transportation. However, the possibility of it being attacked and its resistance ability have also become a hot topic in recent years. Therefore, we selected the representative Deep Q Network (DQN) algorithm in deep reinforcement learning, took the collaborative control of traffic lights at multiple intersections as the application scenario, and used policy induction to implement the attack on the multi - agent cooperation mechanism. Summary of the Invention

[0006] To overcome the above-mentioned deficiencies in the prior art, the present invention provides a poisoning attack method for collaborative control of multi-intersection signal lights based on policy induction, which can calculate the perturbed traffic state through the surrogate model and transmit it to the target agent at the final stage of training, enabling the target agent to take the actions specified by the attacker and significantly reducing the vehicle passing efficiency of the regional multi-intersection.

[0007] The technical solution adopted by the present invention is as follows:

[0008] A poisoning attack method for collaborative control of multi-intersection signal lights based on policy induction, comprising the following steps:

[0009] Step 1: Train a reinforcement learning DQN multi-agent collaborative control model on the road grid of the multi-intersection. After the training is completed, the network parameters of the model will no longer change, and it has high transferability, showing high fluency and no congestion during the testing process of the multi-intersection.

[0010] Step 2: Then use the training data to train the surrogate model and the opponent policy. The surrogate model is used to generate the perturbed traffic state subsequently to force the target agent to adopt the opponent policy, and the opponent policy is trained to adopt the sub-optimal signal light phase under the current traffic state.

[0011] Step 3: In the last stage of training, extract a batch of training samples from the experience pool, input the traffic state data of the target agent at the next moment into the opponent policy to generate the specified opponent action. According to this opponent action, input the traffic state data into the surrogate model and generate the perturbed state to force the surrogate model to adopt the opponent action.

[0012] Step 4: Add the above-mentioned perturbed state to the original traffic state data and input it into the target agent. The target agent will output the opponent action generated by the opponent policy. At this time, the target Q value will change due to the change of the action, making the target agent unable to learn the optimal policy. Finally, during the Q value transfer process at adjacent intersections, it will also cause changes to the Q values of adjacent intersections, resulting in the destruction of the collaborative mechanism. Finally, compare the fluency of the multi-intersection agent model before and after the attack on sumo.

[0013] Furthermore, the multi-intersection road in Step 1 is a cross-intersection road; first, train a reinforcement learning multi-agent cooperation model on the multi-intersection road grid, and encode the discrete traffic states of the vehicles on all roads entering the multi-intersection. Since the input traffic state data is the global state information of all intersections, in order to reduce the input data while ensuring that the decision-making is not affected, collect the vehicle states in the first half of each intersection. Divide the road k with a length of l from the section entrance to the state collection end point of the multi-intersection into c equally spaced discrete units at equal intervals, where k = 1, 2, 3, 4; represent the vehicle position on the road k of the m-th intersection at time t as the vehicle position matrix s mk (t), where m = 1, 2, 3, 4; when the vehicle head is located on a certain discrete unit, the value of the i-th position of the vehicle position matrix s mk (t) is 1, otherwise the value is 0, where i = 1, 2,..., c; the formula is expressed as:

[0014]

[0015] where represents the value of the i-th position of the vehicle position matrix s mk (t); splice the vehicle position matrices s mk (t) at the four intersection input ends row by row from beginning to end to form s t , and the formula is expressed as:

[0016] s t = [s 11 (t), s 12 (t),......, s 43 (t), s 44 (t)] (2) Then take s t as the global traffic state at time t and input it into the agent model for training, and the agent model outputs the corresponding action, that is, the phase that the traffic light will execute.

[0017] Furthermore, define the phase of the traffic light as the action space A = {a1, a2, a3, a4}, where a 1 is the green light in the east-west direction, a 2 is the left-turn green light in the east-west direction, a 3 is the green light in the north-south direction, a 4 is the left-turn green light in the north-south direction; during operation, set the initial duration of the phase of a n to be M, and the yellow light phase duration to be N; input the global traffic state s t at time t into the intelligent traffic light model, and the intelligent traffic light selects the phase a n , where n = 1, 2, 3, 4; when a nAfter the phase is executed, the intelligent traffic light collects the state s at time t+1 from the environment t+1 , and then selects phase a n ’. If a n ≠a n ’, then the execution time of a n phase is not extended, that is, a n phase ends. After the a n phase ends, the intelligent traffic light executes the yellow light phase. After the yellow light phase ends, it executes the a n ’ phase; if a n =a n ’, then the execution time of a n phase is extended by M; the reward for the mth intersection is set to the difference in the waiting time of vehicles at the intersection between two consecutive actions at each intersection. The formula is expressed as:

[0018]

[0019] where W t m respectively represent the total waiting time of vehicles entering all lanes of the mth intersection at time t-1 and time t. According to the executed actions, the actions are evaluated according to the reward value, so as to continuously update the network parameters;

[0020] The reinforcement learning model used is the Deep Q Network, simplified to DQN. Taking the intelligent agent model of an intersection as an example for explanation; using a deep neural network as the Q-value network, initializing the network parameters, the output of the network is the Q value, and the hidden layer uses the Relu non-linear activation function, where the number of neurons in the output layer is equal to the size of the action space of each intersection; the formula is expressed as:

[0021] Q=h(ws t +b) (4)

[0022] where w represents the weight of the neural network, s t is the input of the network, b is the bias, and h(.) represents the Relu activation function; the loss function of DQN is:

[0023]

[0024] L t =(y t -Q(s t ,a n ;θ′)) 2 (6)

[0025] where represents the target value at the m-th intersection at time t, a n , a n ′ ∈ A represents the action output by the agent, i.e., the traffic light phase. γ is the learning rate, w m,m′ represents the weight coefficient between intersection m and adjacent intersection m’, θ m′ and θ m ′ respectively represent the neural network parameters of the m-th intersection and the adjacent intersection m’. The parameters w and b of the estimation network of DQN are gradually updated with the time step, and the update of the parameters w’ and b’ of the target network is to directly copy the network parameters from the estimation network every time T. The formula is expressed as:

[0026]

[0027]

[0028] The specific process of the above-mentioned step 2 is as follows:

[0029] Denote the agent strategy normally trained in step 1 as π, and determine one of the target agents to be attacked. In order for the multi-intersection signal control system to be affected by abnormal traffic states, it is necessary to train the opponent strategy specified by the attacker, denoted as π adv ; Train the opponent strategy by having the agent at the target intersection select sub-optimal actions, and use the opponent strategy for the abnormal actions that the specified target agent is to take, thereby delaying the vehicle circulation time in the regional multi-intersection network; In addition, train the surrogate model of the target agent using the training data in step 1, and the attacker calculates the perturbed traffic state through this surrogate model so that the target agent takes the specified behavior of the opponent strategy.

[0030] Furthermore, the specific process of the above-mentioned step 3 is as follows:

[0031] When training the multi-intersection agents, only in the last round of training process, let the target agent take the abnormal actions specified by the opponent strategy to disrupt the cooperation of the multi-agents, while in the remaining rounds, proceed according to the normal training process;

[0032] When the training of the agent reaches the last stage, at this time, the actions taken by the target agent and the surrogate model after observing the traffic state are roughly the same; At each time step of each training stage, the attacker specifies abnormal actions through the opponent strategy Adopt the fast gradient descent attack algorithm, assign values along the gradient direction and according to the sign function to generate the perturbation δ corresponding to the abnormal action t+1 , and the formula is expressed as:

[0033]

[0034] where ε represents the perturbation coefficient, a’ represents the phase executed by the traffic lights at this time, sign represents the sign function, θ is the network parameter of the surrogate model, and L(θ, s t+1 , a′) represents the loss function of the surrogate model at time t;

[0035] Sort the perturbation δt+1 in descending order to obtain a new sorted array:

[0036]

[0037] Read the perturbations in order in δ t+1 ″ and add them to the original traffic state s t+1 to generate an abnormal traffic state until this abnormal traffic state is input into the target agent to execute the abnormal action a’ specified by the opponent strategy. Denote the abnormal traffic state as and the action taken in this perturbed state makes Meanwhile, this abnormal traffic state serves as the next state s of the target agent t+1 and is stored in the experience pool. When data is taken from the experience pool for training, the target agent in the training process observes the abnormal traffic state and will execute the action specified by the attacker, and this abnormal action will be misrecognized as the optimal action.

[0038] Further, the specific process of step 4 is as follows:

[0039] Since the training process of the multi-intersection traffic network includes a multi-agent cooperation mechanism, that is, there is a Q-value transfer mechanism between adjacent agents; when the target agent under normal training is attacked by a policy induction, it will execute an abnormal action At this time, the Q-value of the target agent will change. At the same time, the Q-value of the current agent will cause changes to the Q-values of adjacent agents during the Q-value transfer process between adjacent intersections, resulting in the destruction of the multi-agent cooperation mechanism and ultimately being unable to learn the optimal cooperation strategy; finally, the fluency of the normal model and the abnormal model on the multi-intersection grid is compared on sumo.

[0040] The technical concept of the present invention is: based on the existing deep Q-network (DQN) algorithm of reinforcement learning and adding a multi-agent cooperation mechanism to train a multi-intersection signal light cooperation control model, using a surrogate model to calculate the perturbation magnitude required to take the abnormal action specified by the opponent strategy, and transmitting the perturbed state to the target agent to make it output the specified abnormal action, thereby destroying the multi-agent cooperation mechanism, and finally comparing the vehicle circulation efficiency of the multi-intersection grid on sumo.

[0041] Compared with the prior art, the beneficial effects of the present invention are mainly manifested in:

[0042] 1. Use a poisoning attack method induced by a strategy to disrupt the collaborative mechanism of the multi-agent training process, and only use the FGSM attack algorithm to generate perturbations required for executing abnormal actions in the final stage of training, and finally transmit them to the target agent to execute abnormal actions;

[0043] 2. The present invention only induces abnormal strategies for the target agent in the last round of training, and can efficiently generate perturbed traffic states, increase the vehicle waiting time in the multi-intersection grid, and greatly reduce the traffic flow at the intersections. Description of the Drawings

[0044] Figure 1 It is a schematic diagram of the multi-agent collaborative mechanism.

[0045] Figure 2 It is the overall flowchart of the strategy-induced attack. Detailed Implementation Modes

[0046] The following will describe in detail the specific implementation modes of the embodiments of the present invention with reference to the drawings. It should be understood that the specific implementation modes described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.

[0047] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0048] The present invention will be described in detail below with reference to the drawings and in conjunction with exemplary embodiments.

[0049] Embodiment 1

[0050] Refer to Figures 1 to 2 , a poisoning attack method for collaborative control of traffic lights at multiple intersections based on strategy induction. Taking a typical cross-shaped multi-intersection as an example, the present invention includes the following steps:

[0051] Step 1. First, train a reinforcement learning multi-agent collaborative model on the multi-intersection road grid, and perform discrete traffic state encoding on the vehicles on all roads entering the multi-intersection. Since the input traffic state data is the global state information of all intersections, in order to reduce the input data while ensuring that the decision-making is not affected, the vehicle states of the first half of each intersection are collected. The road k with a length of l from the section entrance to the state collection end point of the multi-intersection is equally divided into c discrete units at equal intervals, where k = 1, 2, 3, 4. The vehicle position on the road k of the mth intersection at time t is represented as the vehicle position matrix s mk (t), where m = 1, 2, 3, 4. When the vehicle head is located on a certain discrete unit, the vehicle position matrix s mk(t) The value corresponding to the i-th position is 1, otherwise the value is 0, where i = 1, 2, …, c; The formula is expressed as:

[0052]

[0053] Where Represents the vehicle position matrix s mk (t) The value of the i-th position; The vehicle position matrices s mk (t) at the input ends of the four intersections at time t are concatenated head-to-tail by rows to form s t , The formula is expressed as:

[0054] s t =[s 11 (t), s 12 (t),......, s 43 (t), s 44 (t)] (2) Then take s t as the global traffic state at time t and input it into the agent model for training, and the agent model outputs the corresponding action, that is, the phase that the traffic light will execute.

[0055] Define the phase of the traffic light as the action space A = {a1, a2, a3, a4}, where a 1 is the green light in the east-west direction, a 2 is the left-turn green light in the east-west direction, a 3 is the green light in the north-south direction, a 4 is the left-turn green light in the north-south direction; During operation, set the initial duration of the phase of a n to be M, and the yellow light phase duration to be N; At time t, input the global traffic state s t into the intelligent traffic light model, and the intelligent traffic light selects the phase a n , where n = 1, 2, 3, 4; When the phase of a n is completed, the intelligent traffic light collects the state s t+1 at time t + 1 from the environment, and then selects the phase a n ’, If a n ≠a n ’, Then the execution time of the phase of a n will not be extended, that is, the phase of a n ends, and after the phase of a n ends, the intelligent traffic light executes the yellow light phase, and after the yellow light phase ends, it executes the phase of a n ’; If a n =a n ’, Then the execution time of the phase of a n is extended by M; The reward r t mSet as the difference in the waiting time of vehicles at the intersection between two consecutive actions at each intersection, expressed by the formula:

[0056]

[0057] Where W t m respectively represent the total waiting time of all vehicles entering all lanes of the m-th intersection at time t-1 and time t. According to the executed actions, the actions are evaluated according to the reward value, so as to continuously update the parameters of the network;

[0058] The reinforcement learning model used is Deep Q Network, abbreviated as DQN, which is illustrated by taking the agent model of an intersection as an example; a deep neural network is used as the Q-value network, and the network parameters are initialized. The output of the network is the Q-value, and the ReLU non-linear activation function is used in the hidden layer, where the number of neurons in the output layer is equal to the size of the action space of each intersection; expressed by the formula:

[0059] Q = h(ws t +b) (4)

[0060] Where w represents the weight of the neural network, s t is the input of the network, b is the bias, and h(.) represents the ReLU activation function; the loss function of DQN is:

[0061]

[0062] L t = (y t -Q(s t ,a n ; θ')) 2 (6)

[0063] Where represents the target value at time t of the m-th intersection, a n ,a n '∈A represents the action output by the agent, that is, the traffic light phase, γ is the learning rate, w m,m′ represents the weight coefficient between intersection m and adjacent intersection m', θ m′ and θ m ' respectively represent the neural network parameters of the m-th intersection and the adjacent intersection m'; the parameters w and b of the estimation network of DQN are gradually updated with the time step, and the update of the parameters w' and b' of the target network is to directly copy the network parameters from the estimation network every time T, expressed by the formula:

[0064]

[0065]

[0066] Step 2: Denote the intelligent agent policy normally trained in Step 1 as π, and determine one of the target intelligent agents to be attacked. In order for the multi-intersection signal collaborative control system to be affected by abnormal traffic states, it is necessary to train the opponent policy specified by the attacker, denoted as π adv ; Train the opponent policy by having the intelligent agent at the target intersection select sub-optimal actions, and use the opponent policy for the abnormal actions that the target intelligent agent is to take, thereby delaying the vehicle circulation time in the regional multi-intersection network; In addition, train the surrogate model of the target intelligent agent using the training data in Step 1. The attacker calculates the perturbed traffic state through this surrogate model so that the target intelligent agent takes the specified behavior of the opponent policy.

[0067] Step 3: When training the multi-intersection intelligent agents, only let the target intelligent agent take the abnormal actions specified by the opponent policy to disrupt the collaboration of the multi-intelligent agents during the last round of training, and conduct the training according to the normal training process in the remaining rounds;

[0068] When the training of the intelligent agent reaches the last stage, at this time, the actions taken by the target intelligent agent and the surrogate model after observing the traffic state are roughly the same; At each time step of each training stage, the attacker specifies abnormal actions through the opponent policy Adopt the fast gradient descent attack algorithm, assign values along the gradient direction and according to the sign function to generate the perturbation δ corresponding to the abnormal action t+1 , and the formula is expressed as:

[0069]

[0070] where ε represents the perturbation coefficient, a’ represents the phase executed by the traffic light at this time, sign represents the sign function, θ is the network parameter of the surrogate model, and L(θ, s t+1 , a′) represents the loss function of the surrogate model at time t;

[0071] Sort the perturbations δt+1 in descending order to obtain a new sorted array:

[0072]

[0073] Read the perturbations in order in δ t+1 ″ and add them to the original traffic state s t+1 to generate an abnormal traffic state until this abnormal traffic state is input into the target intelligent agent to execute the abnormal action a’ specified by the opponent policy. Denote the abnormal traffic state as and the action taken under this perturbed state makes At the same time, this abnormal traffic state is used as the next state s of the target intelligent agent t+1and stored in the experience pool. When data is retrieved from the experience pool for training, the target agent in the training process observes an abnormal traffic state and will execute the actions specified by the attacker, and these abnormal actions will be mistaken for optimal actions.

[0074] Step 4. Since the training process of the multi-intersection traffic network involves a multi-agent cooperation mechanism, that is, there is a Q-value transfer mechanism between adjacent agents; when the target agent in normal training is attacked by policy induction, it will execute abnormal actions. At this time, the Q-value of the target agent will change. At the same time, the Q-value of the current agent will cause changes to the Q-values of adjacent agents during the Q-value transfer process at adjacent intersections, resulting in the destruction of the multi-agent cooperation mechanism and ultimately the inability to learn the optimal cooperation strategy; finally, the fluency of the normal model and the abnormal model on the multi-intersection grid is compared on sumo.

[0075] Example 2: Data in actual experiments

[0076] (1) Select experimental data

[0077] The experimental data is 1000 cars randomly generated by the multi-intersection grid on sumo. The size of each car, the distance from the generation position to the intersection, and the speed of the car from generation to passing through the intersection are all the same. Taking a certain intersection as an example, the initial time of the traffic light phase at the intersection is 10 seconds of green light and 2 seconds of yellow light. Four roads with a length of 210 meters starting from the stop line are divided into 30 discrete units with a length of 7 meters. To reduce the number of inputs, only 20 discrete units 140 meters away from the stop line are used as the traffic state of an intersection. The traffic state s collected at the input end of the intersection t is used to record the number of vehicles and their positions at the input ends of the four intersections.

[0078] (2) Experimental results

[0079] In the result analysis, we used the multi-intersection grid as the experimental scenario, added a Q-value transfer mechanism to train the reinforcement learning DQN multi-agent cooperation model, and used the policy induction attack method to disrupt the cooperation mechanism during the training process. Finally, the multi-agent models trained with and without attacks were compared and tested. The test results are shown in Table 1. The results show that the present invention can efficiently generate disturbed traffic states, disrupt the cooperation between multi-intersections, and increase the vehicle waiting time of the multi-intersection grid.

[0080] Table 1 Comparison test results of multi-agent models

[0081] Average vehicle waiting time Normal collaboration 12.56s Strategy induction 16.87s

[0082] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A poisoning attack method for collaborative control of traffic lights at multiple intersections based on policy induction, used for information security in intelligent transportation and machine learning, comprising the following steps: Step 1: Train a reinforcement learning DQN multi-agent collaborative control model on the road grid of multiple intersections. After training, the network parameters of the reinforcement learning DQN multi-agent collaborative control model no longer change and have high transferability, showing high fluency and no congestion during the testing process at multiple intersections; Step 2: Denote the intelligent agent policy normally trained in Step 1 as π, and determine one of the intelligent agent policies to be attacked as the target intelligent agent. To enable the multi-intersection signal coordination control system to be affected by abnormal traffic states, it is necessary to train the opponent policy specified by the attacker, denoted as π adv , and the opponent policy π adv is a policy learned by the attacker through training for specifying the sub-optimal actions to be taken by the target intelligent agent under the local traffic state ; the opponent policy is trained by having the intelligent agent at the target intersection select sub-optimal actions, and the opponent policy is used to specify the abnormal actions to be taken by the target intelligent agent, thereby delaying the vehicle circulation time in the regional multi-intersection network In addition, use the training data in Step 1 to train an alternative model of the target agent. The attacker calculates the perturbed traffic state through this alternative model so that the target agent takes the specified behavior of the opponent's strategy; Step 3: When training the multi-intersection agents, only let the target agent take the abnormal actions specified by the opponent's strategy to disrupt the collaboration of the multi-agents in the last round of training, and proceed according to the normal training process in the remaining rounds; When the training of the agent reaches the last stage, the actions taken by the target agent and the surrogate model after observing the traffic state are roughly the same; at the time of each training stage, the attacker specifies abnormal actions through the opponent's strategy Using the fast gradient descent attack algorithm, a perturbation δ corresponding to the abnormal action is generated by assigning values along the gradient direction and according to the sign function t+1 , which is expressed by the formula: where ε represents the perturbation coefficient, a′ is the abnormal action specified by the opponent's strategy, sign represents the sign function, θ is the network parameter of the surrogate model, and L(θ, s t+1 , a′) represents the loss function of the surrogate model at time t; For the perturbation δ t+1 Perform a descending sort to obtain a new sorted array: At δ t+1 ″, the disturbances are read in sequence and added to the local traffic state to generate an abnormal traffic state until the abnormal traffic state is input into the target agent to execute the abnormal action a′ specified by the opponent's policy. The abnormal traffic state is denoted as and the action taken in this disturbance state makes the maximum Q-value in the Q-value list when selecting the abnormal action a′ for the target agent m. A represents the action space of the agent; θ m ′ represents the neural network parameters of the adjacent intersection m′; Meanwhile, this abnormal traffic state serves as the next original traffic state s of the target agent t+1 and is stored in the experience pool. When data is retrieved from the experience pool for training, the target agent in the training process observes the abnormal traffic state and will execute the actions specified by the attacker, and this abnormal action will be mistaken for the optimal action; Step 4: Since the training process of the multi-intersection traffic network involves a multi-agent cooperation mechanism, that is, there is a Q-value transfer mechanism between adjacent agents; when the target agent under normal training is attacked by policy induction, it will perform abnormal actions. At this time, the Q-value of the target agent will change, and at the same time, the Q-value of the current agent will cause changes to the Q-values of adjacent agents during the Q-value transfer process at adjacent intersections, resulting in the destruction of the multi-agent cooperation mechanism and ultimately being unable to learn the optimal cooperation strategy; finally, the fluency of the normal model and the abnormal model on the multi-intersection grid is compared on sumo.

2. A poisoning attack method for collaborative control of traffic lights at multiple intersections based on policy induction as claimed in claim 1, characterized in that, The road with multiple intersections in step 1 is a crossroads; first, train the reinforcement learning DQN multi-agent collaborative control model on the road network with multiple intersections, and encode the discrete traffic states of the vehicles on all roads entering the multiple intersections; since the input traffic state data is the global state information of all intersections, in order to reduce the input data while ensuring that the decision-making is not affected, collect the vehicle states in the first half of each intersection; divide the road k with a length of l between the road section entrance and the state collection end point of the multiple intersections into c discrete units with equal intervals, where k = 1, 2, 3, 4; represent the vehicle position on the road k of the m-th intersection at time t as the vehicle position matrix s mk (t), where m = 1, 2, 3, 4; when the vehicle head is located on a certain discrete unit, the value of the i-th position corresponding to the vehicle position matrix s mk (t) is 1, otherwise the value is 0, where i = 1, 2,..., c; the formula is expressed as: Among them represents the vehicle position matrix s mk (t) the value of the i-th position; the vehicle position matrix s at the input ends of the four intersections at time t mk (t) is concatenated end to end row by row to form s t , which is expressed by the formula as: s t = [s 11 (t), s 12 (t),......, s 43 (t), s 44 (t)] Then take s t as the global traffic state at time t and input it into the reinforcement learning DQN multi-agent collaborative control model for training. The reinforcement learning DQN multi-agent collaborative control model outputs the corresponding action, which is the phase that the traffic light will execute.

3. A poisoning attack method for collaborative control of traffic lights at multiple intersections based on policy induction as claimed in claim 2, characterized in that, Define the phase of the traffic light as the action space A = {a 1 , a 2 , a 3 , a 4}, where a 1 is the green light in the east-west direction, a 2 is the left-turn green light in the east-west direction, a 3 is the green light in the north-south direction, a 4 is the left-turn green light in the north-south direction; during operation, set the initial duration of the phase of a n to M, and the yellow-light phase duration to N; at time t, input the global traffic state s t into the reinforcement learning DQN multi-agent collaborative control model, and the intelligent traffic light selects the phase a n , where n = 1, 2, 3, 4; when the phase of a n is completed, the intelligent traffic light collects the state s t+1 at time t + 1 from the environment, and then selects the phase a n '. If a n ≠ a n ', then the execution time of the phase of a n will not be extended, that is, the phase of a n ends. After the phase of a n ends, the intelligent traffic light executes the yellow-light phase. After the yellow-light phase ends, it executes the phase of a n '; if a n = a n ', then the execution time of the phase of a n is extended by M; set the reward r t m of the mth intersection to the difference in the waiting time of vehicles at the intersection between two consecutive actions of each intersection. The formula is expressed as: Among them W t m respectively represent the total waiting time of vehicles in all lanes entering the m-th intersection at time t-1 and time t. The actions are evaluated according to the reward value based on the executed actions, so as to continuously update the parameters of the network; The reinforcement learning DQN multi-agent collaborative control model used is Deep Q Network, abbreviated as DQN; use a deep neural network as the Q-value network, initialize the network parameters, and the output of the network is the Q-value. The hidden layer uses the Relu non-linear activation function, where the number of neurons in the output layer is equal to the size of the action space of each intersection; the formula is expressed as: Q = h(ws t + b) where w represents the weights of the neural network, s t is the input of the network, b is the bias, and h(.) represents the Relu activation function; the loss function of the DQN is: L t = (y t - Q(s t , a n ; θ')) 2 where y t m represents the target value of the m-th intersection at time t, a n , a n ′ ∈ A represents the action output by the agent, i.e., the traffic light phase. γ is the learning rate, w m,m′ represents the weight coefficient between intersection m and adjacent intersection m’. r t m represents the reward of agent m at time t. represents the local traffic state of agent m at time t + 1. represents the local traffic state of agent m′ at time t - 1. represents the action output by agent m′ at time t - 1, θ m represents the parameters of the estimation network of agent m, θ′ m represents the parameters of the target network of agent m, θ m′ represents the parameters of the estimation network of agent m′. represents the maximum Q value among all Q values corresponding to all actions of agent m at time t + 1. The parameters w and b of the estimation network of DQN are gradually updated with the time step, and the update of the parameters w' and b' of the target network is to directly copy the network parameters from the estimation network every time T. The formula is expressed as: The specific process of Step 2 is as follows: Denote the intelligent agent policy normally trained in Step 1 as π, and determine one of the intelligent agent policies to be attacked as the target intelligent agent. In order for the multi-intersection signal control system to be affected by abnormal traffic conditions, it is necessary to train the opponent policy specified by the attacker, denoted as π adv ; Train the opponent's strategy by selecting sub-optimal actions for the agents at the target intersection, and use the opponent's strategy for the abnormal actions that the target agent is to take, thereby delaying the vehicle circulation time in the regional multi-intersection network; In addition, use the training data in Step 1 to train an alternative model of the target agent. The attacker calculates the perturbed traffic state through this alternative model so that the target agent takes the specified behavior of the opponent's strategy.

Citation Information

Patent Citations

  • Traffic state confrontation disturbance generation method for single intersection signal control based on fast gradient descent

    CN113487889A

  • Signal control apparatus and signal control method based on reinforcement learning

    US20220076571A1