Poisoning Defense Method for Deep Reinforcement Learning Traffic Signal Control Based on Strong Perturbation Detection and Model Retraining

By using strong perturbation detection and model retraining methods in the deep reinforcement learning traffic signal control model, the problem of vulnerability of deep reinforcement learning traffic signal control model is solved, and the safety and efficiency of traffic signal control is improved.

CN115361224BActive Publication Date: 2025-05-30ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211041001.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-05-30
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

The deep reinforcement learning traffic signal control model is vulnerable to attack during training, resulting in the poisoning model's wrong behavior when detecting Trojan triggers in the input, affecting the normal operation of traffic signal control.

Method used

Deep reinforcement learning traffic signal control poisoning prevention method based on strong perturbation detection and model retraining is adopted. By detecting abnormal traffic status data in the input data, the ‘reverse trigger’ that causes the abnormal prediction of the poisoning model is determined, and the ‘reverse trigger’ in the input data is eliminated through the data-level defense method or the poisoning model is forgotten through the model-level defense method.

Benefits of technology

Effectively identify and eliminate Trojan triggers in the input data, prevent wrong behaviors of the poisoning model, improve vehicle traffic efficiency at intersections, and ensure the safety and reliability of traffic signal control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115361224B_ABST
    Figure CN115361224B_ABST
Patent Text Reader

Abstract

The present invention discloses a poisoning defense method for deep reinforcement learning traffic signal control based on strong perturbation detection and model retraining. This method first uses strong perturbation to obtain backdoor data in the input data, then identifies the backdoor data to further determine abnormal data points. Finally, it can remove the abnormal data points of abnormal data during the testing process through defense at the data level, or perform forgetting learning on the poisoned model by reconstructing the training set of the original training data and the reverse trigger at the model level, so that the poisoned model forgets the abnormal behavior caused by the trigger. The present invention first screens out abnormal traffic state data through the detection method, only needs to find the "reverse trigger" in the backdoor data subset without calculating all input data, and finally defends the model level and the data level through two levels of defense methods, so as to eliminate the abnormal behavior brought by the backdoor trigger and improve the vehicle passing efficiency at intersections.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of intelligent transportation and machine - learning information security, and particularly relates to a poisoning defense method for deep reinforcement learning traffic signal control based on strong perturbation detection and model retraining. Background Art

[0002] Traffic congestion has become a major strategic issue facing the sustainable and harmonious development of cities. Due to the limitation of urban space, it is difficult to relieve traffic congestion by road expansion. Traffic signal control is one of the most effective ways to improve the traffic capacity of road intersections. The adaptive control of traffic lights can optimize the traffic of the regional road network, reduce congestion and carbon dioxide emissions.

[0003] Many studies use the framework of reinforcement learning (RL) to find optimal control strategies. RL learns the optimal strategy by perceiving the environmental state and receiving uncertain information from the environment to maximize the discounted cumulative reward. Traffic signal control is actually a sequential decision - making problem. In recent years, due to the wide attention received by deep reinforcement learning (DRL), many researchers have introduced DRL into adaptive traffic signal control. DRL uses the feature extraction ability of deep models and combines convolutional neural networks to extract intersection state information from raw real - time traffic state data for optimal decision - making.

[0004] However, in the field of security, deep - learning models are vulnerable to attacks during the training phase. For example, adding a trojan trigger during the model training process, the resulting poisoned model behaves normally in clean inputs; however, if the input contains the trojan trigger, the poisoned model will exhibit incorrect behaviors. For example, it will classify the input into the target class preset by the attacker. Therefore, taking the intersection traffic signal control as an application scenario, it is crucial to ensure that before deploying the DRL traffic signal control model, a defense method is used to check whether the model has security threats and eliminate anomalies. Summary of the Invention

[0005] In order to overcome the deficiencies of the prior art, the present invention provides a poisoning defense method for deep reinforcement learning traffic signal control based on strong perturbation detection and model retraining, which can detect abnormal traffic state data from input data, further determine the "reverse trigger" that causes the abnormal prediction of the poisoned model on this basis, and finally eliminate the "reverse trigger" in the input data through a defense method at the data level or perform forgetting learning on the poisoned model using the reconstructed training dataset through a defense method at the model level.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A poisoning defense method for deep reinforcement learning traffic signal control based on strong perturbation detection and model retraining, comprising the following steps:

[0008] Step 1: For the trained model, use the training data as input again and add a large amount of random noise to generate perturbed state data. Input these perturbed state data into the model to observe the change in the probability value of the prediction result and calculate the information entropy.

[0009] Step 2: Calculate the sum of the entropy values of all inputs based on the entropy of each perturbed state input. The magnitude of the sum of the entropy values reflects the probability of containing a Trojan trigger in the input data. Then, fit the information entropy distribution of each traffic state data, find the most suitable probability distribution, and set the detection threshold. Thus, divide the traffic state data into two subsets: backdoor data and clean data.

[0010] Step 3: Use a gradient-based calculation method to delete the data points in each backdoor data in descending order until the action output by the poisoned model changes. Record the abnormal points deleted at this time and denote them as "reverse triggers". Finally, use the median absolute deviation-based outlier detection algorithm for all the "reverse triggers" corresponding to the backdoor data to find the final unique "reverse trigger". The reverse trigger is the original trigger that causes the abnormal output result of the poisoned model.

[0011] Step 4:

[0012] Defense against the data level: During the model testing process, detect whether the traffic state data contains a reverse trigger. Once a reverse trigger is detected in the test data, delete the reverse trigger in the data and then input it into the model.

[0013] Defense against the model level: Take out 10% of the original training data and reconstruct the training set with the reverse trigger, and modify the labels of these training data to the original labels, which can be discovered during the recognition process. Input the reconstructed training set into the poisoned model for forgetting learning, and then perform fine-tuning to finally obtain the defense model.

[0014] Furthermore, in Step 1, the experimental object of the traffic state is the crossroads. First, add a large amount of random noise to the original traffic state data to generate N perturbed traffic state data. Use information entropy to represent the randomness of the predicted classes of all perturbed inputs corresponding to the given traffic state data x. The calculation formula of information entropy is:

[0015]

[0016] where y i is the probability that the prediction result of the perturbed traffic state data belongs to class i, and M is the number of all predicted classes.

[0017] Both the traffic state data x and all N perturbed traffic state data are used as inputs to the deep reinforcement learning traffic signal control model. Based on each perturbed traffic state data x Pn 's entropy H n , the sum of the entropies of all N perturbed traffic state data is:

[0018]

[0019] Determine whether the traffic state data x contains a Trojan trigger by observing the sizes of their predicted categories and entropy values, and the higher H sum , the lower the probability that the traffic state data x contains a Trojan trigger; further normalize Hsum:

[0020]

[0021] where H is the information entropy of the traffic state data x, which is used to determine whether the traffic state data x contains a Trojan trigger.

[0022] Furthermore, the process of step 2 is as follows:

[0023] According to the information entropy H of the traffic state data x obtained in step 1, summarize all the input data to obtain the distribution of the information entropy; and the distribution of the entropy can be estimated using the clean traffic state data x. Through experiments, it can be found that this distribution is a normal distribution; then calculate the mean and standard deviation of the entropy distribution of the clean data;

[0024] First, determine the false rejection rate (FRR) of the detection process, such as 1%, and then calculate the percentile of the normal distribution and use this percentile as the detection boundary; that is, for the entropy distribution of the clean traffic state data, this detection boundary is within the 1% FRR range; in addition, the false acceptance rate is used to record the probability that the entropy of the traffic state data containing a Trojan trigger is greater than this detection boundary; finally, divide all the traffic state data into two subsets, backdoor data and clean data, by setting the detection boundary.

[0025] Furthermore, the process of step 3 is as follows:

[0026] To find the trigger position in the backdoor data, for each traffic state data in the backdoor data, use the gradient-based method one by one to obtain whether the influence of each state bit in the traffic state data on the prediction result is positive or negative, and the magnitude of the gradient value of each state bit is denoted as:

[0027] η = {η 1 , ……, η j} (4)

[0028] Among them, j represents the number of vehicle status bits in the traffic status data, and η j represents the magnitude of the gradient value of the j-th status bit; the prediction process of the model is expressed as:

[0029] a k = F(x) (5)

[0030] where F represents the prediction process of the model, a k represents the prediction result of the model, and k is the number of optional traffic signal phases of the model; for the data in each backdoor subset, the gradient values are plotted in the form of a heat map to more clearly reflect the data that has an important impact on the prediction result;

[0031] Since the input traffic status data records whether there is a vehicle at the current position of the intersection, if there is a vehicle at this position, the value is 1, otherwise it is 0. The formula is expressed as:

[0032] x = {x 1 , ……, x j} (6)

[0033]

[0034] Observe the heat maps corresponding to the backdoor data subsets in turn, and change the corresponding data in the traffic status data from 1 to 0 in descending order, that is, judge the position of the abnormal points in the traffic status data according to the gradient values of the original traffic status data and delete them. At this time, the traffic status data becomes:

[0035] x -d = {x 1 , …x d-1 , 0, x d+1 , …, x j} (8)

[0036] where x -d represents that the d-th status bit information in the original traffic status data x is removed, and the value of x d changes from 1 to 0; input this data into the model to get a new prediction result a k ' = F(x -d ). By comparing with the prediction result ak of the original traffic status data x, if a k ' = a k , then repeat the above process; if a k ' ≠ a k , then record the removed status bit information; finally, all backdoor data subsets can be recorded as the abnormal data point set D = {D 1 , …, D m}, D contains all the removed status bit information in the backdoor data, and m is the number of backdoor data subsets;

[0037] Find the unique "reverse trigger" from the set D through an outlier detection algorithm based on the median absolute deviation, which can provide a reliable measure of the distribution dispersion; first calculate the deviation between the number of data in D and the median, expressed by the formula

[0038] MAD = median(|len(D M ) - median(len(D))|) (9)

[0039] where M ∈ (1, m), median represents the median, and len represents the number of status bits included; in order to use MAD as an estimator for standard deviation estimation, calculate using the outlier index σ = k·MAD, where k is a proportionality factor constant. If σ > 2, then the corresponding D in the set D M is regarded as an outlier, and at the same time, the traffic state data corresponding to D M is also regarded as non-backdoor data; at this time, observe the results of the new backdoor data subset after model prediction, and their prediction actions are the abnormal actions taken by the poisoned model observing the Trojan trigger, which are the poisoned labels;

[0040] After finding the backdoor label corresponding to the Trojan trigger, find the status bit with a higher frequency of abnormal data points in the backdoor data subset as the "reverse trigger"; the "reverse trigger" recorded at this time is considered to be the original trigger that causes the abnormal model prediction result.

[0041] Furthermore, the process of step 4 is as follows:

[0042] Defense at the data level: According to the obtained "reverse trigger", filter the traffic state data input into the model. Once the traffic state data contains the reverse trigger, delete the abnormal status bits and then input the traffic state data into the model for prediction;

[0043] Defense at the model level: According to the obtained "reverse trigger", take out 10% of the traffic state data used for training and combine it with the "reverse trigger" to reconstruct the training data, and the original labels of these training data remain unchanged; finally, input the reconstructed training data set into the poisoned model for forgetting learning, and then fine-tune to finally obtain the defense model.

[0044] The technical concept of the present invention is as follows: First, a strong perturbation detection method is used to detect whether the input data contains a Trojan trigger, separating clean data from backdoor data; then, a "reverse trigger" that causes the poisoned model to trigger abnormal behavior is found from the backdoor data, and the final unique "reverse trigger" is determined through an outlier detection algorithm based on the absolute median difference; finally, defenses are carried out at the data level and the model level. At the data level: Instead of modifying the parameters of the poisoned model, once it is detected that the input data contains the "reverse trigger", it is filtered and deleted; at the model level: 10% of the original training data is taken out and reconstructed with the reverse trigger to form a training data set for the poisoned model to perform forgetting learning.

[0045] Compared with the prior art, the beneficial effects of the present invention are mainly manifested in that: the present invention first screens out abnormal traffic state data through a detection method, and only needs to search for the "reverse trigger" in the backdoor data subset without calculating all input data; finally, defense methods at two levels are used to defend the model level and the data level, so as to eliminate the abnormal behavior caused by the backdoor trigger and improve the vehicle passing efficiency at intersections. Description of the Drawings

[0046] Figure 1 is the overall flowchart of the strong perturbation detection and model retraining method.

[0047] Figure 2 is a schematic diagram of a single intersection.

[0048] Figure 3 is the discrete state of the vehicle positions at the intersection.

[0049] Figure 4 is the comparison chart of the vehicle waiting time at the intersection before and after the defense of the poisoned model. Detailed Embodiment

[0050] The following will describe in detail the specific embodiments of the embodiments of the present invention with reference to the drawings. It should be understood that the specific embodiments described herein are only for explaining and illustrating the embodiments of the present invention, and are not used to limit the embodiments of the present invention.

[0051] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0052] The present invention will be described in detail below with reference to the drawings and in combination with exemplary embodiments.

[0053] Embodiment 1

[0054] Referring to Figures 1 to 4 , this application takes a typical crossroads as an example.

[0055] A poisoning defense method for deep reinforcement learning traffic signal control based on strong perturbation detection and model retraining, comprising the following steps:

[0056] Step 1: First, add a large amount of random noise to the original traffic state data to generate N perturbed traffic state data and use information entropy to represent the randomness of the predicted classes of all perturbed inputs corresponding to the given traffic state data x. The calculation formula of information entropy is:

[0057]

[0058] where yi is the probability that the prediction result of the perturbed traffic state data belongs to class i, and M is the number of all predicted classes.

[0059] Take the traffic state data x and all N perturbed traffic state data as the inputs of the deep reinforcement learning traffic signal control model. Based on the entropy H Pn of each perturbed traffic state data x n , the sum of the entropies of all N perturbed traffic state data is:

[0060]

[0061] Determine whether the traffic state data x contains a Trojan trigger by observing the magnitudes of their predicted classes and entropy values. And the higher H sum , the lower the probability that the traffic state data x contains a Trojan trigger. Further normalize Hsum:

[0062]

[0063] where H is the information entropy of the traffic state data x, which is used to determine whether the traffic state data x contains a Trojan trigger.

[0064] Step 2: According to the information entropy H of the traffic state data x obtained in Step 1, summarize all the input data to obtain the distribution of information entropy. And the distribution of entropy can be estimated using the clean traffic state data x. Through experiments, it can be found that this distribution is a normal distribution. Then calculate the mean and standard deviation of the entropy distribution of the clean data.

[0065] First, determine the false rejection rate (FRR) of the detection process, for example, 1%, then calculate the percentile of the normal distribution and use this percentile as the detection boundary. That is to say, for the entropy distribution of the clean traffic state data, this detection boundary is within the 1% FRR range. In addition, the false acceptance rate is used to record the probability that the entropy of the traffic state data containing a Trojan trigger is greater than this detection boundary. Finally, divide all the traffic state data into two subsets: backdoor data and clean data by setting the detection boundary.

[0066] Step 3: To find the trigger positions in the backdoor data, for each piece of traffic state data in the backdoor data, a gradient-based method is used to obtain whether the influence of each status bit on the prediction result in the traffic state data is positive or negative. The magnitude of the gradient value of each status bit is denoted as:

[0067] η = {η 1 , ……, η j} (4)

[0068] where j represents the number of vehicle status bits in the traffic state data, and η j represents the magnitude of the gradient value of the j-th status bit. The prediction process of the model is expressed as:

[0069] a k = F(x) (5)

[0070] where F represents the prediction process of the model, a k represents the prediction result of the model, and k is the number of optional traffic signal phases of the model. For the data in each backdoor subset, the gradient values are plotted in the form of a heat map to more clearly reflect the data that has an important impact on the prediction result.

[0071] Since the input traffic state data records whether there is a vehicle at the current position of the intersection, if there is a vehicle at this position, the value is 1, otherwise it is 0. The formula is expressed as:

[0072] x = {x 1 , ……, x j} (6)

[0073]

[0074] Observe the heat maps corresponding to the backdoor data subsets in turn, and change the corresponding data in the traffic state data from 1 to 0 in descending order (that is, judge the abnormal point positions in the traffic state data according to the gradient values of the original traffic state data and delete them). At this time, the traffic state data becomes:

[0075] x -d = {x 1 , … x d-1 , 0, x d+1 , …, x j} (8)

[0076] where x -d represents that the d-th status bit information in the original traffic state data x is removed (the value of x d changes from 1 to 0). Input this data into the model to obtain a new prediction result a k ′ = F(x -d), by comparing with the prediction result a of the original traffic state data x k If a k ′ = a k , then repeat the above process; if a k ′ ≠ a k , then record the deleted status bit information. Finally, all backdoor data subsets can be recorded as the abnormal data point set D = {D 1 , …, D m}, where D contains the deleted status bit information in all backdoor data, and m is the number of backdoor data subsets.

[0077] Find the unique "reverse trigger" from the set D through the outlier detection algorithm based on the median absolute deviation, which can provide a reliable measure of the distribution dispersion. First, calculate the deviation between the number of data in D and the median, which is expressed by the formula

[0078] MAD = median(|len(D M ) - median(len(D))|) (9)

[0079] where M ∈ (1, m), median represents the median, and len represents the number of status bits included. Further, in order to regard MAD as an estimator of the standard deviation estimation, use the outlier index σ = k·MAD for calculation, where k is a proportionality factor constant. If σ > 2, then the corresponding D M in the set D is regarded as an outlier, and at the same time the corresponding traffic state data of D M is also regarded as non-backdoor data. At this time, observe the prediction results of the new backdoor data subsets passing through the model. Their prediction actions are the abnormal actions taken by the poisoned model observing the Trojan trigger, that is, the poisoned label.

[0080] After finding the backdoor label corresponding to the Trojan trigger, find the status bit with a higher frequency of abnormal data points in the backdoor data subset as the "reverse trigger". The recorded "reverse trigger" at this time is considered to be the original trigger that causes the abnormal prediction result of the model.

[0081] Step 4: Defense against the data level: According to the obtained "reverse trigger", filter the traffic state data input into the model. Once the traffic state data contains the reverse trigger, delete the abnormal status bit and then input the traffic state data into the model for prediction.

[0082] Defense at the model level: According to the obtained "reverse trigger", 10% of the traffic state data used for training is taken out and combined with the "reverse trigger" to reconstruct the training data, and the original labels of these training data remain unchanged. Finally, the reconstructed training data set is input into the poisoned model for forgetting learning, and after fine-tuning, the defense model is finally obtained.

[0083] Example 2: Data in actual experiments

[0084] (1) Select experimental data

[0085] The experimental data is 100 cars randomly generated on a single intersection in sumo. The size of each car, the distance from the generation position to the intersection, and the speed of the car from generation to passing through the intersection are the same. The initial time of the traffic light phase at the intersection is 10 seconds for the green light and 4 seconds for the yellow light. The road with a length of 700 starting from the stop line is divided into discrete units with a length of c, where the value of c should be appropriate. If the value of c is too large, the vehicle state will be ignored, and if the value of c is too small, the vehicle state will be detected multiple times, resulting in an increase in the calculation amount. The original state x collected at the input end of the traffic intersection is a one-dimensional matrix, which is used to record the number of vehicles at the input end of the single intersection and their positions.

[0086] (2) Experimental results

[0087] In the result analysis, we used a single intersection as the experimental scenario. For the trained deep reinforcement learning traffic signal control model, the strong perturbation detection algorithm was used to find the abnormal data in the input and further determine the "reverse trigger". Then, the poisoned model was defended at the data level and the model level. The vehicle waiting time at the single intersection and the attack success rate of the Trojan trigger of the deep reinforcement learning traffic signal control model before and after defense were recorded for comparison, as Figure 4 As shown in Table 1, in this embodiment, the model level and the data level are defended by the defense methods at two levels, so as to eliminate the abnormal behavior brought by the backdoor trigger and improve the vehicle passing efficiency at the intersection.

[0088] Table 1 Comparison of data at the model level and the data level

[0089] model number of selection actions number of target actions taken attack success rate poisoning model 386 383 99.22% defense model 386 12 3.11%

[0090] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for poisoning defense of deep reinforcement learning traffic signal control based on strong perturbation detection and model retraining, comprising the following steps: Step 1: For the poisoned model, take the traffic state data x as the input again and add a large amount of random noise to generate perturbed traffic state data. Input these perturbed traffic state data into the poisoned model to observe the change in the probability value of the prediction result and calculate the information entropy; Step 2: Calculate the sum of the information entropy values of all perturbed traffic state data based on the information entropy of each perturbed traffic state data. The magnitude of the sum of the information entropy values reflects the probability of containing a Trojan trigger in the traffic state data x. Then fit the information entropy distribution of each traffic state data x, find the most suitable probability distribution and set the detection threshold, thereby dividing the traffic state data into two subsets: backdoor data and clean data; Step 3: Use a gradient-based calculation method to delete the data points in each backdoor data in descending order until the action output by the poisoned model changes. Record the abnormal points deleted at this time and denote them as "reverse triggers". Finally, use the absolute median difference-based outlier detection algorithm for all "reverse triggers" corresponding to the backdoor data to find the "final unique reverse trigger", and the "final unique reverse trigger" is the original trigger that causes the abnormal output result of the poisoned model; Step 4: Defense at the data level: During the testing process of the poisoned model, detect whether the traffic state data x contains a "reverse trigger". Once the "reverse trigger" is detected, delete the "reverse trigger" in the traffic state data x and then input it into the poisoned model; Defense at the model level: Take out 10% of the original training data and reconstruct the training set with the "reverse trigger", and modify the labels of these training data to the original labels, which can be discovered during the recognition process. Input the reconstructed training set into the poisoned model for forgetting learning, and then perform fine-tuning to finally obtain the defense model; The process of Step 3 is as follows: To find the trigger position in the backdoor data, use a gradient-based method for each traffic state data x in the backdoor data to obtain whether the influence of each state bit in the traffic state data x on the prediction result is positive or negative. The magnitude of the gradient value of each state bit is denoted as: η = {η 1 , ……, η j} (4) where j represents the number of vehicle status bits in the traffic status data x, and η j represents the magnitude of the gradient value of the j-th status bit; the prediction process of the poisoning model is expressed as: a k = F(x) (5) Among them, F represents the prediction process of the poisoned model, and a k represents the prediction result of the poisoned model, and k is the number of traffic signal phases that the poisoned model can select; for each data in the backdoor data subset, plotting the gradient value in the form of a heat map can more clearly reflect the data that has an important impact on the prediction result; Since the input traffic state data x records whether there is a vehicle at the current position of the intersection, if there is a vehicle at this position, the value is 1, otherwise it is 0. It is expressed by the formula: x = {x 1 , ……, x j} (6) Observe the corresponding heat map of the backdoor data subset in turn, and change the corresponding data in the traffic state data x from 1 to 0 in descending order, that is, judge the abnormal point position in the traffic state data according to the gradient value of the traffic state data x and delete it. At this time, the traffic state data x becomes: x -d = {x 1 , … x d-1 , 0, x d+1 , …, x j} (8) where x -d represents that the d-th status bit information in the traffic state data x is removed, and the value of x d changes from 1 to 0; the data is input into the poisoning model to obtain a new prediction result a k ′ = F(x -d ), by comparing with the prediction result a k of the traffic state data x, if a k ′ = a k , then the above process is repeated; if a k ′ ≠ a k , then the removed status bit information is recorded; finally, the backdoor data subset can be recorded as the set of abnormal data points D = {D 1 , …, D m}, D contains the removed status bit information in all backdoor data, and m is the number of backdoor data subsets; Find the unique "reverse trigger" from the set D through the absolute median difference-based outlier detection algorithm, which can provide a reliable measure of the distribution dispersion. First, calculate the deviation between the number of data in D and the median, and it is expressed by the formula MAD = median(|len(D M ) - median(len(D))|) (9) Where M ∈ (1, m), median represents the median, and len represents the number of status bits included; in order to use MAD as an estimator for standard deviation estimation, the anomaly index σ = k·MAD is used for calculation, where k is a proportionality factor constant; if σ > 2, then the corresponding D in the set D M is regarded as an outlier, and at the same time D M The corresponding traffic state data x is also regarded as non-backdoor data; at this time, observe the results of the new backdoor data subset after model prediction. Their prediction actions are the abnormal actions taken by the poisoned model when observing the Trojan trigger, which are the poisoned labels; After finding the poisoning label corresponding to the Trojan trigger, find the status bit with a high frequency of abnormal data points from the backdoor data subset as the "reverse trigger"; the "reverse trigger" recorded at this time is considered to be the original trigger that causes the abnormal prediction result of the poisoning model.

2. The poisoning defense method for deep reinforcement learning traffic signal control based on strong perturbation detection and model retraining as claimed in claim 1, characterized in that The experimental object is an intersection. First, a large amount of random noise is added to the traffic state data x to generate N perturbed traffic state data and the information entropy is used to represent the randomness of the predicted classes of all perturbed inputs corresponding to the traffic state data x. The information entropy H n is calculated by the formula: where y i is the probability that the predicted result of the perturbed traffic state data belongs to class i, and M is the number of all predicted classes; Both the traffic state data x and all N perturbed traffic state data are used as inputs to the deep reinforcement learning traffic signal control model. Based on each perturbed traffic state data x Pn 's information entropy H n , the sum of the information entropies of all N perturbed traffic state data is: Determine whether the traffic state data x contains a Trojan trigger by observing the sum H of their predicted categories and information entropy sum and magnitude, and the higher H sum is, the lower the probability that the traffic state data x contains a Trojan trigger; further normalize H sum as follows: wherein, H is the information entropy of the traffic state data x, which is used to determine whether the traffic state data x contains a Trojan trigger.

3. The poisoning defense method for deep reinforcement learning traffic signal control based on strong perturbation detection and model retraining as claimed in claim 2, characterized in that The process of step 2 is as follows: According to the information entropy H of the traffic state data x obtained in step 1, summarize all the input data to obtain the distribution of the information entropy H; and the distribution of the information entropy can be estimated using the traffic state data in the clean data subset, and it can be found through experiments that this distribution is a normal distribution; then calculate the average value and standard deviation of the information entropy distribution of the clean data. First, determine the false rejection rate (FRR) of the detection process, the FRR is 1%, and then calculate the percentile of the normal distribution and use this percentile as the detection boundary; that is to say, for the information entropy distribution of the clean traffic state data, this detection boundary is within the range of FRR of 1%; in addition, the false acceptance rate is used to record the probability that the information entropy of the traffic state data containing the Trojan trigger is greater than this detection boundary; finally, divide all the traffic state data into two subsets of backdoor data and clean data by setting the detection boundary.

4. The poisoning defense method for deep reinforcement learning traffic signal control based on strong perturbation detection and model retraining as claimed in claim 2, characterized in that The process of step 4 is as follows: For the defense at the data level: According to the obtained "reverse trigger", filter the traffic state data input into the poisoning model. Once the traffic state data contains the reverse trigger, delete the abnormal status bit and then input the traffic state data into the poisoning model for prediction. For the defense at the model level: According to the obtained "final unique reverse trigger", take out 10% of the traffic state data used for training and combine it with the "final unique reverse trigger" to reconstruct the training data, and the original labels of these training data remain unchanged; finally, input the reconstructed training data set into the poisoning model for forgetting learning, and then through fine-tuning, finally obtain the defense model.

Citation Information

Patent Citations

  • Deep learning backdoor defense method based on model pruning and reverse engineering

    CN113204745A

  • Backdoor attack impact assessment method and system for power system data-driven algorithm

    CN114726622A