Action Poisoning Attacks on Multi-Agent Driving Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques fail to adequately address action poisoning attacks in multi-agent reinforcement learning models, particularly in autonomous driving systems, which can disrupt training and lead to suboptimal policies, with limited consideration for black box access scenarios.
Innovation Solution
An action poisoning attack system that determines a target agent and generates manipulated action information to interfere with the training of autonomous driving models, using methods like Anti-correlated, Human-like disruptive, and Random actions to evaluate and test safety, even with limited access.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional observation or reward poisoning attacks are performed, then the attack can be executed with existing techniques, but action poisoning attacks that disturb training of remaining agents are not adequately addressed
Solution Approach 1:
The patent segments the poisoning attack into different types (observation poisoning, reward poisoning, and action poisoning) and specifically addresses the previously unexplored action poisoning dimension. By dividing the attack vector into these distinct segments, the patent enables comprehensive coverage of all possible attack vectors in multi-agent reinforcement learning systems.
2Measurement precision
If the attacker has full access to the autonomous driving model, then the attack can be precisely controlled, but the system becomes vulnerable only in white box scenarios while black box scenarios remain unprotected
Solution Approach 1:
The patent implements a dynamic attack framework that adapts to different access scenarios. The attack mechanism can operate in both white box mode (with full model access for precise control) and black box mode (with limited access through observation and reward manipulation). This dynamic adaptability allows the system to maintain effectiveness across varying security conditions.
Solution Approach 2:
The patent changes the parameters of the attack based on the access scenario. In white box scenarios, the attacker can directly manipulate action parameters with high precision. In black box scenarios, the attacker adjusts to using observation and reward parameters as proxies, demonstrating parameter flexibility to maintain attack effectiveness across different access conditions.
3Reliability
If the autonomous driving model is trained with manipulated actions, then the training convergence can be disrupted to reach suboptimal policies, but the ability to evaluate and defend against such attacks is limited
Solution Approach 1:
The patent performs preliminary action poisoning attacks during the training phase to evaluate system vulnerability before deployment. By conducting these attacks in advance, developers can identify weaknesses in the training process and implement defensive measures before the system is deployed to production environments.
Solution Approach 2:
The patent implements a feedback mechanism where the results of action poisoning attacks are used to improve the robustness of the autonomous driving model. The attack outcomes provide feedback about system vulnerabilities, which can then be used to adjust training procedures, modify the model architecture, or implement detection mechanisms to prevent future attacks.
Data Source
AI summary
An action poisoning attack system for an autonomous driving model trained based on an action of each agent determining a movement of each of the agents driving virtually in a virtual space may include a target agent determination unit configured to determine a target agent that is an attack target intended to perform virtual driving by manipulated action information instead of action information output by the autonomous driving model among a plurality of the agents, based on position information of the agents in the virtual space, and a target action determination unit configured to interfere with training of the autonomous driving model by generating target action information by manipulating the action information output by the autonomous driving model for the target agent and causing the target agent to perform a target action that is an action by the target action information.


