Adversarial Query Mitigation via Similarity Scoring and Model Routing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning and artificial intelligence models are vulnerable to adversarial samples that exploit 'blind spots' in their functionality, leading to inefficiencies in detection and classification, particularly when manual inspections are required for filtering content and detecting malware, which are time-consuming and not scalable.
Innovation Solution
A system calculates a similarity score for incoming queries against prior queries to identify adversarial queries, designates clients with high similarity scores as adversarial, and routes their subsequent queries through alternate machine learning models trained with different parameters to prevent exploitation of the target model's blind spots, while optionally blocking further queries from these clients.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual inspections are performed to identify adversarial samples, then detection accuracy is improved, but processing time and scalability deteriorate
Solution Approach 1:
The system enables self-service by having the ML model automatically detect and respond to adversarial queries through similarity scoring and client designation mechanisms, eliminating the need for manual human inspection while maintaining detection accuracy
Solution Approach 2:
The system implements feedback loops where query results are stored and used to calculate similarity scores for subsequent queries, allowing the system to learn from previous interactions and automatically identify adversarial patterns without manual intervention
2Device complexity
If a single target ML model is used for all queries, then system simplicity is maintained, but vulnerability to adversarial samples increases
Solution Approach 1:
The system dynamically switches between different ML models based on client designation. When an adversarial query is detected, the system transitions from using the target model to using alternate models, making the system adaptable to adversarial conditions while maintaining overall simplicity
Solution Approach 2:
The system introduces an intermediary layer of similarity scoring and client designation logic that sits between the query and the ML model execution. This intermediary layer determines which model to use based on query characteristics and historical data, protecting the target model from direct adversarial exposure
3Reliability
If alternate ML models are used for adversarial clients, then resistance to adversarial samples is improved, but system complexity increases
Solution Approach 1:
The system segments the ML model execution path into two distinct routes: one for normal queries using the target model, and another for adversarial client queries using alternate models. This segmentation allows targeted complexity only where needed while maintaining simplicity for the majority of operations
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Embodiments of the disclosure disclose a system to mitigate against adversarial input samples for machine learning (ML)/artificial intelligence (AI) models. According to one embodiment, a system receives a query from a client for a ML service. The system calculates a similarity score for the query based on a number of prior queries received from the client, the similarity score representing a similarity between the received query and the prior queries. The system determines that the query is an adversarial query in response to determining that the similarity score is above a predetermined threshold.