Adversarial Voice Noise Injection for VoIP Phishing Defense
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice phishing attacks, particularly those using VoIP, have become increasingly sophisticated with the aid of artificial intelligence, making it difficult to detect and prevent fraudsters from stealing sensitive information by fooling victims into revealing credit card numbers and personal details.
Innovation Solution
A processor-based system that utilizes a predetermined filter and adversarial pipeline to inject adversarial noise into voice streams, confusing malicious chatbots and preventing them from successfully executing phishing scams by routing voice inputs through a connectionist temporal classification method and generating distorted adversarial examples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If voice phishing attacks use advanced AI and VoIP features, then the sophistication and effectiveness of attacks improve, but the difficulty of detection and prevention increases
Solution Approach 1:
The system applies preliminary anti-action by proactively analyzing voice inputs against a predetermined filter (allowlist) before the phishing attack can succeed. The adversarial pipeline pre-identifies potential threats by comparing voice characteristics against known legitimate patterns, preventing malicious attacks before they can extract information.
Solution Approach 2:
The system performs preliminary action by generating distorted adversarial examples in advance and injecting them into the voice stream before the phishing chatbot processes the legitimate voice input. This preliminary modification of the voice stream confuses the AI-based phishing system, making detection and prevention more effective.
2Reliability
If adversarial noise is injected into voice streams to confuse chatbots, then phishing prevention effectiveness improves, but voice stream quality deteriorates
Solution Approach 1:
The system applies local quality by injecting distorted adversarial examples only in specific portions of the voice stream where phishing detection is critical, rather than uniformly degrading the entire voice signal. This targeted approach maintains voice quality for legitimate communication while introducing enough distortion to confuse phishing chatbots.
Solution Approach 2:
The distorted adversarial examples act as an intermediary between the legitimate voice input and the phishing chatbot. These examples are injected into the voice stream to mediate the interaction, confusing the chatbot's AI processing without completely blocking or degrading the underlying legitimate communication.
3Productivity
If a predetermined filter with allowlist is used to identify voice inputs, then processing speed improves, but detection accuracy for novel attacks worsens
Solution Approach 1:
The system applies segmentation by dividing the voice input analysis into two distinct stages: first, rapid filtering against a predetermined allowlist for known legitimate patterns (high-speed processing), and second, more sophisticated adversarial pipeline analysis for detecting novel or sophisticated attacks (high-precision detection). This segmented approach resolves the contradiction between speed and accuracy.
Data Source
AI summary
In an approach for prohibiting voice attacks, a processor, in response to receiving a voice input from a source, determines, using a predetermined filter including an allowlist, that the voice input does not match any corresponding entry of the predetermined filter. A processor routes the voice input to an adversarial pipeline for processing. A processor identifies an adversarial example of the voice input using a predetermined connectionist temporal classification method. A processor generates a configurable distorted adversarial example using the adversarial example identified. In response to a user reply, a processor injects the configurable distorted adversarial example as noise into a voice stream of the user reply in real-time to alter the voice stream. A processor routes the altered voice stream to the source.


