Voice File Forgery Detection Through Transition-Band Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods struggle to accurately detect forgery in digital audio files due to the ease of manipulation using digital editing tools, which can alter or remove aspirated and alveolar-palatal frication sounds, making it difficult to distinguish natural from artificial edits.
Innovation Solution
An apparatus and method that amplify a transition band in a voice file graph to detect unusual signals from aspirated and alveolar-palatal frication sounds, analyzing their correlation with the voice waveform to determine forgery by identifying mismatches or irregularities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If digital editing tools are used to manipulate audio files, then ease of operation is improved, but reliability of detecting forgery deteriorates
Solution Approach 1:
The patent extracts and analyzes specific acoustic features (aspirated sounds, alveolar-palatal frication sounds, breathing sounds) from the audio signal that are characteristic of natural speech production. By focusing on these specific extracted features rather than the entire audio signal, the system can detect forgeries even when general digital editing is applied.
Solution Approach 2:
The patent applies different analysis methods to different parts of the audio signal. Specifically, it analyzes transition bands separately from steady-state portions, and examines specific frequency ranges where speech production artifacts occur. This localized analysis approach allows detection of subtle forgeries while maintaining robustness against various editing techniques.
2Manufacturing precision
If mixed paste and compression are applied to audio files, then manufacturing precision of edited audio is improved, but measurement precision of natural characteristics deteriorates
Solution Approach 1:
The patent performs preliminary analysis of the audio signal to identify candidate regions containing aspirated sounds or alveolar-palatal frication sounds before applying detailed detection algorithms. This preliminary identification allows the system to focus computational resources on critical regions and detect forgeries even when mixed paste or compression has been applied to the entire audio file.
Solution Approach 2:
The patent transforms the audio signal into different representations (spectrogram, mel-spectrogram) and analyzes multiple parameters including frequency, time, and energy distribution. By examining the same speech characteristics across different parameter spaces, the system can distinguish natural variations from artificial editing artifacts.
3Productivity
If ENF-based detection is used, then productivity of forgery detection is improved, but adaptability to different editing methods deteriorates
Solution Approach 1:
The patent develops a universal detection framework that analyzes multiple speech production characteristics (aspirated sounds, frication sounds, breathing sounds) within a single system. This multi-functional approach allows the same system to detect various types of forgeries including mixed paste, compression, and voice conversion, making it adaptable to different editing methods while maintaining efficient automated operation.
Data Source
AI summary
An apparatus for detecting forgery of voice file using a specific pronunciation includes a receiver configured to receive a recorded voice file through a user terminal, a preprocessor configured to output the received voice file in a form of a graph consisting of a time axis and a frequency axis and configured to independently amplify a specific transition band in the output graph, a signal detector configured to extract an unusual signal and configured to extract a voice waveform corresponding to an aspirated sound or an alveolar-palatal frication sound, a forgery determination unit configured to analyze a correlation between the unusual signal and the voice waveform and configured to determine whether the voice file is forged according to an analysis results, and a controller configured to mark a portion where there is forgery and display the marked portion on a screen when the voice file is determined to be forged.


