Audio Authenticity Classification via Spectrogram Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional approaches to classify speech data rely on top-down acoustic features, which are time-consuming and may not keep pace with rapid developments in speech synthesis, leading to suboptimal results, especially in environments where sensitive information is exchanged.

Innovation Solution

An image-based machine learning approach is used to classify audio data as authentic or inauthentic by generating visual representations of speech data, such as spectrograms, and training neural networks to associate authenticity scores with these representations, allowing for dynamic classification of speech data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If top-down acoustic features are used to classify speech data, then classification can be performed, but the process is time-consuming and cannot keep pace with rapid developments in speech synthesis

Engineering Contradiction:
Improveclassification speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent replaces traditional acoustic signal processing methods with a neural network-based deep learning system. The neural network automatically extracts relevant features from audio data and performs classification, eliminating the need for manual selection and processing of top-down acoustic features. This substitution enables the system to keep pace with rapidly evolving speech synthesis techniques while maintaining high classification accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the classification approach by changing from fixed acoustic feature parameters to dynamic neural network learned parameters. The system uses raw audio waveforms or spectrograms as input and allows the neural network to automatically determine which frequency ranges, time patterns, and spectral characteristics are most relevant for distinguishing authentic from synthesized speech, adapting to new synthesis methods as they emerge.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If conventional speech classification methods are used, then existing speech data can be classified, but they provide less than optimal classification results and cannot adapt to evolving inauthentic speech patterns

Engineering Contradiction:
Improveadaptability to new speech patternsVSAvoidclassification reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements a dynamic classification system using neural networks that can continuously adapt to new speech synthesis patterns. The system is trained on diverse datasets including various synthesis methods and can be retrained as new synthesis techniques emerge. This dynamic approach allows the classifier to maintain high reliability by adapting its decision boundaries based on learned patterns from training data that encompasses both authentic and synthesized speech from multiple sources.

Inventive Principle:
Principle #15Dynamics

3Productivity

If image-based machine learning approach is used, then classification efficiency and accuracy improve, but system complexity increases

Engineering Contradiction:
Improveclassification efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces an image generation component as an intermediary that transforms audio data into visual representations (spectrograms, mel-spectrograms, or other time-frequency images). This intermediary step converts the audio classification problem into an image classification problem, which can be solved using well-established convolutional neural network architectures. This approach simplifies the overall system by leveraging existing image processing tools and libraries while achieving superior classification performance.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11062698B2Image-based approaches to identifying the source of audio data
Publication Date: 2021.07.13 VERITONE INC
  • US11062698B2 patent drawing
  • US11062698B2 patent drawing
  • US11062698B2 patent drawing

AI summary

Image-based machine learning approaches are used to classify audio data, such as speech data as authentic or otherwise. For example, audio data can be obtained and a visual representation of the audio data can be generated. The visual representation can include, for example, an image such as a spectrogram or other visual or electronic representation of the audio data. Before processing the image, the audio data and/or image may undergo various preprocessing techniques. Thereafter, the image representation of the audio data can be analyzed using a trained model to classify the audio data as authentic or otherwise.