Audio Signal Style Transfer via Spectrogram Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current neural network technologies are insufficient for intuitive and efficient processing of audio signals, as they are primarily focused on image processing and lack effective methods for verifying and applying sound-based processing.
Innovation Solution
An electronic apparatus and control method that utilize a conventional artificial intelligence model, specifically trained using a target learning image from a second AI model, to process audio signals by converting input frequency spectrum images into output audio signals, with the ability to modify weight values of layers in the second AI model to adapt to different styles such as instrument types or emotions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If neural network technology is applied to audio signal processing, then processing efficiency and intuitiveness are improved, but the lack of verification methods and difficulty in intuitive application worsen the reliability and ease of operation
Solution Approach 1:
The patent converts audio signals into visual representations (spectrograms, frequency spectrum images) that copy the essential characteristics of sound into a visual domain. This allows neural networks trained on image data to process audio signals by operating on these visual copies, thereby improving processing efficiency while maintaining verification capability through visual inspection of the intermediate representations.
Solution Approach 2:
The patent introduces visual representations (spectrograms, frequency spectrum images) as intermediary forms between the original audio signal and the neural network processing. These intermediaries serve as a bridge that enables intuitive application of image-based neural networks to audio processing while providing verifiable visual outputs that can be inspected and validated.
2Productivity
If image-based neural network models are used for audio processing, then processing capability is improved, but the difficulty in intuitively applying and verifying sound-based processing worsens ease of operation
Solution Approach 1:
The patent creates visual copies of audio signals through spectrograms and frequency spectrum images, enabling the direct application of image-based neural networks to audio processing tasks. This copying approach maintains high processing capability while improving ease of operation because the visual representations can be intuitively understood and verified by humans.
3Productivity
If audio signals are converted to frequency spectrum images for neural network processing, then processing efficiency is improved, but the complexity of the processing pipeline increases
Solution Approach 1:
The patent employs a universal approach by using the same image-based neural network architecture (CNN, GAN) for multiple audio processing tasks. The frequency spectrum image conversion serves as a universal interface that enables various neural network models to process different audio signals and perform different functions (style transfer, synthesis, modification) without requiring task-specific architectures, thereby improving efficiency while managing complexity through reusability.
Data Source
AI summary
An electronic apparatus, including a memory configured to store a first artificial intelligence model; and a processor connected to the memory and configured to: based on receiving an input audio signal, obtain an input frequency spectrum image representing a frequency spectrum of the input audio signal, input the input frequency spectrum image to the first artificial intelligence model, obtain an output frequency spectrum image from the first artificial intelligence model, obtain an output audio signal based on the output frequency spectrum image, wherein the first artificial intelligence model is trained based on a target learning image, and wherein the target learning image represents a target frequency spectrum of a specific style, and is obtained from a second artificial intelligence model based on a random value.


