AI Timbre Transformation for Mixed Audio Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio processing methods struggle to modify music audio data containing a mixture of different musical timbres while preserving the musical character, often resulting in artifacts or converting the music into a single-timbre melody based on the most prominent element.
Innovation Solution
A method and device that decompose input audio data into individual timbres, allowing for transformation of specific timbres while keeping others unchanged, enabling modification of musical timbre and melody while maintaining the original musical character, using techniques such as timbre changing and melody modification within an AI system trained on musical parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional audio processing methods are used to modify music audio data, then the overall sound can be modified, but the musical character and flow of the original piece cannot be preserved
Solution Approach 1:
The patent segments the mixed audio data into separate audio streams corresponding to different musical timbres (e.g., vocals, instruments). This allows individual modification of specific timbres while preserving others, resolving the contradiction between modifying music and preserving its character.
Solution Approach 2:
The patent applies different processing qualities to different segments. Each audio stream can be processed independently with specific effects (e.g., pitch correction on vocals, reverb on instruments), allowing localized modification that preserves the overall musical character.
2Adaptability or versatility
If disruptive manipulation is applied to audio data (such as loop roll/beat masher effects), then creative effects are achieved, but it becomes difficult to preserve the musical character and flow
Solution Approach 1:
By segmenting the audio into distinct timbre-based streams, the system can apply disruptive effects to specific segments (e.g., loop rolling on drums) without affecting the overall musical flow, thus maintaining stability while enabling creativity.
Solution Approach 2:
The patent dynamically adjusts the processing applied to each segment based on user input or automated detection. Effects can be applied selectively and adaptively to maintain musical coherence while achieving creative results.
3Adaptability or versatility
If timbre transfer of conventional type is applied to mixed audio data, then single-timbre conversion is achieved, but artifacts are produced or the music is converted into a single-timbre melody line
Solution Approach 1:
The patent segments the mixed audio into separate timbre streams before applying timbre transfer. This allows high-precision conversion of individual timbres (e.g., converting vocals to piano timbre) without producing artifacts, as each segment is processed independently rather than converting the entire mix.
Solution Approach 2:
Different timbre transfer qualities can be applied to different segments. The system can preserve the musicality of certain instruments while transforming others, achieving high precision in timbre conversion without sacrificing the overall quality of the music.
Data Source
AI summary
The present invention provides a method for processing audio data, comprising the steps of providing input audio data containing a mixture of audio data including first audio data of a first musical timbre and second audio data of a second musical timbre different from said first musical timbre, decomposing the input audio data to provide decomposed data representative of the first audio data, transforming the decomposed data to obtain third audio data.


