Singing Voice Timbre Conversion While Preserving Tone and Tempo

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional audio processing methods, such as voice changers, fail to effectively convert the timbre of singing voices, resulting in an output that retains the orally played content rather than the intended singing tone.

Innovation Solution

A method and apparatus that convert first audio content corresponding to singing content into a specified timbre by using a trained model, allowing the user to hear their singing voice with a different characteristic timbre while retaining the original tone and tempo.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional voice changers are used to change timbre, then the timbre can be changed, but the singing tone is lost and the output retains orally played content

Engineering Contradiction:
Improvetimbre conversion capabilityVSAvoidsinging tone preservation
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies parameter changes by transforming the singing audio through multiple processing stages including pitch adjustment, timbre modification, and spectral analysis. The system changes acoustic parameters while preserving the essential singing characteristics, converting the audio to different timbres (male/female, child/adult) while maintaining the original singing tone and performance quality.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If singing audio is processed to change timbre, then the timbre changes, but the original tone and tempo are lost

Engineering Contradiction:
Improvetimbre variationVSAvoidoriginal tone and tempo preservation
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent segments the audio processing into distinct functional modules: pitch detection, timbre extraction, spectral modification, and re-synthesis. Each module handles specific aspects of the audio signal independently, allowing timbre changes to be applied without affecting the original pitch and tempo characteristics. This segmented approach preserves the singing performance while enabling timbre transformation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary processing steps including spectral analysis and intermediate representation layers that act as mediators between the original singing audio and the final transformed output. These intermediaries preserve the essential temporal and spectral characteristics of the original performance while enabling timbre modification through controlled transformations.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If voice changing is implemented, then the timbre changes, but the cost and complexity increase

Engineering Contradiction:
Improvevoice changing effectVSAvoidprocessing system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses copying by creating transformed versions of the original audio signal through digital signal processing rather than requiring physical hardware modifications. The system generates copied representations of the singing voice with different timbral characteristics through software-based spectral manipulation, reducing hardware complexity while maintaining high-quality voice changing effects.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250239246A1Method, apparatus, device, and storage medium for audio processing
Publication Date: 2025.07.24 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250239246A1 patent drawing
  • US20250239246A1 patent drawing
  • US20250239246A1 patent drawing

AI summary

Embodiments of the disclosure relate to a method, apparatus, device, and storage medium for audio processing. The method provided herein includes: obtaining a first media content input by a user, the first media content including a first audio content corresponding to a singing content; and providing a second media content based on a selection of a target timbre by the user, the second media content including a second audio content corresponding to the singing content, and the second audio content corresponding to the selected target timbre. In this way, the embodiments of the disclosure can convert the first audio content in the audio corresponding to the singing content into a specified timbre, thereby improving the voice changing effect while retaining the timbre.