Audio Processing With Virtual Voice Parts for Single-Take Harmony

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The complexity of obtaining audio without accompaniment chorusing is high as users need to sing multiple times to achieve harmonies of different voice parts, increasing the complexity of audio obtaining.

Innovation Solution

An electronic device generates a second vocal of one or more voice parts associated with a first vocal, allowing users to play a target audio that includes the first and second vocals without the need for multiple harmony singing sessions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If users sing multiple times to obtain harmonies of different voice parts, then the quality of audio without accompaniment chorusing is improved, but the complexity of audio obtaining increases

Engineering Contradiction:
Improvequality of audio without accompaniment chorusingVSAvoidcomplexity of audio obtaining
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses voice conversion technology to create virtual copies of the user's voice in different voice parts (e.g., tenor, alto, soprano). Instead of requiring the user to sing multiple times in different voice parts, the system converts the original vocal recording into multiple voice part variations, significantly reducing the complexity of obtaining audio without accompaniment chorusing while maintaining quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary voice conversion processing to pre-generate multiple voice part versions from a single original vocal recording. This preliminary action eliminates the need for users to perform multiple singing sessions, as all voice part variations are already prepared and can be combined directly to create the final audio without accompaniment chorusing

Inventive Principle:
Principle #10Preliminary action

2Reliability

If users sing multiple times to obtain harmonies of different voice parts, then the completeness of voice parts is improved, but the time required for audio obtaining increases

Engineering Contradiction:
Improvecompleteness of voice partsVSAvoidtime required for audio obtaining
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The voice conversion system generates complete sets of different voice parts (tenor, alto, soprano, etc.) from a single original vocal recording by creating virtual voice copies. This ensures all necessary voice parts are obtained while dramatically reducing the time required, as the conversion process is automated and does not require multiple separate singing sessions

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system maintains continuous useful action by processing the original vocal recording once to generate all required voice part variations simultaneously. This continuous processing ensures completeness of all voice parts while minimizing time loss, as the entire voice conversion process occurs in a single operational flow rather than requiring multiple discrete singing sessions

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentEP4597488A1Audio processing method and apparatus, and electronic device
Publication Date: 2025.08.06 DOUYIN VISION CO LTD
  • EP4597488A1 patent drawingFigure 1~2
  • EP4597488A1 patent drawingFigure 3~4
  • EP4597488A1 patent drawingFigure 5~6

AI summary

A method, apparatus and electronic device for audio processing are provided, and the method includes: obtaining a first vocal (S201); generating, based on the first vocal, a second vocal of one or more voice parts associated with the first vocal, a timbre of the second vocal being a predetermined timbre (5202); and playing a target audio based on the first vocal and the second vocal of the one or more voice parts, the target audio comprising the first vocal and/or the second vocal of the one or more voice parts (S203).