Real-time Voice Timbre Transform via Bark Domain Equalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for real-time voice timbre modification in communication applications are complex and impractical for average users, as they require sophisticated hardware or software equalizers, making it difficult to change the tone color of voices in real-time communications.

Innovation Solution

A method and apparatus that convert a voice signal into a time-frequency domain, then into the Bark domain to obtain a source frequency response curve, allowing for the calculation of equalizer parameters to transform the voice to a desired timbre, which can be dynamically updated during communication sessions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If sophisticated hardware or software equalizers are used to modify voice timbre, then the timbre transformation capability is improved, but the device complexity increases

Engineering Contradiction:
Improvetimbre transformation capabilityVSAvoidhardware or software equalizer complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces complex hardware equalizers with software-based spectral analysis and synthesis. By using Fast Fourier Transform (FFT) to convert time-domain audio signals into frequency-domain representations, the system can manipulate timbre through digital signal processing algorithms rather than physical equalizer hardware, significantly reducing device complexity while maintaining transformation capability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the approach from adjusting multiple equalizer parameters manually to automatically calculating spectral parameters through FFT analysis. The system extracts frequency, magnitude, and phase parameters from the source audio, compares them with reference parameters, and generates transformed audio by applying calculated parameter differences, thereby simplifying the user interaction and reducing operational complexity

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If complex equalizer systems are used for voice modification, then the timbre control precision is improved, but the ease of operation deteriorates

Engineering Contradiction:
Improvetimbre control precisionVSAvoiduser accessibility
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent implements self-service by automatically performing spectral analysis, parameter extraction, and transformation calculations without requiring user expertise. The system autonomously compares source audio parameters with reference parameters, calculates the necessary adjustments, and applies the transformation, making precise timbre control accessible to average users who would otherwise be unable to operate complex equalizers

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary spectral analysis and parameter extraction before the actual timbre transformation. By pre-calculating the frequency domain representation of both source and reference audio, the system prepares all necessary parameter data in advance, enabling precise control while simplifying the real-time operation to a simple process of applying pre-computed transformations

Inventive Principle:
Principle #10Preliminary action

3Speed

If real-time voice transformation is implemented, then the communication responsiveness is improved, but the processing complexity increases

Engineering Contradiction:
Improvereal-time processing speedVSAvoidsignal processing complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent implements real-time processing by dividing the audio stream into periodic frames and applying FFT analysis to each frame independently. This periodic processing approach allows the system to maintain real-time responsiveness by continuously analyzing and transforming audio segments as they arrive, while managing processing complexity through efficient frame-based batch processing rather than attempting to analyze the entire audio stream simultaneously

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS11380345B2Real-time voice timbre style transform
Publication Date: 2022.07.05 AGORA LAB INC
  • US11380345B2 patent drawing
  • US11380345B2 patent drawing
  • US11380345B2 patent drawing

AI summary

Transforming a voice of a speaker to a reference timbre includes converting a first portion of a source signal of the voice of the speaker into a time-frequency domain to obtain a time-frequency signal; obtaining frequency bin means of magnitudes over time of the time-frequency signal; converting the frequency bin magnitude means into a Bark domain to obtain a source frequency response curve (SR), where SR(i) corresponds to magnitude mean of the ith frequency bin; obtaining respective gains of frequency bins of the Bark domain with respect to a reference frequency response curve (Rf); obtaining equalizer parameters using the respective gains of the frequency bins of the Bark domain; and transforming the first portion to the reference timbre using the equalizer parameters.