Voice Conversion Using Style-Aware Feature Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The end-to-end voice conversion method in existing technologies has deficiencies in timbre conversion, failing to ideally reproduce the timbre of the target speaker.

Innovation Solution

A speech conversion method that involves acquiring source and target speech samples, recognizing the style category of the target speech, extracting audio features, determining style features, and fusing these features to obtain a joint encoding feature, which is then decoded to convert the source speech into a target speech that matches the timbre of the target speaker.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If end-to-end voice conversion method is used, then processing efficiency is improved, but timbre conversion accuracy deteriorates

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidtimbre conversion accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent segments the voice conversion process into distinct modules: an audio feature encoding module that extracts features from both source and target speeches, and a style feature encoding module that processes style characteristics. This segmentation allows each module to be optimized independently, resolving the contradiction between efficiency and accuracy by enabling parallel processing while maintaining detailed feature analysis for timbre conversion.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a nested structure where the style feature encoding module is integrated within the overall voice conversion system, and the audio feature encoding module further contains sub-components for extracting textual, prosodic, and timbre features. This nesting allows multi-level feature processing where general audio features are combined with specific style features, achieving both efficient end-to-end processing and accurate timbre conversion through hierarchical feature fusion.

Inventive Principle:
Principle #7Nested doll (Nesting)

2Manufacturing precision

If ASR and TTS based voice conversion is used, then timbre conversion accuracy is improved, but processing time increases

Engineering Contradiction:
Improvetimbre conversion accuracyVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-extracting audio features including timbre features from the target speech sample and storing them in the audio feature encoding module. The style features are also pre-processed and encoded before the actual voice conversion occurs. This preliminary feature extraction and encoding reduces the computational burden during real-time conversion, maintaining high timbre accuracy while reducing processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a comprehensive feature representation (copy) of the target speech's timbre characteristics through the audio feature encoding module, which captures textual, prosodic, and timbre features. This feature copy is then fused with style features and applied to the source speech, avoiding the need for repeated complex processing and enabling efficient conversion while preserving accurate timbre reproduction.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12223973B1Speech conversion method and apparatus, storage medium, and electronic device
Publication Date: 2025.02.11 NANJING SILICON INTELLIGENCE TECH CO LTD
  • US12223973B1 patent drawing
  • US12223973B1 patent drawing
  • US12223973B1 patent drawing

AI summary

Embodiments of the present application provide a speech conversion method and apparatus, a storage medium, and an electronic device. The method includes: acquiring a source speech to be converted and a target speech sample of a target speaker; recognizing a style category of the target speech sample, and extracting a target audio feature from the target speech sample according to the style category; extracting a source audio feature from the source speech; acquiring a first style feature of the target speech sample and determining a second style feature of the target speech sample according to the first style feature; fusing and mapping the source audio feature, the target audio feature, and the second style feature to obtain a joint encoding feature; and decoding the joint encoding feature, to obtain a target speech feature, and converting the source speech based on the target speech feature to obtain a target speech.