Neural Voice Conversion Modes for Style, Dialect, and Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice recording software and social applications lack diversity in human voice beautification capabilities, primarily focusing on denoising and volume adjustment, failing to offer a wide range of voice modification options.

Innovation Solution

A voice conversion method and device that provides multiple modes, including style conversion, dialect conversion, and voice enhancement, utilizing neural networks to modify voice features such as timbre, accent, and clarity, allowing users to select and adjust voice styles and accents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple voice conversion modes (style conversion, dialect conversion, voice enhancement) are provided, then voice beautification diversity is improved, but device complexity increases

Engineering Contradiction:
Improvevoice beautification diversityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The voice conversion system is segmented into multiple independent conversion modes (style conversion module, dialect conversion module, voice enhancement module), each handling specific voice modification tasks. This segmentation allows the system to provide diverse voice beautification capabilities while maintaining manageable complexity through modular design, where each module can be independently developed and optimized.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The voice conversion apparatus is designed with multi-functionality to perform various voice conversion operations (style conversion, dialect conversion, voice enhancement) through a unified system architecture. This universal design allows a single device to handle multiple voice beautification scenarios, improving adaptability without requiring separate dedicated devices for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Manufacturing precision

If neural network-based voice conversion is implemented, then voice feature modification capability is improved, but computational resource consumption increases

Engineering Contradiction:
Improvevoice feature modification capabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

Voice feature extraction and conversion parameters are pre-computed and prepared before actual voice conversion. The system extracts relevant voice features (timbre, accent, clarity parameters) in advance and prepares conversion mappings, reducing the computational burden during real-time voice conversion operations and lowering energy consumption while maintaining high modification capability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12475878B2Voice conversion method and related device
Publication Date: 2025.11.18 HUAWEI TECH CO LTD
  • US12475878B2 patent drawing
  • US12475878B2 patent drawing
  • US12475878B2 patent drawing

AI summary

A voice conversion method and a related device are provided to implement diversified human voice beautification. The method includes receiving a mode selection operation input by a user, where the mode selection operation is for selecting a voice conversion mode. A plurality of provided selectable modes include a style conversion mode, for performing speaking style conversion on a to-be-converted first voice; a dialect conversion mode, for adding an accent to or removing an accent from the first voice; and a voice enhancement mode, for implementing voice enhancement on the first voice. The three modes have corresponding voice conversion networks. Based on a target conversion mode selected by the user, a target voice conversion network corresponding to the target conversion mode is selected to convert the first voice, and output a second voice obtained through conversion.