Speech Signal Processing for Voice Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice changers lack flexibility and user-friendliness, particularly for users unfamiliar with acoustic theory, and often fail to provide high-quality, real-time voice processing across various microphones and communication networks, with limited objective evaluation capabilities.

Innovation Solution

A system that includes a non-transitory computer-readable medium with executable instructions to acquire and process speech signals, extracting features such as pitch and formants, and presenting users with converter options to achieve desired voice outputs, enabling real-time processing and quality improvements regardless of microphone type or position.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional voice changers process speech signals using fixed parameters (pitch, formant), then voice conversion function is achieved, but user-friendliness deteriorates for users unfamiliar with acoustic theory

Engineering Contradiction:
Improveuser-friendlinessVSAvoidparameter setting complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system automatically analyzes the user's speech characteristics and selects appropriate conversion parameters without requiring manual input. The processor extracts features from the input speech signal and autonomously determines pitch shift amounts and formant parameters, allowing the device to serve itself in configuring the conversion settings.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system dynamically adjusts conversion parameters based on the analyzed speech characteristics. Instead of using fixed parameters, the pitch shift amount and formant parameters are changed according to the specific user's voice features, enabling adaptive optimization of conversion quality for different users and speech conditions.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If voice conversion is performed with fixed processing methods, then implementation is simple, but adaptability deteriorates across various microphones and communication networks

Engineering Contradiction:
Improvemicrophone and network compatibilityVSAvoidprocessing system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The voice conversion system transitions from static fixed-parameter processing to dynamic adaptive processing. The processor continuously analyzes the input speech signal characteristics and adjusts conversion parameters in real-time, enabling the system to adapt to different microphones, network conditions, and user preferences while maintaining implementation feasibility through automated parameter selection.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If detailed speech feature analysis is performed, then conversion quality is improved, but processing time increases

Engineering Contradiction:
Improveconversion qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of the speech signal to extract essential features (pitch, formant parameters) before the actual voice conversion process. By pre-processing and identifying key characteristics in advance, the system prepares conversion parameters that can be applied efficiently during real-time conversion, reducing overall processing time while maintaining high conversion quality.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12027165B2Computer program, server, terminal, and speech signal processing method
Publication Date: 2024.07.02 GREE HOLDINGS INC
  • US12027165B2 patent drawing
  • US12027165B2 patent drawing
  • US12027165B2 patent drawing

AI summary

A non-transitory computer readable medium stores computer executable instructions which, when executed by at least one processor, cause the at least one processor to acquire a speech signal of speech of a user; perform a signal processing on the speech signal to acquire at least one feature of the speech of the user; and control display of information, related to each of one or more first candidate converters having a feature corresponding to the at least one feature, to present the one or more first candidate converters for selection by the user.