Voice Mixing Conversion with Quality Scoring for Speech Denoising

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis techniques face challenges in obtaining high-quality speech training data, suffer from high annotation costs, and require lengthy evaluation times using single quality evaluation methods.

Innovation Solution

A voice mixing conversion system and method that includes data pre-processing, model training, speech denoising and separation, noise reduction processing, and quality score calculation to enhance speech quality and reduce evaluation time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If traditional speech synthesis techniques are used, then speech synthesis can be achieved, but the cost of speech annotation is high and obtaining speech training data is difficult

Engineering Contradiction:
Improvecost of speech annotationVSAvoidspeech training data
Core Design Contradiction:
Ease of manufactureVSQuantity of substance

Solution Approach 1:

The patent uses voice conversion technology to copy and transform existing voice data into synthetic speech training data. By converting voices from unknown test audio files into target speaker voices through the trained conversion model, the system generates synthetic training data that replicates the characteristics of real speech data, thereby obtaining sufficient training data without requiring expensive manual annotation or recording of diverse speakers

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-service by automatically converting voice data into training data without requiring external manual annotation. The voice conversion model processes unknown test audio files and automatically generates synthetic speech data that can be used for training, making the data generation process autonomous and cost-effective

Inventive Principle:
Principle #25Self-service

2Productivity

If a single speech quality evaluation method is used, then evaluation can be performed, but the evaluation time is long and the best quality synthesized speech may not always be obtained

Engineering Contradiction:
Improveevaluation timeVSAvoidspeech quality evaluation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent merges multiple speech quality evaluation methods into a comprehensive evaluation system. By combining traditional objective evaluation metrics with subjective evaluation and using the multi-person voice mixing output system to generate multiple candidate speeches for evaluation, the system achieves both high measurement precision and improved productivity through automated multi-criteria assessment

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If voice conversion is performed for multiple speakers, then speech diversity is improved, but the complexity of processing and evaluation increases

Engineering Contradiction:
Improvemulti-person voice mixingVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the voice conversion process into distinct stages: (1) extracting voice features from source audio, (2) converting voices using the trained model for each target speaker, (3) generating multiple candidate speeches, and (4) evaluating and selecting the best output. This segmentation allows the system to handle multiple speakers systematically, improving adaptability while managing processing complexity through structured workflows

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250342848A1Voice mixing conversion system and voice mixing conversion method
Publication Date: 2025.11.06 IND TECH RES INST
  • US20250342848A1 patent drawing
  • US20250342848A1 patent drawing
  • US20250342848A1 patent drawing

AI summary

A voice mixing conversion method includes: performing a noise reduction processing on the initial generated speech based on a noise threshold to generate a post-noise reduction speech; calculating a first quality score and a second quality score, wherein in response to a number of the initial generated speech being 1, the first quality score is calculated based on the initial generated speech, and the second quality score is calculated based on the post-noise reduction speech; and determining whether the first quality score is greater than the second quality score, wherein in response to the first quality score being greater than the second quality score, the initial generated speech is output, otherwise the post-noise reduction speech is output.