Voice Mixing Conversion with Quality Scoring for Speech Denoising
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis techniques face challenges in obtaining high-quality speech training data, suffer from high annotation costs, and require lengthy evaluation times using single quality evaluation methods.
Innovation Solution
A voice mixing conversion system and method that includes data pre-processing, model training, speech denoising and separation, noise reduction processing, and quality score calculation to enhance speech quality and reduce evaluation time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If traditional speech synthesis techniques are used, then speech synthesis can be achieved, but the cost of speech annotation is high and obtaining speech training data is difficult
Solution Approach 1:
The patent uses voice conversion technology to copy and transform existing voice data into synthetic speech training data. By converting voices from unknown test audio files into target speaker voices through the trained conversion model, the system generates synthetic training data that replicates the characteristics of real speech data, thereby obtaining sufficient training data without requiring expensive manual annotation or recording of diverse speakers
Solution Approach 2:
The system performs self-service by automatically converting voice data into training data without requiring external manual annotation. The voice conversion model processes unknown test audio files and automatically generates synthetic speech data that can be used for training, making the data generation process autonomous and cost-effective
2Productivity
If a single speech quality evaluation method is used, then evaluation can be performed, but the evaluation time is long and the best quality synthesized speech may not always be obtained
Solution Approach 1:
The patent merges multiple speech quality evaluation methods into a comprehensive evaluation system. By combining traditional objective evaluation metrics with subjective evaluation and using the multi-person voice mixing output system to generate multiple candidate speeches for evaluation, the system achieves both high measurement precision and improved productivity through automated multi-criteria assessment
3Adaptability or versatility
If voice conversion is performed for multiple speakers, then speech diversity is improved, but the complexity of processing and evaluation increases
Solution Approach 1:
The patent segments the voice conversion process into distinct stages: (1) extracting voice features from source audio, (2) converting voices using the trained model for each target speaker, (3) generating multiple candidate speeches, and (4) evaluating and selecting the best output. This segmentation allows the system to handle multiple speakers systematically, improving adaptability while managing processing complexity through structured workflows
Data Source
AI summary
A voice mixing conversion method includes: performing a noise reduction processing on the initial generated speech based on a noise threshold to generate a post-noise reduction speech; calculating a first quality score and a second quality score, wherein in response to a number of the initial generated speech being 1, the first quality score is calculated based on the initial generated speech, and the second quality score is calculated based on the post-noise reduction speech; and determining whether the first quality score is greater than the second quality score, wherein in response to the first quality score being greater than the second quality score, the initial generated speech is output, otherwise the post-noise reduction speech is output.


