Transcript-Based Voice Enhancement Circuitry

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio processing systems require professional recording conditions and equipment to achieve high-quality voice recordings, limiting their effectiveness for non-professional use.

Innovation Solution

An electronic device and method that perform transcript-based voice enhancement using circuitry configured to obtain an enhanced audio signal through processes like forced alignment, gradient descent, and vocoding, based on Automatic Speech Recognition and Deep Neural Network processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If professional recording equipment and indoor recording conditions are used, then voice intelligibility and audio quality are improved, but device complexity and recording conditions requirements increase

Engineering Contradiction:
Improvevoice intelligibilityVSAvoidrecording equipment requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces physical acoustic treatment and professional recording equipment with signal processing algorithms. The voice enhancement is achieved through transcript-based processing that works on recorded audio signals, substituting the need for controlled acoustic environments and expensive equipment with computational methods that can run on standard devices.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the parameters of the audio signal through transcript-based enhancement. By using the transcript information to guide signal processing operations, the system modifies spectral characteristics, temporal features, and amplitude parameters of the audio signal to improve voice intelligibility without requiring changes to the physical recording environment.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If professional recording conditions are required, then audio quality is improved, but ease of operation and accessibility decrease

Engineering Contradiction:
Improveaudio qualityVSAvoidrecording accessibility
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs automatic transcript-based voice enhancement without requiring user intervention for manual audio processing. The transcript is automatically generated or provided, and the enhancement process is automatically applied to the audio signal, making the system easy to operate and accessible to non-professional users.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The voice enhancement processing is performed as a post-processing step after recording, allowing users to capture audio in casual settings and then enhance it automatically. This preliminary enhancement action removes the need for users to understand or control complex recording parameters during the actual recording process.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If transcript-based voice enhancement is implemented, then voice intelligibility is improved without professional equipment, but processing time and computational requirements increase

Engineering Contradiction:
Improvevoice intelligibilityVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The transcript is generated or obtained before the voice enhancement processing begins. This preliminary preparation of the transcript information allows the enhancement algorithm to work efficiently by having the reference data ready, reducing the overall processing time compared to generating the transcript during the enhancement process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11670292B2Electronic device, method and computer program
Publication Date: 2023.06.06 SONY GROUP CORP
  • US11670292B2 patent drawing
  • US11670292B2 patent drawing
  • US11670292B2 patent drawing

AI summary

An electronic device comprising circuitry configured to perform a transcript based voice enhancement based on a transcript to obtain an enhanced audio signal.