Audio Processing System Using Re-trained Synthesis Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio signal processing techniques deteriorate sound quality when modifying sounding conditions such as pitches, volumes, and phonetic identifiers, leading to suboptimal editing results.

Innovation Solution

An audio processing method involving a re-trained synthesis model that generates feature data for acoustic features of an audio signal based on modified sounding conditions, using pre-trained models and additional training with specific condition and feature data from a first audio signal, allowing for precise modification of audio signals while maintaining sound quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional audio signal processing techniques are used to modify sounding conditions (pitch, amplitude, etc.), then the audio signal can be edited according to user instructions, but the sound quality deteriorates

Engineering Contradiction:
Improveaudio signal editing capabilityVSAvoidsound quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system performs preliminary analysis of the audio signal to extract acoustic features and sounding conditions before modification. By pre-processing the signal to identify pitch, amplitude, and other acoustic characteristics, the system prepares the data in a structured format that enables quality-preserving modifications later in the process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediate representation layer between the original audio signal and the modified output. Acoustic features are extracted as intermediate data structures that serve as mediators, allowing modifications to be applied to the feature representation rather than directly to the raw signal, thereby preserving sound quality while enabling editing operations

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If pitch and amplitude of audio signal are analyzed and displayed for editing, then user can modify audio parameters, but sound quality deteriorates due to modification of sounding conditions

Engineering Contradiction:
Improvepitch and amplitude analysis accuracyVSAvoidsound quality after modification
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The system implements feedback by continuously monitoring the acoustic features during the modification process. The extracted pitch and amplitude information feeds back into the synthesis model, allowing the system to adjust modifications in real-time to maintain sound quality while achieving the desired editing results

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

Instead of directly modifying the raw audio signal, the system changes parameters in the acoustic feature domain. By adjusting pitch and amplitude as separate extracted features and then reconstructing the signal from these modified features, the system achieves precise parameter control without the quality degradation that occurs with direct signal manipulation

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11842720B2Audio processing method and audio processing system
Publication Date: 2023.12.12 YAMAHA CORP
  • US11842720B2 patent drawing
  • US11842720B2 patent drawing
  • US11842720B2 patent drawing

AI summary

An audio processing system and a method thereof generate a synthesis model that can input an audio signal to generate feature data that can be used by a signal generator to generate a modified audio signal. Specifically, a pre-trained synthesis model is first generated using training audio data. Thereafter, a re-trained synthesis model is established by additionally training the pre-trained synthesis model. Based on a received instruction to modify at least one of sounding conditions of an audio signal to be processed, feature data is generated by inputting additional condition data into the re-trained synthesis model. The signal generator generates the modified audio signal from the generated feature data.