Generative Sound Synthesis Attack Precision

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional sound synthesis technologies face difficulties in generating musical instrument sounds with appropriate attacks, often resulting in ambiguous attacks that do not accurately reflect the musical characteristics of the note sequence.

Innovation Solution

A sound generation method that acquires both a control data sequence representing the features of a note sequence and a second control data sequence representing performance motion, such as tonguing or bowing parameters, and processes these with trained generative models to produce sound data sequences with accurate attacks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional synthesis technology is used to generate musical sounds from note sequences, then the sound generation process is simple, but the attack characteristics of the synthesized sound become ambiguous and do not reflect musical characteristics

Engineering Contradiction:
Improveattack precisionVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the control data into two distinct sequences: first control data representing note sequence features and second control data representing performance motion features. This segmentation allows the system to process attack characteristics separately from basic note information, thereby improving attack precision without overwhelming system complexity through modular data handling

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by pre-processing performance motion data to extract attack characteristics before sound synthesis. The second control data sequence is prepared in advance to contain specific attack parameters that will be applied during synthesis, allowing the system to achieve precise attack control without increasing real-time processing complexity

Inventive Principle:
Principle #10Preliminary action

2Reliability

If only note sequence features are used for sound synthesis, then the synthesis process is straightforward, but the synthesized sound lacks realistic attack characteristics corresponding to actual performance motions

Engineering Contradiction:
Improvesound fidelityVSAvoiddata quantity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential attack-related features from performance motion data into the second control data sequence. Rather than using all raw performance data, the system selectively extracts parameters directly related to attack characteristics (such as articulation type, attack time, and intensity), thereby improving sound fidelity while minimizing data quantity requirements

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces control data sequences as intermediary representations between raw performance motion and final synthesized sound. These control data sequences act as mediators that translate complex performance motions into standardized attack parameters that the synthesis system can process efficiently, improving reliability without requiring proportional increases in data quantity

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240428760A1Sound generation method, sound generation system, and program
Publication Date: 2024.12.26 YAMAHA CORP
  • US20240428760A1 patent drawing
  • US20240428760A1 patent drawing
  • US20240428760A1 patent drawing

AI summary

A sound generation method that is realized by a computer system includes acquiring a first control data sequence representing a feature of a note sequence and a second control data sequence representing a performance motion for controlling an attack of a musical instrument sound corresponding to each note of the note sequence, and processing the first control data sequence and the second control data sequence with a trained first generative model, thereby generating a sound data sequence representing a musical instrument sound of the note sequence having an attack corresponding to the performance motion represented by the second control data sequence.