Sound Generation Model Coarse-to-Fine Feature Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound generation methods require users to specify detailed time series of musical feature amounts like amplitude, volume, and pitch, making it difficult to generate naturally changing sounds such as those from singing or performing.

Innovation Solution

A sound generation method using a trained model that processes a first feature amount sequence with a lower fineness to generate a sound data sequence with a higher fineness, allowing for the easy acquisition of natural sounds by learning the input-output relationship between the sequences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If detailed time series of musical feature amounts are specified, then sound generation accuracy is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvesound generation accuracyVSAvoidease of operation
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The system performs preliminary extraction of musical feature amounts from reference sound data before the actual sound generation process. By pre-processing and storing feature amount sequences at high fineness, the system eliminates the need for users to manually specify detailed time series data during operation, thus improving ease of operation while maintaining sound generation accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates copies of musical feature amount sequences from reference sound data at different fineness levels. These copied sequences serve as templates that can be directly used for sound generation without requiring users to manually create detailed specifications, resolving the contradiction between accuracy and ease of operation

Inventive Principle:
Principle #26Copying

2Reliability

If fineness of feature amount sequence is increased, then sound naturalness is improved, but data processing complexity increases

Engineering Contradiction:
Improvesound naturalnessVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the feature amount sequence into multiple representations at different fineness levels. By dividing the data into coarse and fine granularity versions, the system can process lower fineness data for basic operations while maintaining high fineness data for natural sound generation, thus reducing overall data processing complexity while preserving sound naturalness

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system adds a fineness dimension to the feature amount data structure. By organizing data in multiple dimensional layers (different fineness levels), the system can efficiently process data at appropriate granularity for each operation, reducing complexity while maintaining the capability to generate natural sounds when needed

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20230386440A1Sound generation method using machine learning model, training method for machine learning model, sound generation device, training device, non-transitory computer-readable medium storing sound generation program, and non-transitory computer-readable medium storing training program
Publication Date: 2023.11.30 YAMAHA CORP
  • US20230386440A1 patent drawing
  • US20230386440A1 patent drawing
  • US20230386440A1 patent drawing

AI summary

A sound generation method that is realized by a computer includes receiving a first feature amount sequence in which a musical feature amount changes over time, and using a trained model that has learned an input-output relationship between an input feature amount sequence in which the musical feature amount changes over time at a first fineness and a reference sound data sequence corresponding to an output feature amount sequence in which the musical feature amount changes over time at a second fineness that is higher than the first fineness, to process the first feature amount sequence, thereby generating a sound data sequence corresponding to a second feature amount sequence in which the musical feature amount changes at the second fineness.