Multi-Encoder Music Processing for Selective Rhythm Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing information processing systems using machine learning struggle to generate songs that match user-designated features, such as changing a melody while maintaining a certain rhythm, as they fail to selectively learn specific characteristics of the content.

Innovation Solution

An information processing apparatus employs multiple encoders to separately extract feature quantities from entire content and specific data, allowing for the generation of a learned model that can selectively learn user-designated features, such as rhythm, by using a first encoder for the entire content and a second encoder for specific data, and a decoder to reconstitute the content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single encoder is used to learn features of entire content, then the system can generate songs automatically, but it cannot selectively learn specific features designated by the user

Engineering Contradiction:
Improveselective feature learningVSAvoidmodel structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent divides the feature learning process into two separate encoders: a first encoder that learns features from entire content and a second encoder that learns features from specific data. This segmentation allows the system to selectively learn specific features designated by the user while maintaining the overall song generation capability, directly resolving the contradiction between adaptability and complexity.

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If machine learning learns all features of content, then comprehensive song generation is achieved, but user-designated feature control is lost

Engineering Contradiction:
Improvefeature controlVSAvoidfeature separation
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent extracts specific features from the entire content by using a second encoder that processes only specific data related to the user-designated feature. This extraction mechanism allows the system to maintain control over user-designated features while preserving the comprehensive information from the first encoder, thereby achieving both ease of operation and preventing information loss.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If the system generates songs with all learned features, then complete song reproduction is achieved, but user-specified feature modification is impossible

Engineering Contradiction:
Improvefeature modificationVSAvoidfeature accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies local quality by allowing different encoders to learn different aspects of the content with different levels of detail. The first encoder learns comprehensive features for complete song reproduction, while the second encoder focuses on specific user-designated features for accurate feature modification. This enables the system to achieve both adaptability in feature modification and precision in maintaining user-specified features.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250266023A1Information processing method, information processing apparatus, and information processing program
Publication Date: 2025.08.21 SONY GROUP CORP
  • US20250266023A1 patent drawing
  • US20250266023A1 patent drawing
  • US20250266023A1 patent drawing

AI summary

An information processing apparatus according to the present disclosure includes an extraction unit that extracts first data from an element constituting first content, and a model generation unit that generates a learned model that has a first encoder that calculates a first feature quantity which is a feature quantity of first content, and a second encoder that calculates a second feature quantity which is a feature quantity of the extracted first data.