Performance Agent Reinforcement Learning for Performer Compatibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating a performance agent suitable for one performer results in high costs due to the variability of attributes such as performance ability and musical instrument, making manual generation inefficient.

Innovation Solution

A performance agent training method that observes a performer's first performance, generates parallel performance data, outputs the data for simultaneous performance, and trains the agent using reinforcement learning with satisfaction as a reward to enhance compatibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a performance agent is manually generated for each performer, then the performance agent can be customized to match the performer's attributes, but the cost of generating performance agents becomes extremely high

Engineering Contradiction:
Improvecompatibility with performerVSAvoidcost of generating performance agent
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent uses machine learning models to automatically copy and adapt performance agent parameters from training data corresponding to the performer's attributes, replacing manual generation. The estimation model infers performance agent parameters by copying patterns from similar performers in the training set, significantly reducing generation cost while maintaining compatibility.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent dynamically changes performance agent parameters based on the performer's attributes (instrument type, performance ability level) by adjusting the estimation model's input parameters. This allows the same base model to generate customized performance agents by simply changing input parameters rather than manually creating each agent.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If machine learning is used to train the performance agent, then the generation cost is reduced, but additional training time and computational resources are required

Engineering Contradiction:
Improvecost of generating performance agentVSAvoidtraining time
Core Design Contradiction:
Ease of manufactureVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-collecting and organizing training data corresponding to multiple performers with different attributes before actual performance agent generation. This preprocessing of training data reduces the computational burden during actual generation, trading some initial setup time for faster subsequent agent creation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a universal estimation model that can handle multiple performer types and instruments through a single trained model. This multi-functional model reduces the need for separate training processes for each performer type, consolidating training efforts into one comprehensive model that serves all performers.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12367854B2Performance agent training method, automatic performance system, and program
Publication Date: 2025.07.22 YAMAHA CORP
  • US12367854B2 patent drawing
  • US12367854B2 patent drawing
  • US12367854B2 patent drawing

AI summary

A performance agent training method realized by at least one computer includes observing a first performance of a musical piece by a performer, generating, by a performance agent, performance data of a second performance to be performed in parallel with the first performance, outputting the performance data such that the second performance is performed in parallel with the first performance of the performer, acquiring a degree of satisfaction of the performer with respect to the second performance performed based on the output performance data, and training the performance agent by reinforcement learning, using the degree of satisfaction as a reward.