Learning Model Hyperparameter Adaptation Against Changing Opponents
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing learning models struggle to adapt effectively when opponents in a competitive environment change, leading to difficulty in generating models that can win against new opponents.
Innovation Solution
A learning device and method that evaluates the strength of opponents and adjusts hyperparameters of the learning model based on this strength, enabling reinforcement learning to be performed on either a search or use side depending on the opponent's strength.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If reinforcement learning is performed using a learning model in the related art, then learning can be executed under a competitive environment, but the learning model cannot adapt effectively when the opponent changes, resulting in low versatility
Solution Approach 1:
The patent applies dynamics by making the learning model's hyperparameter adjustable and adaptive. The learning model dynamically changes its hyperparameter based on the evaluated strength of the opponent, transitioning from a static configuration to a dynamic one that responds to environmental changes. This enables the model to adapt to different opponent strengths and maintain effectiveness across varying competitive environments.
Solution Approach 2:
The patent changes the hyperparameter of the learning model based on the strength evaluation of the opponent. By adjusting the hyperparameter according to the opponent's strength, the learning model can optimize its learning process for different competitive scenarios. This parameter change enables the model to handle both strong and weak opponents effectively, resolving the contradiction between adaptability and reliability.
2Productivity
If the learning model uses a fixed hyperparameter, then the learning process is simple, but the model cannot search for optimal learning models capable of winning against diverse opponents
Solution Approach 1:
The patent transforms the fixed hyperparameter into a dynamic one that changes based on opponent strength evaluation. This dynamic adjustment allows the learning model to optimize its learning efficiency for different scenarios while maintaining versatility. The system can efficiently learn from weak opponents and adaptively search for optimal models when facing strong opponents.
Solution Approach 2:
The patent implements parameter changes by adjusting the hyperparameter according to the evaluated strength of the opponent. This enables the learning model to balance learning efficiency and versatility - using appropriate hyperparameter values for different opponent types to maximize both productivity and adaptability simultaneously.
3Reliability
If learning is performed against a strong opponent, then the learning model can improve its strength, but the learning process becomes more difficult and time-consuming
Solution Approach 1:
The patent applies preliminary action by first evaluating the strength of the opponent before initiating the learning process. Based on this preliminary evaluation, the system determines the appropriate hyperparameter setting in advance. This preliminary action allows the learning model to prepare its learning strategy before facing the actual learning task, reducing the time needed for ineffective learning attempts.
Solution Approach 2:
The patent changes the hyperparameter parameter based on the opponent's strength evaluation to optimize the learning process. When facing strong opponents, the system adjusts the hyperparameter to facilitate more efficient learning, reducing the time required to improve the model's strength while maintaining reliability against the challenging opponent.
Data Source
AI summary
There is provided a learning device including: a processing unit that performs reinforcement learning of a learning model of an agent under a competitive environment in which agents compete against each other, in which the learning model includes a hyperparameter, and the processing unit executes: a step of setting the agent to be an opponent of the agent as a learning target; a step of evaluating a strength of the agent that is the opponent; a step of setting the hyperparameter of the learning model of the agent as the learning target according to the strength of the agent that is the opponent; and a step of executing the reinforcement learning by using the learning model after the setting.


