IoT Device Control Value Generation via Reinforcement Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional reinforcement learning systems require manual specification of external factors affecting reward fluctuations and manual updates to the learning model, making them inefficient in responding to disturbances without human intervention.
Innovation Solution
A device control value generation device that automatically extracts disturbance constituent factors and defines situations by using IoT data to determine external factors, generating device control values, and updating the learning model through reinforcement learning to achieve optimal device control values without human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual specification of external factors and manual updates to the learning model are used, then the system can respond to disturbances, but the efficiency and responsiveness are reduced due to human intervention requirements
Solution Approach 1:
The system automatically extracts disturbance constituent factors from IoT device data and performs reinforcement learning updates without human intervention. The situation classification unit autonomously identifies external factors affecting reward fluctuations and reconstructs the learning model, enabling the system to serve itself in responding to environmental changes.
Solution Approach 2:
The system proactively monitors IoT device data to detect disturbances before they significantly impact system performance. By continuously analyzing data from multiple IoT devices and pre-identifying external factors, the system prepares and updates learning models in advance, enabling faster response to environmental changes.
2Adaptability or versatility
If the learning model is reconstructed frequently to adapt to changing normal state tendencies, then the system can respond to state changes, but the computational cost and time consumption increase
Solution Approach 1:
The situation classification unit focuses computational resources on specific external factors that have the greatest impact on reward fluctuations. By identifying and analyzing only the most significant disturbance constituent factors from IoT device data, the system achieves effective adaptation without requiring complete model reconstruction, thereby reducing computational overhead and time consumption.
Solution Approach 2:
Instead of performing full learning model reconstruction for every state change, the system applies partial updates by reconstructing only with data relevant to the detected external factors. This selective approach maintains adaptability to state changes while significantly reducing the computational burden and time required compared to complete model reconstruction.
3Reliability
If the system monitors and responds to all possible disturbances, then the reliability of meeting predetermined rewards improves, but the system complexity increases
Solution Approach 1:
The situation classification unit extracts and isolates only the disturbance constituent factors that significantly affect reward fluctuations from the comprehensive IoT device data. By separating and focusing on these key external factors rather than monitoring all possible disturbances, the system maintains high reliability in meeting predetermined rewards while reducing the overall system complexity.
Solution Approach 2:
The system dynamically adjusts the parameters and thresholds for disturbance detection based on the specific external factors identified. By changing the monitoring parameters to focus on the most impactful factors rather than uniformly monitoring all disturbances, the system achieves reliable reward fulfillment with reduced computational and operational complexity.
Data Source
AI summary
This device control value generation device comprises: a control value generation unit that generates a device control value for a plurality of control target devices; a learning data management unit that acquires items of learning data represented by the device control values and scores, and stores the same in a learning data DB for each device control factor pattern, which represent the device control values in accordance with the division range of external factors; a situation classification unit that extracts the external factor influencing a reward fluctuation and defines classifications; and a learning model management unit that uses learning data for each defined classification and generates a learning model of each classification by means of carrying out reinforcement learning such that a prescribed reward is fulfilled.


