IoT Device Control Value Generation via Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional reinforcement learning systems require manual specification of external factors affecting reward fluctuations and manual updates to the learning model, making them inefficient in responding to disturbances without human intervention.

Innovation Solution

A device control value generation device that automatically extracts disturbance constituent factors and defines situations by using IoT data to determine external factors, generating device control values, and updating the learning model through reinforcement learning to achieve optimal device control values without human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual specification of external factors and manual updates to the learning model are used, then the system can respond to disturbances, but the efficiency and responsiveness are reduced due to human intervention requirements

Engineering Contradiction:
Improveefficiency of responding to disturbancesVSAvoidlevel of human intervention
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The system automatically extracts disturbance constituent factors from IoT device data and performs reinforcement learning updates without human intervention. The situation classification unit autonomously identifies external factors affecting reward fluctuations and reconstructs the learning model, enabling the system to serve itself in responding to environmental changes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system proactively monitors IoT device data to detect disturbances before they significantly impact system performance. By continuously analyzing data from multiple IoT devices and pre-identifying external factors, the system prepares and updates learning models in advance, enabling faster response to environmental changes.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If the learning model is reconstructed frequently to adapt to changing normal state tendencies, then the system can respond to state changes, but the computational cost and time consumption increase

Engineering Contradiction:
Improveability to respond to state changesVSAvoidtime for model reconstruction
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The situation classification unit focuses computational resources on specific external factors that have the greatest impact on reward fluctuations. By identifying and analyzing only the most significant disturbance constituent factors from IoT device data, the system achieves effective adaptation without requiring complete model reconstruction, thereby reducing computational overhead and time consumption.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Instead of performing full learning model reconstruction for every state change, the system applies partial updates by reconstructing only with data relevant to the detected external factors. This selective approach maintains adaptability to state changes while significantly reducing the computational burden and time required compared to complete model reconstruction.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If the system monitors and responds to all possible disturbances, then the reliability of meeting predetermined rewards improves, but the system complexity increases

Engineering Contradiction:
Improveability to meet predetermined rewardsVSAvoidcomplexity of disturbance monitoring system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The situation classification unit extracts and isolates only the disturbance constituent factors that significantly affect reward fluctuations from the comprehensive IoT device data. By separating and focusing on these key external factors rather than monitoring all possible disturbances, the system maintains high reliability in meeting predetermined rewards while reducing the overall system complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system dynamically adjusts the parameters and thresholds for disturbance detection based on the specific external factors identified. By changing the monitoring parameters to focus on the most impactful factors rather than uniformly monitoring all disturbances, the system achieves reliable reward fulfillment with reduced computational and operational complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20230316133A1Device control value generation apparatus, device control value generation method, program, and learning model generation apparatus
Publication Date: 2023.10.05 NIPPON TELEGRAPH & TELEPHONE CORP
  • US20230316133A1 patent drawing
  • US20230316133A1 patent drawing
  • US20230316133A1 patent drawing

AI summary

This device control value generation device comprises: a control value generation unit that generates a device control value for a plurality of control target devices; a learning data management unit that acquires items of learning data represented by the device control values and scores, and stores the same in a learning data DB for each device control factor pattern, which represent the device control values in accordance with the division range of external factors; a situation classification unit that extracts the external factor influencing a reward fluctuation and defines classifications; and a learning model management unit that uses learning data for each defined classification and generates a learning model of each classification by means of carrying out reinforcement learning such that a prescribed reward is fulfilled.