Action optimization device, method and program
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing optimization systems for controlling environments in buildings, such as air conditioning and cleaning, face challenges with time lags in responding to non-optimal conditions and inability to account for medium-term and long-term changes in people flow, leading to suboptimal energy consumption and comfort issues.
Innovation Solution
An action optimization device that acquires and interpolates environmental data to train environment and exploration models, predicting future states and exploring optimal actions using time-series analysis and data augmentation, allowing for flexible control and consideration of local conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a feedback-type optimization system is used to detect non-optimal states and optimize control, then control optimization is achieved, but a time lag occurs until returning to an optimal state
Solution Approach 1:
The system performs preliminary action by predicting future environmental states and people flow trends before non-optimal conditions occur. The prediction unit forecasts future states, and the optimization unit determines control actions in advance based on these predictions, eliminating the time lag inherent in feedback systems that only react after deviations occur.
2Speed
If a simple feedforward system follows short-term people flow changes, then quick response is achieved, but medium-term and long-term trends cannot be optimized
Solution Approach 1:
The system extends the time dimension by incorporating both short-term and long-term prediction capabilities. The prediction unit analyzes historical data to identify trends across multiple time scales, enabling the system to simultaneously respond to immediate changes while optimizing for medium-term and long-term performance, thus adding a temporal dimensionality to the control strategy.
3Ease of operation
If energy consumption upper limit is simply adjusted without estimating control effects, then easy operation is achieved, but interactions such as heat accumulation and cold/hot air flow cannot be considered
Solution Approach 1:
The system implements feedback by using the prediction unit to estimate the effects of proposed control actions on environmental parameters. The optimization unit receives this predicted effect information and adjusts control strategies accordingly, creating a closed-loop system that considers interactions like heat accumulation and air flow while maintaining operational simplicity through automated prediction and optimization.
4Device complexity
If environmental data is not interpolated in time and space, then data processing is simple, but prediction accuracy for future states is reduced
Solution Approach 1:
The system performs preliminary data processing by interpolating environmental data in time and space before prediction. The prediction unit receives pre-processed, high-resolution data that has been extrapolated to fill temporal and spatial gaps, enabling accurate future state predictions without requiring complex real-time processing during operation.
Data Source
AI summary
Provided is a highly reliable technology for optimizing an action for controlling an environment in a target space. An action optimization device for optimizing an action for controlling an environment: acquires environmental data related to a state of the environment; performs time/space interpolation on the acquired environmental data; trains an environment reproduction model, based on the time/space-interpolated environmental data, such that, when a state of an environment and an action for controlling the environment are input, a correct answer value for an environmental state after the action is output; trains an exploration model such that an action to be taken next is output when an environmental state output from the environment reproduction model is input; predicts a second environmental state corresponding to a first environmental state and a first action by using the trained environment reproduction model; explores for a second action to be taken for the second environmental state; and outputs a result of the exploration.


