Device control method and apparatus
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional control methods for air conditioners, such as PID and model predictive control, fail to achieve ideal operating effects due to their reliance on human-designed rules or accurate system models, leading to high costs and poor generalization in varying environments.
Innovation Solution
A device control method utilizing reinforcement learning models that leverage global data from a cloud to adapt to different environments, enabling autonomous learning and collaborative training with local devices, thereby optimizing energy consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional control methods (PID, model predictive control) are used, then implementation costs are low or model accuracy is high, but control effect is poor and generalization capability is poor
Solution Approach 1:
The system employs reinforcement learning models that enable air conditioners to autonomously learn and optimize control strategies without relying on pre-defined human rules or accurate thermodynamic models. The model independently adapts to different environments through continuous learning from operational data, achieving both effective control and strong generalization capability across varying working conditions
Solution Approach 2:
The patent utilizes reinforcement learning to dynamically adjust control parameters based on learned patterns from data rather than relying on fixed mathematical models. This allows the system to adapt parameters optimally for different environments, resolving the contradiction between achieving good control effect and maintaining generalization capability without requiring accurate system models
2Reliability
If model predictive control based on thermodynamic modeling is used, then control effect depends on model quality, but engineering implementation costs are high
Solution Approach 1:
The patent replaces expensive, complex thermodynamic models with simpler reinforcement learning models that can be trained on operational data. This substitution significantly reduces engineering implementation costs while maintaining or improving control effect, as the RL model learns optimal control strategies directly from data without requiring detailed system modeling
Solution Approach 2:
The patent substitutes traditional model-based control approaches with data-driven reinforcement learning. This replacement eliminates the need for complex thermodynamic modeling and reduces implementation costs, while the learned policies achieve comparable or superior control performance across diverse operating conditions
3Adaptability or versatility
If reinforcement learning model is updated using cloud data, then generalization capability is improved, but data transmission and processing complexity increases
Solution Approach 1:
The patent divides the system into cloud and local components, where the cloud performs centralized training of reinforcement learning models using aggregated data from multiple devices. The trained models are then deployed to local air conditioners for execution. This segmentation allows complex data processing to be performed centrally while keeping individual device complexity low,同时 achieving improved generalization capability
Solution Approach 2:
The patent introduces a cloud platform as an intermediary that aggregates data from multiple air conditioners, performs centralized reinforcement learning training, and distributes optimized models back to devices. This intermediary handles the complex data processing tasks centrally, reducing individual device complexity while improving overall system generalization capability through shared learning
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A device control method and a device control apparatus are provided. The control method includes: receiving a first set from a cloud, where the first set includes at least one of the following: a plurality of pieces of training data, a training parameter of a first reinforcement learning model, or a parameter of the first reinforcement learning model (S410); updating, based on the first set, a second reinforcement learning model deployed in a first device to the first reinforcement learning model, where the first reinforcement learning model is used to adjust energy consumption of the first device (S420); obtaining a first state of the first device (S430); and processing the first state of the first device by using the first reinforcement learning model, to obtain a first control action used to adjust an operating parameter of the first device (S440). The device control method helps implement energy consumption optimization of a device.