Personalized control method and system for electrochromic glass based on migration reinforcement learning

By employing a transfer-based reinforcement learning approach combined with dynamic hybrid rewards and multimodal perception technology, the cold start and personalized adaptation issues of electrochromic glass control technology were resolved, enabling rapid personalized control of electrochromic glass while optimizing energy saving and comfort.

CN121918339APending Publication Date: 2026-04-24GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU UNIVERSITY
Filing Date
2026-03-20
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing electrochromic glass control technology has significant shortcomings in terms of rapid deployment, personalized learning, and real-time optimization. In particular, it suffers from difficulties in cold start, lack of personalized adaptation capabilities, and sparse learning signals, making it difficult to achieve both maximum energy saving and meet the personalized comfort needs of different users.

Method used

A transfer learning-based control method is adopted. A general control policy network for reinforcement learning agents is trained by constructing a building environment simulation model. By combining a dynamic hybrid reward function and a lightweight agent model, transfer learning from general to personalized is achieved. Multimodal perception technology is introduced to capture implicit feedback and construct a closed-loop adaptive system.

Benefits of technology

It enables electrochromic glass to quickly adapt to specific users' personalized habits, balancing energy saving and comfort, and has continuous optimization capabilities. It can respond to user needs in real time on edge computing devices and provide precise personalized control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121918339A_ABST
    Figure CN121918339A_ABST
Patent Text Reader

Abstract

The invention discloses an electrochromic glass personalized control method and system based on migration reinforcement learning, and the method comprises the steps: S1, constructing a building environment simulation model, training a universal control strategy network of a reinforcement learning agent through simulation data, and obtaining a basic universal model; s2, deploying the basic general model into an actual electrochromic glass control system, outputting a suggested control action according to a current environment state, and monitoring whether a user has a manual intervention behavior or not in real time; s3, calculating an instant reward by adopting a dynamic mixed reward function according to a user intervention condition; and S4, based on the instant reward, the current environment state and the actual execution action, carrying out fine adjustment on the strategy network of the reinforcement learning agent, and realizing transfer learning to the personalized preference of the specific user. According to the invention, the intelligent control of the electrochromic glass can be realized, so that the energy-saving goal can be realized to the greatest extent, and the personalized comfort requirements of different users can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent control technology for building environment, specifically relating to a personalized control method and system for electrochromic glass based on transfer reinforcement learning. Background Technology

[0002] Electrochromic glass, as one of the core materials for smart buildings, can dynamically adjust its light transmittance to control solar radiation heat and indoor lighting, playing a crucial role in building energy conservation and indoor comfort regulation. However, how to achieve intelligent control of electrochromic glass so that it can maximize energy conservation goals while meeting the personalized comfort needs of different users remains a significant technical challenge.

[0003] Currently, the mainstream electrochromic glass control technologies can be divided into three categories: First, control methods based on fixed rules. These methods are simple and easy to implement but have poor adaptability and cannot be optimized according to dynamic environments and user needs. Second, control methods based on traditional optimization models. These methods are theoretically better but have high implementation costs, complex calculations, and difficulty in responding to changes in user preferences in real time. Third, control methods based on machine learning. In particular, reinforcement learning technology has shown potential, but it still faces problems such as difficulty in cold start, lack of personalized adaptation capabilities, and sparse learning signals.

[0004] Existing technical solutions have significant shortcomings in areas such as rapid deployment, personalized learning, and real-time optimization. In particular, how to quickly launch a basic strategy with basic performance, how to effectively capture and adapt to users' personalized preferences, and how to achieve continuous optimization with limited computing resources are all technical problems that urgently need to be solved. Summary of the Invention

[0005] In view of this, the present invention proposes a personalized control method and system for electrochromic glass based on transfer reinforcement learning, which can realize intelligent control of electrochromic glass, so as to maximize energy saving and meet the personalized comfort needs of different users.

[0006] To achieve the above objectives, the present invention provides the following technical solution: This invention provides a personalized control method for electrochromic glass based on transfer reinforcement learning, comprising: S1. Construct a building environment simulation model, and use the simulation data to train a general control policy network of a reinforcement learning agent to obtain a basic general model. S2. Deploy the basic general model into the actual electrochromic glass control system, output suggested control actions based on the current environmental conditions, and monitor in real time whether the user has any manual intervention behavior. S3. Calculate immediate rewards using a dynamic hybrid reward function based on user intervention. S4. Based on immediate rewards, current environmental state, and actual actions performed, the policy network of the reinforcement learning agent is fine-tuned to achieve transfer learning to the personalized preferences of specific users.

[0007] Preferably, in step S1, when training the general control policy network, its state space includes the outdoor temperature. Solar radiation intensity Time characteristics Current indoor temperature and illuminance Its action space is the set of discrete transmittance levels of electrochromic glass, and its reward function is... for: in, and These are the energy consumption for cooling and lighting, respectively. The penalty term is based on a general comfort model, including the PMV model or the glare index model. These are the weighting coefficients.

[0008] Preferably, the dynamic hybrid reward function used in step S3 The mathematical expression is: in, Awards are given based on energy efficiency and general comfort. Penalty rewards associated with deviations in suggested actions and user actions. This is a user intervention flag; it is used when the user intervenes. ,otherwise .

[0009] Preferably, in step S3, when the user intervenes, a punitive reward is applied. The calculation method is as follows: in, The penalty coefficient is... Recommend light transmittance for the system With user-defined light transmittance The absolute value of the difference.

[0010] Preferably, before step S1, a surrogate model construction step is included: using building performance simulation software to generate an indoor light and heat environment dataset covering various working conditions, and training a lightweight neural network surrogate model. Used for real-time prediction of indoor environmental indicators under different actions: in, To predict indoor temperature, To predict the probability of glare, To predict the illuminance of the working face, For glass light transmittance, For outdoor temperature, For solar radiation intensity, As a time feature, during the online control phase, the prediction results of the agent model are used to assist in state construction and / or reward calculation.

[0011] Preferably, the policy update in step S4 uses a near-end policy optimization algorithm or a deep Q-network algorithm.

[0012] Preferably, in step S2, after outputting the suggested action, the system will wait for a preset time window. If user intervention is detected within this window, it will be determined as an intervention event.

[0013] Preferably, it also includes an implicit preference learning step based on multimodal awareness: Deploy non-contact biosignal sensors and behavior sensing devices to collect users' physiological response data and behavioral pattern data in real time; A multimodal data fusion network is constructed to spatiotemporally align and fuse physiological response data, behavioral pattern data, and environmental perception data to generate a multidimensional user state feature vector. Based on multidimensional user state feature vectors, a comfort state classification model is trained. The classification model can divide the user state into multiple discrete comfort levels. Based on the output of the comfort state classification model, an implicit feedback reward function is constructed. When the classification model detects that the user's discomfort level has increased, the system strategy provides a negative reward, even if the user does not make an explicit manual intervention at this time. Implicit feedback rewards are weighted and integrated with existing explicit intervention rewards to form a complete personalized learning signal.

[0014] To achieve the above objectives, the present invention also provides a personalized control system for electrochromic glass based on transfer reinforcement learning, comprising: Environmental sensing module: used to collect solar radiation and outdoor temperature through an outdoor weather station, and to collect indoor illuminance, indoor temperature and the status of people indoors through a multi-sensor node installed indoors; User interaction module: used to provide users with a wall control panel or mobile APP to manually and forcibly adjust the light transmittance of the glass; Central control module: This is an embedded computer or cloud server that runs reinforcement learning algorithms. It is responsible for receiving sensor data, calculating the optimal action, and issuing instructions. Execution module: includes electrochromic glass and its driver, used to receive light transmittance commands and change the glass state.

[0015] The present invention has achieved at least the following beneficial effects: 1. Achieving efficient transfer from general to personalized approaches, balancing energy conservation and individualized comfort: This invention employs a technical approach of offline pre-trained general models combined with online personalized fine-tuning. A general control strategy aimed at energy conservation and universal comfort is pre-trained using building simulation data, solving the problems of low efficiency and high cost in the initial exploration of reinforcement learning in real-world scenarios. After online deployment, the system monitors user intervention and uses a dynamic hybrid reward function to transform explicit user preferences into learning signals, efficiently fine-tuning the general model to quickly adapt it to the personalized habits of specific users. This method inherits the energy-saving and stability foundation of the general model while flexibly adapting to individual differences, fundamentally resolving the contradiction that traditional single strategies cannot balance group efficiency and individual preferences.

[0016] 2. A precise personalized learning mechanism based on the fusion of explicit and implicit feedback is proposed: This invention not only utilizes explicit feedback through user intervention but also creatively introduces multimodal perception technology to capture implicit physiological and behavioral feedback from users through non-contact biosignal and behavioral analysis. The system constructs an implicit feedback reward function through multimodal data fusion and comfort state recognition. This enables the system to perceive unspoken discomfort in users (such as irritability caused by excessive light or restlessness due to temperature discomfort) and automatically adjust its strategies. The fusion of explicit and implicit feedback forms a more complete and precise personalized learning signal, enabling the system not only to respond to user "instructions" but also to understand the user's "state," achieving a leap from passive response to proactive consideration.

[0017] 3. Introducing a lightweight agent model to ensure the real-time performance and feasibility of online learning: Addressing the issue that complex building simulation models cannot be used for real-time online decision-making and reward calculation, this invention pre-trains a lightweight neural network agent model. This model can predict future indoor light and heat conditions and energy consumption with high accuracy based on the current environment and the action to be executed, at millisecond speeds. This enables the system to quickly evaluate the potential consequences of different actions (for forward-looking state construction) and instantly calculate reward signals without delay within each control cycle, greatly improving the efficiency and real-time performance of online reinforcement learning decision-making and policy updates, making it possible to stably run complex personalized online learning on edge computing devices.

[0018] 4. Constructing a closed-loop adaptive system with continuous evolution capabilities: This invention designs the control system of electrochromic glass as an intelligent agent with a complete closed loop of "perception-decision-execution-evaluation-learning". The system can not only make decisions based on the current environment, but also learn from the actual effects of each decision (energy consumption results) and user feedback (explicit or implicit), continuously optimizing its decision model. This means that the system's control strategy is not static, but constantly adjusts and evolves with the passage of time, seasonal changes, and the evolution of user habits, truly achieving "becoming smarter with use" and maintaining optimal personalized performance over the long term.

[0019] Other advantages, objectives, and features of the invention will be set forth in the following description and will be apparent to those skilled in the art in some respects, or may be learned by practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0020] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration: Figure 1 This is a logical architecture diagram of the electrochromic glass personalized control system based on transfer reinforcement learning in an embodiment of the present invention; Figure 2 This is an online control flowchart of the electrochromic glass personalized control system based on transfer reinforcement learning in an embodiment of the present invention. Detailed Implementation

[0021] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0022] To achieve the above objectives, the present invention provides the following technical solution: This invention proposes a control framework that combines offline pre-training, online transfer learning, and surrogate model acceleration.

[0023] Its core principle is as follows: First, a "passable" general-purpose agent is trained using physical simulation data (solving the cold start problem); then, this agent is deployed in an actual room, capturing the user's "override" behavior regarding glass transmittance, treating it as expert teaching signals, and using imitation learning mechanisms to fine-tune the control strategy online, thereby quickly converging to the user's personalized preferences. Simultaneously, a lightweight neural network is introduced as a proxy model for the physical environment, achieving millisecond-level environmental prediction.

[0024] System Hardware Architecture This system consists of four core modules: Environmental sensing module: includes an outdoor weather station (collecting solar radiation and outdoor temperature) and an indoor multi-sensor node (collecting indoor illuminance, indoor temperature, and the status of people indoors).

[0025] User interaction module: This is specifically implemented as a wall-mounted control panel or mobile app, allowing users to manually adjust the glass's light transmittance when they feel uncomfortable. This module is the key entry point for the system to obtain personalized tags.

[0026] Central control module: An embedded computer or cloud server that runs reinforcement learning algorithms, responsible for receiving sensor data, calculating the optimal action, and issuing instructions.

[0027] Execution module: Electrochromic glass and its driver, which receive transmittance commands and change the glass state.

[0028] First preferred embodiment: This invention provides a personalized control method for electrochromic glass based on transfer reinforcement learning, referring to... Figure 1 and Figure 2 This includes the following steps: S1: Construct a building environment simulation model, use simulation data to train a general control policy network for a reinforcement learning agent, and obtain a basic general model; S2: Deploy the basic general model into the actual electrochromic glass control system, output suggested control actions based on the current environmental conditions, and monitor in real time whether the user has any manual intervention behavior; S3: Calculate immediate rewards using a dynamic hybrid reward function based on user intervention. S4: Based on immediate rewards, current environmental state, and actual actions performed, the policy network of the reinforcement learning agent is fine-tuned to achieve transfer learning to the personalized preferences of specific users.

[0029] In this embodiment, "constructing a building environment simulation model" refers to using professional building performance simulation software, such as EnergyPlus or TRNSYS, to create a building model that includes electrochromic windows, the building envelope, the air conditioning system, the lighting system, and internal heat sources. This model needs to accurately reflect the building's thermal characteristics, photothermal response, and energy consumption characteristics. The simulation model will serve as a virtual environment for training the reinforcement learning agent. Through exploration and learning in the simulation environment, the agent can quickly acquire preliminary control strategies without interfering with the operation of the real building.

[0030] In this embodiment, step S1, training the reinforcement learning agent using simulation data, is an offline learning process conducted in a virtual environment. The agent learns through extensive interactions with the simulation environment. Its state space is defined as a set of features that describe the current building's light and heat environment, such as outdoor temperature. Solar radiation intensity Time characteristics (Used to represent time of day, such as morning, noon, and evening, which can be in coded form), current indoor temperature. and working surface illuminance The action space is set as a discrete set of transmittance levels that the electrochromic glass can execute, for example, {10%, 30%, 50%, 70%, 90%}. By defining an appropriate "reward function", the agent is guided to learn which transmittance action to choose in a given state.

[0031] In this embodiment, the reward function in step S1 The aim is to balance the two core objectives of energy conservation and comfort. Its general form can be designed as follows: .in, and These are the estimated or actual energy consumption for cooling and lighting calculated by the simulation model (if the model supports it). This is a comfort penalty, which can be calculated based on common comfort models such as the Predicted Average Votes (PMV) model and the Disturbances Gauge (DGP). This item incurs a negative penalty when the indoor thermal or light environment deviates from the comfort zone. The weighting coefficients are used to adjust the relative importance of energy saving and comfort in the total reward. Through training under the guidance of a general reward function, the resulting basic general model possesses a fundamental understanding of the dynamic changes in the built environment and is capable of making preliminary decisions oriented towards energy saving and universal comfort.

[0032] In this embodiment, step S2 loads the trained basic general model into the hardware of the actual control system (such as an embedded computer or server). The system begins to collect sensor data in real time and, based on the current state (such as real-time outdoor temperature, solar radiation, indoor temperature and humidity measured by sensors), outputs a suggested control action, namely a recommended light transmittance level, using the policy network of the basic general model. Simultaneously, the system continuously monitors whether the user has manually intervened through a user interface (such as a wall control panel or a mobile app). When the user manually adjusts the glass light transmittance, it is considered an intervention.

[0033] In this embodiment, the dynamic hybrid reward function in step S3 is a key mechanism for realizing online personalized learning. Its core idea is to dynamically switch the reward calculation mode based on whether the user intervenes. Its mathematical expression can be designed as follows: .in, It is a binary flag: when user intervention is detected, ;otherwise, When the user does not intervene ( When this happens, the system uses a reward similar to the objective during offline training. This reward continues to encourage energy conservation and general comfort, and its calculation can be based on actual energy consumption data or using a lightweight energy consumption and comfort estimation model. When user intervention occurs ( The system's primary learning signal transforms into a punitive reward. Its value is usually negative and is proportional to the deviation between the system's suggested action and the user's actual action. For example, The greater the deviation, the heavier the penalty. This clearly tells the agent that, in the current state, the recommended action deviates from the user's preference.

[0034] In this embodiment, the fine-tuning process in step S4 achieves transfer learning from a general policy to a personal policy. Unlike training from scratch, fine-tuning is performed on a foundational model that already possesses good general knowledge. The system will incorporate the experience (i.e., states) generated during online operation. Actual actions performed The calculated instant reward The data is collected to form an online experience dataset. Then, reinforcement learning algorithms such as Proximal Policy Optimization (PPO), Deep Deterministic Policy Gradient (DDPG), or their online variants are used to train this experience dataset with a small learning rate, updating the parameters of the policy network. Because the base model already has some knowledge, the fine-tuning process converges quickly and can rapidly adapt to the habits of specific users. For example, if a user always prefers to dim the lights in the afternoon, after receiving several penalties for intervention during that time period, the policy network will quickly adjust and proactively recommend lower light transmittance for similar afternoon times in the future, thereby reducing the number of times the user needs to intervene.

[0035] For example, scene setting: A small, south-facing private office equipped with 5-level adjustable electrochromic glass (transmittance: 60%, 40%, 18%, 6%, 1%). User A is a designer who is extremely sensitive to glare.

[0036] Implementation process: Initial stage: The system loads a general pre-trained model.

[0037] Day 1, 10 AM: Solar radiation intensifies. General models indicate it will not be hot at this time; it is recommended that the glass maintain 40% light transmittance.

[0038] User feedback: User A found it too bright and manually adjusted the glass to 6%.

[0039] Algorithm learning: The system captures the intervention.

[0040] Calculation deviation: $|40%-6%|$ is significantly different.

[0041] Generate reward: Give the algorithm a strong penalty of -100.

[0042] Parameter update: The algorithm immediately adjusts the neural network weights to reduce the probability of selecting high transmittance under strong light.

[0043] The next day at 10 AM: similar weather. After the algorithm was updated, it proactively suggested adjusting the glass setting to 18% (attempting to approximate user preferences).

[0044] User feedback: User A thought it was acceptable and made no adjustments.

[0045] Algorithm optimization: The system did not detect any intervention.

[0046] Generate rewards: Positive rewards are given based on the low energy consumption at the time.

[0047] Result: The algorithm confirmed that "darkening in bright light" was the correct strategy for this user.

[0048] Second preferred embodiment: Based on the first preferred embodiment, before step S1, a surrogate model construction step is also included: using building performance simulation software to generate an indoor light and heat environment dataset covering various working conditions, and training a lightweight neural network surrogate model. Used for real-time prediction of indoor environmental indicators under different actions: ,in, To predict indoor temperature, To predict the probability of glare, To predict the illuminance of the working face, For glass light transmittance, For outdoor temperature, For solar radiation intensity, As a time feature, during the online control phase, the prediction results of the agent model are used to assist in state construction and / or reward calculation.

[0049] The lightweight neural network proxy model MenvMenv is introduced in this embodiment because using traditional building simulation software (such as EnergyPlus and TRNSYS) for real-time prediction and optimization is not feasible in terms of computing power and timeliness. Although these high-precision simulation software can accurately simulate the thermal and lighting environment of buildings, their single simulations usually take several minutes or even hours, which cannot meet the millisecond-level decision-making and reward calculation requirements of online control systems. Therefore, to solve this computing power bottleneck in practical engineering, this invention pre-trains a lightweight neural network proxy model using batch data generated by simulation software. This model can complete real-time prediction of key environmental indicators such as indoor temperature, glare probability, and work surface illuminance at millisecond speeds during the deployment phase. This significantly improves the system response speed and online learning efficiency without sacrificing prediction accuracy, making it possible to stably run complex reinforcement learning decisions and real-time personalized optimization on edge devices.

[0050] In this embodiment, the proxy model is a simplified model that replaces high-fidelity but computationally expensive building simulation software. Its construction process employs a model reduction technique. First, extensive parametric simulations are performed using software such as EnergyPlus to generate a dataset containing tens or even hundreds of thousands of different operating conditions (covering various outdoor climates, times, and glass conditions). Each data point records the input conditions. and the corresponding output results Then, using this dataset as training samples, a lightweight neural network (such as a multilayer perceptron, MLP) is trained. The input layer of this network corresponds to the four input features mentioned above, and the output layer corresponds to the three prediction metrics. After sufficient training and validation, this surrogate model can predict with high accuracy, within milliseconds, the indoor temperature, glare, and illuminance of the work surface after performing a certain transmittance action, based on given input conditions.

[0051] In this embodiment, the agent model plays two key roles during the online control phase. First, it assists in state construction. The system's current sensors can only provide the current environmental state. The agent model can be used to predict how the environmental state will evolve in the next few minutes if a candidate action is performed. This provides the agent with more forward-looking state information, helping to make better decisions. Second, it assists in reward calculation. Online rewards... Estimating energy consumption and comfort levels requires knowing the results after an action is performed. Because there is a delay between the execution of an action and the actual change reflected by sensors, directly using sensor data to calculate rewards may be untimely or inaccurate. Using surrogate models, the potential changes in environmental indicators that an action might cause can be predicted immediately after the action is performed. This allows for a quick and lag-free estimation of the corresponding energy consumption. and comfort penalty This makes the calculation of reward signals more real-time and accurate, thereby accelerating the online learning process.

[0052] Third preferred embodiment Based on the first or second preferred embodiment, an implicit preference learning step based on multimodal awareness is also included: Deploy non-contact biosignal sensors and behavior sensing devices to collect users' physiological response data and behavioral pattern data in real time; A multimodal data fusion network is constructed to spatiotemporally align and fuse physiological response data, behavioral pattern data, and environmental perception data to generate a multidimensional user state feature vector. Based on multidimensional user state feature vectors, a comfort state classification model is trained. The classification model can divide the user state into multiple discrete comfort levels. Based on the output of the comfort state classification model, an implicit feedback reward function is constructed. When the classification model detects that the user's discomfort level has increased, the system strategy provides a negative reward, even if the user does not make an explicit manual intervention at this time. Implicit feedback rewards are weighted and integrated with existing explicit intervention rewards to form a complete personalized learning signal.

[0053] In this embodiment, to overcome the limitations of relying solely on explicit manual intervention (users may feel uncomfortable but are too lazy to adjust), the system introduces implicit preference learning. This requires deploying additional sensors to non-invasively sense the user's state. For example, non-contact biosignal sensors can employ millimeter-wave radar to extract heart rate and respiratory rate by analyzing micro-movements in the human chest cavity. Behavioral sensing devices can include piezoelectric film sensors (placed under a chair or mouse pad to sense posture adjustments and activity frequency) and infrared thermal imagers (to sense facial temperature distribution). These devices continuously collect physiological response data (such as elevated heart rate variability (HRV) potentially indicating heat stress or visual discomfort) and behavioral pattern data (such as frequent posture adjustments and leg shaking potentially indicating discomfort) while protecting privacy.

[0054] In this embodiment, the multimodal data fusion network is a specially designed machine learning model tasked with effectively integrating data from different sources, frequencies, and scales. First, data from biosignals, behavioral devices, and environmental sensors need to be timestamped to ensure that the analysis focuses on data from the same time or period. Then, feature extraction is performed on each modality of data (e.g., extracting time-domain features such as mean and variance, and frequency-domain features such as the low-frequency to high-frequency power ratio (LF / HF) from heart rate signals). Finally, a fusion layer (e.g., feature concatenation followed by a fully connected layer, or using an attention mechanism) is used to fuse these heterogeneous features into a unified multidimensional user state feature vector. This vector comprehensively represents the user's physiological and behavioral responses under specific environmental conditions.

[0055] In this embodiment, to identify user comfort levels from fused features, a comfort state classification model needs to be trained. This is a supervised classifier (such as a Support Vector Machine (SVM), Random Forest, or Neural Network classifier). The training data needs to contain a large number of samples labeled with the user's true comfort level. These labels can be obtained through periodic user surveys (such as short questionnaires pushed via a mobile app) or by combining explicit user intervention behaviors (assuming the user is necessarily uncomfortable at the time of intervention). The goal of the model learning is to map the input multidimensional user state feature vector to several discrete comfort levels, such as "comfortable," "mild discomfort," "moderate discomfort," and "severe discomfort." After sufficient training, the model can accurately infer the user's current subjective comfort state based solely on sensor data.

[0056] In this embodiment, an implicit feedback reward function can be constructed based on the output of the classification model. For example, it can be designed as ,in The level of inappropriateness (0, 1, 2, 3...) output by the classification model. The coefficient is used when the system detects that the user is in an uncomfortable state. Even without manual intervention, it will produce a negative effect. Finally, implicit feedback rewards are "weighted and fused" with the original explicit intervention-driven rewards to form a total personalized learning signal: .in It is the weight of implicit feedback. This allows the system to learn not only from what the user "says" (manual settings), but also from how the user "reacts physically," achieving more nuanced, proactive, and human-centered personalized control.

[0057] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.

Claims

1. A personalized control method for electrochromic glass based on transfer reinforcement learning, characterized in that, include: S1. Use building performance simulation software to generate indoor light and heat environment simulation datasets covering various working conditions; S2. Construct a building environment simulation model, and use the simulation data to train a general control policy network for a reinforcement learning agent to obtain a basic general model. S3. Deploy the basic general model into the actual electrochromic glass control system, output suggested control actions based on the current environmental conditions, and monitor in real time whether the user has any manual intervention behavior. S4. Calculate immediate rewards using a dynamic hybrid reward function based on user intervention. S5. Based on immediate rewards, current environmental state, and actual actions performed, the policy network of the reinforcement learning agent is fine-tuned to achieve transfer learning to the personalized preferences of specific users.

2. A personalized control method for electrochromic glass based on transfer reinforcement learning according to claim 1, characterized in that, In step S2, when training the general control policy network, its state space includes the outdoor temperature. Solar radiation intensity Time characteristics Current indoor temperature and illuminance Its action space is the set of discrete transmittance levels of electrochromic glass, and its reward function is... for: in, and These are the energy consumption for cooling and lighting, respectively. The penalty term is based on a general comfort model, including the PMV model or the glare index model. These are the weighting coefficients.

3. A personalized control method for electrochromic glass based on transfer reinforcement learning according to claim 1, characterized in that, In step S4, when the user intervenes, a punitive reward is applied. The calculation method is as follows: in, The penalty coefficient is... Recommend light transmittance for the system With user-defined light transmittance The absolute value of the difference.

4. A personalized control method for electrochromic glass based on transfer reinforcement learning according to claim 1, characterized in that, A lightweight neural network surrogate model was trained by generating an indoor light and heat environment dataset covering various working conditions using building performance simulation software. Used for real-time prediction of indoor environmental indicators under different actions: in, To predict indoor temperature, To predict the probability of glare, To predict the illuminance of the working face, For glass light transmittance, For outdoor temperature, For solar radiation intensity, It is a time-related feature.

5. A personalized control method for electrochromic glass based on transfer reinforcement learning according to claim 1, characterized in that, The policy update in step S5 uses either the near-end policy optimization algorithm or the deep Q-network algorithm.

6. A personalized control method for electrochromic glass based on transfer reinforcement learning according to claim 1, characterized in that, In step S3, after outputting the suggested action, the system will wait for a preset time window. If user intervention is detected within this window, it will be determined as an intervention event.

7. A personalized control method for electrochromic glass based on transfer reinforcement learning according to claim 1, characterized in that, It also includes implicit preference learning steps based on multimodal awareness: Deploy non-contact biosignal sensors and behavior sensing devices to collect users' physiological response data and behavioral pattern data in real time; A multimodal data fusion network is constructed to spatiotemporally align and fuse physiological response data, behavioral pattern data, and environmental perception data to generate a multidimensional user state feature vector. Based on multidimensional user state feature vectors, a comfort state classification model is trained. The classification model can divide the user state into multiple discrete comfort levels. Based on the output of the comfort state classification model, an implicit feedback reward function is constructed. When the classification model detects that the user's discomfort level has increased, the system strategy provides a negative reward, even if the user does not make an explicit manual intervention at this time. Implicit feedback rewards are weighted and integrated with existing explicit intervention rewards to form a complete personalized learning signal.

8. A personalized control system for electrochromic glass based on transfer reinforcement learning, characterized in that, include: Environmental sensing module: used to collect solar radiation and outdoor temperature through an outdoor weather station, and to collect indoor illuminance, indoor temperature and the status of people indoors through a multi-sensor node installed indoors; User interaction module: used to provide users with a wall control panel or mobile APP to manually and forcibly adjust the light transmittance of the glass; Central control module: This is an embedded computer or cloud server that runs reinforcement learning algorithms. It is responsible for receiving sensor data, calculating the optimal action, and issuing instructions. Execution module: includes electrochromic glass and its driver, used to receive light transmittance commands and change the glass state.