Multi-sensor fusion and Q-learning optimization control method for intelligent sunshade system

The intelligent shading system, which integrates multi-sensor fusion and Q-learning algorithms, can perceive environmental data in real time and perform multi-objective optimization. This solves the problems of lagging control strategies and integration difficulties in shading systems, and enables the shading system to adapt and respond quickly in changing environments, thereby improving building energy efficiency.

CN121918399APending Publication Date: 2026-04-24INTELLIGENT TECH CO LTD OF CHINESE CONSTR THIRD ENG BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTELLIGENT TECH CO LTD OF CHINESE CONSTR THIRD ENG BUREAU
Filing Date
2025-12-30
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing shading systems have single and outdated control strategies, making it difficult to respond to environmental changes in real time. They lack multi-objective collaborative optimization, making it difficult to adapt to personalized needs. System integration and coordination are difficult, and they cannot achieve building-level global energy efficiency optimization.

Method used

By using multi-sensor fusion and an improved Q-learning algorithm, environmental data is perceived in real time, a multi-objective reward function is designed, and edge computing is used to achieve dynamic multi-objective optimization of shading devices. Furthermore, the system is linked with other building subsystems and a light-heat coupling model is used for forward-looking control.

Benefits of technology

It achieves multi-objective dynamic optimization of the shading system under changing environments, has adaptive and rapid response capabilities, reduces air conditioning energy consumption, and improves the overall energy efficiency of the building.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121918399A_ABST
    Figure CN121918399A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of building energy conservation and intelligent control, and discloses a multi-sensor fusion and Q-learning optimization control method for an intelligent sunshade system, which comprises the following steps: acquiring and fusing indoor and outdoor illumination, temperature, wind speed and personnel position data in real time through a multi-sensor network, and constructing an environment state vector; the state vector is input into a pre-trained improved Q-learning model, and the model outputs an optimal sunshade control action based on a multi-target reward function integrating energy saving, visual comfort and thermal comfort; a control instruction is executed through the edge intelligent gateway, and sunshade equipment with different orientations and the linkage air conditioning system are coordinated; and in combination with a building light-heat coupling model and short-time weather prediction data, performing prospective optimization on a control decision. According to the invention, the self-adaptive and multi-target cooperative control of the sun-shading system is realized, and the energy-saving efficiency and the indoor environment comfort are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of building energy conservation and intelligent control technology, specifically relating to a multi-sensor fusion and Q-learning optimization control method for an intelligent shading system. Background Technology

[0002] In modern buildings, intelligent shading systems play a crucial role in regulating indoor lighting, reducing air conditioning energy consumption, and improving visual and thermal comfort. However, existing shading control technologies mainly suffer from the following problems: 1. Single and lagging control strategy: Most systems use control based on fixed schedules or single light thresholds, which cannot respond to dynamic changes in environmental factors such as light, temperature and wind in real time and accurately, resulting in lag in regulation.

[0003] 2. Lack of multi-objective collaborative optimization: Traditional methods are difficult to achieve an effective balance among multiple conflicting objectives such as energy saving (reducing air conditioning load), visual comfort (avoiding glare and ensuring uniform illumination) and thermal comfort.

[0004] 3. Difficulty in system integration and coordination: Shading devices often operate independently and lack effective linkage with other building subsystems such as air conditioning and lighting, making it difficult to achieve building-level global energy efficiency optimization.

[0005] 4. Difficulty in adapting to personalized and complex scenarios: Fixed-rule control strategies cannot learn and adapt to the personalized needs of different building orientations, different functional areas (such as conference rooms and office areas), and complex and ever-changing external environments.

[0006] Therefore, there is an urgent need for an optimization method for shading systems that can perceive the environment in real time, make intelligent decisions, and coordinate control. Summary of the Invention

[0007] This invention aims to address the shortcomings of the existing technology by providing a multi-sensor fusion and Q-learning optimization control method for intelligent shading systems. This method achieves dynamic multi-objective optimization of the shading angle by fusing multi-dimensional environmental perception data and utilizing an improved reinforcement learning algorithm for online or offline learning. Furthermore, it achieves rapid response and device collaboration through an edge computing architecture.

[0008] This invention provides a multi-sensor fusion and Q-learning optimization control method for an intelligent shading system, specifically including the following steps: Step 1: Real-time perception and fusion of multi-source environmental data; Real-time collection of environmental data is achieved through a multi-sensor network deployed indoors and outdoors, and the environmental data is fused to construct a state vector S_t representing the current environmental state; Step 2: Decision on the optimal shading action based on the improved Q-learning algorithm; The state vector S_t is input into the pre-trained Q-learning model, which outputs the optimal shading control action A_t in the current state based on the set multi-objective reward function R(S_t, A_t); wherein, the multi-objective reward function R(S_t, A_t) integrates at least the energy-saving reward R_energy, the visual comfort reward R_visual, and the thermal comfort reward R_thermal; Step 3: Perform control and device collaboration through the edge gateway; convert the optimal shading control action A_t into a control command through the edge smart gateway, and send it to the corresponding shading device actuator to drive the shading device to adjust to the target state. Step 4: Optimization based on light-heat coupling model and prediction; Based on the light-heat coupling characteristic model of the building, evaluate the impact of shading actions on the indoor thermal and light environment, and optimize subsequent control decisions based on the evaluation results or combined with environmental prediction data.

[0009] A computer device includes at least: one or more processors; and a memory storing one or more computer programs; wherein the processors invoke the computer programs to implement the steps of the multi-sensor fusion and Q-learning optimization control method of the intelligent shading system.

[0010] A computer storage device stores a computer program that is invoked by a processor to implement the steps of the multi-sensor fusion and Q-learning optimization control method of the intelligent shading system.

[0011] The technical solution provided by this invention has the following beneficial effects: 1. Multi-objective dynamic optimization: By designing a reward function that comprehensively considers energy saving, visual comfort and thermal comfort, the reinforcement learning algorithm can automatically find the optimal control strategy that takes into account multiple objectives in a variable environment.

[0012] 2. Foresight and Adaptability: The Q-learning algorithm can learn through continuous interaction with the environment, constantly optimize its strategy, and thus adapt to seasonal changes, sudden weather changes and the characteristics of the building itself, possessing long-term evolution capabilities.

[0013] 3. Rapid distributed control: Relying on the edge intelligent gateway, local processing of sensor data and rapid issuance of control commands are realized, reducing system response latency and facilitating distributed collaboration of multiple shading devices.

[0014] 4. System Linkage and Integration: By constructing a light-heat coupling model and communicating with the building energy management system (BMS), the shading system can be intelligently linked with subsystems such as air conditioning and lighting to improve the overall energy efficiency of the building. Attached Figure Description

[0015] The present invention will be further described below with reference to the accompanying drawings and examples. In the accompanying drawings: Figure 1 This is a schematic diagram of the overall process of a multi-sensor fusion and Q-learning optimization control method for an intelligent sunshade system according to the present invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention.

[0017] Example 1 Please refer to Figure 1 This invention provides a multi-sensor fusion and Q-learning optimization control method for an intelligent shading system, the main steps of which are as follows: Step 1: Real-time perception and fusion of multi-source environmental data; Real-time collection of environmental data is achieved through a multi-sensor network deployed indoors and outdoors, and the environmental data is fused to construct a state vector S_t representing the current environmental state; Step 2: Decision on the optimal shading action based on the improved Q-learning algorithm; The state vector S_t is input into the pre-trained Q-learning model, which outputs the optimal shading control action A_t in the current state based on the set multi-objective reward function R(S_t, A_t); wherein, the multi-objective reward function R(S_t, A_t) integrates at least the energy-saving reward R_energy, the visual comfort reward R_visual, and the thermal comfort reward R_thermal; Step 3: Perform control and device collaboration through the edge gateway; convert the optimal shading control action A_t into a control command through the edge smart gateway, and send it to the corresponding shading device actuator to drive the shading device to adjust to the target state. Step 4: Optimization based on light-heat coupling model and prediction; Based on the light-heat coupling characteristic model of the building, evaluate the impact of shading actions on the indoor thermal and light environment, and optimize subsequent control decisions based on the evaluation results or combined with environmental prediction data.

[0018] It should be noted that the first step specifically includes: Data including outdoor light intensity E_out, outdoor temperature T_out, and wind speed W_out are collected through outdoor sensors. Data collected through sensors deployed indoors includes at least the illuminance E_in_desk of the indoor work surface, the vertical illuminance E_in_glare used to assess glare, and personnel location information P_occ. The collected indoor and outdoor data are fused with the current time information Time to form the state vector S_t.

[0019] It should be noted that, in the second step, the training and decision-making process of the Q-learning model includes: Define a state space whose elements are the state vector S_t; Define the motion space, whose elements are the adjustable parameter combinations A_t of the shading device, including at least the blade angle θ and / or the deployment displacement H; Define the multi-objective reward function R(S_t, A_t): R(S_t, A_t) = w1 * R_energy + w2 * R_visual + w3 * R_thermal, where R_energy is the energy-saving reward determined by the solar radiation heat gain Q_solar(S_t, A_t) calculated based on the light-heat coupling model, R_visual is the visual comfort reward determined based on the indoor illuminance E_in_desk and glare assessment I_glare, R_thermal is the thermal comfort reward determined based on the deviation between the indoor temperature T_in and the set temperature T_set, and w1, w2, w3 are weighting coefficients; The Q-learning algorithm is used to update the rule Q(S_t, A_t) ← Q(S_t, A_t) + η * [R_t + γ* max_{A}(Q(S_{t+1}, A)) - Q(S_t, A_t)] for iterative learning, to obtain the optimal action value function Q or the optimal policy π, where η is the learning rate, γ is the discount factor, and R_t is the immediate reward; During the control phase, based on the current state S_t, according to Q... or π Select and output the action that maximizes the Q value as the optimal shading control action A_t.

[0020] It should be noted that the Q-learning algorithm adopts the Double Q-Learning algorithm, in which the update process uses two independent action value estimation functions Q1 and Q2 to alternately perform value evaluation and action selection; and / or, during the training process, an experience replay mechanism is used to randomly sample from the replay buffer that stores interaction experience (S_t, A_t, R_t, S_{t+1}) for batch updates.

[0021] It should be noted that the third step also includes: The edge smart gateway coordinates the control of shading devices on facades facing different directions based on the solar azimuth angle α_s information, generating differentiated control commands. After establishing a communication connection with the air conditioning system, the edge smart gateway sends a linkage adjustment signal to the air conditioning system based on the reduction in solar radiation heat gain ΔQ_solar estimated by the light-heat coupling model using the current shading action A_t.

[0022] It should be noted that the construction and calculation of the light-heat coupling characteristic model in the fourth step includes: Input building parameters, which include at least the window orientation Φ, window-to-wall ratio WWR, solar heat gain coefficient SHGC and visible light transmittance VLT of the glass, and the geometric and optical characteristics of the shading device; Calculate the solar altitude angle h_s and azimuth angle α_s based on the time and geographical location; Based on the building parameters, the sun's position, and the shading device status A_t, the solar radiation heat gain Q_solar(S_t, A_t) entering the room is dynamically calculated. The calculation formula includes at least the calculation of the direct radiation heat gain Q_dir: Q_dir = I_dir * A_win * SHGC * cos i * SC(θ), where I_dir is the normal direct radiation intensity, A_win is the window area, i is the angle of incidence of sunlight, and SC(θ) is the shading coefficient of the shading device at angle θ.

[0023] It should be noted that the fourth step, which optimizes subsequent control decisions based on environmental prediction data, specifically includes: Acquire short-term weather forecast data, which includes at least the predicted normal direct radiation intensity I_dir_pred and the predicted outdoor temperature T_out_pred for a future period of time; Based on the predicted data, construct the predicted future environmental state vector S_{t+1|t}; In the second step of the decision-making process, a single-step look-ahead method is adopted, and the evaluation criteria for action A_t are expanded to: A_t = argmax_{A} [R(S_t, A) + γ * max_{A'} Q(S_{t+1|t}, A')], where Q(·, ·) is the trained action value function.

[0024] Example 2 This embodiment provides an optimized control method for an intelligent shading system based on multi-sensor fusion and Q-learning, applied to an office building with east and west glass curtain walls. The system includes outdoor light sensors (measuring illuminance E_out from different directions), outdoor temperature sensors T_out, and wind speed sensors W_out deployed on the building's exterior; indoor light sensors E_in (measuring illuminance on work surfaces and glare index) and infrared human body sensors (monitoring personnel positions P_occ) deployed indoors; and controlled motorized blinds (actions include opening angle θ and lifting position H). All sensors and actuators are connected to a locally deployed edge intelligent gateway via wired or wireless means.

[0025] This method primarily involves an edge smart gateway executing a control loop and collaborating with the cloud for model training. Specifically, it includes the following steps: Step S101: Real-time perception and fusion of multi-source environmental data.

[0026] The edge smart gateway collects and fuses multi-sensor data in real time at a preset frequency (such as 0.5Hz).

[0027] Outdoor data: Collect external horizontal illuminance E_out_e, E_out_s, E_out_w in the east, south, and west directions, outdoor temperature T_out, and instantaneous wind speed W_out.

[0028] Indoor data: collect horizontal illuminance E_in_desk and vertical illuminance E_in_glare (used to estimate glare) on each work surface of each zone, as well as the indoor status and approximate location distribution of personnel determined by human infrared sensors P_occ.

[0029] The fused environment state vector S_t can be represented as: S_t = [E_out_e, E_out_s, E_out_w, T_out, W_out, E_in_desk, E_in_glare, P_occ, Time]. Where Time is the current time (time of day).

[0030] Step S102: Decision on the optimal shading action based on the improved Q-learning algorithm.

[0031] An improved Q-learning model is trained in advance or periodically in the cloud or locally. This model models shading control as a Markov decision process (MDP).

[0032] State Space: The environment state vector S_t constructed in step S201.

[0033] Action Space: Defined as the controllable parameters of the shading device, such as the blind angle θ (0° to 180°, 0° is fully closed, 90° is horizontal, and 180° is fully extended) and the lifting position H (0% to 100%), with the action A_t = (θ, H).

[0034] The reward function (R) is the core of the algorithm and is designed as a multi-objective weighted sum.

[0035] R(S_t, A_t) = w1 * R_energy + w2 * R_visual + w3 * R_thermal The calculation of each sub-item reward is as follows: Energy saving bonus R_energy: Related to the reduced air conditioning cooling load due to shading. It can be calculated using a light-heat coupling model, where solar radiation heat gain Q_solar is calculated as R_energy = -α * Q_solar(S_t, A_t). α is a positive coefficient; the smaller Q_solar (the more shading), the higher the bonus value (the less negative).

[0036] Visual comfort reward R_visual: R_visual = β1 * f(E_in_desk) - β2 * I_glare f(E_in_desk): Illumination comfort function. It can be defined as follows: when E_in_desk is within the comfort range [E_low, E_high] (e.g., 300-750 lux), f(·) = 1; when it is below E_low, f(·) = (E_in_desk / E_low); when it is above E_high, f(·) = (E_high / E_in_desk).

[0037] I_glare: Simplified assessment of glare index. For example, I_glare = 1 when E_in_glare > E_glare_threshold (e.g., 2000 lux) and the people sensor detects someone in the area; otherwise, I_glare = 0.

[0038] β1 and β2 are positive coefficients.

[0039] Thermal comfort bonus R_thermal: Related to the degree to which the indoor temperature deviates from the set value. R_thermal = -γ* (T_in - T_set)^2, where T_in is the indoor temperature (which can be read from the air conditioning system or estimated by the model), T_set is the set temperature (e.g., 24℃), and γ is a positive coefficient.

[0040] w1, w2, and w3 are weighting coefficients that can be adjusted according to building type and user preferences (e.g., w1=0.5, w2=0.3, w3=0.2).

[0041] Core and Improvements of the Q-learning Algorithm: The algorithm aims to learn an optimal action-value function Q*(S, A), representing the maximum expected cumulative discounted reward obtained by taking action A in state S. Its updates follow the Bellman optimality equation: Q(S_t, A_t) ← Q(S_t, A_t) + η * [R_t + γ * max_{A}(Q(S_{t+1}, A)) -Q(S_t, A_t)] Where η is the learning rate, γ is the discount factor (0<γ<1), R_t is the immediate reward, and S_{t+1} is the new state transitioned to after executing action A_t.

[0042] Improvement measures: Double Q-Learning: This method uses two Q-tables, Q1 and Q2, which are alternately used for action selection and value evaluation during updates to mitigate overestimation. The update formula is: Choose to update Q1 with a probability of 0.5: A* = argmax_A Q1(S_{t+1}, A); Q1(S_t, A_t) ← Q1(S_t, A_t) + η * [R_t + γ * Q2(S_{t+1}, A*) - Q1(S_t, A_t)]; Otherwise update Q2: A* = argmax_A Q2(S_{t+1}, A); Q2(S_t, A_t) ← Q2(S_t, A_t)+ η * [R_t + γ * Q1(S_{t+1}, A*) - Q2(S_t, A_t)]; Experience Replay: Experience tuples (S_t, A_t, R_t, S_{t+1}) generated from the agent's interaction with the environment are stored in a fixed-size replay buffer. During training, a batch of experiences is randomly sampled for batch updates, breaking the correlation between data and improving learning stability and sample efficiency.

[0043] After training, the trained Q-value table or neural network model is deployed on the edge smart gateway. During control, the gateway selects the action with the largest Q-value based on the current state S_t: A_t = argmax_{A} Q(S_t, A).

[0044] Step S103: Perform control and device coordination through the edge gateway.

[0045] The edge smart gateway converts the optimal action A_t into specific control commands (such as motor pulse signals) and sends them to the motorized venetian blind actuators in the corresponding areas to drive them to adjust to the target angle and position.

[0046] Meanwhile, the gateway is responsible for coordinating multiple shading devices. For example, based on the real-time calculated solar azimuth angle, it coordinates the different operation strategies of the venetian blinds on the east and west facades (prioritizing protection against direct sunlight in the morning on the east facade and focusing on protection in the afternoon on the west facade).

[0047] In addition, the edge gateway interacts with the building's air conditioning system via standard protocols (such as BACnet / IP). When shading effectively reduces solar radiation heat gain Q_solar, the gateway sends a "load reduction" signal to the air conditioning system, which can then adjust the supply air temperature setpoint or reduce the fan speed accordingly to achieve coordinated energy saving.

[0048] Step S104: Optimization based on the light-heat coupling model and prediction.

[0049] To improve the forward-looking nature of control, a simplified optical-thermal coupling model is maintained at the edge gateway or in the cloud.

[0050] Model inputs: building orientation Φ (south is 0°, east is positive), window-to-wall ratio WWR, solar heat gain coefficient SHGC and visible light transmittance VLT of glass, geometric dimensions of shading louvers (blade width w, spacing p, angle θ) and their optical properties (reflectivity ρ, transmittance τ).

[0051] Core computing module: Calculate the sun's position: Based on the current date, time, and the geographical latitude and longitude of the building's location, calculate the sun's altitude angle h_s and azimuth angle α_s.

[0052] Calculation of heat gain from direct radiation: Calculate the angle of incidence of sunlight on the window i: cos i = sin h_s * cos β + cos h_s * sin β * cos(α_s - Φ), where β is the window tilt angle (usually 90°).

[0053] The shading coefficient SC(θ) of the venetian system can be calculated by ray tracing or empirical formula. It represents the proportion of direct solar radiation blocked by the venetian system at angle θ (0 is complete shading, 1 is no shading).

[0054] Calculate the heat gain from direct solar radiation entering the room, Q_dir: Q_dir = I_dir * A_win * SHGC *cos i * SC(θ), where I_dir is the normal direct radiation intensity (which can be obtained from sensors or meteorological models), and A_win is the window area.

[0055] Heat gain from diffuse radiation and daylighting calculation: Similarly, the heat gain Q_diff from diffuse radiation from the sky and reflected radiation from the ground entering the room can be calculated. Natural daylight illuminance E_daylight can also be calculated using a similar principle, replacing SHGC with VLT.

[0056] Total solar radiation heat gain: Q_solar = Q_dir + Q_diff. This value is used for the reward calculation in step S202 and the air conditioning linkage in step S203.

[0057] Combined with predictive optimization: The edge gateway accesses short-term weather forecast data (such as the predicted normal direct radiation intensity I_dir_pred and outdoor temperature T_out_pred for the next 2 hours). When making decisions, the predicted future environmental state S_{t+1|t} (calculated based on the predictive data) can be taken into consideration. A simplified approach is to use a one-step lookahead, where the evaluation of action A_t is based not only on the immediate reward R_t, but also partly on the expected value of the optimal action in the predicted state S_{t+1|t}, i.e.: A_t = argmax_{A} [R(S_t, A) + γ * max_{A'} Q(S_{t+1|t}, A')].

[0058] Practical applications show that, compared with traditional timing or threshold control methods, this invention can further reduce air conditioning energy consumption caused by the building envelope by about 10%-25% while ensuring indoor environmental comfort.

[0059] Example 3 A computer device includes at least: one or more processors; and a memory storing one or more computer programs; wherein the processors invoke the computer programs to implement the steps of the multi-sensor fusion and Q-learning optimization control method of the intelligent shading system.

[0060] A computer storage device stores a computer program that is invoked by a processor to implement the steps of the multi-sensor fusion and Q-learning optimization control method of the intelligent shading system.

[0061] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementations described. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can better understand and utilize the present invention.

Claims

1. A multi-sensor fusion and Q-learning optimization control method for an intelligent shading system, characterized in that, Includes the following steps: Step 1: Real-time perception and fusion of multi-source environmental data; Real-time collection of environmental data is achieved through a multi-sensor network deployed indoors and outdoors, and the environmental data is fused to construct a state vector S_t representing the current environmental state; Step 2: Decision on the optimal shading action based on the improved Q-learning algorithm; The state vector S_t is input into the pre-trained Q-learning model, which outputs the optimal shading control action A_t in the current state based on the set multi-objective reward function R(S_t, A_t); wherein, the multi-objective reward function R(S_t, A_t) integrates at least the energy-saving reward R_energy, the visual comfort reward R_visual, and the thermal comfort reward R_thermal; Step 3: Perform control and device collaboration through the edge gateway; convert the optimal shading control action A_t into a control command through the edge smart gateway, and send it to the corresponding shading device actuator to drive the shading device to adjust to the target state. Step 4: Optimization based on light-heat coupling model and prediction; Based on the light-heat coupling characteristic model of the building, evaluate the impact of shading actions on the indoor thermal and light environment, and optimize subsequent control decisions based on the evaluation results or combined with environmental prediction data.

2. The multi-sensor fusion and Q-learning optimization control method for the intelligent shading system according to claim 1, characterized in that, The first step specifically includes: Data including outdoor light intensity E_out, outdoor temperature T_out, and wind speed W_out are collected through outdoor sensors. Data collected through sensors deployed indoors includes at least the illuminance E_in_desk of the indoor work surface, the vertical illuminance E_in_glare used to assess glare, and personnel location information P_occ. The collected indoor and outdoor data are fused with the current time information Time to form the state vector S_t.

3. The multi-sensor fusion and Q-learning optimization control method for the intelligent shading system according to claim 1, characterized in that, In the second step, the training and decision-making process of the Q-learning model includes: Define a state space whose elements are the state vector S_t; Define the motion space, whose elements are the adjustable parameter combinations A_t of the shading device, including at least the blade angle θ and / or the deployment displacement H; Define the multi-objective reward function R(S_t, A_t): R(S_t, A_t) = w1 * R_energy + w2 * R_visual + w3 * R_thermal, where R_energy is the energy-saving reward determined by the solar radiation heat gain Q_solar(S_t, A_t) calculated based on the light-heat coupling model, R_visual is the visual comfort reward determined based on the indoor illuminance E_in_desk and glare assessment I_glare, R_thermal is the thermal comfort reward determined based on the deviation between the indoor temperature T_in and the set temperature T_set, and w1, w2, w3 are weighting coefficients; The Q-learning algorithm is used to update the rule Q(S_t, A_t) ← Q(S_t, A_t) + η * [R_t + γ *max_{A}(Q(S_{t+1}, A)) - Q(S_t, A_t)] for iterative learning, to obtain the optimal action value function Q or the optimal policy π, where η is the learning rate, γ is the discount factor, and R_t is the immediate reward; During the control phase, based on the current state S_t, according to Q... or π Select and output the action that maximizes the Q value as the optimal shading control action A_t.

4. The multi-sensor fusion and Q-learning optimization control method for the intelligent shading system according to claim 3, characterized in that, The Q-learning algorithm employs a double Q-learning algorithm, in which the update process uses two independent action value estimation functions, Q1 and Q2, to alternately evaluate value and select actions; and / or, during the training process, an experience replay mechanism is used to randomly sample from the replay buffer storing interaction experience (S_t, A_t, R_t, S_{t+1}) for batch updates.

5. The multi-sensor fusion and Q-learning optimization control method for the intelligent shading system according to claim 1, characterized in that, The third step also includes: The edge smart gateway coordinates the control of shading devices on facades facing different directions based on the solar azimuth angle α_s information, generating differentiated control commands. After establishing a communication connection with the air conditioning system, the edge smart gateway sends a linkage adjustment signal to the air conditioning system based on the reduction in solar radiation heat gain ΔQ_solar estimated by the light-heat coupling model using the current shading action A_t.

6. The multi-sensor fusion and Q-learning optimization control method for the intelligent shading system as described in claim 1, characterized in that, In the fourth step, the construction and calculation of the optical-thermal coupling characteristic model includes: Input building parameters, which include at least the window orientation Φ, window-to-wall ratio WWR, solar heat gain coefficient SHGC and visible light transmittance VLT of the glass, and the geometric and optical characteristics of the shading device; Calculate the solar altitude angle h_s and azimuth angle α_s based on the time and geographical location; Based on the building parameters, the sun's position, and the shading device status A_t, the solar radiation heat gain Q_solar(S_t, A_t) entering the room is dynamically calculated. The calculation formula includes at least the calculation of the direct radiation heat gain Q_dir: Q_dir = I_dir *A_win * SHGC * cos i * SC(θ), where I_dir is the normal direct radiation intensity, A_win is the window area, i is the angle of incidence of sunlight, and SC(θ) is the shading coefficient of the shading device at angle θ.

7. The multi-sensor fusion and Q-learning optimization control method for the intelligent shading system as described in claim 1 or 6, characterized in that, In the fourth step, the subsequent control decisions are optimized based on environmental prediction data, specifically including: Acquire short-term weather forecast data, which includes at least the predicted normal direct radiation intensity I_dir_pred and the predicted outdoor temperature T_out_pred for a future period of time; Based on the predicted data, construct the predicted future environmental state vector S_{t+1|t}; In the second step of the decision-making process, a single-step look-ahead method is adopted, and the evaluation criteria for action A_t are expanded to: A_t = argmax_{A} [R(S_t, A) + γ * max_{A'} Q(S_{t+1|t}, A')], where Q(·, ·) is the trained action value function.

8. A computer device, characterized in that, It includes at least: one or more processors; a memory storing one or more computer programs; wherein the processor calls the computer programs to implement: the steps of the multi-sensor fusion and Q-learning optimization control method of the intelligent shading system according to any one of claims 1-7.

9. A computer storage device, characterized in that, A computer program is stored, which is invoked by a processor to implement the steps of the multi-sensor fusion and Q-learning optimization control method of the intelligent shading system according to any one of claims 1-7.