Building Control Methods and Systems Based on Multi-Source Heterogeneous Data and Pareto Decision
By constructing a multidimensional heterogeneous state space and a non-standardized multi-objective reinforcement learning model, a Pareto optimal instruction set is generated, which solves the problem of insufficient multi-source heterogeneous data processing capability of building control devices. It realizes the synergistic optimization of building energy conservation, indoor comfort and plant health, and has forward-looking and robust features, reducing energy consumption and ensuring system stability and safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-07
AI Technical Summary
Existing building control devices cannot effectively process multi-source heterogeneous data, resulting in a single control strategy that is susceptible to sensor failures, which may lead to the output of incorrect commands, pose a risk of mechanical damage or falling, and cannot adapt to dynamic changes in day and night, seasons and weather conditions, thus increasing energy consumption.
By constructing a multidimensional heterogeneous state space, collecting data from multiple fields in real time, using a non-standardized multi-objective reinforcement learning model for decision inference, generating a Pareto optimal instruction set, and combining preference arbitration and security protection mechanisms, the greening device is driven to adjust its angle. At the same time, a data integrity guarantee mechanism and a photovoltaic power generation module are introduced.
It achieves multi-objective synergistic optimization among building energy conservation, indoor comfort and plant health, and is forward-looking and robust. It can cope with complex environmental changes, reduce energy consumption and ensure system stability and safety.
Smart Images

Figure CN121523064B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent building and automation control technology, and in particular to a building control method and system based on multi-source heterogeneous data and Pareto decision-making. Background Technology
[0002] With the acceleration of global urbanization, the urban heat island effect is becoming increasingly pronounced. Its main contributing factors include artificial heat sources and building materials. Artificial heat sources include the heat generated by transportation, industry, and construction activities, leading to increased urban temperatures. The heat island effect from building materials stems primarily from the high heat capacity of materials like concrete and asphalt, which rapidly absorb and store heat, resulting in urban surface temperatures exceeding those of the natural earth's surface. To mitigate the urban heat island effect and improve the ecological benefits of buildings, vertical greening systems (VGS), commonly known as "green walls," can significantly increase urban green space and reduce the impact of the heat island effect through cooling, thus being widely considered an effective means of mitigating the urban heat island effect. However, while traditional green walls possess certain heat insulation and transpiration cooling functions, their fixed physical form cannot adapt to dynamic changes in day and night, seasons, and weather conditions. This means they may block necessary natural light on cloudy or rainy days, or block beneficial solar radiation in winter. Therefore, improving comfort requires measures such as increased lighting and higher heating power consumption of indoor air conditioning equipment, which actually increases building energy consumption and is detrimental to energy conservation and environmental protection.
[0003] To address these issues, building control devices integrating mechanical drive mechanisms have emerged. These devices allow external building modules carrying vegetation to rotate or move around a specific axis, similar to building shading louvers, aiming to dynamically adjust the building's light and heat environment by changing the angle. However, most current building control devices are reactive, relying solely on single-dimensional real-time sensors such as light intensity or air temperature sensors, and their control strategies are too simplistic to handle diverse data sources. Furthermore, when sensors experience data drift due to aging or communication failures, the building control device may even output erroneous commands, leading to mechanical structural damage or the risk of collapse. Summary of the Invention
[0004] The purpose of this invention is to provide a building control method and system based on multi-source heterogeneous data and Pareto decision-making, so as to solve the problem of insufficient multi-source heterogeneous data fusion capability of existing building control devices.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A building control method based on multi-source heterogeneous data and Pareto decision-making, the method comprising:
[0007] Step 1: Construct a multidimensional heterogeneous state space and collect data in real time;
[0008] By deploying a distributed sensor network within building envelopes, green spaces, and the external environment, heterogeneous data from multiple fields are collected in real time. This data is then combined with third-party meteorological service data to construct a real-time... Multidimensional state vector The state vector Defined as ;
[0009] in, This represents a building thermal state subvector, which includes the heat flux density through the building envelope and indoor thermal environment parameters. This represents a subvector representing the physiological state of organisms, containing soil and light parameters that reflect the vitality of vegetation; This represents an environmental state subvector, containing real-time micrometeorological data and future time windows. The predicted meteorological sequence within.
[0010] Step 2: Decision inference based on non-standardized multi-objective reinforcement learning model;
[0011] The multidimensional state vector The input is fed into a pre-trained non-standardized quantized multi-objective optimization model. This model is built upon a Markov decision process (MDP), and its tuples are defined as follows: .
[0012] In this process, the model does not output a single scalar value, but instead maps a non-dominated action-value vector set through a multi-objective Q-learning algorithm. For the action space... Each candidate angle action The model calculates its corresponding multidimensional Q-vector. This vector represents the state. Execute action Subsequently, the expected cumulative discount returns across multiple target dimensions, including building energy efficiency, comfort, and plant health:
[0013]
[0014] In the formula, As a discount factor, This is a vectorized reward function.
[0015] Step 3: Generate Pareto optimal instruction set and preference arbitration;
[0016] Based on the calculated Q-vector, the Pareto optimal front is selected using a non-dominated sorting algorithm to generate a set of non-dominated instructions. .
[0017] Subsequently, the real-time preference weight vector set by the user or system is received. (corresponding to energy saving, comfort, and plant weight respectively, and) Through scalarization function from Select the optimal action :
[0018]
[0019] Step 4: Implement drive control and safety protection mechanisms;
[0020] Select the optimal action The signal is converted into a pulse width modulation (PWM) signal or a bus command to drive the mechanical actuators of the greening device to adjust to the target angle. Simultaneously, the ambient wind speed is monitored in real time. When detected Exceeding the preset security threshold At this time, the highest priority safety interrupt is triggered, forcibly overriding the optimization command and resetting the device to the minimum wind resistance angle. .
[0021] Preferably, the vectorized reward function The construction specifically includes:
[0022] Define time reward vector .
[0023] Energy Awards Instantaneous heat flux collected by a heat flux sensor (unit: To minimize heat transfer through the building envelope (heat insulation in summer, heat preservation in winter), the following penalty function is defined:
[0024]
[0025] in, This is the normalization coefficient.
[0026] Comfort Bonus : Average evaluation index based on indoor forecasting Build. When Deviating from the preset comfort zone Punishment will be imposed at that time:
[0027]
[0028] in, This is a comfort penalty factor.
[0029] Vegetation health reward items Combined with photosynthetically active radiation intensity and soil moisture content Construct a segmented reward function:
[0030]
[0031] in, This is the optimal light range for plants. For light compensation point, The critical moisture content, The wilting coefficient, It is a positive weighting constant.
[0032] Preferably, the training and updating of the non-standard quantization multi-objective optimization model adopts the following mechanism:
[0033] A vector iterative update rule based on the Bellman equation is adopted. For the state... and actions , The formula for updating a vector is:
[0034]
[0035] in, For learning rate, Operators are used to start from the next state All possibilities From the vector set, representative vectors are selected based on the current preference distribution or hypervolume index to address the propagation problem of multi-objective value. The training process employs a model-in-the-loop simulation architecture, combining the EnergyPlus building energy consumption simulation engine and reinforcement learning agent via the BCVTB interface. Offline pre-training is performed using typical meteorological year (TMY) data until… The rate of change of the average Euclidean distance of the vector set is less than the convergence threshold. .
[0036] Preferably, the acquisition of heterogeneous state data from multiple domains further includes a data integrity assurance step:
[0037] Unsupervised learning algorithms (such as Isolation Forest or Variational Autoencoder (VAE)) are used to detect anomalies in sensor array data. When a certain sensor data is detected... Reconstruction error If the threshold is exceeded, it is determined to be data drift or a fault.
[0038] Subsequently, a time series prediction model based on a Long Short-Term Memory (LSTM) network was used, utilizing historical time series data. In conjunction with other associated sensor data, interpolated values are generated. Replace fault data to ensure input state vector The integrity of.
[0039] Preferably, the greening device integrates a photovoltaic power generation module, and the vectorized reward function... Expanded to a four-dimensional vector:
[0040]
[0041] in, For photovoltaic power generation incentives, defined as , This represents the real-time output power of the photovoltaic module. When making decisions, the model will seek a Pareto optimal solution between shading energy saving and maximizing power generation (which may require tilt angle).
[0042] The present invention also provides a building control system based on multi-source heterogeneous data and Pareto decision-making, the system comprising:
[0043] Multi-domain data acquisition and preprocessing module: Used to acquire real-time data through heat flux sensors, photosynthetically active radiation (PAR) sensors, soil moisture sensors, and ultrasonic anemometers; and equipped with a data cleaning unit for normalizing heterogeneous data and interpolating outliers to generate standardized multidimensional state vectors. .
[0044] Multi-objective intelligent decision-making module: This module incorporates a high-performance edge computing unit to run a pre-trained non-standardized multi-objective optimization model. It receives... As input, a multi-objective Q-learning inference engine is used to compute the set of Q-vectors corresponding to all actions in the action space, and outputs the non-dominated optimal instruction set. This module aims to synergistically optimize building thermal performance targets and plant vitality targets.
[0045] Preference arbitration and security monitoring module: used to receive externally input preference weight vectors. ,from The final action to be performed is selected in the middle. This module runs safety monitoring logic in parallel. When the real-time wind speed or predicted wind speed exceeds a safety threshold... At that time, a high-priority reset interrupt signal is generated, overriding the intelligent decision result.
[0046] Drive control execution module: Connects to the electric linear drive of the greening device via an industrial bus (such as RS485 or CAN), converts angle commands into mechanical displacement, and provides real-time feedback on the position status of the actuator.
[0047] Compared with the prior art, the beneficial effects of the present invention are:
[0048] By employing multi-objective collaborative optimization and nonlinear trade-offs, this invention overcomes the limitations of traditional single-objective control (such as considering only shading or only lighting). It constructs a system that includes... , and By employing a vectorized reward function and utilizing a multi-objective Q-learning algorithm, the system can find a Pareto optimal solution among three often conflicting objectives: building energy conservation, indoor comfort, and plant health. For example, at midday in summer, the system can automatically balance the conflict between "completely closing for heat insulation" and "appropriately opening to prevent plant light inhibition," thereby maximizing overall benefits.
[0049] Based on precise control of biological and physiological characteristics, unlike traditional control logic that ignores plant characteristics, this invention explicitly introduces the photosynthetically active radiation (PAR) range and soil moisture constraints into the reward function. By quantifying the optimal growth zone and stress zone of plants through formulas, it ensures that the dynamic adjustment of the greening device does not come at the expense of plant vitality, thus solving the key technical problems of high plant maintenance costs and low survival rates in dynamic greening.
[0050] Data-driven foresight and robustness, achieved by fusing real-time micrometeorological data with future weather forecasts. This invention possesses the characteristics of Model Predictive Control (MPC), enabling it to proactively address impending severe weather or drastic temperature changes. Simultaneously, it introduces a data integrity guarantee mechanism based on VAE and LSTM, maintaining decision stability through data interpolation even in the event of sensor failure or data loss. Combined with a high-priority wind speed safety interruption mechanism, this significantly enhances the system's survivability and reliability in complex outdoor environments.
[0051] The system exhibits flexible preference adaptability and scalability, generating non-dominated instruction sets rather than single actions, allowing for adjustments to the preference weight vector. The control strategy can be adjusted in real time (such as seamlessly switching from "energy saving priority" mode to "plant care" mode) without retraining the model. Attached Figure Description
[0052] Figure 1 This is a schematic diagram of the building control method based on multi-source heterogeneous data and Pareto decision-making according to the present invention;
[0053] Figure 2 This is a schematic diagram illustrating the definition of the state and vectorized reward function in the multi-objective decision optimization (MDP) of this invention;
[0054] Figure 3 This is a schematic diagram illustrating the principles of multi-objective Q-learning and Pareto front selection in this invention;
[0055] Figure 4This is a flowchart of the security protection mode in this invention;
[0056] Figure 5 This is a block diagram of the building control system based on multi-source heterogeneous data and Pareto decision-making according to the present invention. Detailed Implementation
[0057] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0058] Please refer to Figure 1 As shown, in a first aspect of the present invention, a building control method based on multi-source heterogeneous data and Pareto decision-making is provided, including core steps S101 to S104:
[0059] S101. Acquire heterogeneous state data from multiple domains and generate multi-dimensional state vectors. The system collects three types of data in real time through a deployed sensor network: 1) Building state data. 1) Indoor temperature and humidity, wall heat flux; 2) Biological state data Such as soil moisture, soil temperature, electrical conductivity (EC), and crucial photosynthetically active radiation (PAR) in greening devices; 3) Environmental data For example, real-time wind speed and direction are obtained through ultrasonic anemometers. In addition, the system can also obtain future (e.g., within 6 hours) forecast weather information by accessing third-party weather forecast APIs. All data is fused into a unified multi-dimensional state vector. .
[0060] S102. Based on the multi-dimensional state vector, perform multi-objective decision optimization operations. (The state vector is...) As input, it is passed to a pre-trained non-standardized quantized multi-objective optimization model. This model is based on Markov decision processes (MDPs) and aims to learn a strategy that can synergistically optimize the two core objectives of building thermal performance and plant vitality.
[0061] S103. Generate a set of non-dominated instructions. The core of the model is a multi-objective reinforcement learning algorithm (as described in Example 2). Its output is not a single optimal action, but a set of non-dominated optimal instructions. Each action (i.e., angle command) in this set corresponds to a Pareto optimal solution, which means that without worsening one objective (such as plant health), it is impossible to improve another objective (such as building energy efficiency).
[0062] S104. Based on the preference vector, the optimal action is selected from the set of non-dominated instructions and executed by the building control device. The system selects the optimal action from the set of non-dominated instructions based on a real-time specified preference vector. (For example, during the hot summer months, users might set their preferences to...) These correspond to energy saving, comfort, and plants, respectively (i.e., a high preference for energy saving). Choose the one that best matches your current preference from the set. The optimal action (such as angle command). Finally, control signals are sent to the drive mechanism (such as an electric actuator) of the building's external attachments (such as greening devices), adjusting the device to... The corresponding angle.
[0063] Those skilled in the art will understand that this invention, through artificial intelligence models, particularly multi-objective reinforcement learning, solves the trade-off problem between the conflicting objectives of building energy conservation and plant health in dynamic greening devices. It integrates real-time and predictive data, making control decisions forward-looking and intelligent, overcoming the limitations of traditional static greening and simple threshold control.
[0064] Example 1
[0065] Please refer to Figure 2 In this embodiment, the interaction and execution logic between the non-standardized multi-objective reinforcement learning model and the building physical entity in step S102 are described in detail. This invention utilizes a multi-dimensional physical sensor array deployed in building facades, indoor environments, and plant growth media to transform environmental physical quantities into vectorized reward signals with clear physical meaning. This directly drives the mechanical actuator to produce precise angular displacement.
[0066] 1. Electrical signal conversion and mechanical loss protection rewards for building thermal conditions
[0067] This award is directly coupled to a heat flux sensor array deployed on the wall envelope and window frames. The system acquires the microvolt voltage signals from the heat flux sensors in real time and converts them into instantaneous heat flux density according to a preset electrophysical sensitivity coefficient. To achieve a balance between building energy savings and device self-consumption energy, this embodiment employs a refined physical model for the energy incentive item.
[0068] In terms of the physical energy efficiency trade-off mechanism, the system monitors the instantaneous power consumption of the drive motor during rotation through current / voltage transformers. And obtain the rated power energy consumption of the motor. Hardware fatigue and energy consumption balance formula:
[0069]
[0070] in and The weighting coefficients are used to balance building energy conservation and equipment self-consumption. This is the normalization factor, which takes the value of the historical maximum observed heat flux, and is used to map the heat flux data to... The interval is used to prevent differences in numerical magnitude from affecting the stability of gradient descent. The physical meaning of the formula lies in directly linking the physical insulation requirements of the building envelope with the electrical losses of the mechanical actuator.
[0071] The industrial edge computing unit uses this formula to identify minute fluctuations in heat flux. If the energy-saving benefits of the predicted action are less than the electrical energy consumption of the drive motor, the system will suppress unnecessary mechanical commands, thereby avoiding the risk of overheating caused by frequent motor starts and stops, and significantly reducing the physical wear of mechanical transmission bearings.
[0072] 2. Indoor microenvironment physical feedback and human thermal comfort smoothing mechanism
[0073] The system uses temperature and humidity sensors and average radiation thermometers deployed on the indoor work surface to collect real-time thermal environment signals. Combined with the preset metabolic rate and clothing thermal resistance parameters of the indoor personnel, the edge computing unit calculates the real-time thermal comfort index. To eliminate the step jitter in the mechanical drive mechanism during angle adjustment, a Gaussian kernel feedback control mechanism is adopted, using a Gaussian kernel function to generate the control gradient reward.
[0074]
[0075] In the formula, Based on indoor temperature relative humidity RH Mean radiant temperature The real-time thermal comfort index is calculated based on the set thermal resistance of clothing and the human metabolic rate. The target comfort value is usually set to 0 (i.e., thermal neutrality). The bandwidth parameter of the Gaussian function is used to control the tolerance for comfort deviations. The larger the value, the more smoothly the reward function decreases, indicating a higher tolerance for small deviations; The smaller the value, the more precise the control required; the range of values for this function is... .when Completely equal to At the initial value, the reward is 0 (maximum value); as the deviation increases, the reward value decays exponentially to -1, thus imposing a significant negative feedback penalty on the harsh thermal environment.
[0076] This function takes parameters By controlling the tolerance for comfort deviations, a continuously differentiable gradient signal is provided to the drive unit to guide the electric actuator to perform "micron-level" angle compensation fine-tuning, rather than drastic, discontinuous start-stop actions.
[0077] This formula, through the adjustment of physical gradient descent, effectively avoids instantaneous stress impact on mechanical linkages, significantly extends the service life of precision mechanical transmission systems such as lead screws and push rod bearings, and ensures that indoor occupants will not experience drastic fluctuations in perceived temperature due to sudden changes in the shading angle.
[0078] 3. Hardware sensing of plant physiological characteristics and biomechanical synergistic reward
[0079] This is the core step in achieving the synergy between bioenergy balance and mechanical motion control in this system. The system uses a photosynthetically active radiation (PAR) sensor to monitor the physical light quantum flux on the surface of plant leaves and a frequency domain reflectance (FDR) soil moisture sensor to sense changes in the dielectric constant of the growth medium.
[0080] First, define the photosynthetic efficiency factor. A variant of the Michaelis-Menten equation is used to describe the nonlinear relationship between photosynthetically active radiation (PAR) and plant growth rate:
[0081]
[0082] In the formula, The real-time photosynthetically active radiation value collected by the sensor ( ); This is the plant's maximum photosynthetic rate; This is the light compensation point; when the light intensity is below this value, the plant's respiration consumes more energy than its photosynthetic output. Turns to a negative value, coefficient This signifies a penalty for health decline caused by insufficient sunlight. This is the light saturation point; when light intensity exceeds this value, plants exhibit photoinhibition. It begins to decrease, the coefficient Indicates the severity of punishment for photoburn; It is the half-saturation constant.
[0083] Then, the system introduces soil moisture constraint factors. This factor directly maps the water absorption pressure state of the vegetation growth medium by reading the level or frequency signal returned by the moisture sensor. Soil moisture constraint factor. for:
[0084]
[0085] In the formula, Soil volumetric water content, This is the critical moisture threshold. This is the slope parameter of the Sigmoid function. This factor is... The coefficient represents the plant's ability to fully utilize sunlight when water is plentiful and when water is scarce. When the stomata close, even with sufficient light, the photosynthetic efficiency will decrease significantly.
[0086] The final biological limit alignment mechanism is achieved through a piecewise composite reward function:
[0087]
[0088] When the PAR sensor detects that the ultraviolet intensity or light flux exceeds the plant's tolerance threshold (photoinhibition zone), the reward function value drops sharply and non-linearly. This biosignal is immediately converted into a mechanical drive command, forcing the drive mechanism to rotate the device angle. The physical shadow generated by the building's external modules covers the vegetation surface, achieving direct physical negative feedback closed-loop control from the vegetation's biological survival status to the movement of the mechanical actuator.
[0089] 4. Physical mapping and signal transformation logic of the driver execution layer
[0090] After the multi-objective optimization model generates the non-dominated instruction set in step S103, the preference arbitration module executes the final physical instruction conversion. The system, based on the geometric topology of the four-bar linkage or linear actuator of the building's external greening device, converts the Pareto optimal angle instruction... Converted into physical displacement stroke of electric linear actuator .
[0091] The drive control module sends pulse width modulation (PWM) signals with a specific duty cycle to the DC motor or servo unit, driving the robotic arm to generate angular displacement. Simultaneously, position feedback signals are acquired via integrated Hall effect sensors or potentiometers. It is also connected to a PID controller to ensure the absolute positioning accuracy of mechanical movements, thereby adjusting the device to the physically optimal orientation that can synergistically optimize energy saving, comfort and biological health.
[0092] Example 2
[0093] Please refer to Figure 3 In this embodiment, the non-standardized multi-objective optimization model described in step S102 is fundamentally based on a deep mapping between a virtual reinforcement learning algorithm and a complex building physics environment. Unlike traditional single-objective control systems that simplify all performance indicators such as energy consumption, temperature, and plant growth into a single scalar through weighted summation, the agent in this system maintains a high-dimensional vectorized Q-value table, aiming to approximate the true Pareto optimal solution set in the physical action space through a non-dominated ranking mechanism.
[0094] The following are the specific implementation details of combining algorithms with physical entities:
[0095] S201, Construction of Vectorized Action-Value Function and Physical State Space
[0096] First, the system needs to establish a data structure capable of accommodating multiphysics feedback. Define the state space. With action space The action space The set of adjustable physical discrete angles directly corresponding to greening devices ,
[0097] The agent builds a multidimensional lookup table or a deep neural network approximator to represent the vectorized action-value function. For each physical state And every mechanical movement ,That The value is defined as a 3-dimensional vector (i.e. ):
[0098]
[0099] in, Mapping building heat flux and motor losses, Mapping indoor thermal environment indicators, This maps to the physiological survival rate of the vegetation. During initialization, all state-action pairs... The value is set to zero vector, waiting for feedback stimulus from physical sensors.
[0100] S202, Exploration and Physical Execution Strategies Based on Dynamic Preferences
[0101] Each step during the training phase To enable the model to understand the trade-offs between different weather conditions and user needs, the system employs a dynamic weight sampling method. Strategy.
[0102] In the random exploration phase, using probability Randomly select actions This allows the mechanical actuator to swing within a full range of angles in order to collect sensor data under extreme operating conditions.
[0103] Greedy exploitation and preference mapping: using probability Execution strategy. To traverse different physical regions of the Pareto front, the agent randomly samples a preference weight vector at each step. And satisfy By calculating the scalarized projection value
[0104]
[0105] The agent can simulate decision-making in different scenarios, such as shifting preferences during hot weather. Tilt, and in the dry season towards Tilt the model to ensure it covers the full Pareto front physical surface during training.
[0106] S203, Real-time Acquisition of Environmental Interaction and Multi-Domain Heterogeneous Rewards
[0107] When the intelligent agent drives the greening device to perform actions Afterwards, the physical environment of the building changes from state Transfer to It also provides a vectorized instant reward generated by the physical entity. The increase in wall heat conduction sensed by the heat flux sensor and the motor power consumption measured by the current sensor together constitute a negative reward. (Based on indoor temperature and humidity calculations...) The bias is transformed into a smoothed gradient penalty signal using a Gaussian kernel function. The sensors and soil moisture sensors provide feedback, reflecting whether the current angle adjustment has caused the plant to be in a state of photoinhibition or water deficit.
[0108] S204. Dynamic Maintenance and Pareto Pruning of Non-Dominated Solution Sets
[0109] Before updating the Q value, the system needs to determine the next physical state. The potential value of the next state. Due to the conflict of objectives, the optimal value of the next state is no longer a single scalar, but a non-dominated solution set.
[0110] Define vector dominance relations For two utility vectors and If satisfied and Then it is called Dominate The system iterates through the next state. Lower Action Space Construct a candidate set using the Q-vectors corresponding to all actions. And select rigorous Pareto frontiers :
[0111]
[0112] This set eliminates all dominated inferior solutions and represents the theoretically optimal expected return boundary that different operating modes (e.g., extreme energy saving mode versus plant care mode) can achieve in the future.
[0113] S205, Iteration and Physical Parameter Correction of Vectorized Bellman Equations
[0114] Based on the obtained Pareto frontier The system's current state-action pair Iterative updates are performed. To address the curse of dimensionality caused by multiple objectives, an update mechanism based on random sampling is adopted.
[0115] from The weights used in the current step are determined by the current weights. Select the best matching target vector The Bellman equation in vector form is corrected. value:
[0116]
[0117] In the formula, For learning rate, This is the discount factor. This step, through continuous iteration, gradually converges the Q-vector table and accurately describes the complex nonlinear relationship between the actions of greening devices and building energy consumption and environmental quality.
[0118] S206. Physical Generation and Storage of Non-Dominant Instruction Sets
[0119] After about After the simulation or field training of each step converges, the model will be for each physical state. Maintain a set of non-dominant action instructions Iterate through all possible mechanical angle movements. If the action produces If a vector is not dominated by any other angular motion vector, then store it in the instruction set. .
[0120]
[0121] This set includes multiple "optimal balance points," such as "ultimate shading angle," "optimal photosynthetic angle," and "overall comfort angle." This provides a solid decision-making basis for subsequent steps to make millisecond-level arbitration selections based on the user's real-time preferences (such as energy-saving priority or plant aesthetics priority set via a mobile app).
[0122] Through the detailed steps of S201 to S206 described above, this embodiment realizes an intelligent control logic that can handle high-dimensional target conflicts and does not rely on a single scalar feedback, thus ensuring the multi-objective collaborative optimization capability of the dynamic greening device in complex environments.
[0123] Example 3
[0124] Please refer to Figure 4This embodiment details how the system constructs a "circuit breaker" mechanism independent of the multi-objective intelligent decision-making module through hardware interrupt logic, sensor fusion filtering, and aerodynamic physical modeling. This mode ensures the structural safety of the greening device under extreme weather conditions (such as typhoons and strong convective gusts), preventing the risk of fatigue fracture, plastic deformation, or detachment from the building's exterior wall. The specific implementation process includes the following key steps:
[0125] S301, Real-time fusion and filtering of multi-source wind speed data.
[0126] The system monitors the environmental status in real time. The wind speed component is used. To prevent control jitter caused by instantaneous gusts or sensor noise, the system does not directly use the raw readings of the ultrasonic anemometer. Instead, a sliding time window filter is used to smooth the real-time wind speed. The time window length is defined as... (For example (seconds), smoothed real-time wind speed The calculation formula is:
[0127]
[0128] In the formula, The time decay weighting coefficient satisfies This is used to assign higher weight to data at the current moment. Simultaneously, the system obtains future data via API. Forecast gust extremes within a time window (e.g., the next hour) , as a forward-looking early warning indicator.
[0129] S302. Dynamic threshold determination based on aerodynamic model.
[0130] Safety threshold The setting is not a fixed value, but is calculated based on the structural strength limit and aerodynamic characteristics of the greening device. The aerodynamic drag experienced by the device in the wind field... It is directly proportional to the frontal area and the square of the wind speed, and its physical model can be expressed as:
[0131]
[0132] In the formula, air density, For angle The varying drag coefficient, This represents the total surface area of the device. The system is pre-set with a maximum critical load that the structure can withstand. Therefore, the logic for determining the dynamic wind speed threshold that triggers the safety mode is as follows:
[0133]
[0134] in, For the theoretical limit wind speed, To predict the safety factor (e.g., take 0.8).
[0135] S303, high-priority interrupt and minimum flow resistance reset.
[0136] once The control system immediately activates a safety interrupt. At this time, the system will forcibly suspend the intelligent decision output based on reinforcement learning in step S102, ignore the current energy saving or plant health goals, and directly send the highest priority reset command to the underlying driver.
[0137] Target security perspective This is defined as the state in which the frontal area of the device is minimized, i.e.:
[0138]
[0139] Under normal circumstances, Corresponding to (Complete Level) or (Completely closed and flush against the wall), depending on the mechanical hinge design of the device, to ensure that airflow can pass through or glide over the device surface with minimal resistance, preventing plastic deformation or detachment of the mechanical structure.
[0140] S304, Recovery mechanism based on hysteresis comparator.
[0141] To avoid frequent switching between "safe mode" and "intelligent mode" (i.e., the "ping-pong effect") near critical wind speeds, this embodiment introduces hysteresis comparator logic with a time delay. The system must simultaneously meet the following two conditions to exit safe mode:
[0142] 1. Real-time wind speed decreases to the recovery threshold Below, and (For example ;
[0143] 2. The low wind speed condition lasts longer than the maintenance time. (e.g., 5 minutes).
[0144] The Boolean expression for the recovery logic is:
[0145]
[0146] Only when Only when the condition is true will the system unlock, reactivate the multi-objective intelligent decision-making module, and resume optimized control over building energy consumption and plant status. Through this mechanism, the system possesses extremely high robustness and adaptive survivability in the face of extreme weather such as typhoons and severe convection, and also extends the service life of the precision mechanical drive system through accurate physical modeling.
[0147] Example 4
[0148] This embodiment details the specific implementation process of offline pre-training of a non-standardized quantized multi-objective optimization model based on a model-in-the-loop simulation architecture. Since directly training reinforcement learning models on real buildings involves long training cycles, high trial-and-error costs, and the risk of potential physical damage to equipment, this invention constructs a high-fidelity virtual simulation environment for policy iteration.
[0149] Specifically, the steps include the following:
[0150] S401. Construct a coupled physical model of building thermal engineering and dynamic greening devices.
[0151] In the EnergyPlus building energy consumption simulation engine, a virtual model of the target building is built based on the IDF standard. For dynamic greening devices, they are equivalent to shading components with variable optical and thermal properties.
[0152] Define the greening device at the angle Equivalent solar heat gain coefficient With equivalent thermal resistance Considering the nonlinear superposition of transpiration from plant leaves and mechanical shading, the following physical parameter calculation model is constructed:
[0153]
[0154]
[0155] In the formula, For the device at an angle Effective transmittance at the specified level; and These represent the extreme transmittance values when the device is fully closed and fully open, respectively. This is the blade overlap coefficient; The overall heat transfer coefficient of the building envelope; For the thermal resistance of the base wall; The thermal resistance of the air interlayer varies with angle; This is the equivalent thermal resistance of the plant layer.
[0156] The above equations map the mechanical movements of the dynamic greening device into changes in thermal parameters that EnergyPlus can recognize, ensuring that the simulation environment can accurately reflect the building's thermal state. .
[0157] S402. Build the BCVTB collaborative simulation communication interface.
[0158] A virtual building control test bench (BCVTB) is used as middleware to establish bidirectional data communication between a Python deep learning framework (such as PyTorch or TensorFlow) and the EnergyPlus engine. The BSD Socket protocol is used for data packet transmission, and the simulation step size is set. Define the state exchange vector for co-simulation. :
[0159]
[0160] Among them, the control command vector sent by the agent to the environment Includes the action command (i.e., angle) at the current moment. The state vector fed back to the agent by the environment Includes simulated indoor temperature Wall heat flux and environmental parameters read from meteorological files; This is the simulation synchronization flag.
[0161] To simulate sensor noise in a real-world environment, in the feedback state Inject Gaussian white noise into the middle:
[0162]
[0163] In the formula, This is the standard deviation matrix corresponding to the sensors, to enhance the robustness of the model in real-world deployments.
[0164] S403, Multi-objective Q-vector iterative training based on typical meteorological years.
[0165] The typical meteorological year (TMY3) data file for the target city is loaded as the environmental stimulus. The agent is based on multi-objective... Learn the algorithm for exploration and utilization. At each simulation step... The agent observes the state. Select Action and receive vectorized rewards. .
[0166] The vectorized Bellman optimal operator is used to update the Q-table (or neural network parameters). For the multi-objective case, the Q-value update is no longer scalar convergent, but rather a non-dominated ranking optimization of the Q-vector set. The iterative update formula for the Q-vectors is as follows:
[0167]
[0168] In the formula, The value vector of the state-action pair; The learning rate; Discount factor; A reward vector that includes energy efficiency, comfort, and plant health; A non-dominated sorting operator used to sort from the next state The Pareto front set is selected from all possible Q vectors.
[0169] To balance exploration and utilization, the following approach was adopted. strategy, and With the number of training rounds It exhibits exponential decay:
[0170]
[0171] S404, Model convergence determination and parameter fixation.
[0172] The training termination condition is set as the hypervolume index of the Q-vector set tends to stabilize. Define the first... The iteration and the Q-vector set dissimilarity in the next iteration :
[0173]
[0174] When continuous Each training epoch satisfies ( For example, a large preset convergence threshold. When the model training is complete, it is determined that the training is finished.
[0175] After training, the converged Q-vector table (for discrete state space) or the weight parameters of the deep neural network will be used. (For continuous state space) The data is exported and packaged into a binary inference file. This file is burned into the edge computing unit of the actual control system, serving as the decision-making core for real-time control, thereby enabling policy migration from virtual simulation to physical entities.
[0176] Example 5
[0177] In this embodiment, a photovoltaic (PV) power generation unit (e.g., flexible thin-film solar cell or semi-transparent photovoltaic glass) is integrated on the mechanical structure surface or the back of the blade-bearing module of the greening device. At this point, the system's control logic upgrades from three-dimensional target coordination to four-dimensional target coordination, aiming to find the optimal dynamic balance between building energy conservation, indoor comfort, plant health, and photovoltaic power generation.
[0178] The specific implementation steps and algorithm model construction are as follows:
[0179] S501. Construct a physical coupling model for photovoltaic production capacity.
[0180] In the multi-objective intelligent decision-making module, the relationship between photovoltaic output power and device adjustment angle is first established. A nonlinear coupling model between them. Real-time theoretical output power of the photovoltaic module. It depends not only on the intensity of solar radiation, but also on the effect of the device angle on the angle of solar incidence and the correction for the temperature of the solar panel. The calculation formula is as follows:
[0181]
[0182] In the formula, For at any time The device angle is The predicted value of photovoltaic power generation at that time; The overall conversion efficiency of the photovoltaic system (including inverter efficiency and line loss). The effective light-receiving area of the integrated photovoltaic module; The total radiation intensity on the horizontal surface is derived from meteorological data. Obtain; This is a function of the incident angle correction coefficient, which depends on the angle of the device. Solar altitude angle and solar azimuth This function describes how changes in the device's angle alter the direct radiation component received by the photovoltaic panel. This is the temperature power coefficient of the photovoltaic cell (usually a negative value). The operating temperature of the solar panel. The reference temperature (25°C) is the temperature under standard test conditions.
[0183] Using the above model, the agent can predict and adjust the device to any candidate angle under the current environmental conditions. The exact amount of electricity generated that can be obtained.
[0184] S502, Extended four-dimensional vectorized reward function.
[0185] In order to incorporate production capacity targets into the optimization framework, the original three-dimensional reward function was modified. Expanded to a four-dimensional vector. Time interval defined. The reward vector is as follows:
[0186]
[0187] Among them, the newly added photovoltaic power generation incentive items It is not a simple linear mapping, but rather a normalized nonlinear excitation function to ensure that its numerical magnitude corresponds to the heat flux. The gradient vanishing or dominance problem is avoided due to numerical differences between the gradient index (PMV) and the comfort index (PMV). The calculation formula is:
[0188]
[0189] In the formula, This refers to the rated peak power of the photovoltaic module. The baseline weighting coefficient for the capacity target is used to set the relative importance of the target in the Pareto frontier; This is a sensitivity adjustment factor; This is a hyperbolic tangent activation function used to smooth the reward value and limit its upper limit.
[0190] S503, augmentation of the state space.
[0191] To support accurate forecasting of photovoltaic capacity, the multi-domain data acquisition module constructs a state vector. At this time, it is necessary to augment the photovoltaic-related state sub-vectors. At this point, the total state vector is defined as:
[0192]
[0193] in, Specifically, it includes the following feature parameters: This is the open-circuit voltage of the photovoltaic array; This is the short-circuit current of the photovoltaic array; The temperature of the photovoltaic panel backsheet is collected by a patch sensor; This represents the cumulative power generation for the day.
[0194] S504, Multi-objective conflict resolution and decision execution.
[0195] During the training and inference phases, the agent faces more complex Pareto trade-offs. The system needs to handle the following typical conflict scenarios:
[0196] In the midday summer scenario, the building energy efficiency target is to tend to close the device ( To maximize shading and block solar radiation from entering the room; photovoltaic capacity target: tends to adjust the device at an angle perpendicular to the sunlight (e.g., θ). To maximize power generation, the plant health target is to allow for moderate operation to avoid scorching from strong sunlight while ensuring ventilation and heat dissipation.
[0197] At this point, the non-standardized multi-objective Q-learning model will calculate the four-dimensional Q-vector corresponding to each angle in the action space. Based on the current preference weight vector The system calculates the overall utility value. :
[0198]
[0199] in, For all possible action options The overall utility value can be increased The action that reaches its maximum and is output (i.e., the target angle). The system will select... As the optimal action to take. For example, if the current peak electricity price is high, the user sets... At higher angles, the system may sacrifice some shading effect (allowing a small amount of heat to enter) to adjust the angle to a position more conducive to power generation, using photovoltaic power generation to offset the increased cost of air conditioning energy consumption, thereby achieving optimal overall economic efficiency.
[0200] In overcast, diffused light scenarios, the photovoltaic power generation capacity is reduced, and the model will automatically lower its resolution. The actual influence in decision-making has shifted to a greater focus on the efficiency of plants in capturing diffused light (PAR) and the need for natural lighting indoors.
[0201] Through the above expansion, this system is not only a dynamic shading and vertical greening system, but also an adaptive building-integrated photovoltaic (BIPV) energy management system, achieving comprehensive performance optimization of "energy saving, production capacity, ecology and comfort".
[0202] Example 6
[0203] Please see Figure 5 In a second aspect of the invention, a building control system based on multi-source heterogeneous data and Pareto decision-making is disclosed. This system is a hardware implementation of the aforementioned control method, employing a hierarchical distributed architecture design. It mainly includes a multi-domain data acquisition and preprocessing module, a multi-objective intelligent decision-making module, a preference selection and arbitration module, and a drive control and execution module. Each module achieves data interaction and collaborative control through a high-speed industrial fieldbus and a low-power wide-area network. Specific implementation details are as follows:
[0204] 1. A multi-domain data acquisition and preprocessing module, which acts as the perception layer and is responsible for executing step S101. In terms of hardware configuration, the system integrates a heterogeneous sensor array, including but not limited to heat flux sensors installed on the building facade (for measuring...). Photosynthetically active radiation (PAR) sensor (spectral response range 400-700nm), frequency domain reflectance (FDR) soil moisture sensor, and high-precision ultrasonic anemometer.
[0205] Considering outdoor environmental noise interference, this module incorporates signal conditioning circuitry and an FPGA preprocessing unit. To address the dimensional differences in heterogeneous data, the preprocessing unit employs Z-score normalization to perform real-time normalization of the raw signal. Let's assume a certain sensor... At any moment The original sample value is Its standardized output The calculation formula is:
[0206]
[0207] in, and These represent the mean and standard deviation of the sliding window data from the sensor, respectively. To prevent tiny constants with a denominator of zero.
[0208] Furthermore, for data transmission, the module integrates a LoRaWAN communication module, employing spread spectrum modulation technology for long-distance transmission. To ensure data packet integrity, the module uses a cyclic redundancy check (CRC) mechanism, with its check polynomial set as follows: By utilizing low-power wide-area network (LPWAN) technology, the problems of difficult wiring and signal attenuation for sensors on the facades of high-rise buildings have been effectively solved.
[0209] 2. Multi-objective intelligent decision-making module: This module is the core computing hub of the system, responsible for executing steps S102 and S103. Hardware-wise, it adopts an industrial-grade edge computing gateway with a built-in high-performance embedded processor (such as an NVIDIA Jetson series or an ARM architecture SoC with NPU acceleration).
[0210] This module has a pre-trained, non-standardized multi-objective Q-learning model embedded in its memory. The inference engine does not directly output a single action, but instead computes the current state in parallel through matrix operations. Lower Action Space All candidate actions The corresponding Q-vector matrix :
[0211]
[0212] Each row represents an action in energy saving ( ), comfort ),plant( The expected value across three dimensions. Edge computing units utilize non-dominated sorting algorithms for... Perform row vector filtering to output the Pareto optimal instruction set. This process is completed at the local edge in milliseconds, without continuously relying on cloud computing power, which greatly reduces the impact of network latency on the real-time performance of control.
[0213] 3. Preference Selection and Arbitration Module: This module is responsible for executing the decision arbitration function in step S104. It receives user preference weight vectors from the host computer or mobile terminal App. Furthermore, the instruction set is optimized using a weighted approach combined with security monitoring logic.
[0214] The module internally runs a utility maximization algorithm to compute each action in the Pareto optimal set. scalar utility value :
[0215]
[0216] System selection makes The largest action serves as the initial instruction.
[0217] Simultaneously, this module runs an independent safety monitoring thread in parallel. This thread compares instantaneous wind speeds in real time. With safety threshold The security arbitration logic uses the following Boolean algebra expression to generate the final execution instruction. :
[0218]
[0219] in, The hard-coded mechanical reset angle (such as 0° or 90°, depending on the aerodynamic characteristics of the device) gives the safety interrupt signal the highest system priority, which can directly override the intelligent decision-making results.
[0220] 4. Drive control and execution module: This module is responsible for executing the physical actions in step S104. Control signals are transmitted to the distributed motor controller (MCU) via RS485 or CAN bus.
[0221] Because the mechanical structure of greening devices is usually driven by linkage mechanisms, the target angle Stroke with linear drive There are nonlinear geometric relationships between them. The module integrates a kinematics solution unit, which performs inverse kinematic transformations based on the law of cosines.
[0222]
[0223] in , The length of the link. The initial phase angle, This is the zero-bit length of the driver.
[0224] Calculated target distance The signal is fed into a PID controller to adjust the PWM (Pulse Width Modulation) duty cycle of the electric actuator. :
[0225]
[0226] in, To prevent positional deviation, the actuator uses an industrial-grade electric linear actuator with a protection rating of at least IP66, featuring overload protection and Hall position feedback to ensure accurate and stable adjustment of the greening device to the target angle even under harsh outdoor conditions.
[0227] Through the close collaboration of the above four modules, this system achieves closed-loop control from environmental perception, intelligent decision-making, safety arbitration to precise execution, effectively solving the problem of multi-objective collaborative optimization of dynamic greening devices in complex environments.
[0228] Example 7
[0229] In this embodiment, the multi-domain data acquisition and preprocessing module also integrates a data integrity assurance unit to address the problems of sensor data loss, drift, or noise interference in complex outdoor environments, thereby improving the overall robustness of the control system. The specific implementation process of this unit includes two core stages: data anomaly detection and missing value imputation. The specific technical solution is as follows:
[0230] Phase 1: Unsupervised anomaly detection based on variational autoencoder (VAE)
[0231] In order to accurately identify abnormal data in the sensor network (such as the solidification of readings of soil moisture sensors due to corrosion, and the attenuation of readings of light sensors due to dust coverage), this unit constructs a variational autoencoder (VAE) network to perform probabilistic reconstruction analysis on multidimensional heterogeneous sensor data.
[0232] 1. Data Input and Encoding: Definition The original sensor measurement vector at time is ,in This represents the total number of sensors. The VAE's encoder network will take the input data. Mapping to the latent variable space The distribution parameters, i.e., the mean vector. and standard deviation vector The encoding process is represented as a probability distribution. Its neural network calculation formula is:
[0233]
[0234]
[0235] in, and These are the weight matrix and bias term of the encoder, respectively.
[0236] 2. Resampling and Decoding: Utilizing reparameterization techniques to sample from the latent space. ,in Subsequently, the decoder network latent variables Map back to the data space to generate reconstructed data. :
[0237]
[0238] 3. Anomaly detection criterion: The training objective of VAE networks is to maximize the lower bound of evidence, and its loss function is... Defined as the sum of reconstruction error and KL divergence:
[0239]
[0240] During the detection phase, the reconstruction probability score of the data at the current moment is calculated. If a certain sensor channel Reconstruction residuals Exceeding a dynamic threshold based on historical statistical characteristics (For example If the sensor data is then determined to be... If an abnormality or malfunction occurs, it is marked as a failure state.
[0241] Second stage: Dynamic interpolation based on spatiotemporal correlation LSTM
[0242] When data anomalies or missing data are detected, the system does not discard the sample directly. Instead, it initiates a spatiotemporal correlation prediction model based on a Long Short-Term Memory (LSTM) network for imputation to ensure that the state vector input to the decision module is intact. The integrity of.
[0243] 1. Spatiotemporal Feature Construction: Constructing the input matrix of the interpolation model It not only includes the sensor's time-sliding window historical data It also integrates spatially correlated data from other sensors (such as using air hygrometer and rainfall data to help correct soil hygrometer data) and environmental prediction data.
[0244] 2. LSTM Unit Operations: The LSTM network is used to capture long-range dependencies in time series data. For each time step, the LSTM unit uses a forget gate... Input gate and output gate Control the information flow. The specific state update formula is as follows:
[0245]
[0246]
[0247]
[0248]
[0249]
[0250]
[0251] in, In cellular state, For output of the hidden layer, It is the Sigmoid activation function. It represents the Hadamardi (or Hadama) stack.
[0252] 3. Interpolation value generation and replacement: The final output of the LSTM is mapped through a fully connected layer to generate an estimated value of the failed sensor at the current moment. The system uses this estimate. Replace the original outlier. .
[0253] Through the aforementioned mechanism, for example, if a soil moisture meter reading remains constant for an extended period (variance of 0, high VAE reconstruction error), while meteorological data indicates recent rainfall and a surge in air humidity, the data integrity assurance unit will determine that the sensor is faulty. It will then use an LSTM model to calculate a reasonable soil moisture interpolation value based on rainfall, air humidity, and data from nearby normally functioning soil sensors. Finally, the cleaned and repaired standardized state vector... The data is transmitted to the multi-objective intelligent decision-making module, ensuring that the control system can still make Pareto optimal decisions even in the event of partial hardware failure, thus avoiding plant death or energy waste due to data errors.
[0254] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A building control method based on multi-source heterogeneous data and Pareto decision-making, characterized in that, include: Multi-domain heterogeneous state data is acquired, and cross-validation and anomaly detection are performed on the data using an unsupervised machine learning algorithm based on variational autoencoders to automatically identify and calibrate sensors experiencing data drift or failure. Missing data is then imputed using a time-series prediction model based on long short-term memory networks. Finally, the processed multi-domain heterogeneous state data is fused into a multi-dimensional state vector. The multidimensional state vector At least including: Building status data used to characterize the thermal state of a building. ; Biological state data used to characterize the vitality of internal vegetation on building exteriors. ; Real-time and predicted environmental data used to characterize external environmental conditions ; The multidimensional state vector The input is fed into a pre-trained non-standardized quantized multi-objective optimization model, which is constructed based on a Markov decision process, wherein the tuples of the Markov decision process are... ,in For state space, For the action space, Let be the state transition probability. To vectorize the reward function, The discount factor is used; wherein, the non-standardized multi-objective optimization model constructs a decision space that interacts with the physical environment in the following manner: Configure the state space To map the , and Normalized high-dimensional continuous vectors composed of physical signals are used to characterize the real-time physical conditions of buildings and installations. Configure the action space The set of physically discrete angles that the drive mechanism of the building's external structure can execute. Each element corresponds to a specific mechanical displacement command; Construct the vectorized reward function To quantitatively assess the feedback impact of control actions on the physical environment, in Generate one at a time dimensional reward vector : in, This is an energy reward item calculated based on heat flux and electrical energy consumption. This is a comfort bonus calculated based on thermal environment parameters. This is a vegetation health incentive program based on physiological constraints of light and water. The energy reward items The formula is: in, Indicates the time measured by the heat flux sensor array Instantaneous heat flux density through walls or windows This indicates the suppression of ineffective heat transfer in any direction; The normalization factor is the historical maximum observed heat flux. Indicates the execution of an action The mechanical and electrical energy consumed; Energy consumption corresponding to the rated power of the drive motor; and It is a dimensionless weighting coefficient used to balance the relationship between building energy saving and device self-consumption energy; in, It is a real-time thermal comfort index calculated based on indoor temperature, relative humidity, average radiant temperature, and set clothing thermal resistance and human metabolic rate. The target comfort value; The bandwidth parameter of the Gaussian function is used to control the tolerance for comfort deviations. The larger the value, the more smoothly the reward function decreases, indicating a higher tolerance for small deviations; The smaller the value, the more precise the control required; the range of values for this function is... ; in, To introduce soil moisture constraint factors, The photosynthetic efficiency factor; The multi-objective optimization model outputs the non-dominated optimal instruction set. Each of the actions And every action This corresponds to a Pareto optimal solution; Based on preference vector From the optimal instruction set Select the optimal action Control signals are sent to one or more drive mechanisms of the building exterior to adjust the building exterior to the optimal angle orientation.
2. The building control method based on multi-source heterogeneous data and Pareto decision-making according to claim 1, characterized in that, The acquisition of heterogeneous state data across multiple domains includes: The building status data is acquired by heat flux sensors deployed on the building envelope and temperature and humidity sensors deployed indoors. ; The biological state data is acquired by using soil moisture sensors, soil temperature sensors, conductivity sensors, and photosynthetically active radiation sensors deployed in the growth medium of the internal vegetation within the building's exterior. ; Real-time wind speed and direction data are acquired using an ultrasonic anemometer, and data containing at least future weather forecasts is obtained from an API that connects to a third-party weather forecasting application. Hourly forecast weather information together constitutes the environmental data. .
3. The building control method based on multi-source heterogeneous data and Pareto decision-making according to claim 1, characterized in that, The multi-objective optimization model is configured to execute a multi-objective Q-learning algorithm through a processor to establish a mapping decision logic between the physical state of the building and the control actions; The multi-objective Q-learning algorithm constructs a Q-vector representing the value of the control strategy based on the collected heterogeneous state data from multiple domains. : Wherein, the Q vector Characterized in the current physical state of the building Execute mechanical drive action Subsequently, the cumulative discount rewards that can be obtained in the dimensions of building energy conservation, comfort and plant health; As a discount factor, Vectorized reward function; Utilize real-time sensor feedback as an immediate reward Based on the non-dominant rule, the Q The vector is updated, and the processor's update logic follows the following Bellman equation iterative formula: in, The updated control strategy value assessment value. For learning rate, As a discount factor, In the state The immediate reward obtained after performing action a The next state The set consisting of all the Q vectors mentioned above. It is a vector selection function based on the Pareto front, used to select from the potential physical states at the next time step. Filter out non-dominated control paths; The Pareto pruning algorithm is used in each state. Maintain a non-dominated state A set of vectors Based on this set, an instruction set is generated to drive building attachments, replacing the traditional scalar weighted aggregation.
4. The building control method based on multi-source heterogeneous data and Pareto decision-making according to claim 1, characterized in that, The building control method also includes a safety protection mode configured to respond to extreme weather conditions, which includes: The system monitors the physical sampling values of the wind speed sensor and the predicted wind speed values from external meteorological data in real time, and compares the values with the preset structural safety threshold. When the detected value exceeds the structural safety threshold, a high-priority hardware control signal is generated to physically bypass or block the instruction output channel of the multi-objective optimization model. A forced reset signal is sent directly to the drive mechanism of the building's external structure, driving the mechanical structure of the external structure to rotate to a preset safe angle. This reduces the device's frontal area and air resistance to a physically safe level.
5. The building control method based on multi-source heterogeneous data and Pareto decision-making according to claim 1, characterized in that, The training method for the pre-trained non-standardized multi-objective optimization model is a strategy configuration process for building control devices, executed by a computing platform, and includes: Based on the actual physical and thermal parameters and geometric dimensions of the building's external structure and building envelope, a digital simulation environment that maps the physical entity is constructed. A data communication interface is established between the agent of the non-standardized multi-objective optimization model and the digital simulation environment to simulate the thermodynamic response and mechanical motion state of physical entities under different environmental stimuli. The digital simulation environment is driven by inputting historical meteorological and physical data, and the vectorized reward function is used. Gradient updates and policy optimizations are performed on the decision network of the intelligent agent; When the change in the Q vector set is less than a preset threshold, the training is determined to be converged, the model weight parameters are exported and fixed into the embedded operating unit of the building control device to complete the initial configuration of the building control device.
6. A building control system based on multi-source heterogeneous data and Pareto decision-making, to implement the building control method based on multi-source heterogeneous data and Pareto decision-making as described in claim 1, characterized in that, The building control system includes: A multi-domain data acquisition and preprocessing module is configured to acquire and fuse heterogeneous state data from multiple domains. It performs cross-validation and anomaly detection on the heterogeneous state data using an unsupervised machine learning algorithm based on a variational autoencoder to automatically identify and calibrate sensors experiencing data drift or malfunction. It also imputes missing data using a time-series prediction model based on a long short-term memory network, and fuses the processed heterogeneous state data into a multi-dimensional state vector. The multidimensional state vector It should include at least building status data, biological status data, and real-time environmental data; The multi-objective intelligent decision-making module is configured to be based on the multi-dimensional state vector. The system controls a pre-trained, non-standardized multi-objective optimization model to perform multi-objective decision optimization operations, generating a non-dominated optimal instruction set. ; The preference selection and arbitration module is configured to be used based on preference vectors. From the optimal instruction set Select the optimal action ; The drive control and execution module is configured to perform the optimal action. Control signals are sent to one or more drive mechanisms of the building scaffold to adjust the building scaffold to an optimal angle orientation.
7. The building control system based on multi-source heterogeneous data and Pareto decision-making according to claim 6, characterized in that, The multi-domain data acquisition and preprocessing module includes a heat flux sensor, a photosynthetically active radiation sensor, a soil moisture sensor, an ultrasonic anemometer, and an indoor environment sensor; data from each sensor are wirelessly transmitted to the multi-target intelligent decision-making module via a long-distance wide area network protocol or an NB-IoT protocol. The multi-objective intelligent decision-making module includes an industrial-grade edge computing unit or a single-board computer, on which the non-standard quantitative multi-objective optimization model is embedded and runs.
8. The building control system based on multi-source heterogeneous data and Pareto decision-making according to claim 7, characterized in that, The edge computing unit or single-board computer in the multi-objective intelligent decision-making module is configured to execute a multi-objective Q-learning algorithm. The execution logic of the algorithm includes establishing a Q-vector data structure in memory. : Wherein, the Q vector Characterized in the current physical state of the building Execute mechanical drive action Subsequently, the cumulative discount rewards that can be obtained in the dimensions of building energy conservation, comfort and plant health; As a discount factor, Vectorized reward function; The edge computing unit is configured to invoke an update and action selection procedure to update the Q vector based on non-dominated rules, where the Bellman equation iterative formula is: in, The updated control strategy value assessment value. For learning rate, As a discount factor, In the state The immediate reward obtained after performing action a The next state The set consisting of all the Q vectors mentioned above. It is a vector selection function based on the Pareto front, used to select from the potential physical states at the next time step. Filter out non-dominated control paths; The edge computing unit is also configured to run a Pareto pruning algorithm program, maintaining a set of non-dominated Q vectors in each state s. This replaces the traditional scalar-weighted aggregation.
Citation Information
Patent Citations
Green intelligent building control system
CN109270889A
Building thermal environment and building energy-saving control method for realizing demand side response
CN118818999A