Sunshade and illumination linkage intelligent adjustment method and system based on environment self-learning
By constructing a dynamic environmental characteristic and user preference model, and combining model predictive control and counterfactual inference, the problems of environmental adaptability and user preference understanding in the linkage regulation of shading and lighting are solved, and efficient and personalized energy and comfort optimization is achieved.
Patent Information
- Application Number
- CN202511187888.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-12-02
AI Technical Summary
Existing technologies for the coordinated adjustment of shading and lighting suffer from static and fixed environmental models that cannot adapt to dynamic changes in the external environment. User preference learning remains at the level of superficial behavioral correlations, control strategies lack foresight, and it is difficult to achieve the global optimization of energy and comfort. There is also a cognitive gap between the objective physical model and the user's subjective perception.
By constructing dynamic environmental feature models and user preference models, and combining model predictive control and counterfactual inference mechanisms, predictive and coordinated control of shading and lighting equipment is achieved. A model co-evolution mechanism and an active exploration mechanism are adopted, and user feedback is used to correct the model, ensuring that the system is continuously updated to reflect the external environment and user preference models.
It achieves adaptive learning of the external environment, deeply understands user intervention motivations, and predictively optimizes control decisions, thereby improving the system's robustness and personalized adjustment capabilities in complex environments and ensuring overall optimal user comfort and energy consumption.
Smart Images

Figure CN121050243A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent indoor lighting technology, specifically to a method and system for intelligent adjustment of shading and lighting linkage based on environmental self-learning. Background Technology
[0002] In modern intelligent buildings, the automated control of equipment such as motorized sunshade blinds and dimmable lighting fixtures to meet the visual comfort needs of indoor occupants and optimize building energy consumption has become an important technological application direction.
[0003] However, existing technologies still face several technical bottlenecks in achieving efficient and comfortable integrated regulation of shading and lighting. Current control methods generally rely on pre-input, static building physics models. Once established, these models struggle to adapt to dynamic changes in the external environment, such as construction of nearby new buildings or the growth and shedding of seasonal vegetation, all of which significantly reduce the model's predictive accuracy over time. Furthermore, in achieving personalized adjustments, existing technologies often only learn user preferences at the level of superficial behavioral correlation matching. This learning approach fails to delve into the true causal motivations behind user interventions, leading to frequent misunderstandings of user intentions, especially when faced with new and unfamiliar lighting conditions. In terms of control strategies, many systems employ reactive logic based on the current state, lacking foresight regarding future changes in the lighting environment. This lag in regulation not only makes it difficult to achieve global optimization of energy and comfort but may also lead to frequent switching of control actions, thus disturbing the user. At a deeper level, existing technologies suffer from a disconnect between objective physical models and users' subjective perceptions. An environment deemed comfortable by the system based on physical formulas may still feel uncomfortable to the user, but this subjective feedback cannot be used to correct the system's cognitive model of the objective world. Ultimately, the entire learning process relies entirely on passive user feedback; the system must wait for the user to experience significant discomfort and take action before it can learn. This prevents the system from proactively and efficiently exploring the precise boundaries of the user's comfort zone, resulting in low learning efficiency. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a method and system for intelligent adjustment of shading and lighting based on environmental self-learning. This solves the problems of static and fixed environmental models that cannot adapt to dynamic changes in the external environment; user preference learning that remains at the level of superficial behavioral correlations and fails to deeply understand the true causal motivations of user intervention; control strategies that lack foresight and are difficult to achieve global optimization between energy and comfort; and the cognitive gap between objective physical models and user subjective perception.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] The first aspect of this invention provides a method for intelligent adjustment of shading and lighting linkage based on environmental self-learning. This method achieves predictive and coordinated control of shading and lighting devices by dynamically constructing and continuously evolving an environmental model and a user preference model. The method includes the following steps:
[0007] First, based on the collected ambient light data, an environmental feature model characterizing the optical properties of the external light environment is constructed. Next, based on monitored user manual intervention behaviors with shading or lighting devices, a user preference model reflecting the user's personalized comfort preferences is learned and generated. Then, the system combines the environmental feature model's prediction of the future light environment with the user preference model's inference of user behavior to jointly determine a coordinated control strategy for the shading and lighting devices. Finally, according to this coordinated control strategy, control commands are output to the shading and lighting devices and executed.
[0008] In an optional implementation, the specific steps for constructing the environmental feature model are as follows: In the initial stage of system operation, time-series illumination data collected by the light sensor is continuously acquired, and the solar position information corresponding to the data acquisition time is acquired simultaneously; the combined effects of shading and reflection on light in the external environment are abstracted into a set of environmental feature parameters to be solved; finally, through an optimization algorithm of inverse solution, the environmental feature parameters that minimize the error between the model's predicted illumination and the actual collected time-series illumination data are iteratively calculated, and this set of parameters is used as the initial state of the environmental feature model.
[0009] In one optional implementation, the specific steps for determining the linkage control strategy are as follows: using the constructed environmental feature model, predicting the indoor light environment state within a preset time window; establishing a comprehensive cost function, the objective term of which includes the total system energy consumption, glare discomfort based on the physical model, and the probability of negative user intervention output by the user preference model; finally, using a model predictive control algorithm, continuously solving for the optimal control sequence that minimizes the comprehensive cost function within the time window, and this sequence constitutes the linkage control strategy.
[0010] In an optional implementation, the probability of negative user intervention is learned and output based on a structural causal model. The specific steps for learning the user preference model are as follows: when the system detects a user's manual intervention behavior, a counterfactual inference mechanism is activated. This mechanism uses the structural causal model to analyze the possible user intervention results if other alternative control strategies are implemented under the current environmental conditions, and updates a set of model parameters in the structural causal model that represent the user's potential comfort preferences based on this inference result.
[0011] In an optional implementation, the method further includes a model co-evolution mechanism, which first continuously detects whether there is a preset cognitive conflict between the physical environment prediction results generated by the environmental feature model and the user perception inference results generated by the user preference model; the cognitive conflict is defined as the physical environment prediction results showing that the current environment is comfortable, while the user perception inference results show that the user has a high probability of intervention behavior.
[0012] In an optional implementation, when the cognitive conflict is triggered, the system initiates a correction procedure for the environmental feature model; this correction procedure adjusts the parameters in the environmental feature model by using recent user intervention behaviors as the basis for correction, so as to eliminate the conflict between models.
[0013] In an optional implementation, the specific steps for modifying the environmental feature model are as follows: First, based on the type of conflict, a set of candidate environmental model hypotheses for explaining the conflict phenomenon are generated, where each hypothesis corresponds to a set of adjustments to the original environmental feature model; then, using a posterior probability evaluation algorithm, the likelihood that each candidate environmental model hypothesis can explain historical intervention behaviors is calculated; finally, the candidate environmental model hypothesis with the highest likelihood is selected, and its corresponding adjustments are fixed into the environmental feature model, completing one modification.
[0014] In an optional implementation, the method further includes an active calibration mechanism for user comfort boundaries. The mechanism comprises the following steps: when the model predictive control algorithm predicts that the probability of negative user intervention under all control strategies is lower than a preset exploration threshold within a future time window, the system will superimpose a preset, user-imperceptible exploratory disturbance signal onto the calculated linkage control strategy and execute it; the system will then monitor user feedback and use the feedback information to update the parameters in the user preference model, thereby achieving accurate calibration of the user comfort boundary.
[0015] In an optional implementation, the user preference model includes a potential comfort variable for quantifying the user's current level of comfort. The value of this variable is determined by multiple dimensions of the indoor lighting environment. Whether the user's intervention behavior occurs is determined by the level of this potential comfort variable.
[0016] A second aspect of the present invention provides an intelligent adjustment system for shading and lighting linkage based on environmental self-learning, which is designed to perform the aforementioned method. The system includes:
[0017] The environmental modeling module is configured to construct an environmental feature model that characterizes the optical properties of the external light environment based on the collected ambient lighting data.
[0018] The user learning module is configured to learn and generate a user preference model that reflects the user's personalized comfort preferences based on the monitored user manual intervention behaviors.
[0019] The decision module is configured to combine the environmental feature model and the user preference model to determine the linkage control strategy for the shading device and the lighting device.
[0020] And a control module configured to control the shading device and the lighting device according to the linkage control strategy.
[0021] This invention provides a method and system for intelligent adjustment of shading and lighting linkage based on environmental self-learning.
[0022] It has the following beneficial effects:
[0023] 1. This invention employs a reverse modeling strategy to deduce environmental feature models from actual collected lighting data, achieving adaptive learning of the external lighting environment. This approach eliminates reliance on pre-defined building geometry or 3D models, allowing the system to be deployed without complex manual configuration, significantly reducing the implementation threshold and cost. Furthermore, it enables the system to automatically adapt to long-term or temporary changes in the surrounding environment, improving its environmental adaptability and robustness.
[0024] 2. This invention learns user preferences by introducing a structural causal model and a counterfactual inference mechanism, enabling a deeper exploration of the causal motivations behind users' manual intervention behaviors, rather than merely matching superficial behavioral patterns. This approach allows the system to construct a more accurate and generalizable user comfort model, thereby achieving truly deep personalized adjustment and proactively avoiding lighting conditions that may cause discomfort to specific users.
[0025] 3. This invention employs a model predictive control framework, integrating energy consumption, physical comfort, and the probability of negative user intervention predicted by a user preference model into a unified comprehensive cost function for rolling optimization. This achieves predictive and coordinated control of shading and lighting equipment. This approach ensures that each control decision is an optimal solution made after weighing future environmental changes, system energy consumption, and user subjective experience, avoiding the short-sightedness and lag of traditional control strategies and guaranteeing the overall optimality of system operation.
[0026] 4. This invention, by establishing a cognitive conflict detection and model co-evolution mechanism, endows the system with the ability to discover and correct defects in the objective environment model using the user's subjective perception. When the predictions of the physical model contradict the inferences of the user model, the system can use this conflict as a correction signal to proactively update its perception of the external environment. This high-level feedback loop enables the system to cope with complex or dynamic environmental factors that did not exist during the initial modeling, greatly improving the system's long-term robustness in real complex environments.
[0027] 5. This invention designs a dynamic calibration mechanism for user comfort boundaries based on proactive exploration. This mechanism can proactively probe the unknown boundaries of the user comfort model by applying minute perturbations, ensuring the user remains unaware of the disturbance. Compared to passively waiting for users to experience large-scale discomfort and then intervening, this proactive approach to acquiring high-value information significantly improves the learning efficiency and convergence speed of the user preference model. This allows for the establishment of a more refined and accurate personalized comfort model with less time and interaction costs. Attached Figure Description
[0028] Figure 1 This is a schematic diagram of the system functional module structure according to an embodiment of the present invention;
[0029] Figure 2 This is a schematic diagram of the overall process of the intelligent adjustment method according to an embodiment of the present invention;
[0030] Figure 3 This is a schematic diagram illustrating the principle of the cognitive conflict detection and model co-evolution mechanism in an embodiment of the present invention;
[0031] Figure 4 This is a schematic diagram illustrating the principle of the active exploration and dynamic calibration mechanism for user comfort boundaries in an embodiment of the present invention. Detailed Implementation
[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] Example:
[0034] Please see the appendix Figure 1 -Appendix Figure 4 This invention provides a method for intelligent adjustment of shading and lighting linkage based on environmental self-learning, comprising the following steps:
[0035] Step S1: Perform system initialization and reverse self-construction of the environment model:
[0036] The purpose is to enable the adjustment system of this invention, after deployment, to generate a mathematical model that accurately characterizes the optical properties of the external light environment at its specific location through autonomous learning, without relying on any pre-provided architectural drawings, 3D models, or manual survey data. This step forms the data and model foundation for all subsequent predictive control and intelligent decision-making.
[0037] In one specific implementation, the system acquires environmental data through one or more light sensors deployed inside the building windows. Preferably, a sensor array can be used to obtain richer spatial dimension information. The sensor array represents the comprehensive optical effects of the external light environment on sunlight, abstracted into an environmental feature matrix M to be solved. This matrix is not a direct geometric description of the external physical entities, but rather a mathematical expression of the optical transfer function of the physical environment, the dimensions and structure of which depend on the complexity of the physical effects to be described. This matrix M can comprehensively characterize complex and comprehensive optical effects such as static shading from surrounding buildings, specular reflection from other building glass curtain walls during specific time periods, and diffuse reflection from the ground or vegetation.
[0038] The system internally predefines a forward physics model function f(·), which mathematically describes the theoretically predicted illumination value L that the sensor position inside the window should receive under the combined influence of a given solar position vector S(t) and a set of defined environmental feature matrices M. pred (t). This functional relationship can be expressed as:
[0039] L pred (t)=f(M,S(t));
[0040] During an initial data acquisition cycle after system deployment, the system continuously records the actual acquired temporal illumination observation vector L. pred The solar position vector S(t) and the synchronously calculated solar position vector S(t) form a set of data pairs for model training.
[0041] Based on this dataset, this step transforms the construction process of the environmental feature model into an inverse optimization problem. Its core objective is to find an optimal initial matrix M0 among all possible environmental feature matrices M, such that the cumulative error between the illumination prediction sequence calculated by the forward physics model function f(·) from this matrix M0 and the corresponding solar position vector S(t) is minimized and the actual observed illumination sequence is minimized.
[0042] Specifically, the optimization objective function can be constructed as follows:
[0043]
[0044] In this formula, This represents the square of the Euclidean norm, used to quantify the difference between the predicted and observed vectors at each time step. The summation symbol is ∑. t This means that the differences at all times within the entire initial acquisition period are summed up to evaluate the overall fitting performance of the model over the entire time period.
[0045] To solve this optimization problem, this embodiment can employ one or more numerical optimization algorithms, such as gradient descent, conjugate gradient, or the Levenberg-Marquardt algorithm for optimizing nonlinear least squares problems. Through iterative calculations, the algorithm continuously adjusts the parameters in matrix M until the objective function converges to a local or global minimum.
[0046] Finally, the environmental feature matrix M corresponding to the algorithm's convergence is determined as the initial environmental feature model M0 of the system. This model is generated in a purely data-driven manner. It not only contains static environmental information but also indirectly reflects the regular dynamic optical events that occur during the data acquisition period, providing accurate environmental disturbance prediction capabilities for the model predictive control in the subsequent step S3.
[0047] Step S2: Based on user intervention behavior, perform deep learning of the user preference model:
[0048] Its core objective is to establish a mathematical model that transcends superficial behavioral imitation and delves into the underlying reasons and personalized comfort preferences behind users' manual intervention and control. The establishment of this model aims to solve the technical challenge of traditional control methods failing to effectively handle users' subjective feelings and complex preferences, and is a key step in achieving intelligent and personalized services.
[0049] In one specific implementation, this method does not employ simple behavioral pattern matching or rule induction to learn user preferences, because such methods are difficult to generalize to unfamiliar environmental states and cannot distinguish the differences in motivation for users to perform the same behavior in different contexts. To address this issue, this embodiment introduces a structural causal model (SCM) as a theoretical framework for describing and learning user behavior.
[0050] This structural causal model explicitly defines the causal relationships between variables in the system through a directed acyclic graph. Its nodes include not only directly measurable physical state quantities, such as the window glare index, average illuminance of the indoor work surface, and light uniformity predicted by the environmental model in step S1, but also system control quantities, such as the opening and closing angle of Venetian blinds and the dimming level of artificial lighting.
[0051] Crucially, the model introduces a latent comfort variable C(t) that is not directly observable and is used to characterize the user's overall subjective experience. This variable is a mathematical abstraction of the comfort level of the user's current lighting environment, and its value is directly related to the user's satisfaction. This latent comfort C(t) is modeled as a function h(·) of the current indoor lighting environment state vector x(t), and is further defined by a set of personalized user preference parameters Θ. u This is determined by [the law / regulation]. This relationship can be expressed as:
[0052] C(t) = h(x(t); Θ u );
[0053] In this equation, the state vector x(t) represents the objective physical world that the system can perceive, while the parameter set Θ... u It is an intrinsic quantitative standard that needs to be determined through learning, representing the specific user's attitude towards different dimensions of the light environment (e.g., the degree of aversion to glare, the range of preferred brightness).
[0054] The user's intervention behavior I(t) (e.g., when I(t) = 1, it indicates that the user has manually intervened) is modeled as a probabilistic event determined by their potential comfort level C(t). Specifically, the lower the level of potential comfort level C(t), the higher the probability that the user will take intervention measures to change the status quo. This probabilistic relationship can be established by a monotonic activation function (preferably a sigmoid function), thereby linking unobservable comfort level with observable intervention behavior. The probability P(I(t) = 1) that the user intervenes in state x(t) constitutes the probability D(·) of negative user intervention required for model predictive control in subsequent step S3.
[0055] The learning mechanism in this step is triggered when the system detects a user intervention. Unlike traditional methods, the system does not directly establish a simple mapping from state x(t) to the intervention action. Instead, the system initiates a counterfactual inference procedure. This procedure uses the causal relationship graph of a structural causal model to perform "post-hoc attribution" analysis.
[0056] By accumulating and analyzing each actual user intervention event and using these events as training data, the system iteratively adjusts and updates the user preference parameter set Θ with the optimization objective of maximizing the posterior probability of the entire historical intervention sequence. u The essence of this learning process is to find a set of intrinsic preference parameters that best explain why users choose to intervene at specific times and in specific environments.
[0057] By performing this step, the system ultimately obtains a dynamic, causal user preference model that deeply understands specific user preferences. This model can not only explain past user behaviors, but more importantly, it can predict the probability of a user's possible negative subjective reaction to any potential, unexperienced future lighting environment state. This provides crucial, quantitative input for the subsequent step S3 to achieve truly user-centered predictive control decisions.
[0058] Step S3: Based on model predictive control, make predictive and coordinated decisions regarding shading and lighting.
[0059] The steps aim to treat shading devices and lighting devices as a unified, coupled system for joint optimization, replacing the limitations of traditional methods that separate control or simple linkage between the two, thereby achieving the optimization of the overall system performance while meeting user needs.
[0060] In one specific implementation, to achieve predictability and smoothness in control, this method employs Model Predictive Control (MPC) as the core decision-making framework. The advantage of this framework is that it does not rely solely on reactive control based on the current state, but rather uses a system model to predict the system's behavior over a finite future time domain, and determines the optimal control action by solving an open-loop optimization problem.
[0061] Specifically, at each control decision time t, this step first enters the prediction phase. The controller utilizes the environmental characteristic model M constructed and continuously evolved in step S1, combined with the accurately calculable future solar trajectory vector sequence S(t+k) (where k=0,1,...,N). p -1, N p To predict the future N (in terms of the length of the time domain), p The dynamic changes in the indoor natural light environment caused by the external environment within a given time period. This prediction result constitutes the key external disturbance input in the system model.
[0062] Subsequently, the core task of this step is to construct and solve a comprehensive optimization problem. The controller aims to find a set of optimal control sequences. This sequence enables a pre-defined comprehensive cost function J. MPC It reaches its minimum value throughout the entire prediction time domain. This comprehensive cost function is carefully designed to comprehensively and quantitatively balance multiple, sometimes even conflicting, objectives in the system's operation. Its mathematical expression can be defined as:
[0063]
[0064] In this cost function, each component has a clear physical meaning:
[0065] u t+k It is the control input vector applied to the system at a future time t+k, preferably including the opening and closing angle u of the blinds. shd (t+k) and dimming level u for artificial lighting lgI (t+k).
[0066] x t+k It is the system state vector at a future time t+k, which is jointly determined by the current state, the future control input sequence, and the predicted external disturbances. Preferably, it may include the average illuminance and uniformity of the indoor working surface and the glare index at the window position.
[0067] E(u t+k The item represents the predicted total energy consumption of the system directly related to the control action, such as the energy consumption of the lighting system and the energy consumption of the motor driving the sunshade device.
[0068] G(x t+k The term represents the result calculated based on the physical optics model, in system state x. t+k The glare discomfort index perceived by users.
[0069] D(x t+k ,u t+k The term "(S2)" is a key technical feature of this invention. It represents the predicted value of the probability that the user will generate negative emotions and manually intervene under future states and control actions, output by the user preference model in step S2. The introduction of this term allows the user's personalized subjective preferences to be directly and quantitatively incorporated into the objective function of the optimization decision, thereby enabling the control decision to proactively avoid states that cause user dissatisfaction.
[0070] w e ,w g ,w d These are preset non-negative weighting coefficients used to adjust the relative importance of the three objectives—energy saving, physical comfort, and user subjective satisfaction—in the overall optimization.
[0071] The process of solving this optimization problem also needs to satisfy various physical constraints existing in the system, such as the mechanical limit of the angle of the blinds and the dimming range of the lighting fixtures.
[0072] Based on the rolling time-domain principle of model predictive control, the optimal control sequence is obtained by solving the above optimization problem. Afterward, the system does not execute the entire sequence. Instead, it extracts and executes only the first control action in the sequence. At the next control time t+1, the system will update the current state using the latest sensor measurements and re-execute the entire prediction and optimization process described above to calculate the new optimal control action.
[0073] By performing this step, every control decision made by the system is the optimal choice at present, based on anticipating future environmental changes, weighing multiple performance indicators, and taking into account the user's personalized preferences. This ensures that the entire adjustment process is smooth, efficient, and highly personalized.
[0074] Step S4, based on cognitive conflict, drives the co-evolution of the environment model and the user model:
[0075] This step aims to establish a high-level cognitive feedback and correction loop to address inconsistencies that may arise during long-term system operation between the objective environmental feature model constructed in step S1 and the subjective user preference model learned in step S2. Through this step, the system can leverage users' subjective perceptions, which are difficult to articulate directly, to discover and correct potential flaws in the objective physical model, thereby achieving the co-evolution and mutual verification of the two models.
[0076] In one specific implementation, the system continuously detects "cognitive conflict" during operation. Here, "cognitive conflict" is precisely defined as a specific logical state in which a significant and persistent contradiction arises between the system's physical model predictions and the user model inferences.
[0077] Specifically, a cognitive conflict event is triggered when the following two sub-conditions are met simultaneously within the same time period:
[0078] First, the prediction results based on the physical model show that the current environment is comfortable. This means that the environmental feature model M from step S1 is suitable. t The glare discomfort index G(x(t)) is calculated from the current system state x(t); M t Its value remains below a preset comfort threshold θ g .
[0079] Secondly, the inference results based on the user model show that users feel uncomfortable with the current environment and have a very high tendency to intervene. This means that the user preference model in step S2 (whose parameters are...) The output of the user's negative intervention probability Its value consistently exceeds a preset intervention tendency threshold θ d .
[0080] The mathematical logic expression for this conflict condition can be represented as:
[0081]
[0082] The occurrence of this state indicates that the environmental characteristic model M t There are cognitive blind spots, meaning that certain environmental physical factors that are keenly perceived by users but unknown to the model are not accurately captured. For example, an intermittent strong reflected light from the curtain wall of a nearby newly built building, which did not appear in the initial modeling stage, may not cause a significant increase in the traditional glare index, but its characteristics (such as flicker and angle) may cause subjective visual discomfort to certain users.
[0083] Once the system determines that the aforementioned cognitive conflict is ongoing rather than transient noise, this step will trigger a model correction procedure. The core idea of this procedure is to use the user intervention behavior that caused the conflict itself as the most reliable "real-world label" to correct the physical world model in reverse.
[0084] The correction process first enters the hypothesis generation phase. Based on the specific manifestations of the conflict and historical data, the system generates a set of environmental characteristic models M. t The alternative modified hypotheses form a candidate model set {M}. j Each candidate model M j ′ both represent the original model M t One possible adjustment scheme that could explain this conflict phenomenon is to add a specific orientation and intensity of the reflection source parameter to the original model.
[0085] The program then proceeds to the model selection phase. In this phase, instead of randomly selecting a modification scheme, the system uses a posterior probability-based evaluation method to determine which candidate model best explains the user's intervention history during the conflict. Its goal is to find an optimal correction model M. new This model maximizes the joint likelihood probability of observed historical intervention sequences. This selection process can be described by the following optimization objective:
[0086]
[0087] In this formula, This indicates that when using candidate model M j Given the premise of describing the physical world, the user makes an intervention action I i The probability that = 1.
[0088] Ultimately, the candidate model M with the highest likelihood is selected. best If M is selected as the optimal correction, the system will adopt this correction and use it to update the current environmental characteristic model, that is, let M... t ←M new.
[0089] By performing this step, the system establishes a closed-loop pathway from user subjective perception to objective model correction. This enables the environmental characteristic model to evolve not only through physical data but also through insightful corrections based on user cognitive feedback. This ensures that the system's understanding of the complex "human-environment" system continues to deepen, enhancing the long-term adaptability and robustness of the entire adjustment method.
[0090] Step S5: Based on proactive exploration, dynamically calibrate the user's comfort boundaries:
[0091] The purpose of this step is to address the data sparsity and bias issues encountered when model learning relies solely on passive user intervention (such as step S2). By performing small, goal-oriented exploratory actions without the user's awareness or disturbance, the system can more efficiently and accurately acquire high-value information about the boundaries of the user comfort model, thereby accelerating and optimizing the convergence process of the user preference model.
[0092] In one specific implementation, this step is not continuous, but is triggered by a deliberate "exploration timing decision" mechanism. The system confines proactive exploration behavior within a "safe exploration window" to minimize unnecessary interference with the user.
[0093] The opening condition of this exploration window is based on the predictive capability of model predictive control in step S3. Specifically, during the rolling optimization of MPC, it not only calculates the current optimal control strategy, but also predicts the entire prediction time domain N under this strategy. p The system state and user responses are monitored. The timing decision mechanism continuously monitors the probability D(·) of negative user intervention on the future optimal control path, output by the user preference model. Only when this predicted probability remains consistently below a preset, very conservative exploration threshold θ throughout the entire future time window will the mechanism proceed. exp Only then will the system determine that the user is in a state of deep comfort for the present and a period of time to come, making it suitable to begin exploration. The mathematical expression for this condition is:
[0094]
[0095] in The future optimal control sequence calculated for MPC.
[0096] Once the exploration window is entered, this step will initiate the "exploratory perturbation injection" procedure. This procedure will then apply the optimal control action u calculated in step S3 at the current moment. *Based on (t), a carefully designed, small exploratory perturbation signal δ(t) is superimposed to form the final control command u issued to the hardware actuator. exec (t). This relationship can be represented as:
[0097] u exec (t)=u * (t)+δ(t);
[0098] The perturbation signal δ(t) here is not random noise. Its design follows several principles: First, its amplitude and rate of change are strictly limited below the human perception threshold to ensure that users will not perceive minute changes in the environment; second, its application direction is goal-oriented. Preferably, the system will apply the perturbation in the direction with the highest uncertainty in the current user preference model, aiming to obtain the most informative feedback. For example, if the model has low confidence in the user's response at a specific glare index, the perturbation signal will drive the system state to gently probe in that direction.
[0099] The final step in this process is "boundary learning and model update." The system closely monitors and records the user's feedback after the exploratory perturbation is applied, i.e., whether the intervention behavior I(t+1) occurred in the subsequent time. As a result, the system obtains a valuable "perturbation-feedback" data pair (δ(t), I(t+1)) generated by the active experiment.
[0100] Compared to passively collected data, this actively explored data sample is of extremely high value. It provides the user preference model learning algorithm in step S2 with precise data points in the region where the model is most uncertain, close to the comfort boundary.
[0101] If multiple explorations fail to elicit user intervention, the system will then expand the boundaries of the user's comfort zone based on evidence, reducing the probability of predicted intervention in that area.
[0102] If a small disturbance in a certain direction eventually triggers user intervention, the system learns with extremely high precision a specific, previously unknown comfort threshold point for that user.
[0103] These high signal-to-noise ratio samples are input into the parameter update process of the user preference model to more efficiently iteratively optimize the parameter set Θ. u This will allow us to establish a more refined and robust user comfort model.
[0104] By performing this step, the adjustment method of the present invention can not only learn the behaviors that users have already exhibited, but also explore the boundaries of users' unspoken preferences through safe and intelligent "proactive questioning," thereby achieving a higher level of personalization and adaptability.
[0105] Step S6, execute the linkage control strategy:
[0106] This step forms a bridge from upper-level intelligent decision-making to lower-level physical device actions, and is key to closing the entire "perception-cognition-decision-control" cycle. Its core task is to accurately translate the abstract and mathematical control commands calculated in the preceding steps into physical operations on specific hardware devices (such as sunshade blinds and indoor lighting fixtures).
[0107] In one specific implementation, the input to this step is the control action vector u that should be executed at the current time t, as ultimately determined by the aforementioned decision-making steps. exec (t). It should be noted that this control vector is the result of the system's integrated decision-making. In normal operating mode, it corresponds to the first element u of the optimal control sequence calculated by model predictive control in step S3. * (t); and within the "exploration window" determined in step S5, it corresponds to the optimal control action u. * The superposition of (t) and the exploratory disturbance signal δ(t).
[0108] The control vector u exec The vector (t) is first parsed by the control module inside the system. This vector is a multi-dimensional mathematical entity, where each component corresponds to a specific, controllable physical device. For example, the vector may contain control components ushd(t) for shading devices and ulgt(t) for lighting devices.
[0109] The parsing process involves converting these numerical control components into low-level operation instructions that can be recognized and executed by the corresponding hardware devices.
[0110] For shading devices, such as motorized Venetian blinds, the control component u shd (k) is a numerical value representing the target blade angle (e.g., ranging from 0 to 90 degrees). The control module converts this angle value into corresponding low-level instructions based on the type of drive motor used. Preferably, if a stepper motor is used, the angle value will be converted into the precise number of steps the motor needs to rotate; if a servo motor with a position encoder is used, the angle value will be set as the target position for the motor movement.
[0111] For lighting equipment, such as dimmable LED lights, the control component u lgI(k) is a percentage representing the target brightness or a specific target illuminance value. The control module converts this value into a corresponding electrical signal based on the driving method of the lighting system. Preferably, if pulse width modulation (PWM) dimming is used, the brightness value will be converted into a PWM signal with a specific duty cycle; if the Digital Addressable Lighting Interface (DALI) digital communication protocol is used, the brightness value will be encoded into a digital message conforming to the protocol format.
[0112] Subsequently, these converted low-level operation instructions are sent to the actuators or drivers of various devices through the system's built-in hardware interface or communication bus. For example, the pulse signal of the stepper motor is sent through the driver board, the PWM signal is applied to the lighting driver power supply through the corresponding output pin, and DALI messages are broadcast or sent point-to-point to the designated luminaires through the DALI bus.
[0113] After receiving an instruction, the hardware actuator drives the physical device to complete the corresponding action, such as rotating the blind motor to a specified angle or adjusting the lighting fixtures to the target brightness.
[0114] At this point, a complete linkage control action has been executed. However, the significance of this step is not limited to this. The resulting change in the physical state of the indoor lighting environment will be immediately captured by the light sensors deployed indoors. These latest sensor readings, containing the effect of this control action, will be used as new inputs at the next control time t+1 to update the observed values of the system state, thereby initiating a new round of environmental model prediction, user preference evaluation, and model predictive control optimization, i.e., returning to step S3 (or triggering S4 or S5 as needed), thus forming a continuously running and constantly optimizing dynamic closed-loop control process.
[0115] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for intelligent adjustment of shading and lighting linkage based on environmental self-learning, characterized in that, Includes the following steps: Based on the collected ambient lighting data, an environmental feature model is constructed to characterize the external light environment. Based on user intervention behavior, a user preference model reflecting user comfort preferences is learned; By combining the environmental feature model and the user preference model, a linkage control strategy for shading devices and lighting devices is determined. The sunshade device and the lighting device are controlled according to the linkage control strategy.
2. The intelligent adjustment method for shading and lighting linkage based on environmental self-learning as described in claim 1, characterized in that, The steps for constructing the environmental feature model include: Acquire temporal illumination data and corresponding solar position information; The influence of the external light environment on light is abstracted into environmental characteristic parameters to be solved; By using an inverse optimization algorithm, environmental feature parameters that best fit the time-series illumination data are determined to generate the environmental feature model.
3. The intelligent adjustment method for shading and lighting linkage based on environmental self-learning as described in claim 1, characterized in that, The steps for determining the linkage control strategy include: Predict the light environment state within a future time window based on the aforementioned environmental feature model; Construct a comprehensive cost function that includes energy consumption, glare comfort, and the probability of negative user intervention; The optimal control sequence that minimizes the overall cost function is obtained by using a model predictive control algorithm to generate the linkage control strategy.
4. The intelligent adjustment method for shading and lighting linkage based on environmental self-learning as described in claim 1, characterized in that, The probability of negative user intervention is obtained based on structural causal model learning. The learning steps include: When user intervention behavior is detected, the structured causal model is used to perform counterfactual inference on the intervention behavior in order to update the model parameters representing the user's potential preferences.
5. The intelligent adjustment method for shading and lighting linkage based on environmental self-learning as described in claim 1, characterized in that, The method also includes: The cognitive conflict is detected between the physical environment prediction results of the environmental feature model and the user perception inference results of the user preference model.
6. The intelligent adjustment method for shading and lighting linkage based on environmental self-learning as described in claim 5, characterized in that, The method also includes: When the cognitive conflict is triggered, the environmental feature model is modified based on the user's intervention behavior.
7. The intelligent adjustment method for shading and lighting linkage based on environmental self-learning as described in claim 6, characterized in that, The steps for modifying the environmental feature model include: Generate candidate environmental model hypotheses to explain the cognitive conflict; By evaluating posterior probability, candidate environmental model hypotheses that maximize the likelihood of historical intervention behaviors are selected to update the environmental feature model.
8. The intelligent adjustment method for shading and lighting linkage based on environmental self-learning as described in claim 3, characterized in that, The method also includes: When the predicted probability of negative user intervention is lower than the preset exploration threshold, a preset exploratory disturbance signal is superimposed on the linkage control strategy, and the user preference model is updated according to the user's feedback to calibrate the user comfort boundary.
9. The intelligent adjustment method for shading and lighting linkage based on environmental self-learning according to claim 1, characterized in that, The user preference model includes potential comfort variables for quantifying user comfort, and the user's intervention behavior is determined by the level of these potential comfort variables.
10. A system based on the intelligent adjustment method of claim 1, characterized in that, include: The environment modeling module is used to construct an environmental feature model that characterizes the external light environment based on the collected ambient lighting data. The user learning module is used to learn a user preference model that reflects the user's comfort preferences based on the user's intervention behavior. The decision module is used to combine the environmental feature model and the user preference model to determine the linkage control strategy for the shading device and the lighting device; The control module is used to control the sunshade device and the lighting device according to the linkage control strategy.
Citation Information
Cited By
Intelligent building illumination and sunshade linkage-oriented adaptive adjustment method and system
CN121477672A
A smart building lighting and sunshade linkage adaptive adjustment method and system
CN121477672B