Intelligent door and window cluster monitoring and intelligent scheduling system based on cloud computing platform

CN120030633APending Publication Date: 2025-05-23NANJING KAISHENG DOORS & WINDOW CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411848759.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing door and window management systems lack real-time monitoring and response capabilities to environmental parameters, resulting in waste of energy and poor environmental comfort, and it is difficult to achieve efficient and accurate door and window control.

Method used

The intelligent door and window cluster monitoring and intelligent scheduling system based on the cloud computing platform monitors door and window and environmental parameters in real time through the door and window status monitoring layer, and combines the multi-objective optimization model of the intelligent model construction layer to dynamically adjust the switches and adjustment solutions of doors and windows. The system includes a state definition component, an action definition component, a rule evaluation component and a policy optimization unit, and uses reinforcement learning algorithms to optimize the scheduling strategy.

Benefits of technology

It realizes efficient and precise control of door and window management, improves energy utilization and environmental comfort, extends the service life of the equipment, and reduces maintenance and replacement costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030633A_ABST
    Figure CN120030633A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of door and window management, and discloses an intelligent door and window cluster monitoring and intelligent scheduling system based on a cloud computing platform, which integrates a door and window state monitoring layer and an intelligent model construction layer, monitors door and window states and environmental parameters in real time, and constructs a multi-target optimization model to dynamically adjust door and window opening and closing and an adjustment scheme. The intelligent model building layer comprises a state definition assembly, an action definition assembly, a rule evaluation assembly, a strategy optimization assembly and the like, and a reinforcement learning algorithm is adopted to optimize a scheduling strategy. The cloud control execution layer is responsible for data uploading, intelligent analysis model processing and control instruction issuing, and precise control over doors and windows is achieved. The system makes full use of the Internet of Things, big data and cloud computing technologies, not only improves the automation and intelligence level of door and window management, but also effectively promotes energy conservation and emission reduction, improves the comfort level of a living or working environment, prolongs the service life of doors and windows and related equipment, and meets the urgent demand of modern buildings for intelligent management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of door and window management, and in particular to an intelligent door and window cluster monitoring and intelligent scheduling system based on a cloud computing platform. Background Art

[0002] With the rapid development of modern building technology and the growing demand for intelligentization, the management of doors and windows has gradually shown its limitations. In large buildings or building complexes, there are a large number of doors and windows, and their status monitoring and scheduling management often require a lot of manpower and time, and it is difficult to achieve efficient and accurate control. Especially in the context of pursuing energy conservation and emission reduction, and improving the comfort of living or working environment, how to manage doors and windows intelligently has become an urgent problem to be solved.

[0003] Most of the current door and window management systems use simple switch control, lacking the ability to monitor and respond to environmental parameters (such as temperature, humidity, light intensity, etc.) in real time, and cannot dynamically adjust the control strategy according to the actual use of doors and windows and environmental changes. This not only leads to a large waste of energy, but also may affect the indoor environmental comfort due to improper opening and closing of doors and windows, and even shorten the service life of doors and windows and related equipment.

[0004] In addition, with the rise of technologies such as the Internet of Things, big data, and cloud computing, new possibilities have been provided for the intelligent management of doors and windows. However, existing door and window management systems often simply apply these technologies to remote control or status monitoring of doors and windows, but fail to give full play to the advantages of these technologies in data analysis, model building, and optimized decision-making. Summary of the invention

[0005] The purpose of the present invention is to provide an intelligent door and window cluster monitoring and intelligent scheduling system based on a cloud computing platform to solve the problems raised in the above background technology.

[0006] To achieve the above object, the present invention provides the following technical solution: a smart door and window cluster monitoring and intelligent scheduling system based on a cloud computing platform, the system comprising:

[0007] Door and window status monitoring layer, used to monitor the status of doors and windows and their environmental parameters in real time;

[0008] The intelligent model building layer is used to build a door and window scheduling strategy model. The model models the management of door and window clusters as a multi-objective optimization problem and dynamically adjusts the opening and closing and adjustment schemes of doors and windows through algorithms. The intelligent model building layer further includes:

[0009] The state definition component is used to describe the current state space of the door, window and its environment;

[0010] Action definition component, used to define executable door and window control action space;

[0011] The rule evaluation component is used to define evaluation rules and quantitatively evaluate the results of control actions based on the accuracy of door and window state changes, energy efficiency, environmental comfort, and equipment life;

[0012] The strategy optimization unit uses a reinforcement learning algorithm to select the optimal control action under a given state and optimize the scheduling strategy through continuous iterative learning;

[0013] The cloud control execution layer is used to upload the data collected by the door and window status monitoring module to the cloud. After being processed by the intelligent analysis model, it receives the control instructions issued by the cloud and drives the door and window actuators and related equipment to perform corresponding actions.

[0014] Preferably, the state space includes the open and closed state of doors and windows, location information, indoor temperature, outdoor temperature, humidity, wind speed, light intensity, noise level and air quality.

[0015] Preferably, the action space includes opening doors and windows, closing doors and windows, adjusting the degree of opening of doors and windows, and starting or closing sunshade devices.

[0016] Preferably, the implementation of the rule evaluation component includes:

[0017] The evaluation rule R(s,a) is defined, where s represents the state space of doors, windows and their environment at the current moment; a represents the control action taken, which is selected from the control action space, including opening doors and windows, closing doors and windows, adjusting the opening degree of doors and windows, and starting or closing sunshade devices; the evaluation rule R(s,a) performs weighted summation based on the accuracy of state change C(s,a), energy efficiency E(s,a), environmental comfort H(s,a) and equipment service life D(s,a) to obtain the quantitative evaluation value of the control action result; the specific algorithm expression is:

[0018] R(s,a)=w1*C(s,a)+w2*E(s,a)+w3*H(s,a)-w4*D(s,a)

[0019] Wherein, w1, w2, w3, and w4 are weight coefficients of accuracy, energy efficiency, environmental comfort, and equipment life, respectively, and w1+w2+w3+w4=1; C(s,a) represents the accuracy of the door and window state change after taking action a under state s; E(s,a) represents energy efficiency, which is obtained by calculating the change in energy consumption; H(s,a) represents environmental comfort, which is comprehensively evaluated based on indoor temperature and humidity, light, and air quality; D(s,a) represents the consumption of equipment life, including the wear of doors, windows, and their accessories caused by actions.

[0020] Preferably, the strategy optimization component uses the Deep Q-Network algorithm in the deep reinforcement learning algorithm to train the door and window scheduling strategy model.

[0021] Preferably, various sensors and monitoring devices are deployed in the door and window cluster to collect the actual status data of the doors, windows and their environment in real time; combined with preset scheduling rules and evaluation criteria, simulated control actions are performed on the collected status data, and the corresponding instant rewards are calculated; the actual status data, simulated control actions and instant rewards are combined into a state-action-reward tuple as preliminary training data for the door and window scheduling strategy model.

[0022] Preferably, the door and window operation records and environmental change data in the historical data are used to construct a historical status data set of doors and windows and their environment; for each data in the historical status data set, possible control actions are simulated and the corresponding immediate rewards are calculated according to preset scheduling rules and evaluation criteria; the historical status data, simulated control actions and immediate rewards are combined into additional state-action-reward tuples, which are supplemented into the preliminary training data to form a complete door and window scheduling strategy model training data set.

[0023] Preferably, the step of the strategy optimization component training the door and window scheduling strategy model includes:

[0024] S1: Initialize the experience replay memory and the target Q network. The experience replay memory is used to store the experience of door and window state transfer, and the target Q network is used to stabilize the target value during training. At the same time, initialize the main Q network, whose input is the current state space of the door and window and its environment, and output is the Q value estimation of each control action.

[0025] S2: Set the learning rate α, discount factor γ, exploration rate ε and number of training rounds, where α controls the learning speed, γ controls the discounted value of future returns, ε controls the balance between exploration and exploitation, and the number of training rounds determines the total number of training iterations;

[0026] S3: For each training round, perform the following steps:

[0027] S301: Initialize the current state s of doors, windows and their environmental parameters from the training data set;

[0028] S302: Select a control action a according to the current state s and the ε-greedy strategy, that is, randomly select an action with a probability of ε, and select the action with the largest Q value output by the current main Q network with a probability of 1-ε;

[0029] S303: Execute action a, observe the next state s' and the immediate reward r, where the immediate reward r is calculated according to the evaluation rule R(s,a) of the rule evaluation unit;

[0030] S304: storing the experience (s, a, r, s') into the experience playback memory;

[0031] S305: Randomly sample a batch of experiences from the experience replay memory for updating the main Q network;

[0032] S306: using the sampled experience and the target Q network to calculate the target Q value, and updating the parameters of the main Q network according to the loss function of the Deep Q-Network algorithm, where the loss function is the mean square error between the predicted Q value and the target Q value;

[0033] S307: copying the parameters of the main Q network to the target Q network at regular intervals;

[0034] S4: Repeat step 3 until the preset number of training rounds is reached or the output value of the main Q network converges;

[0035] S5: After the training is completed, the obtained main Q network is used as the door and window scheduling strategy model. For any given door and window and the current state s of its environment, the action a with the largest Q value is selected as the optimal control action through the forward propagation of the main Q network.

[0036] Preferably, before the model training begins, the training data set is further preprocessed, including encoding and standardization.

[0037] Preferably, discrete features are converted to one-hot encoding; continuous features are standardized using the Z-score standardization method, specifically: each numerical data is subtracted from its mean and divided by its standard deviation, so that the processed data conforms to the standard normal distribution, and the standardization formula is: Z = (X-μ) / σ, where X is the original data, μ is the mean of the original data, σ is the standard deviation of the original data, and Z is the standardized data.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] This system obtains comprehensive status information of doors, windows and their environment in real time through the door and window status monitoring layer, and can dynamically adjust the opening and closing and adjustment schemes of doors and windows by combining the multi-objective optimization model in the intelligent model construction layer. Through intelligent management methods, it not only greatly improves the efficiency and accuracy of door and window management, but also ensures the precise matching of door and window status with environmental requirements, thereby effectively improving the overall operational efficiency of the building and the quality of the living or working environment.

[0040] This system can intelligently adjust the opening and closing and adjustment strategies of doors and windows by real-time monitoring of environmental parameters (such as temperature, humidity, light intensity, etc.) and the status of doors and windows, so as to maximize the use of natural light and reduce unnecessary energy consumption. At the same time, the strategy optimization unit adopts a reinforcement learning algorithm, which can find the optimal control strategy in continuous iterative learning, further improving the energy efficiency performance of the system and contributing to the green and sustainable development of buildings. This system fully considers multiple evaluation indicators such as the accuracy of door and window state changes, energy efficiency, environmental comfort and equipment service life. The results of the control actions are quantitatively evaluated through the rule evaluation component to ensure the rationality and effectiveness of the control strategy, which not only improves the indoor environmental comfort, but also reduces equipment damage and shortened life caused by improper operation of doors and windows, and reduces maintenance and replacement costs.

[0041] Relying on the cloud computing platform, this system realizes the centralized storage, efficient processing and intelligent analysis of door and window status data. The cloud control execution layer can quickly respond to the control instructions issued by the cloud, and drive the door and window actuators and related equipment to perform corresponding actions, realizing the remote, intelligent and automated management of doors and windows. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a working principle diagram of the intelligent door and window cluster monitoring and intelligent scheduling system based on the cloud computing platform described in the present invention;

[0043] Figure 2 A schematic diagram of the design of the rule evaluation component;

[0044] Figure 3 Flowchart for building a door and window scheduling strategy model. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] See also Figure 1-3 The present invention provides a technical solution: a smart door and window cluster monitoring and intelligent scheduling system based on a cloud computing platform, the system comprising:

[0047] The door and window status monitoring layer is responsible for real-time monitoring of the status of doors and windows and their environmental parameters. This layer uses various sensors deployed on doors and windows, such as infrared sensors, temperature and humidity sensors, light sensors, wind speed sensors, noise sensors, and air quality sensors, to collect real-time data on the opening and closing status of doors and windows, location information, and indoor and outdoor temperature, humidity, wind speed, light intensity, noise level, and air quality. These data are transmitted to the cloud control execution module via wired or wireless means.

[0048] Intelligent model building layer, which is used to build a door and window scheduling strategy model. This layer models the management of door and window clusters as a multi-objective optimization problem, comprehensively considers multiple factors such as energy efficiency, environmental comfort, and equipment life, and dynamically adjusts the opening and closing and adjustment schemes of doors and windows through algorithms. The intelligent model building layer further includes:

[0049] State definition component: This component is responsible for describing the current state space of doors, windows and their environment. The state space includes but is not limited to the open and closed state of doors and windows, location information, indoor temperature, outdoor temperature, humidity, wind speed, light intensity, noise level and air quality.

[0050] Action definition component: This component defines the executable door and window control action space, including opening doors and windows, closing doors and windows, adjusting the opening degree of doors and windows, and starting or closing sunshade devices. These actions are control instructions output by the model to achieve precise control of doors and windows.

[0051] Rule evaluation component: This component defines evaluation rules for quantitative evaluation of the results of control actions. Evaluation indicators include the accuracy of door and window state changes, energy efficiency improvement, environmental comfort improvement, and equipment life extension. Through rule evaluation, the system can continuously optimize control strategies and improve management efficiency.

[0052] Policy Optimization Unit: This unit uses a reinforcement learning algorithm to select the optimal control action under a given state. Through continuous iterative learning, the policy optimization unit can continuously accumulate control experience and optimize the scheduling strategy. In each iteration, the system selects a control action based on the current state, observes the results after execution and calculates the reward value, and then updates the policy model to maximize the long-term reward.

[0053] Cloud control execution layer: The cloud control execution layer is responsible for uploading the data collected by the door and window status monitoring layer to the cloud and processing it through the intelligent analysis model. The processed data is used to generate control instructions, which are sent to the cloud control execution module, which drives the door and window actuators and related equipment to perform corresponding actions. The cloud control execution module is also responsible for receiving feedback on the execution results and uploading the feedback data to the cloud for subsequent model optimization and decision improvement.

[0054] The present invention will be further described below in conjunction with Examples 1 to 3:

[0055] Embodiment 1:

[0056] The rule evaluation component is used to quantitatively evaluate the results of control actions. The implementation of the rule evaluation component includes:

[0057] The rule evaluation component defines the evaluation rule R(s,a) to quantify the result of taking control action a under a specific state s. State s represents the state space of doors, windows and their environment at the current moment, including the open and closed state of doors and windows, location information, indoor temperature, outdoor temperature, humidity, wind speed, light intensity, noise level and air quality. Control action a is selected from the predefined control action space, including opening doors and windows, closing doors and windows, adjusting the opening degree of doors and windows, and starting or closing sunshade devices.

[0058] The evaluation rule R(s,a) comprehensively considers the accuracy of state change C(s,a), energy efficiency E(s,a), environmental comfort H(s,a) and equipment service life D(s,a), and obtains the quantitative evaluation value of the control action result by weighted summation. The specific algorithm expression is:

[0059] R(s,a)=w1*C(s,a)+w2*E(s,a)+w3*H(s,a)-w4*D(s,a)

[0060] Among them, w1, w2, w3, and w4 are weight coefficients of accuracy, energy efficiency, environmental comfort, and equipment life, respectively, and satisfy w1+w2+w3+w4 = 1. These weight coefficients are set according to actual application scenarios and requirements to reflect the relative importance of each factor in the evaluation.

[0061] Take an office as an example, assuming the current state s is: indoor temperature 25℃, outdoor temperature 30℃, humidity 60%, light intensity 1000 lux, doors and windows are closed. Based on the current state and environmental changes, the system decides to take control action a: open the doors and windows to 50% opening degree.

[0062] Accuracy of state change C(s,a): Assuming that the actual opening degree of the door and window is consistent with the instruction, C(s,a) = 1 (indicating complete accuracy). If there is a deviation, the value of C(s,a) is reduced accordingly according to the degree of deviation.

[0063] Energy efficiency E(s,a): Energy efficiency is evaluated by calculating the change in energy consumption before and after opening doors and windows. Assuming that after opening doors and windows, energy consumption is reduced by 10% due to the use of natural ventilation and light, then E(s,a) = 0.1 (indicating a 10% improvement in energy efficiency).

[0064] Environmental comfort H(s,a): Environmental comfort is evaluated based on the comprehensive changes in indoor temperature, humidity, light, and air quality. Assuming that after opening the doors and windows, the indoor temperature drops slightly, the humidity is moderate, the light becomes more natural, and the air quality improves, the comprehensive evaluation H(s,a) = 0.8 (indicating that the environmental comfort has increased by 80%).

[0065] Equipment service life D(s,a): Consider the wear of doors, windows and their accessories (such as hinges, slide rails, etc.). Assuming that the wear of the equipment caused by opening doors and windows is small, D(s,a) = 0.01 (indicates that the service life of the equipment is consumed by 1%).

[0066] Evaluation value calculation:

[0067] Assuming the weight coefficients are: w1 = 0.3, w2 = 0.2, w3 = 0.4, w4 = 0.1, then the evaluation value R(s,a) = 0.3*1+0.2*0.1+0.4*0.8-0.1*0.01=0.3+0.02+0.32-0.001=0.639.

[0068] Through the above calculation, the quantitative evaluation value R(s,a) of the result after taking control action a in a specific state s is obtained. This value reflects the comprehensive performance of the control action in terms of accuracy, energy efficiency, environmental comfort and equipment service life.

[0069] Embodiment 2:

[0070] The strategy optimization component uses the Deep Q-Network (DQN) algorithm in the deep reinforcement learning algorithm to train the door and window scheduling strategy model. This embodiment is used to elaborate on the specific implementation of the component and illustrate it with specific examples. The steps of training the door and window scheduling strategy model include:

[0071] S1: Initialization

[0072] Initialize the experience replay memory to store the experience of door and window state transition (s, a, r, s'), where s represents the current state, a represents the action taken, r represents the immediate reward, and s' represents the next state.

[0073] Initialize the target Q network, which is used to stabilize the target value during training and reduce fluctuations during training.

[0074] Initialize the main Q network, whose input is the current state space of the doors, windows and their environment, and output is the Q value estimate of each control action. The main Q network is the network that is continuously updated during the training process.

[0075] S2: Parameter setting

[0076] Set the learning rate α to control the learning speed to avoid instability caused by too fast learning or too long convergence time caused by too slow learning.

[0077] Set the discount factor γ to calculate the discounted value of future returns, reflecting the importance of long-term rewards.

[0078] The exploration rate ε is set to balance the relationship between exploration (randomly selecting actions) and utilization (selecting the current optimal action) to promote the exploration ability of the algorithm.

[0079] Set the number of training rounds to determine the total number of training iterations to ensure that the model is fully learned.

[0080] S3: Training iteration

[0081] For each training round, the following steps are performed:

[0082] S301: Initialize the current state s of doors, windows and their environmental parameters from the training data set.

[0083] S302: Select a control action a according to the current state s and the ε-greedy strategy. That is, randomly select an action with a probability of ε, and select the action with the largest Q value output by the current main Q network with a probability of 1-ε. This ensures exploration and utilizes the currently known optimal strategy.

[0084] S303: Execute action a, observe the next state s' and the immediate reward r. The immediate reward r is calculated according to the evaluation rule R(s,a) of the rule evaluation unit, taking into account factors such as the accuracy of state change, energy efficiency, environmental comfort and equipment service life.

[0085] S304: Store the experience (s, a, r, s') in the experience replay memory. The experience replay technology can improve the sample utilization rate, break the correlation between samples, and make the training more stable.

[0086] S305: Randomly sample a batch of experiences from the experience replay memory to update the main Q network. This can reduce the variance in the training process and improve the training efficiency.

[0087] S306: Use the sampled experience and the target Q network to calculate the target Q value. The target Q value is the optimal future reward estimate for the current action. Then, update the parameters of the main Q network according to the loss function of the DQN algorithm. The loss function is the mean square error between the predicted Q value and the target Q value, which is optimized by gradient descent.

[0088] S307: Every certain number of steps (such as every C steps), copy the parameters of the main Q network to the target Q network. This can stabilize the target value and reduce fluctuations during the training process.

[0089] S4: Repeat step S3 until the preset number of training rounds is reached or the output value of the main Q network converges (i.e., the Q value changes very little). This indicates that the model has been fully learned and can stably output the optimal control action.

[0090] S5: After the training is completed, the obtained main Q network is used as the door and window scheduling strategy model. For any given current state s of the door and window and its environment, the Q value of each control action is calculated through the forward propagation of the main Q network, and the action a with the largest Q value is selected as the optimal control action. In this way, the model can make intelligent decisions based on the real-time status and achieve efficient and accurate management of doors and windows.

[0091] Taking the door and window management of a smart office building as an example, the training process is as follows:

[0092] Initialize the experience replay memory, target Q network and main Q network.

[0093] The learning rate α is set to 0.001, the discount factor γ is set to 0.99, the exploration rate ε is initially set to 1.0 and gradually decreases to 0.1, and the number of training rounds is set to 10,000.

[0094] In each training round, a door and window in an office building and its environmental status are randomly selected as the starting point (such as indoor temperature, outdoor temperature, light intensity, etc.).

[0095] Select control actions (such as opening doors and windows, adjusting the opening degree, etc.) according to the current state and the ε-greedy strategy.

[0096] After performing the action, observe the next state and immediate reward (calculated based on energy efficiency, comfort, etc.).

[0097] The experiences are stored in the experience replay memory and a batch of experiences is randomly sampled to update the main Q network.

[0098] Every 100 steps, the parameters of the main Q network are copied to the target Q network.

[0099] After 10,000 rounds of training, the output value of the main Q network converged, and the door and window scheduling strategy model was obtained.

[0100] For any given door, window and its environmental status, this model can quickly select the optimal control action to achieve efficient and accurate door and window management.

[0101] Embodiment 3:

[0102] In order to build an efficient and accurate model, it is first necessary to prepare and preprocess the training data. This embodiment is used to describe the specific implementation of the collection, simulation, combination and preprocessing of the training data.

[0103] ① Real-time data acquisition and simulation control:

[0104] Deploy sensors and monitoring devices: Deploy various sensors and monitoring devices in the door and window cluster, such as temperature sensors, humidity sensors, light sensors, door and window status sensors, etc., to collect the actual status data of doors, windows and their environment in real time.

[0105] Execute simulated control actions: Combine the preset scheduling rules and evaluation criteria to execute simulated control actions on the collected status data. For example, simulate the actions of opening or closing doors and windows, and adjust the degree of opening of doors and windows according to the indoor temperature, outdoor temperature and light intensity.

[0106] Calculate instant rewards: Calculate the corresponding instant rewards based on the state changes after the simulated control action. The reward function can comprehensively consider multiple factors such as energy efficiency, environmental comfort, and equipment life, and give positive or negative rewards to reflect the quality of the action.

[0107] Combined state-action-reward tuple: The actual state data, simulated control actions and immediate rewards are combined into a state-action-reward tuple (s, a, r) ​​as the preliminary training data for the door and window scheduling strategy model.

[0108] ②Historical data utilization and simulation:

[0109] Construct historical status data set: Use the door and window operation records and environmental change data in the historical data to construct the historical status data set of doors, windows and their environment. These data can come from past monitoring records, equipment logs or weather forecast data, etc.

[0110] Simulate control actions and reward calculation: For each piece of data in the historical status data set, simulate possible control actions according to the preset scheduling rules and evaluation criteria, and calculate the corresponding immediate rewards.

[0111] Combine additional tuples: Combine historical state data, simulated control actions, and immediate rewards into additional state-action-reward tuples, add them to the preliminary training data, and form a complete training dataset for the door and window scheduling strategy model.

[0112] ③Preprocessing of training data:

[0113] Before model training begins, the training data set is preprocessed to improve the training efficiency and accuracy of the model.

[0114] The encoding process includes:

[0115] One-hot encoding conversion: Perform one-hot encoding conversion on discrete features in the training dataset. For example, the state of doors and windows (open, closed, partially open) can be converted into three one-hot encoding vectors [1,0,0], [0,1,0] and [0,0,1].

[0116] Standardization includes:

[0117] Z-score standardization: Z-score standardization method is used to standardize the continuous features in the training data set. Specifically, subtract the mean of each numerical data and divide it by its standard deviation so that the processed data conforms to the standard normal distribution. Standardization formula: The formula for standardization is: Z = (X-μ) / σ, where X is the original data, μ is the mean of the original data, σ is the standard deviation of the original data, and Z is the standardized data.

[0118] For example, for temperature data, first calculate the mean and standard deviation of all temperature values ​​in the training data set, and then apply the above formula to each temperature value for standardization. In this way, regardless of the range of the original temperature data, the processed data will have the same distribution characteristics, which is conducive to model training.

[0119] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0120] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. Intelligent door and window cluster monitoring and intelligent scheduling system based on cloud computing platform, characterized by: The system comprises: Door and window status monitoring layer, used to monitor the status of doors and windows and their environmental parameters in real time; The intelligent model building layer is used to build a door and window scheduling strategy model. The model models the management of door and window clusters as a multi-objective optimization problem and dynamically adjusts the opening and closing and adjustment schemes of doors and windows through algorithms. The intelligent model building layer further includes: The state definition component is used to describe the current state space of the door, window and its environment; Action definition component, used to define executable door and window control action space; The rule evaluation component is used to define evaluation rules and quantitatively evaluate the results of control actions based on the accuracy of door and window state changes, energy efficiency, environmental comfort, and equipment life; The strategy optimization unit uses a reinforcement learning algorithm to select the optimal control action under a given state and optimize the scheduling strategy through continuous iterative learning; The cloud control execution layer is used to upload the data collected by the door and window status monitoring module to the cloud. After being processed by the intelligent analysis model, it receives the control instructions issued by the cloud and drives the door and window actuators and related equipment to perform corresponding actions.

2. The intelligent door and window cluster monitoring and intelligent scheduling system based on the cloud computing platform according to claim 1 is characterized in that: The state space includes the open and closed status of doors and windows, location information, indoor temperature, outdoor temperature, humidity, wind speed, light intensity, noise level and air quality.

3. The intelligent door and window cluster monitoring and intelligent scheduling system based on cloud computing platform according to claim 1 is characterized in that: The action space includes opening doors and windows, closing doors and windows, adjusting the opening degree of doors and windows, and starting or closing sunshade equipment.

4. The intelligent door and window cluster monitoring and intelligent scheduling system based on cloud computing platform according to claim 3 is characterized in that: The implementation of the rule evaluation component includes: The evaluation rule R(s,a) is defined, where s represents the state space of doors, windows and their environment at the current moment; a represents the control action taken, which is selected from the control action space, including opening doors and windows, closing doors and windows, adjusting the opening degree of doors and windows, and starting or closing sunshade devices; the evaluation rule R(s,a) performs weighted summation based on the accuracy of state change C(s,a), energy efficiency E(s,a), environmental comfort H(s,a) and equipment service life D(s,a) to obtain the quantitative evaluation value of the control action result; the specific algorithm expression is: R(s,a)=w1*C(s,a)+w2*E(s,a)+w3*H(s,a)-w4*D(s,a) Wherein, w1, w2, w3, and w4 are weight coefficients of accuracy, energy efficiency, environmental comfort, and equipment life, respectively, and w1+w2+w3+w4=1; C(s,a) represents the accuracy of the door and window state change after taking action a under state s; E(s,a) represents energy efficiency, which is obtained by calculating the change in energy consumption; H(s,a) represents environmental comfort, which is comprehensively evaluated based on indoor temperature and humidity, light, and air quality; D(s,a) represents the consumption of equipment life, including the wear of doors, windows, and their accessories caused by actions.

5. The intelligent door and window cluster monitoring and intelligent scheduling system based on the cloud computing platform according to claim 4 is characterized in that: The strategy optimization component uses the Deep Q-Network algorithm in the deep reinforcement learning algorithm to train the door and window scheduling strategy model.

6. The intelligent door and window cluster monitoring and intelligent scheduling system based on cloud computing platform according to claim 1 is characterized by: Through various sensors and monitoring devices deployed in the door and window cluster, the actual status data of doors, windows and their environment are collected in real time; Combined with the preset dispatch rules and evaluation criteria, simulated control actions are performed on the collected status data, and the corresponding instant rewards are calculated; The actual state data, simulated control actions and immediate rewards are combined into state-action-reward tuples as the preliminary training data for the door and window scheduling strategy model.

7. The intelligent door and window cluster monitoring and intelligent scheduling system based on cloud computing platform according to claim 6 is characterized by: The historical status data set of doors and windows and their environment is constructed by using the door and window operation records and environmental change data in the historical data. For each piece of data in the historical status data set, possible control actions are simulated and the corresponding immediate rewards are calculated according to the preset scheduling rules and evaluation criteria. The historical status data, simulated control actions and immediate rewards are combined into additional state-action-reward tuples, which are added to the preliminary training data to form a complete training data set for the door and window scheduling strategy model.

8. The intelligent door and window cluster monitoring and intelligent scheduling system based on cloud computing platform according to claim 7 is characterized in that: The steps of the strategy optimization component training the door and window scheduling strategy model include: S1: Initialize the experience replay memory and the target Q network. The experience replay memory is used to store the experience of door and window state transfer, and the target Q network is used to stabilize the target value during training. At the same time, initialize the main Q network, whose input is the current state space of the door and window and its environment, and output is the Q value estimation of each control action. S2: Set the learning rate α, discount factor γ, exploration rate ε and number of training rounds, where α controls the learning speed, γ controls the discounted value of future returns, ε controls the balance between exploration and exploitation, and the number of training rounds determines the total number of training iterations; S3: For each training round, perform the following steps: S301: Initialize the current state s of doors, windows and their environmental parameters from the training data set; S302: Select a control action a according to the current state s and the ε-greedy strategy, that is, randomly select an action with a probability of ε, and select the action with the largest Q value output by the current main Q network with a probability of 1-ε; S303: Execute action a, observe the next state s' and the immediate reward r, where the immediate reward r is calculated according to the evaluation rule R(s,a) of the rule evaluation unit; S304: storing the experience (s, a, r, s') into the experience playback memory; S305: Randomly sample a batch of experiences from the experience replay memory for updating the main Q network; S306: using the sampled experience and the target Q network to calculate the target Q value, and updating the parameters of the main Q network according to the loss function of the Deep Q-Network algorithm, where the loss function is the mean square error between the predicted Q value and the target Q value; S307: copying the parameters of the main Q network to the target Q network at regular intervals; S4: Repeat step 3 until the preset number of training rounds is reached or the output value of the main Q network converges; S5: After the training is completed, the obtained main Q network is used as the door and window scheduling strategy model. For any given door and window and the current state s of its environment, the action a with the largest Q value is selected as the optimal control action through the forward propagation of the main Q network.

9. The intelligent door and window cluster monitoring and intelligent scheduling system based on cloud computing platform according to claim 7 is characterized in that: Before model training begins, the training data set is further preprocessed, including encoding and standardization.

10. The intelligent door and window cluster monitoring and intelligent scheduling system based on cloud computing platform according to claim 9 is characterized in that: Perform one-hot encoding conversion on discrete features; The Z-score standardization method is used to standardize continuous features. Specifically, each numerical data is subtracted from its mean and divided by its standard deviation so that the processed data conforms to the standard normal distribution. The standardization formula is: Z = (X-μ) / σ, where X is the original data, μ is the mean of the original data, σ is the standard deviation of the original data, and Z is the standardized data.

Citation Information

Cited By

  • Intelligent door and window control method and system

    CN120630768A