Smart city building cluster management method and system

By adopting a hybrid reinforcement learning energy consumption prediction model in the intelligent building management system, combining long and short-term memory networks and Q-Learning algorithms, the existing system has solved the problem of insufficient prediction accuracy and limited real-time response capabilities, and achieved efficient utilization of energy consumption and improved the system's real-time response capabilities.

CN120146403AActive Publication Date: 2025-06-13LIANYUNGANG WANCHANG TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510323780.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The existing intelligent building management system has problems in insufficient prediction accuracy, limited real-time response capabilities and lack of self-learning capabilities, making it difficult to effectively deal with dynamic changing environments and emergencies.

Method used

The hybrid reinforcement learning energy consumption prediction model is adopted, combined with long and short-term memory networks and Q-Learning algorithms, and the model weight is adjusted through dynamic coefficients to achieve accuracy and adaptability of energy consumption prediction, and the model is continuously optimized through online learning and incremental training mechanisms.

Benefits of technology

It significantly improves the prediction accuracy of energy consumption in complex environments such as extreme weather, reduces prediction error by 20%, realizes efficient utilization of energy, reduces comprehensive energy consumption by 10%, and improves the system's real-time response capability and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146403A_ABST
    Figure CN120146403A_ABST
Patent Text Reader

Abstract

The invention discloses a smart city building cluster management method and system, and relates to the technical field of smart cities. The method comprises the following steps: acquiring internal and external environment parameters, energy consumption data and people flow data of a building in real time through an Internet of Things sensor, cleaning and preprocessing the data, eliminating abnormal values and filling missing values; performing big data modeling and intelligent analysis by adopting a hybrid reinforcement learning energy consumption prediction model according to the cleaned data, and adjusting the model weight through a dynamic coefficient; based on the prediction result, the system carries out dynamic resource scheduling and decision support, when the predicted energy consumption is higher than a threshold value, the temperature of an air conditioner is automatically adjusted or unnecessary lighting is closed, and elevator scheduling is optimized in combination with people flow prediction; the system continuously optimizes the model through online learning, updates an action value function according to real-time feedback through a Q-Learning strategy, and performs incremental training on a long-short-term memory network model every 24 hours; and comparing the test verification model effect, and evaluating the prediction error, the energy saving rate and the user satisfaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of smart cities, and particularly relates to a method and system for managing a cluster of smart city buildings. Background Art

[0002] With the acceleration of the urbanization process, the traditional building management mode seems inadequate in meeting the requirements of high efficiency, energy conservation, and safety in modern cities. The traditional method relies on manual monitoring and adjustment, and there are problems such as low efficiency, resource waste, and safety hazards. Manual inspections and paper records lead to information lag and inability to respond to environmental changes in real time; the lack of real-time monitoring and data analysis of the internal and external environmental parameters of buildings results in energy waste; the safety monitoring system relies on manual inspections and simple alarm devices, making it difficult to detect potential hazards in real time, and the emergency response mechanism is imperfect. In recent years, the development of Internet of Things and big data technologies has provided new solutions for intelligent building management systems. By collecting environmental parameters, energy consumption, and pedestrian flow data in real time through sensor networks and combining big data analysis technologies, the intelligence and automation of building management can be achieved. However, existing systems still have problems such as insufficient prediction accuracy, limited real-time response ability, and lack of self-learning ability. Energy consumption prediction models are mostly based on static algorithms and are difficult to adapt to dynamic environments; the response to emergencies relies on fixed strategies and lacks the ability to dynamically adjust; the system cannot continuously optimize models and strategies based on historical data and real-time feedback. Therefore, the present invention proposes a method and system for managing a cluster of smart city buildings. Summary of the Invention

[0003] To solve the above technical problems, the present invention is implemented through the following technical solutions:

[0004] The present invention is a method for managing a cluster of smart city buildings, including the following steps;

[0005] Step 1: Collect the internal and external environmental parameters, energy consumption data, and pedestrian flow data of buildings in real time through Internet of Things sensors, and clean and preprocess the data, removing outliers and filling in missing values;

[0006] Step 2: According to the cleaned data, use a hybrid reinforcement learning energy consumption prediction model for big data modeling and intelligent analysis, and adjust the model weights through dynamic coefficients;

[0007] Step 3: Based on the prediction results, the system performs dynamic resource scheduling and decision support. When the predicted energy consumption is higher than the threshold, automatically adjust the air conditioner temperature or turn off unnecessary lighting, and optimize the elevator scheduling in combination with the pedestrian flow prediction;

[0008] Step 4: The system continuously optimizes the model through online learning. The Q-Learning strategy updates the action value function according to real-time feedback, and the long short-term memory network model is incrementally trained every 24 hours;

[0009] Step Five: Verify the model effect through comparative tests, and evaluate the prediction error, energy saving rate, and user satisfaction.

[0010] Further, in the first step, the building's internal and external environmental parameters, energy consumption data, and pedestrian flow data are collected in real time through the Internet of Things sensor network; the collected data is cleaned, outliers are removed, and missing data is filled. The 3σ principle is used to remove outliers exceeding the mean ± 3 times the standard deviation, and the missing data is completed by the time series linear interpolation method. Finally, the multi-source heterogeneous data is normalized to provide input data for modeling and analysis.

[0011] Further, in the second step, a long short-term memory network module is constructed. The energy consumption data Xt of the historical n time periods is input, and the time series prediction value is output. The formula is as follows:

[0012] LSTM(Xt) = fLSTM(Xt-n, Xt+1,..., Xt);

[0013] In the formula, LSTM(Xt) is the time series prediction value based on the historical data Xt, fLSTM is the input of the energy consumption data Xt of the historical n time periods, and the output is the time series prediction value. Xt-n, Xt+1,..., Xt represents the historical data sequence from time t-n to time t, and n is the length of the historical data used for prediction;

[0014] The Q-Learning module is introduced, and the current state st, the current action at, and the reward function R are defined. The formula of the reward function R is as follows:

[0015] R = -|Eactual(t+1) - Epred(t+1)|;

[0016] In the formula, Eactual(t+1) is the actual energy consumption value at time t+1, and Epred(t+1) is the predicted energy consumption value at time t+1;

[0017] The Q-Learning correction value formula is as follows:

[0018] α(t) × Q(st, at);

[0019] In the formula, α(t) represents the dynamic coefficient, which is used to adjust the weight of the Q-Learning correction value. Q(st, at) represents the output of the Q-Learning module, that is, the Q value of executing the action at in the state st;

[0020] The weight formula of α(t) is as follows:

[0021]

[0022] Wherein, k is a sensitivity parameter used to control the change speed of the dynamic coefficient, and ΔE(t) is the mean value of recent prediction errors, which is used to reflect the deviation of the model prediction;

[0023] Combining the long short-term memory network prediction value and the Q-Learning correction value to obtain the final energy consumption prediction value:

[0024] Epred(t + 1) = LSTM(Xt) + α(t) × Q(st, at);

[0025] Wherein, Epred(t + 1) is the predicted energy consumption value at time t + 1.

[0026] Furthermore, in the third step, based on the prediction result of the hybrid reinforcement learning energy consumption prediction model, it is judged whether the future energy consumption exceeds the preset threshold. If the predicted energy consumption is lower than the threshold, the current device state is maintained. If the predicted energy consumption is higher than the threshold, a resource scheduling strategy is triggered to automatically increase the air conditioner temperature setting value by 1°C and turn off unnecessary lighting. The decision formula is as follows:

[0027]

[0028] Wherein, Ethreshold is the preset energy consumption threshold, Adjust is to execute the energy-saving strategy, and Maintain is to keep the current device state unchanged.

[0029] Furthermore, in the fourth step, by collecting the latest energy consumption data, environmental parameters, and human flow data every 24 hours for model incremental training, the Q-Learning policy table is updated using the reward value. The real-time update formula is as follows:

[0030]

[0031] Wherein, η is the learning rate used to control the update speed of the Q value, and γ is the discount factor used to measure the importance of future rewards. is the maximum Q value for selecting the optimal action a in the next state st + 1.

[0032] Furthermore, in the fifth step, the hybrid reinforcement learning energy consumption prediction model is compared with the traditional ARIMA model to verify the improvement effect of the accuracy of the hybrid reinforcement learning energy consumption prediction and the degree of energy saving. The performance of the updated model is evaluated through verification experiments. If the effect is worse than the traditional ARIMA model, it is fed back to the fourth step to update the model again.

[0033] The smart city building cluster management system includes a data acquisition and processing module, a big data modeling and analysis module, a dynamic resource scheduling and decision-making module, a security monitoring and optimization module, and a system self-learning module;

[0034] The data acquisition and processing module is used to collect building internal and external environmental parameters, energy consumption data, and pedestrian flow data in real time, and perform cleaning, outlier removal, missing value filling, and standardization processing on the original data;

[0035] The big data modeling and analysis module receives the standardized data from the data acquisition and processing module. Based on the hybrid reinforcement learning energy consumption prediction model, combined with the long short-term memory network module and the Q-Learning algorithm, it is used to achieve dynamic and adaptive energy consumption prediction, and provide energy consumption demand prediction and pedestrian flow distribution prediction for a period of time in the future;

[0036] The dynamic resource scheduling and decision-making module dynamically adjusts the states of air conditioners, lighting equipment, and elevator operation strategies based on the energy consumption and pedestrian flow prediction results, generates energy-saving strategies and emergency evacuation plans, assists managers in decision-making, receives the prediction results from the big data modeling and analysis module, and outputs control instructions to building equipment;

[0037] The safety monitoring and optimization module receives real-time safety data from the sensor network, outputs alarm information to managers and the emergency system, monitors fires and safety hazards in real time through devices such as smoke detectors and cameras, triggers emergency responses, analyzes user behavior data, and optimizes elevator operation modes and water supply and power supply strategies;

[0038] The system self-learning module continuously collects historical and real-time data, trains machine learning models, and optimizes prediction accuracy. It receives the prediction errors from the big data modeling and analysis module and the scheduling effects of the dynamic resource scheduling and decision-making module, and outputs optimized model parameters.

[0039] The present invention has the following beneficial effects:

[0040] 1. Through the hybrid reinforcement learning energy consumption prediction model, combined with the long short-term memory network and the Q-Learning algorithm, the present invention can dynamically adjust the model weights, accurately predict future energy consumption demands. Compared with traditional static prediction models, the prediction accuracy of this model is significantly improved in complex environments such as extreme weather, and the prediction error is reduced by 20%. Based on the prediction results, the system can automatically adjust the air conditioner temperature, turn off unnecessary lighting equipment, and optimize elevator scheduling in combination with pedestrian flow prediction, thereby achieving efficient utilization of energy and reducing the comprehensive energy consumption by 10%.

[0041] 2. Through online learning and incremental training mechanisms, the system can continuously optimize the model according to real-time feedback and historical data. The Q-Learning strategy updates the action value function according to the real-time energy consumption error. The long short-term memory network model performs incremental training every 24 hours to ensure that the model can adapt to the dynamically changing environment. In addition, the system can monitor the environmental parameters, energy consumption data and passenger flow information inside and outside the building in real time, and dynamically adjust the resource scheduling strategy according to the prediction results, thereby improving the real-time response capability of the system, reducing the elevator waiting time by 15%, and improving user satisfaction.

[0042] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0044] Figure 1 It is a flow chart of the smart city building cluster management method of the present invention;

[0045] Figure 2 It is a flow chart of the smart city building cluster management system of the present invention. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] See also Figure 1 - Figure 2 As shown, the present invention is a smart city building cluster management method, comprising the following steps:

[0048] Step 1: Use IoT sensors to collect real-time building internal and external environmental parameters, energy consumption data, and pedestrian flow data, and clean and pre-process the data to remove outliers and fill in missing values;

[0049] Step 2: Based on the cleaned data, a hybrid reinforcement learning energy consumption prediction model is used to perform big data modeling and intelligent analysis, and the model weight is adjusted through dynamic coefficients;

[0050] Step 3: Based on the prediction results, the system performs dynamic resource scheduling and decision support. When the predicted energy consumption is higher than the threshold, it automatically adjusts the air conditioner temperature or turns off unnecessary lighting, and optimizes the elevator scheduling in combination with the crowd flow prediction.

[0051] Step 4: The system continuously optimizes the model through online learning. The Q-Learning strategy updates the action value function according to the real-time feedback, and the long short-term memory network model is incrementally trained every 24 hours.

[0052] Step 5: Verify the model effect through comparative tests, and evaluate the prediction error, energy saving rate, and user satisfaction.

[0053] In Step 1, the system collects the building's internal and external environmental parameters, energy consumption data, and crowd flow data in real time through the Internet of Things sensor network; cleans the collected data, removes outliers and fills in the missing data. The 3σ principle is used to remove outliers that exceed the mean ± 3 times the standard deviation, and the missing data is filled in by the time series linear interpolation method. Finally, the multi-source heterogeneous data is normalized to provide input data for modeling and analysis.

[0054] In Step 2, a long short-term memory network module is constructed. The energy consumption data Xt of the historical n time periods is input, and the time series prediction value is output. The formula is as follows:

[0055] LSTM(Xt) = fLSTM(Xt-n, Xt+1,..., Xt);

[0056] In the formula, LSTM(Xt) is the time series prediction value based on the historical data Xt, fLSTM is the input of the energy consumption data Xt of the historical n time periods and the output of the time series prediction value, Xt-n, Xt+1,..., Xt is the historical data sequence from time t-n to time t, and n is the length of the historical data used for prediction.

[0057] Introduce the Q-Learning module, define the current state st, the current action at, and the reward function R. The formula of the reward function R is as follows:

[0058] R = -|Eactual(t+1) - Epred(t+1)|;

[0059] In the formula, Eactual(t+1) is the actual energy consumption value at time t+1, and Epred(t+1) is the predicted energy consumption value at time t+1.

[0060] The Q-Learning correction value formula is as follows:

[0061] α(t) × Q(st, at);

[0062] In the formula, α(t) represents the dynamic coefficient, which is used to adjust the weight of the Q-Learning correction value. Q(st, at) represents the output of the Q-Learning module, that is, the Q value of executing action at in state st.

[0063] The weight formula of α(t) is as follows:

[0064]

[0065] In the formula, k is the sensitivity parameter, which is used to control the change speed of the dynamic coefficient, and ΔE(t) is the mean value of the recent prediction error, which is used to reflect the deviation of the model prediction.

[0066] Combining the long short-term memory network prediction value and the Q-Learning correction value, the final energy consumption prediction value is obtained:

[0067] Epred(t + 1) = LSTM(Xt) + α(t) × Q(st, at);

[0068] In the formula, Epred(t + 1) is the predicted energy consumption value at time t + 1.

[0069] In step three, based on the prediction result of the hybrid reinforcement learning energy consumption prediction model, it is judged whether the future energy consumption exceeds the preset threshold. If the predicted energy consumption is lower than the threshold, the current device state is maintained. If the predicted energy consumption is higher than the threshold, the resource scheduling strategy is triggered to automatically increase the air conditioner temperature setting value by 1°C and turn off unnecessary lighting. The decision formula is as follows:

[0070]

[0071] In the formula, Ethreshold is the preset energy consumption threshold, Adjust is to execute the energy-saving strategy, and Maintain is to keep the current device state unchanged.

[0072] In step four, by collecting the latest energy consumption data, environmental parameters, and pedestrian flow data every 24 hours for model incremental training, the Q-Learning policy table is updated using the reward value. The real-time update formula is as follows:

[0073]

[0074] In the formula, η is the learning rate, which is used to control the update speed of the Q value, and γ is the discount factor, which is used to measure the importance of future rewards. is the maximum Q value of selecting the optimal action a in the next state st + 1.

[0075] In Step 5, the hybrid reinforcement learning energy consumption prediction model is compared with the traditional ARIMA model to verify the improvement effect of the accuracy of hybrid reinforcement learning energy consumption prediction and the degree of energy saving. The performance of the updated model is evaluated through verification experiments. If the effect is worse than that of the traditional ARIMA model, it is fed back to Step 4 to update the model again.

[0076] The intelligent city building cluster management system includes a data acquisition and processing module, a big data modeling and analysis module, a dynamic resource scheduling and decision-making module, a safety monitoring and optimization module, and a system self-learning module.

[0077] The data acquisition and processing module is used to collect real-time environmental parameters inside and outside the building, energy consumption data, and pedestrian flow data, and perform cleaning, outlier removal, missing value filling, and standardization processing on the original data.

[0078] The big data modeling and analysis module receives the standardized data from the data acquisition and processing module, and based on the hybrid reinforcement learning energy consumption prediction model, combines the long short-term memory network module and the Q-Learning algorithm to achieve dynamic and adaptive energy consumption prediction, and provides energy consumption demand prediction and pedestrian flow distribution prediction for a period of time in the future.

[0079] The dynamic resource scheduling and decision-making module dynamically adjusts the states of air conditioners, lighting devices, and elevator operation strategies based on the energy consumption and pedestrian flow prediction results, generates energy-saving strategies and emergency evacuation plans, assists managers in decision-making, receives the prediction results from the big data modeling and analysis module, and outputs control instructions to building equipment.

[0080] The safety monitoring and optimization module receives real-time safety data from the sensor network, outputs alarm information to managers and the emergency system, monitors fires and safety hazards in real time through devices such as smoke detectors and cameras, triggers emergency responses, analyzes user behavior data, and optimizes elevator operation modes and water supply and power supply strategies.

[0081] The system self-learning module continuously collects historical and real-time data, trains machine learning models, optimizes prediction accuracy, receives the prediction errors from the big data modeling and analysis module and the scheduling effects of the dynamic resource scheduling and decision-making module, and outputs optimized model parameters.

[0082] A specific application of this embodiment is:

[0083] 1. Scenario description

[0084] The large commercial complex includes an office area, a shopping mall, and a hotel, and needs to achieve intelligent energy consumption management. The complex is equipped with devices such as temperature and humidity sensors, smart meters, water meters, and pedestrian flow statistical cameras to collect environmental parameters, energy consumption data, and pedestrian flow information in real time, and realizes dynamic energy consumption prediction and optimization through the hybrid reinforcement learning energy consumption prediction model of the present invention.

[0085] 2. Implementation Steps

[0086] Step 1: Data Collection and Preprocessing

[0087] Data Collection:

[0088] Temperature and Humidity Sensor: Collect indoor and outdoor temperature and humidity data every 5 minutes;

[0089] Smart Meter: Record power consumption data every 15 minutes;

[0090] People Flow Statistics Camera: Count the number of people entering and leaving every 10 minutes;

[0091] Data Cleaning:

[0092] Outlier Removal: The power consumption data at 12:00 noon on a certain day was 5000 kW, far exceeding the normal range (mean ± 3 times standard deviation), and was determined as an outlier and removed;

[0093] Missing Value Filling: The temperature and humidity data was missing in a certain period, and linear interpolation was used to complete it;

[0094] Data Standardization: Perform Z-Score standardization on temperature and humidity, energy consumption, and people flow data to eliminate the dimension difference;

[0095] Step 2: Hybrid Reinforcement Learning Energy Consumption Prediction

[0096] Model Input:

[0097] Historical Data: Temperature and humidity, energy consumption, and people flow data for the past 24 hours;

[0098] Real-Time Data: Environmental parameters and people flow density at the current time point;

[0099] Long Short-Term Memory Network Module:

[0100] Input historical 24-hour data and predict the energy consumption value for the next 1 hour;

[0101] Output: Preliminary prediction value Epred(t+1);

[0102] Q-Learning Module:

[0103] State st: Current energy consumption value, temperature and humidity, and people flow density;

[0104] Action at: Adjust the correction amplitude of the prediction value (±5%, ±10%);

[0105] Reward Function R: R = -|Eactual(t+1) - Epred(t+1)|;

[0106] Dynamic Coefficient α(t):

[0107] where Δt is the mean value of recent prediction errors, and k = 0.5;

[0108] Final predicted value: Epred(t+1) = LSTM(Xt) + α(t) × Q(st, at);

[0109] Step 3: Dynamic resource scheduling and decision support

[0110] Energy consumption optimization:

[0111] If it is predicted that the energy consumption in the next 1 hour will exceed the threshold (5000 kW), the system will automatically perform the following operations:

[0112] Raise the air-conditioning temperature setting by 1°C (adjusted from 24°C to 25°C);

[0113] Turn off non-essential lighting equipment in the mall (such as corridor lights, advertising light boxes);

[0114] Elevator scheduling optimization:

[0115] According to the passenger flow prediction model, it is predicted that 6 pm is the peak off-work period, and the system will increase the elevator operation frequency in advance to reduce the waiting time;

[0116] Step 4: System self-learning and upgrade

[0117] Q-Learning strategy update:

[0118] Update the Q-value table according to the error between the actual energy consumption and the predicted value:

[0119]

[0120] LSTM model incremental training:

[0121] Every 24 hours, fine-tune the LSTM network parameters with the latest data to ensure that the model adapts to dynamic changes;

[0122] Step 5: Verification and effectiveness evaluation

[0123] Evaluation indicators:

[0124] Prediction error: The MAE of the hybrid reinforcement learning energy consumption prediction model is 120 kW, which is 20% lower than that of the traditional ARI MA model (MAE = 150 kW);

[0125] Energy saving rate: Through dynamic scheduling, the comprehensive energy consumption is reduced by 8%;

[0126] User satisfaction: The elevator waiting time is reduced by 15%, and the indoor comfort score is increased by 10%;

[0127] Comparative test:

[0128] Compare the hybrid reinforcement learning energy consumption prediction model with the traditional model. The prediction accuracy of the hybrid reinforcement learning energy consumption prediction model is significantly improved under extreme weather conditions;

[0129] In the description of this specification, the description referring to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0130] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A smart city building cluster management method, characterized by: The steps include: Step 1: Use IoT sensors to collect real-time building internal and external environmental parameters, energy consumption data, and pedestrian flow data, and clean and pre-process the data to remove outliers and fill in missing values; Step 2: Based on the cleaned data, a hybrid reinforcement learning energy consumption prediction model is used to perform big data modeling and intelligent analysis, and the model weight is adjusted through dynamic coefficients; Step 3: Based on the prediction results, the system performs dynamic resource scheduling and decision support. When the predicted energy consumption is higher than the threshold, the system automatically adjusts the air conditioning temperature or turns off non-essential lighting, and optimizes elevator scheduling in combination with passenger flow prediction. Step 4: The system continuously optimizes the model through online learning. The Q-Learning strategy updates the action value function based on real-time feedback, and the long short-term memory network model is incrementally trained every 24 hours. Step 5: Verify the model effect through comparative testing and evaluate the prediction error, energy saving rate and user satisfaction.

2. The smart city building cluster management method according to claim 1, characterized in that: In the step 1, the environmental parameters inside and outside the building, energy consumption data and pedestrian flow data are collected in real time through the Internet of Things sensor network; the collected data is cleaned, outliers are removed and missing data are filled, the 3σ principle is used to remove outliers that exceed the mean ±3 times the standard deviation, and the missing data is filled by the time series linear interpolation method, and finally the multi-source heterogeneous data is normalized to provide input data for modeling and analysis.

3. The smart city building cluster management method according to claim 1, characterized in that: In the step 2, a long short-term memory network module is constructed, the energy consumption data Xt of the historical n time periods is input, and the time series prediction value is output. The formula is as follows: LSTM(Xt)=fLSTM(Xt-n,Xt+1,...,Xt); Where, LSTM(Xt) is the time series prediction value based on historical data Xt, fLSTM is the input of energy consumption data Xt of n historical time periods, and outputs the time series prediction value, Xt-n, Xt+1,..., Xt represents the historical data sequence from time tn to time t, and n is the length of historical data used for prediction; The Q-Learning module is introduced to define the current state st, the current action at and the reward function R. The formula of the reward function R is as follows: R=-Eactual(t+1)-Epred(t+1)|; Where, Eactual(t+1) is the actual energy consumption value at time t+1, and Epred(t+1) is the predicted energy consumption value at time t+1; The Q-Learning correction value formula is as follows: α(t)×Q(st,at); Where α(t) represents the dynamic coefficient, which is used to adjust the weight of the Q-Learning correction value, and Q(st,at) represents the output of the Q-Learning module, that is, the Q value of executing action at in state st; The α(t) weight formula is as follows: In the formula, k is the sensitivity parameter, which is used to control the change speed of the dynamic coefficient, and ΔE(t) is the mean of the recent prediction error, which is used to reflect the deviation of the model prediction; Combine the LSTM prediction value with the Q-Learning correction value to get the final energy consumption prediction value: Epred(t+1)=LSTM(Xt)+α(t)×Q(st,at); Where Epred(t+1) is the predicted energy consumption value at time t+1.

4. The smart city building cluster management method according to claim 1, characterized in that: In step 3, based on the prediction results of the hybrid reinforcement learning energy consumption prediction model, it is determined whether the future energy consumption exceeds the preset threshold. If the predicted energy consumption is lower than the threshold, the current device state is maintained. If the predicted energy consumption is higher than the threshold, the resource scheduling strategy is triggered to automatically increase the air conditioning temperature setting value by 1°C and turn off non-essential lighting. The decision formula is as follows: In the formula, Ethreshold is the preset energy consumption threshold, Adjust is to execute the energy-saving strategy, and Maintain is to keep the current device status unchanged.

5. The smart city building cluster management method according to claim 1, characterized in that: In step 4, the latest energy consumption data, environmental parameters and traffic data are collected every 24 hours for incremental model training, and the Q-Learning strategy table is updated using the reward value. The real-time update formula is as follows: In the formula, η is the learning rate, which is used to control the speed of Q value update, and γ is the discount factor, which is used to measure the importance of future rewards. The maximum Q value of selecting the optimal action a in the next state st+1.

6. The smart city building cluster management method according to claim 1, characterized in that: In the step 5, the hybrid reinforcement learning energy consumption prediction model is compared with the traditional ARIMA model to verify the accuracy improvement effect of the hybrid reinforcement learning energy consumption prediction and the degree of energy saving. The performance of the updated model is evaluated through verification experiments. If the effect is worse than that of the traditional ARIMA model, feedback is given to step 4 to re-update the model.

7. Smart city building cluster management system, characterized by: It includes data collection and processing module, big data modeling and analysis module, dynamic resource scheduling decision module, security monitoring and optimization module and system self-learning module; The data acquisition and processing module is used to collect real-time building internal and external environmental parameters, energy consumption data and human flow data, and clean the original data, remove abnormal values, fill in missing values ​​and perform standardization processing; The big data modeling and analysis module receives the standardized data from the data acquisition and processing module, and is used to realize dynamic and adaptive energy consumption prediction based on the hybrid reinforcement learning energy consumption prediction model, combined with the long short-term memory network module and the Q-Learning algorithm, and provide energy consumption demand prediction and crowd flow distribution prediction in the future period; The dynamic resource scheduling decision module dynamically adjusts the status of air conditioners and lighting equipment and elevator operation strategies based on the energy consumption and crowd flow prediction results, generates energy-saving strategies and emergency evacuation plans, assists management personnel in making decisions, receives prediction results from the big data modeling and analysis module, and outputs control instructions to building equipment; The safety monitoring and optimization module receives real-time safety data from the sensor network, outputs alarm information to management personnel and emergency systems, monitors fires and safety hazards in real time through smoke detectors, cameras and other equipment, triggers emergency responses, analyzes user behavior data, and optimizes elevator operation modes and water and power supply strategies; The system self-learning module continuously collects historical and real-time data, trains the machine learning model, receives the prediction error from the big data modeling and analysis module and the scheduling effect from the dynamic resource scheduling decision module, and outputs the optimized model parameters.

Citation Information

Patent Citations

  • Automatic driving positioning method and system based on LSTM deep reinforcement learning

    CN115840240A

  • Automatic filing method and system for personnel archives

    CN117827750A

  • Multi-mode reinforcement learning vehicle decision planning method with compensation feedback

    CN118917179A

  • Smart city energy optimization management method and system based on deep learning

    CN119313118A

  • Building energy consumption control system based on deep learning and control method thereof

    CN119575802A

Cited By

  • Energy consumption prediction and optimization system and method based on artificial intelligence

    CN120764744A

  • An artificial intelligence-based energy consumption prediction and optimization system and method

    CN120764744B

  • Urban risk management decision optimization method and system based on reinforcement learning

    CN121543779A

  • A method and system for optimizing urban risk governance decisions based on reinforcement learning

    CN121543779B