Urban centralized heating network optimization control system and method based on reinforcement learning
By introducing multi-source perception modules, edge computing nodes and cloud platform reinforcement learning engines into the urban centralized heating network, the problem of uneven heating is solved, the uniformity and energy efficiency of heating is improved, and user satisfaction is improved.
Patent Information
- Application Number
- CN202510485329.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, urban central heating systems have imbalance in heat distribution, resulting in significant temperature gradient differences in different buildings in the covered area of the same heating pipeline network, resulting in low energy utilization efficiency and reduced user comfort.
A system combining multi-source perception module, edge computing node and cloud platform reinforcement learning engine is adopted to dynamically adjust the heat source output, pump and valve opening and heat exchange station flow distribution through real-time data acquisition, preprocessing and global optimization strategy generation to achieve uniformity and energy efficiency optimization of the heating network.
It improves the uniformity of heating, reduces excessive heating, reduces energy consumption, and improves the prediction accuracy and user satisfaction of pump and valve failures.
Smart Images

Figure CN120335302A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of urban heating, and specifically to an optimized control system and method for urban central heating networks based on reinforcement learning. Background Art
[0002] In northern China, the operation of winter central heating systems often faces a technical challenge - the phenomenon of unbalanced heat distribution. Specifically, within the area covered by the same heat supply network, there are significant temperature gradient differences among different buildings: in some areas, the room temperature is too high due to excessive heat supply, while in other areas, the temperature is low due to insufficient heat transmission. This uneven heating not only destroys the comfortable experience of heat users but also causes low energy utilization efficiency, resulting in a large amount of heat energy being dissipated ineffectively.
[0003] Therefore, this application proposes an optimized control system and method for urban central heating networks based on reinforcement learning. Summary of the Invention
[0004] Therefore, this application provides an optimized control system and method for urban central heating networks based on reinforcement learning to solve the problem of uneven heating in the prior art.
[0005] To achieve the above object, this application provides the following technical solutions:
[0006] In a first aspect, an optimized control system for an urban central heating network based on reinforcement learning, characterized by comprising:
[0007] A multi-source perception module, deployed at key nodes of the heat supply network, for collecting water temperature, pressure, flow rate, room temperature, and equipment vibration data;
[0008] An edge computing node, connected to the multi-source perception module, for performing data cleaning, feature extraction, and local prediction;
[0009] A cloud platform reinforcement learning engine, generating a global optimization strategy based on the deep deterministic policy gradient algorithm, with input parameters including meteorological data, user load, and real-time electricity price;
[0010] A distributed actuator, dynamically adjusting the heat source output, pump valve opening, and heat exchange station flow distribution according to the optimization strategy.
[0011] Preferably, the multi-source perception module includes:
[0012] A high-precision temperature sensor, with a measurement error ≤ ±0.1°C and a sampling frequency of 1 Hz, covering the main pipeline and the user end;
[0013] A vibration monitoring unit, using a MEMS accelerometer and an acoustic emission sensor to detect abnormal vibrations of pump valves;
[0014] Wireless transmission unit, supporting LoRa and NB-IoT dual-mode communication to ensure that the data backhaul success rate is ≥99.9%.
[0015] Preferably, the edge computing node includes:
[0016] Local prediction model, predicting the future 1-hour pipe section heat load based on a lightweight LSTM network;
[0017] Anomaly detection algorithm, using the isolation forest algorithm to identify abnormal sensor data;
[0018] Data compression module, using Huffman coding and PCA dimensionality reduction technology to compress the original data to 30% of its volume and then upload it to the cloud.
[0019] Preferably, the design of the deep deterministic policy gradient algorithm of the cloud platform reinforcement learning engine includes:
[0020] State space, pipe network topology, real-time water temperature / pressure matrix, future 24-hour temperature prediction, time-of-use electricity price interval;
[0021] Action space, heat source output adjustment, pump valve opening adjustment, heat exchange station flow distribution weight;
[0022] Reward function, comprehensive heating cost, user comfort, equipment life loss.
[0023] Preferably, the distributed actuator includes:
[0024] Adaptive fuzzy PID controller, dynamically adjusting the proportional-integral-derivative parameters according to the cloud instructions, with a control accuracy of ±0.5%;
[0025] Redundant execution unit, main and standby dual-loop adjustment mechanism, single-point failure switching time ≤100ms;
[0026] Safety limiter module, forcibly restricting the execution instructions within the physical safety range.
[0027] An urban central heating network optimization control method based on reinforcement learning includes the following steps:
[0028] Step S1, the multi-source perception module collects pipe network data in real time, preprocesses it through the edge computing node, and then uploads it to the cloud;
[0029] Step S2, the cloud platform reinforcement learning engine generates a whole-network optimization strategy based on the deep deterministic policy gradient algorithm;
[0030] Step S3, the distributed actuator receives the policy instructions and adjusts the parameters of the heat source, pump valve, and heat exchange station;
[0031] Step S4: Cyclically update the state-action-reward data and online optimize the deep deterministic policy gradient policy network.
[0032] Preferably, in step S2, the Actor network of the deep deterministic policy gradient algorithm adopts a 4-layer fully connected structure, and the Critic network introduces an attention mechanism to preferentially process the state information of key pipeline segments.
[0033] Preferably, step S3 further includes a fault self-healing strategy;
[0034] When a leakage of a certain pipeline segment is detected, automatically reduce the pressure of the upstream pump valve, start the standby heat source and mark the maintenance area;
[0035] When the communication of a certain heat exchange station is interrupted, switch to the conservative operation mode preset by the edge node.
[0036] Preferably, the method further includes an energy efficiency report generation function:
[0037] Generate multi-dimensional energy efficiency reports on a daily / weekly / monthly basis, including heat consumption intensity, pipeline network loss rate, and equipment health index;
[0038] Visualization interface, marking hot spots, real-time warning information and optimization suggestions through a GIS map.
[0039] Compared with the prior art, the present application has at least the following beneficial effects:
[0040] 1. When the present invention is implemented, it can effectively improve the uniformity of heat supply, reduce the occurrence of excessive heat supply, and at the same time reduce energy consumption;
[0041] 2. Through the multi-source perception module, the information of the pump valve can be collected more accurately, the faults of the pump valve can be accurately predicted, and the user satisfaction is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more intuitively illustrate the prior art and the present application, exemplary drawings are given below. It should be understood that the specific shapes and structures shown in the drawings generally should not be regarded as limiting conditions when implementing the present application; for example, those skilled in the art are capable of making routine adjustments or further optimizations to the addition / deletion / attribution division of certain units (components), specific shapes, positional relationships, connection methods, dimensional proportional relationships, etc. based on the technical concept disclosed in the present application and the exemplary drawings.
[0043] Figure 1 It is a module diagram of the urban central heating network optimization control system based on reinforcement learning provided in the first embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The following further details the present application through specific embodiments in conjunction with the drawings.
[0045] As Figure 1 shown, an optimized control system for urban central heating network based on reinforcement learning includes:
[0046] Multi-source perception module: Deployed at key nodes of the heating pipe network, it collects water temperature, pressure, flow rate, room temperature and equipment vibration data;
[0047] Edge computing node: Connected to the multi-source perception module, it performs data cleaning, feature extraction and local prediction;
[0048] Cloud platform reinforcement learning engine: Generates a global optimization strategy based on the Deep Deterministic Policy Gradient (DDPG) algorithm, and the input parameters include meteorological data, user load, and real-time electricity price;
[0049] Distributed actuator: Dynamically adjusts the heat source output, pump valve opening and heat exchange station flow distribution according to the optimization strategy.
[0050] Through the edge-cloud collaborative architecture, the system response can be greatly shortened, the heating uniformity of the whole network can be improved, and the overheating rate can be reduced. The edge computing node preprocesses the data to reduce the cloud load; the DDPG algorithm dynamically balances energy efficiency and stability, avoiding the oscillation problem of traditional PID control.
[0051] The multi-source perception module includes: a high-precision temperature sensor with a measurement error ≤ ±0.1°C and a sampling frequency of 1Hz, covering the main pipeline and the user end;
[0052] Vibration monitoring unit, using MEMS accelerometer and acoustic emission sensor to detect abnormal vibration of pumps and valves, and the frequency range of the sensor is 20Hz - 10kHz;
[0053] Wireless transmission unit, supporting LoRa and NB-IoT dual-mode communication, ensuring that the data backhaul success rate ≥ 99.9%;
[0054] The integrity rate of multi-source data acquisition ≥ 98%, the vibration monitoring unit can predict pump and valve failures 2 hours in advance (accuracy rate ≥ 90%), and the wireless transmission unit ensures communication stability in complex pipe network environments and reduces the risk of data loss.
[0055] The edge computing node includes:
[0056] Local prediction model: Based on the lightweight LSTM network to predict the heat load of the pipe section in the next 1 hour;
[0057] Anomaly detection algorithm: Adopts the isolation forest algorithm to identify abnormal sensor data (such as sudden drop in temperature, sudden change in pressure);
[0058] The data compression module uses Huffman coding and PCA dimensionality reduction technology to compress the original data to 30% of its volume and then upload it to the cloud.
[0059] When the edge computing node is implemented, the edge prediction model reduces the cloud computing pressure, and the anomaly detection delay ≤ 10 seconds. The data compression saves bandwidth occupancy and is suitable for large-scale pipe network deployment.
[0060] The DDPG algorithm design of the cloud platform reinforcement learning engine includes:
[0061] State space: pipe network topology, real-time water temperature / pressure matrix, 24-hour future temperature prediction, time-of-use electricity price interval;
[0062] Action space: heat source output adjustment (±10%), pump valve opening adjustment (0% - 100%), heat exchange station flow distribution weight;
[0063] Reward function: comprehensive heating cost (weight 0.5), user comfort (PMV index deviation, weight 0.3), equipment life loss (weight 0.2); for example, in the reward function, the weight of the comprehensive heating cost is 0.5, the weight of user comfort (PMV index deviation) is 0.3, and the weight of equipment life loss is 0.2.
[0064] The convergence speed of the DDPG algorithm is 3 times faster than that of traditional algorithms. The design of the reward function greatly reduces the heating cost and extends the life of key equipment by 20%.
[0065] The distributed actuator includes:
[0066] Adaptive fuzzy PID controller: dynamically adjusts the proportional-integral-derivative parameters according to the cloud instructions, and the control accuracy of the controller is ±0.5%;
[0067] Redundant execution unit: master-slave dual-loop regulation mechanism, single-point failure switching time ≤ 100ms;
[0068] Safety limiter module: forcibly restricts the execution instructions within the physical safety range (such as water temperature ≤ 95°C).
[0069] The response time of the adaptive PID control ≤ 1 second. The redundant design makes the system availability ≥ 99.99%. The safety limiter avoids the risk of pipe network overload caused by human misoperation or algorithm anomalies.
[0070] An optimization control method for urban central heating network based on reinforcement learning includes the following steps:
[0071] Step S1: The multi-source perception module collects pipe network data in real time, preprocesses it through the edge computing node, and then uploads it to the cloud;
[0072] Step S2: The cloud platform reinforcement learning engine generates an optimization strategy for the entire network based on the DDPG algorithm;
[0073] Step S3: The distributed actuator receives the policy instruction and adjusts the parameters of the heat source, pump valve, and heat exchange station;
[0074] Step S4: Cyclically update the state-action-reward data and online optimize the DDPG policy network.
[0075] When the above method is implemented, the entire process realizes minute-level dynamic closed-loop optimization, significantly improves the comprehensive energy efficiency ratio of the heating network, reduces carbon emissions, and is applicable to large-scale heating networks.
[0076] In the step S2, the Actor network of the DDPG algorithm adopts a 4-layer fully connected structure (256-128-64-32 nodes), and the Critic network introduces an attention mechanism to preferentially process the state information of key pipe sections (such as the main pipe network).
[0077] The attention mechanism enables the algorithm to improve the regulation accuracy of the main pipe network and improves the network training efficiency.
[0078] The step S3 further includes a fault self-healing strategy:
[0079] If a leak is detected in a certain pipe section, automatically reduce the pressure of the upstream pump valve, start the standby heat source, and mark the maintenance area;
[0080] If the communication of a certain heat exchange station is interrupted, switch to the preset conservative operation mode (such as constant temperature water supply) of the edge node to ensure the basic heating demand and avoid a sudden drop in the room temperature of a large number of users.
[0081] The method further includes an energy efficiency report generation function:
[0082] Generate multi-dimensional energy efficiency reports on a daily / weekly / monthly basis, including heat consumption intensity (GJ / m2), pipeline network loss rate, and equipment health index;
[0083] Visualization interface: Mark hot spots, real-time warning information, and optimization suggestions through the GIS map.
[0084] The energy efficiency report helps managers quickly locate high-loss pipe sections (location accuracy ±50 meters), improve the efficiency of operation and maintenance decision-making, and the visualization interface reduces the technical threshold and is suitable for non-professional personnel to use.
[0085] The technical features of the above embodiments can be combined arbitrarily (as long as there is no contradiction in the combination of these technical features). For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described; these embodiments not explicitly written out should also be considered to be within the scope described in this specification.
Claims
1. An optimized control system for urban central heating networks based on reinforcement learning, characterized in that, It includes: A multi-source perception module, deployed at key nodes of the heat supply network, to collect water temperature, pressure, flow rate, room temperature, and equipment vibration data; An edge computing node, connected to the multi-source perception module, to perform data cleaning, feature extraction, and local prediction; A cloud platform reinforcement learning engine, which generates a global optimization strategy based on the deep deterministic policy gradient algorithm, and the input parameters include meteorological data, user load, and real-time electricity price; A distributed actuator, which dynamically adjusts the heat source output, pump valve opening, and heat exchange station flow distribution according to the optimization strategy, The multi-source perception module includes: A high-precision temperature sensor, with a measurement error ≤ ±0.1°C, a sampling frequency of 1Hz, covering the main pipeline and the user end; A vibration monitoring unit, using a MEMS accelerometer and an acoustic emission sensor to detect abnormal vibrations of pumps and valves; A wireless transmission unit, supporting LoRa and NB-IoT dual-mode communication, ensuring that the data backhaul success rate ≥ 99.9%; The edge computing node includes: A local prediction model, which predicts the heat load of the pipeline section in the next 1 hour based on a lightweight LSTM network; An anomaly detection algorithm, using the isolation forest algorithm to identify abnormal sensor data; A data compression module, using Huffman coding and PCA dimensionality reduction technology to compress the original data to 30% of its volume and then upload it to the cloud; The design of the deep deterministic policy gradient algorithm of the cloud platform reinforcement learning engine includes: The state space, the pipe network topology structure, the real-time water temperature / pressure matrix, the temperature prediction for the next 24 hours, and the electricity price time-sharing interval; The action space, the adjustment of the heat source output, the adjustment of the pump valve opening, and the flow distribution weight of the heat exchange station; The reward function, comprehensively considering the heating cost, user comfort, and equipment life loss.
2. The optimized control system for urban central heating network based on reinforcement learning according to claim 1, characterized in that: The distributed actuator includes: An adaptive fuzzy PID controller, which dynamically adjusts the proportional-integral-derivative parameters according to the cloud instructions, with a control accuracy of ±0.5%; A redundant execution unit, with a main-backup dual-loop adjustment mechanism, and the single-point failure switching time ≤ 100ms; A safety limiter module, which forcibly restricts the execution instructions within the physical safety range.
3. An optimized control method for urban central heating networks based on reinforcement learning, characterized in that, It includes the following steps: Step S1, the multi-source perception module collects the pipe network data in real time, and uploads it to the cloud after preprocessing by the edge computing node; Step S2, the cloud platform reinforcement learning engine generates a whole-network optimization strategy based on the deep deterministic policy gradient algorithm; Step S3, the distributed actuator receives the policy instructions and adjusts the parameters of the heat source, pump valve, and heat exchange station; Step S4, circularly update the state-action-reward data, and online optimize the deep deterministic policy gradient policy network.
4. The optimized control method for urban central heating network based on reinforcement learning according to claim 3, characterized in that: In step S2, the Actor network of the deep deterministic policy gradient algorithm adopts a 4-layer fully connected structure, and the Critic network introduces an attention mechanism to preferentially process the key pipe section state information.
5. The optimized control method for urban central heating network based on reinforcement learning according to claim 3, characterized in that: Step S3 also includes a fault self-healing strategy; When a leak in a certain pipe section is detected, automatically reduce the pressure of the upstream pump valve, start the standby heat source, and mark the maintenance area; When the communication of a certain heat exchange station is interrupted, switch to the conservative operation mode preset by the edge node.
6. The optimization control method for urban central heating network based on reinforcement learning according to claim 3, characterized in that: The method also includes an energy efficiency report generation function: Generate multi-dimensional energy efficiency reports daily / weekly / monthly, including heat consumption intensity, pipe network loss rate, and equipment health index; Visual interface, marking hotspots, real-time warning information and optimization suggestions on the GIS map.