Centralized heating secondary network regulation and control method and system based on reinforcement learning

By using reinforcement learning technology, key nodes and bottlenecks in the heating network are dynamically identified, and control parameters are optimized, which solves the problems of uneven heating and energy waste in the secondary heating network of centralized heating and achieves efficient and stable heating control.

CN121876503APending Publication Date: 2026-04-17YANTAI 500 HEATING LTD CO +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YANTAI 500 HEATING LTD CO
Filing Date
2025-11-17
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, the control methods for centralized heating secondary networks rely on static models and fixed rules, which leads to uneven heating and energy waste, and makes it impossible to achieve efficient operation.

Method used

By employing a reinforcement learning-based approach, key nodes and bottleneck nodes are dynamically identified through data collection, processing, and analysis. Path resistance and flow changes are calculated, user demand correlation analysis and branch path priority ranking are performed, control parameters are optimized, an energy transfer path model is generated, and system stability is assessed.

Benefits of technology

It enables precise control of the heating network, improves energy efficiency and thermal balance stability, solves the problems of uneven heating and energy waste caused by traditional control methods, and enhances the system's adaptive optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121876503A_ABST
    Figure CN121876503A_ABST
Patent Text Reader

Abstract

The invention discloses a centralized heating secondary network regulation and control method and system based on reinforcement learning, and relates to the technical field of reinforcement learning, and the method comprises the steps: collecting flow, temperature and pressure data in a pipe network, carrying out the denoising and standardization processing, and obtaining a pipe network data matrix; extracting a key node set, and screening out dynamic bottleneck nodes; calculating path resistance and flow variation in combination with the pipe network data matrix, generating energy flow characteristics, extracting branch influence factors, and performing user demand association analysis and branch path priority ranking; performing regulation and control parameter optimization and pipe network overall response simulation according to path priority ranking, and calculating an energy waste index to perform optimization node verification to obtain an optimization node set; performing energy transfer path model calculation according to the optimized node set to obtain an energy transfer path model; and extracting a regulation and control instruction sequence from the energy transfer path model, and determining system stability improvement configuration. According to the method, efficient operation of the heat supply pipe network is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of reinforcement learning technology, and in particular to a method and system for regulating a centralized heating secondary network based on reinforcement learning. Background Technology

[0002] Currently, the centralized heating secondary network is a crucial link in urban energy transmission, and its stable and efficient operation is vital for ensuring people's livelihoods and reducing energy consumption. By enhancing learning and continuous interaction with the operating environment, and optimizing strategies based on feedback, a feasible solution is provided for achieving precise control of the system.

[0003] In one existing technology, temperature and flow sensors are first deployed at key branches of the pipeline network to continuously collect data such as return water temperature. The system has a built-in database of rated parameters corresponding to the design conditions, which specifies the target flow and temperature values ​​for each branch under different outdoor temperatures. The control center compares the real-time data read by the sensors with this database of rated parameters. Once it detects that the return water temperature of a branch is lower than the target value under the current conditions, the system generates a control command to linearly increase the opening of the electric regulating valve on that branch to increase the hot water flow. The entire process is a feedback control based on a fixed mapping relationship, namely "temperature deviation → valve opening adjustment amount," and its control logic and regulation coefficients are set and fixed at the initial stage of system commissioning. Traditional control methods rely on static models and fixed rules for local compensation. This isolated adjustment will disrupt the original hydraulic balance of the pipeline network, causing operational disturbances and leading to new uneven heating and energy waste.

[0004] Therefore, efficient operation of heating pipe networks cannot be achieved using existing technologies. Summary of the Invention

[0005] This invention provides a method and system for regulating a centralized heating secondary network based on reinforcement learning, so as to achieve efficient operation of the heating network.

[0006] In a first aspect, to address the aforementioned technical problems, this invention provides a centralized heating secondary network control method based on reinforcement learning, comprising: Collect flow, temperature, and pressure data from the pipeline network to obtain the raw dataset; Based on the original dataset, denoising and standardization processes are performed to obtain the pipeline network data matrix; Based on the pipeline network data matrix, extract the set of key nodes containing node pressure fluctuations, and filter out dynamic bottleneck nodes; Based on the dynamic bottleneck nodes and the pipeline data matrix, path resistance and flow rate changes are calculated to generate energy flow characteristics. Based on the energy flow characteristics, branch influence factors are extracted, user demand correlation analysis is performed, and branch path priority is ranked to obtain path priority ranking. Based on the path priority sorting, the control parameters are optimized and the overall pipeline network response is simulated. Energy waste indicators are calculated to verify the optimized nodes, resulting in an optimized node set. Based on the optimized node set, the energy transfer path model is calculated to obtain the energy transfer path model; Extract the control command sequence from the energy transfer path model, determine the degree of coverage of the control command sequence to dynamic operating conditions, and determine the system stability improvement configuration.

[0007] In one optional implementation, based on the pipeline network data matrix, a set of key nodes containing node pressure fluctuations is extracted, and dynamic bottleneck nodes are screened out, including: Based on the pipeline data matrix, a pressure data matrix is ​​extracted. Combined with the pre-stored pipeline topology data, a graph attention network is used to score the importance of nodes, resulting in a set of key nodes. Based on the set of key nodes, extract the node pressure time series, and quantify the pressure fluctuation amplitude by calculating the standard deviation within a preset sliding time window to obtain a set of key nodes containing node pressure fluctuations. Based on the set of key nodes, key nodes whose pressure fluctuations exceed a preset pressure fluctuation threshold are extracted to obtain dynamic bottleneck nodes.

[0008] In one optional implementation, based on the dynamic bottleneck node and the pipeline data matrix, path resistance and flow rate changes are calculated to generate energy flow characteristics, including: Extract the pipeline data matrix corresponding to the dynamic bottleneck node, combine it with the pipeline topology data, and perform steady-state flow field simulation using the SIMPLE algorithm to obtain a set of path resistance distributions including pipeline segment pressure drop, pipeline flow rate and resistance value. Extract abnormal pipeline paths from the path resistance distribution set whose resistance values ​​exceed a preset resistance threshold, and calculate the unsteady flow field using the finite volume method to obtain the time series data of the flow variation of the abnormal pipeline paths. Based on the path resistance distribution set and the flow change time series data, the flow channel connectivity and energy transfer efficiency are extracted using Laplace eigenmaps to obtain energy flow characteristics.

[0009] In one optional implementation, branch influence factors are extracted based on the energy flow characteristics, user demand correlation analysis is performed, and branch path priority is ranked to obtain path priority ranking, including: Based on the energy flow characteristics and the pipeline topology data, the dynamic bottleneck nodes are divided into regions using the Kernighan-Lin graph segmentation algorithm to obtain bottleneck node partitions. Based on the bottleneck node partition and the path resistance distribution set, the resistance and flow of the branch path are simulated and calculated using the PISO algorithm to obtain the branch influence factor composed of the resistance coefficient and the average flow. By obtaining user load demand and combining it with the aforementioned branch influencing factors, nonlinear correlation analysis is performed by calculating the Spearman correlation coefficient to obtain the nonlinear correlation results. Based on the nonlinear correlation results, the priority of branch paths is ranked by combining the entropy weight method and the TOPSIS method to obtain the path priority ranking.

[0010] In one optional implementation, based on the path priority ranking, control parameters are optimized and the overall pipeline network response is simulated. Energy waste indicators are calculated to verify the optimized nodes, resulting in an optimized node set, including: Based on the path priority sorting and the traffic, the absolute difference between the preset target traffic and the traffic is calculated to obtain traffic deviation data; Based on the flow deviation data, the PISO algorithm is used to solve the coupled solution of non-isothermal flow and heat transfer, resulting in a set of heating distribution deviations composed of temperature deviation values ​​of each region. Based on the set of heating distribution deviations, a deep Q-network algorithm is used to iteratively optimize the strategy with the deviation value as the state and the valve opening as the action, to obtain the optimized sequence of control parameters. The control parameter optimization instruction is generated and executed according to the control parameter optimization sequence. Based on the flow rate, the heat supply area is divided using the METIS algorithm to obtain a heat balance adjustment scheme. Based on the control parameter optimization instructions in the thermal balance adjustment scheme, a transient simulation of the hydraulic and thermal coupling of the pipeline network is performed using the finite element method to obtain a simulated pipeline network response sequence that includes changes in node pressure and flow rate. The total power consumption of the system is obtained, and the simulated flow rate and simulated pressure are extracted by combining the simulated pipeline response sequence. The energy waste index is quantified by calculating the area of ​​the difference between the total power consumption of the system and the preset design power consumption using the trapezoidal numerical integration method, and the energy waste index value is obtained. When the energy waste index value is lower than the preset optimization target threshold, the K-means clustering algorithm is used to select performance improvement nodes from the simulated pipeline response sequence to obtain an optimized node set.

[0011] In one optional implementation, an energy transfer path model is calculated based on the optimized node set to obtain the energy transfer path model, including: Based on the optimized node set, the corresponding flow rate and pipeline topology data are extracted, and the node connection importance weight is calculated using the PageRank algorithm to obtain the node connection weight matrix. By acquiring time-series temperature data and combining it with the node connection weight matrix, a path optimization parameter set including path transmission efficiency is obtained through multiple rounds of message passing and state iteration updates via a gated graph neural network. Energy consumption data is acquired, and combined with the path optimization parameter set, the energy allocation coefficient of the path is calculated through linear weighted fusion to obtain the energy allocation coefficient; When the energy allocation coefficient is outside the preset equilibrium threshold range, the energy allocation coefficient is fed back to the gated graph neural network for a new round of iterative optimization. When the energy distribution coefficient is within a preset equilibrium threshold range, the node connection weights are solidified through a graph attention network to obtain an energy transfer path model.

[0012] In one optional implementation, a control command sequence is extracted from the energy transfer path model, the coverage of the control command sequence to dynamic operating conditions is determined, and a system stability improvement configuration is identified, including: The optimized sequence of control parameters is extracted based on the energy transfer path model, and the command pattern is matched with the pre-stored operating condition database using a dynamic time warping algorithm to obtain the operating condition matching command. Based on the operating condition matching command, the stability index is quantified by calculating the maximum eigenvalue of the system state matrix using the Lyapunov stability analysis method, and the system stability evaluation value is obtained. When the system stability assessment value is lower than the preset stability threshold, the Kalman filter algorithm is used to fuse the operating condition database and the node connection weight matrix to obtain an enhanced system state estimate. Based on the enhanced system state estimation, a stable set of control parameters is obtained by iterating the node states and optimizing the parameters through a graph attention network. Based on the set of stable control parameters, the optimal matching between energy allocation tasks and node carrying capacity is performed using the Hungarian algorithm to determine the system stability enhancement configuration.

[0013] Secondly, the present invention provides a centralized heating secondary network control system based on reinforcement learning, comprising: The data acquisition module is used to collect flow, temperature and pressure data in the pipeline network to obtain the raw dataset; The data processing module is used to perform denoising and standardization processing on the original dataset to obtain the pipeline network data matrix; The bottleneck node screening module is used to extract a set of key nodes containing node pressure fluctuations based on the pipeline network data matrix, and to screen out dynamic bottleneck nodes. The energy flow analysis module is used to calculate path resistance and flow rate changes based on the dynamic bottleneck nodes and the pipeline data matrix, and generate energy flow characteristics. The path ranking module is used to extract branch influence factors based on the energy flow characteristics, perform user demand correlation analysis, and rank the branch paths to obtain the path priority ranking. The optimization node analysis module is used to optimize control parameters and simulate the overall response of the pipeline network according to the path priority sorting, and to calculate energy waste indicators to verify the optimization nodes and obtain the optimization node set. The energy transfer path construction module is used to calculate the energy transfer path model based on the optimized node set, and obtain the energy transfer path model. The output module is used to extract the control command sequence from the energy transfer path model, determine the coverage of the control command sequence to the dynamic operating conditions, and determine the system stability improvement configuration.

[0014] Compared with the prior art, the present invention has the following beneficial effects: (1) A graph attention network is used to score the importance of nodes, and the pressure fluctuation is quantified by calculating the standard deviation within the sliding time window. Key nodes and bottleneck nodes are dynamically screened to solve the problem that traditional static models are difficult to capture the dynamic changes of the pipeline network, and improve the real-time identification capability and control response speed of key links.

[0015] (2) The SIMPLE algorithm is used to simulate the steady flow field, and the finite volume method is used to calculate the unsteady flow field. The energy flow characteristics are extracted by the Laplace eigenmap, and the branch influence factor is calculated based on the Kernighan-Lin diagram segmentation and the PISO algorithm. The path priority is sorted by combining the entropy weight method and the TOPSIS method to solve the control lag problem caused by the non-uniform energy distribution, and to achieve accurate sorting of branch paths and efficient energy allocation.

[0016] (3) The strategy is iteratively optimized by using the deep Q-network algorithm with the heating deviation as the state and the valve opening as the action. The balanced heating area is divided by combining the METIS algorithm. The hydraulic-thermal coupling transient simulation is carried out by the finite element method. The energy waste index is quantified by the trapezoidal numerical integration method. The optimized nodes are selected by K-means clustering to solve the problem of uneven heating and energy waste caused by local regulation, and improve the overall energy efficiency and thermal balance stability of the pipeline network.

[0017] (4) The gated graph neural network is used to perform multi-round message passing and state iteration update. The PageRank algorithm is combined to calculate the node connection weight. The dynamic time warping algorithm is used to perform instruction pattern matching. The Lyapunov stability analysis method and Kalman filter algorithm are used to evaluate and enhance the system state estimation. The optimal matching of energy allocation task and node carrying capacity is performed based on the Hungarian algorithm to solve the problem of insufficient system stability under dynamic working conditions and realize the adaptive optimization of control instructions and effective improvement of system stability. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the centralized heating secondary network control method based on reinforcement learning provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a centralized heating secondary network control system based on reinforcement learning, provided in the second embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Reference Figure 1 The first embodiment of the present invention provides a centralized heating secondary network control method based on reinforcement learning, comprising the following steps: S11, collect flow, temperature and pressure data in the pipeline network to obtain the raw dataset; S12, Based on the original dataset, perform denoising and standardization processing to obtain the pipeline network data matrix; S13, Based on the pipeline data matrix, extract the set of key nodes containing node pressure fluctuations, and filter out dynamic bottleneck nodes; S14, Calculate path resistance and flow rate change based on the dynamic bottleneck node and the pipeline data matrix, and generate energy flow characteristics; S15, extract branch influence factors based on the energy flow characteristics, perform user demand correlation analysis, and sort the branch path priorities to obtain the path priority ranking. S16. Based on the path priority sorting, optimize the control parameters and simulate the overall response of the pipeline network, and calculate the energy waste index to verify the optimized nodes, thereby obtaining the optimized node set. S17. Based on the optimized node set, calculate the energy transfer path model to obtain the energy transfer path model. S18, extract the control command sequence from the energy transfer path model, determine the coverage of the control command sequence to the dynamic operating conditions, and determine the system stability improvement configuration.

[0021] In step S11, flow rate, temperature and pressure data in the pipeline network are collected to obtain the raw dataset.

[0022] Specifically, temperature, flow, and pressure sensors deployed at key branches of the centralized heating secondary network are used to collect real-time flow, temperature, and pressure data from the pipeline network, forming a raw dataset. This raw dataset contains the flow rate (e.g., cubic meters per hour), temperature (e.g., degrees Celsius), and pressure (e.g., kilopascals) values ​​for each monitoring point at each sampling time point. This data is stored in time-series format and includes monitoring point identifiers and timestamps. Collecting this raw data provides the foundation for subsequent data processing and analysis, ensuring that control methods can make decisions based on the real-time operating status of the pipeline network, thereby providing accurate input information for the reinforcement learning model.

[0023] In step S12, the original dataset is denoised and standardized to obtain the pipeline data matrix.

[0024] Specifically, in step S12, the original dataset obtained from step S11 is denoised and standardized to obtain a pipeline data matrix. The original dataset contains three time-series variables: original flow rate, original temperature, and original pressure. Each variable contains sampled values ​​from monitoring points at consecutive time points. The data processing module first uses a sliding window averaging method for denoising. This method sets a fixed-length sliding time window, the size of which is determined based on typical noise cycles in historical data statistics, for example, selecting five consecutive sampling points as a window. For each data point in the original dataset, the arithmetic mean of the values ​​of that point and its five adjacent points is calculated. This mean is used to replace the value of the original data point at its current position, thereby generating the denoised flow rate, denoised temperature, and denoised pressure time series. This denoising process effectively smooths out random fluctuations introduced by instantaneous sensor disturbances or transmission interference, providing a more stable data foundation for subsequent analysis. Subsequently, the denoised data is standardized to eliminate the influence of different physical dimensions of flow rate, temperature, and pressure. Standardization processing is performed separately for each variable: denoised flow rate, denoised temperature, and denoised pressure. The calculation process uses the global mean (e.g., mean flow rate, mean temperature, mean pressure) and global standard deviation (e.g., standard deviation of flow rate, standard deviation of temperature, standard deviation of pressure) of each variable, obtained from historical normal operation data. For each value in the denoised sequence, the corresponding global mean is subtracted, and then divided by the corresponding global standard deviation to obtain the standardized flow rate, standardized temperature, and standardized pressure values. Finally, the standardized data from all monitoring points are aligned by time and organized into a two-dimensional pipeline data matrix. The rows of the matrix correspond to the sampling time points, and the columns correspond to the standardized flow rate, standardized temperature, and standardized pressure variables of different monitoring points. This pipeline data matrix provides standardized input with unified dimensions and noise suppression for subsequent feature extraction, node selection, and model calculation. It is a key preprocessing step to ensure that the reinforcement learning model accurately identifies the system state and makes effective control decisions.

[0025] In step S13, based on the pipeline data matrix, a set of key nodes containing node pressure fluctuations is extracted, and dynamic bottleneck nodes are selected.

[0026] In one specific implementation, the step of extracting a set of key nodes containing node pressure fluctuations based on the pipeline network data matrix, and filtering out dynamic bottleneck nodes, includes: Based on the pipeline data matrix, a pressure data matrix is ​​extracted. Combined with the pre-stored pipeline topology data, a graph attention network is used to score the importance of nodes, resulting in a set of key nodes. Based on the set of key nodes, extract the node pressure time series, and quantify the pressure fluctuation amplitude by calculating the standard deviation within a preset sliding time window to obtain a set of key nodes containing node pressure fluctuations. Based on the set of key nodes, key nodes whose pressure fluctuations exceed a preset pressure fluctuation threshold are extracted to obtain dynamic bottleneck nodes.

[0027] Specifically, in step S13, based on the pipeline data matrix obtained from step S12, a set of key nodes containing node pressure fluctuations is extracted, and dynamic bottleneck nodes are screened out. The data source for this step is a pipeline data matrix containing standardized pressure data from all monitoring points, as well as pipeline topology data pre-stored in the system database. This topology data defines the adjacency relationships of all nodes and pipe connections in the pipeline network in a graph structure.

[0028] First, standardized pressure data is extracted from the pipeline network data matrix to form a pressure data matrix. This pressure data matrix, along with the pipeline network topology data, is then input into a pre-constructed graph attention network. This graph attention network is a pre-trained model. Its training process uses historical pressure data and the pipeline network topology as input, with node importance scores calculated using the PageRank algorithm as training labels. By minimizing the mean squared error loss between the predicted score and the label, and using the Adam optimizer to update the network weights, a model capable of evaluating node importance is finally obtained. This network treats each node in the pipeline network topology as a graph node, pipeline connections as edges, and the standardized pressure values ​​of the corresponding nodes in the pressure data matrix as node features. Graph attention networks determine weights by calculating attention coefficients between nodes. For each target node, its attention coefficient with each neighbor node is obtained by concatenating the feature vectors of the target node and the neighbor nodes, multiplying the concatenated vector by a learnable weight vector, and then processing it through the non-linear activation function LeakyReLU. The attention coefficients of all neighbor nodes are then normalized using the Softmax function, resulting in the normalized attention weight of each neighbor node relative to the target node. Using the calculated normalized attention weights, the features of all neighbor nodes are weighted and summed, then linearly transformed using the ReLU non-linear activation function to update the feature representation of the target node. This process is iterated over multiple layers in the network, with attention weights calculated and node features updated at each layer. The network ultimately outputs a node importance score for each node. All node importance scores are sorted in descending order, and the top-ranked nodes are selected to form the initial set of key nodes.

[0029] Subsequently, from the initial set of key nodes, the standardized pressure time series corresponding to each node in the pressure data matrix is ​​extracted, i.e., the node pressure time series. For the pressure time series of each node, a sliding time window of a preset width is used to quantify the pressure fluctuation amplitude. The size of this sliding window is determined by analyzing historical pressure data: the autocorrelation function of historical pressure data at different time lags is calculated, and the value of the autocorrelation function exceeding a threshold calculated based on the historical autocorrelation function sequence is set as the 95th percentile of the historical autocorrelation function sequence. The time lag length corresponding to this peak is taken as the typical hydraulic fluctuation period of the system, and this period length is set as the size of the sliding window. The window is gradually slid along the time axis, and within each window, the standard deviation of the node pressure time series data for that segment is calculated as the node pressure fluctuation amplitude value corresponding to that time point. After traversing the entire time series, each key node obtains a node pressure fluctuation amplitude sequence corresponding to its node pressure time series. The calculated node pressure fluctuation amplitude is used as a new node attribute and attached to the initial set of key nodes to form a key node set containing node pressure fluctuations.

[0030] Finally, dynamic bottleneck node screening is performed based on the set of critical nodes containing node pressure fluctuations. The system sets a pressure fluctuation threshold, which is determined by calculating the distribution of pressure fluctuation amplitude data for all nodes under historical normal operating conditions and the distribution of pressure fluctuation amplitude data for all nodes under known bottleneck conditions, and then averaging the 95th percentile of the distribution under normal operating conditions and the 50th percentile of the distribution under bottleneck conditions. The average node pressure fluctuation amplitude of each critical node (i.e., the mean of its node pressure fluctuation amplitude sequence) is compared with this pressure fluctuation threshold. Critical nodes whose average node pressure fluctuation amplitude is consistently greater than the pressure fluctuation threshold are selected; these nodes are the dynamic bottleneck nodes, forming the dynamic bottleneck node set.

[0031] By combining real-time pressure data with pipeline topology, nodes with significant pressure fluctuations and located in critical positions are dynamically identified. This provides crucial input for accurately locating energy transfer bottlenecks and implementing targeted regulation, overcoming the shortcomings of traditional methods that rely on static analysis and cannot adapt to dynamic changes in operating conditions.

[0032] In step S14, path resistance and flow rate change are calculated based on the dynamic bottleneck node and the pipeline data matrix to generate energy flow characteristics.

[0033] In one specific implementation, the step of calculating path resistance and flow rate changes based on the dynamic bottleneck node and the pipeline data matrix to generate energy flow characteristics includes: Extract the pipeline data matrix corresponding to the dynamic bottleneck node, combine it with the pipeline topology data, and perform steady-state flow field simulation using the SIMPLE algorithm to obtain a set of path resistance distributions including pipeline segment pressure drop, pipeline flow rate and resistance value. Extract abnormal pipeline paths from the path resistance distribution set whose resistance values ​​exceed a preset resistance threshold, and calculate the unsteady flow field using the finite volume method to obtain the time series data of the flow variation of the abnormal pipeline paths. Based on the path resistance distribution set and the flow change time series data, the flow channel connectivity and energy transfer efficiency are extracted using Laplace eigenmaps to obtain energy flow characteristics.

[0034] Specifically, firstly, a subset of the pipeline network data matrix related to dynamic bottleneck nodes is extracted, and combined with the pipeline network topology data, a steady-state flow field simulation is performed using the SIMPLE algorithm. The specific execution process of the SIMPLE algorithm is as follows: a computational grid is constructed based on the pipeline network topology, dividing the pipeline into several computational units; the predicted pressure field and predicted velocity field of all computational units are initialized; then the iterative solution process begins. In each iteration, the momentum discrete equation is first solved. This equation is the fluid motion equation derived based on Newton's second law, describing the relationship between the rate of change of fluid momentum and the pressure gradient, viscous force, and source term. After solving, the velocity prediction value based on the current predicted pressure field is obtained. Next, a pressure correction equation is constructed based on the law of conservation of mass, which establishes the relationship between the velocity correction value and the pressure correction value. The pressure correction equation is solved to obtain the pressure correction value, which is used to correct the current predicted pressure field and velocity prediction value, resulting in the updated pressure field and updated velocity field for this iteration. The above steps constitute a complete iteration. The convergence criterion for the iterative process is set as follows: the maximum relative change of the updated pressure field and the maximum relative change of the updated velocity field in all computational units compared to the previous cycle must be less than a preset convergence threshold. This threshold is determined by statistical analysis of historical successfully converged simulation data, taking the 99th percentile of the relative changes in the field quantities between adjacent iteration steps for all computational units. The iteration terminates when this convergence criterion is met. Through this steady-state simulation calculation, a path resistance distribution set is finally obtained, containing the pressure drop of each pipeline segment, the pipeline flow rate, and the pipeline resistance value calculated based on Darcy's formula.

[0035] Next, abnormal pipeline paths are identified from the path resistance distribution set. The system sets a resistance threshold, and the criteria and steps for setting this threshold are as follows: First, all historical data explicitly marked as normal and stable operation are collected and filtered to form a clean baseline dataset. Then, the resistance values ​​of all pipelines in this baseline dataset are calculated and sorted from smallest to largest. Next, based on the statistical process control principles widely used in engineering practice, to effectively identify anomalies while preserving the normal operating volatility of the system, the 95th percentile is selected as the threshold. This means that this threshold covers 95% of the pipeline resistance values ​​in the baseline dataset within the normal range. Finally, in this sorted resistance value sequence, the resistance value corresponding to the 95th percentile of the total data is taken as the final resistance threshold. Pipe segments with resistance values ​​exceeding this resistance threshold are filtered out and marked as abnormal pipeline paths. For these abnormal pipeline paths, unsteady flow fields are calculated using the finite volume method: a set of transient flow control equations based on the mass conservation equation and the momentum conservation equation is established; the computational units of each abnormal pipeline path and its associated region are spatially discretized, transforming the partial differential equations into a set of algebraic equations; a time-progression method is used to solve the problem, dividing the simulation time into small time steps, and solving the discretized momentum equation and continuity equation sequentially within each time step to update the velocity and pressure values ​​of each computational unit; the flow rate values ​​of key sections of each abnormal pipeline path are recorded as a function of time throughout the entire simulation period to form time-series data of flow rate changes in abnormal pipeline paths.

[0036] Finally, based on the path resistance distribution set and the time-series data of abnormal pipeline path flow changes, features are extracted using the Laplace eigenmap method. A graph structure is constructed with network nodes as vertices and pipeline connections as edges, using path resistance values ​​as edge weights and the principal components of the flow change time-series data as node features. The normalized Laplace matrix of the graph is calculated, and eigenvalue decomposition is performed on the matrix. Eigenvectors corresponding to the smallest eigenvalues ​​(excluding zero) are selected. These eigenvectors constitute a low-dimensional embedding space, where the distance between adjacent nodes in the embedding space reflects the channel connectivity, and the distribution pattern of the eigenvectors reflects the energy transfer efficiency. These eigenvectors are concatenated in node order to form the final energy flow feature vector.

[0037] By combining steady-state and unsteady flow analysis, the inherent laws of energy flow in pipeline networks are revealed, providing a quantitative basis for subsequent branch path priority ranking and control parameter optimization, and solving the problem that traditional methods are difficult to accurately characterize the dynamic energy transmission characteristics of complex pipeline networks.

[0038] In step S15, branch influence factors are extracted based on the energy flow characteristics, user demand correlation analysis is performed, and branch path priority is sorted to obtain path priority ranking.

[0039] In one specific implementation, the step of extracting branch influence factors based on the energy flow characteristics, performing user demand correlation analysis, and ranking branch paths to obtain a path priority ranking includes: Based on the energy flow characteristics and the pipeline topology data, the dynamic bottleneck nodes are divided into regions using the Kernighan-Lin graph segmentation algorithm to obtain bottleneck node partitions. Based on the bottleneck node partition and the path resistance distribution set, the resistance and flow of the branch path are simulated and calculated using the PISO algorithm to obtain the branch influence factor composed of the resistance coefficient and the average flow. By obtaining user load demand and combining it with the aforementioned branch influencing factors, nonlinear correlation analysis is performed by calculating the Spearman correlation coefficient to obtain the nonlinear correlation results. Based on the nonlinear correlation results, the priority of branch paths is ranked by combining the entropy weight method and the TOPSIS method to obtain the path priority ranking.

[0040] Specifically, the data sources for this step include: energy flow characteristic vectors, dynamic bottleneck node sets, path resistance distribution sets, pipeline topology data, and user load demand data obtained through data interface communication with the centralized heating system monitoring and data acquisition system. This data is provided in the form of a structured data table containing timestamps, area identifiers, and instantaneous heat load values, and the data update frequency is consistent with the system control cycle. This data includes the heat load values ​​of each heating area at different times.

[0041] First, based on energy flow characteristics and pipeline topology data, the dynamic bottleneck nodes are partitioned into regions using the Kernighan-Lin graph segmentation algorithm. The specific execution process is as follows: A weighted graph is constructed using pipeline nodes as vertices, pipeline connections as edges, and the Euclidean distance of the energy flow feature vectors as edge weights. The set of dynamic bottleneck nodes is randomly divided into two initial partitions of approximately equal size. Then, an iterative optimization process begins. In each iteration, the cut weight gain that moving all nodes to the other partition might bring is calculated. The node with the largest gain is selected for movement and marked as moved. This process is repeated until all nodes have been moved once. During the movement process, the solution with the optimal cut weight is recorded, ultimately yielding the bottleneck node partitioning result that minimizes the cut weight.

[0042] Next, based on the bottleneck node partitions and path resistance distribution sets, the resistance and flow rate of the branch paths are simulated and calculated using the PISO algorithm. Specifically, within each partition, a set of governing equations describing the fluid flow is established based on the pipe resistance values ​​obtained from the path resistance distribution set. This set of governing equations includes the mass conservation equation and the momentum conservation equation. The mass conservation equation describes the continuity of the fluid in the pipe network, while the momentum conservation equation is based on Newton's second law, its core being the balancing of pressure gradient, inertial force, viscous force, and friction loss characterized by the pipe resistance values. The pipe resistance values ​​are correlated with the square of the fluid velocity using Darcy's formula, and are introduced as source terms in the momentum conservation equation. The PISO algorithm is then used for solving the problem: First, in the prediction step, the discretized momentum equation is solved to obtain the predicted velocity value. Then, in the first correction step, the pressure Poisson equation derived from the mass conservation equation is solved to obtain the preliminary pressure correction value and to make the first correction to the velocity and pressure fields. Next, a second correction is performed, solving the pressure Poisson equation again and updating the flow field to improve the accuracy of pressure-velocity coupling. This process is repeated multiple times until the flow field converges. Based on the converged flow field data, the drag coefficient and average flow rate of each branch path are calculated, and these values ​​are combined to form the branch influence factor.

[0043] Then, user load demand data is acquired, and combined with branch influence factors, nonlinear correlation analysis is performed by calculating the Spearman correlation coefficient. Specifically, the drag coefficient sequence and average flow sequence in the branch influence factors are sorted with their corresponding user load demand sequences; the sum of squares of the grade differences for each pair of sequences is calculated; and the Spearman correlation coefficient is calculated based on this sum of squares. This coefficient reflects the degree of monotonic correlation between the branch influence factors and user demand, forming a nonlinear correlation result.

[0044] Finally, based on the nonlinear correlation results, the priority ranking of branch paths is performed by combining the entropy weight method and the TOPSIS method. First, the weight of each evaluation indicator is determined using the entropy weight method: an evaluation matrix including drag coefficient, average flow rate, and Spearman correlation coefficient is constructed; the information entropy value of each indicator is calculated; and the weight coefficient of each indicator is calculated based on the information entropy. Then, the TOPSIS method is applied for ranking: a weighted normalized matrix is ​​constructed based on the weight coefficients; the positive and negative ideal solutions of each indicator are determined; the Euclidean distance from each branch path to the positive and negative ideal solutions is calculated; and the paths are ranked according to their relative proximity to obtain the path priority ranking.

[0045] By comprehensively considering the relationship between the physical characteristics of the pipeline network and user needs, accurate decision-making basis is provided for subsequent optimization of control parameters, thus resolving the contradiction between system operating efficiency and user needs that is difficult to balance using traditional methods.

[0046] In step S16, the control parameters are optimized and the overall response of the pipeline network is simulated according to the path priority sorting, and the energy waste index is calculated to verify the optimized nodes, thereby obtaining the optimized node set.

[0047] In one specific implementation, the process of optimizing control parameters and simulating the overall pipeline network response based on the path priority ranking, and calculating energy waste indicators to verify the optimized nodes, results in an optimized node set, including: Based on the path priority sorting and the traffic, the absolute difference between the preset target traffic and the traffic is calculated to obtain traffic deviation data; Based on the flow deviation data, the PISO algorithm is used to solve the coupled solution of non-isothermal flow and heat transfer, resulting in a set of heating distribution deviations composed of temperature deviation values ​​of each region. Based on the set of heating distribution deviations, a deep Q-network algorithm is used to iteratively optimize the strategy with the deviation value as the state and the valve opening as the action, to obtain the optimized sequence of control parameters. The control parameter optimization instruction is generated and executed according to the control parameter optimization sequence. Based on the flow rate, the heat supply area is divided using the METIS algorithm to obtain a heat balance adjustment scheme. Based on the control parameter optimization instructions in the thermal balance adjustment scheme, a transient simulation of the hydraulic and thermal coupling of the pipeline network is performed using the finite element method to obtain a simulated pipeline network response sequence that includes changes in node pressure and flow rate. The total power consumption of the system is obtained, and the simulated flow rate and simulated pressure are extracted by combining the simulated pipeline response sequence. The energy waste index is quantified by calculating the area of ​​the difference between the total power consumption of the system and the preset design power consumption using the trapezoidal numerical integration method, and the energy waste index value is obtained. When the energy waste index value is lower than the preset optimization target threshold, the K-means clustering algorithm is used to select performance improvement nodes from the simulated pipeline response sequence to obtain an optimized node set.

[0048] Specifically, firstly, based on path priority ranking and standardized flow data, the absolute difference between the preset target flow and the standardized flow is calculated to obtain flow deviation data. The target flow is the ideal flow value of each branch determined through heat balance calculation based on user load demand data and heating system design parameters.

[0049] Next, based on the flow deviation data, the PISO algorithm is used to solve the coupled non-isothermal flow and heat transfer. The specific process includes: establishing a set of governing equations containing mass conservation, momentum conservation, and energy conservation equations, where the energy conservation equations describe the coupling relationship between fluid temperature distribution and flow parameters; discretizing the governing equations and employing a separate solution strategy; solving the momentum equation in the prediction step to obtain the velocity field; and sequentially solving the pressure correction equation and energy equation in the correction step, where the energy equation considers convective heat transfer between the fluid and the pipe wall, as well as the fluid's own heat conduction; and iterating multiple times until both the flow field and temperature field reach convergence, resulting in a set of heat distribution deviations composed of temperature deviation values ​​from each region.

[0050] Then, based on the set of heating distribution deviations, the strategy is iteratively optimized using a deep Q-network algorithm. This deep Q-network requires offline training before deployment, and the training process is as follows: First, a large-scale historical dataset containing historical heating distribution deviations, corresponding valve actions, and subsequent system responses and energy consumption data is prepared. Then, a Q-network structure with fully connected layers is constructed, whose reward function is defined as the weighted sum of the decrease in total system energy consumption and the improvement in heating uniformity. During training, an empirical replay mechanism is used to store state transition samples, and a fixed target network is set to stabilize the learning process. The network weights are updated by minimizing the temporal difference error between the current Q-value and the target Q-value until the network converges. The algorithm uses the heating distribution deviation value as the state input and the valve opening adjustment amount as the action output, constructing a value function to evaluate the long-term cumulative reward of different state-action pairs. An ε-greedy strategy is used to balance exploration and exploitation. The Q-network parameters are trained using the empirical replay mechanism, and the Q-value estimate is iteratively updated using the Bellman equation, ultimately obtaining the optimized sequence of control parameters that maximizes the long-term cumulative reward.

[0051] The system generates and executes control parameter optimization instructions based on the control parameter optimization sequence. Simultaneously, it uses the METIS algorithm to divide the heating balance zone based on standardized flow data. The execution process of the METIS algorithm includes: constructing a graph model based on the pipeline topology, using pipeline flow differences as edge weights; employing a multi-level partitioning strategy, through three stages—graph coarsening, initial partitioning, and refinement—the pipeline network is divided into several flow-balanced zones, resulting in a heat balance adjustment scheme.

[0052] Based on the optimization instructions for control parameters in the thermal balance adjustment scheme, transient simulation of the hydraulic and thermal coupling of the pipeline network is performed using the finite element method. The specific process includes: establishing a set of transient control equations considering fluid compressibility and temperature changes; spatially discretizing the computational domain and dividing the pipeline into a finite number of elements; using the Galerkin weighted residual method to establish the element equations and assembling them into the overall system equations; and using the implicit Euler method in the time domain for progressive solving to obtain the simulated pipeline network response sequence including nodal pressure and flow rate changes.

[0053] The system's total power consumption data is obtained, and combined with the simulated flow rate and simulated pressure extracted from the simulated pipeline response sequence, the area of ​​the difference between the system's total power consumption and the preset design power consumption is calculated using the trapezoidal numerical integration method. Specifically, the system's total power consumption curve and the design power consumption curve are discretized into several equally spaced data points within the same time interval. The area of ​​the trapezoid formed by the difference in the ordinates of the two curves at each adjacent time point is calculated, and these trapezoidal areas are accumulated to obtain the energy waste index value.

[0054] The system sets an optimization target threshold. Once sufficient historical operating data is available, the threshold is set through the following steps: First, all operating periods marked as 'optimal operation' are selected from the system's historical database. Then, the corresponding energy waste index values ​​for these periods are calculated, forming a data sample set. Next, the values ​​in this data set are arranged in ascending order. Finally, the value at the 20th percentile of this sorted sequence is taken as the official optimization target threshold. During system initialization or in the absence of such historical data, this threshold is temporarily set as a conservative initial value determined based on the system's rated heating load, the design circulating water pump power consumption, and reference to typical energy efficiency benchmarks of similar heating systems, ensuring safe and stable operation in the initial stage. As system operating data accumulates, when the collected optimal operating data reaches the preset minimum statistical sample requirement, the system automatically initiates a threshold update procedure, switching to the threshold setting method based on historical data statistical analysis. When the calculated energy waste index value is lower than the optimization target threshold, performance improvement nodes are selected from the simulated pipeline response sequence using a K-means clustering algorithm. The execution process of the K-means clustering algorithm includes: randomly selecting initial cluster centers from the feature space composed of the pressure improvement rate and flow improvement rate of all nodes; calculating the Euclidean distance of each node to each cluster center and assigning it to the nearest cluster; recalculating the mean of each cluster as the new cluster center; repeating the above process until the cluster centers no longer change; and finally selecting the nodes belonging to the cluster with the most significant performance improvement to form the optimized node set.

[0055] By comprehensively utilizing various optimization algorithms and simulation technologies, the system achieves systematic optimization of control parameters and accurate prediction of pipeline network response, thereby reducing system energy consumption, improving heating uniformity, and providing technical support for building an efficient and stable heating system.

[0056] In step S17, the energy transfer path model is calculated based on the optimized node set to obtain the energy transfer path model.

[0057] In one specific implementation, the step of calculating the energy transfer path model based on the optimized node set to obtain the energy transfer path model includes: Based on the optimized node set, the corresponding flow rate and pipeline topology data are extracted, and the node connection importance weight is calculated using the PageRank algorithm to obtain the node connection weight matrix. By acquiring time-series temperature data and combining it with the node connection weight matrix, a path optimization parameter set including path transmission efficiency is obtained through multiple rounds of message passing and state iteration updates via a gated graph neural network. Energy consumption data is acquired, and combined with the path optimization parameter set, the energy allocation coefficient of the path is calculated through linear weighted fusion to obtain the energy allocation coefficient; When the energy allocation coefficient is outside the preset equilibrium threshold range, the energy allocation coefficient is fed back to the gated graph neural network for a new round of iterative optimization. When the energy distribution coefficient is within a preset equilibrium threshold range, the node connection weights are solidified through a graph attention network to obtain an energy transfer path model.

[0058] Specifically, firstly, standardized flow data and pipeline topology data are extracted from the optimized node set, and the importance weights of node connections are calculated using the PageRank algorithm. The specific execution process is as follows: The pipeline topology is abstracted into a directed graph, where nodes represent key locations in the pipeline, edges represent pipeline connections, and the direction of the edges represents the main flow direction of standardized flow; the standardized flow values ​​are normalized and used as the edge weights; each node is assigned the same initial PageRank value; the PageRank value of each node is updated through iterative calculation. In each iteration, the PageRank value of a node is distributed and summed from the PageRank values ​​of all its incoming neighbor nodes according to the weight ratio of its outgoing edges, while a damping factor is introduced to handle dangling nodes; the iteration process continues until the maximum change in the PageRank value of all nodes in two adjacent rounds is less than a preset convergence threshold, which is a value determined through multiple experimental tests; finally, a stable node connection weight matrix is ​​obtained.

[0059] Next, temperature time-series data is acquired and combined with the node connection weight matrix. A gated graph neural network is then used for multiple rounds of message passing and state iteration updates. The specific process includes: constructing a graph structure using pipeline nodes as graph nodes and node connection weights as edge weights; using the temperature time-series data after moving average processing as the initial node features; in each layer of the network, for each target node, aggregating the features of its neighboring nodes. The aggregation weight is calculated by multiplying the node connection weights by the attention coefficient, which is obtained by multiplying the target node features and neighboring node features by the trainable weight matrix and then calculating the dot product of their feature vectors; the resulting product is normalized using the Softmax function to obtain the final aggregation weight; the aggregated neighboring features and the target node's own features are input into the gated recurrent unit, where updating and resetting the gate controls the retention and forgetting of information, generating a new feature representation for the node; after multiple layers of such message passing and state updates, a path optimization parameter set containing path transmission efficiency is finally obtained. The gated graph neural network is also an offline trained model. Its training data consists of historical temperature time series, node connection weights, and corresponding true values ​​of path transmission efficiency verified by hydrothermal simulation. The training objective is to minimize the error between the network's output path transmission efficiency prediction and the simulation true value. Gradient descent algorithm is used for parameter optimization during training.

[0060] Then, energy consumption data is acquired and, combined with the path optimization parameter set, the energy allocation coefficient of the path is calculated through linear weighted fusion. Specifically, the transmission efficiency parameters in the path optimization parameter set are normalized along with their corresponding energy consumption data; the weight ratio of each parameter is determined by principal component analysis based on the significance of each parameter's impact on system energy efficiency in historical operating data; for each path, its normalized transmission efficiency parameters and energy consumption parameters are linearly combined according to this weight vector, and the summation yields the energy allocation coefficient for that path.

[0061] The system sets an equilibrium threshold range, which is determined by statistically analyzing the energy allocation coefficient distribution of all paths under historical optimal operating conditions and taking the 40th to 60th percentile range. When the calculated energy allocation coefficient is outside this equilibrium threshold range, the energy allocation coefficient is fed back to the gated graph neural network to adjust the weight parameters in the network and perform a new round of iterative optimization.

[0062] When the energy distribution coefficient is within a preset equilibrium threshold range, the node connection weights are solidified through a graph attention network. The specific process includes: calculating the attention coefficients between nodes based on the current graph structure and node features; normalizing the attention coefficients using the Softmax function; weighting and aggregating the features of neighboring nodes using the normalized attention coefficients; generating the final node representations through a nonlinear transformation of the aggregated features; and integrating these node representations with the corresponding connection weights to form a stable energy transfer path model.

[0063] By comprehensively considering the characteristics of the pipeline network topology, environmental factors, and energy consumption, a computational model that can accurately describe the energy transfer law is established, providing a reliable theoretical basis for subsequent system stability improvement configuration, and realizing accurate modeling and optimization of energy flow in the heating system.

[0064] In step S18, a control command sequence is extracted from the energy transfer path model, the degree of coverage of the control command sequence to the dynamic operating conditions is determined, and the system stability improvement configuration is determined.

[0065] In one specific implementation, the step of extracting a control command sequence from the energy transfer path model, determining the coverage of the control command sequence to dynamic operating conditions, and determining the system stability improvement configuration includes: The optimized sequence of control parameters is extracted based on the energy transfer path model, and the command pattern is matched with the pre-stored operating condition database using a dynamic time warping algorithm to obtain the operating condition matching command. Based on the operating condition matching command, the stability index is quantified by calculating the maximum eigenvalue of the system state matrix using the Lyapunov stability analysis method, and the system stability evaluation value is obtained. When the system stability assessment value is lower than the preset stability threshold, the Kalman filter algorithm is used to fuse the operating condition database and the node connection weight matrix to obtain an enhanced system state estimate. Based on the enhanced system state estimation, a stable set of control parameters is obtained by iterating the node states and optimizing the parameters through a graph attention network. Based on the set of stable control parameters, the optimal matching between energy allocation tasks and node carrying capacity is performed using the Hungarian algorithm to determine the system stability enhancement configuration.

[0066] Specifically, the data sources for this step include: the energy transfer path model, the optimized sequence of control parameters obtained from step S16, a pre-stored operating condition database, and a node connection weight matrix. The operating condition database is constructed by collecting and organizing historical operating data over a long period of time, and contains system state parameters and corresponding control command modes under various typical operating conditions.

[0067] First, an optimized sequence of control parameters is extracted based on the energy transfer path model. Then, a dynamic time warping algorithm is used to match the sequence with a pre-stored operating condition database. The specific execution process is as follows: The optimized sequence of control parameters to be matched is compared with the standard control command sequences stored in the operating condition database. The dynamic time warping algorithm constructs a cumulative distance matrix to find the optimal curved path between the two sequences, minimizing the overall distance between them. By comparing the matching distances of all standard control command sequences, the standard command with the smallest distance is selected as the matching result, thus obtaining the operating condition matching command.

[0068] Next, based on the operating condition matching command, the stability index is quantified by calculating the maximum eigenvalue of the system state matrix using Lyapunov stability analysis. The method for constructing the system state-space model is as follows: the pressure values ​​of key nodes in the pipeline network, the flow values ​​of the main pipeline, and the opening values ​​of the control valves are collectively used as the system's state variables; the model input is the valve opening adjustment amount included in the operating condition matching command; based on the linearized dynamic relationship established by mass conservation, momentum conservation, and the pipeline topology, the law of evolution of the state variables over time is determined, thereby constructing the system's state-space model. The specific process includes: establishing the system's state-space model based on the operating condition matching command, which describes the law of evolution of the system's state variables over time; constructing a Lyapunov function, which is a positive definite scalar function; obtaining the system state matrix by solving the derivative of this function along the system trajectory; calculating all eigenvalues ​​of the state matrix, and taking the eigenvalue with the largest real part as the stability criterion; and converting the largest eigenvalue into a system stability evaluation value after normalization.

[0069] The system sets a stability threshold, which is determined by statistically analyzing the distribution of system stability assessment values ​​under historical stable operating conditions and taking the 95th percentile value. When the system stability assessment value is lower than this stability threshold, the Kalman filter algorithm is used to fuse the operating condition database and the node connection weight matrix. The execution process of the Kalman filter algorithm includes: establishing the system state equation and observation equation, where the system state is composed of operating condition parameters and node connection weights; in the prediction step, predicting the system state and error covariance at the next moment based on the state equation; in the update step, correcting the predicted state using actual observation data and balancing the weights of the predicted and observed values ​​by calculating the Kalman gain; after multiple iterations, an enhanced system state estimate containing more accurate state information is finally obtained.

[0070] Based on the enhanced system state estimation, node state iteration and parameter optimization are performed using a graph attention network. The specific process includes: using network nodes as graph nodes and the enhanced system state estimation as the initial node features; in each layer of the network, for each target node, calculating its attention coefficients with all neighboring nodes, which are obtained by multiplying the target node features and neighboring node features by a trainable weight matrix, followed by processing with a nonlinear activation function; normalizing the attention coefficients using the Softmax function; weighting and aggregating the neighboring node features using the normalized attention coefficients; fusing the aggregated features with the target node's own features and generating a new feature representation for the node through a nonlinear transformation; and finally, after multiple iterations, obtaining a stable set of control parameters.

[0071] Based on the set of stable control parameters, the Hungarian algorithm is used to achieve optimal matching between energy allocation tasks and node carrying capacity. The specific execution process includes: constructing a bipartite graph, where one set of vertices represents energy allocation tasks and the other set represents network nodes; calculating the matching cost between each task and each node based on the set of stable control parameters, forming a cost matrix; finding the optimal match using the Hungarian algorithm, which employs row reduction, column reduction, and covering zero elements to ultimately find the task-node allocation scheme that minimizes the total matching cost; and determining the system stability enhancement configuration based on this optimal allocation scheme. Before executing this matching and configuration process, the system first performs data quality and system status verification: if key sensor data loss or anomalies, network communication interruption, or incomplete state variables required for calculation are detected, the application of the optimal matching scheme is suspended, and a preset degradation control strategy is executed instead. The specific implementation steps of this degradation strategy are as follows: First, the system ignores all abnormal or unreachable nodes; then, it obtains the rated heating capacity of all nodes still in normal operation; next, it directly distributes the current total heat load to be allocated according to the ratio of the rated heating capacity of these normal nodes; finally, it issues fixed valve opening commands to these normal nodes corresponding to the allocated load. These fixed opening commands are conservative safety values ​​pre-mapped between the node's rated capacity and the allocated load. This process aims to maintain the most basic hydraulic balance and heat supply in abnormal situations, preventing system instability.

[0072] By comprehensively utilizing a variety of advanced algorithms, intelligent matching of control commands and accurate assessment of system stability are achieved. Furthermore, through data fusion and optimized matching technologies, the operational stability and energy efficiency of the heating system under dynamic operating conditions are improved.

[0073] Reference Figure 2 The second embodiment of the present invention provides a centralized heating secondary network control system based on reinforcement learning, comprising: The data acquisition module is used to collect flow, temperature and pressure data in the pipeline network to obtain the raw dataset; The data processing module is used to perform denoising and standardization processing on the original dataset to obtain the pipeline network data matrix; The bottleneck node screening module is used to extract a set of key nodes containing node pressure fluctuations based on the pipeline network data matrix, and to screen out dynamic bottleneck nodes. The energy flow analysis module is used to calculate path resistance and flow rate changes based on the dynamic bottleneck nodes and the pipeline data matrix, and generate energy flow characteristics. The path ranking module is used to extract branch influence factors based on the energy flow characteristics, perform user demand correlation analysis, and rank the branch paths to obtain the path priority ranking. The optimization node analysis module is used to optimize control parameters and simulate the overall response of the pipeline network according to the path priority sorting, and to calculate energy waste indicators to verify the optimization nodes and obtain the optimization node set. The energy transfer path construction module is used to calculate the energy transfer path model based on the optimized node set, and obtain the energy transfer path model. The output module is used to extract the control command sequence from the energy transfer path model, determine the coverage of the control command sequence to the dynamic operating conditions, and determine the system stability improvement configuration.

[0074] It should be noted that the reinforcement learning-based centralized heating secondary network control system provided in this embodiment of the invention is used to execute all the process steps of the reinforcement learning-based centralized heating secondary network control method in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.

[0075] This invention also provides an electronic device. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a reinforcement learning-based centralized heating secondary network control program. When the processor executes the computer program, it implements the steps described in the various reinforcement learning-based centralized heating secondary network control method embodiments above, for example... Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above system embodiments, such as a centralized heating secondary network control module based on reinforcement learning.

[0076] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.

[0077] The electronic device may be a desktop computer, laptop, handheld computer, or smart tablet, etc. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0078] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting all parts of the electronic device via various interfaces and lines.

[0079] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0080] Wherein, if the modules / units integrated in the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0081] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0082] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A centralized heating secondary network control method based on reinforcement learning, characterized in that, include: Collect flow, temperature, and pressure data from the pipeline network to obtain the raw dataset; The original dataset is denoised and standardized to obtain the pipeline data matrix; Based on the pipeline network data matrix, extract the set of key nodes containing node pressure fluctuations, and filter out dynamic bottleneck nodes; Based on the dynamic bottleneck nodes and the pipeline data matrix, the path resistance and flow rate change are calculated to generate energy flow characteristics. Based on the characteristics of energy flow, branch influencing factors are extracted, user demand correlation analysis is performed, and branch path priority is ranked to obtain path priority ranking. Based on the path priority, the control parameters are optimized and the overall pipeline network response is simulated. Energy waste indicators are calculated to verify the optimized nodes, resulting in an optimized node set. Based on the optimized node set, the energy transfer path model is calculated to obtain the energy transfer path model; Extract the control command sequence from the energy transfer path model, determine the coverage of the control command sequence to dynamic operating conditions, and determine the system stability improvement configuration.

2. The centralized heating secondary network control method based on reinforcement learning according to claim 1, characterized in that, Based on the pipeline network data matrix, a set of key nodes containing node pressure fluctuations is extracted, and dynamic bottleneck nodes are selected, including: The pressure data matrix is ​​extracted from the pipeline data matrix, and combined with the pre-stored pipeline topology data, the importance of nodes is scored through a graph attention network to obtain the set of key nodes. Extract the node pressure time series based on the key node set, and quantify the pressure fluctuation amplitude by calculating the standard deviation within a preset sliding time window to obtain the key node set containing node pressure fluctuations. Based on the set of key nodes, key nodes whose pressure fluctuations exceed a preset pressure fluctuation threshold are extracted to obtain dynamic bottleneck nodes.

3. The centralized heating secondary network control method based on reinforcement learning according to claim 2, characterized in that, Based on the dynamic bottleneck nodes and pipeline network data matrix, path resistance and flow rate changes are calculated to generate energy flow characteristics, including: Extract the pipeline data matrix corresponding to the dynamic bottleneck node, combine it with the pipeline topology data, and perform steady-state flow field simulation using the SIMPLE algorithm to obtain a set of path resistance distributions including pipeline segment pressure drop, pipeline flow rate and resistance value. Extract abnormal pipeline paths from the path resistance distribution set whose resistance values ​​exceed a preset resistance threshold, and use the finite volume method to calculate the unsteady flow field to obtain the time series data of flow variation of the abnormal pipeline paths. Based on the path resistance distribution set and flow change time series data, the energy flow characteristics are obtained by extracting features of channel connectivity and energy transfer efficiency through Laplace eigenmaps.

4. The centralized heating secondary network control method based on reinforcement learning according to claim 3, characterized in that, Based on the energy flow characteristics, branch influence factors are extracted, user demand correlation analysis is performed, and branch path priority is ranked to obtain the path priority ranking, including: Based on energy flow characteristics and pipeline topology data, the dynamic bottleneck nodes are divided into regions using the Kernighan-Lin graph segmentation algorithm to obtain bottleneck node partitions. Based on the bottleneck node partition and path resistance distribution set, the resistance and flow of the branch path are simulated and calculated using the PISO algorithm to obtain the branch influence factor composed of resistance coefficient and average flow. By obtaining user load demand and combining it with branch influence factors, nonlinear correlation analysis is performed by calculating the Spearman correlation coefficient to obtain nonlinear correlation results. Based on the nonlinear correlation results, the priority of branch paths is ranked by combining the entropy weight method and the TOPSIS method, resulting in a path priority ranking.

5. The centralized heating secondary network control method based on reinforcement learning according to claim 4, characterized in that, Based on path priority, control parameters are optimized and the overall pipeline network response is simulated. Energy waste indicators are calculated to verify the optimized nodes, resulting in a set of optimized nodes, including: Based on the path priority sorting and the flow rate in the pipeline network, the absolute difference between the preset target flow rate and the flow rate is calculated to obtain the flow deviation data; Based on the flow deviation data, the PISO algorithm is used to solve the coupled non-isothermal flow and heat transfer, resulting in a set of heating distribution deviations composed of temperature deviation values ​​in each region. Based on the set of heating distribution deviations, a deep Q-network algorithm is used to iteratively optimize the strategy with deviation values ​​as states and valve opening as actions, resulting in an optimized sequence of control parameters. The control parameter optimization sequence is used to generate and execute control parameter optimization instructions. Based on the flow rate, the METIS algorithm is used to divide the heating area into balanced areas to obtain a thermal balance adjustment scheme. Based on the optimization instructions of the control parameters in the thermal balance adjustment scheme, the transient simulation of the hydraulic and thermal coupling of the pipeline network is carried out by the finite element method to obtain the simulated pipeline network response sequence including the changes in node pressure and flow rate; The total power consumption of the system is obtained, and the simulated flow and simulated pressure are extracted by combining the simulated pipeline response sequence. The energy waste index is quantified by calculating the area of ​​the difference between the total power consumption of the system and the preset design power consumption using the trapezoidal numerical integration method, and the energy waste index value is obtained. When the energy waste index value is lower than the preset optimization target threshold, the K-means clustering algorithm is used to select performance improvement nodes from the simulated pipeline response sequence to obtain the optimized node set.

6. The centralized heating secondary network control method based on reinforcement learning according to claim 5, characterized in that, Based on the optimized node set, an energy transfer path model is calculated to obtain the energy transfer path model, including: Based on the optimized node set, the corresponding flow rate and the pipeline topology data are extracted, and the node connection importance weight is calculated using the PageRank algorithm to obtain the node connection weight matrix. By acquiring time-series temperature data, combining it with node connection weight matrices, and using a gated graph neural network to perform multiple rounds of message passing and state iteration updates, a path optimization parameter set including path transmission efficiency is obtained. By acquiring energy consumption data and combining it with a path optimization parameter set, the energy allocation coefficient of the path is calculated through linear weighted fusion. When the energy allocation coefficient is outside the preset equilibrium threshold range, the energy allocation coefficient is fed back to the gated graph neural network for a new round of iterative optimization. When the energy distribution coefficient is within the preset equilibrium threshold range, the node connection weights are solidified through a graph attention network to obtain the energy transfer path model.

7. The centralized heating secondary network control method based on reinforcement learning according to claim 6, characterized in that, Extract the control command sequence from the energy transfer path model, determine the coverage of the control command sequence to dynamic operating conditions, and determine the system stability improvement configuration, including: The optimized sequence of control parameters is extracted based on the energy transfer path model, and the command pattern is matched with the pre-stored operating condition database using a dynamic time warping algorithm to obtain the operating condition matching command. Based on the operating condition matching command, the stability index is quantified by calculating the maximum eigenvalue of the system state matrix using the Lyapunov stability analysis method, and the system stability evaluation value is obtained. When the system stability assessment value is lower than the preset stability threshold, the Kalman filter algorithm is used to fuse the operating condition database and the node connection weight matrix to obtain an enhanced system state estimate. Based on the enhanced system state estimation, a stable set of control parameters is obtained by iterating the node states and optimizing the parameters through a graph attention network. Based on the set of stable control parameters, the optimal matching between energy allocation tasks and node carrying capacity is determined using the Hungarian algorithm to determine the configuration for improving system stability.

8. A centralized heating secondary network control system based on reinforcement learning, characterized in that, include: The data acquisition module is used to collect flow, temperature and pressure data in the pipeline network to obtain the raw dataset; The data processing module is used to perform denoising and standardization on the obtained raw dataset to obtain the pipeline network data matrix; The bottleneck node screening module is used to extract a set of key nodes containing node pressure fluctuations based on the pipeline network data matrix, and to screen out dynamic bottleneck nodes. The energy flow analysis module is used to calculate path resistance and flow rate changes based on the dynamic bottleneck nodes and the pipeline data matrix, and generate energy flow characteristics. The path ranking module is used to extract branch influence factors based on the energy flow characteristics, perform user demand correlation analysis, and rank the branch paths to obtain the path priority ranking. The optimization node analysis module is used to optimize control parameters and simulate the overall response of the pipeline network according to the path priority sorting, and to calculate energy waste indicators to verify the optimization nodes and obtain the optimization node set. The energy transfer path construction module is used to calculate the energy transfer path model based on the optimized node set, and obtain the energy transfer path model. The output module is used to extract the control command sequence from the energy transfer path model, determine the coverage of the control command sequence to the dynamic operating conditions, and determine the system stability improvement configuration.