Operation control methods and devices for courtyard heating networks
By adopting a strategy parameter fusion training mechanism based on valley value evaluation update in the courtyard heating network, accurate perception of multi-dimensional real-time status is achieved, solving the problems of increased energy consumption and regulation lag caused by static or weak dynamic control, and improving the operating efficiency and energy consumption management of the courtyard heating network.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TIANJIN UNIV
- Filing Date
- 2026-03-20
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, static or weakly dynamic control is difficult to adapt to the rapid changes in building needs, resulting in a small temperature difference between secondary supply and return water and excessive flow of circulating pumps, which increases energy consumption; traditional fusion prediction models rely on accurate prediction, which is complex and inflexible in engineering implementation.
A training mechanism based on valley value evaluation value to update strategy parameters and parameter fusion is adopted. By fusing multi-dimensional real-time state information such as the opening degree of the first valve, heat medium flow rate and temperature, dynamic operation control of the courtyard heating network is realized, and intelligent agents are used to adaptively adjust the valve opening degree.
It improves the policy robustness and training stability of the control agent in complex environments, reduces system energy consumption, enhances the dynamic response capability of heating regulation, and avoids the regulation lag problem of traditional control methods.
Smart Images

Figure CN121897961B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart heating technology, specifically to a method and device for controlling the operation of a courtyard heating network. Background Technology
[0002] District heating systems mainly consist of heat sources, heating stations, courtyard pipe networks, and building terminals. They exhibit strong hydraulic and thermal coupling, with high correlation between different operating parameters. Joint control of valves in different zones plays a crucial role in improving the overall efficiency of the heating system, enhancing indoor comfort, and reducing energy consumption of heat sources and equipment. Related technologies mainly include: controlling equipment valves within a preset time period by predicting building load information and combining it with the water and heat status of the pipe network and target information (e.g., using cost and equipment fluctuations as penalties); or, implementing static or weakly dynamic control of valves in different zones, such as linear feedback control for some zone valves and fixed opening degrees for valves in other zones.
[0003] However, static or weak dynamic control is difficult to adapt to the rapid changes in building needs, which can easily lead to a small temperature difference between secondary supply and return water and an excessive flow of circulating pumps, resulting in an unnecessary increase in energy consumption. Traditional fusion prediction of load information, hydrothermal state and target information relies on accurate prediction models and hydrothermal models, which are complex to implement and maintain in engineering, and the multi-dimensional target information has complex constraints, resulting in poor flexibility in actual engineering. Summary of the Invention
[0004] In view of the above problems, the present invention provides a method and apparatus for controlling the operation of a courtyard heating network.
[0005] According to a first aspect of the present invention, a method for operating and controlling a courtyard heating network is provided, comprising: using sample state information and sample action information, sequentially updating the initial evaluation parameters of each of a plurality of parallel initial evaluation networks, the initial strategy parameters of an initial strategy network, and the initial temperature parameters for controlling the randomness of action information in an initial intelligent agent, to obtain an intermediate intelligent agent, wherein the initial strategy parameters are updated with the valley value among the multiple evaluation values of the plurality of initial evaluation networks; updating the intermediate parameters based on an update factor for controlling the update speed of the intermediate parameters of the intermediate intelligent agent and the parameter fusion result; and determining the intermediate intelligent agent as a control intelligent agent when the performance information of the intermediate intelligent agent meets preset conditions, wherein the parameter fusion result is obtained by fusing the initial evaluation parameters and the intermediate evaluation parameters; using the control intelligent agent to fuse the opening degree of the first valve between the heating station and the heat source, the heat medium flow rate and temperature of the courtyard heating network between the heating station and the building, the heat medium flow rate and temperature at the building entrance, the temperature of the target area in the building, and equipment operation information to obtain the current operating state information of the courtyard heating network, and using the current operating state information to control the opening degree of the first valve and the valve opening at the building entrance to obtain the subsequent operating state of the courtyard heating network.
[0006] A second aspect of the present invention provides an operation control device for a courtyard heating network, comprising: an update module, configured to sequentially update the initial evaluation parameters of each of multiple parallelly configured initial evaluation networks, the initial strategy parameters of an initial strategy network, and the initial temperature parameters for controlling the randomness of action information in an initial agent using sample state information and sample action information, to obtain an intermediate agent, wherein the initial strategy parameters are updated using the valley value evaluation value among the multiple output values of the multiple initial evaluation networks; and a determination module, configured to update the intermediate parameters based on an update factor for controlling the update speed of the intermediate parameters of the intermediate agent and a parameter fusion result, wherein the intermediate agent... If the performance information of the agent meets the preset conditions, the intermediate agent is determined as the control agent. The parameter fusion result is obtained by fusing the initial evaluation parameters and the intermediate evaluation parameters. The control module is used to use the control agent to fuse the opening of the first valve between the heating station and the heat source, the flow rate and temperature of the heat medium in the courtyard heating network between the heating station and the building, the flow rate and temperature of the heat medium at the building entrance, the temperature of the target area in the building, and the equipment operation information to obtain the current operating status information of the courtyard heating network. The current operating status information is then used to control the opening of the first valve and the valve opening at the building entrance to obtain the subsequent operating status of the courtyard heating network.
[0007] A third aspect of the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0008] A fourth aspect of the present invention also provides a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions, when executed by a processor, implement the steps of the above-described method.
[0009] A fifth aspect of the present invention also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0010] According to embodiments of the present invention, by updating strategy parameters based on valley evaluation values and a training mechanism that integrates parameters, the robustness of the control agent in complex heating network environments and the stability of training are effectively improved. This overcomes the problems of complex modeling, difficult maintenance, and model mismatch caused by traditional reliance on prediction models and hydro-thermal models. By integrating multi-dimensional real-time states such as the opening degree of the first valve, the flow and temperature of the heat medium at multiple nodes, the temperature of the building area, and equipment operation information, accurate perception of the dynamic operation of the courtyard heating network is achieved. This avoids the problems of small secondary supply and return water temperature difference and excessive circulation pump flow caused by adjustment lag in traditional static or weak dynamic control. This enables the control agent to output valve opening control commands online in real time, directly optimizing the operating state of the courtyard heating network. Without the need for complex prediction models, it adaptively balances the constraints between energy consumption and the opening degree of multiple valves, reducing system operating energy consumption and improving the dynamic response capability of heating regulation. Attached Figure Description
[0011] The above-described features, other objects, and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:
[0012] Figure 1 An application scenario diagram of the operation control method and apparatus for a courtyard heating network according to an embodiment of the present invention is shown;
[0013] Figure 2 A flowchart of a method for controlling the operation of a courtyard heating network according to an embodiment of the present invention is shown;
[0014] Figure 3 A structural block diagram of an operation control device for a courtyard heating network according to an embodiment of the present invention is shown. Detailed Implementation
[0015] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0016] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0017] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0018] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0019] In related technologies, static or weak dynamic control is difficult to adapt to the rapid changes in building needs, which can easily lead to a small temperature difference between secondary supply and return water and an excessive flow rate of circulating pumps, resulting in an unnecessary increase in energy consumption. Traditional fusion prediction of load information, hydrothermal state and target information relies on accurate prediction models and hydrothermal models, which are complex to implement and maintain. Moreover, the multi-dimensional target information has complex constraints and is not very flexible in actual engineering.
[0020] In some examples, load forecasting information is obtained by nonlinearly transforming the acquired monitoring data using gating algorithms, and then combined with the water and heat information of the pipeline network and target information to determine the start-up and shutdown combinations of heating equipment to control the heating system. However, load forecasting relies on meteorological and historical load data, and pipeline water and heat data both depend on prediction models. Actual heating systems have complex networks, and parameters may change with time and operating conditions. Model mismatch can lead to optimization results deviating from the actual optimum, causing system instability. Solving constrained optimization problems (such as nonlinear programming or similar) involves optimization variables including the start-up and shutdown of multiple devices (discrete) and load distribution (continuous), resulting in high computational complexity. When the system is large, the time window is long, and there are many devices, the solution is time-consuming, making it difficult to achieve high-frequency real-time control.
[0021] In view of this, embodiments of the present invention provide a method and apparatus for operating control of a courtyard heating network, comprising: using sample state information and sample action information, sequentially updating the initial evaluation parameters of each of the multiple parallel initial evaluation networks, the initial strategy parameters of the initial strategy network, and the initial temperature parameters for controlling the randomness of action information in an initial intelligent agent to obtain an intermediate intelligent agent, wherein the initial strategy parameters are updated with the valley value among the multiple evaluation values of the multiple initial evaluation networks; updating the intermediate parameters based on the update factor for controlling the update speed of the intermediate parameters of the intermediate intelligent agent and the parameter fusion result; and determining the intermediate intelligent agent as a control intelligent agent when the performance information of the intermediate intelligent agent meets preset conditions, wherein the parameter fusion result is obtained by fusing the initial evaluation parameters and the intermediate evaluation parameters; using the control intelligent agent to fuse the opening degree of the first valve between the heating station and the heat source, the heat medium flow rate and temperature of the courtyard heating network between the heating station and the building, the heat medium flow rate and temperature at the building entrance, the temperature of the target area in the building, and equipment operation information to obtain the current operating state information of the courtyard heating network, and using the current operating state information to control the opening degree of the first valve and the valve opening at the building entrance to obtain the subsequent operating state of the courtyard heating network.
[0022] According to embodiments of the present invention, by updating strategy parameters based on valley evaluation values and a training mechanism involving parameter fusion, the policy robustness and training stability of the control agent in complex heating network environments are effectively improved. This overcomes the problems of complex modeling, difficult maintenance, and model mismatch caused by traditional reliance on prediction models and hydrothermal models. By integrating multi-dimensional real-time states such as the opening degree of the first valve, the flow and temperature of the heat medium at multiple nodes, the temperature of the building area, and equipment operation information, accurate perception of the dynamic operation of the courtyard heating network is achieved. This avoids the problems of small secondary supply and return water temperature difference and excessive circulation pump flow caused by adjustment lag in traditional static or weak dynamic control. This enables the control agent to output valve opening control commands online in real time, directly optimizing the operating state of the courtyard heating network. Without the need for complex prediction models, it adaptively balances the constraints between energy consumption and the opening degrees of multiple valves, reducing system operating energy consumption and improving the dynamic response capability of heating regulation.
[0023] Figure 1 The diagram illustrates an application scenario of the operation control method and apparatus for a courtyard heating network according to an embodiment of the present invention.
[0024] like Figure 1 As shown, application scenario 100 according to this embodiment may include an information acquisition device 110, a network 120, and a server 130. The network 120 is used as a medium to provide a communication link between the information acquisition device 110 and the server 130. The network 120 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0025] The information acquisition device 110 can be a variety of sensors used to monitor the heating system, including but not limited to pressure sensors, temperature sensors, and flow sensors. For example, a pressure sensor can be used to monitor the pressure difference across valves in different areas. A flow sensor can be used to monitor the flow rate of the heating medium in the pipeline. A temperature sensor can be used to collect the temperature of the heating medium in different areas. It is understood that the locations of different information acquisition devices can be set according to actual needs and are not limited here.
[0026] Server 130 can be a server or intelligent control system that provides operation control for the courtyard heating network. For example, it can train the mechanical energy of the initial intelligent agent based on the collected historical data related to the courtyard heating network, and after obtaining a control intelligent agent that meets the control accuracy, it can use the control intelligent agent to control the operation status of the courtyard heating network in real time.
[0027] It should be noted that the operation control method for the courtyard heating network provided in this embodiment of the invention can generally be executed by server 130. Correspondingly, the operation control device for the courtyard heating network provided in this embodiment of the invention can generally be located in server 130. The operation control method for the courtyard heating network provided in this embodiment of the invention can also be executed by a server or server cluster that is different from server 130 and capable of communicating with information acquisition device 110 and / or server 130. Correspondingly, the operation control device for the courtyard heating network provided in this embodiment of the invention can also be located in a server or server cluster that is different from server 130 and capable of communicating with information acquisition device 110 and / or server 130.
[0028] It should be understood that Figure 1 The number of information collection devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of information collection devices, networks, and servers can be included.
[0029] Figure 2 A flowchart of a method for controlling the operation of a courtyard heating network according to an embodiment of the present invention is shown.
[0030] like Figure 2 As shown, the operation control method of the courtyard heating network in this embodiment includes operations S210 to S230.
[0031] In operation S210, using sample state information and sample action information, the initial evaluation parameters of each of the multiple parallel initial evaluation networks, the initial policy parameters of the initial policy network, and the initial temperature parameter used to control the randomness of action information in the initial agent are updated sequentially to obtain the intermediate agent. Specifically, the initial policy parameters are updated using the valley value among the multiple evaluation values from the multiple initial evaluation networks.
[0032] In embodiments of the present invention, the sample state information is historical sensing data acquired using multiple sensing devices, serving as observations (e.g., flow rate and temperature). The sample action information can be action information actually performed under the sample state information (e.g., valve opening adjustment). Multiple parallel initial evaluation networks can have the same network structure but different parameters, and each initial evaluation network can evaluate the expected long-term value information under a given "state-action pair" in parallel and independently.
[0033] The initial policy network can output suggested actions (e.g., valve opening adjustment) based on the input state information. The initial temperature parameter can be a learnable scalar parameter used to adjust the randomness (exploratory nature) of the initial policy network's output actions; a higher temperature parameter value results in more random actions, while a lower value results in more deterministic actions. The valley evaluation value can be the minimum value selected from multiple evaluation values calculated by all parallel evaluation networks for the same state-action pair at each step of the policy network update.
[0034] For example, during the training loop, sample state information and sample action information are sampled from the experience replay pool; the evaluation values of the two initial evaluation networks are calculated, the valley evaluation value is used to calculate the loss function of the policy network, and the initial policy parameters of the initial policy network are updated; at the same time, the initial evaluation parameters of the two initial evaluation networks can be updated using loss functions such as mean squared error, and the initial temperature parameters are also updated.
[0035] In operation S220, based on the update factor used to control the update rate of intermediate parameters of the intermediate agent and the parameter fusion result, the intermediate parameters are updated. If the performance information of the intermediate agent meets the preset conditions, the intermediate agent is determined as the control agent. The parameter fusion result is obtained by fusing the initial evaluation parameters and the intermediate evaluation parameters.
[0036] In embodiments of the present invention, the update factor can be used to control the speed of intermediate parameter updates, and can be a hyperparameter between 0 and 1. The parameter fusion result can be a set of updated parameters obtained by weighting the initial evaluation parameters and intermediate evaluation parameters according to the proportion specified by the update factor. The intermediate evaluation parameters can be the updated parameters obtained after multiple parallel initial evaluation networks (online evaluation networks) complete gradient updates in a single iteration.
[0037] Performance information can be used as a quantitative indicator to evaluate the merits of an intermediate agent's current strategy. For example, running the agent multiple times in a test environment and calculating its average cumulative reward, or evaluating its control effect (such as average temperature deviation or average energy consumption). Preset conditions can be convergence thresholds or conditions set for the performance information. The control agent can be an agent whose performance meets the preset conditions after training, and can be deployed to a real-world courtyard heating network system to achieve adaptive joint control of valves on different ends.
[0038] For example, in a simulation training platform, a digital twin system of a courtyard heating network can be built on a server. Soft actions and evaluation strategies can be run within this simulation environment, allowing the agent to interact with the environment, generating massive amounts of sample data, and completing the training. Performance information can be evaluated directly in the simulation environment by comparing energy consumption and comfort indicators under different control strategies to determine whether preset conditions are met.
[0039] In operation S230, the control agent integrates the opening degree of the first valve between the heating station and the heat source, the flow rate and temperature of the heat medium in the courtyard heating network between the heating station and the building, the flow rate and temperature of the heat medium at the building entrance, the temperature of the target area in the building, and equipment operation information to obtain the current operating status information of the courtyard heating network. The current operating status information is then used to control the opening degree of the first valve and the valve opening degree at the building entrance to obtain the subsequent operating status of the courtyard heating network.
[0040] In embodiments of the present invention, the fusion process can be a process of integrating and extracting features from multi-source heterogeneous data. The current operating status information can be a high-dimensional feature vector obtained through multi-source data fusion, capable of fully characterizing the real-time operating status of the courtyard heating network. Valve opening control can include the opening adjustment actions of the first valve (primary side) and the building entrance valve. The subsequent operating status can be the new operating state that the courtyard heating network enters after the control action is executed.
[0041] In one feasible embodiment, a joint control method for heating network valves based on the Soft Actor-Critic (SAC) strategy can be used to construct an integrated adaptive optimization operation system for courtyard heating systems, encompassing "heating station—courtyard network—building terminal—intelligent agent." First, an equal percentage regulating valve hydraulic model can be established at the heating station side. The primary side high-temperature water mass flow rate is calculated using the valve opening and pressure difference as inputs. A heat transfer model is then used to determine the secondary side supply water temperature based on the flow between the heating station and the heat source (primary side), the courtyard heating network flow rate (secondary side), and the building inlet temperature. Second, at the courtyard heating network level, the network can be treated as a quasi-steady-state system, constructing a single-pipe hydraulic model. This model combines the correlation matrix and the basic loop matrix to form nodal flow balance and loop pressure balance equations. Newton iteration is used to solve for the branch flow rate and pressure drop of the heating network. A first-order implicit upwind discretization method is used for the energy equation, calculating the temperature decay along the flow path. Thus, the flow rate and temperature at each building inlet are determined from the heating station outlet conditions. At the building and terminal sides, a building geometric model is established using a 3D model based on computer-aided design drawings. Building structure and system information can be defined using building energy modeling software and imported into building energy consumption simulation software. A real-time coupled simulation framework with the courtyard heating network model is built with the help of application programming interface, realizing the two-way linkage between building thermal response and pipeline hydraulic-thermal processes.
[0042] Given the strong nonlinear coupling between valve opening and hydraulic-thermal response in heating system control, and the fact that the control variables are continuous, the primary-side regulating valves and the valves at each building inlet can be modeled as a continuous control decision problem. A deep reinforcement learning algorithm based on policy gradients for continuous control can be chosen. For example, within a single-agent framework, a joint valve control policy can be learned using soft actions and evaluation policies to achieve distributed collaborative control of the opening of the first valve and the valves at the building inlets. By employing a maximum entropy deep reinforcement learning (SAC) policy oriented towards the continuous action space, and using a parallel approach with stochastic policies (Actors), dual evaluation networks (Critics), and an automatic entropy adjustment mechanism, high stability and high sample efficiency can be achieved. This maximizes not only long-term rewards but also policy entropy, improving exploration capabilities and preventing the policy from prematurely falling into local optima.
[0043] For example, in a deep reinforcement learning control framework, state information It can be all the information that the controlling agent can observe from the environment at each control step, which can characterize the system's operating status and determine the next state evolution. In the embodiments of the present invention, in order to satisfy the Markov property, the thermodynamic and hydraulic coupling characteristics can be covered simultaneously, and the following state information is constructed as shown in the following formula (1):
[0044] (1);
[0045] in, It can be Current state It can be Secondary flow rate at any given time; It can be Secondary side temperature at any given time; It can be Traffic flow vectors at each building entrance at any given time; It can be Temperature vectors at each building entrance at any given time; It can be The indoor temperature vector of the most unfavorable user in each building at any given time; It can be outdoor temperature at all times; It can be The opening degree of the valve on each side at any given time; It can be The opening degree of each building valve at all times; It can be The power of the secondary circulation pump at all times.
[0046] Action information It can be a control command applied to the environment by the control agent. The control target includes adjusting the primary regulating valve and the valves at each building entrance. Therefore, the action information is as shown in the following formula (2):
[0047] (2);
[0048] in, It can be the opening degree of all controlled valves in the entire heating or fluid transport system at time t. It can be The opening degree of the valve on each side at any given time; It can be The valve opening degree of building 1 at any given time, where [1, 2, ..., Y] are the building identifiers; It can be Momentary action; It can be set to the maximum permissible valve opening variation. It can be the change in the valve opening state vector from time t to the next time. This could be a maximum amplitude limit for a single action. It's understandable that the sample state information and sample action information during the training phase are the same as or similar to the state information and action information during the application and inference phase; details will not be elaborated here.
[0049] Traditional reinforcement learning agents (such as single Actor-Critic agents) are prone to training instability or policy divergence due to overestimation of evaluation values when updating policies. Embodiments of this invention introduce a parallel evaluation network and use its valley (minimum value) as the benchmark for policy updates. This effectively suppresses optimistic bias in value estimation, providing a more conservative and robust value guide, making the policy network's update direction more reliable, improving the stability and convergence speed of the training process, smoothing the training curve, and steadily improving policy performance while avoiding oscillations. The resulting control agent policy is superior, more accurately assessing action risks, thus making more robust decisions in online control, reducing control risks, and improving system reliability. Compared to traditional agents, this valley-value-based evaluation mechanism significantly reduces the risk of training failure, ensures the global optimality of the policy, improves training efficiency, and reduces system resource consumption.
[0050] According to embodiments of the present invention, by updating strategy parameters based on valley evaluation values and a training mechanism involving parameter fusion, the policy robustness and training stability of the control agent in complex heating network environments are effectively improved. This overcomes the problems of complex modeling, difficult maintenance, and model mismatch caused by traditional reliance on prediction models and hydrothermal models. By integrating multi-dimensional real-time states such as the opening degree of the first valve, the flow and temperature of the heat medium at multiple nodes, the temperature of the building area, and equipment operation information, accurate perception of the dynamic operation of the courtyard heating network is achieved. This avoids the problems of small secondary supply and return water temperature difference and excessive circulation pump flow caused by adjustment lag in traditional static or weak dynamic control. This enables the control agent to output valve opening control commands online in real time, directly optimizing the operating state of the courtyard heating network. Without the need for complex prediction models, it adaptively balances the constraints between energy consumption and the opening degrees of multiple valves, reducing system operating energy consumption and improving the dynamic response capability of heating regulation.
[0051] According to an embodiment of the present invention, the performance information of the intermediate agent satisfies the preset conditions including: the rate of change of the average cumulative reward of multiple training iterations is less than a change threshold; the loss function value of the intermediate evaluation network and the loss function value of the intermediate policy network are less than their respective loss thresholds; and the temperature of the target area is greater than or equal to a preset temperature threshold.
[0052] In embodiments of the present invention, the cumulative reward average can refer to the average of the total rewards obtained by the control agent in a complete training round. The rate of change is the relative magnitude of change of this average over multiple consecutive rounds (e.g., the slope or standard deviation of a moving average).
[0053] The loss function value of the intermediate evaluation network being less than the loss threshold can be used to determine whether the accuracy of the value function estimation meets the standard. A loss value consistently below the threshold indicates that the intermediate evaluation network can make a fairly accurate and stable value assessment of the current policy, providing a reliable foundation for policy optimization in the policy network. The loss function value of the intermediate policy network being less than the loss threshold can determine whether the optimization of the policy itself has reached stability. When the policy network loss value converges to a low level, it indicates that the improvement in the policy itself has become very small, and the policy network parameters tend to stabilize.
[0054] According to an embodiment of the present invention, the above method further includes: obtaining value evaluation information of sample action information based on the value evaluation value of sample information pairs composed of sample state information and sample action information, and the action probability of the initial policy network outputting sample action information under sample state information; and combining reward information and value evaluation information in sample information to obtain an updated target value for updating the initial evaluation parameters.
[0055] In embodiments of the invention, the value assessment value can be an estimate of the long-term expected reward for a given "state-action pair" calculated by multiple initial assessment networks, for example, answering the question: "How much cumulative reward can be obtained in the future by performing action a in state s?" The value assessment information can be a quantitative assessment value that combines the value assessment value and the action probability. The reward information can be an immediate scalar signal fed back after the controlling agent performs action a; for example, in a yard heating network, the reward information can be information that integrates comfort (e.g., room temperature deviation), energy consumption (e.g., pump consumption), and control smoothness (e.g., valve action penalty). The updated target value can be a target value used to update the parameters of the current policy network.
[0056] For example, while ensuring indoor thermal comfort, the energy consumption of circulating pumps and heat consumption of pipeline distribution can be reduced, and the valve operation range can be decreased. Therefore, the reward function is constructed as shown in the following formula (3):
[0057] (3);
[0058] in, It can be the reward function at time t; It could be the indoor temperature of the most unfavorable user in the i-th building at time t; It could be the indoor design temperature of a building for winter heating; It could be the outdoor temperature at time t; It can be the calculated outdoor temperature for winter heating; It could be the electrical power of the circulating pump at time t; It could be the specific heat of water; It could be the entrance flow of the i-th building at time t; It could be the temperature of the secondary side at any given time; It could be the entrance temperature of the i-th building at time t; It can be the opening degree of each control valve at time t, including primary side valves and each building valve; It can be the set of opening degrees of all controlled valves. It can be a weighting factor.
[0059] Objective function of soft action and evaluation strategy It can be achieved by maximizing the sum of future rewards and policy entropy, as shown in the following formula (4):
[0060] (4);
[0061] in, It can be a strategy; It can be a discount factor (0–1) to control the impact of future rewards; It can be the reward function at time t; It could be a temperature coefficient, controlling the trade-off between "reward" and "entropy"; It can be policy entropy, describing the randomness of actions. , Indicates the strategy The expected reward for the entire trajectory resulting from interaction with the environment.
[0062] Policy Network The Gaussian distribution parameters of the output action are shown in the following formula (5):
[0063] (5);
[0064] in, These can be policy network parameters; It can be the mean of the policy network output; It can be the standard deviation of the policy network output. This can be achieved through reparameterized sampling actions. As shown in the following formula (6):
[0065] (6);
[0066] in, It can be random noise; It can be a Gaussian distribution.
[0067] Traditional technologies struggle to directly handle continuous action spaces (such as valve opening), typically requiring discretization. This can lead to the curse of dimensionality or loss of control precision, and the efficiency of exploration strategy adjustment is low, making it difficult to effectively balance exploration and utilization in complex environments. To address these technical issues, embodiments of this invention output action distribution through a policy network. The calculation of value assessment information can accurately evaluate the value of continuous actions, thereby supporting high-precision continuous control. This achieves smooth and fine adjustment of valve opening, making it more suitable for practical engineering scenarios such as heating systems that require continuous and subtle adjustments, and avoiding control abrupt changes and performance losses caused by discretization.
[0068] According to an embodiment of the present invention, an intermediate agent is obtained by sequentially updating the initial evaluation parameters of multiple parallel initial evaluation networks, the initial policy parameters of the initial policy network, and the initial temperature parameters for controlling the randomness of action information using sample state information and sample action information. This includes: using the error between the value evaluation value and the updated target value as an evaluation loss function, and updating the initial evaluation parameters by optimizing the evaluation loss function to obtain the intermediate evaluation parameters of the intermediate agent; using the weighted sum of the action probability of the sample action information and the valley evaluation value as a policy loss function, and updating the initial policy parameters by optimizing the policy loss function to obtain the intermediate policy parameters of the intermediate agent; determining the policy entropy representing the randomness level of the policy in the current state based on the action probability, with the goal of reducing the difference between the policy entropy and the reference entropy, constructing a temperature loss function, and updating the initial temperature parameters by optimizing the temperature loss function to obtain the intermediate temperature parameters; and updating the initial agent based on the intermediate evaluation parameters, intermediate policy parameters, and intermediate temperature parameters to obtain the intermediate agent.
[0069] In embodiments of the present invention, the evaluation loss function can be a function that measures the difference between the output value (value evaluation value) of the evaluation network and the target value (updated target value). The action probability can be the natural logarithm of the probability that the policy network will select the action information of the evaluated sample given sample state information. The valley evaluation value can be the minimum value taken from the outputs of the two evaluation networks (or their target networks) for stable training in soft action and evaluation policy. The policy loss function can be a function used to update the parameters of the policy network.
[0070] Policy entropy can be an indicator of the randomness of a probability distribution. In embodiments of this invention, it refers to the level of randomness of a policy under a given state, to quantify the policy's exploratory capability. The temperature loss function can be used to update the temperature parameters. An intermediate agent can be an agent whose parameters have been partially optimized after completing one round of updates to the evaluation parameters, policy parameters, and temperature parameters; it represents an intermediate state in the evolution from the initial agent to the final convergent control agent.
[0071] For example, two evaluation networks (double Q network) can be used to reduce overestimation bias. The target Q network in the double Q network takes the form of the following formula (7):
[0072] (7);
[0073] in, It can be Current state; It can be Momentary action; These can be parameters of the target Q-network.
[0074] Intelligent agents interact with the environment to generate experience information sets. Stored in the experience replay pool, from which a batch can be sampled during training. Sample information.
[0075] In traditional techniques, during online updates, there is a strong correlation between continuous sample data, which can easily lead to unstable, oscillating, or even divergent training of the agent network. Traditional tabular methods face the problem of state space explosion and are not suitable for continuous state / action spaces. Furthermore, online learning usually discards data after it has been used once, resulting in low data utilization efficiency. In environments with high computational costs, such as heating simulation, low sample efficiency can lead to long training times.
[0076] To address this issue, embodiments of the present invention break down the correlation between data through experience replay, smooth gradient estimation through batch training, and avoid constantly chasing changing targets by calculating target values and introducing a target network, thereby improving the stability and convergence of training. The experience replay pool allows the agent to repeatedly learn from past experiences, significantly improving sample utilization efficiency. The agent can learn robust policies from relatively few simulation interactions, accelerating training speed and reducing the demand for simulation resources, making it economical and feasible to train practical control policies on a digital twin platform.
[0077] For updating the initial policy network, the update target value (soft target value) can be calculated first, as shown in the following formula (8):
[0078] (8);
[0079] in, It can be the first Update target value for each sample; It can be the first Rewards for each sample; It can be a discount factor; It could be the action that the policy samples in the next state; It can be a target Q-network; It can be the temperature coefficient; It could be the logarithmic probability of the action sampled in the next state. It can represent the expected value of the average value of all possible actions of the policy in the next state when constructing the target Q value of SAC.
[0080] loss function of Q network As shown in the following formula (9):
[0081] (9);
[0082] in, These can be parameters of a Q-network, which can be updated via gradient descent. ; This could be the batch size of the sampled sample; It can be the first The status of each sample; It can be the first The actions of the sample; It could be the action value of the Q network. It can be the first Update the target value for each sample.
[0083] The parameters of the dual-Q network are updated as shown in the following formula (10):
[0084] (10);
[0085] in, It can be the initial policy network parameters, which are updated through backpropagation. This could be the batch size of the sampled sample; It can be the temperature coefficient; It could be the logarithmic probability of the action; It could be the action value of the Q network.
[0086] For example, at the start of training, a small batch (e.g., 256) of samples can be randomly sampled from the experience replay pool, each containing state information, action information, reward information, and next state information.
[0087] First, an updated target value is calculated for each sample in the batch. This calculation may include: inputting the next state into the current policy network, sampling actions, and calculating the expected value using the target Q-network. Then, the evaluation loss function is calculated, which is the mean squared error between the current two Q-networks' value assessments of the outputs and the updated target value. This loss is minimized using the gradient descent algorithm, and the initial evaluation parameters of the two Q-networks are updated. Thus, the updated intermediate evaluation parameters (denoted as) are obtained. ).
[0088] For each sample in the batch, the two Q-networks are used to calculate the valley evaluation value of the information pair consisting of state information and action information. Simultaneously, the policy network calculates the logarithmic probability of choosing the action given the state information. Then, a policy loss function is constructed, and this loss function is minimized using gradient descent to update the initial policy parameters of the policy network, resulting in the updated intermediate policy parameters.
[0089] Finally, the temperature parameters are updated based on the behavior of the intermediate policy parameters. For each sample in the batch, the action probability of the policy network under the state information is calculated, and the policy entropy is estimated. Then, a temperature loss function is constructed, and this loss is minimized by gradient descent to update the initial temperature parameters, obtaining the intermediate temperature parameters. Based on this, the initial parameters of the initial agent have been updated to the intermediate parameters, resulting in the intermediate agent.
[0090] According to an embodiment of the present invention, the temperature parameter can be automatically adjusted to make the strategy entropy close to the target entropy, as shown in the following formula (11):
[0091] (11);
[0092] in, It could be a loss of temperature parameters; It can be the temperature coefficient; This could be the batch size of the sampled sample; It could be the logarithmic probability of the action; It could be the target entropy.
[0093] The update rule for the temperature coefficient can be shown in the following formula (12):
[0094] (12);
[0095] in, It could be the learning rate of the temperature coefficient; It could be the gradient of the loss with respect to the temperature coefficient.
[0096] Traditional technologies often lead to agents getting stuck in local optima during training, resulting in insufficient policy exploration, difficulty in discovering globally superior cooperative control patterns, and inaccurate value estimation, which can cause policy improvement to go in the wrong direction. To address this problem, embodiments of this invention reduce overestimation of value by using two parallel initial evaluation networks (dual-Q networks) and valley value evaluations. This makes the policy optimization benchmark more conservative and reliable, and the goal of maximizing entropy encourages thorough exploration. This allows the agent to discover valve cooperative strategies with better global performance than traditional experience, achieving a better balance between energy saving and comfort. Ultimately, the learned control strategy outperforms traditional methods, enabling truly optimal operation of the courtyard heating network.
[0097] According to an embodiment of the present invention, the method further includes: weighting the ambient temperature deviation using the weight of the temperature deviation between the building ambient temperature and the preset temperature to obtain an ambient temperature weighted result, which serves as ambient temperature difference penalty information; weighting the power information using the weight of the power information in the equipment operation information to obtain a power weighted result, which serves as equipment power penalty information; weighting the temperature deviation using the weight of the temperature deviation between the heat medium temperature of the courtyard heating network and the heat medium temperature at the building entrance to obtain a heat medium temperature weighted result, which serves as heat medium temperature difference penalty information; weighting the opening deviation using the weight of the opening deviation between the valve opening at the previous time and the valve opening at the current time to obtain an opening weighted result, which serves as opening deviation penalty information; and combining the ambient temperature difference penalty information, the equipment power penalty information, the heat medium temperature difference penalty information, and the opening deviation penalty information to obtain reward information.
[0098] In embodiments of the present invention, the building ambient temperature may include indoor temperature and outdoor temperature. Correspondingly, the ambient temperature difference penalty information may include indoor temperature difference penalty information and outdoor temperature difference penalty information. The ambient temperature difference penalty information can be a scalar value obtained by multiplying the ambient temperature deviation by its weight; for example, the indoor temperature deviation multiplied by weight d1 is the indoor temperature penalty information, and the outdoor temperature deviation multiplied by weight d2 is the outdoor temperature penalty information. The electrical power information may be the operating electrical power of the secondary side circulation pump of the courtyard heating network, which can be obtained in real time through a smart meter or the pump's frequency converter. The equipment electrical power penalty information can be a scalar value obtained by multiplying the electrical power information by its weight d3.
[0099] The heat transfer medium temperature difference penalty information can be a heat loss penalty term calculated by combining the heat transfer medium temperature deviation, fluid specific heat capacity, flow rate, and weight d4. The opening deviation penalty information can be a scalar value obtained by multiplying the opening deviation by its weight d5. By adding the above weighted penalty information, the reward information can be obtained, and then the reward information can be obtained through the reward function in the above formula (3).
[0100] In embodiments of the present invention, the weights d1 to d5 corresponding to each sub-item in the reward function can respectively reflect the optimization focus of different operational objectives. Weight d1 may be the indoor temperature deviation term corresponding to the most unfavorable user, used to characterize heating comfort constraints; weight d2 may be the outdoor meteorological condition correction term, used to enhance the adaptability of the control strategy to different outdoor temperature conditions; weight d3 may be the circulating pump power penalty term, used to suppress system power consumption; weight d4 may be the heat loss penalty term corresponding to the transmission and distribution process, used to improve the overall thermal efficiency of the system; and weight d5 may be the valve opening change penalty term, used to constrain the control action amplitude and improve the system's operational stability.
[0101] Under different operational requirements, the above weighting coefficients can be configured or adjusted to achieve different optimization objectives. For example, in an operational scenario prioritizing heating comfort, d1 and d2 can be appropriately increased, while the values of d3 and d4 can be relatively decreased; in a scenario where energy-saving operation is the primary objective, the values of d3 and d4 can be increased; and in a scenario emphasizing system operational stability and equipment protection, the value of d5 can be increased to suppress frequent or drastic valve adjustment actions.
[0102] Traditional technologies focus on a relatively singular objective. Even when multiple objectives are considered, adjusting the optimization objective (such as shifting from prioritizing comfort to prioritizing energy conservation) requires redesigning or significantly modifying the control logic and parameters, resulting in a large workload and inflexibility. The embodiments of this invention incorporate multiple objectives into the optimization framework simultaneously, flexibly reflecting their relative importance through weight coefficients, thereby achieving optimal overall performance. Under different operational requirements, the optimization focus can be changed simply by adjusting the weight coefficients in the reward function.
[0103] According to an embodiment of the present invention, the intermediate parameters are updated based on an update factor for controlling the update speed of intermediate parameters of the intermediate agent and the parameter fusion result, including: using the update factor to perform weighted fusion of the updated value of the initial evaluation parameter and the intermediate value of the intermediate evaluation parameter to obtain a weighted result; and determining the weighted result as the updated value of the intermediate evaluation parameter so that the intermediate evaluation parameter approaches the target evaluation parameter infinitely.
[0104] In embodiments of the present invention, the updated value of the initial evaluation parameter can refer to the new value of the evaluation network parameter, which is considered superior after gradient descent optimization. Weighted fusion can refer to the linear interpolation of the two parameter values using an update factor. The weighted result can refer to the result obtained by weighted linear combination of the target evaluation parameter (current target network parameter) and the intermediate evaluation parameter (currently updated evaluation network parameter) using the update factor. By repeatedly performing soft updates, the intermediate evaluation parameter is made to approach the target evaluation parameter.
[0105] To ensure training stability, a soft update can be performed on the evaluation network, as shown in the following formula (13):
[0106] (13);
[0107] in, These can be the target Q-network parameters; It can be the current Q network parameters; It can be a soft update factor.
[0108] When the average cumulative reward change is less than a preset threshold for several consecutive rounds during the training of the agent, and the changes in the evaluation loss function, policy loss function and temperature loss function all converge to the stable range, and the operating indicators of the heating system (indoor temperature) all reach the steady-state tolerance range, it can be considered that the action and evaluation policy have converged.
[0109] In traditional techniques, when training complex networks to handle high-dimensional state spaces (such as the full state of a heating system), hard updates can easily lead to policies failing to converge to meaningful results, or the convergence process being extremely fragile and having a low success rate. The embodiments of this invention, through stable target values, make the learning and policy optimization processes of the evaluation network smoother and more controllable. The agent can steadily improve along a more defined gradient direction, significantly improving the success rate and reliability of converging to a high-performance policy. This ensures that the agent can reliably learn effective control policies from interactions with the digital twin environment, reducing the risk of training failure and time costs.
[0110] According to an embodiment of the present invention, a control agent is used to fuse the opening degree of the first valve between the heating station and the heat source, the flow rate and temperature of the heat medium in the courtyard heating network between the heating station and the building, the flow rate and temperature of the heat medium at the building entrance, the temperature of the target area in the building, and equipment operation information to obtain the current operating status information of the courtyard heating network. This includes: performing time-series alignment processing on the opening degree of the first valve, the flow rate and temperature of the heat medium in the courtyard heating network, the flow rate and temperature of the heat medium at the building entrance, the temperature of the target area, and equipment operation information relative to the current time to obtain aligned multi-source time-series data; dividing the multi-source time-series data into data segments of preset duration based on the current time to obtain a data sequence; determining the time-series feature vector based on the mean, variance, and trend slope obtained by processing the data sequence; and weighting the time-series feature vector according to the weights determined by the importance of each of the multi-source time-series data to obtain a weighted result, and using the weighted result as the current operating status information.
[0111] In embodiments of the present invention, the control agent can be an agent trained based on soft actions and evaluation strategies. The first valve opening can be the opening of the primary-side regulating valve of the heating station, which determines the flow rate of the high-temperature water on the primary side. The heat medium flow rate and temperature of the courtyard heating network can refer to the total flow rate and supply water temperature at the secondary-side outlet of the heating station. The target area can be the indoor area within each building that is most unfavorable for heating. Equipment operation information can include energy consumption information such as the power consumption of the circulating pump.
[0112] For sensors from different locations, their acquisition timestamps may have slight offsets or transmission delays. Time series alignment refers to data preprocessing that unifies these multi-source time series data onto the same time base. The data sequence can be a multi-dimensional data sequence formed by extracting historical data of a fixed duration (e.g., the past 30 minutes) from the aligned multi-source time series data, with the current control time as the endpoint. The mean, variance, and trend slope reflect the static level, dynamic fluctuation, and direction of change of the parameters, respectively. The time series feature vector can be a comprehensive vector composed of the mean, variance, trend slope, and other features of all parameters concatenated in sequence. The current operating state information can be determined by using the weighted feature vector as the final state representation input to the control agent, i.e., the state vector in reinforcement learning.
[0113] According to embodiments of the present invention, interactive environment information of an intelligent agent can be constructed through multiple physical models of heating stations, pipelines, and buildings.
[0114] For example, using a regulating valve model, the primary flow rate of a heating station can be calculated by the valve opening. In the primary and secondary flow control of a heating station, an equal percentage regulation strategy can be used for valve regulation. Based on the inherent characteristics of the equal percentage regulating valve, the flow rate-opening relationship of the valve can be derived as shown in the following formula (14):
[0115] (14);
[0116] in, It can be the mass flow rate of the pipeline; It can be used to describe the flow capacity of a valve, reflecting the maximum flow rate of the valve under a unit pressure difference; The pressure difference across the valve can be measured using a sensor; It can be the valve's adjustable ratio, representing the flow rate ratio between the valve's maximum opening and minimum controllable opening; The valve opening (usually 0-1) reflects the proportion of the valve core's current position to its maximum stroke. Because the inherent flow characteristics of an equal percentage control valve exhibit a significant non-linear relationship, it can provide high regulation accuracy within a small opening range and greater flow capacity at large openings, making it suitable for heating systems with large load variations.
[0117] A plate heat exchanger model can be used to calculate the outlet temperature of each component based on the primary and secondary flow rates and inlet water temperature of the heating station. In regulation analyses with timescales ranging from minutes to hours, the internal heat capacity is much smaller than that of the pipe network, resulting in extremely fast dynamic response. Therefore, it can be considered a quasi-steady-state element. The thermodynamic process of the plate heat exchanger can be described using the Effectiveness-NTUMethod (ε-NTU) strategy, yielding the heat capacity flow rates on both sides. and As shown in the following formula (15):
[0118] (15);
[0119] in, It can be the heat capacity flow rate of the primary fluid. It can be the heat capacity flow rate of the secondary fluid. It can be the density of water; It can be the specific heat of water; The mass flow rate of the primary side pipeline can be calculated using a regulating valve model. The mass flow rate of the secondary side pipeline can be measured by a sensor. It can be the smaller value of the heat capacity flow rate on both sides; This can be the larger of the heat capacity flow rates on both sides. Calculate the heat capacity ratio. and number of heat transfer units As shown in the following formula (16):
[0120] (16);
[0121] in, It can be the overall heat transfer coefficient of the heat exchanger; This can be the effective heat transfer area of the heat exchanger. Calculate the heat exchanger's effectiveness. As shown in the following formula (17):
[0122] (17);
[0123] in, It can be the heat capacity ratio; This can be the number of heat transfer units. And it can calculate the heat transferred by the heat exchanger. As shown in the following formula (18-1):
[0124] (18-1);
[0125] in, The primary inlet temperature can be measured by a sensor. The secondary side inlet temperature can be measured by a sensor. Secondary water supply temperature (secondary side outlet temperature) The calculation method is shown in the following formula (18-2):
[0126] (18-2);
[0127] in, This can be the secondary side inlet temperature; It can be the secondary side heat capacity flow rate.
[0128] Using a pipe network model, the inlet flow rate and water temperature for each user are calculated based on the outlet flow rate and water temperature at the secondary side of the heating station. In actual operation, the secondary pipe network for courtyard heating exhibits significant weak transient characteristics: large water volume, low flow velocity, and strong heat capacity and storage capacity, resulting in slow changes in the hydraulic state over time. The circulating pumps employ constant differential pressure control or slowly adjusted variable frequency control, making the rate of change of hydraulic variables much lower than the adjustment solution step size. Therefore, in engineering calculations and operational analysis, the courtyard heating network can be considered a quasi-steady-state system, and a steady-state hydraulic model can be used for solution. This model not only reduces computational complexity but also effectively reflects the hydraulic distribution law of the pipe network under design and typical operating conditions.
[0129] We can take two cross sections (section 1 and section 2) of any pipe segment, and its hydraulic characteristics can be expressed by Bernoulli's equation, as shown in the following formula (18-3):
[0130] (18-3);
[0131] in, , It can be the elevation at both cross-sections of the pipeline; , It can be the static pressure at both ends of the pipeline; It can be the density of water; It can be gravitational acceleration; , It can be the average flow velocity at both cross-sections of the pipe; This can be considered as the drag loss of a fluid between two cross sections. Drag loss This includes friction loss and local resistance loss. Friction loss can be... The calculation is performed using Darcy's formula, as shown in formula (19):
[0132] (19);
[0133] in, It can be the friction loss coefficient, which is determined by the Reynolds number, the relative roughness of the pipe wall, and the pipe diameter. This can be the pipe length; This can be the pipe diameter; This can be the average flow velocity across the cross section. Local resistance loss. As shown in the following formula (20):
[0134] (20);
[0135] in, It can be the local resistance loss coefficient; The average velocity across the cross section can be used. The fluid velocity within the same horizontal pipe is considered constant, and the pressure difference across the pipe is shown in formula (21):
[0136] (twenty one);
[0137] in, It can reduce pipeline pressure drop; It can be the mass flow rate of the pipeline; This can be the impedance of the pipe. Pipe impedance. As shown in the following formula (22):
[0138] (twenty two);
[0139] in, This can be the pipe diameter; It can be the friction loss coefficient along the friction path; This can be the local resistance loss coefficient of the i-th pipe component. Assume a centralized heating network has M nodes and N branches; its structure can be represented by the correlation matrix. Description, dimension is , elements Represents a node With branches The relationship is shown in the following formula (23):
[0140] (twenty three);
[0141] Correlation Matrix The rank is equal to ,Right now 1-th order matrix any The rows are linearly independent. From the incidence matrix... Remove the main heat source The corresponding row yields The fundamental correlation matrix of order The basic correlation matrix is a full-rank matrix. The vector composed of the mass flow rates of each branch of the pipeline network is G, as shown in the following formula (24):
[0142] (twenty four);
[0143] in, The mass flow rate of each branch can be calculated. Let Q' be the vector composed of the net flow rates of the nodes in the pipeline network, as shown in the following formula (25):
[0144] (25);
[0145] in, The mass flow rate of each node can be calculated. According to the node flow balance, that is, the net flow of a node is equal to the total flow inflow from the branches adjacent to that node minus the total flow outflow, as shown in the following formula (26):
[0146] (26);
[0147] in, It can be a vector composed of the mass flow rates of each node in the pipeline network; G can be the basic correlation matrix; G can be a vector composed of the mass flow rates of each branch of the pipeline network. If the heating pipeline network includes both supply and return water networks, it also includes heat sources and heating stations. Without considering pipeline leakage, the net flow rate of all nodes in the pipeline network diagram is 0, i.e. As shown in the following formula (27):
[0148] (27);
[0149] The centralized heating network includes Each of the three independent basic loops has a structure that can be derived from a basic loop matrix. Description, dimension is , elements Representing a basic circuit With branches The relationship between them is shown in the following formula (28):
[0150] (28);
[0151] A basic loop matrix composed of independent basic loops is a full-rank matrix. The vector formed by the pressure drops of each branch in the pipeline network is... As shown in the following formula (29):
[0152] (29);
[0153] in, It can be used to measure the pressure drop of each branch in the pipeline network. According to the circuit pressure balance, that is, the circuit pressure drop of each basic circuit is 0, as shown in the following formula (30):
[0154] (30);
[0155] in, It can be a basic loop matrix; It can be a vector composed of the pressure drops of each branch in the pipeline network; This can be the pump head vector for each branch. The steady-state hydraulic model of the pipe network is solved using Newton's iteration method. As shown in the following formula (31):
[0156] (31);
[0157] in, It can be the basic correlation matrix; G can be a vector composed of the mass flow rates of each branch of the pipeline network. It can be a basic loop matrix; It can be a vector composed of the pressure drops of each branch in the pipeline network; It can be the pump head vector for each branch. This indicates that a steady-state hydraulic equilibrium has been reached. Perform a first-order Taylor expansion, as shown in formula (32):
[0158] (32);
[0159] in, It can be a vector increment composed of the mass flow rates of each branch of the pipeline network. .
[0160] (33);
[0161] in, For steady-state hydraulic model Differentiate the vector G composed of the mass flow rates of each branch of the pipeline network on its boundary. This can be the pump head vector for each branch. Setting the right-hand side of equation (32) to 0, we obtain the following equation (34):
[0162] (34);
[0163] in, It can be the first The vector increment of the mass flow rate of each branch of the pipeline network during the next iteration; It can be the first The mass flow vector of each branch of the pipeline network at the next iteration. The termination condition for the iteration is shown in formula (35):
[0164] (35);
[0165] in, It can be the first The mass flow vector of each branch of the pipeline network at the next iteration; It can be a given residual limit. When secondary network fluid (e.g., hot water) flows along a pipe, it can be considered as incompressible flow, and the heat conduction along the pipe length and axial direction can be ignored. Its temperature field is shown in the following formula (36):
[0166] (36);
[0167] in, It can be the temperature of the fluid in the pipe; It can be the mass flow rate of the pipeline; It can be the density of water; It can be the cross-sectional area of the pipe; It can be the distance the fluid flows along the axis of the pipe; It can be based on ambient temperature; It can be the specific heat of water; The total heat transfer resistance from the fluid to the external environment can be used. It is discretized using a first-order implicit upwind differential calculation scheme, as shown in the following formula (37):
[0168] (37);
[0169] in, It can be the first control body in the Temperature at any moment; It can be the time step; It can be the first The mass flow rate of the pipeline at any given time; It can be the spatial step size; It can be based on ambient temperature; It can be the specific heat of water; The total thermal resistance can be considered as the heat transfer resistance from the fluid to the external environment. The recursive formulas for the temperatures of each control volume are obtained, as shown in formula (38) below:
[0170] (38);
[0171] in, It can be the first control body in the Temperature at any moment; It can be the time step; It can be the first The mass flow rate of the pipeline at any given moment. time, , Since all the temperatures are known, the temperatures of all control volumes in the pipeline can be solved.
[0172] Using the above method, the inlet flow rate and water temperature of each building can be calculated from the outlet flow rate and water temperature of the secondary side of the heating station. For all users in each building, the water supply temperature of that user is taken as the building inlet water temperature, and the flow rate is taken as the average of the building inlet flow rate. The outlet water temperature of the building is obtained by mixing the return water of each user, as shown in the following formula (39):
[0173] (39);
[0174] in, This can be the outlet water temperature of the i-th building; It can be the quality flow for the j-th user in the i-th building; It can be the outlet water temperature of the j-th user in the i-th building; It can be the mass flow rate of the i-th building.
[0175] The hydraulic regulation of courtyard heating networks usually relies on circulating pumps to provide the necessary head. In order to accurately describe the performance of circulating pumps under different operating conditions in a steady-state hydraulic model, a flow-head characteristic model of the circulating pump can be established. When running at a fixed speed (e.g., rated 50Hz), the head-flow relationship of the circulating pump can be fitted with a binomial formula using data points provided by the manufacturer or field test data to obtain its characteristic curve, as shown in the following formula (40):
[0176] (40);
[0177] in, This can be used to determine the head of the circulating pump; This can be the flow rate of the circulating pump; These can be the head-flow quadratic coefficient, linear coefficient, and constant coefficient of the circulating pump, respectively. To improve energy efficiency, variable frequency control is commonly used for circulating pumps in courtyard heating networks. According to the similarity law of centrifugal pumps, the main performance parameters of the same pump at different speeds have the following relationship (the subscript 0 indicates the parameter at the rated frequency of 50 Hz), as shown in the following formula (41):
[0178] (41);
[0179] in, This can be used to determine the head of the circulating pump; This can be the initial flow rate of the circulating pump; This can be the rotational speed of the circulating pump; This can be the initial speed of the circulating pump. This can be the frequency of the circulating pump; The frequency ratio can be represented by the ratio of the operating frequency to the rated frequency of the circulating pump. By introducing the similarity law into the constant frequency characteristic equation (40), the flow-head relationship of the circulating pump at any operating frequency can be obtained, as shown in the following formula (42):
[0180] (42);
[0181] in, The head of the circulating pump can be set at any operating frequency. The flow-efficiency characteristic curve of the circulating pump can be binomially fitted using data points provided by the manufacturer or field test data, as shown in the following formula (43):
[0182] (43);
[0183] in, This can be used to assess the efficiency of the circulating pump; This can be the flow rate versus efficiency coefficient of the circulating pump. The electrical power of the circulating pump. The calculation method is shown in the following formula (44):
[0184] (44);
[0185] in, It can be the electrical power of the circulating pump; It can be the density of water; This can be the flow rate of the circulating pump; This can be used to determine the head of the circulating pump; This can be used to measure the efficiency of the circulating pump.
[0186] For building and heating terminal models, outlet water temperature and indoor temperature can be calculated using the inlet flow rate and water temperature of each user. To link the hydraulic and thermal calculations of the courtyard heating network with the building's heating demand, a high-precision building-terminal-pipeline coupled simulation platform is constructed using a co-simulation framework of building energy consumption simulation software and a programming language. This platform enables dynamic feedback between node supply and return water temperatures, indoor temperatures, and system operating parameters. Traditional building energy consumption simulation software possesses mature simulation capabilities for building thermal environments and equipment systems; however, its internal real-time control functions rely on the runtime language of the software, limiting its expressive power and making it unsuitable for complex control strategies and reinforcement learning algorithms. Therefore, secondary development is performed using the programming language interface provided by the building energy consumption simulation software. This allows for real-time data interaction with the courtyard heating network model within the programming language, constructing a flexibly expandable dynamic control framework.
[0187] For example, a geometric model can be built in 3D software based on architectural computer-aided design drawings. The 3D geometry of the building can be reconstructed based on the building floor plans, elevations, and sections to ensure geometric closure and meet energy balance requirements. Building energy modeling software plugins can be used to annotate building components, define spatial and thermal zones based on building information, and synchronize this information to the building energy modeling software. Building physical attributes and system information can be defined in the building energy modeling software, including meteorological files and operating condition settings, building component annotations and material property settings, heating loops, and heating terminal equipment. The output is used as an input file for building energy consumption simulation software, and dynamic feedback between building simulation and the courtyard heating network model is achieved through the software's interface: real-time inlet flow and water temperature of each user are obtained from the courtyard pipe network model; the building energy consumption simulation software can calculate the outlet water temperature and indoor temperature of each user in real time and output the results to a programming language; the outlet flow and water temperature of each user will be used as the inlet conditions of the courtyard heating network return water network to calculate the secondary inlet flow and water temperature of the heating station, thus forming a complete system-level full-process simulation.
[0188] According to an embodiment of the present invention, controlling the opening degree of the first valve and the valve opening degree on the building entrance side using current operating state information to obtain the subsequent operating state of the courtyard heating network includes: determining target action information from multiple initial action information based on the current operating state information; parsing and processing the target action information to obtain updated information of the opening degree of the first valve and the valve opening degree on the building entrance side, and combining the updated information with the valve opening degree at a previous time to obtain the valve opening command at the current time; sending the valve opening command to the actuator corresponding to the opening degree of the first valve and the valve on the building entrance side to maintain or update the valve opening degree, thereby obtaining the subsequent operating state of the courtyard heating network.
[0189] In embodiments of the present invention, multiple initial action information can refer to a set of candidate actions or action distribution parameters that the agent's policy network may output given the current state, such as a multidimensional Gaussian distribution. Target action information can be a determined multidimensional vector obtained by sampling (or directly taking the mean) from the action distribution, where each dimension corresponds to the opening change of the controlled valve (primary valve, valves at each building entrance).
[0190] The valve opening at the previous moment refers to the valve opening value maintained by each valve actuator in the previous control cycle (time t-1). The actuator can refer to the electric regulating valve or its drive unit installed in the primary side pipeline and at each building entrance, capable of receiving control signals and driving the valve core to a designated position. The subsequent operating state of the courtyard heating network refers to the new state the courtyard heating system enters after the control agent issues a valve opening command and it is executed by the actuator; it is the result of the current control action acting on the physical system.
[0191] For example, changes in valve opening can disrupt hydraulic and thermal operating conditions. After a dynamic process involving one sampling cycle, the system enters a new steady-state or dynamic process. Sensors collect data at this point, which is processed to form the subsequent operating state of the courtyard heating network. At this time, the control system clock enters the next control cycle (time t+1), and the subsequent operating state becomes the new current operating state information. The above steps are repeated to achieve continuous closed-loop optimization control.
[0192] Traditional reactive control technologies, which adjust only after a deviation occurs, suffer from severe response lag in heating systems with high thermal inertia. This can easily lead to overshoot, oscillation, or slow adjustment. Furthermore, independent control of each loop lacks coordination, making it prone to coupling interference and hindering global optimization. The embodiments of this invention, through a control agent that calculates the coordinated actions of all valves based on system-level state information, achieve system-level performance optimization by considering the coupling effects between valve actions. This solves the hydraulic imbalance problem, enables precise on-demand heat distribution and coordinated optimized system operation, significantly improving heating quality and system energy efficiency.
[0193] According to an embodiment of the present invention, determining target action information from multiple initial action information based on current operating state information includes: performing multi-level nonlinear transformation on the current operating state information to obtain feature parameters representing the probability distribution of actions; determining multiple preliminary action information by probability sampling based on the mean vector and standard deviation vector in the feature parameters, and performing numerical range restriction processing on the multiple initial action information to obtain target action information.
[0194] In embodiments of the present invention, the current operating state information refers to the comprehensive state vector of the courtyard heating system obtained by the agent at a specific time (e.g., time t). Multi-layer nonlinear transformation may include extracting abstract features from high-dimensional, complex system states and learning complex nonlinear functions that map states to actions. The feature parameters of the action probability distribution may refer to the final output of the policy network, defining the probability distribution of the actions the agent should take in a given state. The mean vector may represent the optimal or most likely effective action that the agent believes based on learned experience. The standard deviation vector can be used to control the agent's exploratory behavior; a larger standard deviation means that the agent tries more different actions near the mean to explore potentially better strategies.
[0195] Probabilistic sampling can be a method of randomly generating samples based on a given probability distribution, such as randomly sampling from a multivariate Gaussian distribution defined by a mean vector and a standard deviation vector, to obtain a specific action vector. Multiple preliminary action information can be one or more candidate action vectors obtained through probabilistic sampling, each vector representing a set of control commands (such as the opening changes of various valves) that may be executed under the current strategy. Numerical range constraint processing can be a post-processing step on the sampled preliminary action information, limiting the values of each dimension to a reasonable range acceptable to the physical system or actuator. Target action information can refer to the specific action vector that, after probabilistic sampling and numerical range constraint processing, is finally determined and will be sent to the environment for execution.
[0196] Traditional techniques are feasible in discrete action spaces, but inefficient in continuous spaces because random actions are likely meaningless, parameter space noise adds noise to the policy network parameters, the perturbation is not smooth, the exploration direction is blind and unrelated to the state, and each state corresponds to only one action, making it difficult to express multiple optimal or suboptimal actions that may exist in a specific state, and also difficult to handle tasks that require random policies (such as incomplete information games).
[0197] Embodiments of this invention set the direction and magnitude of exploration to be related to the current state, exploring more in uncertain states and less in familiar states. The exploration is guided and adaptive, and the stochastic policy can represent complex action preferences and naturally handle environments with randomness. The maximum entropy framework encourages policies to cover multiple possible good actions, not just one, enabling the learning of more robust and expressive policies. Especially in environments like heating systems with multiple feasible control schemes, it can discover smoother and safer control strategies, improving the intelligence and efficiency of exploration. This allows the agent to discover high-performance policies with fewer interactions, accelerating the learning process.
[0198] Based on the above-described method for controlling the operation of a courtyard heating network, this invention also provides a device for controlling the operation of a courtyard heating network. The following will be combined with... Figure 3 The device is described in detail.
[0199] Figure 3 A structural block diagram of an operation control device for a courtyard heating network according to an embodiment of the present invention is shown.
[0200] like Figure 3 As shown, the operation control device 300 of the courtyard heating network in this embodiment includes an update module 310, a determination module 320 and a control module 330.
[0201] The update module 310 is used to sequentially update the initial evaluation parameters of each of the multiple parallel initial evaluation networks, the initial policy parameters of the initial policy network, and the initial temperature parameters used to control the randomness of the action information in the initial agent, using sample state information and sample action information, to obtain an intermediate agent. The initial policy parameters are updated using the valley value among the multiple output values of the multiple initial evaluation networks. In one embodiment, the update module 310 can be used to perform the operation S210 described above, which will not be repeated here.
[0202] The determining module 320 is used to update the intermediate parameters based on the update factor used to control the update speed of the intermediate agent and the parameter fusion result. If the performance information of the intermediate agent meets preset conditions, the intermediate agent is determined as the controlling agent. The parameter fusion result is obtained by fusing the initial evaluation parameters and the intermediate evaluation parameters. In one embodiment, the determining module 320 can be used to perform the operation S220 described above, which will not be repeated here.
[0203] The control module 330 is used to fuse information such as the opening degree of the first valve between the heating station and the heat source, the flow rate and temperature of the heat medium in the courtyard heating network between the heating station and the building, the flow rate and temperature of the heat medium at the building entrance, the temperature of the target area in the building, and equipment operation information using a control agent. This results in the current operating status information of the courtyard heating network. The control module then uses this information to control the opening degree of the first valve and the valve at the building entrance, thus obtaining the subsequent operating status of the courtyard heating network. In one embodiment, the control module 330 can be used to execute the operation S230 described above, which will not be repeated here.
[0204] According to an embodiment of the present invention, based on the update module 310, determination module 320, and control module 330 in the operation control device 300, the strategy robustness and training stability of the control agent in complex heating network environments are effectively improved through the training mechanism of updating strategy parameters based on valley evaluation values and parameter fusion. This overcomes the problems of complex modeling, difficult maintenance, and model mismatch caused by traditional reliance on prediction models and hydraulic-thermal models. By integrating multi-dimensional real-time states such as the opening degree of the first valve, the flow and temperature of the heat medium at multiple nodes, the temperature of the building area, and equipment operation information, the control agent can accurately perceive the dynamic operation of the courtyard heating network. This avoids the problems of small secondary supply and return water temperature difference and excessive circulation pump flow caused by adjustment lag in traditional static or weak dynamic control. This allows the control agent to output valve opening control commands online in real time, directly optimizing the operating state of the courtyard heating network. Without the need for complex prediction models, it adaptively balances the constraints between energy consumption and the opening degree of multiple valves, reducing system operating energy consumption and improving the dynamic response capability of heating regulation.
[0205] According to an embodiment of the present invention, the above-described apparatus further includes: an evaluation information determination module and a combination module. The evaluation information determination module is used to obtain the value evaluation information of the sample action information based on the value evaluation value of the sample information pair composed of sample state information and sample action information, and the action probability of the initial policy network outputting the sample action information under the sample state information; the combination module is used to combine the reward information and the value evaluation information in the sample information to obtain an updated target value for updating the initial evaluation parameters.
[0206] According to an embodiment of the present invention, the update module 310 includes: a parameter determination submodule, a parameter update submodule, a construction submodule, and an agent update submodule. The parameter determination submodule is used to use the error between the value assessment value and the update target value as an evaluation loss function, and to update the initial evaluation parameters by optimizing the evaluation loss function to obtain intermediate evaluation parameters of the intermediate agent. The parameter update submodule is used to use the weighted sum of the action probabilities of sample action information and the valley value assessment value as a policy loss function, and to update the initial policy parameters by optimizing the policy loss function to obtain intermediate policy parameters of the intermediate agent. The construction submodule is used to determine the policy entropy representing the randomness level of the policy in the current state based on the action probabilities, with the goal of reducing the difference between the policy entropy and the reference entropy, to construct a temperature loss function, and to update the initial temperature parameters by optimizing the temperature loss function to obtain intermediate temperature parameters. The agent update submodule is used to update the initial agent based on the intermediate evaluation parameters, intermediate policy parameters, and intermediate temperature parameters to obtain the intermediate agent.
[0207] According to an embodiment of the present invention, the above-mentioned device further includes: a first weighting module, a second weighting module, a third weighting module, a fourth weighting module, and a combination module. The first weighting module is used to weight the ambient temperature deviation using the weights of the temperature deviation between the building ambient temperature and a preset temperature, obtaining an ambient temperature weighted result as ambient temperature difference penalty information; the second weighting module is used to weight the electrical power information using the weights of the electrical power information in the equipment operation information, obtaining an electrical power weighted result as equipment electrical power penalty information; the third weighting module is used to weight the temperature deviation using the weights of the temperature deviation between the heat medium temperature of the courtyard heating network and the heat medium temperature at the building entrance, obtaining a heat medium temperature weighted result as heat medium temperature difference penalty information; the fourth weighting module is used to weight the opening deviation using the weights of the opening deviation between the valve opening at the previous time and the valve opening at the current time, obtaining an opening weighted result as opening deviation penalty information; the combination module is used to combine the ambient temperature difference penalty information, the equipment electrical power penalty information, the heat medium temperature difference penalty information, and the opening deviation penalty information to obtain reward information.
[0208] According to an embodiment of the present invention, the determining module 320 includes a fusion submodule and an update value determining submodule. The fusion submodule is used to perform weighted fusion of the updated value of the initial evaluation parameter and the intermediate value of the intermediate evaluation parameter using an update factor to obtain a weighted result; the update value determining submodule is used to determine the weighted result as the updated value of the intermediate evaluation parameter, so that the intermediate evaluation parameter approaches the target evaluation parameter infinitely.
[0209] According to an embodiment of the present invention, the control module 330 includes: an alignment submodule, a partitioning submodule, a vector determination submodule, and a vector weighting submodule. The alignment submodule is used to perform time-series alignment processing on the opening degree of the first valve, the heat medium flow rate and temperature of the courtyard heating network, the heat medium flow rate and temperature at the building entrance, the temperature of the target area, and equipment operation information relative to the current time, to obtain aligned multi-source time-series data; the partitioning submodule is used to divide the multi-source time-series data into data segments of preset duration based on the current time, to obtain a data sequence; the vector determination submodule is used to determine the time-series feature vector based on the mean, variance, and trend slope obtained by processing the data sequence; the vector weighting submodule is used to weight the time-series feature vector according to the weights determined by the importance of each of the multi-source time-series data, to obtain a weighted result, and uses the weighted result as the current operating status information.
[0210] According to an embodiment of the present invention, the control module 330 includes a parsing submodule and an opening update submodule. The parsing submodule is used to determine target action information from multiple initial action information based on current operating state information; parse and process the target action information to obtain updated information on the opening of the first valve and the valve opening on the building entrance side; and combine the updated information with the valve opening at a previous time to obtain the valve opening command at the current time. The opening update submodule is used to send the valve opening command to the actuators corresponding to the first valve opening and the valve on the building entrance side to maintain or update the valve opening, thereby obtaining the subsequent operating state of the courtyard heating network.
[0211] According to an embodiment of the present invention, the parsing submodule includes a conversion unit and an action information determination unit. The conversion unit is used to perform multi-level nonlinear transformation on the current running state information to obtain feature parameters representing the probability distribution of actions; the action information determination unit is used to determine multiple preliminary action information based on the mean vector and standard deviation vector in the feature parameters using probability sampling, and to perform numerical range restriction processing on the multiple preliminary action information to obtain target action information.
[0212] According to an embodiment of the present invention, the performance information satisfies the preset conditions including: the rate of change of the cumulative reward average of multiple training iterations is less than a change threshold; the loss function value of the intermediate evaluation network and the loss function value of the intermediate policy network are less than their respective loss thresholds; and the temperature of the target area is greater than or equal to a preset temperature threshold.
[0213] According to embodiments of the present invention, any plurality of modules among the update module 310, the determination module 320, and the control module 330 may be combined into one module, or any one of these modules may be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules may be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of the present invention, at least one of the update module 310, the determination module 320, and the control module 330 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the update module 310, the determination module 320, and the control module 330 may be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.
[0214] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0215] Those skilled in the art will understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention can be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.
[0216] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.
Claims
1. A method for operating and controlling a courtyard heating network, characterized in that, The method includes: Using sample state information and sample action information, the initial evaluation parameters of each of the multiple parallel initial evaluation networks, the initial policy parameters of the initial policy network, and the initial temperature parameter for controlling the randomness of action information in the initial agent are updated sequentially to obtain an intermediate agent. The initial policy parameters are updated with the valley value among the multiple evaluation values of the multiple initial evaluation networks. The valley value is the minimum value selected from the multiple evaluation values calculated by all parallel evaluation networks for the same state-action pair at each step of updating the policy network. Based on the update factor and parameter fusion result used to control the update speed of intermediate parameters of the intermediate agent, the intermediate parameters are updated. If the performance information of the intermediate agent meets the preset conditions, the intermediate agent is determined as the control agent. The parameter fusion result is obtained by fusing the initial evaluation parameters and the intermediate evaluation parameters. The control agent integrates the opening degree of the first valve between the heating station and the heat source, the flow rate and temperature of the heat medium in the courtyard heating network between the heating station and the building, the flow rate and temperature of the heat medium at the building entrance, the temperature of the target area in the building, and equipment operation information to obtain the current operating status information of the courtyard heating network. The current operating status information is then used to control the opening degree of the first valve and the valve at the building entrance to obtain the subsequent operating status of the courtyard heating network.
2. The method according to claim 1, characterized in that, The method further includes: Based on the value assessment value of the sample information pair composed of the sample state information and the sample action information, and the action probability of the initial policy network outputting the sample action information under the sample state information, the value assessment information of the sample action information is obtained. By combining the reward information from the sample information with the value assessment information, an updated target value is obtained for updating the initial assessment parameters.
3. The method according to claim 2, characterized in that, Using sample state information and sample action information, the initial evaluation parameters of each of the multiple parallel initial evaluation networks, the initial policy parameters of the initial policy network, and the initial temperature parameter used to control the randomness of action information in the initial agent are updated sequentially to obtain the intermediate agent, including: The error between the value assessment value and the updated target value is used as the evaluation loss function, and the initial evaluation parameters are updated by optimizing the evaluation loss function to obtain the intermediate evaluation parameters of the intermediate agent. The weighted sum of the action probability of the sample action information and the valley value evaluation value is used as the policy loss function, and the initial policy parameters are updated by optimizing the policy loss function to obtain the intermediate policy parameters of the intermediate agent. Based on the action probability, the policy entropy representing the randomness level of the policy in the current state is determined. With the goal of reducing the difference between the policy entropy and the reference entropy, a temperature loss function is constructed, and the initial temperature parameter is updated by optimizing the temperature loss function to obtain the intermediate temperature parameter. The initial agent is updated based on the intermediate evaluation parameters, the intermediate policy parameters, and the intermediate temperature parameters to obtain the intermediate agent.
4. The method according to claim 2, characterized in that, The method further includes: The ambient temperature deviation is weighted by the weight of the temperature deviation between the building ambient temperature and the preset temperature, and the ambient temperature weighted result is used as ambient temperature difference penalty information. The power information is weighted using the weights of the power information in the equipment operation information to obtain a power weighting result, which is used as the penalty information for the equipment power. The temperature deviation is weighted by the weight of the temperature difference between the heat medium temperature of the courtyard heating network and the heat medium temperature at the building entrance, and the weighted result of the heat medium temperature difference is used as the heat medium temperature difference penalty information. The opening deviation is weighted by the weight of the opening deviation between the valve opening at the previous time and the valve opening at the current time, and the weighted opening result is used as the penalty information for the opening deviation. The reward information is obtained by combining the environmental temperature difference penalty information, the equipment power penalty information, the heat medium temperature difference penalty information, and the opening deviation penalty information.
5. The method according to claim 1, characterized in that, The intermediate parameters are updated based on the update factor used to control the update rate of the intermediate agent and the parameter fusion result, including: The updated values of the initial evaluation parameters and the intermediate values of the intermediate evaluation parameters are weighted and fused using the update factor to obtain a weighted result; The weighted result is determined as the updated value of the intermediate evaluation parameter, so that the intermediate evaluation parameter approaches the target evaluation parameter infinitely.
6. The method according to claim 1, characterized in that, The control agent integrates the opening degree of the first valve between the heating station and the heat source, the flow rate and temperature of the heat medium in the courtyard heating network between the heating station and the building, the flow rate and temperature of the heat medium at the building entrance, the temperature of the target area in the building, and equipment operation information to obtain the current operating status information of the courtyard heating network, including: The opening degree of the first valve, the flow rate and temperature of the heat medium of the courtyard heating network, the flow rate and temperature of the heat medium at the building entrance, the temperature of the target area, and the equipment operation information are time-series aligned with the current time to obtain aligned multi-source time-series data. Based on the current time, the multi-source time-series data is divided into data segments of preset duration to obtain a data sequence; Based on the mean, variance, and trend slope obtained from processing the data sequence, the time series feature vector is determined; The time series feature vector is weighted according to the weights determined by the importance of each of the multi-source time series data, and the weighted result is used as the current running status information.
7. The method according to claim 1, characterized in that, By controlling the opening degree of the first valve and the valve opening degree on the building entrance side using the current operating status information, the subsequent operating status of the courtyard heating network is obtained, including: The target action information is determined from multiple initial action information based on the current operating status information; The target action information is parsed and processed to obtain updated information on the opening degree of the first valve and the opening degree of the valve on the building entrance side. The updated information is then combined with the valve opening degree at the previous moment to obtain the valve opening command at the current moment. The valve opening command is sent to the actuator corresponding to the first valve opening and the valve on the building entrance side to maintain or update the valve opening, thereby obtaining the subsequent operating state of the courtyard heating network.
8. The method according to claim 7, characterized in that, Based on the current operating status information, the target action information is determined from multiple initial action information, including: The current operating state information is subjected to multi-layer nonlinear transformation to obtain feature parameters representing the probability distribution of actions; Based on the mean vector and standard deviation vector in the feature parameters, multiple initial action information is determined by probability sampling, and the numerical range of the multiple initial action information is restricted to obtain the target action information.
9. The method according to claim 1, characterized in that, The performance information meets the preset conditions, including: The rate of change of the cumulative average reward across multiple training iterations is less than the change threshold; The loss function values of the intermediate evaluation network and the intermediate policy network are both less than their respective loss thresholds; The temperature of the target area is greater than or equal to a preset temperature threshold.
10. An operation control device for a courtyard heating network, characterized in that, The device includes: An update module is used to sequentially update the initial evaluation parameters of each of the multiple parallel initial evaluation networks, the initial policy parameters of the initial policy network, and the initial temperature parameters for controlling the randomness of the action information in the initial agent using sample state information and sample action information, to obtain an intermediate agent. The initial policy parameters are updated with the valley evaluation value among the multiple output values of the multiple initial evaluation networks. The valley evaluation value is the minimum value selected from the multiple evaluation values calculated by all parallel evaluation networks for the same state-action pair at each step of updating the policy network. The determination module is used to update the intermediate parameters based on the update factor and parameter fusion result used to control the update speed of the intermediate agent. When the performance information of the intermediate agent meets the preset conditions, the intermediate agent is determined as the control agent. The parameter fusion result is obtained by fusing the initial evaluation parameters and the intermediate evaluation parameters. The control module is used to integrate the opening degree of the first valve between the heating station and the heat source, the flow rate and temperature of the heat medium in the courtyard heating network between the heating station and the building, the flow rate and temperature of the heat medium at the building entrance, the temperature of the target area in the building, and equipment operation information using the control agent to obtain the current operating status information of the courtyard heating network. The module then uses the current operating status information to control the opening degree of the first valve and the valve at the building entrance to obtain the subsequent operating status of the courtyard heating network.