Cold station operation control method, device, equipment and medium

Through the cold station operation control method combined with deep learning and deep reinforcement learning, the problems of high complexity and insufficient adaptability of cold station operation control are solved, intelligent control of cold stations and efficient energy utilization are achieved, and operation efficiency and stability are improved.

CN120332887AActive Publication Date: 2025-07-18BEIJING ZHONGHE ZHILIAN TECHNOLOGY CO LTD
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510375565.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-18
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The existing cold station operation control methods are complex and have low computing efficiency, difficult to cope with variable working conditions, lack online adaptability, and cannot adjust optimization strategies online.

Method used

Using a combination of deep learning and deep reinforcement learning, by obtaining the operating status and environmental status data of the cold station, determining the probability distribution method based on the data type of the control variable, data sampling is performed to obtain control actions, and a decision-making control model is constructed to realize intelligent control of the cold station.

Benefits of technology

It improves the intelligence level and energy utilization efficiency of cold station operations, can better cope with complex and changeable situations, improves the operating efficiency and stability of cold stations, and reduces operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120332887A_ABST
    Figure CN120332887A_ABST
Patent Text Reader

Abstract

The invention discloses a cold station operation control method, device, equipment and medium, and relates to the technical field of energy management, artificial intelligence and the like, the cold station operation control method comprises the steps that state variables are acquired, and the state variables comprise cold station operation state data and environment state data; when a decision is made based on the state variable, determining a matched probability distribution mode based on the data type of the control variable, and obtaining control variable probability distribution matched with the probability distribution mode; performing data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable; and controlling the cold station to operate based on the control action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of energy management, artificial intelligence, etc., and particularly relates to a cold station operation control method, device, equipment and medium. Background Art

[0002] As an important part of building energy consumption, the energy consumption of the cold station accounts for a large proportion of the total building energy consumption. According to statistics, building energy consumption accounts for about 40% of the global energy demand, and the energy consumption of central air conditioners accounts for 40% - 50% of the total building energy consumption. Among them, the energy consumption of the cold station is the main part of the air conditioner energy consumption, accounting for more than 30% of the total building electricity consumption. Therefore, optimizing the operation control of the cold station and improving energy efficiency are of great significance for reducing building energy consumption and operating costs. The optimization control of the cold station not only helps to save energy and reduce emissions, but also improves the comfort and operation efficiency of the building, and promotes sustainable development.

[0003] The cold station operation control methods proposed by related technologies have high complexity, resulting in low calculation efficiency; and they are not adaptable to the input, making it difficult to cope with changing working conditions; in addition, they lack online self - adaptability and cannot adjust optimization strategies online to deal with emergencies. Summary of the Invention

[0004] The embodiments of the present invention aim to at least solve one of the technical problems in related technologies to some extent. For this reason, an object of the present invention is to provide a cold station operation control method, device, equipment and medium, which realizes the efficient utilization of energy and improves the intelligent level of the cold station.

[0005] The embodiments of the present invention provide a cold station operation control method, which includes: obtaining state variables, where the state variables include cold station operation state data and environmental state data; when making a decision based on the state variables, determining a matching probability distribution method based on the data type of the control variable, and obtaining a control variable probability distribution matching the probability distribution method; performing data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable; and controlling the operation of the cold station based on the control action.

[0006] Exemplarily, the data type of the control variable includes at least one of continuous type, discrete type, and switch type. Determining a matching probability distribution method based on the data type of the control variable includes: when the data type of the control variable is continuous, determining the probability distribution method as a normal probability distribution method; when the data type of the control variable is discrete, determining the probability distribution method as a categorical probability distribution method; when the data type of the control variable is switch type, determining the probability distribution method as a Bernoulli probability distribution method.

[0007] Exemplarily, data sampling is performed on the probability distribution of the control variable to obtain a control action corresponding to the control variable, including: using the preset safety rule data as a sampling constraint, performing data sampling on the probability distribution of the control variable to obtain a control action corresponding to the control variable.

[0008] Exemplarily, the cold station operation state data and the environmental state data include current data, historical data, and future data. The current data includes at least one of the refrigeration load, equipment parameters, system parameters, and weather. The historical data includes at least one of the refrigeration load and weather. The future data includes at least one of the refrigeration load and weather.

[0009] Exemplarily, the cold station operation control method is applied to a decision control model. The decision control model includes an input layer, an intermediate layer, and an output layer, where: the input layer is used to input state variables; the intermediate layer is used to make decisions based on the state variables; the output layer is used to determine a matching probability distribution method based on the data type of the control variable, obtain a control variable probability distribution matching the probability distribution method, perform data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable, and output the control action.

[0010] Exemplarily, the method further includes: obtaining training samples, where the training samples include historical state variables; training the decision control model based on the training samples; where training the decision control model based on the training samples includes: making a decision based on the historical state variables to obtain a target control action; according to the gradient loss function, obtaining a loss function value based on the probability of the target control action and the reward of the target control action; obtaining parameter gradient information based on the loss function value; and updating the model parameters of the decision control model based on the parameter gradient information.

[0011] Exemplarily, the reward is accumulated and processed to obtain an accumulated discounted reward. The gradient loss function represents the sum of the products of the accumulated discounted reward and the logarithm of the probability.

[0012] Exemplarily, training the decision control model based on the training samples further includes: determining the probability of the target control action corresponding to the historical state variables based on the policy function; and determining the reward of the target control action corresponding to the historical state variables based on the revenue function.

[0013] Another embodiment of the present invention provides a cold station operation control device, which includes: an acquisition module for acquiring state variables, where the state variables include cold station operation status data and environmental status data; a first acquisition module for determining a matching probability distribution method based on the data type of the control variable and obtaining a control variable probability distribution matching the probability distribution method when making a decision based on the state variables; a second acquisition module for performing data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable; and a control module for controlling the operation of the cold station based on the control action.

[0014] An embodiment of the present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method according to any one of the above embodiments are implemented.

[0015] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method according to any one of the above embodiments are implemented.

[0016] In the above embodiment, the cold station operation control method includes: acquiring state variables, where the state variables include cold station operation status data and environmental status data; when making a decision based on the state variables, determining a matching probability distribution method based on the data type of the control variable and obtaining a control variable probability distribution matching the probability distribution method; performing data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable; and controlling the operation of the cold station based on the control action. By acquiring state variables including cold station operation status data and environmental status data, this method can comprehensively grasp the real-time operation status of the cold station and external environmental conditions, which helps to make more practical decisions. Determining a matching probability distribution method based on the data type of the control variable fully considers the characteristics of the control variable, making the decision-making process more scientific and reasonable, improving the accuracy and reliability of the decision. By performing data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable, the control action is more flexible and adaptable, can better handle various complex and changeable situations during the operation of the cold station, improve the operation efficiency and stability of the cold station, and achieve efficient utilization of energy.

[0017] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings

[0018] Figure 1 It is a flowchart of the cold station operation control method provided by the embodiment of the present invention;

[0019] Figure 2Schematic diagram of the decision control model structure provided by the embodiments of the present invention;

[0020] Figure 3 Inference decision flow chart of the decision control model provided by the embodiments of the present invention;

[0021] Figure 4 Training logic flow chart of the decision control model provided by the embodiments of the present invention;

[0022] Figure 5 Effect comparison diagram of the decision control model provided by the embodiments of the present invention;

[0023] Figure 6 Block diagram of the cold station operation control device provided by another embodiment of the present invention;

[0024] Figure 7 Block diagram of the electronic device provided by another embodiment of the present invention. Detailed implementation manners

[0025] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.

[0026] As an important part of building energy consumption, the energy consumption of the cold station accounts for a relatively large proportion of the total building energy consumption. According to statistics, building energy consumption accounts for about 40% of the global energy demand, and the energy consumption of central air-conditioning accounts for 40% - 50% of the total building energy consumption. Among them, the energy consumption of the cold station is the main part of the air-conditioning energy consumption, accounting for more than 30% of the total building electricity consumption. Therefore, optimizing the operation control of the cold station and improving energy efficiency are of great significance for reducing building energy consumption and operating costs. Optimizing the control of the cold station not only helps to save energy and reduce emissions, but also can improve the comfort and operation efficiency of the building, and promote sustainable development.

[0027] In the field of cold station optimization control, the proposed related technologies demonstrate the diversity and frontiers of the technological development in this field. A related technology proposes an energy-saving control method based on multi-agent of the cold station. This method constructs a building agent and an air-conditioning unit agent to globally optimize the control of the cold station, comprehensively considers the building load, environmental parameters and the operating state of the air-conditioning unit, and realizes the minimization of the overall energy consumption. Another related technology shows an optimization control method based on event-driven genetic algorithm. This method regularly obtains the historical data of the central air-conditioning refrigeration station equipment, constructs a power model, and performs real-time optimization calculations based on the event trigger points, effectively avoiding the problems of long global optimization calculation time and many parameters, and improving the optimization efficiency.

[0028] With the development of artificial intelligence technology, another related technology has proposed a data center energy-saving control method based on Transformer. This method constructs a PUE (Power Usage Effectiveness) prediction model, comprehensively analyzes the cold station, terminal equipment, and environmental indicators of the data center, realizes the online optimization of operating parameters, and significantly improves the energy-saving effect. Another related technology has proposed a solution based on an adaptive algorithm. This solution collects environmental data and cold station equipment data, uses multiple machine learning models for cold load prediction, and selects the optimal prediction model through an adaptive algorithm to achieve precise energy-saving control of the chiller unit.

[0029] In terms of model fusion and self-learning, another related technology has proposed an optimized control method for a central air-conditioning refrigeration station based on a self-learning fusion model. This method combines a mechanism model and a data model, and uses a self-learning architecture to continuously optimize the model accuracy and adaptability, realizing the overall energy efficiency optimization of the refrigeration station system. Another related technology provides an efficient cold station control system. This system uses sensors to monitor cold station parameters in real time and uses an iterative optimization module to optimize control strategies and parameter configurations. At the same time, the system also includes a temperature adjustment device, a data storage module, a fault detection and protection module, etc., realizing the intelligent control and efficient operation of the cold station.

[0030] In terms of the comprehensive optimization of the air-conditioning system, another related technology has proposed an optimized control method for the energy consumption of the air-conditioning air system and water system. This method collects sensor data, uses a long short-term memory neural network for load prediction, and synchronously controls the air system and water system in a coordinated manner to achieve the comprehensive optimization and high-efficiency energy saving of the system. Another related technology demonstrates an energy-saving control method for a cooling water system based on the overall optimal energy efficiency. This method constructs a historical database, monitors the refrigeration station environment and system operating parameters in real time, and uses historical data matching and parameter optimization steps to achieve the overall optimal energy efficiency of the cooling water system. At the same time, the system also includes an evaluation module to quantitatively evaluate the energy-saving effect.

[0031] Although related technologies have made significant progress in the optimized control of cold stations, there are still some deficiencies. For example, the complexity of some models is relatively high, which may lead to low computational efficiency; some systems have poor adaptability to inputs and are difficult to cope with changing working conditions; in addition, some systems lack online self-adaptability and cannot adjust optimization strategies online to cope with emergencies. Therefore, future cold station optimized control technologies need to further explore and innovate in aspects such as reducing model complexity, improving input adaptability, and enhancing online self-adaptability.

[0032] In view of this, an embodiment of the present application provides a cold station operation control method. This method deeply integrates the latest progress of deep learning and deep reinforcement learning with the profound professional knowledge in the field of heating, ventilation, and air conditioning (HVAC). Specifically targeting a series of complex challenges encountered in cold station energy efficiency optimization, based on deep neural networks and stochastic policy gradients, it significantly enhances the intelligence level of cold station operation, providing strong technical support for achieving efficient energy utilization and reducing operating costs.

[0033] Figure 1 It is a flowchart of the cold station operation control method provided by the embodiment of the present invention.

[0034] As Figure 1 shown, the cold station operation control method 100 includes steps S110 to S140.

[0035] Step S110, obtain state variables, where the state variables include cold station operation state data and environmental state data.

[0036] Exemplarily, the cold station operation state data may include, for example, the refrigeration load of the cold station, equipment parameters (start-stop status, frequency, load rate, cooling efficiency, etc. of chillers, chilled water pumps, cooling water pumps, cooling towers), and system parameters (temperature, pressure, flow rate, etc. on the cooling side and the chilled water side). The environmental state data may include, for example, weather data (temperature, humidity, air enthalpy value, rainfall, etc.).

[0037] Step S120, when making a decision based on the state variables, determine a matching probability distribution method based on the data type of the control variables, and obtain the control variable probability distribution matching the probability distribution method.

[0038] Exemplarily, the data types of the control variables may include, for example, continuous control variables, discrete control variables, switch control variables, etc. The probability distribution methods may include, for example, normal probability distribution method, categorical probability distribution method, Bernoulli probability distribution method, etc. Making a decision based on the state variables means inputting the state variables into a decision control model, and the decision control model includes a model based on a deep neural network.

[0039] Step S130, perform data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable.

[0040] Exemplarily, for the control variable probability distribution (normal probability distribution, categorical probability distribution, Bernoulli probability distribution, etc.), each possible value has a corresponding probability. Data sampling can be to randomly select some specific values as control actions according to the rules of the probability distribution. The control actions may include, for example, the set values of the freezing and cooling temperatures, the number of operating units of chillers, pumps, and cooling towers, the operating frequencies of pumps and cooling towers, the start-stop of equipment, etc.

[0041] Step S140: Control the operation of the chilled water station based on the control action.

[0042] Exemplarily, based on the obtained control actions (set values of freezing and cooling temperatures, number of operating units of chillers, pumps, and cooling towers, operating frequencies of pumps and cooling towers, start / stop of equipment), corresponding control is performed on the chilled water station equipment, and further the operation of the chilled water station system is controlled.

[0043] According to the embodiments of the present application, by acquiring state variables including the operating status data and environmental status data of the chilled water station, the real-time operating conditions of the chilled water station and the external environmental conditions can be comprehensively grasped, which helps to make more practical decisions. By determining the matching probability distribution method according to the data type of the control variables, the characteristics of the control variables are fully considered, making the decision-making process more scientific and reasonable, improving the accuracy and reliability of the decision. By performing data sampling on the probability distribution of the control variables to obtain the control actions corresponding to the control variables, the control actions are more flexible and adaptable, can better cope with various complex and changeable situations during the operation of the chilled water station, improve the operation efficiency and stability of the chilled water station, and achieve efficient utilization of energy.

[0044] The data type of the control variables includes at least one of continuous type, discrete type, and switch type. Determining the matching probability distribution method based on the data type of the control variables includes: when the data type of the control variable is continuous, determining the probability distribution method as the normal probability distribution method; when the data type of the control variable is discrete, determining the probability distribution method as the categorical probability distribution method; when the data type of the control variable is switch type, determining the probability distribution method as the Bernoulli probability distribution method.

[0045] Specifically, for continuous control variables, such as temperature setting, the normal distribution can be used to describe. The normal distribution is a continuous probability distribution, symmetric about the mean, with the peak of the curve corresponding to the mean, and the curve gradually decreases as the distance from the mean increases. This distribution is completely determined by two parameters: the mean (μ): representing the average level of the data, corresponding to the expected value of the temperature setting. Variance (σ 2 ) or standard deviation (σ): representing the degree of dispersion of the data, that is, the fluctuation range of the temperature setting value around the mean.

[0046] For discrete control variables, such as the operating status or working mode of equipment, the categorical distribution can be used to describe. The categorical distribution is suitable for representing random variables with a finite number of possible values, and each value has a corresponding probability. In the chilled water station control system, the number of equipment units and working modes can be regarded as finite categories, and each category corresponds to a certain probability distribution.

[0047] For binary switch variables, such as the start-stop state of a device, etc., the Bernoulli distribution can be used to describe them. The Bernoulli distribution is a discrete distribution, also known as the "0-1 distribution" or "two-point distribution". It has only two possible values: 0 and 1, corresponding to the off and on states of the device respectively. The parameter of the Bernoulli distribution is the probability of success p, that is, the probability of the device being turned on. In the cold station control system, this probability can be adjusted according to the operation strategy and requirements of the device.

[0048] The probability distribution method of this application is illustrated by the normal probability distribution method, the categorical probability distribution method, and the Bernoulli probability distribution method, and is not a specific limitation of this application. The description of the probability distribution of the control variable will not be limited to the above variables, as long as it can clearly describe the probability of the control action to be taken.

[0049] In the above embodiments, it can adapt to various control variables, covering common data types such as continuous, discrete, and switch types, meeting the data analysis requirements in a variety of actual scenarios, and having strong generality and applicability.

[0050] Perform data sampling on the probability distribution of the control variable to obtain the control action corresponding to the control variable, including: using the preset safety rule data as the sampling constraint, performing data sampling on the probability distribution of the control variable to obtain the control action corresponding to the control variable.

[0051] Specifically, based on the obtained probability distribution of the control variable, random variable sampling can be performed to obtain the output of the optimized control variable. For example, performing one sampling on a normal distribution with a mean of 5.6 °C and a standard deviation of 0.2 °C can obtain an optimized chilled water set temperature of 5.7 degrees, etc., and it may also be other temperatures, but it is very likely to be close to 5.6 °C; for the number of operating chilled water pumps, sampling is performed, and it is very likely to obtain that 3 chilled water pumps need to be turned on in the optimized state; for the action of the No. 1 cooling water pump, the result is very likely to be that the No. 1 cooling water pump is turned on, but there is also about a 20% probability that the output is not turned on. An example of the output is that in the optimized state, the output should be: the chilled water outlet temperature should be set to 5.6 °C, 3 chilled water pumps should be turned on, and the No. 1 cooling water pump should be turned on.

[0052] Filter the output control actions. For example, if the predefined safety rule stipulates that when the outdoor temperature is 30 °C or above, 4 chilled water pumps must be turned on, then the above output does not meet the predefined safety rule, and the above output distribution needs to be randomly sampled again until it meets the safety rule.

[0053] In the above embodiments, the safety rule data reflects the reasonable boundaries and limiting conditions of system operation. Sampling is carried out under such constraints, so that the sampled control actions are based on the premise of safety and reasonableness, which helps to make more scientific and reasonable decisions on the basis of meeting safety requirements and avoid unreasonable control actions caused by blind sampling.

[0054] The cold station operation status data and environmental status data include current data, historical data, and future data. The current data includes at least one of refrigeration load, equipment parameters, system parameters, and weather. The historical data includes at least one of refrigeration load and weather. The future data includes at least one of refrigeration load and weather.

[0055] For example, the cold station operation status data includes at least one of refrigeration load, equipment parameters, and system parameters. The environmental status data includes weather. The refrigeration load of the cold station includes the current value, historical value, and future value of the current refrigeration load of the cold station. Among them, the current value is the cold station load measured in real time, usually expressed as the cooling capacity per unit time (such as tons of refrigeration per hour or kilowatts), which can be directly measured or calculated by sensors such as flow meters and thermometers. The short-term historical value is the load data in the past period (such as a few minutes or a few hours), which is used to analyze the change trend of the load and may be used to predict short-term load fluctuations. The predicted value for a future period is the future load prediction obtained based on historical data and weather forecast information using prediction models (such as time series analysis and machine learning algorithms), which is crucial for energy scheduling and optimizing operation strategies.

[0056] Equipment parameters: Include the operation parameters of the main equipment, including the start-stop status of the chiller, indicating whether the chiller is running; frequency, the running frequency of the chiller compressor, which affects the cooling capacity and energy consumption; load rate, the ratio of the actual load of the chiller to its rated load; power, the electric energy consumed by the chiller. The start-stop status of the chilled water pump and cooling water pump, indicating whether the pump is running; frequency, the rotational speed of the pump, which affects the water flow and energy consumption. The start-stop status of the cooling tower, indicating whether the cooling tower is running; fan frequency, the rotational speed of the cooling tower fan, which affects the heat dissipation effect and energy consumption, and cooling efficiency, the ability of the cooling tower to cool hot water to the set temperature.

[0057] System parameters include the cooling side temperature, the temperature of the cooling water, which affects the cooling effect of the chiller; pressure, the pressure of the cooling system, which needs to be maintained within a safe range; flow rate, the flow rate of the cooling water, to ensure sufficient cooling capacity. The chilled water side temperature, the temperature of the chilled water, determines the output temperature of the cooling system; pressure, the pressure of the chilled water system, also needs to be maintained within a safe range; flow rate, the flow rate of the chilled water, which affects the overall efficiency of the cooling system.

[0058] Weather includes key state parameters of the current value, history, and future value of the outdoor temperature at the cold station. Among them, the current value is the outdoor temperature, humidity, air enthalpy value, etc. measured in real time, and these parameters directly affect the cooling demand of the cold station. The weather forecast values are future weather prediction data from weather stations, including temperature, humidity, rainfall, etc., which are used to influence and adjust the cold station operation strategy to cope with upcoming weather changes. The historical values for a period of time are weather data (such as the past few hours) in the past, which are used to input the impact of weather patterns on the cold station load to improve the accuracy of optimization decisions.

[0059] The state variables (cold station operation status data and environmental status data) of this application are described as above, which are not specific limitations of this application. It may not be limited to the above several types. As long as the variables that can provide information for optimization control decisions can be used as input variables of the system. The above variables can be read from sensors, building automation systems, cold station group control systems, etc.

[0060] In the above embodiments, integrating current, historical, and future data provides a rich information basis for decision-making of the cold station. It not only considers the current operation status of the equipment but also refers to the life cycle and failure conditions of the equipment in historical data, and combines future load demands and development trends, enabling more scientific and reasonable decisions to be made.

[0061] The cold station operation control method is applied to a decision control model. The decision control model includes an input layer, an intermediate layer, and an output layer, where: the input layer is used to input state variables; the intermediate layer is used to make decisions based on state variables; the output layer is used to determine a matching probability distribution method based on the data type of control variables, obtain the control variable probability distribution matching the probability distribution method, perform data sampling on the control variable probability distribution, obtain the control actions corresponding to the control variables, and output the control actions.

[0062] Exemplarily, the decision control model contains a neural network module that can input the state variables of the cold station and output optimized control actions.

[0063] The input layer of the neural network: After preprocessing the above state variables of the cold station, they are input. Typical preprocessing methods include normalization of continuous variables and encoding of discrete variables, such as applying one-hot encoding or label encoding.

[0064] The intermediate layer of the neural network: Common neural network structures can be adopted. A typical implementation can be defined using several fully connected layers.

[0065] Output layer of the neural network: A series of probability distribution functions are used to describe the probability of taking a certain action. Different probability distributions are used to describe continuous control variables, discrete control variables, and switch control variables, etc. For continuous control variables, a normal probability distribution is used to describe; for discrete control variables, a CategoryDistribution is used to describe; for switch control variables, a Bernoulli probability distribution is used to describe. According to the probability distribution parameters, random sampling can be performed to obtain the output (control action) of the control variable.

[0066] The output of the control variable of the decision control model includes the freezing temperature set value: which determines the output temperature of the freezing system and affects the cooling effect and temperature control at the user end. The decision control model needs to dynamically adjust the set value of the freezing temperature according to factors such as load demand, outdoor weather conditions, and system efficiency to achieve the best cooling effect and energy consumption balance; the cooling temperature set value: which affects the cooling efficiency and energy consumption of the chiller. The model needs to optimize the set value of the cooling temperature according to the performance characteristics of the chiller, the heat dissipation capacity of the cooling tower, and the overall heat balance requirements of the system; the number of operating chillers: according to the magnitude and change trend of the cooling load, the model needs to determine how many chillers to start to meet the cooling demand, while considering the energy efficiency ratio and operating cost of the chillers; the number of operating water pumps: the water pumps are responsible for circulating the chilled water and cooling water, and need to meet the water flow demand and the hydraulic balance requirements of the system; the number of operating cooling towers: the cooling towers are used for heat dissipation, and the model determines how many cooling towers to start to achieve the best heat dissipation effect; the operating frequency of the water pumps: by adjusting the rotation speed (i.e., the operating frequency) of the water pumps, the water flow can be controlled, thereby affecting the efficiency and energy consumption of the cooling system; the operating frequency of the cooling tower fans: the rotation speed of the cooling tower fans directly affects the heat dissipation effect and energy consumption, and the decision control model needs to adjust the operating frequency of the cooling tower fans to achieve the best heat dissipation effect and energy consumption balance; the start and stop of equipment: such as the start and stop of a specific chiller, water pump, etc.

[0067] The above output control variables are used as examples in this application for illustration, and are not intended as specific limitations of this application. The output variables of the decision control model are not limited to the above output variables, and all variables that can affect the cold station state and can be output can be defined as the output control actions.

[0068] Figure 2 It is a schematic structural diagram of the decision control model provided by the embodiment of the present invention.

[0069] As Figure 2 shown, the input layer of the decision control model inputs state variables such as outdoor temperature, outdoor humidity, and cold station load. After the reasoning and decision-making of the intermediate layer, the output layer outputs the distribution μ of the chilled water temperature set value, the standard deviation σ of the chilled water temperature set value, the probability distribution p1 of starting the No. 1 water pump, and the probability distribution q of starting n water pumps netc.

[0070] Figure 3 It is the inference decision flow chart of the decision control model provided by the embodiment of the present invention.

[0071] Such as Figure 3 As shown, the inference decision process of the decision control model includes step S310 to step S350.

[0072] Step S310, organize the operating state variables of the cold station and input them into the decision neural network.

[0073] Step S320, the neural network outputs to obtain a series of probability distributions of control variables.

[0074] Step S330, sample the probability distribution to obtain a series of control variables.

[0075] Step S340, determine whether the combination of control variables conforms to the constraint rules. If it conforms, go to step S350; otherwise, repeat step S330 until the constraint rules are met.

[0076] Step S350, output the control variables, end this inference, and enter the next cycle.

[0077] Specifically, when the neural network performs control inference, the input state variables are mapped to the output probability distribution, and the output is calculated using the input of the neural network. For example, the outdoor temperature is 30°C and the outdoor humidity is 60%. After combining other state variables and performing preprocessing such as normalization, it is input into the neural network for decision-making to obtain the parameters of the control variable distribution. For example, the chilled water outlet temperature distribution in the optimized state is a random variable that conforms to a normal distribution with a mean of 5.6°C and a standard deviation of 0.2°C; the probability of turning on 4 chilled water pumps and 3 pumps in the optimized state is 70%, and the probabilities of turning on 1, 2, and 4 pumps are 5%, 15%, and 10% respectively, which is a random variable that conforms to a categorical distribution; the probability of turning on the No. 1 cooling water pump in the optimized state is 20%, etc., and the probability of not turning on the No. 1 cooling water pump is 80%, etc., which is a random variable that conforms to a Bernoulli distribution. Sampling according to the above probability distributions can obtain the output of the optimized control variables, and filter the output control actions. For example, if it is predefined in the safety rules that 4 chilled water pumps must be turned on when the outdoor temperature is 30°C or above, then the above output does not conform to the predefined safety rules, and the above output distribution needs to be randomly sampled again until it conforms to the safety rules.

[0078] The method further includes: obtaining training samples, where the training samples include historical state variables; training a decision control model based on the training samples; where training the decision control model based on the training samples includes: making a decision based on the historical state variables to obtain a target control action; according to the gradient loss function, obtaining a loss function value based on the probability of the target control action and the reward of the target control action; obtaining parameter gradient information based on the loss function value; and updating the model parameters of the decision control model based on the parameter gradient information.

[0079] Exemplarily, the training samples include historical state variables, which is a quantity that changes over time and is recorded as s t , s t is a set of outdoor weather, cold station, and equipment status at a certain moment t. The target control actions include temperature setting, number of units to be started, set frequency, start / stop of a certain device, etc., which is a set of quantities that change over time, and the control action at moment t is recorded as a t .

[0080] When training a neural network, the definition of the gradient loss function is shown in formula (1):

[0081] loss(t)=-G t ·logπ θ (a t |s t ) (1)

[0082] In formula (1), π θ (a t |s t ) is the probability of taking action a t in state s t . This gradient loss function encourages the policy network to assign higher probabilities to state-action pairs with high rewards (gains), thereby reducing the loss value. G t is the reward.

[0083] Training the decision control model based on the training samples further includes: determining the probability of the target control action corresponding to the historical state variable based on the policy function; and determining the reward of the target control action corresponding to the historical state variable based on the reward function.

[0084] For example, the policy function outputs the probability distribution of the control action through a neural network, probability sampling, and rule filtering, denoted as π θ , where θ is a set of model parameters, which includes not only the weight parameters of each layer of the neural network but also the distribution parameters of the random variables in the output layer, such as μ and σ of the normal distribution, and the probabilities p of each option in the categorical distribution and Bernoulli distribution. π θ (a t |s t)The probability of taking action a t in state s t .

[0085] Reward function: At a certain moment t, in state s t , using the policy function π θ , the obtained reward is denoted as G t (i.e., the reward). The definition of G t can include but is not limited to the following forms: the COP (Coefficient of Performance) of the system at this moment and in a future period, the negative of the total power of the system at this moment and in future moments, and the reward function considering the sum of comfort and energy conservation, etc. The definition of the function only needs to conform to the overall indicators of energy consumption, energy efficiency, and comfort, and the larger the reward value, the better.

[0086] The reward is accumulated and processed to obtain the cumulative discounted reward, and the gradient loss function represents the sum of the product of the cumulative discounted reward and the logarithm of the probability.

[0087] The cumulative discounted reward (usually denoted as G t ) is the weighted sum of all future rewards starting from the current time step t, and its mathematical form is shown in formula (2):

[0088]

[0089] r t+k is the immediate reward obtained at time step t + k, γ is the discount factor (0 ≤ γ ≤ 1), which is used to adjust the importance of future rewards. The smaller γ is, the more attention is paid to short-term rewards; the larger γ is, the more attention is paid to long-term rewards. The cumulative discounted reward Gt reflects the total possible future rewards starting from the current state s t .

[0090] In practical applications, the gradient loss function is usually the sum of the product of the cumulative discounted reward and the logarithm of the action probability for all time steps T. For example, for T time steps in a trajectory (episode), the gradient loss function can be expressed as formula (3):

[0091]

[0092] The objective of the gradient loss function in formula (3) is to maximize the cumulative discounted reward of the entire trajectory.

[0093] Figure 4 This is the training logic flow chart of the decision control model provided by the embodiment of the present invention.

[0094] As Figure 4 shown, the training logic flow of the decision control model includes steps S410 to S440.

[0095] Step S410, Status Input, input the operating status of the cold station system and the environment.

[0096] Step S420, Optimization Decision Model, a model based on a deep neural network.

[0097] Step S430, Control Output, output the control actions of the cold station system.

[0098] Step S440, Model Training and Update, based on historical data, according to the defined optimization objective, use the stochastic gradient descent algorithm to update the model parameters.

[0099] Using the above training logic, train the policy function π θ The training process specifically includes the following steps:

[0100] (a) Initialize the policy network: Initialize the parameters of the policy network with random weights.

[0101] (b) Data collection and interaction: In each training iteration, select an action according to the current policy function, execute the action, and observe the system's feedback (state and reward).

[0102] (c) Store the trajectory: Store the state, action, and reward information in each episode to form a trajectory.

[0103] (d) Calculate the total return (cumulative discounted reward): Use the Monte Carlo method to calculate the total return for each state-action pair.

[0104] (e) Calculate the loss and gradient: Calculate the loss for each state-action pair according to the loss function, and use the backpropagation algorithm to calculate the gradient of the loss with respect to the policy network parameters.

[0105] (f) Update the policy network: Use an optimizer (such as Adam) to update the parameters of the policy network according to the gradient.

[0106] (g) Repeat the iteration: Repeat the above steps until the policy network converges or reaches the preset number of training epochs.

[0107] Through the above training process, an optimized control policy function π θ (a t |s t ) can be obtained.

[0108] The decision control model proposed in this application has a simple and efficient structure. Different from traditional complex models based on mechanisms, it only needs to abstract key state variables and control action variables according to the actual operating characteristics of the cold station, and construct a simple and practical model framework. This feature reduces the complexity of model construction and improves the generality and practicality of the model. It can adapt to various control variable inputs, showing extremely high flexibility, and can interface with and process various types of control variables. Whether it is the fine adjustment of continuously changing control set values, such as the temperature, flow rate, and pressure, or discrete control decisions, such as the start-stop control of equipment such as chillers, cold stations, and cooling towers and the dynamic optimization of the number of operating units, it can accurately respond and achieve optimal control. The online update ability, in order to cope with the dynamic changes in the operating environment of the cold station and the challenges of different working conditions, incorporates the stochastic policy gradient algorithm, enabling the system to utilize the historical data accumulated during the operation in real time to optimize and update the deep neural network model online. This self-learning and self-adaptive ability ensures that the model always remains in the best state, effectively improving the persistence and accuracy of the energy efficiency optimization of the cold station.

[0109] For ease of understanding, this application gives a specific embodiment, which uses simplified input and output and is not a limitation to this application.

[0110] Inputs of the decision control model: The environmental state of the cold station, specifically including the wet bulb temperature (Twb) and the cooling load (CL). These two state parameters are obtained in real time through sensors and used as the inputs of the neural network.

[0111] Outputs of the decision control model: The control actions of the cold station, namely the set temperature (Tchws_set) of the chiller and the operating frequency (t_f) of the cooling tower. These two control parameters will be directly applied to the control system of the cold station to adjust the operating state of the cold station.

[0112] A simple fully connected neural network is adopted, which includes two layers: Input layer: Receives OBS_SIZE (set to 2, corresponding to Twb and CL) input nodes and outputs 128 nodes. Middle layer: Three fully connected layers with 128 nodes each, serving as the middle layer. Output layer: Receives 128 input nodes and outputs ACT_SIZE * 2 (set to 4, corresponding to the mean and standard deviation of two actions, that is, the mean and standard deviation of Tchws_set, and the mean and standard deviation of t_f) nodes. The parameters of the output layer are used to represent the mean (μ) and standard deviation (σ) of each action, and then form a normal distribution for action sampling. Transfer function: (Activation function) is the ReLU function, which is used to increase the nonlinear ability of the network.

[0113] Revenue function: It is calculated based on the operating efficiency of the cold station (such as COP (Coefficient of Performance)) and energy consumption. Higher revenue will be obtained with efficient operation and lower energy consumption. The loss function is the policy gradient loss, which is the sum of the products of negative cumulative discounted rewards and the logarithmic probabilities of actions. The policy parameters are optimized by maximizing the cumulative discounted rewards.

[0114] The training process is as follows: The deep learning framework used is PyTorch; the optimizer selected is the Adam optimizer with a learning rate set to 0.001. In each training episode, the model selects an action based on the current state, executes the action, and obtains a new state and reward. The rewards are accumulated and processed into cumulative discounted rewards. Then, the loss is calculated and backpropagated to update the network parameters. The training process continues until a preset stopping condition (such as the number of training episodes or convergence criteria) is reached.

[0115] Through policy-gradient-based training, the policy model can learn to adjust the control parameters (Tchws_set and t_f) in real time according to the historical operation data of the cold station, so as to optimize the operating efficiency of the cold station and reduce energy consumption. In practical applications, the model can make decisions based on the real-time environmental state to achieve intelligent control of the cold station.

[0116] Figure 5 This is the comparison chart of the decision control model provided by the embodiment of the present invention.

[0117] As Figure 5 shown, the system can automatically output the frequency of the cooling tower operation according to the outdoor wet-bulb temperature, achieving the effect of improving the system COP. The yellow curve is the optimized effect, and the green is the baseline effect. Compared with the baseline situation during the test period, the energy consumption can be reduced by 3% - 10%.

[0118] Through continuous training and optimization, the decision-making ability of the model will gradually improve, and the operating efficiency and energy-saving effect of the cold station will also be significantly improved. At the same time, through data recording and analysis, the optimization effect and practical application value of the model can be further verified.

[0119] This application not only provides a brand-new and efficient solution for optimizing the energy efficiency of cold stations, but also significantly enhances the intelligent level of cold station operation through its model simplicity, wide adaptability of control variable input, and online update ability, providing strong technical support for achieving efficient energy utilization, reducing operating costs, and promoting sustainable development goals. The model structure is simple and clear, without relying on complex mechanism models, and can be updated online in real time according to historical operation data, continuously adapting to the actual operation conditions of cold stations. Through this optimization process, the decision control model can achieve instant regulation and optimization of cold stations, effectively improving the operation efficiency of cold stations, thereby achieving significant energy conservation and consumption reduction effects.

[0120] Figure 6 It is a block diagram of a cold station operation control device provided in another embodiment of the present invention.

[0121] An embodiment of the present invention provides a cold station operation control device 600. Please refer to Figure 6 , the cold station operation control device 600 includes: an acquisition module 610, a first acquisition module 620, a second acquisition module 630, and a control module 640.

[0122] Exemplarily, the acquisition module 610 is configured to acquire state variables, where the state variables include cold station operation state data and environmental state data.

[0123] Exemplarily, the first acquisition module 620 is configured to determine a matching probability distribution method based on the data type of the control variable when making a decision based on the state variables, and obtain a control variable probability distribution that matches the probability distribution method.

[0124] Exemplarily, the second acquisition module 630 is configured to perform data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable.

[0125] Exemplarily, the control module 640 is configured to control the operation of the cold station based on the control action.

[0126] Exemplarily, the data type of the control variable includes at least one of continuous type, discrete type, and switch type. The first acquisition module 620 is further configured to determine that the probability distribution method is a normal probability distribution method when the data type of the control variable is continuous; determine that the probability distribution method is a categorical probability distribution method when the data type of the control variable is discrete; and determine that the probability distribution method is a Bernoulli probability distribution method when the data type of the control variable is switch type.

[0127] Exemplarily, the second acquisition module 630 is further configured to perform data sampling on the control variable probability distribution with preset safety rule data as the sampling constraint to obtain a control action corresponding to the control variable.

[0128] Exemplarily, the cold station operation status data and environmental status data include current data, historical data, and future data. The current data includes at least one of refrigeration load, equipment parameters, system parameters, and weather. The historical data includes at least one of refrigeration load and weather. The future data includes at least one of refrigeration load and weather.

[0129] Exemplarily, the cold station operation control device is applied to a decision control model. The decision control model includes an input layer, an intermediate layer, and an output layer, where: the input layer is used to input state variables; the intermediate layer is used to make decisions based on the state variables; the output layer is used to determine a matching probability distribution method based on the data type of the control variables, obtain the control variable probability distribution matching the probability distribution method, perform data sampling on the control variable probability distribution, obtain the control actions corresponding to the control variables, and output the control actions.

[0130] Exemplarily, the cold station operation control device further includes: obtaining training samples, where the training samples include historical state variables; training the decision control model based on the training samples; where training the decision control model based on the training samples includes: making decisions based on the historical state variables to obtain target control actions; according to the gradient loss function, obtaining the loss function value based on the probability of the target control actions and the rewards of the target control actions; obtaining the parameter gradient information based on the loss function value; and updating the model parameters of the decision control model based on the parameter gradient information.

[0131] Exemplarily, the rewards are accumulated and processed to obtain the cumulative discounted rewards. The gradient loss function represents the sum of the products of the cumulative discounted rewards and the logarithms of the probabilities.

[0132] Exemplarily, training the decision control model based on the training samples further includes: determining the probability of the target control actions corresponding to the historical state variables based on the policy function; and determining the rewards of the target control actions corresponding to the historical state variables based on the revenue function.

[0133] It can be understood that for the specific description of the cold station power prediction device 500, reference can be made to the description of the cold station power prediction method in the above text, which will not be elaborated here.

[0134] Figure 7 It is a block diagram of an electronic device provided in another embodiment of the present invention.

[0135] An embodiment of the present application provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above method is implemented.

[0136] As Figure 7 shown, for the sake of easy understanding, an embodiment of the present application shows a specific electronic device 700.

[0137] The electronic device 700 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0138] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes in accordance with a computer program stored in a read only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0139] A plurality of components in the electronic device 700 are connected to the I / O interface 705, and the plurality of components include: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0140] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods described above. For example, in some embodiments, any one or more of the above-described methods can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of any one or more of the above-described methods can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute any one or more of the above-described methods in any other suitable manner (e.g., by means of firmware).

[0141] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method according to any one of the above embodiments are implemented.

[0142] Note that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of the present invention, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0143] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.

[0144] In the description of the present invention, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In the present invention, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0145] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention.

[0146] In addition, the terms "first", "second", etc. used in the embodiments of the present invention are only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated in this embodiment. Thus, the features defined by the terms "first", "second", etc. in the embodiments of the present invention may explicitly or implicitly indicate that at least one such feature is included in this embodiment. In the description of the present invention, the meaning of the word "plurality" is at least two or more, such as two, three, four, etc., unless otherwise explicitly and specifically defined in the embodiment.

[0147] In the present invention, unless otherwise explicitly specified or limited in the embodiment, the terms "mounted", "connected", "connected" and "fixed" etc. appearing in the embodiment should be understood in a broad sense. For example, the connection can be a fixed connection, a detachable connection, or integrated. It can be understood that it can also be a mechanical connection, an electrical connection, etc.; of course, it can also be directly connected, or indirectly connected through an intermediate medium, or it can be the communication inside two elements, or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific implementation situation.

[0148] In the present invention, unless otherwise explicitly specified and limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "below" and "beneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.

[0149] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A cold station operation control method, characterized in that The method includes: Obtaining state variables, where the state variables include cold station operation status data and environmental status data; When making a decision based on the state variables, determining a matching probability distribution method based on the data type of the control variable, and obtaining a control variable probability distribution matching the probability distribution method; Performing data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable; Based on the control action, controlling the operation of the cold station.

2. The method according to claim 1, wherein The data type of the control variable includes at least one of continuous type, discrete type, and switch type. The determining of the matching probability distribution method based on the data type of the control variable includes: When the data type of the control variable is continuous, determining the probability distribution method as a normal probability distribution method; When the data type of the control variable is discrete, determining the probability distribution method as a categorical probability distribution method; When the data type of the control variable is switch type, determining the probability distribution method as a Bernoulli probability distribution method.

3. The method according to claim 1, wherein The performing data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable includes: Using preset safety rule data as a sampling constraint to perform data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable.

4. The method according to claim 1, characterized in that, The cold station operation status data and environmental status data include current data, historical data, and future data. The current data includes at least one of refrigeration load, equipment parameters, system parameters, and weather. The historical data includes at least one of refrigeration load and weather. The future data includes at least one of refrigeration load and weather.

5. The method according to any one of claims 1-4, characterized in that, The method is applied to a decision control model, and is characterized in that the decision control model includes an input layer, an intermediate layer, and an output layer, where: The input layer is used to input the state variables; The intermediate layer is used to make a decision based on the state variables; The output layer is used to determine a matching probability distribution method based on the data type of the control variable, obtain a control variable probability distribution matching the probability distribution method, perform data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable, and output the control action.

6. The method according to claim 5, characterized in that, The method further includes: Obtaining training samples, where the training samples include historical state variables; Based on the training samples, training the decision control model; Wherein, the training the decision control model based on the training samples includes: Making a decision based on the historical state variables to obtain a target control action; According to the gradient loss function, obtaining a loss function value based on the probability of the target control action and the reward of the target control action; Obtaining parameter gradient information based on the loss function value; Updating the model parameters of the decision control model based on the parameter gradient information.

7. The method according to claim 6, characterized in that, The reward is accumulated and processed to obtain an accumulated discounted reward, and the gradient loss function represents the sum of the products of the accumulated discounted reward and the logarithm of the probability.

8. The method according to claim 6 or 7, characterized in that, The training the decision control model based on the training samples further includes: Based on the policy function, determining the probability of the target control action corresponding to the historical state variables; Based on the reward function, determine the reward for the target control action corresponding to the historical state variable.

9. A cold station operation control device, characterized in that, The device includes: An acquisition module, configured to acquire state variables, where the state variables include cold station operation state data and environmental state data; A first obtaining module, configured to, when making a decision based on the state variables, determine a matching probability distribution method based on the data type of the control variable, and obtain a control variable probability distribution matching the probability distribution method; A second obtaining module, configured to perform data sampling on the control variable probability distribution to obtain a control action corresponding to the control variable; A control module, configured to control the operation of the cold station based on the control action.

10. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1-8 are implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Cold station control method and system of central air-conditioning system

    CN113739368A

  • Model training method, method and device for controlling heating and ventilation system, equipment and medium

    CN117494556A

  • Heating ventilation air conditioner control method based on random probability weighted composite sampling strategy

    CN118669937A

  • Indoor temperature control method and device, electronic equipment and readable storage medium

    CN118998950A

  • Air conditioner control method, system and equipment based on deep reinforcement learning and medium

    CN119532918A