A Refrigeration Room Energy-Saving Optimization Control Method Based on Deep Reinforcement Learning
Through self-organizing thermal grid division and air conditioning leakage path monitoring based on deep reinforcement learning, combined with infrared cameras and micro wind pressure sensors, the air volume and angle are dynamically adjusted, solving the problem of mixing of hot and cold air flows in traditional computer rooms, achieving improved air conditioning utilization and temperature uniformity, and forming an intelligent closed-loop control.
Patent Information
- Application Number
- CN202510984129.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-07-17
AI Technical Summary
When vacant spaces in traditional computer rooms are not sealed in a timely manner, hot and cold air flows mix, resulting in reduced air conditioning utilization. In addition, there is a lack of intelligent and dynamic automatic blocking and compensation mechanisms, and the room relies on manual intervention, making it difficult to form closed-loop control.
Using a method based on deep reinforcement learning, through self-organizing thermal grid division, cold air leakage path monitoring and air volume configuration algorithm, it can automatically identify cold air leakage paths and perform precise compensation. Combined with infrared cameras and micro wind pressure sensors, it can dynamically adjust the air volume and angle to optimize the cooling zone.
It improves the utilization rate of cooling air, avoids energy waste, ensures temperature uniformity, reduces manual intervention, and forms an intelligent closed-loop control.
Smart Images

Figure CN120494453B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer room refrigeration, and in particular to a method for energy-saving optimization control of a refrigeration room based on deep reinforcement learning. Background Art
[0002] Traditional computer rooms often use a "cold aisle / hot aisle" separation method, with airflow organized by arranging the cabinets so that cold air enters from the front and hot air exits from the back. However, in actual operation, not every U-slot in the cabinet is occupied by servers or network equipment. Especially during the initial stages of equipment rollout or capacity expansion, a large number of vacant U-slots may appear. If these vacant spaces are not promptly sealed with baffles, the air flow path inside the cabinet will become seriously disrupted, resulting in the following symptoms:
[0003] The high-temperature airflow exhausted by the IT equipment at the rear of the cabinet will quickly flow back to the front cold channel along the empty spaces in the cabinet, mixing with the low-temperature cold air sent in by the air-conditioning unit, causing the temperature at the cold channel inlet to rise, and the cooling system needs to work extra to compensate for the temperature deviation; the "wind tunnel effect" caused by the lack of blank baffles causes a large amount of cold air to leak directly into the hot channel or the computer room floor through the empty spaces, significantly reducing the utilization rate of cold air; the blank baffles in the cabinet need to be manually installed or removed U by U, and maintenance and operation personnel often find it difficult to follow up in a timely manner. When equipment is frequently plugged in and out, blank baffles are easily forgotten or misplaced. In large-scale deployment scenarios, the problem of missing installation is common.
[0004] Although some manufacturers have launched prototypes of intelligent air baffles with temperature / flow sensing, most of them only have single-point alarm functions and cannot be linked with the air conditioning system or air valve execution strategy. They cannot automatically adjust the sealing degree or air volume compensation according to the real-time temperature gradient. There are still a lot of manual intervention links, making it difficult to form closed-loop control.
[0005] In summary, the current solutions to the backflow problem of empty cabinet spaces mostly rely on manual installation, fixed structures and passive monitoring, and lack intelligent and dynamic automatic blocking and compensation mechanisms.
[0006] To this end, the present invention provides a refrigeration room energy-saving optimization control method based on deep reinforcement learning. Summary of the Invention
[0007] The purpose of the present invention is to provide a refrigeration room energy-saving optimization control method based on deep reinforcement learning to solve the existing problems raised in the above background technology.
[0008] To achieve the above objectives, the present invention provides the following technical solution: a refrigeration room energy-saving optimization control method based on deep reinforcement learning, comprising the following steps:
[0009] S1. Divide the computer room into self-organizing thermal grids based on the cabinet spacing to obtain initial cooling zones and calculate the load status within the zones in real time.
[0010] S2: Check whether there is a device in each U position of each cabinet. If there is an empty position, run S3; if there is a device, run S4 directly;
[0011] S3. Set up a cold air leakage path monitoring strategy to monitor the heat distribution of the vacant U-space at the front end of each cooling zone cabinet in real time, automatically identify the cold air leakage path, and obtain the compensation air volume for each cooling zone;
[0012] S4. Construct an air volume configuration algorithm, take the load status and compensation air volume in the cooling zone as inputs of the air volume configuration algorithm, and combine it with the cable density in the cooling zone to obtain the air volume and angle of each air outlet.
[0013] A further improvement of the present invention is that the self-organizing thermal grid division process includes:
[0014] S11. Divide the computer room plane into M×N two-dimensional grids of uniform size;
[0015] S12: For each cabinet j, release evenly in the r grid units around it. Each grid unit is continuously moved on the grid by a virtual thermal particle. Each virtual thermal particle moves on the grid according to the temperature potential energy iteration rule. The grid point with the highest energy becomes the thermal core of the cabinet. The thermal core attracts the surrounding cabinets, and adjacent thermal cores naturally converge to form a partition boundary. For each thermal core partition, the cabinets it covers are collected to complete the initial cooling partition.
[0016] S13: Extract the number of cabinets corresponding to each initial cooling zone, and calculate the load status in the zone according to the number of U positions in the cabinets.
[0017] A further improvement of the present invention is that the temperature potential energy iteration rule specifically includes:
[0018] S121, in the current grid cell Calculate the temperature potential energy of each adjacent grid cell in the 8 adjacent grid cells around it ,in Represents a grid cell To the power consumption value of the nearest cabinet, represents the Euclidean distance between grid cells, and represents weight;
[0019] S122, the virtual thermal particle selects the grid cell with the maximum temperature potential energy from the neighborhood and moves to it. When the maximum temperature potential energy is less than 0 or the virtual thermal particle reaches the set maximum number of steps, the current grid cell is determined to be a precipitation point and the iteration stops.
[0020] S123. Count the number of virtual particle deposits in all grid cells , as the core energy level of the corresponding grid unit, the grid with the highest core energy level of each cabinet is defined as the thermal core of the cabinet;
[0021] S124. For each hot core , calculate the potential energy mountain spreading outward from the core ,in, Represents a grid cell to the hot core The number of grid steps, represents the diffusion damping coefficient, for the grid element Belongs to the partition represented by the hot core that makes its potential energy mountain the largest.
[0022] A further improvement of the present invention is that the cold air leakage path monitoring strategy specifically comprises the following steps:
[0023] S31. Place an infrared camera on the top or in front of the cabinet to output a thermal image sequence. , Represents the U coordinate within the partition, and denoises the heat map sequence to obtain a smooth temperature field ;
[0024] S32, synchronously collect the static differential pressure of each U position With dynamic flow rate Signal, the temperature gradient and pressure gradient are used as leakage threat input, and the leakage threat of each U position is defined by setting the target inlet temperature ;
[0025] S33: The top 5% of leakage threat U positions are used as the starting point set, and the return air outlet or cold channel outlet at the corresponding partition boundary is used as the end point;
[0026] S34. Setting leakage cost function ,in represents the distance between U-bits, and l represents the number of U-bits in the partition;
[0027] S35. Output the path corresponding to the minimum leakage cost function of each cooling partition , get the compensation air volume for each cooling zone.
[0028] The present invention is further improved in that the cooling partition compensation air volume is On each node The threat of leakage is The proportion of the total leakage threats of all nodes on the path is used as the path weight of the node , and then get the compensation air volume in the cooling zone , Indicates the target inlet air temperature, Indicates the system calibration coefficient, L indicates the path The number of all U positions on the .
[0029] A further improvement of the present invention is that the specific steps of the air volume configuration algorithm include:
[0030] S41, randomly generate particle air volume and angle, and use a vector to represent all configurations of each particle;
[0031] S42, according to the air volume and angle configuration of each particle, taking the load state and compensation air volume in the cooling zone as input, creating a temperature balance model and obtaining a fitness function fit;
[0032] S43. After the particle is updated, the wind volume and direction configuration corresponding to the current solution is accepted according to the particle acceptance mechanism;
[0033] S44. When the fitness function fit reaches the set maximum number of iterations or reaches the set convergence threshold, the iteration is stopped and the current air volume and angle configuration is output. Otherwise, return to step S42 to continue iteration.
[0034] A further improvement of the present invention is that the temperature balance model in the fitness function is expressed as ,in, Indicates the Temperature balance of each cooling zone, Indicates the reference supply air temperature, Indicates the cable density of the partition, Indicates the air volume cooling item, Indicates the load status, Indicates cooling zones Total air supply volume, cooling zones The total air volume, c represents the cooling coefficient per unit air volume, Indicates the set proportional coefficient;
[0035] The fitness function is obtained by combining the temperature balance model, the compensation air volume penalty and the air volume demand compensation. ,in, Indicates the temperature balance of all cooling zones, represents the first set of penalty terms, represents the second set of penalty terms, Indicates the minimum air volume requirement within the set cooling zone.
[0036] A further improvement of the present invention is that the particle acceptance mechanism includes:
[0037] If the new fitness is less than the particle fitness before the update, the newly generated solution is accepted, that is, the set of wind volumes corresponding to the particle after this iteration and wind direction Configuration, taking the new solution as the current solution for the next step;
[0038] When the new fitness is greater than or equal to the fitness of the particle before the update, the probability Accept the air volume and angle configuration corresponding to the current solution.
[0039] The present invention is further improved in that the load state in step S13 is obtained by calculating the product of the total number of U bits in the cooling zone and the average power consumption. Load status of cooling zones .
[0040] A further improvement of the present invention is that step S2 includes installing a miniature wind pressure sensor and a wind pressure fine-tuning device on the rear side or bottom of each U position of each cabinet, and independently deploying a group of sensing units at each U position to synchronously collect static differential pressure and dynamic flow velocity signals. A second-order low-pass filter is performed inside the sensor to remove high-frequency mechanical noise and then perform sliding averaging. The actuator is triggered to close only after KL consecutive judgments are that the position is empty; if it is determined that there is equipment at any time, the judgment count is reset.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] The present invention first solves the problems of "rough zoning" and "mismatched hot and cold channels" caused by traditional fixed thresholds or manual zoning by refining the computer room plane into an M×N grid with cabinet spacing. It then releases virtual thermal particles in each grid cell, applies temperature potential iterative precipitation, and attracts domain diffusion. This solves the problems of "rough zoning" and "mismatched hot and cold channels" caused by traditional fixed thresholds or manual zoning. It can adaptively form cooling zones based on the actual heat load of the cabinets and accurately capture the distribution of heat sources.
[0043] Through the leakage path monitoring strategy, the main leakage channel is searched and identified, and the zone-level compensation air volume is calculated based on the path node weight and temperature difference mapping. This solves the problems of unclear cold air return channel, delayed compensation and blind air increase, and achieves precise cooling compensation for the most serious leakage points, avoiding energy waste caused by large-scale or excessive air increase.
[0044] By incorporating the load of each partition, compensation air volume and cable density into the air volume configuration algorithm, and performing global search and local disturbance on the air volume and angle of the adjustable air outlet, the problems of temperature uniformity being difficult to take into account, search easily falling into local optimum, and optimization taking a long time in complex scenarios with multiple air outlets, adjustable wind direction, and equal emphasis on load and leakage compensation are solved. The air volume and wind direction configuration can not only meet the compensation requirements but also achieve the most uniform temperature, while eliminating premature algorithm maturation and improving convergence speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a flow chart of a refrigeration room energy-saving optimization control method based on deep reinforcement learning in the present invention;
[0046] Figure 2 This is a flow chart of a cold air leakage path monitoring strategy for a refrigeration room energy-saving optimization control method based on deep reinforcement learning in the present invention;
[0047] Figure 3 This is a flow chart of the air volume configuration algorithm of the energy-saving optimization control method for a refrigeration room based on deep reinforcement learning in the present invention. DETAILED DESCRIPTION
[0048] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0049] The term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0050] Example 1
[0051] Figure 1 A flowchart of a refrigeration room energy-saving optimization control method based on deep reinforcement learning disclosed in this embodiment is shown. The steps are as follows:
[0052] S1. Self-organizing thermal grid division is performed on the computer room based on the cabinet spacing to obtain the initial cooling zones, and the load status within the zones is calculated in real time. The self-organizing thermal grid division process includes:
[0053] S11. Divide the computer room plane into M×N two-dimensional grids of uniform size. The grid spacing is set to 1 / 4 of the minimum cabinet spacing, ensuring that at least a 4×4 grid can cover the area with the two most densely populated cabinets. The fine granularity helps capture subtle differences in cabinet spacing while also controlling the computational effort.
[0054] S12: For each cabinet j (i.e. each U position), release the power evenly in the r grid units around it. A virtual hot particle, , Indicates rounding up. represents the power consumption of cabinet j, Represents the baseline power consumption unit. At this point, the greater the power consumption, the more particles there are, and the more pronounced the core effect. Each grid unit is continuously moved on the grid by virtual thermal particles. Each virtual thermal particle moves on the grid according to the temperature potential energy iteration rule. The grid point with the highest energy becomes the thermal core of the cabinet. The thermal core attracts surrounding cabinets, and adjacent thermal cores naturally converge to form partition boundaries. For each thermal core partition, the cabinets it covers are collected to complete the initial cooling partition.
[0055] The temperature potential energy iteration rule specifically includes:
[0056] S121, in the current grid cell Calculate the temperature potential energy of each adjacent grid cell in the 8 adjacent grid cells around it ,in Represents a grid cell The power consumption value of the nearest cabinet can be calculated by the difference. represents the Euclidean distance between grid cells, and Represents weight; taking into account both high power consumption attraction (moving towards the heat source) and diffusion resistance (avoiding excessive aggregation);
[0057] S122, the virtual thermal particle selects the grid cell with the maximum temperature potential energy from the neighborhood and moves to it. When the maximum temperature potential energy is less than 0 or the virtual thermal particle reaches the set maximum number of steps, the current grid cell is determined to be a precipitation point and the iteration stops.
[0058] S123. Count the number of virtual particle deposits in all grid cells , as the core energy level of the corresponding grid unit, the grid with the highest core energy level of each cabinet is defined as the thermal core of the cabinet; the more particles, the more concentrated the thermal core, which can accurately mark the high-load location;
[0059] S124. For each hot core , calculate the potential energy mountain spreading outward from the core ,in, Represents a grid cell to the hot core The number of grid steps, represents the diffusion damping coefficient, for the grid element It belongs to the partition represented by the hot core that makes its potential energy mountain the largest; potential energy contour lines are formed naturally, and the boundaries are where the potential energy of each core is equal. The partition is smooth and adaptive.
[0060] S13, extract the number of cabinets corresponding to each initial cooling zone, and calculate the load state in the zone by the number of U positions in the cabinet; the load state is obtained by calculating the product of the total number of U positions in the cooling zone and the average power consumption. Load status of cooling zones .
[0061] The self-organizing thermal grid division of the present invention is used to characterize the overall layout density of the computer room, ensuring that the threshold value is random and room scale is adaptive.
[0062] S2: Check whether there is a device in each U position of each cabinet. If there is an empty position, run S3; if there is a device, run S4 directly;
[0063] A miniature wind pressure sensor and wind pressure fine-tuning device are installed on the rear side or bottom of each U-position of each cabinet. A set of sensing units are independently deployed at each U-position to synchronously collect static differential pressure and dynamic flow velocity signals. A second-order low-pass filter is performed inside the sensor to remove high-frequency mechanical noise and then perform sliding averaging (L sampling points). The actuator is triggered to close only after KL consecutive determinations that the position is empty. If it is determined that there is equipment at any time, the determination count is reset.
[0064] S3. Set up a cold air leakage path monitoring strategy to monitor the heat distribution of the vacant U-space at the front end of each cooling zone cabinet in real time, automatically identify the cold air leakage path, and obtain the compensation air volume for each cooling zone;
[0065] S4. Construct an air volume configuration algorithm, take the load status and compensation air volume in the cooling zone as inputs of the air volume configuration algorithm, and combine it with the cable density in the cooling zone to obtain the air volume and angle of each air outlet.
[0066] Example 2
[0067] Based on the technical solution of Example 1, the present invention proposes a specific implementation method of the cold air leakage path monitoring strategy in step S3. Figure 2 The flow chart of the cold air leakage path monitoring strategy disclosed in this embodiment is shown, and the steps are as follows:
[0068] S31, including placing an infrared camera on the top or in front of the cabinet to output a thermal image sequence , Represents the U coordinate within the partition, and denoises the heat map sequence to obtain a smooth temperature field ;
[0069] S32, synchronously collect the static differential pressure of each U position With dynamic flow rate Signal, the temperature gradient and pressure gradient are used as leakage threat input, and the leakage threat of each U position is defined by setting the target inlet temperature ; ,in Indicates the target inlet air temperature. When the point is considered as a potential leakage source;
[0070] S33: The top 5% of leakage threat U positions are used as the starting point set, and the return air outlet or cold channel outlet at the corresponding partition boundary is used as the end point;
[0071] S34. Setting leakage cost function ,in represents the distance between U-bits, and l represents the number of U-bits in the partition;
[0072] The distance penalty is a penalty for the distance "crossed" at each step of the path. Air flow tends to take the shortest path. The longer the distance, the greater the flow resistance, and the less likely the path will become a major leakage channel. This term can be added to avoid searching for overly circuitous routes.
[0073] Indicates that a penalty is imposed on situations where the leakage threat between adjacent nodes changes dramatically. If the leakage threat (i.e., temperature difference + pressure difference) between two points differs too much, it usually means that the segment is not a smooth and continuous leakage channel, but a sudden change boundary or noise. Adding this penalty can make the algorithm prefer a path along a gently rising potential energy gradient, improving the stability and physical rationality of recognition. When the channel space in the computer room is limited, it can be improved. , more emphasis on path length; if the temperature fluctuation is large, you can increase , which emphasizes the continuity of potential energy.
[0074] S35. Output the path corresponding to the minimum leakage cost function of each cooling partition , get the compensation air volume of each cooling zone; the compensation air volume of each cooling zone is calculated by dividing the On each node The threat of leakage is The proportion of the total leakage threats of all nodes on the path is used as the path weight of the node , and then get the compensation air volume in the cooling zone , Indicates the target inlet air temperature, Indicates the system calibration coefficient, L indicates the path The number of all U positions on the .
[0075] Using potential energy difference and flow pressure to jointly construct the cost can more accurately reflect the actual airflow deviation; the weight maps the temperature difference distribution of the path nodes to the compensation distribution, ensuring that the compensated air volume is focused on the most serious leakage point; the calibration coefficient is obtained through on-site testing of the heat-flow response curve to match the characteristics of the computer room air volume system.
[0076] Example 3
[0077] Based on the technical solutions of Example 1 and Example 2, the present invention proposes a specific implementation of the air volume configuration algorithm in step S4. Figure 3The flow chart of the air volume configuration algorithm disclosed in this embodiment is shown, and the steps are as follows:
[0078] S41, randomly generate particle volume and angle, each particle uses a vector to represent all configurations; for example, The dimension is the air volume of each adjustable air outlet ,back The dimension corresponds to the air outlet angle The particle dimensions change dynamically with the number of air outlets: if a new air outlet is added, the corresponding air volume and angle variables are appended to the end of the particle; if an air outlet is removed, the corresponding dimension is deleted and the original optimal value is inherited. In this way, the particle encoding can flexibly adapt to changes in the number and location of air outlets;
[0079] S42, according to the air volume and angle configuration of each particle, taking the load state and compensation air volume in the cooling zone as input, creating a temperature balance model and obtaining a fitness function fit;
[0080] The temperature balance model in the fitness function is expressed as ,in, Indicates the Temperature balance of each cooling zone, Indicates the reference supply air temperature, Indicates the cable density of the partition. Cables will hinder the flow of cold air and generate a little heat themselves, so they are introduced. Cable density is expressed as the ratio of the area of the current partition cable to the total area of the partition. Indicates the air volume cooling item, Indicates the load status, Indicates cooling zones The total air volume, c represents the cooling coefficient per unit air volume, representing 1m 3 / sHow much heat can the air conditioner take away in this zone? The corresponding cooling range. Indicates the set proportional coefficient;
[0081] Assuming that the heat source and cold source in the same partition can be approximately linearly superimposed, the heat load With cold load The difference determines the temperature deviation, It reflects the ideal supply air temperature. Any heat surplus will "raise" the partition temperature. High cable density not only physically hinders airflow, but also brings wiring heat, so cable density needs to be used for equivalent compensation.
[0082] The fitness function is obtained by combining the temperature balance model, the compensation air volume penalty and the air volume demand compensation through the penalty term in reinforcement learning:
[0083] ;
[0084] in, represents the standard deviation of temperature uniformity across all cooling zones, Indicates the first set of penalty items. If the actual air supply Less than , indicating that the area is short of air conditioning and needs compulsory compensation, so the penalty is increased by the square of the shortfall. It should be selected to be larger, which is greater than 0.5, to ensure that the algorithm prioritizes the compensation requirements. represents the second set of penalty terms, Indicates the minimum air volume requirement within the set cooling zone. If , also penalized in square form, comparable Smaller to ensure that the base load cooling requirement is considered after the compensation is met.
[0085] The fitness function of this embodiment ensures that a penalty is only applied when the air volume is insufficient. Otherwise, the fitness is not affected. The greater the deviation, the heavier the penalty, guiding the optimization to make up for the deficiency as quickly as possible. This is also consistent with the general least squares approach. Make "compensation air volume" a hard constraint, while "original load demand" is slightly less important. While pursuing temperature uniformity, it is also strictly guaranteed that the air volume in each zone is not lower than the key threshold to avoid "over-optimization" that may cause severe overheating in a certain area.
[0086] S43. After the particle is updated, the wind volume and direction configuration corresponding to the current solution is accepted according to the particle acceptance mechanism; the particle acceptance mechanism includes:
[0087] If the new fitness is less than the particle fitness before the update, it means that the new solution performs better on the objective function, and the newly generated solution is accepted, that is, the set of wind volumes corresponding to the particles after this iteration. and wind direction Configuration, taking the new solution as the current solution for the next step;
[0088] When the new fitness is greater than or equal to the fitness of the particle before the update, it is usually discarded to avoid performance degradation. However, in a complex, multi-peak search space, if you only "greedily" accept the better solution, it is easy to fall into a local optimum and cannot get out; balance exploration and utilization: in the early stage where a large amount of compensation air volume is required due to high temperature, a larger degree of random jump is allowed to expand the search area; in the later stage where the compensation air volume is reduced, strict screening is carried out and fine convergence is made to the global optimum or approximate optimum; therefore, according to the probability Accept the air volume and angle configuration corresponding to the current solution; this probability is related to the compensation air volume. The larger the compensation air volume, the greater the acceptance probability. As the iteration proceeds, the temperature gradually decreases, the required compensation air volume also gradually decreases, and the probability of accepting the "inferior solution" also gradually decreases. When the algorithm converges, it tends to strictly accept only the better solution.
[0089] S44. When the fitness function fit reaches the set maximum number of iterations or reaches the set convergence threshold, the iteration is stopped and the current air volume and angle configuration is output. Otherwise, return to step S42 to continue iteration.
[0090] The thresholds, weights and other setting values may be set by default according to the present invention, or may be set by an operator.
[0091] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0092] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0093] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0095] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which are all protected by the present invention.
Claims
1. A refrigeration room energy-saving optimization control method based on deep reinforcement learning, characterized by: The following steps are involved: S1. Divide the computer room into self-organizing thermal grids based on the cabinet spacing to obtain initial cooling zones and calculate the load status within the zones in real time. S2: Check whether there is a device in each U position of each cabinet. If there is an empty position, run S3; if there is a device, run S4 directly; S3. Set up a cold air leakage path monitoring strategy to monitor the heat distribution of the vacant U-space at the front end of each cooling zone cabinet in real time, automatically identify the cold air leakage path, and obtain the compensation air volume for each cooling zone; S4. Construct an air volume configuration algorithm, taking the load status and compensation air volume in the cooling zone as inputs of the air volume configuration algorithm, and combining the cable density in the cooling zone to obtain the air volume and angle of each air outlet; The self-organizing thermal grid division process includes: S11. Divide the computer room plane into M×N two-dimensional grids of uniform size; S12: For each cabinet j, release evenly in the r grid units around it. Each grid unit is continuously moved on the grid by a virtual thermal particle. Each virtual thermal particle moves on the grid according to the temperature potential energy iteration rule. The grid point with the highest energy becomes the thermal core of the cabinet. The thermal core attracts the surrounding cabinets, and adjacent thermal cores naturally converge to form a partition boundary. For each thermal core partition, the cabinets it covers are collected to complete the initial cooling partition. S13, extracting the number of cabinets corresponding to each initial cooling zone, and calculating the load status in the zone according to the number of U positions in the cabinets; The specific steps of the cold air leakage path monitoring strategy include: S31. Place an infrared camera on the top or in front of the cabinet to output a thermal image sequence , Represents the U coordinate within the partition, and denoises the heat map sequence to obtain a smooth temperature field ; S32, synchronously collect the static differential pressure of each U position With dynamic flow rate Signal, the temperature gradient and pressure gradient are used as leakage threat input, and the leakage threat of each U position is defined by setting the target inlet temperature ; S33: The top 5% of leakage threat rankings in U-positions are used as the starting point set, and the return air outlet or cold aisle outlet at the corresponding partition boundary is used as the end point; S34. Setting leakage cost function ,in represents the distance between U-bits, and l represents the number of U-bits in a partition; S35. Output the path corresponding to the minimum leakage cost function of each cooling partition , get the compensation air volume of each cooling zone; The specific steps of the air volume configuration algorithm include: S41, randomly generate particle air volume and angle, and use a vector to represent all configurations of each particle; S42, according to the air volume and angle configuration of each particle, taking the load state and compensation air volume in the cooling zone as input, creating a temperature balance model and obtaining a fitness function fit; S43. After the particle is updated, the wind volume and direction configuration corresponding to the current solution is accepted according to the particle acceptance mechanism; S44. When the fitness function fit reaches the set maximum number of iterations or reaches the set convergence threshold, the iteration is stopped and the current air volume and angle configuration is output. Otherwise, return to step S42 to continue iteration.
2. The method for energy-saving optimization control of a refrigeration room based on deep reinforcement learning according to claim 1 is characterized in that: The temperature potential energy iteration rule specifically includes: S121, in the current grid cell Calculate the temperature potential energy of each adjacent grid cell in the 8 adjacent grid cells around it ,in Represents a grid cell To the power consumption value of the nearest cabinet, represents the Euclidean distance between grid cells, and represents weight; S122, the virtual thermal particle selects the grid cell with the maximum temperature potential energy from the neighborhood and moves to it. When the maximum temperature potential energy is less than 0 or the virtual thermal particle reaches the set maximum number of steps, the current grid cell is determined to be a precipitation point and the iteration stops. S123. Count the number of virtual particle deposits in all grid cells , as the core energy level of the corresponding grid unit, the grid with the highest core energy level of each cabinet is defined as the thermal core of the cabinet; S124. For each hot core , calculate the potential energy mountain spreading outward from the core ,in, Represents a grid cell to the hot core The number of grid steps, represents the diffusion damping coefficient, for the grid element Belongs to the partition represented by the hot core that makes its potential energy hill the largest.
3. The method for energy-saving optimization control of a refrigeration room based on deep reinforcement learning according to claim 2 is characterized in that: The cooling zone compensation air volume will be increased by On each node The threat of leakage is The proportion of the total leakage threats of all nodes on the path is used as the path weight of the node , and then get the compensation air volume in the cooling zone , Indicates the target inlet air temperature, Indicates the system calibration coefficient, L indicates the path The number of all U positions on the .
4. The method for energy-saving optimization control of a refrigeration room based on deep reinforcement learning according to claim 3 is characterized in that: The temperature balance model in the fitness function is expressed as ,in, Indicates the Temperature balance of each cooling zone, Indicates the reference supply air temperature, Indicates the cable density of the partition, Indicates the air volume cooling item, Indicates the load status, Indicates cooling zones Total air supply volume, cooling zones The total air volume, c represents the cooling coefficient per unit air volume, Indicates the set proportional coefficient; The fitness function is obtained by combining the temperature balance model, the compensation air volume penalty and the air volume demand compensation. ,in, Indicates the temperature balance of all cooling zones, represents the first set of penalty terms, represents the second set of penalty terms, Indicates the minimum air volume requirement within the set cooling zone.
5. The method for energy-saving optimization control of a refrigeration room based on deep reinforcement learning according to claim 4 is characterized in that: The particle acceptance mechanism includes: If the new fitness is less than the particle fitness before the update, the newly generated solution is accepted, that is, the set of wind volumes corresponding to the particle after this iteration and wind direction Configuration, taking the new solution as the current solution for the next step; When the new fitness is greater than or equal to the fitness of the particle before the update, the probability Accept the air volume and angle configuration corresponding to the current solution.
6. The method for energy-saving optimization control of a refrigeration room based on deep reinforcement learning according to claim 1, characterized in that: In step S13, the load state is obtained by calculating the product of the total number of U bits in the cooling zone and the average power consumption. Load status of cooling zones .
7. The method for energy-saving optimization control of a refrigeration room based on deep reinforcement learning according to claim 1, characterized in that: Step S2 includes installing a miniature wind pressure sensor and wind pressure fine-tuning device on the rear side or bottom of each U position of each cabinet. A group of sensing units are independently deployed at each U position to synchronously collect static differential pressure and dynamic flow rate signals. A second-order low-pass filter is performed inside the sensor to remove high-frequency mechanical noise and then perform sliding averaging. The actuator is triggered to close only after KL consecutive judgments are that the position is empty. If it is judged that there is equipment at any time, the judgment count is reset.
Citation Information
Patent Citations
Evaporative cooling air conditioner energy-saving optimization control method and system based on mechanism model
CN119508965A
Shell-and-tube heat exchanger abnormity diagnosis method based on multi-parameter monitoring
CN120160776A