Compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning
By introducing adaptive multi-strategy deep reinforcement learning in the compressor control system and dynamically adjusting the compressor operating parameters, the problem that traditional control methods are difficult to adapt to variable operating conditions is solved, and more efficient energy efficiency optimization and intelligent control are achieved.
Patent Information
- Application Number
- CN202510208103.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional compressor control methods are difficult to effectively adapt to variable working conditions and load requirements, resulting in the inability to maximize energy efficiency.
The compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning is adopted. By initializing equipment, function construction equipment and policy execution equipment, dynamically adjusting the operating parameters of the compressor to optimize energy efficiency.
It greatly enhances the performance of the compressor under complex control needs, reduces local optimal risks, improves control efficiency, and makes the control strategy more intelligent.
Smart Images

Figure CN119957472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of compressor control technology, and in particular to a compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning. Background Art
[0002] Air compressors are widely used in industrial production, especially in compressed air supply, driving of pneumatic tools and equipment, etc.; however, due to the high energy consumption characteristics and complex operating environment of compressors, how to optimize the energy efficiency of compressors has always been a challenge in the industrial field. In traditional control methods, such as adjustment methods based on fixed compression strategies and load control, they often cannot effectively adapt to changing working conditions and changes in load demand, resulting in the inability to maximize energy efficiency.
[0003] In recent years, with the rapid development of machine learning, especially reinforcement learning, data-driven optimization methods have gradually been applied to compressor control. Adaptability is the core characteristic of reinforcement learning, which can select the optimal action under different states and control the compressor.
[0004] Therefore, “how to introduce an adaptive control strategy to dynamically adjust the compressor operating parameters” is the technical problem that the present invention needs to solve. Summary of the invention
[0005] The purpose of the present invention is to provide a compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning, so as to solve the problem of "how to introduce adaptive control strategy to dynamically adjust the compressor operating parameters" raised in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning, the system comprising:
[0008] The system includes: an initialization device, a function construction device and a strategy execution device;
[0009] The initialization device is used to determine the operating parameters of the compressor, wherein the operating parameters include at least: rotation speed, exhaust temperature, gas flow rate and load, configure a fluctuation range, integrate the fluctuation range, generate an action space, use the sensor pre-integrated in the compressor, collect the real-time value of the operating parameters, and define a state space;
[0010] The function construction device is used to construct a reward function, wherein the reward function is composed of an energy efficiency reward function, a temperature control reward function and an equipment health reward function, and all reward functions are integrated to obtain a comprehensive function;
[0011] The strategy execution device is used to configure the initial operating parameters of each compressor, create several strategy networks, establish the corresponding relationship between the strategy network, action space and compressor, use the existing data set to train all the strategy networks, calculate the Q value of each strategy network through the comprehensive function, traverse the execution strategy based on the Q value, send the execution strategy to the compressor, and modify the state space.
[0012] Further, the initialization device includes:
[0013] A range defining module is used to determine the operating parameters of the compressor, wherein the operating parameters at least include: power consumption, exhaust temperature, gas flow rate and load, and configure the fluctuation range of the operating parameters;
[0014] The space definition module is used to integrate the fluctuation range, generate an action space, collect the real-time value of the operating parameter by using the sensor pre-integrated in the compressor, and define it as a state space.
[0015] Furthermore, the policy execution device includes:
[0016] The strategy network building module is used to configure the initial operating parameters of each compressor, create several strategy networks, and establish the corresponding relationship between the strategy network, action space and compressor;
[0017] A parallel training module, used to train all of the policy networks using existing data sets;
[0018] The strategy selection module is used to calculate the Q value of each strategy network through the comprehensive function, define the strategy network with the largest Q value as the execution strategy, send the execution strategy to the compressor, and modify the state space.
[0019] Furthermore, the space definition module includes:
[0020] The action space definition module is used to establish the mapping between the operating parameters and the control actions, and obtain the speed control action, pressure threshold control action, cooling system control action and load regulation control action, which are recorded as , , and ;
[0021] Integration generates action space, denoted as ; ;
[0022] in , and
[0023] , and are the maximum and minimum speeds of the compressor, respectively. and are the maximum and minimum pressures of the compressor, Indicates that the cooling system is shut down. Indicates that the cooling system is running at full power. Indicates that the compressor stops running. Indicates that the compressor starts running and adjusts the load;
[0024] The state space definition unit is used to define the state space by using the control action corresponding to the real-time value.
[0025] Furthermore, the function construction device includes:
[0026] Function building module, used to create energy efficiency reward function, temperature control reward function and equipment health reward function respectively;
[0027] The energy efficiency reward function is: , For energy efficiency rewards, is the weight parameter, is the power of the compressor at time t, is the ideal power;
[0028] The temperature control reward function is: , Reward for temperature control, is the weight parameter, is the temperature of the compressor at time t, is the ideal temperature;
[0029] The device health reward function is: , is the weight parameter, It is the health status of the equipment and is used to characterize the durability of the compressor;
[0030] An integration module is used to integrate the energy efficiency reward function, the temperature control reward function and the equipment health reward function to generate a comprehensive function, wherein the comprehensive function is: .
[0031] Furthermore, the strategy network construction module includes:
[0032] An output unit, used to determine a test state, and input the test state into all strategy networks, and output a number of initial strategy groups;
[0033] The obtaining unit is used to calculate the Q value of the initial strategy group, and sort them in descending order to obtain a queue. A preset number of initial strategy groups are selected from the front of the queue and defined as excellent strategies. The excellent strategies are mutated and cross-examined to obtain execution strategies.
[0034] Furthermore, the system also includes:
[0035] A monitoring device, used to record the change in the Q value after the compressor is adjusted according to the execution strategy;
[0036] The sending device is used to send the change amount to a preset terminal.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] By constructing a strategy network, the parameters of the compressor can be optimized for different operating states or environmental conditions, greatly enhancing the compressor's ability to cope with complex control requirements. Through mutation and crossover, the strategy network can be globally optimized. By calculating the Q value of each strategy network, the optimal control of the compressor under different operating conditions can be ensured. At the same time, it also greatly enhances the global search capability of the strategy, reduces the risk of falling into local optimality, and greatly improves the compressor control efficiency. At the same time, the compressor can be continuously optimized and continuously adapted to new operating conditions by constructing a strategy network, making the control strategy more intelligent. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A schematic diagram of the operation flow of a compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention;
[0040] Figure 2 A block diagram of a compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention;
[0041] Figure 3 A block diagram of the composition of an initialization device in a compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention;
[0042] Figure 4 A block diagram of the composition of a function construction device in a compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention;
[0043] Figure 5 A block diagram of the composition of a strategy execution device in a compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention;
[0044] Figure 6A block diagram of the composition of a space definition module in a compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention;
[0045] Figure 7 A block diagram of the composition of the strategy network building module in the compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0047] Figure 1 and Figure 2 The structure block diagram of the compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention is shown. The compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning 1 includes: an initialization device 11, a function construction device 12 and a strategy execution device 13;
[0048] The initialization device 1 is used to determine the operating parameters of the compressor, wherein the operating parameters include at least: rotation speed, exhaust temperature, gas flow rate and load, configure a fluctuation range, integrate the fluctuation range, generate an action space, use sensors pre-integrated in the compressor to collect real-time values of the operating parameters, and define a state space.
[0049] During the actual operation of the compressor, its operating status and energy efficiency are affected by the following factors: 1. Current power consumption: indicates the current power usage of the compressor; 2. Exhaust temperature: indicates the temperature of the compressed gas. Too high an exhaust temperature will affect the long-term performance of the compressor; 3. Output gas flow: The gas output flow of the compressor is closely related to the compressor speed and load; 4. Load conditions: The current system load conditions reflect the working pressure and actual usage of the compressor; 5. Equipment health: The operating efficiency of the compressor is affected by factors such as equipment aging and maintenance status.
[0050] Using the above variables, the state space is defined, namely ;in ,in is the current power consumption (unit: watt); is the exhaust temperature (unit: Celsius); is the gas flow rate (unit: cubic meter / second); is the load condition (unit: kW); It is the historical status information, including the power consumption, temperature and other information at the last moment.
[0051] Furthermore, the action space needs to include all parameters that may be adjusted during the operation of the compressor, such as compressor speed, pressure threshold, cooling system and load adaptation and start-stop control.
[0052] The function construction device 12 is used to construct a reward function, wherein the reward function is composed of an energy efficiency reward function, a temperature control reward function and an equipment health reward function, and all reward functions are integrated to obtain a comprehensive function.
[0053] The reward function directly affects the learning objective of the compressor. In this example, the reward function takes into account energy efficiency, temperature control, and equipment health:
[0054] Energy efficiency rewards: and ideal power The goal is to minimize energy consumption;
[0055]
[0056] in For energy efficiency rewards, is a weight parameter, which is pre-determined by professionals. is the power of the compressor at time t, is the ideal power;
[0057] Temperature Control Bonus: With ideal temperature The goal is to maintain the operating temperature of the equipment.
[0058]
[0059] Reward for temperature control, is a weight parameter, which is pre-determined by professionals. is the temperature of the compressor at time t, For the ideal temperature.
[0060] Equipment health rewards: The long-term stability and health of the equipment also need to be considered to avoid over-operation.
[0061]
[0062] in It is the health status of the device and represents the durability of the device.
[0063] Integrating the above energy efficiency reward function, temperature control reward function and equipment health reward function, the final comprehensive reward function is:
[0064]
[0065] The strategy execution device 13 is used to configure the initial operating parameters of each compressor, create several strategy networks, establish the corresponding relationship between the strategy network, action space and compressor, use the existing data set to train all the strategy networks, calculate the Q value of each strategy network through the comprehensive function, traverse the execution strategy based on the Q value, send the execution strategy to the compressor, and modify the state space.
[0066] The initial moment is determined and the operating parameters at the initial moment are defined as the initial operating parameters. Several policy networks are created, and multi-strategy parallel learning methods, policy optimization methods and adaptive policy selection methods are introduced to enhance the capabilities of the policy network, so that it can quickly optimize the control strategy under multi-task and multi-environment settings and avoid local optimality.
[0067] Specifically, each policy network Optimize for different compressor operating conditions or environmental conditions. Next, for each policy network Output a specific action ;
[0068]
[0069] in, Represents different policy networks, It is The parameters of the policy network, For the current system state, each strategy network is trained independently and optimized through the objective function. The diversity of the algorithm is enhanced through parallel strategy networks, so as to better cope with the complex control requirements of the compressor.
[0070] The current environment status The inputs to each policy network are associated with each policy network, and each policy network is optimized for operation within a specific range of states; for example, when the compressor is under high load, some policy networks focus on increasing power output, while others focus on maintaining the compressor in a stable and energy-saving mode.
[0071] These policy networks are According to the input status Output different control actions; guide the policy network to learn how to select the optimal control strategy in the current state during training, so as to dynamically select among multiple strategies; the training of each policy network adopts a similar method to traditional deep optimization, but the training objectives of each strategy are different; the policy network updates parameters by optimizing the following objectives :
[0072]
[0073] in, It is a critic network Value function, is the target critic network Value estimation; by minimizing this loss function, the policy network can be optimized through gradient descent to find the optimal control strategy; in order to ensure the stability of policy network training, the policy network uses a target network mechanism similar to deep reinforcement learning. Each policy network has a corresponding target network , the target network is updated using a soft update strategy to smooth the training process:
[0074]
[0075] in, is the step size of the soft update (usually a small value, such as 0.001). This strategy ensures that the strategy update will not fluctuate drastically, thereby improving the stability of training.
[0076] The policy network uses evolutionary strategies to globally optimize multiple parallel strategies. Evolutionary strategies can optimize control strategies without gradients, avoiding the problem of gradient vanishing or local optimality. Specifically, we implement evolutionary strategies through the following steps:
[0077] 1. Generate an initial strategy group: For each control task, initialize a strategy group , each strategy Corresponding to a policy network, the initial parameters are randomly generated by Gaussian distribution.
[0078] 2. Evaluate strategy groups: Each strategy Execute several steps in the environment and reward Evaluate the strategy. The reward function can be energy saving performance, compressor efficiency, etc.
[0079] 3. Select the best strategy: Select the best performing strategy based on the evaluation results strategies and use them as the basis for the next generation of strategies.
[0080] 4. Mutation operation: mutate the selected strategy to generate a new strategy. The mutation operation modifies the parameters of the strategy network by adding noise:
[0081]
[0082] in, is the mutation step length, is a random change in strategy.
[0083] Crossover operation: Cross the selected strategies to generate new strategies. The crossover operation generates new strategies by combining some parameters of multiple excellent strategies, thereby achieving strategy fusion and optimization. Through evolutionary strategies, the policy network can effectively search for the optimal strategy in complex, high-dimensional, and nonlinear tasks, avoiding the local optimal solution problem that may occur in traditional gradient descent methods.
[0084] Based on the parallel learning of multiple strategies, the policy network also designs an adaptive strategy selection mechanism, which enables the system to automatically select the optimal strategy according to the current environmental status and task requirements. This mechanism decides which strategy to select by calculating the Q value of each strategy and can be dynamically adjusted during operation.
[0085] In the framework of multi-strategy parallel learning, the current state of the environment Will be input into each policy network to generate different actions , the policy network will evaluate the performance of each policy by calculating the Q value of each policy, and select the policy with the largest Q value for execution. The specific policy selection formula is as follows:
[0086]
[0087] in, Representation strategy In the current state The policy network selects the policy with the largest Q value to perform the control action.
[0088] The strategy network is used to adaptively adjust the compressor. The working environment of the compressor will change continuously (for example, load fluctuations, external temperature changes, etc.). The strategy network can automatically select the most adaptable strategy based on the current state information, avoiding the performance degradation caused by relying on a single strategy. Specifically, each strategy network will generate a corresponding Q value for each possible environmental state. When the environment changes, the strategy network can automatically select the most suitable strategy based on the current state.
[0089] In addition, in a multi-compressor system, the strategy network achieves collaborative control among multiple compressors through collective Q learning. Each compressor acts as an independent compressor, and the collective Q value function is used to evaluate the energy efficiency of the overall system and collaboratively optimize their respective control strategies.
[0090] The core idea of collective Q-learning is to learn from the contribution of each compressor (each compressor) to the global reward. Each compressor adjusts its own control strategy by calculating the impact of its actions on the global energy efficiency. To achieve this, we need to design a global Q-value function strategy network for multiple compressor systems. , which represents the combined efficiency of all compressors.
[0091] The key steps of collective Q-learning include a global reward signal: each compressor is assigned a reward based on its local environment state policy network Policy Network Execution Action Policy Network The policy network is a self-learning network with the goal of maximizing the overall energy efficiency of the system. Therefore, each compressor receives a shared global reward signal during training, which reflects the energy efficiency of the entire system. For example, when multiple compressors work in coordination, the overall energy efficiency of the system may be improved, and the global reward signal increases.
[0092]
[0093] Among them, the policy network The policy network is the first policy network The local reward of the policy network compressor, the overall reward policy network The strategy network can be an energy efficiency function of multiple compressor systems, which integrates factors such as the efficiency and working status of each compressor.
[0094] Each compressor of the collective Q learning calculates its policy network Q policy network value policy network at the current state Strategy network, and optimize these Q values through collaborative learning. Each compressor evaluates the impact of its actions on the overall energy efficiency of the system by referring to the global reward signal, and updates its own strategy network based on this evaluation. Policy network value; specifically, the collective policy network Policy Network in Policy Network Learning The strategy network value update formula is:
[0095]
[0096] The Q value of each compressor is updated by the following formula:
[0097]
[0098] Among them, the policy network The policy network is the learning rate, which represents the response speed of the compressor to the global reward signal. Each compressor adjusts its own strategy according to the global Q value to improve the overall energy efficiency of the system.
[0099] Figure 3 The structure block diagram of the compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention is shown, and the initialization device 11 includes:
[0100] The range defining module 111 is used to determine the operating parameters of the compressor, wherein the operating parameters at least include: power consumption, exhaust temperature, gas flow rate and load, and configure the fluctuation range of the operating parameters;
[0101] The space definition module 112 is used to integrate all the fluctuation ranges to generate an action space, collect the real-time values of the operating parameters using sensors pre-integrated in the compressor, and define it as a state space.
[0102] Real-time data of the compressor is obtained through sensors, including parameters such as power consumption, exhaust temperature, gas flow, load conditions, and equipment health status; these data constitute the state space of the strategy network, and the state of each compressor will be passed as input to the reinforcement learning model in the strategy network; based on the input of the current state space, the system chooses to adjust the compressor's speed, pressure threshold, cooling system, load adaptation and other control parameters according to the defined action space; set initial operating parameters for each compressor to ensure that the compressor can start working under basic conditions.
[0103] Figure 5 The structure block diagram of the compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention is shown, and the strategy execution device 13 includes:
[0104] A strategy network construction module 131 is used to configure the initial operating parameters of each compressor, create a number of strategy networks, and establish a corresponding relationship between the strategy network, the action space and the compressor;
[0105] A parallel training module 132, used to train all of the policy networks using existing data sets;
[0106] The strategy selection module 133 is used to calculate the Q value of each strategy network through the comprehensive function, define the strategy network with the largest Q value as the execution strategy, send the execution strategy to the compressor, and modify the state space.
[0107] Strategy network generation: The system generates multiple strategy networks for different states of each compressor. These strategy networks are optimized for different working conditions (such as high load, low load, over temperature, etc.).
[0108] Parallel training: The policy network uses parallel learning to adaptively adjust the operating status of the compressor using multiple policy networks. Each policy network is trained independently to optimize the control strategy by calculating the reward function (energy efficiency, temperature control, equipment health).
[0109] Adaptive strategy selection: According to the changes in the environment, the strategy network will dynamically select the optimal strategy network. When the compressor status changes, the strategy is automatically adjusted to improve energy saving and system stability.
[0110] Figure 6 The structure block diagram of the compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention is shown, and the space definition module 112 includes:
[0111] The action space definition module 1121 is used to establish a mapping between the operating parameters and the control actions, and obtain the speed control action, the pressure threshold control action, the cooling system control action and the load adjustment control action, which are respectively recorded as , , and ;
[0112] Integration generates action space, denoted as ; ;
[0113] in , and , and are the maximum and minimum speeds of the compressor, respectively. and are the maximum and minimum pressures of the compressor, Indicates that the cooling system is shut down. Indicates that the cooling system is running at full power. Indicates that the compressor stops running. Indicates that the compressor starts running and adjusts the load;
[0114] The state space definition unit 1122 is used to define a state space using the control action corresponding to the real-time value.
[0115] Figure 4 The structure block diagram of the compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention is shown, and the function construction device 12 includes:
[0116] A function building module 121, used to create an energy efficiency reward function, a temperature control reward function and a device health reward function respectively;
[0117] The energy efficiency reward function is: , For energy efficiency rewards, is the weight parameter, is the power of the compressor at time t, is the ideal power;
[0118] The temperature control reward function is: , For energy efficiency rewards, is the weight parameter, is the temperature of the compressor at time t, is the ideal temperature;
[0119] The device health reward function is: , is the weight parameter, It is the health status of the equipment and is used to characterize the durability of the compressor;
[0120] The integration module 122 is used to integrate the energy efficiency reward function, the temperature control reward function and the equipment health reward function to generate a comprehensive function, wherein the comprehensive function is: .
[0121] Figure 7 The structure block diagram of the compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention is shown. The strategy network construction module 131 includes:
[0122] An output unit 1311 is used to determine a test state, and input the test state into all strategy networks to output a number of initial strategy groups;
[0123] Obtaining unit 1312, used to calculate the Q value of the initial strategy group, and sort them in descending order to obtain a queue, select a preset number of initial strategy groups from the front of the queue, and define them as excellent strategies, mutate the excellent strategies, and cross-examine to obtain execution strategies.
[0124] Multi-compressor collaborative control: In the application scenario of multiple compressors, the collective Q learning mechanism is introduced. Each compressor acts as an independent compressor and works together with other compressors to optimize the energy efficiency of the entire system.
[0125] Global reward signal: The system calculates a global reward function for multiple compressors, and the operation of all compressors is optimized through collective learning to improve the energy saving effect of the overall system.
[0126] Strategy update and optimization: Each compressor evaluates the impact of its own actions on system efficiency by calculating the Q value related to the global energy efficiency, and optimizes the strategy based on the collective Q value.
[0127] Figure 2 The structure block diagram of the compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning provided by an embodiment of the present invention is shown, and the system also includes:
[0128] A monitoring device 14, for recording a change in the Q value after the compressor is adjusted according to the execution strategy;
[0129] The sending device 15 is used to send the change amount to a preset terminal.
[0130] Continuous training and evaluation: The system continuously trains the policy network during operation, and minimizes power consumption, maintains temperature stability, and extends device life by constantly adjusting the control strategy.
[0131] Execution strategy: Once the strategy is optimized and verified, the system will output the optimal control action based on real-time data, and adjust the working status of each compressor (such as speed, pressure, cooling, etc.) in real time to achieve energy saving.
[0132] Real-time monitoring and feedback: The system monitors the actual operating results of each compressor and compares the energy efficiency performance before and after optimization. Through the reward function, the system will further adjust the strategy based on feedback information such as equipment health, energy efficiency, and temperature.
[0133] Strategy adjustment and evolution: The strategy network further optimizes multiple strategy networks through evolutionary strategies to avoid local optimality and ensure long-term stable operation of the compressor under different working conditions.
[0134] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0135] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
[0136] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning, characterized in that: The system includes: an initialization device, a function construction device and a strategy execution device; The initialization device is used to determine the operating parameters of the compressor, wherein the operating parameters include at least: rotation speed, exhaust temperature, gas flow rate and load, configure a fluctuation range, integrate the fluctuation range, generate an action space, use the sensor pre-integrated in the compressor, collect the real-time value of the operating parameters, and define a state space; The function construction device is used to construct a reward function, wherein the reward function is composed of an energy efficiency reward function, a temperature control reward function and an equipment health reward function, and all reward functions are integrated to obtain a comprehensive function; The strategy execution device is used to configure the initial operating parameters of each compressor, create several strategy networks, establish the corresponding relationship between the strategy network, action space and compressor, use the existing data set to train all the strategy networks, calculate the Q value of each strategy network through the comprehensive function, traverse the execution strategy based on the Q value, send the execution strategy to the compressor, and modify the state space.
2. The compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning according to claim 1 is characterized in that: The initialization device comprises: A range defining module is used to set the operating parameters of the compressor, wherein the operating parameters include at least: power consumption, exhaust temperature, gas flow rate and load, and configure the fluctuation range of the operating parameters; The space definition module is used to generate an action space through the fluctuation range, collect the real-time value of the operating parameter by using the sensor pre-integrated in the compressor, and define it as a state space.
3. The compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning according to claim 2 is characterized in that: The function construction device comprises: The action space definition module is used to establish the mapping between the operating parameters and the control actions, and obtain the speed control action, pressure threshold control action, cooling system control action and load regulation control action, which are recorded as , , and ; Integrate all control actions to generate action space, denoted as ; ; in , and , and are the maximum and minimum speeds of the compressor, respectively. and are the maximum and minimum pressures of the compressor, Indicates that the cooling system is shut down. Indicates that the cooling system is running at full power. Indicates that the compressor stops running. Indicates that the compressor starts running and adjusts the load; The state space definition unit is used to define the state space by using the control action corresponding to the real-time value.
4. The compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning according to claim 1 is characterized in that: The policy execution device comprises: The strategy network construction module is used to configure the initial operating parameters of each compressor, build several strategy networks, and establish the corresponding relationship between the strategy network, action space and compressor; A parallel training module, used to train all of the policy networks using existing data sets; The strategy selection module is used to calculate the Q value of each strategy network through the comprehensive function, define the strategy network with the largest Q value as the execution strategy, send the execution strategy to the compressor, and modify the state space.
5. The compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning according to claim 3 is characterized in that: The space definition module includes: Function building module, used to create energy efficiency reward function, temperature control reward function and equipment health reward function respectively; The energy efficiency reward function is: , For energy efficiency rewards, is the weight parameter, for The power of the compressor at the moment, is the ideal power; The temperature control reward function is: , For energy efficiency rewards, is the weight parameter, is the temperature of the compressor at time t, is the ideal temperature; The device health reward function is: , is the weight parameter, It is the health status of the equipment and is used to characterize the durability of the compressor; An integration module is used to integrate the energy efficiency reward function, the temperature control reward function and the equipment health reward function to generate a comprehensive function, wherein the comprehensive function is: .
6. The compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning according to claim 3 is characterized in that: The policy network building module includes: An output unit, used for traversing the test state, inputting the test state into all the strategy networks, and outputting a number of initial strategy groups; The obtaining unit is used to calculate the Q value of the initial strategy group, and sort them in descending order to obtain a queue. A preset number of initial strategy groups are selected from the front of the queue and defined as excellent strategies. The excellent strategies are mutated and cross-examined to obtain execution strategies.
7. The compressor energy-saving control system based on adaptive multi-strategy deep reinforcement learning according to claim 6 is characterized in that: The system further comprises: A monitoring device, used to record the change in Q value after the compressor is adjusted according to the execution strategy; The sending device is used to send the change amount to a preset terminal.
Citation Information
Cited By
Frequency conversion speed regulation and energy efficiency optimization integrated system for air conditioner compressor of new energy automobile
CN120889735A
Intelligent grading compressor control method and system based on load self-adaption
CN121142983A
Modulation mode optimization method, controller, electronic equipment and refrigeration equipment
CN122149120A
Modulation optimization methods, controllers, electronic devices, and refrigeration equipment.
CN122149120B