A compressor energy-saving operation control method and system based on reinforcement learning
Patent Information
- Application Number
- CN202511291797.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-09-10
AI Technical Summary
现有的压缩机难以解决内部及外部环境突发情况而导致的压缩机运行状态受限,在以往压缩机与管网阻力适配性差且未建立完善的闭环控制链条
[0009]本申请实施例提供的一种基于强化学习的压缩机节能运行控制方法及系统的有益效果在于:本发明通过多参数传感器网络全面采集多维数据,结合状态预测模型对未来运行参数的精准预判,使强化学习模型能基于更全面的信息制定最优控制策略,可动态适配管网负载与环境变化,最大限度减少无效能耗,显著提升压缩机运行的能源利用效率,长期应用能带来节能效益。在控制精准性与适应性上,采用近端策略优化框架的强化学习模型,能通过持续的训练样本更新不断优化控制逻辑,针对复杂多变的运行工况具备强大的自适应调节能力,确保控制动作始终贴合实际需求,有效避免传统控制方法在动态环境下的滞后性与局限性,提升运行稳定性。从安全运行角度看,预设的安全运行约束对最优控制动作进行校验,可从源头规避可能导致设备超压、超温等危险情况的控制指令,保障压缩机在高效运行的同时不偏离安全阈值,延长设备使用寿命,降低故障维修成本。
Smart Images

Figure CN121165467B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of compressor control technology, and in particular to a compressor energy-saving operation control method and system based on reinforcement learning. Background Technology
[0002] Currently, in the field of industrial compressor operation control, dynamic adaptation and stable, efficient operation under complex working conditions remain a technological bottleneck. Existing compressors struggle to address the limitations imposed on their operation by unforeseen internal and external environmental factors. Furthermore, the compressors have historically exhibited poor compatibility with pipeline resistance and lacked a robust closed-loop control chain. Addressing these technical issues results in additional losses during compressor operation.
[0003] Therefore, a method and system for energy-saving operation control of compressors based on reinforcement learning is needed to improve the stability, energy efficiency and energy-saving effect of compressor operation. Summary of the Invention
[0004] To address the aforementioned technical problems, this application provides a compressor energy-saving operation control method and system based on reinforcement learning.
[0005] A first aspect of this application provides a compressor energy-saving operation control method based on reinforcement learning, comprising: A multi-parameter sensor network is used to collect multi-dimensional data during the compressor's operation; the multi-dimensional data includes compressor operating status data, environmental data, and pipeline load data. The multidimensional data is preprocessed to obtain target feature data; The target feature data is input into the state prediction model to predict the operating parameter values for the next control cycle. The target feature data is concatenated and fused with the predicted values of the running parameters to construct the state representation of the reinforcement learning model; The state representation is input into a target reinforcement learning model based on a proximal policy optimization framework to obtain the optimal control action. The optimal control action is verified based on preset compressor safety operation constraints to determine the execution instruction; The compressor's operating parameters are adjusted based on the executed instructions.
[0006] A second aspect of this application provides a compressor energy-saving operation control system based on reinforcement learning, comprising: The data acquisition module is used to collect multi-dimensional data during the operation of the compressor through a multi-parameter sensor network; the multi-dimensional data includes compressor operating status data, environmental data, and pipeline load data. The data processing module is used to preprocess the multidimensional data to obtain target feature data; The data prediction module is used to input the target feature data into the state prediction model to predict the predicted values of the operating parameters within a future control cycle. The splicing and fusion module is used to splice and fuse the target feature data with the predicted values of the running parameters to construct the state representation of the reinforcement learning model. The action generation module is used to input the state representation into a target reinforcement learning model based on a proximal policy optimization framework to obtain the optimal control action. The safety verification module is used to perform safety verification on the optimal control action based on preset compressor safety operation constraints, and to determine the execution instruction; The execution instruction module is used to adjust the compressor's operating parameters based on the execution instructions.
[0007] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described reinforcement learning-based compressor energy-saving operation control method.
[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described reinforcement learning-based compressor energy-saving operation control system.
[0009] The beneficial effects of the energy-saving operation control method and system for compressors based on reinforcement learning provided in this application are as follows: This invention comprehensively collects multi-dimensional data through a multi-parameter sensor network and combines this with a state prediction model to accurately predict future operating parameters. This enables the reinforcement learning model to formulate optimal control strategies based on more comprehensive information, dynamically adapting to changes in pipeline load and environment, minimizing ineffective energy consumption, significantly improving the energy utilization efficiency of the compressor, and bringing energy-saving benefits in the long term. Regarding control accuracy and adaptability, the reinforcement learning model using a near-end strategy optimization framework can continuously optimize the control logic through continuous training sample updates. It possesses strong adaptive adjustment capabilities for complex and changing operating conditions, ensuring that control actions always conform to actual needs, effectively avoiding the lag and limitations of traditional control methods in dynamic environments, and improving operational stability. From a safety operation perspective, preset safety operation constraints verify the optimal control actions, avoiding control commands that may lead to dangerous situations such as overpressure and overtemperature, ensuring that the compressor operates efficiently without deviating from safety thresholds, extending equipment lifespan, and reducing fault repair costs. Attached Figure Description
[0010] Figure 1 A flowchart illustrating a reinforcement learning-based compressor energy-saving operation control method provided in an embodiment of this application; Figure 2 A schematic diagram of the hardware composition of a compressor energy-saving operation control system based on reinforcement learning provided in an embodiment of this application; Figure 3 A structural block diagram of a compressor energy-saving operation control system based on reinforcement learning provided in an embodiment of this application. Figure 4 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0012] To make the purpose, technical solution, and advantages of this application clearer, the following will be described in conjunction with the appendix. Figure 1-4 The following is an explanation using specific examples.
[0013] Please refer to Figure 1 , Figure 1 A flowchart illustrating a reinforcement learning-based compressor energy-saving operation control method according to an embodiment of this application is provided. The method includes: S101: Collects multi-dimensional data during compressor operation through a multi-parameter sensor network; the multi-dimensional data includes compressor operating status data, environmental data, and pipeline load data.
[0014] In this embodiment, the multidimensional data includes compressor operating status data, environmental data, and pipeline load data; the compressor operating status data includes core operating parameters and equipment health parameters; the environmental status data includes meteorological parameters and geographical parameters; and the pipeline parameters include target parameters and dynamic characteristic parameters.
[0015] The core operating parameters include intake pressure, exhaust pressure, motor speed, shaft power, intake temperature, and exhaust temperature; equipment health parameters include bearing vibration amplitude; meteorological parameters include ambient temperature and relative humidity; geographical parameters include atmospheric pressure; target parameters include target flow rate and outlet pressure setpoint; and dynamic characteristic parameters include pipeline resistance coefficient.
[0016] S102: Preprocess the multidimensional data to obtain the target feature data.
[0017] In this embodiment, preprocessing includes data cleaning, outlier handling, dynamic normalization, and feature extraction of multidimensional data. Data cleaning removes high-frequency noise using sliding window filtering, with the window size dynamically adjusted according to the compressor's operating cycle. Outlier handling uses the 3σ rule and isolated forest to remove abnormal data. Dynamic normalization maps the noise- and outlier-free multidimensional data to the [0,1] interval. Feature extraction uses principal component analysis to reduce the dimensionality of the normalized multidimensional data to obtain target feature data. Feature extraction retains principal components with a cumulative contribution rate ≥95% of the normalized multidimensional data, while also retaining key physical features such as motor power and unit energy consumption.
[0018] S103: Input the target feature data into the state prediction model to predict the operating parameter values for the next control cycle.
[0019] In this embodiment, the state prediction model is constructed using a Long Short-Term Memory (LSTM) network. The input layer dimension is consistent with the target feature data dimension. The hidden layer consists of two fully connected layers plus a multi-head self-attention mechanism. Each fully connected layer has 256 neurons with the LeakyReLU activation function, and the attention score dynamically focuses the pressure. The system incorporates parameters such as power to accelerate the strategy's response to pressure fluctuations and suppress noise parameter interference. The output layer contains predicted operating parameters for future control cycles, including compressor energy consumption and pipeline resistance coefficients. Training samples are generated through a sliding time window, and the model is periodically fine-tuned using newly acquired data to maintain prediction accuracy.
[0020] S104: Concatenate and fuse the target feature data with the predicted values of the running parameters to construct the state representation of the reinforcement learning model.
[0021] In this embodiment, the state representation, i.e., the state vector dimension, is the sum of the target feature data dimension and the predicted value dimension. During the fusion process, dynamic weights are assigned to historical features and predicted features respectively. In addition, the state vector needs to be subject to boundary constraint processing to ensure that the state input to the reinforcement learning model conforms to the actual running scenario.
[0022] S105: Input the state representation into the target reinforcement learning model based on the proximal policy optimization framework to obtain the optimal control action.
[0023] In this embodiment, the reinforcement learning model adopts an actor-critic structure. The actor network outputs control actions, which include a continuous action space of compressor speed adjustment, intake valve opening adjustment, and return valve opening adjustment. The critic network evaluates the value of the actions. Policy optimization adopts a proximal policy optimization algorithm, which ensures the stability of policy updates by pruning the objective function. The reward function is designed as a multi-objective weighted form, including energy-saving reward, stability reward, and response reward. By dynamically adjusting the priority of the objectives, global optimal control under complex operating conditions is achieved.
[0024] The multi-objective reward function calculation formula in this embodiment is as follows: , in, Let be the weight coefficient, and satisfy... Each weighting coefficient is based on the system load rate. Make adjustments. For target traffic, Maximum flow rate; For multi-objective reward functions; As an energy-saving reward; As a stability reward; In response to the reward.
[0025] Specifically, low load Used for priority energy-saving rewards; Variable load Used to balance the three rewards; High load Used for priority response rewards S106: Based on preset compressor safety operation constraints, perform safety verification on the optimal control action and determine the execution command; In this embodiment, safety constraints include hard constraints and fluctuation constraints. For example, when exhaust temperature and shaft vibration exceed hard constraints, emergency actions are triggered; when pressure standard deviation and speed standard deviation exceed fluctuation constraints, smoothing corrections are performed. The verification process is implemented using a rule engine, which performs gradient projection corrections on actions that do not meet the constraints to ensure that the adjusted actions are within the safe operating range of the equipment. At the same time, the number of constraint triggers is recorded as a feedback signal for the optimization of the reinforcement learning model.
[0026] S107: Adjust the compressor's operating parameters based on the executed command.
[0027] In this embodiment, the execution process employs a closed-loop training model that involves executing actions, providing feedback data, and updating the model. Action commands are input into the compressor system, and the system adjusts the compressor's operating parameters, including speed regulation, suction valve opening regulation, and compressor start / stop status commands. During adjustment, the sampling frequency is synchronized with the control frequency, and the operating parameters are collected in real time for subsequent status updates and reward calculations. Furthermore, for scenarios with multiple compressors operating in parallel, a load allocation algorithm is used to dynamically distribute commands, ensuring optimal overall unit operating efficiency.
[0028] As can be concluded from the above, this invention comprehensively collects multi-dimensional data through a multi-parameter sensor network and combines it with a state prediction model to accurately predict future operating parameters. This enables the reinforcement learning model to formulate optimal control strategies based on more comprehensive information, dynamically adapting to changes in pipeline load and environment, minimizing ineffective energy consumption, significantly improving the energy utilization efficiency of the compressor, and bringing energy-saving benefits in the long term. Regarding control accuracy and adaptability, the reinforcement learning model using a near-end strategy optimization framework can continuously optimize the control logic through continuous training sample updates. It possesses strong adaptive adjustment capabilities for complex and changing operating conditions, ensuring that control actions always conform to actual needs, effectively avoiding the lag and limitations of traditional control methods in dynamic environments, and improving operational stability. From a safety operation perspective, preset safety operation constraints verify the optimal control actions, avoiding control commands that may lead to dangerous situations such as overpressure and overtemperature, ensuring that the compressor operates efficiently without deviating from safety thresholds, extending equipment lifespan, and reducing fault repair costs. This embodiment can also use the environmental state when executing the execution instruction, the actual execution action, the environmental feedback reward, and the environmental state after executing the execution instruction as training samples to update the target reinforcement learning model.
[0029] In this embodiment, the sample storage adopts an experience replay pool and a priority experience replay mechanism. The priority experience replay mechanism includes a first-level priority and a second-level priority. The first-level priority refers to marking the sparse reward scenario. When the response reward is triggered only when the traffic reaches the target, the sampling weight is increased by the corresponding multiple (3 times) when the absolute value of the reward is greater than the preset reward threshold (0.5). The second-level priority refers to: based on the gradient contribution effect, the learning efficiency is improved by 2 times in the sparse reward scenario compared with ordinary replay.
[0030] In summary, the state representation constructed by the preprocessing and fusion of multidimensional data provides high-quality input information for reinforcement learning, improving the reliability of decision-making. Secondly, the closed-loop training and update mechanism can continuously learn and improve in actual operation, constantly optimizing the control effect, and is applicable to compressors of different types and application scenarios.
[0031] In one embodiment of this application, the optimal control action is safety-verified based on preset compressor safety operation constraints to determine the execution instruction, including: If the optimal control action passes the verification, then the optimal control action will be used as the execution instruction. If the optimal control action fails the verification, the optimal control action that violates the safety constraints is projected to the boundary point of the action space through the action projection algorithm to obtain the adjustment control action as the execution instruction. The compressor safety operation constraints include at least one of the following: the safety threshold range of operating parameters, the limit on the rate of change of operating parameters, and anti-surge logic; the action projection algorithm is used to project the action vector that violates the constraints to the nearest boundary point of the available action space.
[0032] In this embodiment, the compressor's safe operation constraints include: the safe threshold range of operating parameters is set with reference to equipment safety standards and process red lines, such as exhaust temperature ≤120℃, shaft vibration amplitude <5mm / s, motor current ≤1.1 times rated current, and suction pressure / exhaust pressure fluctuation within ±0.5%FS; the operating parameter change rate limit is designed for continuous control actions, such as speed adjustment ≤±15% of rated speed, and suction valve / anti-surge return valve opening adjustment step ≤1% / time, to avoid mechanical shock or system oscillation caused by sudden parameter changes; the anti-surge logic is modified based on the API617 standard, for example, when the compression ratio K = exhaust pressure / suction pressure >1.2 times the safe compression ratio, the constraint is triggered, and the anti-surge return valve opening is limited to ≤30% to balance energy saving and anti-surge requirements.
[0033] In this embodiment, the optimal control action is verified based on safe operation constraints: if the optimal control action meets all constraints, it is directly used as the execution command and transmitted to the execution layer via the OPCUA industrial bus; the optimal control action includes: continuous actions, such as speed adjustment and valve opening adjustment; discrete actions, such as multi-machine start / stop commands, with specific speed adjustment within ±15% of rated speed and return valve opening ≤30%), ensuring that the command execution error is ≤0.5%; In this embodiment, if the optimal control action fails the verification, the action vector is projected to the nearest boundary point using an action projection algorithm to generate an adjustment control action that satisfies all constraints, which is then output as the execution command. The action projection algorithm calculates the Euclidean distance between the constraint-violating action vector and the feasible action space (defined by safe operation constraints). This verification mechanism reduces the probability of compressor safety constraint breach from 1.2% to 0.3%, preventing unplanned shutdowns while ensuring the continuity and stability of control actions, thus adapting to the safe operation requirements of the compressor under complex conditions such as variable load, sudden environmental changes, and fluctuations in pipeline resistance.
[0034] In this embodiment, the optimal control action is verified using preset compressor safety operation constraints. If the verification fails, an action projection algorithm is used for adjustment to ensure that the compressor's operating parameters are within the constraints of safety threshold range, rate of change limit, and anti-surge logic, effectively guaranteeing the safe and stable operation of the compressor. Furthermore, the action projection algorithm projects actions that violate constraints to the boundary points of the feasible action space, resulting in adjusted control actions. This ensures both safety and effective compressor control, avoiding the problem of poor control performance that might result from completely discarding actions that do not meet the conditions.
[0035] In one embodiment of this application, the training method for a target reinforcement learning model includes: A reinforcement learning state space is constructed based on historical feature data and their corresponding historical operational parameter predictions. Construct a hybrid action space, which includes continuous action dimensions and discrete action dimensions. The continuous action dimension includes the compressor speed regulation, the suction valve opening regulation, and the anti-surge return valve opening regulation. The discrete action dimension includes the start-stop state commands of one or more compressors in a multi-machine parallel operation. Design a policy network that supports mixed action outputs, wherein the hidden layer includes a multi-head self-attention module, the discrete action dimension output layer adopts a Softmax distribution, and the continuous action dimension output layer adopts a Gaussian distribution parameterization. A multi-objective reward function is constructed by calculating energy-saving reward items, stability reward items, and response reward items based on historical feature data; A reinforcement learning model is trained using a proximal policy optimization algorithm with the state space, mixed action space, multi-objective reward function, and experience pool sample pairs to obtain a well-trained objective reinforcement learning model.
[0036] In this embodiment, a multi-dimensional reinforcement learning state space is constructed based on historical feature data and corresponding predicted values of historical operating parameters. The learning state space includes compressor operating parameters, environmental parameters, and pipeline load parameters; among which, the compressor operating parameters include suction pressure. Exhaust pressure Motor speed n, shaft power W, intake air temperature Exhaust temperature and bearing vibration amplitude Environmental parameters include ambient temperature. relative humidity and atmospheric pressure Pipeline load parameters include target flow rate. Export pressure setpoint Pipeline resistance coefficient and outlet pressure setpoint In this embodiment, all parameters are standardized using online Z-score (the mean is dynamically updated every 1000 training steps). with standard deviation (To eliminate dimensional differences and to characterize operational status.)
[0037] This embodiment constructs a continuous-discrete hybrid motion space. In the continuous motion dimension, the speed adjustment Δn is limited to ±15% of the rated speed, with an adjustment resolution of 0.05%; the intake valve opening adjustment... Coverage ranges from 0-100%, with an adjustment resolution of 0.1%; anti-surge return valve opening adjustment range. The surge is controlled within 0-30% based on the API617 surge boundary model gradient adjustment; the discrete motion dimension is controlled by instructions. ∈{0,1}, where 0 represents running and 1 represents stopping; used to realize the start-stop scheduling of one or more compressors in a multi-machine parallel system, and the start-stop decision is triggered based on the system load rate, used to shut down redundant units or wake up standby units.
[0038] In this embodiment, a policy network adapted to mixed action outputs is designed. The input layer performs robust normalization on the 13-dimensional state vector. The hidden layer adopts a structure of two fully connected layers plus an 8-head self-attention module. Each fully connected layer deploys 256 neurons with LeakyReLU activation. The multi-head self-attention module focuses on core parameters such as stress and power (the historical attention ratio is stable at 65%±5%), strengthens key feature interactions and suppresses noise interference. The output layer is designed differently for action types. Discrete action branches use a Softmax distribution combined with 0.01 entropy regularization to maintain policy flexibility, while continuous action branches are parameterized using a Gaussian distribution. Subsequently, a dynamically weighted multi-objective reward function is constructed based on historical feature data.
[0039] The formula for calculating energy-saving rewards is as follows:
[0040] in, For real-time shaft power, For the rated shaft power, in this embodiment Energy consumption incentive strategies are quantified based on the percentage of rated power to reduce power consumption. For example, when the power of a single unit is reduced from 90kW to 75kW, the reward is increased by 0.167. It is a Boolean state, and This indicates that the energy-saving idle mode has been entered; otherwise, it is 0. For example, when the load rate L < 15% and lasts for 10 seconds, the idle mode is triggered, and the speed drops to 30% of the rated speed to maintain the minimum circulation flow. The idle speed bonus coefficient can be set based on experience as follows: For example, idling mode saves 35% more energy than regular low-load operation, corresponding to a 0.2 increase in reward.
[0041] The formula for calculating stability rewards is:
[0042] in, The standard deviation of exhaust pressure is expressed in MPa. The standard deviation of exhaust temperature is expressed in °C. The standard deviation of rotational speed is given in rpm; real-time fluctuations are calculated using a sliding window based on the above parameters. The fluctuation penalty coefficient can be set empirically to... For example, when the standard deviation of the exhaust pressure increases from 0.05 MPa to 0.1 MPa, the penalty increases by 0.05, keeping the constraint parameter fluctuation within the process allowable range to avoid equipment surge caused by pressure oscillations. The process range can be set as follows:
[0043] The formula for calculating the response reward is:
[0044] in, For real-time traffic, For target traffic, For deviation threshold; exponential function The smaller the incentive deviation, the higher the reward; This is the reward coefficient; for example, the reward is 0.3 when there is no deviation in traffic flow. For example, when the deviation between real-time traffic and target traffic exceeds a certain threshold... When a fixed penalty is triggered. Specifically, when the traffic deviation is 10%, 0.2 reward will be deducted directly, and the forced strategy will prioritize meeting the load demand.
[0045] In this embodiment, a reinforcement learning state space is constructed based on historical feature data and predicted operating parameters. Simultaneously, a hybrid action space containing continuous and discrete action dimensions is constructed, which can comprehensively and accurately describe the compressor's operating state and possible control actions, providing rich input information for the reinforcement learning model. A policy network supporting hybrid action output is designed, with multi-head self-attention modules in the hidden layers to better capture the relationships between features. The discrete action dimension output layer uses a Softmax distribution, while the continuous action dimension output layer uses a Gaussian distribution parameterization, enabling the network to flexibly output different types of actions and improving the model's control capability. A multi-objective reward function is constructed based on historical feature data to calculate energy-saving, stability, and response reward terms, guiding the reinforcement learning model to achieve energy saving while also considering the compressor's operational stability and responsiveness, thus realizing multi-objective optimization.
[0046] This embodiment uses the Proximal Policy Optimization (PPO) algorithm to train the model, based on a dual-priority experience pool with a capacity of 100,000 samples. Priority is given to replaying samples with sparse rewards and large TD errors, which improves learning efficiency by 2 times. The clip range [0.8, 1.2] restricts the policy update step size, and L2 clipping (norm threshold of 1.0) is performed on the gradients of continuous action branches. Iterative optimization is carried out through offline pre-training and online fine-tuning to finally obtain a well-trained target reinforcement learning model.
[0047] For example, the offline training mode is a digital twin-driven policy preview; specifically, it includes: Simulation environment construction: Integrating the AMESim compressor thermodynamic model and Flownex pipeline fluid simulation, a digital twin was built, encompassing eight typical operating conditions, and supporting dynamic interpolation of operating parameters. These typical operating conditions include: low load, variable temperature, pipeline scaling, and multi-unit switching.
[0048] Pre-training objective: To generate an initial strategy library by iterating the PPO algorithm for 100,000 steps in the simulation environment, so that the energy saving rate in the simulation scenario can reach 15%±2% compared with traditional PID control, and the strategy's response delay to the scenario of a 15% change in pipeline resistance is ≤3 seconds.
[0049] Domain Adaptation Enhancement: By employing adversarial domain adaptation technology, the distribution difference between the simulation and the real environment is reduced, which improves the convergence speed of the pre-trained initial strategy by 3 times in the early stage of its deployment.
[0050] Specifically, the online optimization mode is a real-time self-correction strategy based on operating conditions, including: Data acquisition and triggering: Based on the EtherCAT bus, 13-dimensional status data is synchronously acquired at the 50ms level, and the strategy update is triggered every 10 control cycles.
[0051] Optimize algorithm configuration: Employ Adaptive-LR gradient descent, with the learning rate dynamically adapted to changing operating conditions. Operating conditions are stable: the learning rate has decreased. To prevent strategy oscillations, the standard deviation of speed fluctuations was reduced from 5 rpm to 3 rpm. A change rate of more than 15% in the load mutation causes the learning rate to increase to [a certain value]. This is used to accelerate the response, reducing the traffic step response latency from 8 seconds to 5 seconds.
[0052] Strategy stability constraint: Set the strategy update pruning threshold to 0.2, corresponding to a control action change rate ≤10%. For example, the speed adjustment is narrowed from ±15% of the rated value to ±12% to avoid equipment overshoot due to drastic strategy updates. For example, the pressure overshoot is reduced by 40%.
[0053] This embodiment also constructs a closed-loop control logic of real-time monitoring, tiered response, and reward / penalty. Performance constraint: exhaust temperature. (Level 1 response triggered by temperature exceeding 5°C), shaft vibration amplitude (20% over-amplitude triggers secondary response); Equipment constraint: motor current (A 10% overload triggers a Level 3 response, resulting in immediate shutdown). The graded response mechanism (time ≤ 200ms, meeting SIL2 safety level) is shown in Table 1.
[0054] Table 1
[0055] Specifically, the reward function is linked to the penalty; among them, Level 1 constraint: Add dynamic penalty term For example, for every 5°C of overheating, a linear increase of 0.1 is penalized to force the model to learn a temperature suppression strategy; Level 2 / Level 3 constraints: Add rigid penalties For example, directly deducting 50% / 100% of the core rewards can prevent dangerous strategies from recurring.
[0056] The effect of this embodiment is to reduce the probability of policy update by 80% when constraints are breached, so that the model prioritizes learning the decision logic of optimization within the safe boundary.
[0057] The iterative enhancement of virtual-real collaboration in this embodiment includes a simulation feedback mechanism: during online operation, constraint triggering events are reinjected into the digital twin, automatically generating 30% of the anti-disturbance training samples to simulate similar fault conditions, thereby improving the digital twin model's defense capability against such conditions by 50%, for example, reducing the number of over-temperature triggers by 60% per month; strategy version management: a rolling update + rollback mechanism is adopted, retaining the top 3 effective strategy versions, and automatically rolling back to the historical best version when the weekly number of safety constraint breaches increases by more than 20%; and combining Bayesian optimization to select update timing to avoid system risks caused by strategy degradation.
[0058] In one embodiment of this application, the training method for the target reinforcement learning model further includes: Based on the state vector within a preset time window, the average Euclidean distance change rate relative to the mean state value within the window is calculated and divided by the historical maximum change rate, which serves as a volatility indicator. Adjust the upper limit parameter of the trust domain range of the near-end strategy optimization algorithm based on the volatility index; When the volatility indicator is less than the benchmark value, increase the upper limit parameter of the trust domain range; When the volatility indicator is greater than the benchmark value, reduce the upper limit parameter of the trust domain range; The baseline value is the average rate of change of the Euclidean distance under historical stable operating conditions.
[0059] In this embodiment, continuous 13-dimensional state data is collected based on a preset time window. First, the Euclidean distance of all state vectors within the window relative to the mean of the states within the window is calculated. Then, the real-time rate of change of the Euclidean distance is calculated with a sliding step of 1 second. The mean of the rate of change within the window is taken to obtain the average rate of change of the Euclidean distance. Finally, it is divided by the maximum average rate of change of the Euclidean distance recorded in typical operating conditions in history to obtain a standardized volatility index, which can accurately capture the volatility characteristics of operating conditions such as sudden changes in ambient temperature of ±15℃ and sudden changes in pipeline resistance of 15%.
[0060] In this embodiment, the volatility index is used as the core feedback signal to dynamically adjust the upper limit parameter of the trust domain range of the near-end strategy optimization algorithm. The algorithm's default clip range is [0.8, 1.2]. The upper limit of the trust domain range directly determines the strategy update step size: when the volatility index is less than the benchmark value, it indicates that the system is in a stable operating condition. At this time, the upper limit parameter of the trust domain range is expanded to improve the flexibility and exploration efficiency of the strategy update and push the real-time power to approach below 90% of the rated power. When the volatility index is greater than the benchmark value, it means that the operating condition fluctuates violently. The upper limit parameter of the trust domain range needs to be reduced to limit the strategy update amplitude and avoid equipment surge or shutdown caused by sudden changes in speed of ±20% or valve opening mismatch.
[0061] The benchmark value needs to be determined based on historical stable operating condition data: select stable operating condition segments in historical operation that meet the requirements of exhaust temperature ≤120℃, shaft vibration amplitude <5mm / s, and pressure and speed standard deviation within the process allowable range, calculate the average Euclidean distance change rate of the time window corresponding to each segment, and take the mean of all segments as the benchmark value to ensure that it can accurately represent the fluctuation benchmark of the stable operation of the system, and finally realize the dynamic adaptation of algorithm parameters and operating conditions.
[0062] In this embodiment, the upper limit parameter of the trust domain range of the near-end strategy optimization algorithm is adjusted according to the volatility index of the state vector within a preset time window. When the volatility index is less than the benchmark value, the trust domain is expanded; when it is greater than the benchmark value, the trust domain is shrunk. This allows the model to adaptively adjust the learning step size under different operating conditions, improving learning efficiency and model stability. This adaptive adjustment mechanism avoids the model from overexploring under stable conditions, which could lead to low learning efficiency, and also prevents the model from becoming overly conservative due to an excessively large trust domain when operating conditions change significantly, thus hindering its ability to adapt to environmental changes in a timely manner.
[0063] In one embodiment of this application, the training method for the target reinforcement learning model further includes: The prediction error index of the state prediction model for the predicted values of operating parameters is calculated; the prediction error index is the weighted average of the absolute values of the differences between the predicted values of operating parameters and the actual values of operating parameters within a preset window. When the prediction error index exceeds the preset error threshold, the discount factor will be reduced by the first step size. When the prediction error index is less than or equal to the preset error threshold, the discount factor will be increased by a second step. In this case, the length of the first step is less than the length of the second step, and the adjustment range of the discount factor is between the preset minimum and maximum values.
[0064] In this embodiment, the prediction error index is calculated based on the core operating parameters in the state space, including the suction pressure of the compressor body. Exhaust pressure Motor speed n, shaft power W, and target flow rate of the pipeline load. Pipeline resistance coefficient A preset time window is used to simultaneously acquire the predicted values of operating parameters output by the state prediction model and the actual values measured by the sensors. First, the absolute value of the difference between the predicted and actual values of each parameter within the time window is calculated. Then, weights are assigned according to the importance of the parameters, and finally, a weighted average is obtained as the prediction error index. The preset error threshold is determined based on the historical training target reinforcement learning model.
[0065] For example, the initial value of the discount factor is set to 0.9 to balance short-term and long-term rewards. The preset error threshold is based on the average prediction error under historical stable operating conditions. For example, the shaft power threshold is set to 3% of the rated power and the flow rate threshold is set to 2% of the target flow rate. The first step size is set to 0.02 and the second step size is set to 0.05, where the first step size is smaller than the second step size to avoid excessive decrease of the discount factor when the error increases. The adjustment range is limited to 0.7-0.95. When the prediction error index is greater than the preset error threshold, it indicates that the model prediction accuracy is insufficient. The discount factor is reduced from 0.9 to 0.88 with a step size of 0.02 to reduce the model's dependence on uncertain long-term prediction results and prioritize the optimization of short-term control based on real-time actual values. When the prediction error index is less than or equal to the preset error threshold, it indicates that the model prediction is reliable. The discount factor is increased from 0.9 to 0.95 with a step size of 0.05 to strengthen long-term rewards and promote the strategy to converge toward the global optimum, further approaching the optimal strategy, i.e., the optimal control action.
[0066] In this embodiment, the discount factor is adjusted based on the prediction error. The prediction error index of the state prediction model is calculated, and the discount factor is adjusted accordingly. When the prediction error exceeds a preset threshold, the discount factor is decreased; when it is less than or equal to the threshold, the discount factor is increased. This allows the model to focus more on long-term or short-term rewards, thus adaptively adjusting the learning strategy based on prediction accuracy. This adjustment mechanism helps incentivize the model to continuously improve the accuracy of the state prediction model, because the magnitude of the prediction error directly affects the adjustment of the discount factor, thereby influencing the model's learning and optimization direction.
[0067] In one embodiment of this application, multidimensional data is preprocessed to obtain target feature data, including: Missing data is filled by interpolation, and outlier data is filtered by threshold. Continuous data is standardized, and discrete data is encoded using one-hot encoding. The mutual information method is used to select features that are strongly correlated with compressor energy consumption and safety, and target feature data is generated.
[0068] In this embodiment, preprocessing requires linear interpolation to fill missing data, threshold filtering of abnormal data with reference to process safety boundaries, online Z-score standardization for continuous data, one-hot encoding for discrete data, and finally selection of features strongly correlated with compressor energy consumption and safety using mutual information method to generate target feature data.
[0069] In summary, this embodiment performs preprocessing operations on multidimensional data, including missing value imputation, outlier filtering, standardization, and one-hot encoding. This improves data quality and consistency, making subsequent feature selection and model training more effective. Secondly, by selecting features strongly correlated with compressor energy consumption and safety using the mutual information method to generate target feature data, data dimensionality is reduced, model complexity is lowered, and the focus is placed on the most important factors for compressor energy saving and safe operation, thereby improving model performance and efficiency.
[0070] In one embodiment of this application, target feature data is concatenated and fused with predicted values of running parameters to construct a state representation of the reinforcement learning model, including: The attention score of each feature in the target feature data is calculated based on the attention mechanism. The target feature data is weighted based on the attention score to obtain a weighted feature vector; The weighted feature vector is concatenated and fused with the predicted values of the running parameters to generate a state representation.
[0071] In this embodiment, the attention score of each feature in the target feature data is calculated based on a multi-head self-attention mechanism. The target feature data originates from core parameters that are strongly correlated with energy consumption and safety after multi-dimensional preprocessing, such as inhalation pressure. Shaft power (W), target flow rate Pipeline resistance coefficient Etc. This embodiment utilizes an attention mechanism to assign scores based on the correlation strength between various features and the real-time control objectives (energy saving, stability) of the compressor; for example, shaft power W, pipeline resistance coefficient, etc. The proportion of attention to features that significantly impact energy consumption optimization remained stable at 65% ± 5%, and the ambient humidity was also a factor. Low-correlation features should account for ≤10% to ensure that the core features have prominent weights.
[0072] The target feature data is weighted based on the attention score, amplifying the weight of features with high attention scores and suppressing the weight of features with low attention scores, resulting in a weighted feature vector that accurately reflects the key operating state and avoids irrelevant features interfering with model decision-making.
[0073] The weighted feature vector is concatenated and fused with the predicted values of operating parameters to obtain a weighted feature vector. The predicted values of operating parameters include predicted values of speed and exhaust pressure for multiple (3-7) control cycles in the future, as well as outputs based on the state prediction model.
[0074] The weighted feature vector corresponds to the number of target features, and the dimension of the predicted values of the running parameters matches the type of the core predicted parameters. After concatenation, the state representation of the reinforcement learning model is generated. This state representation retains the current key operating condition information and incorporates the prediction of future operating trends, providing a comprehensive and dynamic decision-making basis for the mixed actions output by the multi-head self-attention module in the hidden layer of the policy network, thus adapting to the real-time control requirements of the compressor to cope with sudden changes in operating conditions.
[0075] In this embodiment, the attention score of each feature in the target feature data is calculated based on the attention mechanism. This highlights features that are more important to the compressor's operating state and control, enabling the model to pay more attention to these key information and improve the accuracy of decision-making. Secondly, by concatenating and fusing the weighted feature vector with the predicted values of the operating parameters to generate the state representation of the reinforcement learning model, current feature information is combined with future prediction information, providing a more comprehensive and accurate state description for the reinforcement learning model, thereby helping the model make better control decisions.
[0076] In one embodiment of this application, the calculation of the attention score includes a feature grouping competition mechanism: The target feature data are divided into mechanical, thermal, and flow groups according to their physical attributes. Perform intra-group attention calculation within each group to obtain intra-group weighted feature vectors; The weighted feature vectors of each group are input into the inter-group attention layer to calculate the inter-group weights. The feature vectors of each group are merged based on the inter-group weights to generate a weighted feature vector.
[0077] In this embodiment, the target feature data is divided into three feature groups according to physical attributes: mechanical group, thermal group, and flow group; wherein, the mechanical group includes motor speed n, bearing vibration amplitude, etc. Shaft power (W) is used to reflect the core indicators of equipment mechanical operation and energy consumption; the thermal group includes: intake temperature. Exhaust temperature Ambient temperature This is used to correlate the thermodynamic efficiency of the compression process with the environmental coupling effect; the flow rate group includes: intake pressure Exhaust pressure Target traffic Pipeline resistance coefficient It is used to characterize the load demand and fluid dynamics of the pipeline network.
[0078] Within each group, intra-group attention is calculated. The correlation strength of features of the same dimension is quantified through an intra-group self-attention layer. The intra-group features are weighted according to the scores to obtain an intra-group weighted feature vector that highlights the core parameters of the group. The mutual information value is used to measure the dependency between the two.
[0079] Three sets of weighted feature vectors are input into the inter-group attention layer. Combined with the real-time compressor control objective, inter-group weights are calculated, with the dynamic weight allocation logic linked to the multi-objective reward function weight scheduling mechanism. For example, under low load, energy saving is prioritized, increasing the weight of the mechanical group; under high load, response is prioritized, increasing the weight of the flow group. For instance, under low load, the mechanical group has a weight of 0.4, the thermal group 0.2, and the flow group 0.4; under high load, the mechanical group has a weight of 0.3, the thermal group 0.2, and the flow group 0.5.
[0080] By fusing the feature vectors of each group according to the inter-group weights, the information gain of the high-weight recombination feature (mechanical group under low load) is amplified, while the low-weight recombination feature is appropriately suppressed, generating a weighted feature vector that takes into account key information from multiple physical dimensions.
[0081] In another possible embodiment of this application, the attention mechanism employs grouped attention computation: The target feature data is grouped according to physical attributes into a subset of operating status features, a subset of environmental features, and a subset of pipeline load features; Calculate the local attention score for each feature within each subset; The local attention scores of each subset are weighted and merged into a global attention score, the weight of which is determined by the correlation between the subset and the compressor energy-saving target.
[0082] In this embodiment, the target feature data is divided into three feature subsets according to physical attributes: operating status feature subset, environmental feature subset, and pipeline load feature subset. The operating status feature subset includes: intake pressure. Exhaust pressure Motor speed n, shaft power W, bearing vibration amplitude This directly reflects the compressor's operating efficiency and mechanical health, and is the core basis for energy consumption optimization; the environmental characteristic subset includes: ambient temperature. relative humidity Atmospheric pressure (Pa) indirectly changes the compression work by affecting gas density and moisture content, for example... For every 10°C increase, the gas density decreases by 3.5%. For every 10% increase, power consumption rises by 1.2%-1.5%; the pipeline load characteristic subset includes: target flow rate. Pipeline resistance coefficient Export pressure setpoint This characterizes downstream demand and the dynamic characteristics of the pipeline network. An annual increase of 3%-5% will directly lead to an increase in energy consumption.
[0083] Local attention scores are calculated for features within each subset. The correlation strength of features with the same attribute is quantified using a self-attention module within the subset. For example, the local attention ratio between shaft power W and rotational speed n in the operating status feature subset reaches 50% ± 5%, and the local attention ratio within the pipeline load feature subset is also high. Local attention accounts for over 40%, ensuring that the core feature weights are prominent within each subset.
[0084] The local attention scores of each subset are weighted and merged into a global attention score. The weights are dynamically determined based on the correlation between the subset and the compressor's energy-saving target. For example, under low load conditions, the weight of the operating status feature subset is set to 0.5, the weight of the pipeline load feature subset is 0.3, and the weight of the environmental feature subset is 0.2; under high load conditions, the weights of the operating status feature subset, pipeline load feature subset, and environmental feature subset are 0.4, 0.4, and 0.2.
[0085] This embodiment strengthens the contribution of energy-saving related features to global attention through weight allocation, and finally generates a global attention score that can accurately focus on energy-saving targets. This provides support for weighting target feature data and subsequently splicing it with predicted values of operating parameters to construct a state representation, adapting to the energy-saving optimization needs of compressors under different loads, environments, and pipeline conditions.
[0086] In one embodiment of this application, a state representation is generated by concatenating and fusing a weighted feature vector with predicted values of running parameters, including: A feature correction vector is generated based on the predicted values of operating parameters and the target feature data through cross-attention. The weighted feature vector and the correction vector are fused using a gating mechanism to generate a state representation.
[0087] In this embodiment, a cross-attention module is constructed based on the predicted values of operating parameters and the target feature data. By calculating the correlation strength between the predicted values and the target feature data, a feature correction vector is generated. This feature correction vector can compensate for future operating condition trend information not covered in the target feature data, and at the same time correct the deviation between the predicted values and the current operating conditions.
[0088] This embodiment employs a gating mechanism constructed using the Sigmoid activation function to input the weighted feature vector and the correction vector. This gating mechanism dynamically allocates fusion weights based on the compressor's real-time operating conditions. For example, when the operating conditions are stable, the weighted feature vector is set to 0.6 to emphasize the current stable conditions, while the correction vector is set to 0.4 to account for future trends. When the operating conditions change abruptly, the weighted feature vector is adjusted to 0.4, and the correction vector is increased to 0.6 to strengthen the support of future predictions for decision-making. This embodiment, through gating fusion, retains the reliability of current key operating condition information while incorporating the predictive value of future operating trends, ultimately generating a state representation that matches the input requirements of the reinforcement learning model.
[0089] In one specific embodiment, refer to Figure 2 A schematic diagram of the hardware composition of a compressor energy-saving operation control system based on reinforcement learning. The hardware architecture design of the compressor energy-saving operation control system is a three-layer structure consisting of a sensor layer, a control layer, and an execution layer, which is a closed-loop collaborative design of perception, decision-making, and execution. The sensor layer employs a high-precision sensing network for multiple physical quantities. Specifically, the pressure / temperature sensor uses a combination of diffused silicon piezoresistive type and platinum resistance PT100 to ensure that the error of the acquired data is ≤±0.0125MPa or ±0.75℃. The electromagnetic flowmeter is a Coriolis mass flowmeter, which achieves 1kHz high-frequency sampling through Modbus / TCP protocol, providing 50ms-level time accuracy for response reward items, and buffering 50 sampling points per control cycle to support dynamic load response. The vibration sensor is an IEPE type accelerometer, deployed in the orthogonal direction of bearing housing 3. In addition, this embodiment also monitors shaft vibration in real time, providing physical basis for stability rewards and safety constraints through millisecond-level early warning.
[0090] The data acquisition module in this embodiment employs an industrial-grade time-series data engine. It utilizes a 16-bit ADC synchronous acquisition module, constructing a dual anti-interference system of hardware filtering and software noise reduction. The sampling performance is 1kHz sampling frequency, synchronously capturing 13 state variables every 1ms, matched with a 50ms control cycle for reinforcement learning, extracting statistical features from 50 sampling points per cycle. The communication protocol uses Modbus / TCP with a transmission latency ≤20ms to ensure the real-time performance of the edge controller's state input. A built-in 100Hz hardware low-pass filter, combined with 3σ anomaly detection, suppresses sensor noise, reducing state fluctuations caused by sensor noise.
[0091] In this embodiment, the control layer serves as an edge intelligent decision-making hub. The edge computing controller, based on NVIDIA Jetson AGX Orin, performs model inference, with a single card supporting PPO model inference time ≤10ms. The remaining 40ms is used for data preprocessing and command issuance, meeting hard real-time control requirements. Safety redundancy is achieved through a HIMAF8650 safety module, enabling dual-controller hot backup with fault switching <20ms and control deviation ≤±2% of the set value during switching.
[0092] The execution layer in this embodiment is a high-precision motion execution unit, specifically including: a frequency converter, an electric regulating valve, and a fault shutdown solenoid valve; Among them, the frequency converter is the ABBACS880 series, which supports vector control of permanent magnet synchronous motors. The speed control accuracy is ±0.1% of the rated speed, and it can accurately execute the speed adjustment amount with a deviation of <0.5% from the continuous action output of reinforcement learning.
[0093] The electric control valve uses a Siemens SIPARTPS2 positioner with an opening resolution of 0.1% (adjustment step ≤ 0.05% / time), and receives signals via a 4-20mA input. The command has a response time of ≤200ms, matching the dynamic compensation requirements of pipeline resistance.
[0094] The fault-stop solenoid valve is a direct-acting solenoid valve with a response time of <10ms, used in a series safety circuit. The safety trigger condition is the breach of a level 3 constraint.
[0095] Corresponding to the reinforcement learning-based compressor energy-saving operation control method in the above embodiment, Figure 3 This is a structural block diagram of a compressor energy-saving operation control system based on reinforcement learning, provided as an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 3 The compressor energy-saving operation control system 20 based on reinforcement learning includes: a data acquisition module 21, a data processing module 22, a data prediction module 23, a splicing and fusion module 24, an action generation module 25, a safety inspection module 26, and an execution instruction module 27. The data acquisition module 21 is used to acquire multi-dimensional data during the operation of the compressor through a multi-parameter sensor network; the multi-dimensional data includes compressor operating status data, environmental data, and pipeline load data. Data processing module 22 is used to preprocess the multidimensional data to obtain target feature data; Data prediction module 23 is used to input the target feature data into the state prediction model to predict the predicted values of operating parameters within a future control cycle; The splicing and fusion module 24 is used to splice and fuse the target feature data with the predicted values of the running parameters to construct the state representation of the reinforcement learning model; Action generation module 25 is used to input the state representation into a target reinforcement learning model based on a proximal policy optimization framework to obtain the optimal control action; Safety verification module 26 is used to perform safety verification on the optimal control action based on preset compressor safety operation constraints and determine the execution instruction; The execution instruction module 27 is used to adjust the operating parameters of the compressor based on the execution instructions.
[0096] This embodiment may also include: The model update module 28 is used to update the target reinforcement learning model by using the environmental state when executing the execution instruction, the actual execution action, the environmental feedback reward, and the environmental state after executing the execution instruction as training samples.
[0097] See Figure 4 , Figure 4This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 4 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the aforementioned device embodiments, for example... Figure 3 The functions of the data acquisition module 21, data processing module 22, data prediction module 23, splicing and fusion module 24, action generation module 25, safety inspection module 26, execution instruction module 27, and model update module 28 are shown.
[0098] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0099] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.
[0100] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.
[0101] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in any embodiment of the compressor energy-saving operation control method based on reinforcement learning provided in the embodiments of this application, or they can execute the implementation methods of the electronic devices described in the embodiments of this application, which will not be repeated here.
[0102] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0103] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0104] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0105] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0106] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.
[0107] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0108] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0109] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A compressor energy-saving operation control method based on reinforcement learning, characterized in that, include: Multi-dimensional data during compressor operation are collected through a multi-parameter sensor network; The multidimensional data includes compressor operating status data, environmental data, and pipeline load data; The multidimensional data is preprocessed to obtain target feature data; The target feature data is input into the state prediction model to predict the operating parameter values for the next control cycle. The target feature data is concatenated and fused with the predicted values of the operating parameters to construct the state representation of the reinforcement learning model, including: Attention scores are calculated for each feature in the target feature data based on an attention mechanism; the target feature data is weighted according to the attention scores to obtain a weighted feature vector; the weighted feature vector is concatenated and fused with the predicted values of the running parameters to generate the state representation. The state representation is input into a target reinforcement learning model based on a proximal policy optimization framework to obtain the optimal control action. The optimal control action is safety-verified based on preset compressor safety operation constraints to determine the execution instructions, including: If the optimal control action passes the verification, then the optimal control action will be used as the execution instruction. If the optimal control action fails the verification, the optimal control action that violates the safety constraints is projected to the boundary point of the movable action space through the action projection algorithm to obtain the adjustment control action as the execution instruction. The compressor safety operation constraints include at least one of the following: safety threshold range of operating parameters, limit on the rate of change of operating parameters, and anti-surge logic; The action projection algorithm is used to project action vectors that violate constraints to the nearest boundary point in the action space. The compressor's operating parameters are adjusted based on the executed instructions.
2. The compressor energy-saving operation control method based on reinforcement learning according to claim 1, characterized in that, The training method for the target reinforcement learning model includes: A reinforcement learning state space is constructed based on historical feature data and their corresponding historical operational parameter predictions. Construct a hybrid action space, which includes a continuous action dimension and a discrete action dimension. The continuous action dimension includes the compressor speed regulation, the intake valve opening regulation, and the anti-surge return valve opening regulation. The discrete action dimension includes the start-stop state commands of one or more compressors in a multi-machine parallel operation. Design a policy network that supports mixed action outputs, wherein the hidden layer includes a multi-head self-attention module, the discrete action dimension output layer adopts a Softmax distribution, and the continuous action dimension output layer is parameterized with a Gaussian distribution; Based on the historical feature data, an energy-saving reward item, a stability reward item, and a response reward item are calculated to construct a multi-objective reward function; The state space, hybrid action space, multi-objective reward function, and experience pool sample pairs are used to train a reinforcement learning model using a proximal policy optimization algorithm, resulting in a well-trained target reinforcement learning model.
3. The compressor energy-saving operation control method based on reinforcement learning according to claim 2, characterized in that, The training method for the target reinforcement learning model further includes: Based on the state vector within a preset time window, the average Euclidean distance change rate relative to the mean state value within the window is calculated and divided by the historical maximum change rate, which serves as a volatility indicator. Adjust the upper limit parameter of the trust domain range of the near-end strategy optimization algorithm according to the volatility index; When the volatility indicator is less than the benchmark value, the upper limit parameter of the trust domain range is increased. When the volatility indicator is greater than the benchmark value, the upper limit parameter of the trust domain range is reduced. The benchmark value is the average rate of change of the Euclidean distance under historical stable operating conditions.
4. The compressor energy-saving operation control method based on reinforcement learning according to claim 3, characterized in that, The training method for the target reinforcement learning model further includes: Calculate the prediction error index of the state prediction model for the predicted values of operating parameters; the prediction error index is the weighted average of the absolute values of the differences between the predicted values of operating parameters and the actual values of operating parameters within a preset window; When the prediction error index is greater than the preset error threshold, the discount factor will be reduced by the first step length. When the prediction error index is less than or equal to the preset error threshold, the discount factor will be increased by a second step. Wherein, the first step length is less than the second step length, and the adjustment range of the discount factor is between a preset minimum value and a preset maximum value.
5. The compressor energy-saving operation control method based on reinforcement learning according to claim 1, wherein the preprocessing of the multidimensional data to obtain target feature data includes: Missing data is filled by interpolation, and outlier data is filtered by threshold. Continuous data is standardized, and discrete data is encoded using one-hot encoding. The mutual information method is used to select features that are strongly correlated with compressor energy consumption and safety, and target feature data is generated.
6. The compressor energy-saving operation control system based on reinforcement learning according to claim 1, characterized in that, include: The data acquisition module is used to collect multi-dimensional data during the operation of the compressor through a multi-parameter sensor network; The multidimensional data includes compressor operating status data, environmental data, and pipeline load data; The data processing module is used to preprocess the multidimensional data to obtain target feature data; The data prediction module is used to input the target feature data into the state prediction model to predict the predicted values of the operating parameters within a future control cycle. The splicing and fusion module is used to splice and fuse the target feature data with the predicted values of the running parameters to construct the state representation of the reinforcement learning model, including: Attention scores are calculated for each feature in the target feature data based on an attention mechanism; the target feature data is weighted according to the attention scores to obtain a weighted feature vector; the weighted feature vector is concatenated and fused with the predicted values of the running parameters to generate the state representation. The action generation module is used to input the state representation into a target reinforcement learning model based on a proximal policy optimization framework to obtain the optimal control action. The safety verification module is used to perform safety verification on the optimal control action based on preset compressor safety operation constraints, and to determine the execution instruction, including: If the optimal control action passes the verification, it is used as the execution instruction; if the optimal control action fails the verification, the optimal control action that violates the safety constraints is projected to the boundary point of the movable action space using the action projection algorithm to obtain the adjustment control action as the execution instruction; the compressor safety operation constraints include at least one of the following: safety threshold range of operating parameters, limit on the rate of change of operating parameters, and anti-surge logic; the action projection algorithm is used to project the action vector that violates the constraints to the nearest boundary point of the movable action space; The execution instruction module is used to adjust the compressor's operating parameters based on the execution instructions.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Compressor control method and system
CN119957473A
Compressor vibration self-adaptive suppression method and device based on reinforcement learning
CN119982484A