Compressor start-stop optimization method based on reinforcement learning
By constructing a compressor start-stop optimization model based on reinforcement learning, and combining it with the real-time operating conditions of the oil and gas storage and transportation pipeline network and engineering boundary conditions, the problems of accuracy and adaptability of compressor start-stop control were solved. This enabled adaptive and precise control of compressor start-stop and extended equipment life, reduced the number of start-stop cycles and energy consumption, and improved pipeline network operating efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XI'AN PETROLEUM UNIVERSITY
- Filing Date
- 2026-04-18
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies have failed to effectively address the accuracy and adaptability of compressor start-stop control in oil and gas storage and transportation pipelines, especially when facing dynamic nonlinear operating conditions, making it difficult to achieve multi-objective collaborative optimization.
A compressor start-stop optimization model is constructed by using a reinforcement learning-based approach and combining real-time operating parameters of oil and gas storage and transportation pipelines with engineering boundary conditions. By combining offline training with online feedback, a start-stop control strategy adapted to dynamic characteristics is constructed, and specific boundary conditions are set during the model training process to ensure safety and accuracy.
It achieves adaptive and precise control of compressor start-up and shutdown, significantly reducing the number of start-ups and shutdowns and unit energy consumption, extending equipment lifespan, and improving pipeline network operation efficiency and intelligence level.
Smart Images

Figure CN122082973A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of oil and gas storage and transportation engineering technology, and in particular to a compressor start-up and shutdown optimization method based on reinforcement learning. Background Technology
[0002] Compressors are the core power equipment in oil and gas storage and transportation pipeline networks, playing a crucial role in pressurizing and transporting oil and gas. The rationality of their start-stop control directly affects the operational economy, equipment lifespan, and transportation stability of the pipeline network. In recent years, the scale of my country's oil and gas storage and transportation pipeline network has continued to expand, and cross-regional, long-distance gas and oil transmission pipelines have become increasingly complex. Operating parameters such as flow rate, pressure, and temperature within the pipeline network exhibit dynamic and nonlinear changes due to factors such as upstream feed supply, downstream demand, ambient temperature, and pipeline leakage. This places higher demands on the accuracy and adaptive capability of compressor start-stop control.
[0003] Optimizing compressor start-up and shutdown using intelligent control methods is an important direction for the intelligent development of oil and gas storage and transportation engineering. Determining a reasonable compressor start-up and shutdown control strategy requires closed-loop management encompassing pipeline network operating parameter sensing, multi-objective optimization decision-making, and execution feedback correction. The core of this approach is establishing an optimization model adapted to the dynamic characteristics of the pipeline network, ensuring precise matching between start-up and shutdown decisions and real-time operating conditions.
[0004] Reinforcement learning, as an end-to-end intelligent optimization algorithm, achieves goal optimization through trial and error in the interaction between the agent and the environment. It possesses adaptive and self-learning characteristics and has been applied in fields such as industrial control and intelligent scheduling. Its core is to describe the interaction between the agent and the environment through Markov Decision Processes (MDPs) and to maximize rewards by optimizing policy functions. This characteristic is highly compatible with the dynamic operating conditions, multiple constraints, and multiple objectives of compressor start-up and shutdown in oil and gas storage and transportation pipeline networks. However, research on combining reinforcement learning algorithms with compressor start-up and shutdown control in oil and gas storage and transportation pipeline networks, improving algorithms for the dynamic nonlinear characteristics of pipeline networks, and constructing models based on actual engineering boundary conditions has not yet formed mature engineering application methods.
[0005] Therefore, how to construct a compressor start-stop closed-loop control method that can adapt to changes in operating conditions and achieve multi-objective collaborative optimization, based on fully considering the dynamic nonlinear characteristics of oil and gas storage and transportation pipelines, is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0006] This invention addresses the technical problems existing in the background art by proposing a compressor start-stop optimization method based on reinforcement learning. It integrates data acquisition and monitoring, reinforcement learning model training, start-stop decision execution, and online model updates to construct a start-stop optimization model adapted to the dynamic characteristics of oil and gas storage and transportation pipeline networks. This method, based on reinforcement learning algorithms and combined with real-time pipeline network operating parameters and engineering boundary conditions, achieves adaptive and precise start-stop of the compressor. By combining offline model training with online operating condition feedback, the optimal start-stop strategy under different operating conditions is determined, clarifying the applicable scenarios and parameter adjustment methods of the optimization model, thus laying the foundation for intelligent control of compressors in oil and gas storage and transportation pipeline networks.
[0007] To solve the technical problem, the technical solution of the present invention is as follows:
[0008] A compressor start-stop optimization method based on reinforcement learning, the method comprising:
[0009] Real-time operating parameters of oil and gas storage and transportation pipelines and compressors are collected, preprocessed to obtain standardized operating condition data, and then the standardized operating condition data is split into an offline training dataset and a real-time inference data stream.
[0010] A Markov decision process model is constructed, defining a state space, an action space, and a multi-objective weighted immediate reward function, and setting specific boundary conditions, including physical condition boundaries, state space and action space boundaries, model training constraint boundaries, and decision execution hard limit boundaries.
[0011] The Proximal Policy Optimization (PPO) algorithm is used to perform offline training on the offline training dataset under the constraints of the specific boundary conditions to obtain a pre-trained optimized model.
[0012] The real-time inference data stream is input into the pre-trained optimization model for online inference, and a start / stop decision is output. After the start / stop decision is subjected to hard limit boundary verification, the compressor is controlled to perform the corresponding action.
[0013] As can be understood, this invention, based on reinforcement learning algorithms, combines real-time operating parameters of oil and gas storage and transportation pipeline networks with actual engineering boundary conditions to construct a multi-objective optimization model for compressor start-up and shutdown. This allows compressor start-up and shutdown decisions to return to the essence of adapting to dynamic pipeline operating conditions and complying with engineering constraints. By combining offline model training with online operating condition feedback, the impact of different methods such as reinforcement learning optimization, threshold control, and fuzzy control on the compressor start-up and shutdown control effect is determined. The applicable scenarios, parameter adjustment methods, and optimization strategies of the reinforcement learning-based compressor start-up and shutdown optimization model are clarified, providing a robust technical solution for intelligent start-up and shutdown control of compressors in oil and gas storage and transportation pipeline networks.
[0014] Furthermore, the method also includes: collecting working condition data after decision execution, evaluating model decision deviation, and when the deviation exceeds a preset threshold, incrementally fine-tuning the pre-trained optimization model under the constraints of the model training constraint boundary to form a closed-loop correction.
[0015] Furthermore, the construction of the Markov decision process model, defining the state space, action space, and multi-objective weighted immediate reward function, specifically includes:
[0016] The compressor start-stop control scenario is fitted as a Markov decision process, and the mathematical expression is a quintuple:
[0017]
[0018] In the formula, For state space, For the action space, Let be the state transition probability. For instant reward function, As a discount factor, For Markov decision process models;
[0019] The state space It is a high-dimensional continuous space, composed of key parameters of the pipeline network and compressor, and its mathematical expression is:
[0020]
[0021] In the formula, The average pressure at the inlet and outlet of the pipeline compressor. To deliver flow to the pipeline network, For the temperature of the medium in the pipeline network, Real-time energy consumption of the compressor. This represents the vibration value of the compressor body. This refers to the cumulative operating time of the compressor. Represents the n-dimensional real space;
[0022] The action space For a discrete finite set, the mathematical expression is:
[0023]
[0024] The multi-objective weighted instantaneous reward function The mathematical expression is:
[0025]
[0026] In the formula, This indicates the agent's current state. Next action Then, the environmental state transitions to the next state. At that time, the instantaneous reward value that the agent obtains from the environment; This is the current state, belonging to the state space. ; The action to be performed belongs to the action space. ; The next state after performing an action satisfies the state transition probability. ,in, Indicates time The state is given by P(⋅∣⋅), which represents the conditional probability. Indicates time The action;
[0027] These are weighting coefficients. This is an energy consumption reward item, negatively correlated with the real-time energy consumption of the compressor. This is a bonus item for equipment wear and tear, negatively correlated with the number of compressor start-ups and shutdowns and vibration levels. This is a pipeline stability incentive item, negatively correlated with pipeline pressure and flow fluctuation rates. For the first Weighting coefficients for each reward item; For indexing, These correspond to energy consumption incentives, equipment loss incentives, and pipeline stability incentives, respectively.
[0028] Furthermore, the optimization objective of the Markov decision process model is to maximize the agent's cumulative discounted reward from the initial state, mathematically expressed as:
[0029]
[0030] In the formula, For strategy The expected value of the cumulative discount reward under the following conditions The policy of the agent is defined as follows: , indicating the state Select action The probability of the strategy is a random strategy; Indicating in strategy The expected value of the following; Discount factor of Power; This is an instant reward function.
[0031] Furthermore, the specific boundary conditions include:
[0032] Physical operating condition boundaries are the inherent safe operating ranges for pipeline networks and compressors in engineering design. They are used to identify and eliminate outliers during the data acquisition phase and to trigger boundary pre-protection during the model inference phase.
[0033] The boundary between the state space and the action space is determined by linearly normalizing the state space parameters based on the physical condition boundary, thus limiting the input of the reinforcement learning optimization model to a uniform numerical range and constraining the action space of the model output to be a discrete finite set.
[0034] The model training constraint boundary refers to the numerical constraints followed by the reinforcement learning optimization model during the offline training and online fine-tuning stages, including the learning rate boundary, the policy network parameter update magnitude boundary, and the cumulative discount reward convergence boundary.
[0035] The decision execution hard limit boundary serves as a safety fallback boundary during the compressor start-stop decision execution phase. The priority of the decision execution hard limit boundary is higher than the decision output of the reinforcement learning optimization model. When the real-time monitored pipeline or compressor parameters exceed the hard limit boundary, an emergency start-stop command is directly triggered, and the model inference results are masked.
[0036] Furthermore, the mathematical expression for the physical operating condition boundary is a range of parameter values, and the core boundary includes:
[0037]
[0038] In the formula, As the pressure boundary of the pipeline network, For pipeline flow boundary, For the temperature boundary of the medium, For the compressor energy consumption boundary, For compressor vibration boundary, This refers to the compressor speed boundary; The minimum / maximum operating pressure allowed by the pipeline network; The minimum / maximum allowable flow rate of the pipeline network; The minimum / maximum operating temperature allowed by the medium; This represents the lower / upper limit of normal fluctuations in compressor energy consumption. The minimum / maximum allowable vibration amplitude of the compressor; The minimum / maximum permissible speed of the compressor;
[0039] The normalized mathematical expression for the state space parameters at the boundary between the state space and the action space is:
[0040]
[0041] In the formula, For the first in the state space The normalized values of the parameters, and This represents the physical boundary extreme value of the parameter. For the first The raw real-time measurement values of each parameter;
[0042] The mathematical expression for the model training constraint boundary is:
[0043]
[0044] In the formula, The learning rate boundary; To update the magnitude boundaries of the policy network parameters, This is the preset maximum update range for a single operation; To accumulate the expected value of the discount reward, and This is the preset convergence boundary;
[0045] The constraint logic of the model training constraint boundary is as follows: During model training, if the learning rate... or parameter update magnitude If the value exceeds the corresponding boundary, parameter correction will be automatically triggered; if the expected value of the accumulated discount reward is... Failed to converge to the preset boundary If the condition is met, continue training until it is met.
[0046] The mathematical expression for the hard boundary of the decision execution is:
[0047]
[0048] in, This is represented as the lower limit threshold of emergency pressure. This is represented as the upper limit threshold for emergency energy consumption. This is represented as the upper limit threshold for emergency vibration.
[0049] Furthermore, the offline training using the near-end strategy optimization algorithm specifically includes:
[0050] Set the pruning factor for the near-end policy optimization algorithm. The policy update magnitude is limited by a pruning objective function, the mathematical expression of which is:
[0051]
[0052] In the formula: The ratio of policy updates, i.e., the ratio of the new policy to the old policy in the state. Select Action The probability ratio; Indicates the expectation of a time step; This is the current strategy; This is the old strategy; The advantage function estimate represents the action. The relative merits of the average strategy; For the clipping function, Limited to Within the scope; the model training process is to minimize the pruning objective function. The process continues until the function converges to a preset threshold, yielding the optimal policy network parameters. .
[0053] Furthermore, the evaluation model's decision bias involves incrementally fine-tuning the pre-trained optimization model when the bias exceeds a preset threshold, specifically including:
[0054] The evaluation indicators for the model optimization effect are set, including the number of compressor start-stop cycles, unit delivery energy consumption, and pipeline pressure fluctuation rate.
[0055] The actual performance metrics in the feedback dataset are compared with the predicted metrics of the pre-trained optimization model to calculate the decision bias rate.
[0056] When the decision bias rate exceeds the preset threshold, the model is determined to have a working condition drift. The feedback dataset is used as incremental training data, and the policy network parameters of the pre-trained optimization model are fine-tuned in small batches under the constraints of the model training constraint boundary. The fine-tuned model is then redeployed to the online inference unit.
[0057] This application has the following advantages:
[0058] This invention constructs a lightweight PPO reinforcement learning model with compressor- and pipeline-specific boundary constraints, establishes a closed-loop intelligent control system encompassing "acquisition-inference-execution-feedback-correction," and employs an incremental online update method to adapt to operating condition drift, achieving adaptive and precise control of compressor start-up and shutdown in oil and gas storage and transportation pipeline networks. Compared to traditional methods, this invention significantly reduces the number of compressor start-ups and shutdowns and unit delivery energy consumption, extending equipment lifespan. Simultaneously, the full-process boundary constraints ensure decision-making safety, the system has low modification costs, strong generalization capabilities, and can adapt to various pipeline network operating conditions, effectively improving the overall operating efficiency and intelligence level of the pipeline network. Attached Figure Description
[0059] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0060] Figure 1This is a simplified process flow diagram of the compressor start-stop optimization system based on reinforcement learning proposed in an embodiment of this application;
[0061] Figure 2 This is a structural diagram of the near-end strategy optimization algorithm framework proposed in the embodiments of this application;
[0062] Figure 3 This application proposes a technical roadmap for a compressor start-stop optimization method based on the reinforcement learning PPO algorithm in its embodiments.
[0063] Figure 4 This is a comparison chart of the number of compressor start-stop cycles proposed in the embodiments of this application.
[0064] Figure label:
[0065] 1-Sensor group; 2-Data transmission module; 3-Data preprocessing unit; 4-Reinforcement learning model training server; 5-Offline training dataset; 6-Online inference unit; 7-Start-stop decision execution module; 8-Compressor control system; 9-Compressor body; 10-Model performance evaluation unit; 11-Online model parameter update unit; 12-Pipeline network operating condition feedback module. Detailed Implementation
[0066] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0067] Example 1:
[0068] This embodiment first provides a compressor start-stop optimization system based on reinforcement learning, comprising four core parts: a data acquisition and monitoring system, a reinforcement learning optimization model training module, a start-stop decision execution module, and a model online update and correction module, forming a closed-loop optimization system of "data acquisition - model inference - decision execution - feedback correction". A simplified process flow diagram of this embodiment is shown below. Figure 1As shown, sensor group 1 collects real-time operating parameters of the oil and gas storage and transportation pipeline network and the compressor body, and transmits them to data preprocessing unit 3 through data transmission module 2. Data preprocessing unit 3 divides the processed standard operating condition data into two paths: one path is sent to online inference unit 6 for real-time decision-making, and the other path is stored in offline training dataset 5. Reinforcement learning model training server 4 reads offline training dataset 5 to train the model, and deploys the trained pre-trained optimized model to online inference unit 6. Online inference unit 6 outputs start-stop decision commands based on real-time data, sends them to start-stop decision execution module 7, and then drives compressor body 9 to perform corresponding actions through compressor control system 8. At the same time, the actual operating parameters of compressor body 9 and pipeline network are fed back to model effect evaluation unit 10 through pipeline network operating condition feedback module 12. Model effect evaluation unit 10 compares the deviation between actual operating effect and model prediction result, triggers model parameter online update unit 11 to fine-tune the model, and sends the updated model parameters back to online inference unit 6 to form closed-loop optimization.
[0069] For example, the data acquisition and monitoring system is the perception layer of the optimization model. It is used to collect real-time operating parameters of key nodes in the oil and gas storage and transportation pipeline network and the compressor body 9, providing a data foundation for model inference and training. It consists of sensor group 1, data transmission module 2 and data preprocessing unit 3.
[0070] Sensor group 1: Pressure, flow and temperature sensors are installed at the compressor inlet and outlet, key distribution nodes and pressure regulating stations in the pipeline network. Speed, vibration, energy consumption and running time sensors are installed on the compressor body 9 to realize real-time acquisition of pipeline network operating parameters and compressor operating parameters in all dimensions. The acquisition frequency can be adjusted according to the characteristics of pipeline network operating conditions.
[0071] Data transmission module 2: Adopting a dual-mode transmission method of industrial Ethernet + 5G, it transmits the raw data collected by sensor group 1 to the data preprocessing unit to ensure the real-time performance and stability of data transmission, and adapts to the field deployment requirements of long-distance oil and gas storage and transportation pipelines.
[0072] Data preprocessing unit 3: performs noise reduction, normalization, outlier removal, and data completion on the raw data to eliminate data distortion caused by sensor errors, transmission interference, and other factors. The processed standardized data is divided into two paths: one path is transmitted to the online inference unit 6 of the reinforcement learning model for real-time start and stop decisions; the other path is stored in the offline training dataset 5 for offline training and online fine-tuning of the model.
[0073] For example, the reinforcement learning optimization model training module is used to construct a compressor start-up and shutdown optimization model adapted to the dynamic characteristics of oil and gas storage and transportation pipeline networks and with boundary constraints. It consists of an offline training dataset (5), a reinforcement learning model training server (4), and a model algorithm library, enabling offline training and online initialization of the model. The core elements of the reinforcement learning model are defined through precise mathematical formulas, and specific boundary conditions under the compressor environment are clearly defined to ensure the theoretical rigor and engineering practicality of the model.
[0074] Example 2:
[0075] This embodiment applies to Embodiment 1 and provides a compressor start-stop optimization method based on reinforcement learning. The method specifically includes the following steps:
[0076] Step 1: Core mathematical modeling of reinforcement learning models (Markov Decision Process MDP);
[0077] The dynamic characteristics of compressor start-stop control scenarios in oil and gas storage and transportation pipeline networks can be accurately fitted to a Markov decision process. Its core is the interaction between the agent and the environment, satisfying the Markov property: the next state is only related to the current state and actions, and is independent of history. The mathematical expression is a quintuple:
[0078] (1)
[0079] In the formula: S is the state space, a high-dimensional vector composed of preprocessed pipeline operating parameters and compressor operating parameters, which is the input of the model and is constrained by the compressor environmental boundary conditions; A is the action space, the set of compressor start-stop decisions, which is a discrete finite set and is constrained by the engineering operation boundary. Let be the state transition probability, representing the agent's state transition probability. Execute action After that, the transition state The probability of this is determined by the dynamic nonlinear characteristics of the pipeline network. ; Let be the immediate reward function, representing the transition of an agent from state s to state s by performing action a. The reward value obtained is a multi-objective weighted form, and the calculation of the reward value is constrained by boundary conditions. The discount factor, ranging from [0,1], represents the degree of importance attached to future rewards. This application considers both short-term operating condition response and long-term optimization objectives, and takes... .
[0080] This application employs a model-free reinforcement learning approach to handle the state transition probability P, eliminating the need to pre-establish or identify precise mathematical expressions for pipeline network state transitions. Through online interactive sampling between the agent and the virtual pipeline environment, the policy gradient and advantage function are directly estimated using empirical data. This avoids the problems of difficult accurate modeling of high-dimensional nonlinear pipeline systems and the difficulty in analytically solving state transition probabilities, making the model more adaptable to the dynamic uncertainty characteristics of oil and gas storage and transportation pipeline networks.
[0081] The optimization objective of a reinforcement learning model is to maximize the agent's cumulative discounted reward from the initial state, expressed mathematically as:
[0082] (2)
[0083] In the formula: The agent's strategy, i.e. Eπ represents the probability that the agent chooses action a in state s. In this application, the policy is a stochastic policy to adapt to the randomness of pipeline network conditions; Eπ represents the mathematical expectation under policy π.
[0084] Step 2: Mathematical definition of the core elements of the model;
[0085] For the practical engineering considerations of compressor start-up and shutdown control in oil and gas storage and transportation pipeline networks, the core elements of the MDP quintuple are defined mathematically and linked to subsequent boundary conditions to form a constraint relationship:
[0086] State space S;
[0087] The state space is a high-dimensional continuous space, consisting of n key parameters of the pipeline network and compressor, and its mathematical expression is:
[0088] (3)
[0089] In the formula: p is the average inlet and outlet pressure of the compressor in the pipeline network, q is the pipeline flow rate, T is the temperature of the medium in the pipeline network, E is the real-time energy consumption of the compressor, and v is the vibration value of the compressor body. This refers to the cumulative operating time of the compressor. It represents an n-dimensional real number space; in this application, n=6 is taken, and the parameter dimension can be expanded according to the actual needs of the pipeline network.
[0090] Action space A;
[0091] The action space is a discrete finite set, adapted to the engineering operation logic of compressor start-up and shutdown, and its mathematical expression is:
[0092] (4)
[0093] The actions are mutually exclusive discrete values with no intermediate transition states, which conforms to the actual operating rules of the compressor.
[0094] Multi-objective weighted instantaneous reward function ;
[0095] Constructing a multi-objective weighted reward function that takes into account energy consumption costs, equipment losses, and pipeline stability is the core guiding principle for model optimization. Its mathematical expression is:
[0096] (5)
[0097] In the formula: The weighting coefficient can be dynamically adjusted according to the pipeline network operation requirements (energy saving priority / equipment protection priority / stability priority);
[0098] Energy consumption bonus The reward value is negatively correlated with the real-time energy consumption of the compressor; the lower the energy consumption, the higher the reward value. The mathematical expression is: This represents the energy consumption boundary of the compressor.
[0099] Equipment Loss Rewards The reward value is negatively correlated with the number of compressor start-stop cycles and vibration levels; the lower the loss, the higher the reward value. The mathematical expression is: v is the real-time vibration value of the compressor. max The upper limit boundary for vibration is given by N, k is the start / stop penalty coefficient, and N is the upper limit boundary for vibration. start The number of starts and stops per unit time;
[0100] Pipeline stability bonus It is negatively correlated with the fluctuation rate of pipeline pressure / flow; the more stable it is, the higher the reward value. The mathematical expression is: , These are the fluctuation rates of pressure and flow rate, respectively, and their values are constrained by the boundary conditions of the pipeline network.
[0101] The objective function of the PPO algorithm;
[0102] This embodiment selects the PPO algorithm, adapted for industrial real-time control, as the basic algorithm. It achieves the goal of selecting the optimal action in the environment to maximize cumulative reward by iteratively optimizing the agent's policy. The policy is a mapping that defines how the agent selects actions given a state. In the PPO algorithm, the agent's policy is represented by a parameterized probability distribution, denoted as […]. Where θ is the policy parameter and s represents the current state; the basic flowchart of the algorithm is as follows: Figure 2 As shown.
[0103] To address the high-dimensional and nonlinear characteristics of pipeline network operations, a lightweight improvement is implemented. The PPO algorithm avoids excessive policy update amplitude through a pruned objective function, thereby enhancing the stability of model training. Its core mathematical expression is:
[0104] (6)
[0105] In the formula: The parameters of the policy network are the optimization variables for model training; The ratio of policy updates, i.e., the ratio of the new policy to the old policy in the state. Select Action The probability ratio; The advantage function estimate represents the action. The relative merits of the average strategy; The cutting factor is taken as the cutting factor in this application. =0.2, limiting the policy update range to within ±20%; For the clipping function, Limited to Within the scope, avoid policy mutations.
[0106] The process of model training is to minimize the pruning objective function. The process continues until the function converges to a preset threshold, thus obtaining the optimal policy network parameters θ.
[0107] Step 3: Define specific boundary conditions for the compressor environment;
[0108] This embodiment addresses the practical engineering challenges of compressors in oil and gas storage and transportation pipeline networks. It establishes specific boundary conditions for model training and decision execution. All model inputs, inferences, and decisions are strictly constrained by these boundaries to prevent model outputs from deviating from engineering safety limits. The boundary conditions are categorized into four main types: physical operating condition boundaries, state / action space boundaries, model training constraint boundaries, and decision execution hard limits, covering the entire process from "data input to model training to decision execution." The mathematical definitions and physical meanings of each boundary condition are as follows:
[0109] The physical operating condition boundary is the basic boundary, including the inherent safe operating range of the pipeline network and compressor;
[0110] These boundaries are inherent engineering design boundaries for oil and gas storage and transportation pipelines and compressors, determined by equipment manufacturers and pipeline design specifications. They serve as the fundamental constraints for all parameters, and their mathematical expressions represent the range of parameter values. The core boundaries are as follows:
[0111] (7)
[0112] In the formula: p is the pipeline pressure boundary; q is the pipeline flow boundary; T is the medium temperature boundary; E is the compressor energy consumption boundary; v is the compressor vibration boundary; n is the compressor speed boundary.
[0113] Constraint logic: During the data acquisition phase, parameters exceeding the boundary are judged as outliers and are directly removed; during the model inference phase, if the real-time parameters are close to the boundary threshold, the model prioritizes outputting the "start compressor" action to achieve boundary pre-protection.
[0114] State / action space boundaries, i.e. model input / output constraints;
[0115] These boundaries serve as input and output constraints for reinforcement learning models. They are normalized based on physical condition boundaries to ensure the model inputs are within a uniform numerical range and the outputs conform to engineering operational logic. The mathematical expression is:
[0116] (8)
[0117] In the formula: Let be the normalized value of the i-th parameter in the state space. This represents the physical boundary extreme value of the parameter.
[0118] Constraint logic: In the data preprocessing stage, all state parameters are mapped to the [0,1] interval to avoid the impact of dimensional differences on model training; in the model output stage, only one action can be selected from the discrete action set, and there is no invalid output.
[0119] Model training constraint boundaries are the numerical constraints in the model training process.
[0120] These boundaries serve as numerical constraints during the training phase of reinforcement learning models, preventing issues such as gradient explosion and policy divergence during training and ensuring training stability. The core boundaries are as follows:
[0121] (9)
[0122] In the formula: .
[0123] Constraint logic: During model training, if the learning rate or parameter update magnitude exceeds the boundary, parameter correction is automatically triggered; if the cumulative reward does not converge to the preset boundary, training continues until the condition is met.
[0124] The hard limit boundary for decision execution, i.e., the safety net boundary for the final decision;
[0125] This type of boundary serves as a safety fallback boundary during the compressor decision-making and execution phase. It has a higher priority than model-based decisions. When pipeline and compressor parameters exceed this boundary, an emergency start / stop command is directly triggered without model reasoning. It acts as the last line of defense for equipment and the pipeline network. The mathematical expression is:
[0126] (10)
[0127] Constraint logic: Hard boundary conditions are embedded in the start / stop decision execution module, which monitors parameters in real time. Once triggered, the model decision is directly blocked, and an emergency control signal is output to ensure the safe operation of equipment and pipelines.
[0128] Step 4: Offline training of the model;
[0129] Historical operating condition data from the offline training dataset is input into the reinforcement learning model training server. Model hyperparameters, including the number of training iterations, learning rate, and discount factor, are set according to the aforementioned mathematical formulas and boundary conditions. The virtual pipeline environment is constructed based on the oil and gas storage and transportation pipeline mechanism model and noise injection method: First, steady-state mechanism equations for the pipeline network are established based on the pipeline topology, pipe diameter, length, medium properties, and compressor performance curves. These equations include pressure-flow coupling equations, energy conservation equations, and compressor characteristic equations, enabling the mechanism deduction of operating condition parameters. Then, Gaussian white noise and operating condition disturbance noise are superimposed on the mechanism output to simulate random disturbances such as on-site transmission interference, downstream energy fluctuations, and ambient temperature changes, ensuring the virtual environment matches the actual on-site operating condition distribution. The virtual environment can quickly output the next state based on real-time state input, providing a stable, safe, and efficient interactive training scenario for the reinforcement learning agent, avoiding the safety risks and cost losses associated with training with physical equipment. The model is trained through trial and error between the agent and the virtual network environment with boundary constraints until the cumulative reward function J(π) of the model converges to the preset boundary and the policy network parameter θ is stable. The pre-trained optimized model is then deployed to the online inference unit.
[0130] For example, the start-stop decision execution module is the execution layer of the optimization model, consisting of an online inference unit 6, a start-stop decision execution module 7, and a compressor control system 8. It realizes the real-time decision-making and execution of the start-stop of the compressor body 9, and all execution processes strictly follow the decision execution hard limit boundary.
[0131] Online inference: The online inference unit 6 receives real-time standardized operating condition data mapped to the boundary of the state space from the data preprocessing unit, inputs it into the pre-trained optimization model, and the model outputs the optimal action space A, i.e., the start and stop decision, based on the real-time state space S. The inference process takes less than 1 second, which meets the real-time control requirements of industrial sites.
[0132] Decision execution: The start / stop decision execution module 7 converts the optimal decision output by the model into an industrial control signal, a 4~20mA analog quantity / Modbus digital quantity, and verifies the hard limit boundary in real time. If the parameter does not exceed the hard limit boundary, the control signal is transmitted to the original control system of the compressor body 9; if the parameter exceeds the hard limit boundary, the emergency control signal is directly triggered.
[0133] Command execution: The compressor control system 8 executes the operation of starting, stopping or maintaining the current state of the compressor according to the received control signals, so as to realize intelligent and safe control of the start and stop of the compressor body 9. The control signals include model decision / emergency commands.
[0134] For example, the online model update and correction module is the self-learning layer of this application, which consists of the pipeline network operating condition feedback module 12, the model effect evaluation unit 10, and the online model parameter update unit 11. It realizes the closed-loop correction of the model, ensures the optimization accuracy of the model in scenarios such as operating condition drift and pipeline network transformation, and the update process still follows the model training constraint boundary.
[0135] Operating condition feedback: The pipeline operating condition feedback module 12 collects pipeline operating condition parameters and compressor operating parameters in real time after the decision is executed, compares the parameter changes before and after the decision is executed, forms a feedback dataset, and maps the data to the physical operating condition boundary.
[0136] Performance Evaluation: The model performance evaluation unit 10 sets evaluation indicators for the model optimization effect, such as the number of start-stop cycles, unit energy consumption, and pipeline pressure fluctuation rate. It compares the actual indicators in the feedback dataset with the model's predicted indicators to calculate the decision deviation rate. When the deviation rate exceeds a preset threshold of 5%, the threshold is determined based on field operation statistics and the 95th percentile method: through distribution statistics of historical operating data, when the model decision deviation is less than 5%, the pipeline pressure, flow rate, and energy consumption indicators are all within a stable range; exceeding 5% indicates that the degree of operating condition drift has affected the control effect, requiring model correction. This threshold balances control accuracy and update frequency, and is the optimal value determined jointly by engineering experience and statistical analysis. If operating condition drift is detected in the model, the online model update process is triggered.
[0137] Online Update: The online model parameter update unit uses the feedback dataset as incremental training data to perform small-batch fine-tuning on the pre-trained optimized model, updating the model's policy network parameters θ. The fine-tuning process strictly follows the model training constraints, namely the learning rate and parameter update magnitude, without requiring full offline retraining, ensuring the efficiency of the update process. The fine-tuned model is redeployed to the online inference unit to achieve adaptive correction of the model, forming a closed-loop optimization of "decision-feedback-correction".
[0138] like Figure 3 As shown, this embodiment also provides the following overall technical approach:
[0139] 1. Collect historical / real-time operating data of pipeline network and compressor, preprocess it and map it to preset boundary conditions to build offline training dataset and real-time inference dataset;
[0140] 2. Construct a multi-objective reinforcement learning optimization model with boundary constraints based on the PPO algorithm. The core elements of the model are precisely defined by mathematical formulas. The model training constraint boundary is strictly followed during offline training to obtain a pre-trained model and deploy it to the online inference unit.
[0141] 3. Input real-time operating condition data that conforms to the state space boundary into the pre-trained model, output the optimal start-stop decision, verify the hard boundary of decision execution, and transmit the legal control signal to the compressor control system for execution;
[0142] 4. Provide real-time feedback on working condition data that conforms to the physical working condition boundary after the decision is executed, evaluate the model decision deviation, and when the deviation exceeds the threshold, perform small-batch fine-tuning of the model within the model training constraint boundary;
[0143] 5. Repeat steps 3-4 to achieve adaptive, precise, and safe control of compressor start-up and shutdown, while continuously accumulating interactive data and optimizing model performance.
[0144] Example 3:
[0145] This example provides an application scenario for a compressor start-stop optimization method based on reinforcement learning. Focusing on compressor start-stop optimization, an intelligent control scheme is designed based on a near-end policy optimization algorithm, and its superiority over traditional constant pressure control is verified through comparative experiments. The core objective is to reduce the number of start-stop cycles and energy consumption during compressor operation. The specific implementation process and results are as follows:
[0146] Experimental scenario and parameter settings;
[0147] The experimental scenario used a pipeline compressor operating continuously for 7 days (168 hours). The core parameters were set as follows: target operating pressure of 0.8 MPa, safe operating range of 0.75 MPa to 0.85 MPa; basic energy consumption of 50 kW, and upper limit of 100 kW. The near-end policy optimization algorithm parameters were set as follows: learning rate 4e-4, discount factor 0.96, pruning factor 0.2, and network width 64. The training configuration was: 600 epochs of offline training, 200 steps per epoch, and an online update frequency of once every 50 steps.
[0148] Core control logic;
[0149] Near-end strategy optimization intelligent control: The model learns the optimal start-stop strategy through offline training. When running online, it dynamically decides the compressor's actions based on real-time pressure, flow, energy consumption and other statuses. The actions include maintaining operation, increasing pressure, and decreasing pressure, with priority given to operations with low energy consumption and fewer start-stop cycles.
[0150] Traditional constant pressure control: It adopts a fixed threshold start-stop logic, that is, the pressure increases when it is below 0.78MPa and decreases when it is above 0.82MPa. It only triggers the action based on the pressure threshold, without any consideration for energy consumption and the number of start-stop cycles.
[0151] Experimental verification and visualization;
[0152] The compressor start-stop count comparison chart was obtained after programming and training using MATLAB software, as shown below. Figure 4 As shown. By Figure 4 It can be seen that after optimization control using reinforcement learning's proximal strategy, the number of compressor start-stop cycles is significantly reduced compared to the traditional control mode. Furthermore, as the operating time increases, the fluctuation in the number of start-stop cycles becomes smaller, and the curve becomes smoother. The reduction in the number of start-stop cycles further reduces the compressor's energy consumption, thereby improving the pipeline's transport efficiency.
[0153] Through simulation comparison, the PPO optimized control of this invention has significant advantages over traditional constant pressure control in three aspects: compressor unit delivery energy consumption is reduced by 12.3%, pipeline pressure fluctuation rate is reduced from ±3.2% to ±1.1%, and the average number of compressor start-ups and shutdowns per day is reduced by 68.5%. It has significant advantages in energy saving, equipment protection and pipeline stability.
[0154] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0155] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A compressor start-stop optimization method based on reinforcement learning, characterized in that, The method includes: Real-time operating parameters of oil and gas storage and transportation pipelines and compressors are collected, preprocessed to obtain standardized operating condition data, and then the standardized operating condition data is split into an offline training dataset and a real-time inference data stream. A Markov decision process model is constructed, defining the state space, action space, and multi-objective weighted immediate reward function, and setting specific boundary conditions, including physical condition boundaries, state space and action space boundaries, model training constraint boundaries, and decision execution hard limit boundaries. The near-end strategy optimization algorithm is adopted, and the offline training dataset is used to perform offline training under the constraints of the specific boundary conditions to obtain a pre-trained optimized model; The real-time inference data stream is input into the pre-trained optimization model for online inference, and a start / stop decision is output. After the start / stop decision is subjected to hard limit boundary verification, the compressor is controlled to perform the corresponding action.
2. The compressor start-stop optimization method based on reinforcement learning according to claim 1, characterized in that, The method further includes: collecting working condition data after decision execution, evaluating model decision deviation, and when the deviation exceeds a preset threshold, incrementally fine-tuning the pre-trained optimization model under the constraints of the model training constraint boundary to form a closed-loop correction.
3. The compressor start-stop optimization method based on reinforcement learning according to claim 1, characterized in that, The construction of the Markov decision process model, defining the state space, action space, and multi-objective weighted immediate reward function, specifically includes: The compressor start-stop control scenario is fitted as a Markov decision process, and the mathematical expression is a quintuple: , In the formula, For state space, For the action space, Let be the state transition probability. For instant reward function, As a discount factor, For Markov decision process models; The state space It is a high-dimensional continuous space, composed of key parameters of the pipeline network and compressor, and its mathematical expression is: , In the formula, The average pressure at the inlet and outlet of the pipeline compressor. To deliver flow to the pipeline network, For the temperature of the medium in the pipeline network, Real-time energy consumption of the compressor. This represents the vibration value of the compressor body. This refers to the cumulative operating time of the compressor. Represents the n-dimensional real space; The action space For a discrete finite set, the mathematical expression is: , The multi-objective weighted instantaneous reward function The mathematical expression is: , In the formula, This indicates the agent's current state. Next action Then, the environmental state transitions to the next state. At that time, the instantaneous reward value that the agent obtains from the environment; This is the current state, belonging to the state space. ; The action to be performed belongs to the action space. ; The next state after performing an action satisfies the state transition probability. ,in, Indicates time The state is given by P(⋅∣⋅), which represents the conditional probability. Indicates time The action; in, These are weighting coefficients. This is an energy consumption reward item, negatively correlated with the real-time energy consumption of the compressor. This is a bonus item for equipment wear and tear, negatively correlated with the number of compressor start-ups and shutdowns and vibration levels. This is a pipeline stability incentive item, negatively correlated with pipeline pressure and flow fluctuation rates. For the first Weighting coefficients for each reward item; For indexing, These correspond to energy consumption incentives, equipment loss incentives, and pipeline stability incentives, respectively.
4. The compressor start-stop optimization method based on reinforcement learning according to claim 3, characterized in that, The optimization objective of the Markov decision process model is to maximize the cumulative discounted reward of the agent starting from the initial state, and the mathematical expression is: , In the formula, For strategy The expected value of the cumulative discount reward under the following conditions The policy of the agent is defined as follows: , indicating the state Select action The probability of the strategy is a random strategy; Indicating in strategy The expected value of the following; Discount factor of Power of; This is an instant reward function.
5. The compressor start-stop optimization method based on reinforcement learning according to claim 3, characterized in that, The specific boundary conditions include: Physical operating condition boundaries are the inherent safe operating ranges for pipeline networks and compressors in engineering design. They are used to identify and eliminate outliers during the data acquisition phase and to trigger boundary pre-protection during the model inference phase. The boundary between the state space and the action space is determined by linearly normalizing the state space parameters based on the physical condition boundary, thus limiting the input of the reinforcement learning optimization model to a uniform numerical range and constraining the action space of the model output to be a discrete finite set. The model training constraint boundary refers to the numerical constraints followed by the reinforcement learning optimization model during the offline training and online fine-tuning stages, including the learning rate boundary, the policy network parameter update magnitude boundary, and the cumulative discount reward convergence boundary. The decision execution hard limit boundary serves as a safety fallback boundary during the compressor start-stop decision execution phase. The priority of the decision execution hard limit boundary is higher than the decision output of the reinforcement learning optimization model. When the real-time monitored pipeline or compressor parameters exceed the hard limit boundary, an emergency start-stop command is directly triggered, and the model inference results are masked.
6. The compressor start-stop optimization method based on reinforcement learning according to claim 5, characterized in that, The mathematical expression for the physical operating condition boundary is a range of parameter values, and the core boundary includes: , In the formula, For pipeline pressure boundary, For pipeline flow boundary, For the temperature boundary of the medium, For the compressor energy consumption boundary, For compressor vibration boundary, This refers to the compressor speed boundary; The minimum / maximum operating pressure allowed by the pipeline network; The minimum / maximum allowable flow rate of the pipeline network; The minimum / maximum operating temperature allowed by the medium; This represents the lower / upper limit of normal fluctuations in compressor energy consumption. The minimum / maximum allowable vibration amplitude of the compressor; The minimum / maximum permissible speed of the compressor; The normalized mathematical expression for the state space parameters at the boundary between the state space and the action space is: , In the formula, For the first in the state space The normalized values of the parameters, and This represents the physical boundary extreme value of the parameter. For the first The raw real-time measurement values of each parameter; The mathematical expression for the model training constraint boundary is: , In the formula, The learning rate boundary; To update the magnitude boundaries of the policy network parameters, This is the preset maximum update range for a single operation; To accumulate the expected value of the discount reward, and This is the preset convergence boundary; The constraint logic of the model training constraint boundary is as follows: During model training, if the learning rate... or parameter update magnitude If the value exceeds the corresponding boundary, parameter correction will be automatically triggered; if the expected value of the accumulated discount reward is... Failed to converge to the preset boundary If the condition is met, continue training until it is met. The mathematical expression for the hard boundary of the decision execution is: , in, This is represented as the lower limit threshold of emergency pressure. This is represented as the upper limit threshold for emergency energy consumption. This is represented as the upper limit threshold for emergency vibration.
7. The compressor start-stop optimization method based on reinforcement learning according to claim 1, characterized in that, The offline training using the near-end strategy optimization algorithm specifically includes: Set the pruning factor for the near-end policy optimization algorithm. The policy update magnitude is limited by a pruning objective function, the mathematical expression of which is: , In the formula: The ratio of policy updates, i.e., the ratio of the new policy to the old policy in the state. Select Action The probability ratio; Indicates the expectation of a time step; This is the current strategy; This is the old strategy; The advantage function estimate represents the action. The relative merits of the average strategy; For the clipping function, Limited to Within the scope; the model training process is to minimize the pruning objective function. The process continues until the function converges to a preset threshold, yielding the optimal policy network parameters. .
8. The compressor start-stop optimization method based on reinforcement learning according to claim 2, characterized in that, The evaluation model's decision bias involves incremental fine-tuning of the pre-trained optimization model when the bias exceeds a preset threshold. Specifically, this includes: The evaluation indicators for the model optimization effect are set, including the number of compressor start-stop cycles, unit delivery energy consumption, and pipeline pressure fluctuation rate. The actual performance metrics in the feedback dataset are compared with the predicted metrics of the pre-trained optimization model to calculate the decision bias rate. When the decision bias rate exceeds the preset threshold, the model is determined to have a working condition drift. The feedback dataset is used as incremental training data, and the policy network parameters of the pre-trained optimization model are fine-tuned in small batches under the constraints of the model training constraint boundary. The fine-tuned model is then redeployed to the online inference unit.